惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
T
The Blog of Author Tim Ferriss
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
B
Blog RSS Feed
人人都是产品经理
人人都是产品经理
J
Java Code Geeks
Apple Machine Learning Research
Apple Machine Learning Research
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Google Online Security Blog
Google Online Security Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
酷 壳 – CoolShell
酷 壳 – CoolShell
The GitHub Blog
The GitHub Blog
N
News and Events Feed by Topic
Security Archives - TechRepublic
Security Archives - TechRepublic
D
DataBreaches.Net
阮一峰的网络日志
阮一峰的网络日志
H
Heimdal Security Blog
PCI Perspectives
PCI Perspectives
A
About on SuperTechFans
Forbes - Security
Forbes - Security
博客园 - 三生石上(FineUI控件)
博客园_首页
Vercel News
Vercel News
H
Hacker News: Front Page
WordPress大学
WordPress大学
Hacker News: Ask HN
Hacker News: Ask HN
MyScale Blog
MyScale Blog
Recent Announcements
Recent Announcements
N
News | PayPal Newsroom
T
Threatpost
Hacker News - Newest:
Hacker News - Newest: "LLM"
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
SecWiki News
SecWiki News
D
Docker
G
Google Developers Blog
L
LINUX DO - 最新话题
O
OpenAI News
S
Schneier on Security
AI
AI
T
The Exploit Database - CXSecurity.com
S
Security Affairs
I
Intezer
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Security Latest
Security Latest
博客园 - 司徒正美
Know Your Adversary
Know Your Adversary
爱范儿
爱范儿
T
Troy Hunt's Blog
L
LINUX DO - 热门话题

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Why Coding Agents Need Two Halves of Infrastructure: Control Plane + Fast Data Plane
Paul Twist · 2026-06-18 · via DEV Community

You've probably seen the benchmarks by now. Bifrost does 11 microseconds. LiteLLM Python does 40-50ms. The messaging is simple: latency matters for gateways. But this misses what teams actually building with Claude Code and Codex have discovered: the real problem isn't gateway latency alone. It's that coding agents need two completely different infrastructure layers, and teams are treating them like one.

The Two Layers of Coding Agent Infrastructure

When you deploy Claude Code or Codex into a team, you're actually solving two separate problems:

1. The Control Plane: Agent Management & Governance

What it does: Manage agent lifecycle, sessions, memory, permissions, scheduling, tool access, and audit trails.

What teams need here:

  • Create agents without touching provider consoles
  • Run agents on a schedule or via webhook
  • Persist session state across restarts
  • Control which teams/users can access which tools
  • Log every agent action for compliance
  • Manage MCP (Model Context Protocol) tool access across a team
  • Integrate agents with internal systems (databases, git, APIs)

Why a gateway can't do this: A gateway sees request-response pairs. It doesn't know that Agent-A is trying to access Customer Database and needs to be blocked, or that you want to run a code review agent every night at 2 AM. These are control decisions that live above the request layer.

2. The Data Plane: Fast LLM Routing & Reliability

What it does: Route LLM requests to the right provider, handle fallbacks, track costs, log traffic, enforce budgets.

Why it matters for coding agents: Each code edit, test run, or tool invocation is an LLM call. Claude Code can make 30-50 calls per task. If your gateway adds 1ms per call, that's 30-50ms of compounded overhead. With 40-50ms per Python gateway call, you're looking at 1.2-2.5 seconds of pure gateway latency on a single 30-call task. That's the latency you actually feel.


The Gap Teams Are Hitting

Most gateway discussions treat these as one problem. A fast gateway that routes LLM requests. That's necessary, but not sufficient.

Here's what I'm seeing in production teams running Claude Code:

  1. Someone needs to create agents without coding → Control plane job
  2. Agents need to persist across pod restarts → Control plane job
  3. Tool definitions need to stay in sync → Control plane job
  4. Agent A needs MCP server X, Agent B can't have it → Control plane job
  5. We need to route to Claude when latency matters, fallback to Gemini on rate limits → Data plane job
  6. We need sub-millisecond overhead so 30 calls don't add 1.5 seconds → Data plane job
  7. We need cost tracking per agent → Data plane job
  8. We need audit logs of every tool call → Hybrid: agent control + data logging

You can't do #1-4 with a pure gateway. You can't efficiently do #5-8 with a control platform that doesn't understand routing.

Teams are currently solving this by bolting together:

  • A managed service for agent orchestration (Bedrock, Anthropic Hosted)
  • A separate gateway for LLM routing (Bifrost, Portkey, Kong)
  • Custom scripts to sync them
  • Extra glue for observability

That works, but it creates operational friction.


What Production Teams Actually Need

The teams I've talked to who are scaling Claude Code and Codex beyond a single developer describe it like this:

"Claude Code is incredible when it's one engineer using it locally. The moment we try to run it on a team, we need:

  • A place to define 'these are our agents' (not 30 copies in 30 notebooks)
  • A way to say 'this agent runs on a schedule'
  • Control over which tools each agent can access
  • Visibility into what each agent is doing
  • The ability to route to Claude or Gemini based on task type without editing the agent
  • Sub-millisecond gateway latency so 30 LLM calls don't turn into 1.5 seconds of overhead"

That's control plane + data plane thinking. Two distinct layers.


Where the Separation Matters

Control Plane (Agent Platform)

  • Harness abstraction: swap Claude Code ↔ Codex ↔ OpenCode without rewriting agents
  • Session persistence: pause an agent, restart, pick up from where it left off
  • Scheduling: cron, webhooks, API triggers
  • Tool management: centralized MCP registry, per-agent capability matrix
  • Memory: persistent context across runs
  • Access control: who can create/run/modify which agents

Example: You define an agent in the platform UI, attach it to Claude Code, give it GitHub + AWS MCPs, schedule it to run every night on your codebase's failing tests, and it automatically creates PRs with fixes. All without anyone touching provider consoles.

Data Plane (Fast Gateway)

  • Sub-millisecond routing: multiple calls per task, latency compounds
  • Multi-provider routing: route to Claude for complex tasks, Gemini for simple ones
  • Fallback chains: if Claude rate-limits, automatically try Gemini
  • Cost tracking: per-provider, per-model, per-agent visibility
  • Budget enforcement: hard caps that prevent runaway costs
  • Observability: structured logs of every LLM call

Example: Your coding agent makes 40 calls. Gateway adds <1ms per call (40 microseconds total). Control plane tracks that this agent used Claude on 25 calls, Gemini on 15, cost $0.23. If Claude hits rate limits, gateway transparently retries on Gemini without the agent knowing.


Recent Signals That This Separation Is Hardening

Claude Code: Now supports Hooks for policy enforcement at key lifecycle events (TaskStarted, ToolCall, TaskCompleted). That's control-plane-level thinking inside the agent harness.

Codex: Added Managed Agents API for creating/running agents from your own infrastructure. That's recognizing the control plane problem.

LiteLLM-Rust: Launched June 2026 specifically for "coding agent workloads" with <1ms target on Claude Code calls, integrated sandbox support (E2B, Daytona), and durable sessions on the roadmap. That's explicitly targeting the data plane + agent runtime integration.

TrueFoundry, Kong, Portkey: All shipping "agent gateway" features that blur the line — they're trying to build control + data in one platform.

The market is recognizing that governance + routing are different concerns, even if some platforms try to unify them.


The Practical Decision Framework

If you're running Claude Code on a single developer or small team:

  • Local Claude Code + direct API calls to Claude/OpenAI
  • No gateway needed yet

If you're scaling Claude Code or Codex across a team:

  1. Pick your control plane: self-hosted (LiteLLM Agent Platform), managed (Bedrock), or API-first (Codex API)
  2. Pick your data plane: fast gateway (LiteLLM-Rust, Bifrost) or managed routing (OpenRouter, Portkey)
  3. Ensure they can talk: same config format, shared database, compatible APIs

If you're mixing multiple harnesses (Claude Code + Codex + OpenCode):

  • Control plane becomes critical: you need abstraction over harness differences
  • Data plane must support all providers those harnesses call

If you need compliance/audit/data residency:

  • Control plane must self-host
  • Data plane must self-host
  • Consider platforms that do both (LiteLLM Agent Platform + LiteLLM-Rust + LiteLLM core)

What I'd Measure

When evaluating control + data infrastructure for coding agents:

Control Plane:

  • Can I create an agent in a UI? How much YAML/JSON?
  • Does it understand session state? Can agents pause and resume?
  • Can I attach MCPs to agents and control per-agent access?
  • Do audit logs show me every tool call and agent action?
  • Can I schedule agents on cron/webhook/API?

Data Plane:

  • What's the measured overhead per request? (target: <1ms for coding agents)
  • Does it support all the providers my harnesses call?
  • Can I route based on task type or cost? (not just round-robin)
  • Does it fail gracefully when a provider is rate-limited?
  • Is cost/token tracking per-agent visible?

Integration:

  • Do the two layers use the same config format?
  • Do they share the same database or API?
  • Can I swap data planes without redeploying agents?

The Unsexy Part

The sexiest discussion is always about latency. But in teams I talk to running Claude Code at scale, the conversation goes:

"Okay, gateway overhead is solved. Now: how do we keep prod from running experiments? How do we keep Agent-A from accessing the customer database? How do I know what happened yesterday when Agent-B deleted something? Can I run this every night? Can I give the junior engineer the ability to create agents without giving them API access?"

That's all control plane work. And it's unglamorous, but it's what stops you from sleeping at 3 AM.


Wrapping Up

Coding agents (Claude Code, Codex, OpenCode) need two halves:

  1. Control Plane: Agent creation, sessions, memory, scheduling, tool governance, audit trails. This is where you keep your team safe.
  2. Data Plane: Sub-millisecond LLM routing, fallbacks, cost tracking, multi-provider support. This is where you keep your latency low.

A fast gateway solves problem #2. A control platform solves problem #1. Both are table stakes for production.

The platforms that will win here are the ones that make the separation clear and let teams pick the right tool for each job — or that do both well without unnecessary coupling.

What's your experience been? Are you running coding agents on a team? What's the first thing that broke when you tried to scale from one developer to five?


Resources


Paul Twist is an AI infrastructure engineer based in Berlin, focused on production agent systems and multi-provider LLM routing.