惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
P
Privacy International News Feed
Vercel News
Vercel News
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
博客园 - 叶小钗
F
Fortinet All Blogs
Security Archives - TechRepublic
Security Archives - TechRepublic
L
LINUX DO - 最新话题
AWS News Blog
AWS News Blog
Engineering at Meta
Engineering at Meta
Attack and Defense Labs
Attack and Defense Labs
Recent Announcements
Recent Announcements
Recent Commits to openclaw:main
Recent Commits to openclaw:main
PCI Perspectives
PCI Perspectives
Cloudbric
Cloudbric
AI
AI
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
IT之家
IT之家
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
J
Java Code Geeks
M
MIT News - Artificial intelligence
Cisco Talos Blog
Cisco Talos Blog
V2EX - 技术
V2EX - 技术
Webroot Blog
Webroot Blog
Microsoft Security Blog
Microsoft Security Blog
Cyberwarzone
Cyberwarzone
博客园 - 聂微东
G
Google Developers Blog
W
WeLiveSecurity
罗磊的独立博客
P
Privacy & Cybersecurity Law Blog
阮一峰的网络日志
阮一峰的网络日志
A
About on SuperTechFans
WordPress大学
WordPress大学
The GitHub Blog
The GitHub Blog
T
Tailwind CSS Blog
V
Visual Studio Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
S
Secure Thoughts
Apple Machine Learning Research
Apple Machine Learning Research
Hugging Face - Blog
Hugging Face - Blog
Google DeepMind News
Google DeepMind News
Google DeepMind News
Google DeepMind News
雷峰网
雷峰网
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
F
Full Disclosure
Blog — PlanetScale
Blog — PlanetScale
The Last Watchdog
The Last Watchdog
P
Proofpoint News Feed

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
AI Harnesses: Why DevOps Principles Are the Missing Piece in Agentic Development
Hector Flore · 2026-05-19 · via DEV Community

The Breakthrough That Changed How I Think About Agents

Last night I watched a talk from AI.engineer that crystallized something I've been building toward for months. The speaker demonstrated fixing a series of agent failures — not by changing the prompt, not by upgrading the model, not by adding more context. They fixed it by improving the harness.

The agent was the same. The model was the same. The instructions were identical. But the harness — the infrastructure that controls how the agent operates — made the difference between a broken tool and a production-ready system.

That's when the parallel hit me like a freight train:

DevOps was the tool we gave humans to control their workflows. AI Harnesses are the tool we give agents to control their workflows.

The AI.engineer talk that crystallized the harness thesis — fixing agent failures through infrastructure, not prompts.

This isn't a loose analogy. It's a direct architectural parallel — and understanding it is the key to building agent systems that actually work in production.

What Is an AI Agent Harness?

An agent harness is a set of computer science primitives that govern agent behavior regardless of model. It's the runtime infrastructure that sits between "I have an LLM" and "I have a production agent system."

Think of it this way: a model is a brain. A harness is the nervous system, skeleton, and reflexes that turn that brain into a functioning organism.

The 5 core harness primitives — Tool Registry, Context Management, Guardrails, Agent Loop, and Compaction & Memory — arranged around a central hub
The five primitives that make up a complete AI agent harness. Each solves a different dimension of agent governance.

A complete harness includes five core primitives:

1. Tool Registry

The tool registry defines what an agent can do — the complete set of capabilities available at runtime. But it's more than a list. A well-designed registry includes:

  • Discovery — agents can find tools by capability, not just name
  • Schema validation — every tool call is validated against its input schema before execution
  • Access control — not every agent gets every tool. The registry enforces boundaries.
  • Versioning — tools evolve. The registry manages backward compatibility.

In DevOps terms, this is your artifact registry meets your IAM policy. You wouldn't give every developer root access to production. You shouldn't give every agent access to every tool.

2. Context Management

Context management is the assembly pipeline that determines what information an agent receives before making a decision. This includes:

  • Memory tiers — structured context (core identity → working state → long-term patterns → event streams) loaded at the right time
  • Skill injection — reusable capability definitions loaded on-demand when relevant
  • Compaction — intelligent summarization when context windows fill up
  • Priority — not all context is equal. Critical instructions survive compaction; nice-to-haves don't.

The DevOps parallel: this is your configuration management. Ansible, Terraform, Puppet — all of them solve the same problem for infrastructure that context management solves for agents. The right configuration, at the right time, to the right target.

3. Guardrails

Guardrails are the moat. They're pre-execution interceptors that prevent agents from doing the wrong thing — not by instruction, but by architecture.

  • Pre-tool hooks — intercept dangerous operations before they execute (like hookflows that redirect git commit to dev_commit)
  • Post-tool validation — verify outputs match expected patterns
  • Dynamic enforcement — rules that can change per-prompt, per-context, at runtime
  • Sandboxing — file system restrictions, network controls, tool allowlists

Here's the key insight: a guardrail doesn't add instructions about what not to do. It removes the wrong option entirely. The agent can't make the mistake because the mistake path doesn't exist. That's fundamentally different from telling an agent "please don't do X" and hoping it listens.

Comparison of instruction-based governance vs architecture-based governance — showing how guardrails remove wrong paths entirely rather than relying on instructions
Guardrails don't instruct against bad behavior — they architecturally remove it. The wrong path simply doesn't exist.

4. Agent Loop

The agent loop is the core execution cycle — observe, think, act, repeat. But a harness-engineered loop includes:

  • Termination conditions — when to stop (max iterations, goal reached, error threshold)
  • Sub-agent orchestration — spawning specialized agents for sub-tasks
  • Error recovery — retry strategies, fallback paths, graceful degradation
  • Progress tracking — observability into where the agent is in its workflow

This maps directly to your CI/CD pipeline runner. Jenkins, GitHub Actions, Azure DevOps — they all implement a task loop with retries, parallelism, error handling, and observability. An agent loop is the same pattern applied to cognitive work.

5. Compaction and Memory

Long-running agents hit context limits. Compaction is the strategy for managing this:

  • Checkpoint summaries — periodic snapshots of progress
  • Selective retention — keep critical decisions, discard routine operations
  • Persistence hierarchy — what goes to short-term memory vs. long-term storage
  • Cross-session recall — enabling agents to build on prior work

The DevOps equivalent: log rotation and data lifecycle management. You don't keep every debug log forever. You tier your data — hot storage for recent, warm for searchable history, cold for compliance archives.

Why Harnesses Matter More Than Model Choice

Here's the uncomfortable truth the AI industry doesn't want you to hear: the model is increasingly commodity. GPT-5, Claude Opus, Gemini Ultra — they're all converging on similar capability levels. The performance gap between top models shrinks every quarter.

But the performance gap between a well-harnessed agent and a raw model with a prompt? That gap is enormous and growing.

The demo I watched proved it empirically. Same model. Same prompt. Different harness quality. The results went from "broken and unreliable" to "production-ready" — purely through harness improvements.

This maps perfectly to DevOps history. In 2010, the "developer" was the bottleneck everyone focused on. Hire better developers. Give them better tools. Train them more. But the organizations that won weren't the ones with the best individual developers — they were the ones with the best delivery systems. CI/CD, IaC, observability, feature flags — the infrastructure that made even average developers reliably productive.

Same pattern. Same lesson. Invest in the harness, not just the model.

Static vs. Dynamic: The Frontier

Most harnesses today are static. You define your tools, write your system prompt, set up guardrails, and deploy. The configuration is the same for every execution.

Visual showing the evolution from rigid static harnesses to adaptive dynamic harnesses that adjust governance per-prompt based on risk level
Static harnesses apply the same rules to every execution. Dynamic harnesses adapt governance per-prompt — loose for low risk, tight for high risk.

This is like having a single CI/CD pipeline for every project. It works — until it doesn't. Until you need different governance for different contexts, different risk levels, different domains.

The future is dynamic harnesses — where guardrails, tool access, context assembly, and execution policy are defined per-prompt at runtime.

Imagine an agent that, for low-risk tasks, operates with minimal guardrails and maximum autonomy. But for the same agent handling a financial transaction or a production deployment, the harness tightens: human-in-the-loop gates activate, additional validation layers engage, the tool registry narrows to only approved operations.

This isn't science fiction. This is what GitHub Copilot Extensions enable right now — the ability to dynamically inject tools, context, and governance into an agent's runtime based on the specific task. It's also what I've been building toward with my own agent-harness project.

The DevOps Parallel (It's Not a Metaphor — It's a Pattern)

Let me make this explicit. Every major DevOps innovation has a direct parallel in harness engineering:

Side-by-side mapping showing DevOps concepts on the left connected by arrows to their AI harness equivalents on the right
The DevOps → Harness parallel isn't a loose analogy — it's the same architectural pattern applied to cognitive work instead of infrastructure.

DevOps Gave Humans Harnesses Give Agents
CI/CD pipelines Agent loops with termination and retry
Infrastructure as Code Context-as-code (skills, constitutions, memory tiers)
Deployment gates Autonomy levels and approval gates
RBAC and least privilege Tool registry access control
Observability (logs, metrics, traces) Agent event streams, checkpoints, memory
Git hooks Pre-tool hookflows
Sandboxed environments Agent sandboxing (file, network, tool boundaries)
Feature flags Dynamic guardrails per-prompt

The lesson DevOps taught us was: don't rely on discipline alone — build systems where the right thing is the easy thing. CI/CD didn't succeed because developers became more careful. It succeeded because the pipeline made quality automatic.

The same principle applies to agents. You don't make agents trustworthy by writing better prompts. You make them trustworthy by engineering harnesses where the right behavior is the default path and the wrong behavior is architecturally impossible.

"Make the right thing to do the easy thing to do" — this principle built the DevOps movement. Now it's building the agentic one.

My Journey: Building Before the Industry Named It

I started building what I now call a harness before the industry had a name for it. My agent-harness repo (TypeScript) emerged from running a production multi-agent system — 30+ agents, 60+ skills, 4-tier memory, hookflows, constitutions, and sandboxing.

Every pattern I described above? I learned it the hard way. Agents going rogue because there were no guardrails. Context windows overflowing because there was no compaction strategy. Agents fighting each other because there was no orchestration layer.

The harness primitives crystallized through production pain, not academic theory.

Now I'm building a Go implementation focused on maximum determinism and testability. Why Go?

  • Compiled binary — no runtime dependencies, no "works on my machine"
  • Strong typing — harness configuration errors caught at compile time
  • Concurrency primitives — goroutines map naturally to parallel agent execution
  • Testability — interfaces and table-driven tests for every harness primitive
  • Performance — sub-millisecond hook execution for real-time guardrails

The goal: a harness so deterministic that you can write unit tests for agent governance. If you can test your CI/CD pipeline, you should be able to test your agent harness.

The 2027 Vision: Every Prompt Ships With Its Governance

Here's where this is heading. By 2027, I believe we'll see:

Timeline roadmap showing the 2027 vision: Dynamic Harnesses as Infrastructure, Harness-as-Code, Harness Marketplaces, and the Harness Engineer role
The roadmap to 2027: harness engineering becomes a first-class discipline with its own tooling, marketplaces, and career paths.

Dynamic harnesses as first-class infrastructure. Every prompt that enters an agent system will carry its own governance metadata — tool permissions, guardrail configuration, context assembly rules, termination conditions. Not as an afterthought. Not as a bolted-on safety layer. As first-class harness engineering.

Harness-as-Code. Just as Infrastructure-as-Code made deployments reproducible, reviewable, and testable — Harness-as-Code will do the same for agent governance. Your agent's behavior policy will live in version control, go through code review, and have automated tests.

Harness marketplaces. The same way Terraform modules and GitHub Actions created ecosystems of reusable infrastructure — harness components (guardrail packs, context strategies, tool registries) will become composable building blocks.

The "Harness Engineer" role. Just as DevOps created the DevOps Engineer, Platform Engineer, and SRE roles — harness engineering will create a new discipline. People who specialize in making agents trustworthy, deterministic, and governed.

What You Should Do Today

If you're building with agents — whether it's a single Copilot workspace or a fleet of autonomous systems — start thinking about your harness:

  1. Audit your tool access. What can your agents do? Should they be able to do all of it? Implement a registry with boundaries.
  2. Add pre-execution hooks. Before an agent calls a tool, validate the call. Block dangerous patterns. Redirect to governed alternatives.
  3. Structure your context. Don't dump everything into the system prompt. Tier it. Load what's needed when it's needed.
  4. Define termination. How does your agent know when to stop? What's the max iteration count? What's the error budget?
  5. Make it testable. If you can't test your agent's governance, you can't trust it.

The Bottom Line

DevOps taught us that you don't build reliable software by hiring more careful developers. You build it by engineering systems where reliability is the default. The same principle — exactly the same principle — applies to agentic development.

The model is the developer. The harness is the DevOps infrastructure. And just like the 2010s proved that investing in delivery infrastructure beats investing in individual developer heroics, the 2020s will prove that investing in harness engineering beats investing in prompt engineering alone.

The right thing to do should be the easy thing to do — for humans AND agents. That's Agentic DevOps. That's what I'm building.


Want to go deeper on harness engineering and Agentic DevOps? Subscribe to the htek.dev newsletter for weekly deep-dives, or explore the blueprints for hands-on implementation guides.