惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
Darknet – Hacking Tools, Hacker News & Cyber Security
NISL@THU
NISL@THU
S
Securelist
O
OpenAI News
S
Security Affairs
Cyberwarzone
Cyberwarzone
T
Threatpost
Simon Willison's Weblog
Simon Willison's Weblog
The Last Watchdog
The Last Watchdog
L
LINUX DO - 最新话题
C
Cisco Blogs
PCI Perspectives
PCI Perspectives
SecWiki News
SecWiki News
S
Secure Thoughts
GbyAI
GbyAI
I
Intezer
AWS News Blog
AWS News Blog
F
Fortinet All Blogs
I
InfoQ
阮一峰的网络日志
阮一峰的网络日志
Google Online Security Blog
Google Online Security Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
A
About on SuperTechFans
S
Schneier on Security
P
Proofpoint News Feed
雷峰网
雷峰网
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
小众软件
小众软件
H
Heimdal Security Blog
Microsoft Security Blog
Microsoft Security Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
T
The Exploit Database - CXSecurity.com
T
Threat Research - Cisco Blogs
V
V2EX
L
Lohrmann on Cybersecurity
Security Latest
Security Latest
A
Arctic Wolf
Apple Machine Learning Research
Apple Machine Learning Research
H
Hacker News: Front Page
Cisco Talos Blog
Cisco Talos Blog
Webroot Blog
Webroot Blog
T
Tenable Blog
MyScale Blog
MyScale Blog
博客园 - 司徒正美
S
SegmentFault 最新的问题
Y
Y Combinator Blog
腾讯CDC
Hacker News: Ask HN
Hacker News: Ask HN
M
MIT News - Artificial intelligence
G
GRAHAM CLULEY

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
Why Your AI Agents Keep Breaking Your Workflows
kdb1008 · 2026-04-28 · via Hacker News - Newest: "AI"

Your AI investment isn’t paying off the way you expected. You added agents to your workflows, and now your team spends more time debugging the AI than the AI saves them. So you write better prompts. Add more guardrails. Spell out every constraint. The agents still break things, just in new ways.

The prompts aren’t the problem. The architecture is.

I build and operate multi-agent systems where AI agents coordinate across multi-step workflows, handling tasks from analysis and planning through execution and verification. In one of those systems, an agent recently skipped two entire workflow phases, bypassing review, tests, and isolation checks. A single line in the implementation plan said “no worktree needed,” and the agent interpreted that as permission to shortcut the whole process. Its reasoning was locally coherent. The decision was globally catastrophic. Nothing in the prompt prevented it.

That experience confirmed something I’d been seeing across every multi-agent system I’ve worked on: instructions cannot enforce workflow structure. Only architecture can.

Before I explain why agents fail this way, here’s the mental model that makes everything else in this post click.

Think of it like a restaurant kitchen. The chef handles creative decisions: how to balance flavors, how to adapt when an ingredient is missing, how to plate something beautifully. The kitchen manager controls which stations are open, what’s available, and when service begins. The chef works within the structure the kitchen manager defines. Nobody asks the chef to also manage the schedule.

In an agentic system, this maps to two layers: a deterministic control plane and a probabilistic data plane.

The control plane owns the workflow. It manages the execution graph, state persistence, timeouts, and retry logic. It decides what happens next and enforces that decision. Agents cannot skip a step the control plane hasn’t authorized.

The data plane is where agents live. They receive bounded context from the control plane, execute a discrete reasoning step, and return structured output. They don’t manage state. They don’t decide what comes next. They process and respond.

Your agents are acting as both chef and kitchen manager, and they’re not equipped for the second job.

This is a different class of failure than hallucination, and it’s harder to catch. The agent reasons its way to a wrong decision. The logic looks sound when you read the transcript. The outcome is wrong because the agent has no awareness of the larger workflow it’s operating inside.

A refund agent bypasses the 30-day return window because a customer’s message conveyed urgency. An order processing agent skips inventory verification because the previous step returned success. The phase-skipping failure I described in the opening is the same pattern: the agent found a locally reasonable shortcut that violated architectural constraints it couldn’t see. Every one of those skipped steps existed for a reason. The agent couldn’t know that, because its context window only contained the immediate task, not the architectural rationale for the workflow.

The instinct is to add more rules to the prompt. It doesn’t work. The failures are structural, not informational.

Prompt-driven state loss is the most common: as conversations grow, tool outputs and system messages fill the context window, pushing early constraints out or diluting them. The agent continues operating on an incomplete picture of its own rules.

Context overflow compounds the damage. When models compact context into summaries to stay within limits, specifics disappear. An agent that knows it’s in phase 4 of a multi-phase workflow may, after compaction, only know it’s “working on a task.” Both failures happen silently, without throwing errors that a monitoring system could catch.

Those are the accidental failures. The adversarial ones are worse. Vibe hacking exploits the model’s responsiveness to emotional signals: a customer who expresses urgency or authority can cause an agent to skip validation steps, because the model is designed to be responsive to tone. Indirect prompt injection is more deliberate: a document the agent reads (a support ticket, an invoice, a code comment) contains instructions that redirect its behavior. No amount of prompt engineering fully prevents either, because both exploit the same context-sensitivity that makes the model useful in the first place.

Every one of these failure modes shares the same root cause: the workflow’s integrity depends on the agent remembering and respecting constraints. That’s a bet against the architecture of how language models process context. And it’s exactly why the control plane, not the agent, has to own state.

Prodigal’s analysis of multi-step workflows quantifies what anyone running agents in production already suspects.

At 95% per-step accuracy, a 5-step workflow succeeds 77% of the time. At 10 steps, 60%. At 20 steps, 36%.

Even at 99% per-step accuracy, a 20-step workflow fails nearly 1 in 5 runs.

Deterministic software doesn’t work this way. When a function call fails, you get an error. When an agent makes a locally reasonable but globally wrong decision, you get an HTTP 200—meaning that the browser responds by saying that the page was found—and a corrupted business process. The response tells you nothing about whether the right thing happened.

The two-layer separation I described earlier eliminates the core failure mode. State lives in deterministic code, not in a context window. If a model hallucinates or fails, the control plane catches the schema validation error and triggers a retry or escalation. Not an unconstrained recovery loop.

Two additional patterns address specific failure modes that show up when agents interact with multiple systems or with each other.

Transaction recovery. A workflow that processes a refund, updates inventory, and sends a confirmation email modifies three independent systems. If the email step fails, nothing rolls back the first two. The Saga pattern pairs every forward action with a predefined compensating action. If any step fails, the orchestrator fires compensating commands for everything that already succeeded. This matters for agents specifically because their failures are often silent: a hallucinated parameter might cause a downstream API to reject a request, but the agent won’t recognize that as a failure requiring compensation. The control plane does.

Typed handoffs. Agent-to-agent handoffs are where workflows fracture. Passing natural language between agents creates ambiguity. “Handle the customer ticket” could mean close it, escalate it, or email the customer. All reasonable interpretations. None predictable. Typed schemas eliminate this by enforcing that agent output is serialized into a predefined structure before it’s passed anywhere. The receiving agent gets structured data, not prose.

Action-selector patterns go further. Instead of letting the model output a command, it outputs an action identifier:

The orchestrator maps that identifier to a hardcoded function. The model’s output is treated as data, not as executable instructions. This is the agentic equivalent of parameterized queries: it closes off an entire class of injection and bypass vulnerabilities.

Here’s how these patterns work in the multi-agent development workflows we operate, where agents coordinate across phases from specification through deployment.

The problem is familiar: agents skip review steps, bypass worktree checks, or commit directly to the main branch. Each bypass is locally reasonable from the agent’s perspective. None are acceptable from the workflow’s perspective.

Each phase transition runs a gate check that returns an exit code:
0: context valid, proceed
3: wrong session, stop immediately
4+: phase-specific validation failure

Build completion markers like READY_FOR_TEST_VERIFY appear before phase transitions. The archive gate runs a multi-condition pre-flight check before finalization. The workflow cannot proceed without satisfying all validation requirements.

The agent doesn’t get a vote.

Exit codes handle individual transitions. But the deeper principle is that state cannot live in a context window. Context windows are volatile, lossy, and invisible to the control plane. State has to live in files.

Think of it as the difference between checking a single door lock and running a building-wide security sweep. Exit codes are the door locks. File-based state management is the sweep.

In practice, this works at several levels. Volatile execution state is tracked through schema-validated edits. Frozen requirements and architecture documents carry status markers that prevent modification. A three-tier rule hierarchy — SYSTEM, AGENT, COMMAND — keeps agents from overriding system-level constraints. And manifest integrity is tracked with SHA256 checksums to detect drift or corruption.

Templates are the source of truth. Runtime files derive from templates and are never edited directly. If the agent wants to know what phase it’s in, it reads a file. If the control plane wants to verify what happened, it reads a file. No one asks the model to remember.

The impact is measurable. Before implementing this architecture, a multi-agent workflow with 15+ phases had a 36% success rate per run. After adding deterministic gates and file-based state, the same workflow runs at 95%+ reliability, with debugging time dropping from hours of transcript analysis to minutes of log review.

Every deterministic intervention has a cost. The question is whether the reliability gain justifies it.

Enforce hard where the cost of bypass is high: security validations and data sanitization, human-in-the-loop approvals for financial or production changes, multi-system transactions where partial completion creates inconsistency, and compliance workflows that require audit trails.

Allow flexibility where the cost of bypass is low: brainstorming and exploratory research, read-only operations where no state is modified, single-system workflows where rollback is trivial, and development/testing environments.

Start with standard enforcement for production workloads. Gate at phase transitions and cross-system boundaries, not at every step. Over-gating is the most common mistake. Too many checkpoints slow the workflow without improving reliability.

When something unexpected happens, ask the agent to explain its reasoning. “Why did you skip the commit?” often surfaces a gap in the constraints that rules alone can’t catch. Use those answers to tune your gates.

The problem with agents in deterministic workflows isn’t the agents. It’s the assumption that instructions alone can enforce workflow structure. They can’t. That’s not a bug. It’s how probabilistic systems work.

Start with one gate at the most critical phase transition in your workflow. Measure the consistency improvement. Add gates incrementally as you identify where bypass is causing problems. The goal isn’t maximum enforcement. It’s the minimum enforcement that produces reliable outcomes.

The agents stay probabilistic. The workflow becomes deterministic. That’s the combination that works.

These patterns come from building and operating multi-agent systems at Keryx Solutions, where we help companies make AI work in production, not just in demos. If your AI investment is creating more work than it saves, that’s usually an architecture problem.

Share