惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
AWS News Blog
AWS News Blog
Y
Y Combinator Blog
云风的 BLOG
云风的 BLOG
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
F
Full Disclosure
H
Help Net Security
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
A
About on SuperTechFans
J
Java Code Geeks
Jina AI
Jina AI
GbyAI
GbyAI
酷 壳 – CoolShell
酷 壳 – CoolShell
爱范儿
爱范儿
美团技术团队
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Latest news
Latest news
Vercel News
Vercel News
博客园 - 【当耐特】
P
Privacy & Cybersecurity Law Blog
P
Proofpoint News Feed
阮一峰的网络日志
阮一峰的网络日志
V
Vulnerabilities – Threatpost
Stack Overflow Blog
Stack Overflow Blog
Hugging Face - Blog
Hugging Face - Blog
D
Docker
Microsoft Security Blog
Microsoft Security Blog
博客园_首页
S
Securelist
WordPress大学
WordPress大学
S
Secure Thoughts
博客园 - 聂微东
Cloudbric
Cloudbric
Help Net Security
Help Net Security
腾讯CDC
T
Threat Research - Cisco Blogs
T
Tor Project blog
L
LINUX DO - 热门话题
Project Zero
Project Zero
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - Franky
N
Netflix TechBlog - Medium
小众软件
小众软件
Cyberwarzone
Cyberwarzone
量子位
MyScale Blog
MyScale Blog
W
WeLiveSecurity
MongoDB | Blog
MongoDB | Blog
I
InfoQ
M
MIT News - Artificial intelligence

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
5 ways AI agents quietly die inside n8n production
Mirza Iqbal · 2026-05-27 · via DEV Community
[ERROR] node "GPT Decision" execution_id=9f...3a status=failure
  cause: structured_output_schema_violation
  retries: 6  total_runtime_ms: 184302
  workflow: invoice-router-v3 owner: ap-team

Enter fullscreen mode Exit fullscreen mode

That node ran 47 times today.

Each run burned three retries before n8n gave up.

Cost on the OpenAI side was real money.

Cost on the human side was a finance ops lead manually re-routing 47 invoices because the agent never told anyone it was looping.

The article most teams read this week is "agents hallucinate". That problem is solved by structured output. The real failure modes in n8n agent production are different. Here are the five that actually fire on weeknights.

1. The silent retry storm

n8n retries on error by default. An LLM node that 429s under load retries with the same prompt, same model, same payload. Each retry costs money and produces the same failure.

The fix is to gate retries on the error class.

// In a Code node before the LLM call
const lastError = $input.first().json.error;
if (lastError?.code === 'rate_limit_exceeded') {
  // exponential backoff, not n8n's flat retry
  await new Promise(r => setTimeout(r, 1000 * Math.pow(2, $runIndex)));
  return $input.all();
}
if (lastError?.code === 'invalid_request') {
  // schema problem. retrying will not help.
  throw new Error('halt invalid_request, escalate to human queue');
}
return $input.all();

Enter fullscreen mode Exit fullscreen mode

The agent now distinguishes "transient" from "terminal". The retry storm dies.

2. Tool-call drift across long workflows

A multi-step agent flow calls tool A, then tool B, then tool C. Each tool returns slightly different JSON shapes. By step C the agent is reasoning over a structure that no longer matches its system prompt.

I have seen this in 6-step Clay-to-n8n-to-Salesforce flows. The Salesforce step fails because the contact object got mutated three steps back and nobody normalized it.

The fix is a normalization node between every tool call.

Tool A => Set node (rename + strip) => Tool B => Set node (rename + strip) => Tool C

Enter fullscreen mode Exit fullscreen mode

It looks redundant. It is not. The Set node enforces the schema your agent's system prompt promised. If schema and reality diverge, you get a typed error early, not a wrong invoice routed at midnight.

3. Silent payload truncation inside n8n's HTTP wrapper

n8n's OpenAI node and the HTTP Request node both have request body size limits that are NOT documented anywhere obvious. When the prompt plus tool history plus retrieval results cross about 950 KB, the HTTP body gets truncated by the n8n proxy, the LLM sees a malformed request, and the agent returns a vague refusal.

This one bit me twice on two different client projects. The agent worked fine on test inputs and failed mysteriously on real production payloads that had longer chat history.

The fix is to chunk payload BEFORE the LLM node, not inside the LLM provider.

// In a Code node sized for n8n's 1MB practical limit
const MAX_KB = 800;
const history = $input.first().json.history || [];
let total = 0;
const trimmed = [];
for (let i = history.length - 1; i >= 0; i--) {
  total += JSON.stringify(history[i]).length;
  if (total > MAX_KB * 1024) break;
  trimmed.unshift(history[i]);
}
return [{ json: { history: trimmed } }];

Enter fullscreen mode Exit fullscreen mode

Trim from the tail. Newest messages survive. Oldest get summarized in a separate node and re-injected as a single system message.

4. The credentials-rotation blackout

Enterprise rotates API keys quarterly. n8n credentials are encrypted at rest and decrypted by the credential service on every workflow run. If a key rotates and the credential update has not propagated, every active workflow fails silently to a 401, and n8n's default error path swallows the auth failure as a "node error" with no alert.

You find out when revenue dashboards stop refreshing.

The fix is a credentials health check workflow that runs every hour and pings every active integration.

Cron (every 1h) =>
  HTTP GET /credentials/all (n8n REST API) =>
  Loop over each credential =>
    Trigger a dry-run call against the provider =>
      If 401, send to ntfy.sh on the on-call channel

Enter fullscreen mode Exit fullscreen mode

That single workflow saved one of my clients 11 hours of debugging in March when their Slack OAuth key rotated and the marketing team's lead-routing flow went dark.

5. Memory poisoning across runs

If you store conversation memory in a Postgres or Redis-backed n8n credential and reuse it across runs, one bad agent output can poison every subsequent run.

I saw this happen with a customer service flow. A user typed a prompt-injection payload. The agent's "memory" node wrote the payload into Redis. Every subsequent customer for that ticket inherited the injection. Three hours later, the agent was telling people their refund was approved when it was not.

The fix is to validate memory on read, not only on write.

// In a Code node before the agent's memory-recall step
const recalled = $input.first().json.memory;
const SUSPICIOUS = /(ignore.*previous|you are now|system\W|admin\W)/i;
if (SUSPICIOUS.test(recalled)) {
  // memory is contaminated. drop it and start fresh.
  return [{ json: { memory: '', alert: 'memory_quarantined' } }];
}
return $input.all();

Enter fullscreen mode Exit fullscreen mode

The memory pattern works fine right up until the day someone feeds a poisoned input. The validate-on-read step is two lines and prevents a class of failure that costs trust to recover from.

What dies in production is not what you tested

Hallucination shows up in dev. These five patterns show up in production. The split matters because the dev-time fixes (structured output, retries, evals) do not catch any of the five above.

If you are running agents in n8n right now, the cheapest thing you can do this week is add a normalization node between every tool call and a credentials health-check workflow. Those two changes alone caught roughly 70 percent of the silent failures in the last enterprise rollout I audited.

What is the failure mode that bit you that you do not see written about anywhere? Drop a snippet in the comments. The pattern library only grows when more people share the n8n flows that actually broke.