惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Privacy & Cybersecurity Law Blog
Engineering at Meta
Engineering at Meta
Forbes - Security
Forbes - Security
MongoDB | Blog
MongoDB | Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
A
About on SuperTechFans
量子位
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
雷峰网
雷峰网
腾讯CDC
P
Proofpoint News Feed
S
Schneier on Security
S
Secure Thoughts
V
Visual Studio Blog
Help Net Security
Help Net Security
The Hacker News
The Hacker News
C
Cyber Attacks, Cyber Crime and Cyber Security
P
Privacy International News Feed
SecWiki News
SecWiki News
S
SegmentFault 最新的问题
T
Threatpost
小众软件
小众软件
MyScale Blog
MyScale Blog
F
Fortinet All Blogs
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
P
Proofpoint News Feed
T
Tailwind CSS Blog
I
Intezer
C
CERT Recently Published Vulnerability Notes
U
Unit 42
V
V2EX
Cyberwarzone
Cyberwarzone
Recorded Future
Recorded Future
O
OpenAI News
Project Zero
Project Zero
有赞技术团队
有赞技术团队
Google DeepMind News
Google DeepMind News
Last Week in AI
Last Week in AI
Hugging Face - Blog
Hugging Face - Blog
Know Your Adversary
Know Your Adversary
C
Cybersecurity and Infrastructure Security Agency CISA
Scott Helme
Scott Helme
V2EX - 技术
V2EX - 技术
博客园 - 叶小钗
S
Securelist
A
Arctic Wolf
The Cloudflare Blog
W
WeLiveSecurity
T
Threat Research - Cisco Blogs
博客园 - Franky

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
What Is an AI Agent? A Production Definition From Running Multi-Agent Systems
Elena Revich · 2026-05-06 · via DEV Community

Originally published on AIdeazz — cross-posted here with canonical link.

Most definitions of AI agents are either too academic ("autonomous entities that perceive and act") or too marketing-driven ("ChatGPT but with buttons!"). After building and deploying multiple agent systems in production — from Telegram bots handling thousands of daily queries to multi-agent workflows on Oracle Cloud — I've developed a more practical definition.

The Core Loop: Observe → Decide → Act → Persist

An AI agent is software that runs this loop continuously:

  1. Observe: Gather context from multiple sources (messages, APIs, database state, other agents)
  2. Decide: Use LLMs or other models to determine next actions based on observations
  3. Act: Execute those actions (send messages, call APIs, update databases, trigger workflows)
  4. Persist: Maintain state across interactions for continuity

This differs fundamentally from chat-only wrappers that simply pipe user input to an LLM and return the response. The key distinction? Agents do things beyond returning text.

Here's a concrete example from one of our production systems: A user messages our Telegram agent asking about their order status. The agent:

  • Observes the message and retrieves the user's ID from Telegram metadata
  • Decides it needs order information, checking its permission scope
  • Acts by querying our Oracle database for order records
  • Persists the interaction context for follow-up questions

The user might then ask "Can you expedite shipping?" The agent already has the order context, checks business rules, and could actually modify the order priority in the system — not just explain how expediting works.

Architecture Patterns That Actually Scale

When people ask "what is an AI agent," they often imagine a single monolithic system. In practice, production agents are usually specialized components in larger systems.

Our typical architecture:

  • Router Agent: Analyzes incoming requests and delegates to specialized agents
  • Task Agents: Handle specific domains (customer service, data analysis, document processing)
  • Coordinator Agent: Manages multi-step workflows across task agents
  • Monitor Agent: Tracks system health and intervenes when needed

This isn't arbitrary complexity. Single-agent systems hit walls quickly:

  • Context windows overflow with state management
  • One prompt template can't handle diverse tasks well
  • Failure in one area cascades everywhere
  • Testing becomes impossible

With specialized agents, each maintains focused state, uses optimized prompts, and fails independently. Our router agent uses Groq for fast classification (under 200ms), then delegates complex reasoning to Claude-3.5-Sonnet agents that might take 2-3 seconds but handle nuanced tasks.

The tradeoff: coordination overhead. Agents must pass context efficiently, handle partial failures, and avoid infinite delegation loops. We've found explicit state schemas (JSON) work better than natural language for inter-agent communication.

State Management: The Difference Between Toy and Production

Chat wrappers maintain conversation history. Agents maintain operational state. This distinction separates demos from production systems.

Consider our WhatsApp scheduling agent:

User: "Book a meeting with Sarah next Tuesday at 2pm"
Agent: "I'll check availability..."
[Agent queries calendar API, finds conflict]
Agent: "Sarah has a conflict at 2pm. She's free at 10am or 3pm. Which works?"
User: "Actually make it Wednesday instead"

Enter fullscreen mode Exit fullscreen mode

A chat wrapper would need the entire conversation to understand "it" refers to the meeting. Our agent maintains structured state:

{
  "pending_action": "schedule_meeting",
  "participants": ["user_123", "sarah_456"],
  "proposed_time": null,
  "constraints": ["tuesday_2pm_conflict"],
  "alternatives": ["tuesday_10am", "tuesday_3pm"]
}

Enter fullscreen mode Exit fullscreen mode

When the user says "Wednesday instead," the agent updates the specific field rather than reinterpreting everything. This approach:

  • Reduces token usage by 60-80%
  • Enables resuming conversations after connection drops
  • Allows other agents to understand ongoing tasks
  • Supports compliance logging

We persist this state in Oracle Autonomous JSON Database, which handles concurrent updates and provides ACID guarantees — critical when multiple agents might update the same user's state.

The LLM Is Just One Component

A common misconception: AI agents are just LLMs with extra steps. In our production systems, LLM calls represent maybe 20-30% of execution time.

Real agent loop timing breakdown (WhatsApp order processing):

  • Message decryption/validation: 50ms
  • State retrieval from cache/DB: 80-120ms
  • LLM decision call: 200-800ms (Groq) or 1-3s (Claude)
  • Business logic validation: 100ms
  • External API calls: 200ms-5s
  • State persistence: 50-100ms
  • Response encryption/sending: 50ms

The LLM provides reasoning capability, but agents need:

  • Message queue integration for reliable async processing
  • Caching layers to avoid repeated LLM calls
  • Circuit breakers for external dependencies
  • Retry logic with exponential backoff
  • Monitoring/alerting for production issues

Our Oracle Cloud infrastructure provides much of this — OCI Queue service for message handling, Redis for caching, and built-in monitoring. But even with good infrastructure, agent complexity lives in orchestration logic, not LLM prompts.

Multi-Agent Coordination: Beyond Pipeline Thinking

Single agents hit complexity ceilings. Multi-agent systems break through but introduce coordination challenges. The naive approach — agents calling each other like functions — creates brittle pipelines.

Our production pattern uses event-driven coordination:

  1. Agents publish state changes to a shared event bus
  2. Other agents subscribe to relevant event types
  3. A coordinator agent manages workflow-level concerns
  4. Each agent maintains local state, syncing through events

Example from our document processing system:

  • Upload agent receives PDF, publishes document_received event
  • OCR agent subscribes to this event, processes, publishes text_extracted
  • Classification agent takes extracted text, publishes document_classified
  • Multiple specialized agents handle different document types in parallel

This architecture handles partial failures gracefully. If the classification agent crashes, documents queue up but OCR continues. When classification recovers, it processes the backlog without losing work.

The challenge: event ordering and consistency. We use Oracle Streaming Service with exactly-once semantics and explicit sequence numbers. Agents checkpoint their progress, enabling clean recovery from any point.

Common Failure Modes and Mitigation

Production agents fail in predictable ways:

Context corruption: Agents lose track of conversation state or mix up users. Mitigation: Explicit session IDs, regular state validation, automatic reset after idle periods.

Infinite loops: Agent A delegates to Agent B who delegates back to Agent A. Mitigation: Loop detection via request IDs, maximum delegation depth, circuit breakers on agent communication.

Prompt injection: Users manipulate agents into unintended behaviors. Mitigation: Structured output formats (JSON schema validation), privilege separation between agents, sanitization of user inputs before prompt inclusion.

Cost explosion: Recursive agent calls or large context accumulation. Mitigation: Token budgets per interaction, cost attribution to user/session, automatic fallback to cheaper models.

Latency cascades: Slow responses compound in multi-agent flows. Mitigation: Aggressive timeouts, parallel processing where possible, caching of intermediate results.

Our monitoring tracks these failure modes explicitly. We measure not just success rates but loop detection triggers, context reset frequency, and cost per interaction. This data drives architectural improvements.

Building Your First Production Agent

Start with a single, focused agent that does one thing well. Our recommendation based on what works:

  1. Choose a narrow scope: "Schedule meetings via Telegram" beats "AI assistant for everything"
  2. Design state schema first: What must persist between interactions?
  3. Build the non-LLM parts: Message handling, state storage, external integrations
  4. Add LLM decision-making: Start with simple prompts, iterate based on real usage
  5. Implement monitoring early: Track decisions, not just errors

Avoid these common mistakes:

  • Starting with multi-agent systems before mastering single agents
  • Putting everything in prompts instead of code
  • Ignoring state management until it's too late
  • Optimizing LLM costs before validating the use case

The Reality of Production AI Agents

What is an AI agent? It's not a chatbot with API access or an LLM with a for-loop. It's a system that observes its environment, makes decisions, takes actions, and maintains state — reliably, at scale, with production constraints.

Our agents handle thousands of daily interactions across Telegram and WhatsApp, coordinate complex workflows, and integrate with enterprise systems. They're not perfect. They require constant monitoring, regular prompt tuning, and occasional manual intervention. But they deliver real value by automating tasks that would otherwise require human attention.

The key insight from running these systems: agents are software engineering challenges more than AI challenges. The LLM provides reasoning capability, but production value comes from reliable orchestration, state management, and system integration. Focus there, and agents become powerful tools rather than impressive demos.

— Elena Revicheva · AIdeazz · Portfolio