惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Recent Commits to openclaw:main
Recent Commits to openclaw:main
N
News | PayPal Newsroom
TaoSecurity Blog
TaoSecurity Blog
Google Online Security Blog
Google Online Security Blog
NISL@THU
NISL@THU
T
Threatpost
C
CXSECURITY Database RSS Feed - CXSecurity.com
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Engineering at Meta
Engineering at Meta
AWS News Blog
AWS News Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
P
Privacy International News Feed
B
Blog
PCI Perspectives
PCI Perspectives
Martin Fowler
Martin Fowler
Spread Privacy
Spread Privacy
P
Proofpoint News Feed
T
Tenable Blog
F
Fortinet All Blogs
G
GRAHAM CLULEY
V2EX - 技术
V2EX - 技术
C
Check Point Blog
Project Zero
Project Zero
P
Palo Alto Networks Blog
J
Java Code Geeks
W
WeLiveSecurity
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
The Exploit Database - CXSecurity.com
博客园 - 司徒正美
P
Privacy & Cybersecurity Law Blog
S
SegmentFault 最新的问题
Last Week in AI
Last Week in AI
Forbes - Security
Forbes - Security
C
Cybersecurity and Infrastructure Security Agency CISA
Security Latest
Security Latest
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Vercel News
Vercel News
Recent Announcements
Recent Announcements
博客园 - Franky
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Recorded Future
Recorded Future
The Last Watchdog
The Last Watchdog
MongoDB | Blog
MongoDB | Blog
人人都是产品经理
人人都是产品经理
酷 壳 – CoolShell
酷 壳 – CoolShell
Cisco Talos Blog
Cisco Talos Blog
量子位
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
GitHub - audi0417/agent-engineering-roadmap: Bilingual hands-on roadmap for production-aware AI agents: MCP, memory, RAG, workflows, evaluation, safety, and agent colonies.
Audi0417 · 2026-06-26 · via Hacker News - Newest: "AI"

Agent Engineering Roadmap cover

Traditional Chinese English GitHub Stars GitHub Forks Last Commit License Runnable examples status

Agent Engineering MCP Ready Memory Systems Multi-Agent Workflow Agent Colony Status

A hands-on roadmap for building production-ready AI Agents, MCP Servers, Memory Systems, Multi-Agent Workflows, and Agent Colonies.

繁體中文 · Website · Course · Roadmap · Examples · Showcases · Benchmarks · Labs · Teaching · Templates · Architecture · Healthcare · Finance


Agent Engineering Course Map


flowchart LR
    User((User)) --> Agent[AI Agent]
    Agent --> Tools[Tool Use]
    Tools --> MCP[MCP Layer]
    MCP --> Memory[Memory System]
    Memory --> Workflow[Agent Workflow]
    Workflow --> MultiAgent[Multi-Agent Team]
    MultiAgent --> Colony[Agent Colony]
    Colony --> Production[Production AI App]
Loading

Why this roadmap exists

Most AI tutorials stop at prompts, RAG, or simple tool calling.

Real agentic products require more than that:

  • agents that can use tools safely
  • MCP servers that connect agents to real systems
  • memory layers that persist useful context
  • workflows that are observable and controllable
  • multi-agent teams that can specialize and collaborate
  • evaluation, security, and production guardrails

This repository is a practical learning path for builders who want to move from chatbot demos to real agent engineering.


Teaching approach

This roadmap teaches agents like an engineering course, not a tool catalog.

Each major topic follows the same pattern:

  1. Start with the problem: what breaks if you only use a chatbot?
  2. Build the intuition: what is the simplest mental model?
  3. Open the box: what components are actually involved?
  4. Run a minimal example: what can you inspect locally?
  5. Add production judgment: what needs evaluation, observability, approval, or safety gates?

In one sentence: an agent is not magic. It is context, tools, memory, workflow, evaluation, and human judgment arranged around a useful task.


What you will learn

Level Topic Outcome
0 AI & LLM Fundamentals Understand LLM apps, embeddings, RAG, and structured output
1 Single Agent Build a task-focused agent with a clear role and output format
2 Tool Use Connect agents to external tools and APIs
3 MCP Build and use MCP clients, servers, tools, resources, and prompts
4 Agent Memory Design short-term, episodic, semantic, user, and shared memory
5 Agent Workflow Build reliable planning, execution, review, retry, and approval flows
6 Multi-Agent Systems Coordinate specialized agents using supervisor, debate, and reflection patterns
7 Agent Colony Build shared-memory colonies with domain agents and evaluation loops
8 Production & Safety Deploy agents with observability, evaluation, security, and cost control

Course materials

Section Purpose
Course Complete syllabus and graduation criteria
Curriculum Concept chapters from foundations to production
Visual Assets SVG diagrams for teaching and slides
Roadmap Level-by-level learning milestones
Examples Runnable minimal implementations
Benchmarks Lightweight checks for tool use, RAG, workflow, security, and observability
Showcases Dependency-free demos for healthcare, finance, and enterprise workflows
Domain Casebooks Healthcare, finance, and enterprise case studies with eval cases
Labs Guided exercises for each stage
Teaching Layer Teaching audit, misconceptions, deliverables, and module blueprint
Lab Solution Guides Solution shapes and grading direction for hands-on labs
Lesson Plans Instructor-ready teaching plans for each module
Study Group Kit 4-week, 8-week, and workshop formats for cohorts
Patterns Reusable agent architecture patterns
Templates Agent specs, memory policies, evals, and safety gates
Papers Research papers, reading roadmap, and engineering notes
Open Source Projects Curated ecosystem map for frameworks, MCP, RAG, evals, observability, and ops
Framework Selection Matrix Choose agent frameworks by engineering tradeoff
Open Source Reading Guide Learn how to study real agent repositories
DeepEval And RAGAS Practical guide to LLM and RAG evaluation frameworks
Release Checklist v1 release verification and project hygiene
Assessments Quiz bank and rubrics
Capstone Final project for building a production-aware colony
Portfolio Projects Project ideas with deliverables, evals, and open-source references
Capstone Starter Runnable starter scaffold for the final project
Glossary Core terms and definitions

The learning path

AI Fundamentals
      ↓
Single Agent
      ↓
Tool Use
      ↓
MCP Integration
      ↓
Agent Memory
      ↓
Agent Workflow
      ↓
Multi-Agent Systems
      ↓
Agent Colony
      ↓
Production, Evaluation & Safety

Try it in 60 seconds

Run a showcase without API keys:

python showcases/enterprise-support-agent/main.py
python showcases/finance-research-agent/main.py
python showcases/healthcare-agent-colony/main.py

Then run the evaluation harness:

python examples/07-evaluation-harness/main.py
python examples/08-mini-rag/main.py
python benchmarks/benchmark_runner.py
python scripts/verify_examples.py

Production readiness artifacts

Artifact Use
Agent Registry Template Register owner, scopes, tools, data, evals, and operations
Risk Assessment Template Classify agent risk before launch
Deployment Review Template Check release gates and operational readiness
Release Checklist Prepare a public course release
v1.0 Readiness Track stable release readiness

Showcase demos

Demo Shows
Enterprise Support Agent Ticket routing, risk classification, approval gates
Finance Research Agent Research support, assumptions, risk boundaries
Healthcare Agent Colony Safety boundaries, escalation, medical-advice avoidance

Runnable examples

Example Shows No API key
01 Single Agent Role, task boundary, structured output Yes
02 Tool-Using Agent Local tool call and validation Yes
03 MCP-style Agent Client/server tool boundary Yes
04 Memory Agent Memory write/retrieve policy Yes
05 Multi-Agent Workflow Planner, researcher, writer, reviewer Yes
06 Agent Colony Supervisor, domain agent, evaluator Yes
07 Evaluation Harness Regression eval suite Yes
08 Mini RAG Retrieval, grounded answer, RAG eval Yes
09 Graph Approval Agent Graph transitions, approval gate, production eval Yes
10 Observable Agent Trace events, guardrail logs, replayable debugging Yes
11 Prompt Injection Defense Untrusted retrieval filtering and security eval Yes
12 Cost-Aware Agent Model routing, budget, latency, fallback eval Yes
13 Durable Workflow Agent Checkpoint, resume, durable workflow eval Yes
14 Modern MCP Gateway Tools, resources, prompts, auth, elicitation Yes
15 Memory Governance Agent Memory redaction, merge, decay, deletion, audit Yes
16 Agent Permission System Agent identity, scopes, access review, audit Yes
17 Advanced Eval Harness Regression, safety, adversarial, golden trace release gate Yes
Capstone Starter Starter colony demo and regression eval Yes

Run every dependency-free example with:

python scripts/verify_examples.py

README widgets used

This README uses lightweight visual widgets commonly seen in popular GitHub projects:

  • Local cover image for the top hero banner
  • shields.io for stars, forks, language, status, and topic badges
  • Mermaid for architecture diagrams

Plugin ecosystem

Agent Engineering is not only about prompts. A production agent needs a plugin ecosystem around it.

Category Purpose Example Plugins / Tools
MCP Servers Standardized access to tools and data filesystem, database, browser, GitHub, Slack, Google Drive
Memory Persistent context and retrieval Qdrant, LanceDB, Chroma, PostgreSQL, Redis
Orchestration Workflow and multi-agent control LangGraph, CrewAI, AutoGen, OpenAI Agents SDK
RAG Knowledge retrieval and grounding LlamaIndex, LangChain, Haystack
Observability Tracing, debugging, monitoring Langfuse, OpenTelemetry, Helicone, Phoenix
Evaluation Quality and safety testing DeepEval, RAGAS, promptfoo, custom eval suites
Guardrails Safety and structured validation Guardrails AI, Pydantic, JSON Schema, policy checkers
UI / App Layer User-facing agent applications Streamlit, Gradio, Next.js, FastAPI
Domain Tools Industry-specific integrations healthcare records, finance data, CRM, ERP, ticketing systems

Core architecture

graph TD
    User[User] --> Supervisor[Supervisor Agent]
    Supervisor --> Planner[Planner]
    Planner --> MemoryAgent[Memory Agent]
    Planner --> ResearchAgent[Research Agent]
    Planner --> ToolAgent[Tool Agent]
    Planner --> DomainAgent[Domain Agent]
    MemoryAgent --> SharedMemory[Shared Memory]
    ToolAgent --> MCP[MCP Servers]
    DomainAgent --> MCP
    ResearchAgent --> MCP
    MCP --> PluginLayer[Plugin Ecosystem]
    PluginLayer --> Databases[Databases]
    PluginLayer --> Documents[Documents]
    PluginLayer --> APIs[External APIs]
    PluginLayer --> SaaS[SaaS Apps]
    Supervisor --> Evaluator[Evaluator Agent]
    Evaluator --> Final[Final Response]
    Final --> User
    Evaluator --> SharedMemory
Loading

Repository structure

agent-engineering-roadmap/
├── README.md
├── README_zh.md
├── COURSE.md
├── assets/           # Visual diagrams and teaching images
├── roadmap/          # Level 0-8 learning path
├── curriculum/       # Full course chapters
├── examples/         # Hands-on examples
├── benchmarks/       # Lightweight behavior checks
├── security/         # Prompt injection and agent security labs
├── study-groups/     # Cohort and workshop facilitation kit
├── showcases/        # Shareable demos with sample outputs
├── labs/             # Guided exercises
├── lesson-plans/     # Instructor-ready lesson plans
├── patterns/         # Architecture pattern catalog
├── architecture/     # System design patterns
├── templates/        # Reusable agent and MCP templates
├── assessments/      # Quiz bank and rubrics
├── projects/         # Capstone and portfolio projects
├── glossary/         # Agent engineering terms
├── healthcare/       # Healthcare agent engineering track
├── finance/          # Finance and quantitative research track
├── resources/        # Curated learning resources
├── docs/             # GitHub Pages site
└── launch-kit/       # Launch copy, topics, and checklist

Real-world tracks

Healthcare Agent Engineering

Build agent systems for care management, nutrition tracking, personal health memory, and healthcare workflow automation.

Example colony:

Care Manager Agent
├── Nutrition Agent
├── Vital Sign Agent
├── Psychology Agent
├── Medication Agent
├── Memory Agent
└── Safety Evaluator Agent

Finance Agent Engineering

Build research agents, factor-analysis agents, portfolio agents, risk agents, and trading research workflows.

Example colony:

Research Agent
├── Market Data Agent
├── Factor Analysis Agent
├── Portfolio Agent
├── Risk Agent
└── Report Agent

Enterprise Agent Engineering

Build customer support agents, internal knowledge agents, document agents, workflow automation agents, and evaluation pipelines.


Design principles

  1. Agents should be useful before they are autonomous.
  2. Memory should be intentional, auditable, and safe.
  3. MCP should be treated as an integration layer, not just a plugin mechanism.
  4. Multi-agent systems should reduce complexity for users, not create complexity for developers.
  5. Production agents need evaluation, observability, cost control, and human approval gates.

Project roadmap

  • Initialize bilingual repository structure
  • Add Level 0-8 roadmap skeleton
  • Add architecture documents
  • Add healthcare and finance tracks
  • Add README badges and hero banner
  • Expand each roadmap level into handbook chapters
  • Add minimal runnable examples
  • Add MCP server templates
  • Add memory system examples
  • Add agent colony demo
  • Add evaluation and safety templates
  • Add full course syllabus
  • Add observable agent and prompt injection defense examples
  • Add benchmark runner and study group kit
  • Add cost, durable runtime, and modern MCP gateway modules
  • Add memory governance, identity permission, and incident response modules
  • Add advanced eval, product UX, and enterprise operating model modules
  • Add guided labs
  • Add instructor-ready lesson plans
  • Add pattern catalog
  • Add quiz bank, rubrics, glossary, and capstone
  • Add full healthcare agent colony application
  • Add full finance research agent application

Who this is for

  • AI engineers
  • LLM application developers
  • Startup builders
  • Researchers building agent systems
  • Product teams moving from chatbot demos to real workflows
  • Developers interested in MCP, memory, and multi-agent systems

License

This project is licensed under the MIT License.