惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Martin Fowler
Martin Fowler
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
A
About on SuperTechFans
Apple Machine Learning Research
Apple Machine Learning Research
The Register - Security
The Register - Security
Vercel News
Vercel News
H
Hackread – Cybersecurity News, Data Breaches, AI and More
人人都是产品经理
人人都是产品经理
MyScale Blog
MyScale Blog
云风的 BLOG
云风的 BLOG
博客园_首页
U
Unit 42
T
Tailwind CSS Blog
G
GRAHAM CLULEY
F
Full Disclosure
V
Vulnerabilities – Threatpost
T
Tenable Blog
月光博客
月光博客
P
Privacy & Cybersecurity Law Blog
P
Privacy International News Feed
K
Kaspersky official blog
Scott Helme
Scott Helme
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
N
News and Events Feed by Topic
T
The Exploit Database - CXSecurity.com
N
News and Events Feed by Topic
有赞技术团队
有赞技术团队
Recent Commits to openclaw:main
Recent Commits to openclaw:main
L
LINUX DO - 最新话题
Recorded Future
Recorded Future
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Help Net Security
Help Net Security
The GitHub Blog
The GitHub Blog
Cisco Talos Blog
Cisco Talos Blog
SecWiki News
SecWiki News
P
Proofpoint News Feed
Security Latest
Security Latest
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
罗磊的独立博客
S
Security Affairs
M
MIT News - Artificial intelligence
L
LINUX DO - 热门话题
美团技术团队
Simon Willison's Weblog
Simon Willison's Weblog
T
Threat Research - Cisco Blogs
Stack Overflow Blog
Stack Overflow Blog
Forbes - Security
Forbes - Security
Hugging Face - Blog
Hugging Face - Blog
博客园 - Franky
V
Visual Studio Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Hermes Agent: How Nous Research Built an AI That Actually Learns from Its Own
Ajay Mourya · 2026-06-01 · via DEV Community

If you've been following the AI agent ecosystem, you've probably noticed that most agent frameworks are running into the same limitation: memory.

The majority of today's agents are effectively stateless. The moment a session ends, they forget everything, including bugs they helped solve, architectural decisions, coding preferences, and workflow patterns. As a result, developers spend an increasing amount of time rebuilding context by pasting logs, re-explaining projects, and managing ever-expanding context windows.

Nous Research's Hermes Agent takes a fundamentally different approach.

Rather than treating every interaction as an isolated conversation, Hermes is built around a continuous learning loop. Designed to run locally or on lightweight server infrastructure, it can distill successful workflows into reusable skills, maintain long-term user preferences through its dialectic memory system, curate and refine knowledge in the background, and compress runtime experiences into high-quality training trajectories.

The result is an agent that doesn't simply execute tasks; it accumulates experience.

Instead of wrapping a language model inside a conventional chatbot interface, the Hermes team has built a highly extensible agent platform that actively learns from usage. It generates procedural skills from completed work, audits and organizes its own knowledge, and constructs a persistent model of the user over time.

In this article, we'll skip the installation walkthroughs and introductory demos. Instead, we'll dive directly into the hermes-agent codebase and perform a file-by-file audit of the architecture to understand how these learning systems work under the hood, how memory is implemented, and how Hermes attempts to solve one of the biggest limitations of modern AI agents.


1. Navigating the Codebase: The Big Picture

When you clone the repository, you will see a codebase that separates the user interface, execution runtime, tool integrations, and background automation:

hermes-agent/
├── run_agent.py               # AIAgent Class (The main engine and conversation loop)
├── cli.py                     # HermesCLI (The classic terminal interface)
├── model_tools.py             # Tool discovery, schema compilation, and call dispatching
├── toolsets.py                # Predefined bundles of permitted agent capabilities
├── hermes_state.py            # SessionDB (SQLite FTS5-backed local session store)
├── hermes_constants.py        # Path helpers (profile-aware get_hermes_home())
│
├── agent/                     # Modular Agent Internals
│   ├── conversation_loop.py   # Main multi-turn tool execution loop
│   ├── curator.py             # Background skill curation and consolidation daemon
│   ├── memory_manager.py      # Local vector recall and context injection
│   └── prompt_builder.py      # System prompts, soul-personas, and environment hints
│
├── tools/                     # Modular Tool Implementations
│   ├── registry.py            # Central self-registering tool registry
│   └── environments/          # Execution backends (Local, Docker, SSH, Modal, Daytona)
│
├── gateway/                   # Messaging Gateway (Telegram, Discord, Slack, WeChat)
│   └── run.py                 # Gateway server loop and command router
│
└── plugins/                   # Extensible Plugin Subsystem
    ├── hermes-achievements/   # Gamified local badge and share-card engine
    └── memory/                # Memory backends (Honcho, mem0, supermemory)

Enter fullscreen mode Exit fullscreen mode

The Unidirectional Tool Chain: No More Circular Imports

If you have ever built a complex Python application, you know how quickly import chains can turn into a messy spiderweb.

To solve this, Hermes implements a self-registering tool registry inside tools/registry.py. Instead of the main agent runner importing fifty different tool files, it reverses the flow:

[tools/registry.py] (Defines the ToolRegistry singleton; no external imports)
         ▲
         │ (Calls registry.register() at import-time)
  [tools/*.py]
         ▲
         │ (Static syntax scan via ast.parse() dynamically imports files)
 [model_tools.py]
         ▲
         │ (Queries registry for schema generation and dispatch)
[run_agent.py, cli.py]

Enter fullscreen mode Exit fullscreen mode

At startup, every python file inside the tools/ folder executes a module-level registry.register(...) call to declare its JSON schema, handler function, and environmental requirements.

Then, model_tools.py runs a fast Abstract Syntax Tree (ast.parse) scan over the files, dynamically loading only the modules that are registered. This keeps the core engine lightweight and lets you add a new capability by dropping a single file into the tools/ directory.


2. Under the Hood of the Agent Loop (run_agent.py)

When you send a prompt, the AIAgent class initiates a synchronous conversation loop inside run_conversation(). It is a classic tool-calling loop, but with a few clever engineering guardrails:

                  AIAgent.run_conversation(user_message)
                                     │
                                     ▼
                      [Session state initialization]
                  - Pull system prompts & Soul profiles
                  - Inject workspace file context
                  - Trigger Memory Provider recall
                                     │
                                     ▼
                ┌────────────────────────────────────────┐
                │        Standard LLM API Invocation     │
                └───────────────────┬────────────────────┘
                                    │
                         Is there a Tool Call?
                       ◄─────────────────────►
                       Yes                  No
                        │                    │
                        ▼                    ▼
             [Parallel execution]    [Deliver final response]
             - Check environment     - Record trajectory log
             - Execute handlers      - End loop iteration
             - Return results        
                        │
                        ▼
            [Increment api_call_count]
            - Check budget constraints
            - Recurse back to LLM Call

Enter fullscreen mode Exit fullscreen mode

Preventing the Surrogate Pair Crash

LLMs can get messy when dealing with raw terminal outputs or binary file dumps. If a shell tool outputs non-ASCII symbols, wild terminal escape sequences, or incomplete surrogate pairs, cloud API endpoints (like OpenAI or Anthropic) will often reject the payload, causing your entire run to crash.

Hermes handles this defensively in agent/message_sanitization.py. Before any API call goes over the wire, it sweeps the message array, dynamically stripping out raw ANSI terminal colors, sanitizing surrogate blocks, and automatically truncating giant stdout outputs into external log files.

If it truncates something, it leaves a clean text pointer, such as: Output truncated. Full logs written to local file path. This lets the agent know the file exists but does not waste precious context tokens reading it.


3. The Skills Curator: How Hermes Tidies Its Own Mind

Let's talk about how Hermes learns. If you walk the agent through a complex, multi-step debugging flow, like configuring a specific database connection, you can tell it to save that workflow as a permanent Skill. The agent runs the workflow-skill-creator tool and writes a clean, structured Markdown folder under .hermes/skills/.

But here is the catch: if your agent creates a new file for every single bug it solves, its directory will quickly become cluttered. This leads to slow search queries and redundant instructions.

Hermes fixes this using its background Curator (agent/curator.py).

       [Skills Library] (~/.hermes/skills/)
              │
      Is the Agent idle?
      Was the last Curator run > 7 days ago?
              │
              ▼
    [Apply Automatic Transitions]
    - Mark untouched skills as STALE (>30 days inactive)
    - Move STALE skills to ARCHIVE (>90 days inactive)
              │
              ▼
    [Spawn Background Review Agent]
    - Read the remaining active skills
    - Scan for name overlaps and prefix clusters
    - Reorganize skill assets via consolidation
              │
              ▼
    ┌──────────────────────────────────────────────┐
    │       Umbrella Skill Synthesis               │
    │  - Patches sibling instructions into one     │
    │  - Demotes support scripts to scripts/       │
    │  - Demotes raw notes to references/          │
    │  - Archives the original micro-skills        │
    └──────────────────────────────────────────────┘

Enter fullscreen mode Exit fullscreen mode

The Weekly Spring Cleaning

When your agent is completely idle, a weekly background timer triggers apply_automatic_transitions(). First, it runs a fast metadata audit to mark skills untouched for 30 days as STATE_STALE. If a skill sits untouched for 90 days, the engine moves the entire folder to a .archive/ directory.

Consolidating into Umbrellas

Next, it boots an auxiliary model pass to sweep the active library for redundant clusters, like multiple files matching mcp-* or git-*. The CURATOR_REVIEW_PROMPT directs the LLM to consolidate these into Umbrella Skills:

  1. Merging Instructions: It extracts the core steps of similar micro-skills and merges them into a single, master SKILL.md umbrella document.
  2. Sorting Assets: It organizes supporting files, demoting raw documentation to B's references/ folder and helper scripts to scripts/.
  3. Forwarding Links: It archives the original narrow files and tells the SQLite database to point future queries directly to the parent umbrella.

This background curation means the agent's procedural memory stays clean, organized, and cheap to search.


4. Dialectic Memory: Evolving Developer Profiles

For long-term memory, many frameworks just run a simple vector database lookup over past messages. The problem is that developer goals change. If you were working on a Python project last month, but you are writing Rust today, a basic search might pollute the context window with old Python snippets.

Hermes tackles this by integrating Honcho (plugins/memory/honcho/), a memory backend that uses a two-layer, dialectic reasoning system.

                      [User Message Received]
                                │
                 Injected every N turns (contextCadence)
                                ▼
         ┌──────────────────────────────────────────────┐
         │            Layer 1: Base Context             │
         │ - Session Summary                            │
         │ - Evolving User Representation (Honcho profile)│
         │ - Factual User/AI Peer cards                 │
         └──────────────────────┬───────────────────────┘
                                │
                 Injected every M turns (dialecticCadence)
                                ▼
         ┌──────────────────────────────────────────────┐
         │          Layer 2: Dialectic Supplement       │
         │ - Evolving summary of active session topics │
         │ - Multi-pass dialectic audit output          │
         └──────────────────────┬───────────────────────┘
                                ▼
         Injected into USER message wrapped in XML tags

Enter fullscreen mode Exit fullscreen mode

Saving Prompt Cache Budgets

Updating the system prompt on every single turn invalidates the KV prompt cache on modern LLM endpoints. This slows down response times and spikes costs.

Hermes side-steps this by injecting memory context directly into the user message wrapped in <memory-context> XML tags. The system prompt remains static and the cache stays warm.

The Dialectic Reflection Loop

Honcho runs an active reflection loop over your chat logs using three levels of depth (dialecticDepth):

  • Depth 1 (Fast Summary): Writes a quick summary of active session topics.
  • Depth 2 (Self-Audit): Evaluates the summary to check for accuracy. If the summary is strong, it finishes the run early to save tokens.
  • Depth 3 (Reconciliation): Resolves contradictions. If you suddenly pivot from writing React to Vanilla CSS, Depth 3 spots the change, flags your old React preferences as stale, and rewrites the context injection to favor Vanilla CSS.

5. Trajectory Compression: Squeezing Logs into Gold

AI models excel at tool-calling when they are fine-tuned on real-world developer runs, which are also known as trajectories. But developer sessions are incredibly verbose, easily stretching past standard context limits.

To solve this, Hermes packages a high-performance Trajectory Compressor inside trajectory_compressor.py. It uses a clever sandwich compression strategy to shrink historic runs to fit tight token budgets while preserving crucial training signals:

Original Trajectory Logs:
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ System & Setup  │ │ Middle Turns    │ │ Middle Turns    │ │ Conclusion      │
│ (Turns 1 - 3)   │ │ (Turns 4 - 20)  │ │ (Turns 21 - 40) │ │ (Last 4 Turns)  │
└────────┬────────┘ └────────┬────────┘ └────────┬────────┘ └────────┬────────┘
         │                   │                   │                   │
         ▼                   └─────────┬─────────┘                   ▼
      PROTECTED                        │                          PROTECTED
    (Keep intact)                      ▼                        (Keep intact)
                              [AUXILIARY MODEL]
                        Compresses middle turns into
                         a factual context summary
                                       │
                                       ▼
Compressed Trajectory File:
┌─────────────────┐ ┌─────────────────────────────────────┐ ┌─────────────────┐
│ System & Setup  │ │ [CONTEXT SUMMARY]: Unified summary  │ │ Conclusion      │
│ (Turns 1 - 3)   │ │ of all intermediate terminal calls  │ │ (Last 4 Turns)  │
└─────────────────┘ └─────────────────────────────────────┘ └─────────────────┘

Enter fullscreen mode Exit fullscreen mode

  1. Protecting Key Boundaries: The compressor locks the setup turns (the system prompt, initial human question, first tool choice) and the final conclusion turns (last $N$ steps showing the working code and check results) in place.
  2. Token Sweeper: It tokenizes the intermediate turns using the moonshotai/Kimi-K2-Thinking tokenizer. If the payload is over the target threshold, it marks the middle turns for compression.
  3. Context Synthesizer: The middle turns are compiled and sent to an auxiliary model. The prompt instructs the model to act as a neutral summarizer, writing a dense, factual summary containing the exact variables checked, tools executed, and files modified.
  4. Re-Assembling the Sandwich: The original middle turns are replaced with a single, highly compressed message containing the [CONTEXT SUMMARY]: prefix.

This compressed format preserves perfect semantic continuity. A training run studying this log sees the initial problem setup, a dense overview of the intermediate actions, and the exact final execution result. This makes these outputs incredibly valuable for Supervised Fine-Tuning (SFT) and Reinforcement Learning (RLHF) to train future tool-calling models.


6. Gamifying Your Terminal: Hermes Achievements

A great agent is not just about robust backends, it is also about developer experience. Hermes bundles a native Achievements Plugin under plugins/hermes-achievements/ that parses the local SQLite SessionDB and rewards you with tiered badges:

  • Let Him Cook / Toolchain Maxxer: Earned when you let the agent execute long, autonomous multi-step tool runs to solve complex programming challenges.
  • Red Text Connoisseur: Unlocked when the agent encounters system/compiler errors in the terminal and successfully edits files to recover without developer intervention.
  • Port 3000 Is Taken: Triggered when the agent diagnoses blocked network ports during local web server setups and dynamically re-routes configurations.

Snapshot Caching

To keep the CLI fast, the plugin uses a snapshot caching system with incremental checkpoints. Once a badge is unlocked, it writes the state to state.json. Future sweeps only scan new session logs generated since the last checkpoint, keeping dashboard load times under 50 milliseconds. You can then render these badges as beautiful 1200×630 OpenGraph share cards via a local HTML5 canvas, ready to share on social channels.


The Verdict: A Blueprint for What's Next

Taking a look under the hood of hermes-agent reveals an engine built for real-world development. By shifting past stateless wrappers, Nous Research has created a robust blueprint for self-improving systems:

  1. Logical Separation: Separating the CLI, React Ink terminal TUI, and messaging Gateway keeps execution clean and persistent.
  2. Mental Hygiene: The Curator and Skills system ensure the agent's procedural library remains highly accurate and organized over time.
  3. Smart Personalization: The Honcho provider maps platform IDs to evolving user profiles across devices without losing prompt cache performance.
  4. Data Generation: The Trajectory Compressor turns daily work sessions into rich fine-tuning datasets, creating a true self-improving loop.

Hermes Agent is a glimpse into the future of software development: a world where our tools don't just run code, but actively learn how to build it alongside us.