惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

N
Netflix TechBlog - Medium
G
Google Developers Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
T
The Blog of Author Tim Ferriss
Microsoft Azure Blog
Microsoft Azure Blog
GbyAI
GbyAI
L
LangChain Blog
云风的 BLOG
云风的 BLOG
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
aimingoo的专栏
aimingoo的专栏
P
Proofpoint News Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
小众软件
小众软件
WordPress大学
WordPress大学
A
About on SuperTechFans
大猫的无限游戏
大猫的无限游戏
C
Check Point Blog
月光博客
月光博客
Stack Overflow Blog
Stack Overflow Blog
美团技术团队
Jina AI
Jina AI
T
Tailwind CSS Blog
Google DeepMind News
Google DeepMind News
D
Docker

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
How to Add Caching to Any AutoGen Workflow in 2 Lines
Mahika jadhav · 2026-06-06 · via DEV Community

Mahika jadhav

AutoGen doesn't have a built-in execution cache. Every GroupChat, every ConversableAgent run starts fresh. If your multi-agent workflow runs similar tasks repeatedly — research pipelines, code review agents, scheduled reports — you're paying full LLM price every time.

Here's how to fix it without touching your AutoGen code.


The setup


bash
pip install mnemon-ai

import mnemon
mnemon.init()

# your existing AutoGen code — completely unchanged
import autogen

assistant = autogen.AssistantAgent(
    name="assistant",
    llm_config={"model": "gpt-4o", "api_key": "..."},
)
user_proxy = autogen.UserProxyAgent(
    name="user_proxy",
    human_input_mode="NEVER",
)

user_proxy.initiate_chat(
    assistant,
    message="Analyze Q2 sales data for Acme Corp and generate a summary report",
)
# second run with same or similar message: 2.66ms · 0 tokens · $0.00

Mnemon's MOTH layer patches AutoGen at startup. No agent changes, no conversation changes.

---
What gets cached

Every LLM call your agents make is intercepted. On repeat runs:

- Exact match — same message, instant response from cache
- Semantic match — "Analyze Q2 sales for Acme Corp" matches "Generate Q2 sales analysis for Acme" — same task, different phrasing

For multi-agent workflows where agents pass messages between each other, common sub-tasks (data parsing, formatting, summarization) hit the cache across different top-level goals.

---
For structured recurring workflows

If your AutoGen setup runs the same workflow repeatedly with varying inputs, use m.run() for segment-level caching:

import autogen, mnemon

m = mnemon.init()

def run_analysis(goal, inputs, context, capabilities, constraints):
    user_proxy.initiate_chat(assistant, message=goal)
    return user_proxy.last_message()["content"]

result = m.run(
    goal="Q2 sales analysis for Acme Corp",
    inputs={"quarter": "Q2", "client": "Acme Corp"},
    generation_fn=run_analysis,
)

print(result["tokens_saved"])   # tokens saved on this run
print(result["cache_level"])    # "system1" | "system2" | "miss"

---
Numbers

┌─────────┬────────────┬────────────┐
│         │ First run  │ Cached run │
├─────────┼────────────┼────────────┤
│ Tokens  │ ~1,250     │ 0          │
├─────────┼────────────┼────────────┤
│ Latency │ ~20s       │ 2.66ms     │
├─────────┼────────────┼────────────┤
│ Cost    │ full price │ $0.00      │
└─────────┴────────────┴────────────┘

At 80% hit rate on recurring workflows: 93% token reduction.

---
Install

pip install mnemon-ai           # exact match only
pip install mnemon-ai[full]     # + semantic matching (local, no API key)

import mnemon
mnemon.init()