惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

人人都是产品经理
人人都是产品经理
Stack Overflow Blog
Stack Overflow Blog
L
LINUX DO - 最新话题
Google Online Security Blog
Google Online Security Blog
Schneier on Security
Schneier on Security
Spread Privacy
Spread Privacy
www.infosecurity-magazine.com
www.infosecurity-magazine.com
雷峰网
雷峰网
Google DeepMind News
Google DeepMind News
Microsoft Azure Blog
Microsoft Azure Blog
IT之家
IT之家
V
Vulnerabilities – Threatpost
K
Kaspersky official blog
S
Schneier on Security
B
Blog
The Register - Security
The Register - Security
SecWiki News
SecWiki News
Hacker News: Ask HN
Hacker News: Ask HN
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
S
Security Affairs
T
The Blog of Author Tim Ferriss
G
Google Developers Blog
T
Tenable Blog
P
Proofpoint News Feed
Apple Machine Learning Research
Apple Machine Learning Research
D
DataBreaches.Net
S
Secure Thoughts
Security Latest
Security Latest
H
Heimdal Security Blog
The Hacker News
The Hacker News
O
OpenAI News
AWS News Blog
AWS News Blog
量子位
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
腾讯CDC
U
Unit 42
L
Lohrmann on Cybersecurity
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
L
LangChain Blog
阮一峰的网络日志
阮一峰的网络日志
T
The Exploit Database - CXSecurity.com
NISL@THU
NISL@THU
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Hugging Face - Blog
Hugging Face - Blog
The Last Watchdog
The Last Watchdog
Recorded Future
Recorded Future
V2EX - 技术
V2EX - 技术
爱范儿
爱范儿
F
Full Disclosure

ByteByteGo Newsletter

A Beginner’s Guide to Clocks, Causality, and Ordering in Distributed Systems Best Practices for Building AI Agents That Work in Production Inside Roblox’s Bet on World Models A Guide to Multi-Tenancy: Benefits and Challenges AI Customer Support at Scale: The Travel Industry’s $Billion Bet How LLMs Learn to Be Helpful (RLHF vs DPO) How Microsoft Ships AI Agents at Enterprise Scale EP221: How Docker Works Under the Hood LAST CALL FOR ENROLLMENT: Become an AI Engineer - Cohort 7 Streaming vs Batch: Two Philosophies of Data Processing The Agent Loop: How AI Goes From Answering Questions to Doing Things ChatGPT vs Gemini vs Claude: How They Differ LAST CALL FOR ENROLLMENT: Become an AI Engineer - Cohort 7 Proof of Human: How to Verify a Person Is Real and Unique Multi-Region Architecture: Going Global Without Going Broke How OpenAI Delivers Low-Latency Voice AI for 900M Users Inside Thinking Machines’ Interaction Models How AI Agents Manage Memory and Avoid Forgetfulness EP220: RAG vs Graph RAG vs Agentic RAG Top Anti-Patterns to Avoid in Service Architecture Large Language Models vs Small Language Models An Ex-Meta L8’s Agentic Engineering Setup AI-Native Leaders: The Organizational Playbook for Engineering Transformation at Scale EP219: 12 Open-source LLMs Observability for Beginners: Logs, Metrics, Traces, and Everything Around Them LAST CALL FOR ENROLLMENT: Build with Claude Code - Cohort 2 How Open-Weight Models Changed the AI Landscape A Guide to AI Inference Engineering EP218: The Typical AI Agent Stack, Explained Must- Know Deployment Strategies: From Big-Bang to Progressive Delivery Love Teaching? ByteByteGo Is Hiring Part-Time AI & Engineering Instructors What Salesforce Learned from 20,000 Enterprise Agent Deployments Token Spend Out of Control? The Case for Smarter Routing EP217: Latency vs Throughput vs Bandwidth The Path of a Request: A Tour of Modern Web Architecture How OpenAI Built Its Data Agent A Practical Guide to Becoming an AI-Native Engineer How DoorDash Built a Testing System to Evaluate LLMs Must-Know Failure Modes in Distributed Systems How Airtable Built the Search Layer Behind Their AI Features How Vercel Cut Build Wait Times From 90 Seconds To 5 How CockroachDB Built Vector Indexing at Scale EP216: RAGs vs Agents 🚀 New cohort based course launch: Build with Claude Code A Guide to Async Patterns in API Design How Netflix is Using Multimodal AI to Power Video Search How Snapchat Serves a Billion Predictions Per Second How Grab is Using AI Agents to Boost Team Productivity EP215: The Anatomy of an AI Agent LAST CALL FOR ENROLLMENT: Become an AI Engineer - Cohort 6 A Guide To Event-Driven Architectural Patterns High Performance Rate Limiting at Databricks How Figma Upgraded Data Pipeline from Multi-Day Latency to Real-Time How Pinterest Built a Production MCP Ecosystem EP214: Claude Code vs. OpenClaw: 5 Design Dimensions Become an AI Engineer | Enrollment Ends Soon Container Design Patterns for Distributed Systems How Instacart Built a Search for Billions of Products Connecting LLMs to the Real World: Tool Use, Function Calling, and MCP EP213: MCP vs Skills, Clearly Explained A Beginner’s Guide to Kubernetes The Tech Stack Powering Wise How Stripe Detects Fraudulent Transactions Within 100 ms How Amazon Uses LLMs to Recommend Products EP212: Data Warehouse vs Data Lake vs Data Mesh B-Trees vs LSM Trees: Comparison and Trade-Offs How DoorDash Launches a New Country in One Week The Security Architecture of GitHub Agentic Workflow EP211: How the JVM Works A Guide to Relational Database Design Figma Design to Code, Code to Design: Clearly Explained How LinkedIn Feed Uses LLMs to Serve 1.3 Billion Users EP210: Monolithic vs Microservices vs Serverless Must-Know Cross-Cutting Concerns in API Development How Spotify Ships to 675 Million Users Every Week Without Breaking Things Nextdoor’s Database Evolution: A Scaling Ladder A Guide to Context Engineering for LLMs EP209: 12 Claude Code Features Every Engineer Should Know Our New Book on Behavioral Interviews Is Now Available on Amazon Database Performance Strategies and Their Hidden Costs How Datadog Redefined Data Replication How Meta Turned Debugging Into a Product How Roblox Uses AI to Translate 16 Languages in 100 Milliseconds EP208: Load Balancer vs API Gateway LAST CALL FOR ENROLLMENT: Become an AI Engineer - Cohort 5 How to Implement API Security How Anthropic’s Claude Thinks How Netflix Live Streams to 100 Million Devices in 60 Seconds How Agentic RAG Works? Last Chance to Enroll | Become an AI Engineer | Cohort-Based Course EP207: Top 12 GitHub AI Repositories Event Sourcing Explained: Benefits and Use Cases How OpenAI Codex Works
MCP vs A2A vs ACP: How AI Agents Actually Talk to Each Other
ByteByteGo · 2026-07-19 · via ByteByteGo Newsletter

As AI agents take on decision-making, they need awareness of real-world conditions.

A weather delay can reroute deliveries. Lightning can pause field operations. Road conditions can affect autonomous vehicles. These signals shape decisions, yet most AI systems can’t observe them.

But context is only as valuable as the data behind it. Enterprise applications need trusted, real-time weather intelligence—not just generic forecasts. Signals like lightning activity, road conditions, hail risk, and weather impacts often matter more than temperature alone.

Xweather’s MCP-ready weather API gives AI agents access to trusted weather intelligence and the real-world context behind weather-driven decisions. Learn how it works in this technical guide.

Get your free API key

This week’s system design refresher:

  • MCP vs A2A vs ACP: How AI Agents Actually Talk to Each Other

  • 8 Frontier Open Models I’m Most Excited About

  • LLM vs RAG vs Agent evals

  • How Distributed Tracing Works at the High Level?

  • The Life of a Redis Query

Agents are capable on their own. Combined with tools and other agents, their capabilities compound. But how should they communicate?

Image
  1. MCP: agent to tool communication.
    The host app receives the user request, its embedded MCP client formats and routes it to the right MCP server, the server executes the tool call and returns a structured response. The agent uses the result to continue reasoning.

  2. A2A: agent to agent communication.
    An agent that can't complete a task alone discovers a capable peer via its Agent Card (published at a well-known URL), delegates the task, and receives a structured result back. If the second agent needs more input mid-task, it pauses in the input-required state and loops back to the first.

  3. ACP: agent to agent communication over REST (merged into A2A).
    ACP took a REST-first approach. Peers were discovered through an Agent Manifest, called directly over HTTP, and responded to sync for low-latency tasks or via async SSE stream.

In production, MCP and A2A are complementary. MCP handles tool access, A2A handles agent communication.

Are you running MCP and A2A together in your agent stack?

Attio, the agentic CRM, makes it incredibly easy for anyone to run workflows for any GTM play they need.

Describe what you want, and Attio builds it. I just built a workflow that runs every morning, surfaces the deals that need my attention today, like anything with a stage change or a new signal in the last 24 hours.

Hundreds of thousands of automations already run on Attio every day. Ready to try now?

Try it today

LLMs, RAG pipelines, and agents are different systems, but the recipe for evaluating them is the same: pick a task, collect eval data, develop a grader.

The diagram below breaks it down across four popular AI systems.

  1. LLMs: Input is a prompt, output is text. Tasks focus on what a raw model can do on its own: safety, code, instruction-following. Grader is often LLM-as-judge on the final answer.

  2. RAG: Adds a retriever before the LLM. So you grade two things: retrieval (right docs?) and generation (faithful answer?).

  3. Coding Agents: The LLM gets tools and a loop. Tasks become end-to-end like bug fixing and long-horizon planning. Grading is mostly code-based: run unit tests on the final patch.

  4. Multi-Agent Systems: Multiple agents coordinating through an orchestrator. Tasks shift to coordination and role adherence. Grading blends code tests, LLM-as-judge, and human review.

Every new component in the pipeline is a new place for things to go wrong, and a new thing your evals need to catch.

Over to you: What's your go-to evaluation metric for multi-agent systems?

table
  • Inkling (Thinking Machines): released this week, now the strongest American open model. Text, image, and audio input.

  • Nemotron 3 Ultra (NVIDIA): a solid choice for long-running agents. The Mamba hybrid keeps long-context inference cheap.

  • GLM-5.2 (Z.ai): currently the best open model for coding.

  • Kimi K2.6 (Moonshot): strong on long agent tasks. Holds up over hundreds of tool calls.

  • DeepSeek-V4 Pro: a very cheap way to get frontier-level quality over an API.

  • Qwen3.6-35B (Alibaba): the best model to run on your own machine. A single 24 GB GPU is enough.

  • Gemma 4 31B (Google): the best choice for on-device multimodal. Takes image and audio input on a gaming GPU.

  • MiniMax M3: the only open model with native video input.

Over to you: Which open model are you most excited about?

  1. Services generate telemetry data (traces, logs, metrics) as they handle requests.

  2. The OpenTelemetry Collector receives this data from all services in a unified format.

  3. The collector splits the data into three streams: traces, logs, and metrics.

  4. Each stream is sent to a Receive & Process unit that prepares it for storage and analysis.

  5. Processed data is stored in a Log Database for querying and long-term access.

  6. Data from the database is visualized through a Visualization dashboard for monitoring and debugging.

Over to you: What else will you add to better understand distributed tracing?

Redis is an in-memory database, which means all data lives in RAM for speed. However, if the server crashes or restarts, data could be lost.

Image

To solve this problem, Redis provides two persistence mechanisms to write data to disk:

  1. AOF (Append-Only File)

    When a client sends a command, Redis first executes it in memory (RAM). After that, Redis logs the command by appending it to an AOF file on disk. This ensures every operation can be replayed later to rebuild the dataset. Since the command is executed first and logged afterward, writes are non-blocking. The recovery process uses the event log to replay the recorded commands.

  2. RDB (Redis Database)

    Instead of writing every command, Redis can periodically take snapshots of the entire dataset.

    The main thread forks a subprocess (bgsave) that shares all the in-memory data of the main thread. The bgsave subprocess reads the data from the main thread and writes it to the RDB file.

    Redis uses copy-on-write. When the main thread modifies data, a copy of the data is created, and the process works on that so that writes don’t get blocked. The snapshot is then written as an RDB file on disk, allowing Redis to quickly reload the snapshot into memory when needed.

  3. Mixed Approach

    In production, Redis often uses both AOF and RDB. RDB provides fast reloads with compact snapshots. AOF guarantees durability by recording every operation since the last snapshot.

Over to you: Have you used Redis in your project?

Discussion about this post

Ready for more?