惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
Engineering at Meta
Engineering at Meta
GbyAI
GbyAI
MongoDB | Blog
MongoDB | Blog
Blog — PlanetScale
Blog — PlanetScale
腾讯CDC
U
Unit 42
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Apple Machine Learning Research
Apple Machine Learning Research
M
MIT News - Artificial intelligence
人人都是产品经理
人人都是产品经理
Hugging Face - Blog
Hugging Face - Blog
MyScale Blog
MyScale Blog
小众软件
小众软件
博客园 - 三生石上(FineUI控件)
N
Netflix TechBlog - Medium
阮一峰的网络日志
阮一峰的网络日志
博客园 - Franky
Recent Announcements
Recent Announcements
A
About on SuperTechFans
Stack Overflow Blog
Stack Overflow Blog
The GitHub Blog
The GitHub Blog
D
Docker
H
Hackread – Cybersecurity News, Data Breaches, AI and More

Google Developers Blog

Why client SDK generation belongs in the open- Google Developers Blog Agent Anomaly Detection, now in Private Preview on the Gemini Enterprise Agent Platform- Google Developers Blog Build zero-trust AI agents that judge intent, not just syntax- Google Developers Blog Autonomous LLM post-training with Tunix on TPUs- Google Developers Blog The Anatomy of Harness Engineering: How to Evaluate, Iterate, and Guard AI Coding Agents- Google Developers Blog Announcing ADK for Kotlin 1.0: Building Production-Ready AI Agents in Kotlin, Android, and Beyond- Google Developers Blog Driving Developer Excellence: Inside the Program Sprints- Google Developers Blog 4 engineering patterns behind the strongest AI Agents Challenge submissions- Google Developers Blog Decoding cosmic signals with deep learning and Keras- Google Developers Blog Enterprise-Grade Precision for Long-Context Multimodal Embedding Inference on Cloud TPU- Google Developers Blog How to Evaluate Live & Voice Agents in ADK- Google Developers Blog Build zero-trust AI agents with Google's Agent Development Kit- Google Developers Blog Introducing Credentio: Open Source C++ Library for C2PA Content Credentials from Google- Google Developers Blog HeyGen x Google Cloud: Bringing Avatar IV to TPUs- Google Developers Blog Why Go is an Ideal Language for AI-Assisted Software Engineering- Google Developers Blog Mastering Edge AI on Raspberry Pi with LiteRT and Gemma- Google Developers Blog Agent Plugins package your skills, tools, and more- Google Developers Blog Scaling AI Agent Infrastructure with the MCP Stateless updates- Google Developers Blog A unified API for AI model routing- Google Developers Blog Scaling real-time AI agents with session-aware load balancing- Google Developers Blog Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA- Google Developers Blog Enable on-demand expertise with Agent Skills in Genkit Go- Google Developers Blog How to use Google microbenchmarks for evaluating TPU performance- Google Developers Blog Run Ray on TPU, Part 2: Ray AI libraries- Google Developers Blog Scaling Agentic RL: High-Throughput Agentic Training with Tunix- Google Developers Blog Run Ray on TPU, Part 1: The foundations- Google Developers Blog Expanding Choice in Gemini Enterprise Agent Platform: Introducing Grounding with Parallel Web Search- Google Developers Blog Evolving Spec-Driven Development: Conductor Now Supports Antigravity- Google Developers Blog Systems Engineering Playbook: Optimizing Qwen 3.5-397B MoE on Ironwood (TPU7x)- Google Developers Blog Unlocking the Next Era of On-Device AI with Google Tensor and Pixel- Google Developers Blog
Building scalable AI agents with modular prompt transpila...
Simerus Mahesh · 2026-07-16 · via Google Developers Blog

When you’re first building an AI agent, a single, monolithic system prompt is usually fine. You have a few instructions, maybe a tool definition or two, and everything lives in one readable file.

But as you start using them for production purposes, that format simply just breaks down. Teams start layering on safety policies, domain-specific rules, formatting requirements, and escalation behaviors. Suddenly, you have your entire agent’s control plane within a single instruction file which is exactly where the trouble starts.

This is a classic software engineering scaling problem. When you push every concern into a single file, you lose the ability to reason about the system. Collaboration becomes a nightmare, testing gets finicky, and a small change meant to improve one workflow can quietly break another.

At production scale, prompt maintainability becomes agent reliability.

Why monolithic prompts break down

We typically see three main failure modes when prompts grow beyond a certain size:

  1. Obscured blast radius: In standard software engineering, it’s easy for reviewers to reason about the scope of a change through module boundaries, call sites, and tests. System prompt diffs are harder though. Adding a sentence could have unintended side effects across the entire agent, which is often hard to predict or test.
  2. Copy-paste drift: As organizations scale, many teams end up duplicating shared logic for various applications such as internal service usage instructions, PII handling, safety policies, or escalation protocols. This leads to copy-pasting or multiple versions of the same functionality leading to inconsistencies.
  3. Deferred runtime errors: To manage the sprawl, teams often resort to ad-hoc string formatting or simple templates. While this helps with authoring, it pushes error detection to runtime. You might deploy a prompt that only fails when a specific, rarely-used workflow is triggered because of a missing variable or an invalid import path.

Templates are a good start, but they aren't enough. Production systems require deterministic builds, static validation, and CI/CD integration.

Treat prompts like software artifacts

The solution here is to treat prompts like build artifacts as opposed to just static text.

Instead of maintaining one monolithic prompt file, you can author modular skill files. This allows you to reduce the scope of each file and encapsulate a specific behavior, which allows teams to separate concerns and iterate on components individually.

A top-level agent prompt template might look something like this:

# agents/sre_agent.prompt.md (prompt template file)

{% include "shared/safety.prompt.md" %}
{% include "shared/tool_usage.prompt.md" %}

You are an SRE triage agent operating in the {{ environment }} environment.

{% if allow_remediation %}
You may recommend remediation steps, but destructive actions require human approval.
{% else %}
You may inspect, summarize, and explain the issue, but do not recommend remediation actions.
{% endif %}

{% macro bullet_section(title, items) %}
## {{ title.rstrip() }}
{% for item in items %}
- {{ item.rstrip() }}
{% endfor %}
{% endmacro %}

{{ bullet_section("Required investigation steps", [
  "Inspect recent deployment events",
  "Check service metrics for latency or error-rate changes",
  "Review logs for repeated failure patterns"
]) }}

Plain text

Copied

This gives you the best of both worlds. The templating layer lets you compose shared instructions, inject environment-specific values, and make use of macros. But for the build system, every include is a dependency, and every variable is a requirement. The result is a deterministic, fully rendered artifact that you can test, audit, and diff before it ever reaches the model. We can then use a transpiler to resolve the template imports to generate a file that is ready to be ingested by an agent.

For example, if environment = production and allow_remediation = true, the transpiled artifact would look like this:

You are an SRE triage agent operating in the production environment.

You may recommend remediation steps, but destructive actions require human approval.

## Required investigation steps

- Inspect recent deployment events
- Check service metrics for latency or error-rate changes
- Review logs for repeated failure patterns

Plain text

Copied

A high-level transpilation pipeline would look something like this:

Figure1

Build-time validation is mandatory

A production-grade transpiler should catch errors before runtime.

We should be running validation checks for missing imports, undefined variables, and circular dependencies during the build process. Dependency graphs are invaluable here, reinforcing the need for a solid template engine. If you treat each prompt fragment as a node in a directed graph, you can easily catch recursive imports that would otherwise cause a silent failure in production.

Figure2

This also enables drift checking. You can set your CI pipelines to be able to regenerate the transpiled prompt from source (referred to as the golden file) and compare it against the currently committed artifact. If the outputs differ, the build fails. This ensures that the code in your repo is exactly what’s running in production, eliminating the gap between source files and deployed artifacts.

Figure3

Dynamic skills and agent-authored updates

As your skill library of modular prompt fragments grows, you don't necessarily want every agent to load every skill every time. Doing so consumes tokens and introduces noise that can interfere with the agent's task-specific performance.

A better architectural pattern would be to leverage progressive disclosure. This is where we separate the stable control plane from task-specific context. The compiled base prompt should enforce non-negotiable behaviors like identity and safety boundaries. Then, at runtime, the agent can use a tool to dynamically retrieve only the specific skill modules required for the task at hand; this reduces context exhaustion and helps to keep the agent focused on its task.

Figure4

Once you have this modular system, you unlock a powerful workflow: agents can help maintain their own instruction layer, helping to create a self-sustaining agentic system. When an agent resolves a new type of incident, it could theoretically draft a new skill module, update the relevant imports, and open a pull request.

The agent isn't mutating its own instructions in real-time; it's proposing a code change. The transpiler then subjects that proposal to the same validation and review rigors as any other code change. A human reviewer can inspect the PR, run the evals, and merge the change.

Figure5

Conclusion

A production prompt transpiler reframes prompt engineering as a build-system problem.

When we build modular skill files, we can resolve dependencies, validate imports, and enforce drift checks just as we do with our standard software infrastructure. Agents become capable of suggesting improvements to their own logic, provided those changes pass through our existing validation and review processes.

As AI agents become deeply integrated into critical workflows, their instruction layers need the same reliability standards we demand of our software. Prompts shouldn't just be edited, they should be built, validated, versioned, and deployed.