惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Troy Hunt's Blog
Blog — PlanetScale
Blog — PlanetScale
Engineering at Meta
Engineering at Meta
F
Full Disclosure
Recorded Future
Recorded Future
The GitHub Blog
The GitHub Blog
Microsoft Security Blog
Microsoft Security Blog
GbyAI
GbyAI
博客园_首页
博客园 - 叶小钗
MongoDB | Blog
MongoDB | Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Recent Commits to openclaw:main
Recent Commits to openclaw:main
H
Hacker News: Front Page
人人都是产品经理
人人都是产品经理
The Cloudflare Blog
博客园 - 司徒正美
Webroot Blog
Webroot Blog
Google DeepMind News
Google DeepMind News
Help Net Security
Help Net Security
Cloudbric
Cloudbric
PCI Perspectives
PCI Perspectives
有赞技术团队
有赞技术团队
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
TaoSecurity Blog
TaoSecurity Blog
L
Lohrmann on Cybersecurity
量子位
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Tailwind CSS Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
B
Blog RSS Feed
Apple Machine Learning Research
Apple Machine Learning Research
大猫的无限游戏
大猫的无限游戏
P
Proofpoint News Feed
N
News and Events Feed by Topic
罗磊的独立博客
T
Threat Research - Cisco Blogs
Schneier on Security
Schneier on Security
T
Tor Project blog
IT之家
IT之家
M
MIT News - Artificial intelligence
S
Security @ Cisco Blogs
O
OpenAI News
AI
AI
S
Securelist
Simon Willison's Weblog
Simon Willison's Weblog
The Last Watchdog
The Last Watchdog
月光博客
月光博客
Security Archives - TechRepublic
Security Archives - TechRepublic
L
LINUX DO - 热门话题

Sanity.io

A Board Game agent built using Sanity Context and Vercel's AI SDK | Sanity Build a prototype with Claude Code that your whole team can edit | Sanity What’s New - May 2026 | Sanity I built a London pub guide with v0 and the Sanity MCP in six hours. Here's what I learned. | Sanity Build a conference concierge with Agent Context and Anthropic | Sanity Build a content-aware Telegram agent with Vercel AI SDK and Chat SDK | Sanity How I used Agent API to generate photos for my family’s recipes | Sanity What’s New April - 2026 | Sanity Better context, better matches: An AI love story (for dogs) | Sanity How to write for an agent | Sanity Content Agent, meet Slack: AI content operations in your workflow | Sanity Structure powers intelligence | Sanity Your agent needs better content. Here's how to give it. | Sanity Sanity TypeGen GA: Automatic TypeScript types for content and GROQ | Sanity Sanity is now available on the Vercel Marketplace | Sanity The logo soup problem (and how to solve it) | Sanity Content Releases: From scattered updates to coordinated publishing | Sanity What's New - February 2026 | Sanity How we solved the agent memory problem | Sanity v0 Builder Challenge: The winners | Sanity Introducing: Sanity Agent Skills | Sanity Content Agent: Days of work in one conversation | Sanity Our Sanity Values | Sanity Open Source Pledge 2025: Stepping up when it matters | Sanity v0 builder challenge: $3000 in prizes | Sanity Why AI Breaks Without Structured Content Operations | Sanity What’s New January - 2026 | Sanity BFCM 2025: What teams built when infrastructure stopped being the problem | Sanity How AI shaped holiday shopping and what it means for content in 2026 | Sanity Sanity Studio v5: Embracing React 19 | Sanity You’ll need a CMS eventually. Let your agent set it up. | Sanity “You should never build a CMS” | Sanity AI Content Operations: A 30-Day Implementation Guide | Sanity What’s New December - 2025 | Sanity Scheduled Drafts: Stop manually publishing content at midnight | Sanity What’s New November - 2025 | Sanity Everything *[NYC] 2025 recap: A day of AI, Content Operations, and Culture | Sanity Clankers and content operations | Sanity Content Agent: AI that understands your structured content is here | Sanity Why design-driven content modeling creates technical debt, not velocity | Sanity What's New October - 2025 | Sanity From studio to inbox: How Kevin Green eliminated email campaign friction | Sanity The content editor's guide to content operations [E-commerce edition] | Sanity styled-components maintenance mode: A 40% faster fork | Sanity From zero code to a live website in 7 hours (thanks, Cursor!) | Sanity First attempt will be 95% garbage: A staff engineer's 6-week journey with Claude Code | Sanity Internationalization is more than translating words | Sanity What's New - September 2025 | Sanity We just deleted our 35k-member community Slack | Sanity What's New - August 2025 | Sanity The engineer's guide to content operations [E-commerce edition] | Sanity SEO for AI: Evolving from Web Pages to the Content Lake | Sanity What's New - July 2025 | Sanity Sanity Studio v4: A major version bump for a minor reason | Sanity What's New - June 2025 | Sanity Dashboard and Insights: Your New Content HQ | Sanity Canvas: AI-accelerated, context-aware, freeform authoring | Sanity Agent Actions: AI building blocks for structured content | Sanity Functions: Life beyond pressing publish | Sanity A new era for content applications with Sanity App SDK | Sanity The end of CMS era and our $85M Series C. | Sanity What's New – May 2025 | Sanity Introducing the Sanity Model Context Protocol (MCP) server | Sanity What's New – April 2025 | Sanity Pushing all the envelopes with ambitious content | Sanity Self-hosting is only free if your time is worth nothing | Sanity Content that lasts: Scaling beyond your frontend | Sanity The Live Content API is now Generally Available | Sanity The future beyond AI chat bots | Sanity Learning the new skill of working with AI | Sanity What's New - March 2025 | Sanity Give it in plain text: Making your content AI-Ready | Sanity No More 'DO NOT PUBLISH': Introducing Content Releases | Sanity React in 2025, what's next? | Sanity The final boss of front-end: block editors | Sanity Introducing Sanity for Startups | Sanity A block content editor that loves you back | Sanity A Black Friday Snooze Fest: Massive Traffic, No Drama | Sanity How to make a recipe site that scales well | Sanity The Sanity Winter Release 2024 | Sanity AVIF Arrives, Sanity’s Promise Fulfilled | Sanity Sanity joins the Open Source Pledge | Sanity Your content is now Live by default | Sanity Begin Team to Join Sanity | Sanity Sanity Digest - September '24 Edition | Sanity Sanity partners with Google. Now live on the Google Cloud Marketplace. | Sanity Sanity Digest - August ‘24 Edition | Sanity Now playing: the latest Mux Video Input plugin for Sanity | Sanity Community Digest - June ‘24 Edition | Sanity Community Digest - May ‘24 Edition | Sanity Guide to Sanity's newest product announcements | Sanity AI and Content Creation: A Leader's Guide | Sanity Of course, you should be able to type your content quickly! | Sanity New to AI Assist: translation, reference suggestions, image generation | Sanity Speak the language of your editors: Sanity Studio UI localization | Sanity Introducing the new Sanity Growth plan to serve collaborative teams | Sanity Presentation: Work faster than ever with structured content | Sanity Goodbye Feedback Frenzy, Hello Sanity Studio Comments! | Sanity Easing into the App Router with the Sanity Toolkit for Next.js | Sanity Making website updates easier with structured content | Sanity
How to serve content to agents (a field guide) | Sanity
Knut Melvær · 2026-02-19 · via Sanity.io

"AI-ready content." Everyone agrees you need it. Nobody agrees on what it means. AEO strategies (or GEO, or 'SEO for AI'), llms.txt debates, Cloudflare shipping markdown at the edge, agents that negotiate content types. The conversation is getting louder, and most of it conflates at least three different questions.

How do you get AI to cite and recommend your content? That's the positioning question, and it has two parts. Namely, what your content looks like when models are trained on it, and when agents retrieve it to answer a prompt.

The other question is about when an agent does show up, how do you serve your content without bloating its context window or losing meaning in translation? That's the content consumption question. Your content now has consumers you didn't design for: agents requesting your pages, humans copying your docs into Claude, RAG pipelines pulling your content into retrieval systems. None of them want your cookie banners, navigation chrome, or ad scripts. They want the content. And probably only the content that actually provides the answers they’re looking for.

And then: does agentic content consumption affect positioning? Does serving cleaner content to agents also improve how they represent you?

If you're a developer, the content consumption section is where the actionable stuff lives. If you're a content strategist, the positioning question is more your territory. Read both. You'll need to explain the other one to your team.

Can you optimize for AI citations? (here's what we know)

Let's start with Agent Engine Optimization (AEO, sometimes called GEO or "SEO for AI"), since that's what got everyone's attention: can you optimize your content so AI models cite you more?

Maybe. But the honest answer is: we don't really know yet.

One group that's been looking closely at this is Profound, a company that tracks how AI platforms cite and recommend content across ChatGPT, Perplexity, and Gemini. They've been publishing primary research on the topic.

In their latest study, they took 381 pages across 6 websites, randomly assigned half to serve markdown and half to serve HTML, and watched for three weeks. The result? They found no statistically significant increase in bot traffic from serving markdown. Their recommendation? Focus on fundamentals: quality content, clear structure, fast load times. "The format you serve them in? Probably not the leverage point you're looking for."

Their most important finding for anyone building an AEO strategy is that AI citations can shift by up to 60% in a single month. The page ChatGPT recommends today might not be the one it recommends next month.

If you're trying to "optimize" for a system that volatile, you're chasing a moving target.

There's academic work too. A paper from Princeton (published at KDD 2024, which feels like three decades ago in the AI-timeline) found that certain strategies (adding statistics, using authoritative language, citing sources) could boost visibility by up to 40% in their benchmark. Worth noting: those numbers come from lab conditions, not from testing against ChatGPT or Perplexity in the wild. The strategies themselves are basically good writing advice, which is worth doing regardless of AI.

Why web pages are the wrong mental model for AI-ready content

We've also published AEO/GEO: Evolving from Web Pages to the Content Lake, which covers the strategic evolution of search and what it means for organizing content systems. This guide focuses on the tactical implementation.

Meanwhile, the models themselves are getting better at handling whatever you throw at them. Anthropic just released dynamic filtering for Claude's web search. The model now writes Python code to parse and filter HTML results before they hit the context window. The result: 11% better accuracy, 24% fewer input tokens. The models are investing heavily in solving the "finding you" problem on their end.

So while the models will get better at finding you, they won't necessarily get that much better at your content being clean. Which brings us to the part where you have agency.

What to serve agents when they show up

Profound (the AI citation tracking company) measured bot visits, how often agents show up. They found format doesn't change that. Agents show up regardless. Cool.

If agents are already showing up (and they are, check your server logs), then the question isn't how to attract them. It's what you serve them when they arrive. And not just agents: humans are copying your docs into AI tools, RAG pipelines are pulling your content into retrieval systems, edge services are converting your pages on the fly.

The strategies here are practical, the evidence that this might matter is strong, and you control the outcome.

Here's what the options look like, from zero effort to full infrastructure investment.

Do nothing

Agents will convert your HTML to markdown themselves. Every major AI tool (Claude, ChatGPT, Gemini, Perplexity) does this internally. Most AI crawlers don't even execute JavaScript: they see your raw HTML, not your rendered page. Claude Code uses a library called Turndown. It works. It's also lossy and token-expensive. Your 100K-token HTML page becomes maybe 3K tokens of useful content after the agent strips out navigation, footers, scripts, and cookie banners. That's a 97% waste of context window. Even the agents that do render JavaScript (like Google's crawler or ChatGPT's Operator) still get the full DOM with all the navigation chrome. The token waste problem doesn't go away just because JS executes.

It gets worse if your site relies on client-side rendering. Vercel's research found that ChatGPT and Claude crawlers fetch JavaScript files but don't execute them, the aforementioned Google's Gemini (via Googlebot) and AppleBot being the exceptions.

Add an llms.txt file

llms.txt is a markdown file at a known URL (/llms.txt) that gives agents an overview of your site with links to detailed content. Over 2,000 sites have adopted it, including Next.js, shadcn/ui, TanStack, Cloudflare, and Hugging Face. It's simple to implement and useful as a discovery layer. Anthropic uses theirs as a lightweight sitemap: brief descriptions and links organized by section, 892 tokens total. Even the company building the agents treats llms.txt as an index, not a content dump.

But isn't llms.txt becoming a standard? It's becoming adopted, which isn't the same thing. llms.txt is a proposal from Jeremy Howard's FastHTML project, not a ratified standard. The GitHub issues show debates about merging it with other proposals (AGENTS.md), and scope creep into things like crypto wallet addresses and "emotional brand positioning extensions." More practically: it's all-or-nothing. An agent gets your entire corpus or nothing. No per-page granularity, no per-agent control, no governance over what gets consumed. For developer docs, that's probably fine. For anything you'd rather serve selectively, you've just made it trivially easy to copy-paste everything.

The spec also proposes llms-full.txt, a companion file containing your entire corpus as markdown. In theory, agents can grab everything at once. In practice, the numbers work against you. Cloudflare's developer docs produce an llms-full.txt of 46.6MB, roughly 12 million tokens, about 60x Claude's context window. Even when the file fits, longer context degrades model performance regardless of content quality. An agent that needs one answer doesn't benefit from receiving your entire library.

If you want a quick maybe win, add an llms.txt. If you want control over what agents get and when, content negotiation gives you more options (more on that below).

Turn on Cloudflare's edge conversion (if you host on Cloudflare)

Cloudflare launched Markdown for Agents in February 2026: a dashboard toggle that converts your HTML to markdown at the edge when agents request it. No code changes. 80% token reduction on their own blog. It's a good default if you can't touch your content layer.

The tradeoff: reverse-engineering HTML back to markdown is inherently lossy. A generic parser doesn't know which parts of your page are content and which are chrome. Custom components, structured relationships, section-level meaning: none of it survives the round trip.

Serve markdown routes with content negotiation

Agents have started to request markdown from you using the Accept header. So when an agent sends Accept: text/markdown, your server could respond with markdown. Same URL, different representation. The same content negotiation pattern HTTP has supported for decades. I built a Sanity course around this, and we use it for Sanity Learn itself.

It doesn’t take a lot of code if you use Sanity already. With the @portabletext/markdown library you can take the same content that renders to HTML, and render it to markdown as well:

Most frameworks and hosting platforms lets you define URL rewrites too, here is how content negotiation looks like in Next.js:

On our learning platform, the same lesson page goes from 392KB of HTML (~100K tokens) to 13KB of markdown (~3,300 tokens). That's a 97% reduction, and it's not lossy, because the markdown is generated from the structured source, not reverse-engineered from HTML.

And because the content is structured, you can serve it at multiple levels of granularity. An agent landing on any lesson finds links to the full course, a sitemap of all content, and the complete corpus. It picks the level that matches its task: one lesson for a specific question, a full course for context, the sitemap for discovery. Same content, four levels of access, all from the same structured source.

The granularity matters because of how agents actually work. AI coding agents are already requesting markdown. Remotion, Bun, and nuqs serve it. (For the curious: Claude Code currently sends Accept headers that omit text/html. The markdown preference is already baked in.) When a server responds with Content-Type: text/markdown, Claude Code skips its summarization step entirely. Your content goes straight to the model, verbatim. That's a fast path you earn by serving clean content.

Expose content via APIs or MCP

The furthest end of the spectrum: skip the web entirely. Rely on agents query your content through tools: GROQ queries, GraphQL, the MCP server. No HTML-to-markdown conversion needed because there's no HTML involved. This is the most powerful approach but requires the most integration work. It's real and growing, but early for most sites.

And it also requires your users to install or instruct agents to do it this way, which might be fine if you are a SaaS company with users who are proficient in AI tools, but not realistic for most of the rest of us.

Copy-as-markdown buttons

Worth mentioning because it shows the demand from the human side too. Cursor, Vercel, and Remotion docs all have "copy page as markdown" buttons. The popular documentation platform Mintlify comes with it out of the box too. Cursor has "Open in ChatGPT" and "Open in Claude" buttons. We also have them on our docs and learning platform. Developers are manually doing what agents should get automatically. The frugality argument isn't just about bots.

What all of this requires from your content system

Each strategy in that spectrum asks something different from your content system.

StrategyWhat your system needs
Do nothingNothing. Agents handle conversion.
llms.txtA way to export all your content as markdown
Cloudflare edgeCloudflare hosting (conversion is lossy)
Markdown routesStructured content that serializes to markdown
MCP / APIStructured content + query language + governance

Notice the pattern. The further right you go, the more your system needs to treat content as queryable data rather than documents to be parsed. If your content lives as markdown files in a repo, you can serve markdown, but that's all you can serve. You can't query it ("give me all articles tagged X published this quarter"). You can't serve HTML to browsers and markdown to agents and JSON to your mobile app from the same source without building a (gnarly) pipeline for each. We covered this deeper in “You should never build a CMS.”

Markdown was designed as a “writing format” back in the days when HTML was the only destination. (Gruber even offered a .text suffix to view the markdown source of each page. Content negotiation before we called it that.) We see it as a serialization format, not a content storage format that scales. (LLMs are great at SQL too, that doesn’t mean you store your database as .sql files in git.)

The systems that serve agents best are the ones where content is structured, governed, and queryable. That's true for markdown routes, and it's even more true as agents move from fetching web pages to accessing content directly through tools and APIs. That's a bigger conversation, one that touches how organizations manage knowledge for both humans and machines. But it starts with the same infrastructure question.

What's changing fast vs. what's durable

Half of the specific tools I just mentioned will look different in six months. New standards will emerge. Agents will get smarter. Claude's dynamic filtering shipped today. Chrome is previewing WebMCP, a proposal that lets websites expose structured tools directly to in-browser AI agents, so they can call functions instead of scraping pages. By the time you read this, there's probably something newer. The landscape is genuinely moving fast.

But the mental model holds. The difference between "positioning yourself for AI discovery" and "being intentional about how your content gets consumed" isn't going away. They're different questions with different evidence bases and different levels of maturity. Knowing which one you're solving for changes what you should invest in.

If this sounds familiar, it should. Serving the right format to the right consumer is multichannel distribution. Developers have been doing this for years: HTML for browsers, JSON for mobile apps, RSS for feed readers. AI agents are just the newest channel. The durable investments are the same ones that have always mattered: clean content structure, standards-based content negotiation, governance over what gets published where. These pay off regardless of which specific tools win.

What I'd actually do

Here's what I'd tell a team asking me where to start.

Start with your server logs. Are agents already visiting? What are they requesting? You might be surprised. Then serve markdown to agents that ask for it. If an agent sends Accept: text/markdown, respond with markdown. It's a decades-old HTTP pattern, and it's the highest-ROI move for most sites right now.

Beyond that:

Probably don’t do llms.txt. Dumping your whole corpus into a markdown file is probably just going to lead to context bloat and “semantic collapse.” It seems like agents prefer more curated and specific content depending on the intent they’re resolving for.

Watch the positioning space, and start to bet on it. Follow Profound's research. They're doing the most focused work in this space. When there's proven methodology for getting cited by AI search, invest in it. Until then, assume that the answer is going to be “serve quality content that fulfills what your audience might be looking for.”

Think about governance. Know which bots are visiting and what they're doing. Dark Visitors tracks 80+ AI-specific bots across 15 categories. Your robots.txt probably needs updating.

If you're building with Sanity, we've published SEO and AEO best practices as an agent skill that agents can use directly when implementing these patterns. It includes EEAT principles, structured data guidance, and technical SEO checklists, as well as the emerging best practices for AEO.

And invest in your content structure. It's the unsexy one, but it's the one that compounds. If your content is structured, markdown is just another output. So is HTML. So is JSON. So is whatever format the next generation of agents will want. You're not optimizing for AI. You're doing what good content infrastructure has always done: serving the right thing to whatever shows up asking for it.

Which brings us back to the third question: does serving cleaner content to agents also help you get cited? We don't know yet. Profound's data suggests format alone isn't the lever, but the research is young and the landscape is moving fast. What we do know is that the content consumption work pays off on its own. And if the connection to positioning turns out to be real, you're already there.

The next time someone tells you to make your content "AI-ready," ask them which kind. The answer changes the conversation.