惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Know Your Adversary
Know Your Adversary
阮一峰的网络日志
阮一峰的网络日志
V
Visual Studio Blog
H
Help Net Security
博客园 - Franky
博客园_首页
博客园 - 【当耐特】
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
腾讯CDC
人人都是产品经理
人人都是产品经理
T
Tailwind CSS Blog
博客园 - 三生石上(FineUI控件)
爱范儿
爱范儿
博客园 - 聂微东
小众软件
小众软件
宝玉的分享
宝玉的分享
美团技术团队
WordPress大学
WordPress大学
L
LINUX DO - 热门话题
S
Secure Thoughts
IT之家
IT之家
博客园 - 叶小钗
Apple Machine Learning Research
Apple Machine Learning Research
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
有赞技术团队
有赞技术团队
V
V2EX
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
The Cloudflare Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
量子位
G
GRAHAM CLULEY
Attack and Defense Labs
Attack and Defense Labs
Jina AI
Jina AI
罗磊的独立博客
Security Archives - TechRepublic
Security Archives - TechRepublic
W
WeLiveSecurity
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
T
Tenable Blog
L
Lohrmann on Cybersecurity
S
SegmentFault 最新的问题
J
Java Code Geeks
Last Week in AI
Last Week in AI
Cyberwarzone
Cyberwarzone
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
大猫的无限游戏
大猫的无限游戏
T
Tor Project blog
酷 壳 – CoolShell
酷 壳 – CoolShell
月光博客
月光博客
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
H
Heimdal Security Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Claude Code Costs, Act II — Where the big hidden costs are
Sumedh Bala · 2026-06-26 · via DEV Community

A single-model session that stays well-cached is cheap. The biggest swing in a multi-model bill comes from one move — switching models — because the prompt cache belongs to a single model. The instant you switch, the cache you already paid for is thrown away.

Whether that helps or hurts comes down to how you switch:

  • Sticky — keep a whole conversation (or a sub-agent) on one model. The cheaper model builds and reuses its own warm cache, so the bill goes down.
  • Per-turn bouncing — flip between models mid-conversation. It's not a full cold rebuild each time: only the first call to each model is fully cold, and returning to one you've already used re-reads that model's own warm prefix and writes only the catch-up since you left (so long as you're back within the ~20-block / 1-hour reuse window — see first call cold, then warm catch-up below). But you're now keeping two caches warm and paying a catch-up write on every flip, so at the 2× write rate it can still cost more than never switching.

Our 25-turn run shows both ends: bouncing 20% of turns to Sonnet lost ~2%, while keeping the whole run on Sonnet saved ~53% (on Haiku, ~85%) — see Act III. So the goal isn't to avoid switching; it's to switch the right way. The costly version is the one that sneaks in by accident: a router that picks a model per request, or a manual swap partway through a conversation.

What about the model's prior reasoning — its thinking blocks? It's tempting to count that as a second switching cost, but mechanically it's just context: the client re-sends it every turn, switch or not (Mental model 1). A switch only changes what the new model does with it — it either re-bills the reasoning as input (a cost), or strips it before the model sees it (a behavioral change).

The strip case is the one to watch, and it's about behavior, not money. Reasoning travels as an opaque, encrypted signature: you can't read it or edit it — you can only carry it whole or lose it. If a switch drops it, the model continues without its earlier chain of thought, and may no longer behave the way the previous turns set it up to.

So this part covers the cache cost of a switch first, then what happens to reasoning across one.

Switching models: the model-scoped cache

Mental model 3: the model-scoped key

A cache entry belongs to exactly one model. Another model cannot read it.

This is the rule that makes "just route the easy turns to a cheaper model" so often backfire. There's really one fundamental reason a switch can't reuse the cache you already paid for — proven below: the cache key is model-scoped. And even once you accept that, a switch costs more than a clean cold start would, because the token counts shift between models — a separate effect that compounds the bill, covered after.

Why the cache can't carry: caches are model-scoped [measured]

It falls straight out of What the cache stores (Mental model 2). A cache hit reuses the key/value vectors a model computed for the prefix — and those vectors are produced by that model's own weights. Run the identical tokens through a different model and you get different queries, keys, and values, a different attention computation, and therefore a different KV state. So a cached entry is meaningful only to the exact model that produced it: hand Sonnet's KV cache to Haiku and it's noise. That's why the cache can't carry across a switch — not a policy choice but a consequence of attention itself; the saved state simply isn't the state the new model would have computed.

The clean proof — a byte-identical 63K-token prompt, varying only the model:

call model cache_read cache_creation
sonnet #1 sonnet-4-6 0 63,422
sonnet #2 sonnet-4-6 63,422 0
haiku #1 haiku-4-5 0 64,031
haiku #2 haiku-4-5 64,031 0

Sonnet's second call reads its own warm entry. Haiku's first call — same bytes — reads 0 and cold-writes its own copy. The two models cannot share an entry.

The live consequence — duplicated cache writes in a mixed session (a real 50/50 Sonnet/Haiku session, per-model totals) [measured]:

model requests cache_creation cache_read
sonnet 9 27,374 536,662
haiku 6 57,739 321,744

Haiku cold-wrote a 57,739-token duplicate of the shared prefix it could never read from Sonnet — about 85,113 total cache_creation tokens of pure duplication a single-model session would never pay. At the 2× write rate you pay the write premium twice and lose the 0.1× read discount on the rerouted slice. The savings from a cheaper model are eaten by cache duplication. This is the central caveat of the whole guide.

The bytes happen to differ too — diffing the system prompt across models, two lines change (the model-name line and the knowledge-cutoff line, e.g. Opus 4.7 / January 2026Haiku 4.5 / February 2025). But that's moot for caching: the model-scoped key already settled it. No amount of byte-matching would let Haiku read Sonnet's KV state — so don't think of the differing text as a second cause; it's the same wall.

And the cost shifts too: the tokenizers differ [measured]

This one isn't about the cache at all — it's a separate cost effect that rides along with every switch. Token counts shift between models, so even "the same text" bills as a different number of tokens:

  • Opus 4.8 = 1.0 (reference)
  • Haiku 4.5 = 0.775 (measured: 64,316 vs 83,004 tokens for the same content)
  • Sonnet 4.6 = 0.775 (assumed — same pre-4.7 tokenizer family, unmeasured)

Anthropic publishes no official ratio (use count_tokens per model); it documents the 4.7/4.8/Fable tokenizer as ~1×–1.35× an older one, putting older models at roughly 0.74–1.0× of Opus.

Mental model refinement: first call cold, then warm catch-up

The good news is that a switch is not cold forever:

Only the first call to a given model is fully cold. Each subsequent call to that model is partially warm: it reads that model's own cache and cold-writes only the catch-up diff — the content added by intervening turns on the other model. Cost = one cold start + a recurring catch-up write per re-entry, not a cold start every call.

Two bounds apply: the TTL (1 hour, refreshed on each read of that model's entry) and the 20-block lookback (~7–10 turns). Beyond the lookback the tail can't re-link, but the front breakpoints (tools+system) still hit — so you re-read the system prefix warm and only cold-write the message history. At the 2× write rate, both the cold start and the catch-up writes hurt twice as much, which is exactly what flips routing economics in Act III.

⚠️ Mistake — routing per turn to save money. Bouncing between models mid-conversation pays a cold prefix write on the first switch plus a catch-up write on every re-entry. For a higher-write-rate model (Sonnet), this can cost more than just staying on Opus.

✅ Fix — Route sticky: pick a model per conversation or per sub-agent and stay there. (Act III quantifies it: 20%-Sonnet per-turn bouncing loses 2.1%, while all-Sonnet sticky saves 53%.)

What to do: Default to treating the model as a per-conversation decision, not a per-turn one — and if you use multiple tiers, isolate them into separate sub-agents/conversations so each keeps its own warm cache. Plenty of commercial routers and gateways will route per request for you automatically, and they can save money — but the win is workload-dependent, and (as the numbers above show) per-request routing can quietly cost more, especially for a higher-write-rate model. So don't switch one on blindly. Adopt it only once you can see the evidence on your traffic: measure cache_read vs cache_creation and your actual billed cost, understand why the cache is model-scoped, and confirm a real net saving before you rely on it. And weigh more than cost: a cheaper tier can carry an older knowledge cutoff (e.g. Haiku 4.5's is ~11 months behind Opus's) — a behavioral difference that routing-for-price quietly inherits.

Thinking blocks across a model switch

Act I established that the client re-sends everything each turn, including the model's prior reasoning — the thinking blocks. Carrying them is ordinary context, but a switch forces a choice with two kinds of consequence: they're either re-rendered into the target model's prompt and billed as input (a cost), or stripped before they get there (a behavioral change — the new model loses the prior chain of thought), depending on the target model's class. To weigh either, you first need to know what a thinking block is.

First: what model "thinking" is

If you haven't worked with reasoning models, start here; if you have, skip to Mental model 4.

A reasoning model doesn't answer immediately. Given a hard prompt, it first generates a run of intermediate tokens — working the problem out step by step — and only then writes its reply. That working-out is the model's thinking (also called reasoning, extended thinking, or chain of thought): a scratch pad, the "let me work through this" pass a person makes before answering, except the model does it by emitting tokens.

Why it does this: spending tokens on reasoning before answering measurably improves accuracy on anything multi-step — math, code, planning — where one wrong early step dooms the result. It's test-time compute — trade tokens (and latency, and money) for a better answer. Modern models use adaptive thinking: the model decides per request whether a problem is worth thinking about and how hard, so a trivial lookup gets none and a hard puzzle gets thousands of tokens.

In the response, thinking isn't blended into the answer. The reply is a list of typed content blocks, and thinking is its own block type, emitted before the text answer — the wire keeps the model's private working-out separate from the words meant for the user. That separate block is what gets carried back each turn, and what a model switch has to make a decision about. The next question is what's inside it.

Mental model 4: the encrypted envelope

A thinking block is not readable text you're carrying around. It's a sealed envelope.

The model seals its reasoning into an encrypted signature and hands you an envelope you can't open. You carry it back each turn (stateless — you hold it, not the server). The server has the key: it decrypts the signature to reconstruct the reasoning for the model. The result is private (you can't read it), stateless (the content rides in your request), and continuous (the server reconstructs it each turn).

It wasn't always sealed. The first generation of extended thinking handed the chain of thought back as plain, readable text — you got the model's working-out verbatim and could log it, diff it, even hand-edit it before resending. That openness is gone. Current Claude 4.x returns reasoning only in protected form: a summary written by a separate model, or nothing but the encrypted signature. The motive is anti-distillation — a raw chain of thought is exactly the training signal a competitor needs to clone the reasoning into their own model, so the readable text was replaced by a signature you can carry but not inspect. (summarized is the protected form, not a peek behind it — see Why you can't just read the chain of thought below.)

What you can still control — and what you can't. A few knobs shape the envelope; none of them open it:

  • Whether the model thinks, and how hard. The thinking parameter (adaptive, or enabled with a budget_tokens ceiling) plus the effort setting decide whether a block is produced and how long the reasoning runs. More reasoning → a bigger signature (the depth table below shows ~45× across difficulty).
  • How much of it you see. display has exactly two values — "summarized" (a paraphrase) and "omitted" (empty text, signature only) — and the default flips by model (newer models default to omitted). Neither returns the raw reasoning, and display is visibility-only: it does not change billing.
  • What you can't touch. You cannot read the raw thinking, edit it, reorder it, or author your own block. The sequence is integrity-checked — any modification is a hard 400. You take the envelope whole, or not at all.

Carry-over: required in some places, impossible in others. Because the content is sealed and integrity-locked, several moves that were fair game when thinking was open text are now off the table:

  • You can't edit the reasoning across a switch — though the block itself still travels. When thinking was readable you could take one model's chain of thought, trim or translate it, and feed it to another. Now it's a sealed, integrity-locked envelope: on a switch the previous blocks are re-rendered into the new model's prompt and billed as-is (across the Opus/Sonnet/Haiku family) — carried whole, never translated, trimmed, or merged — or, for the Fable/Mythos family, dropped. Encryption seals the contents from you; it does not stop the block from crossing the boundary and billing. (The replay matrix below measures which, and what it costs.)
  • You can't curate it to save tokens. Tempting as it is to prune a long chain before resending, editing the block is rejected outright, and removing it — though accepted as a 200 — silently breaks continuity and can convert cheap cache reads into cold writes.
  • You sometimes must carry it. During tool use, the thinking blocks between tool calls are part of the integrity-checked sequence: drop them and the model loses the thread it built across the tool loop. Here carry-over isn't an optimization you opt into — it's required for the model to keep behaving the way the earlier turns set it up to.

The measured detail behind each of these — what an omitted block contains, how signature size tracks reasoning depth, how it's billed, and what survives a switch — follows.

What an omitted-display block actually contains [measured]

Opus 4.7 with display:"omitted" emits: { "type":"thinking", "thinking":"" (empty), "signature": "<360–732 chars>" }. Nothing else. The readable thinking text is empty; the signature is the payload.

The signature is the encrypted full thinking — not a tag [doc-confirmed]

Verbatim from Anthropic's extended-thinking documentation:

"The signature field still carries the encrypted full thinking for multi-turn continuity."
"The server decrypts the signature to reconstruct the original thinking for prompt construction."

It also enforces integrity — blocks may not be edited or reordered:

"the entire sequence of consecutive thinking blocks must match the outputs generated by the model… you can't rearrange or modify the sequence of these blocks."

Modifying a block returns 400 invalid_request_error ("thinking … blocks in the latest assistant message cannot be modified").

Signature length scales with reasoning depth [measured]

What this shows: a thinking block isn't a fixed-size tag — it grows with how hard the model actually thought. Same model and settings (Sonnet 4.6, forced thinking, display:"summarized"), five prompts from trivial to hard. The column that matters is the last one, signature size:

prompt output_tokens summary text (chars) signature (chars)
trivial 20 1 276
easy 35 45 332
medium 444 34 320
hard (12-coin puzzle) 5,595 3,352 12,524
very_hard 64 129 448

Read down the signature column: the hard puzzle's signature is ~45× the trivial one, and larger than the visible summary (12,524 vs 3,352 chars). So the signature carries the full thinking; the summary is just a condensation. (Adaptive thinking decides per prompt whether to think at all, which is why the jump tracks actual reasoning — note very_hard happened to reason little — not the difficulty label.)

How thinking is billed [measured + docs]

  • At generation: thinking is billed as output tokens — even when display:"omitted" and you never see the text.
  • On replay: resent thinking bills as input tokens. Sonnet visible-text blocks run ≈ 3,085 tokens each; Opus omitted blocks were small (~85–155 tokens) only because those turns reasoned little — a deep omitted block is large (a 12-coin puzzle ran +1,522 on replay with zero visible text). Replay cost tracks reasoning depth (signature length), not the empty display.
  • display does not change billing. Thinking bills the same under omitted, summarized, or full.

Why you can't just read the chain of thought [docs]

There are only two allowed display values, and neither exposes the raw reasoning:

  • "summarized" — a separate model's paraphrase. "Summarization is processed by a different model than the one you target… The thinking model does not see the summarized output."
  • "omitted" — empty text; the signature carries the encrypted full thinking.

There is no value that returns verbatim chain of thought ("In rare cases where you need access to full thinking output for Claude 4 models, contact Anthropic sales."). So switching to summarized does not bypass anti-distillation — summarized is the protected form.

Defaults by model [measured]

Which models hide the thinking text out of the box: newer models default to omitted (signature only — Opus 4.7, Opus 4.8, Fable 5, Mythos 5, Mythos Preview), while older ones default to summarized (you also get the paraphrase — Opus 4.6, Sonnet 4.6, earlier Claude 4).

The cross-model replay matrix [measured]

The key cost question for switching: when the next turn goes to a different model, does it re-read the previous model's thinking blocks (and bill them as input), or silently drop them? The API never tells you which — so the test is to send the same request twice, once with the blocks kept and once with them removed, and diff the prompt-token count. Costs more with them kept → rendered and billed. Identical → dropped.

One concrete run, 3 Sonnet blocks replayed to Haiku: kept = 73,252 tokens vs removed = 63,997 — a 9,255-token gap (~3,085 per block), so they were rendered into Haiku's prompt and billed. (A naive model-only swap 400s first — Haiku rejects Sonnet's adaptive thinking param — so the params have to be fixed before the keep-vs-removed comparison is even valid.)

Run that same keep-minus-removed diff across every Opus/Sonnet/Haiku pairing and the verdict is uniform — always rendered, never dropped. Each row is one source model's blocks replayed to one target; the number is the extra tokens billed when you keep them (positive = billed):

source blocks → target keep − removed verdict
Sonnet (visible text) Opus 4.7 +10,950 rendered & billed
Sonnet Haiku 4.5 +9,253 rendered & billed
Sonnet Sonnet (control) +9,077 rendered & billed
Opus (omitted/empty text) Sonnet 4.6 +125 rendered & billed
Opus Haiku 4.5 +308 rendered & billed
Opus Opus (control) +170 rendered & billed

The Sonnet rows are large (~3,085 per block of real visible reasoning); the Opus rows are tiny (+125 to +308) — but that gap is block size, not display: those captured Opus blocks were shallow (short signatures), not cheap because their text is empty. Measured on a deep prompt [2026-06-25]: an Opus 4.8 omitted block — zero readable text, 4,300-char signature — cost +1,522 tokens to replay (pure signature), and the same prompt on summarized Sonnet cost +6,106. The replay bill tracks serialized block size, dominated by the signature, not whether you can read it. Every number above is positive — nothing was dropped.

The Fable/Mythos exception [docs]: the Fable/Mythos family's thinking blocks are dropped (unbilled) when replayed to a different model. Not reproducible here (those models 404 on the test host), but the contrast is documented: the entire Opus/Sonnet/Haiku family replays freely.

Key distinction: omitting the thinking text (Opus 4.7/4.8 + Fable) is not the same as dropping blocks cross-model (Fable/Mythos only). Encryption is orthogonal to transfer: a sealed, empty-text Opus block still replays across models — confirmed live by resuming an Opus session on Sonnet, where the sealed Opus block was carried verbatim, accepted, and billed. Caveat: carried + accepted + billed is what's measured — it does not prove the target model semantically reuses another model's reasoning; only that the block crosses the boundary and costs you.

The documented strip rule, and Claude Code's override

There is a documented rule about when previous thinking blocks are stripped from context [docs]:

"When a non-tool-result user block is included: on Opus 4.5+ and Sonnet 4.6+, previous thinking blocks are kept; on earlier Opus/Sonnet models and all Haiku models, all previous thinking blocks are ignored and stripped from context."

So by default: stripped on all Haiku + earlier Opus/Sonnet; kept on Opus 4.5+ and Sonnet 4.6+. The trigger is a normal (non-tool-result) user turn; inside tool loops, thinking is kept either way.

Does that strip disturb the cache? Only if the model counts thinking as part of its cached prefix in the first place — and that differs by model. How to read the next table: the same warm cache, measured with thinking kept vs forcibly stripped, on each model. Identical rows = thinking was never in that model's cache key (stripping is free); a cache_read that collapses on STRIP = it was in the key (stripping re-keys it) [measured]:

model variant cache_read cache_create reading
Haiku 4.5 KEEP 64,202 0 rows identical → thinking isn't in Haiku's cache…
Haiku 4.5 STRIP 64,202 0 …so stripping it costs nothing
Sonnet 4.6 KEEP 72,499 15 cache_read collapses on STRIP → thinking is in Sonnet's cache…
Sonnet 4.6 STRIP 54,080 9,342 …so stripping re-keys ~9.3K tokens (read drops ~18K)

Haiku strips thinking before it reaches the cache or the bill, so a strip is free. Sonnet keeps it in the prefix, so removing it cold-rewrites that slice — on a keep-model, stripping thinking hurts.

But inside real Claude Code, neither default bites — because Claude Code always sends one field [measured]:

"context_management": {"edits": [{"type": "clear_thinking_20251015", "keep": "all"}]}

keep:"all" overrides Haiku's default strip and forces every model to retain thinking:

default API inside Claude Code (keep:"all")
Haiku 4.5 strips (excluded from cache + billing) keeps (in cache + billed)
Sonnet 4.6 keeps keeps

So in genuine Claude Code usage every model keeps thinking; the Haiku-strip behavior only appears in raw API usage that omits keep:"all".

Switching Haiku → Sonnet stacks three costs [measured + inferred]

  1. Full cache miss — model-scoped caches; Sonnet cold-writes its whole prefix.
  2. Sonnet keeps and bills the thinking that default-Haiku had been ignoring — work that was "free" on Haiku now costs input tokens on Sonnet's cold write.
  3. Any prefix-leanness benefit from Haiku's strip evaporates the moment Sonnet serves a turn.

Editing thinking blocks in a proxy: position-dependent cache cost [measured]

If you run a proxy that rewrites or strips thinking blocks, the cost depends sharply on where the touched block sits — because everything before the edit stays cached and everything from the edit onward must be rewritten. What this shows: on a keep-model, warm the cache, then remove a single thinking block from a 13-message conversation and watch how much warm cache_read survives, depending on the removed block's position:

variant block removed at cache_read cache lost vs baseline
baseline 82,065
strip last msg 11 of 13 81,625 440
strip first msg 1 of 13 73,961 8,104

Same total prompt size either way — the loss is purely cheap cache_read turning into expensive cache_creation. Touch the last block and you lose almost nothing (440); touch the first and you re-key 8,104 tokens. The earlier the edit, the bigger the bill (~18× here).

⚠️ Mistake — assuming HTTP 200 means your cache survived. A 200 means "valid request," not "cache preserved." A proxy that removes a thinking block is accepted (200) but can silently convert tens of thousands of cheap cache reads into expensive cold writes. (Editing a block's content or signature is outright rejected — 400, integrity — but removing is tolerated and quietly costly.)

✅ Fix — Never mutate the prefix in a proxy. If you must touch messages, do surgical byte-fragment replacement on the tail only and confirm cache_read is unchanged on the next request.

What to do (end of Act II): Switching models and carrying thinking are the two transitions that blow up a bill. Both reward the same discipline: pick a model per conversation and leave the request body — including its thinking blocks and its interlocked thinking/effort/context-management parameters — exactly as Claude Code assembled it.