












I honestly don’t understand ALL of the linked thread on prompt caching either 😃 but I provided it to Claude with your questions… Scenario 1: Fewer cached tokens than total Caching starts at 1024 tokens and increases in 128-token blocks Maximum cached tokens will be the largest multiple of 128 that fits your total Example: With 5672 total tokens, you’ll see 5432 cached (42 blocks of 128 + 1024) Scenario 2: Large cache drop with small changes KV cache requires exact prefix matches Eve...
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。