
















Thanks for the replies, it is a bit clear now. Regarding the token sequence priority in gpt-4o-mini: Is prefix of input tokens counted in same order System prompts → Tools → User messages ? Looks like there are 2 separate caches being used one for system prompts and one for tools? when changing system or tools prompts in subsequent requests, getting cached token count as zero if there are less than 1024 tokens and 1024 + incremented by 128 if there are more than 1024 tokens.
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。