惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
量子位
D
DataBreaches.Net
博客园 - 司徒正美
J
Java Code Geeks
博客园 - 【当耐特】
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
aimingoo的专栏
aimingoo的专栏
B
Blog
The Cloudflare Blog
D
Docker
I
InfoQ
爱范儿
爱范儿
MongoDB | Blog
MongoDB | Blog
腾讯CDC
月光博客
月光博客
Hugging Face - Blog
Hugging Face - Blog
Microsoft Azure Blog
Microsoft Azure Blog
Vercel News
Vercel News
阮一峰的网络日志
阮一峰的网络日志
小众软件
小众软件
S
SegmentFault 最新的问题
GbyAI
GbyAI
有赞技术团队
有赞技术团队

Release Notes on DigitalOcean Documentation

Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note Release Note
Release Note
DigitalOcean · 2026-07-01 · via Release Notes on DigitalOcean Documentation

Last verified 1 Jul 2026

Prompt caching for open-source models in serverless inference chat completions and responses API is now in public preview. Open-source models cache context automatically, so you do not need to set the cache_control or prompt_cache_retention parameters.

Prompt caching is available for the following open-source models:

  • DeepSeek V3.2
  • DeepSeek V4 Pro
  • DeepSeek V4 Flash
  • Kimi K2.5
  • Kimi K2.6
  • GLM 5
  • GLM-5.1
  • GLM-5.2
  • gpt-oss-120b
  • MiMo V2.5
  • MiMo V2.5 Pro
  • MiniMax M2.5
  • Qwen 3.5
  • Qwen3 Coder Flash

For more information, see Use Prompt Caching.