惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Recent Announcements
Recent Announcements
J
Java Code Geeks
U
Unit 42
GbyAI
GbyAI
大猫的无限游戏
大猫的无限游戏
L
LangChain Blog
D
Docker
F
Fortinet All Blogs
N
Netflix TechBlog - Medium
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
罗磊的独立博客
I
InfoQ
The Cloudflare Blog
小众软件
小众软件
V
Visual Studio Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Engineering at Meta
Engineering at Meta
S
SegmentFault 最新的问题
爱范儿
爱范儿
Hugging Face - Blog
Hugging Face - Blog
P
Proofpoint News Feed
V
V2EX
月光博客
月光博客
Martin Fowler
Martin Fowler

Sid's Blog

When Thinking Goes Sedentary | Sid's Blog I Just Want to Search Don't Let AI Fill in All the Important Blanks Double Entry Programming | Sid's Blog The Newest Instagram "Exploit" is the Goofiest I've Seen Google's Antigravity Bait and Switch Agentic Coding is Burning Me Out An Audience of One: Cutting Corners on Unscalable Personal Software No, I Won't Download Your App. The Web Version is A-OK. The 667MHz Machine | Sid's Blog Never Buy A .online Domain Accelerated FOMO in the Age of AI ai;dr | Sid's Blog App Store Review Feels Like RNG, and That’s the Problem Welcome to the Machine | Sid's Blog
On DeepSeek V4 Flash and Cheap Intelligence
2026-08-03 · via Sid's Blog

The cost of tokens and intelligence seems to be plunging, despite what my own internet bubble led me to believe was going to happen. Between DeepSeek V4 Flash going toe to toe with many SOTA models at a very, very small fraction of the cost and GPT-5.6 Luna getting a massive price cut, the narrative that intelligence would remain expensive, if not increase over time, is looking increasingly difficult to defend in my head.

Full disclaimer on my workloads though: my intelligence needs are very prosaic. I mainly use AI for code: a lot of C/C++ for embedded devices plus generic web endpoints and dashboards to ingest and present data. I also end up needing a ton of Swift. So not exactly (or exclusively) webslop but not cutting edge research work either. (I’m not using it to disprove the Jacobian conjecture, that’s for sure.)

The release of DeepSeek V4 Flash has upended tokenomics and has caught a lot of people off guard with its performance and cost. Artificial Analysis’ analysis shows it costing 1/100th the cost of Fable while being fairly competitive in various benchmarks. Yes, not a typo. 1/100th. 1/60th the cost of GPT-5.6 Sol, 1/80th the cost of Opus 5. Arena’s leaderboard paints the same picture.

Maybe it’s recency bias but never before could you do so much for so little. Intelligence that’s so cheap and so good that it’s too cheap to meter. I no longer find myself model switching with Claude Code just to protect my 5-hour, and weekly, quota.

And yes, it meanders around on long horizon tasks, is slower, and not very token efficient so I just end up spinning up way more subagents and have something else orchestrate and coordinate. Sure, it doesn’t have the taste of Opus, but those are areas where I can step in and fill in the blanks. The weaknesses are things I could live with and engineer around.

When the cost of a workflow, any workflow, drops from a few dollars to a few cents, many ideas that were previously only viable for high-value enterprises, or just untenable altogether, can suddenly become practical for everyday use. I’m not downplaying the impact of true frontier intelligence, but intelligence at mass scale is where things start to get really interesting.

Between DeepSeek V4 Flash, MiniMax H3, and Qwen 3.8 27B, this has been an incredible week in open source LLM history. I’ve been cautiously optimistic for the longest time but these last few weeks have nudged me deep into pure meliorism.


If you've reached this far, thank you for reading! :)

I thought retiring in my mid 30s after a few exits would be fun but I've just been bored and a bit undersocialized without morning Slacks and emails to wake up to. If you’re building something interesting and could use an extra set of hands to ship, or just want to say hi, feel free to reach out. My inbox is open.