惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Google DeepMind News
Google DeepMind News
B
Blog RSS Feed
The GitHub Blog
The GitHub Blog
Recent Announcements
Recent Announcements
A
About on SuperTechFans
G
Google Developers Blog
aimingoo的专栏
aimingoo的专栏
U
Unit 42
WordPress大学
WordPress大学
Y
Y Combinator Blog
罗磊的独立博客
J
Java Code Geeks
Microsoft Azure Blog
Microsoft Azure Blog
腾讯CDC
博客园 - 叶小钗
Stack Overflow Blog
Stack Overflow Blog
Engineering at Meta
Engineering at Meta
Microsoft Security Blog
Microsoft Security Blog
GbyAI
GbyAI
V
V2EX
雷峰网
雷峰网
H
Hackread – Cybersecurity News, Data Breaches, AI and More
S
SegmentFault 最新的问题
酷 壳 – CoolShell
酷 壳 – CoolShell

Martin Alderson

Have the frontier labs mixed up AI safety and security? Reducing codebase cognitive debt through... quizzes? What GLM-5.3 Flash running on Chinese hardware actually means How I think about reducing AI costs Watch out for cache read costs I'm (mostly) picking models on speed now, not intelligence The first known runaway AI agent - or a very bad marketing stunt? Winners and losers in the coming AI margin collapse (part 2) GLM 5.2 and the coming AI margin collapse (part 1) Expert-aware quantisation: near-Q4 quality at near-Q2 size? A brief history of KV cache compression developments xAI is looking more like a datacentre REIT than a frontier lab Is datacentre sovereignty really that important? I went on the Built for Turbulence podcast What's going on with Gemini? Managed agents are the new Lambda Open weights are quietly closing up - and that's a problem 29th August 2026: a scenario Figma's woes compound with Claude Design A little tool to visualise MoE expert routing Has Mythos just broken the deal that kept the internet safe? What next for the compute crunch? Telnyx, LiteLLM and Axios: the supply chain crisis Using agents and Wine to move off Windows Why Claude's new 1M context length is a big deal How to use the Qwen 3.5 LLMs to OCR documents No, it doesn't cost Anthropic $5k per Claude Code user Is the AI Compute Crunch Here? Why on-device agentic AI can't keep up Using OpenCode in CI/CD for AI pull request reviews
The summer of open weights
Martin Alderson · 2026-08-23 · via Martin Alderson

Over the winter of 2025, after the release of Opus 4.5, coding agents grew tremendously and usage exploded. I think this summer is proving itself to be a similar tipping point for open weight models.

The compute crunch and pricing

As I argued in my margin collapse blogs (part 2 here), we're starting to see some very aggressive moves on pricing. We've seen OpenAI cut the cost of 5.6 Luna - its fast, cheapest tier - by 80%, and now Sol - its flagship - by 20%.

Meta is also offering its open model Muse Spark 1.2 for an almost-free price of $0.10/$0.20 per MTok on its contributor tier (where Meta may train on your data - standard pricing is $1.25/$4.25), with currently the cheapest API price for cache reads of $0.002 per MTok on that tier (!).

Anthropic hasn't matched this pricing yet, but the FT is leading with a story about the poor uptake of Fable 5 - Anthropic's $10/$50 frontier model, its most expensive tier - (tl;dr: it's too expensive) - headline: "Anthropic's best AI model struggles to attract users as cheaper tools thrive". And their Claude Developer social media account is suggesting they are still (extremely?) compute starved, wanting to make their weekly limit increases permanent but struggling for capacity:

ClaudeDevs tweet: extending 50% increase to weekly Claude Code limits through August 31, noting capacity may be tight

By no means am I suggesting that Anthropic is in real trouble here - they have very impressive market share, but if the market starts to move towards much cheaper models, (currently) I believe they're the lab with the least ability to respond price wise because of their lack of available compute.

A plethora of alternative models are here

We've now got at least five AI labs outside of OpenAI, Anthropic and Google[1] offering very good models - Z.AI, DeepSeek and Kimi - plus Meta and Grok. It's been strongly suggested that Meta is going to release their frontier models as open weights, which leaves us with four open weight models of good enough quality to drive agentic sessions.

No doubt there will be more - the Ox Alpha stealth model has been getting a lot of hype[2] - but it really indicates to me that there is a huge amount of competition for this inference.

In my eyes there are two possible scenarios that play out here:

The first one is that the gap between frontier and challenger/open weights models continues to decrease substantially - to the point where it becomes almost a commodity between these models. This is extremely bad news for labs built around proprietary models. Right now, this seems to very much be the path ahead.

However, the other scenario I wouldn't discount is a huge leap from the frontier labs, which would then expand the gap. While historically this has been what happens - open weights close the gap, then just when it looks like they are about to catch up, OpenAI/Anthropic puts a new release out which expands the gap again. This time I do feel it's different - the gap has never been this small, and I'm struggling where to see this huge jump would come from. But regardless, it's definitely possible - and with many trillions of dollars of IPO market cap riding on this - I wouldn't rule out any surprises like this.

It's important to note as well that this leap could come through token efficiency too, not just pure "intelligence". The flagship models from Anthropic and OpenAI are ~5-10x more expensive than the best open weight models per token. But, it's fair to say the open weight models tend to use quite a lot more tokens per task. So if a hypothetical future Fable 6 could achieve similar intelligence to Fable 5, but use 10x less tokens to achieve the same end goal, it'd still be a very competitive model.

The lack of compute is a wildcard, though

I recently saw this very interesting interview with Gavin Baker, who makes the salient point that the industry has a lot of compute that was reserved for say $2/GPU-hour in 2-3 year commitments going to roll "off contract" and therefore going to be repriced.

Given current rates for Blackwell GPUs are significantly above that, he makes the point that these people are hoping to pay $4/GPU-hour - the win would be a doubling of underlying costs.

I think this really makes the token efficiency angle even stronger. There is enormous competitive advantage - more so than pure intelligence I think right now - in being able to serve these models more efficiently.

So what happens next?

Winter 2025 was about agents needing frontier intelligence at any price. This summer is about good enough intelligence at a tenth of the price - and whether the frontier labs can keep charging a premium for being slightly better.

Right now the pricing power isn't with the smartest model, it's with whoever has the megawatts to spare. OpenAI can afford to cut Luna 80% and put Sol on a three-month 20% promo because it has the capacity and the efficiency gains to back it up. Anthropic - renting 300MW from SpaceX at $1.25bn a month precisely because it doesn't - has to extend limits with a caveat that "capacity may be tight".

That flips the usual tech story. For decades software captured the margin and hardware was the commodity. Here the hardware is the margin, as I argued in xAI's new rental business. The handful of open-weight hosts - Fireworks, Together, Cloudflare and a dozen others - all have the same incentive: squeeze more tokens per GPU, because whoever does wins the price war regardless of who trained the model.

If that efficiency race keeps going, then cheap inference keeps pulling demand forward. The frontier labs' two escape routes are the ones I laid out in the margin collapse series: stay meaningfully ahead on intelligence, or make the model so much more token-efficient that the sticker price stops mattering. Fable 5 being called "too expensive" at $10/$50 tells you neither is guaranteed.

I wouldn't bet against a surprise leap - we've seen the gap close and re-widen before, and trillions in IPO market cap is a strong motivator. But this is the first summer where the open weights are close enough that most agentic work just doesn't need the frontier. That's a genuine tipping point, just like agents were last winter.


  1. Though I do question Google's addition to this list, given their very obvious troubles (what's going on with Gemini) at the frontier ↩︎

  2. It looks like this is another GLM model, but if not it could be another significant challenger in the market ↩︎