惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
J
Java Code Geeks
H
Hackread – Cybersecurity News, Data Breaches, AI and More
D
Docker
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 聂微东
MyScale Blog
MyScale Blog
H
Help Net Security
Last Week in AI
Last Week in AI
T
The Blog of Author Tim Ferriss
M
MIT News - Artificial intelligence
大猫的无限游戏
大猫的无限游戏
酷 壳 – CoolShell
酷 壳 – CoolShell
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
P
Proofpoint News Feed
博客园 - 叶小钗
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Y
Y Combinator Blog
Recent Announcements
Recent Announcements
F
Fortinet All Blogs
Martin Fowler
Martin Fowler
Microsoft Security Blog
Microsoft Security Blog
T
Tailwind CSS Blog
aimingoo的专栏
aimingoo的专栏

The Decoder

Google files first joint lawsuit with FBI over Chinese AI scam network, OpenAI blocks PRC influence clusters The AI industry's platform trap is starting to look a lot like Microsoft's OpenAI buys Ona to push Codex toward long-running, autonomous coding tasks Jeff Bezos' AI startup Prometheus closes $12 billion round at a $41 billion valuation Free Deezer tool lets users on any streaming service check their playlists for AI music OpenAI vs. Anthropic: A price war over API tokens is brewing Dario Amodei's new essay reads like a Cold War playbook for the AI age Claude Fable 5: Anthropic admits "wrong tradeoff" after invisibly throttling rival AI researchers Google's new open model DiffusionGemma generates text from noise instead of word by word OpenAI's IPO slips as Altman tells staff to expect a public offering "within the next year" Anthropic study shows AI needs hours, not weeks, to build exploits from security patches OpenAI wants its biggest data center yet, and Nvidia would back the bill Claude Fable 5: The first Mythos model is powerful, expensive, and heavily filtered Germany's National Security Council greenights an AI Safety Institute modeled after the UK's AISI Google's NotebookLM now runs its own cloud computer with code execution and agent-based research Anthropic releases Claude Fable 5 and Mythos 5 with major gains in coding and science Google's Gemini 3.5 Live Translate delivers real-time voice translation across 70+ languages SpaceX wants to put data centers in orbit, and Musk says it's no big deal Landmark German ruling declares Google's AI Overviews are Google's own words and makes it liable for false answers Beijing's $295 billion AI buildout would require 80 percent domestic chips, locking out US suppliers Apple Intelligence gets a second shot with help from Google and Nvidia OpenAI now says "entirely automating everything is not the future we want" OpenAI says going public is "a complicated set of tradeoffs" and is unsure about the timing Microsoft Research's Lens proves detailed captions matter more than raw scale for training efficient image generators Intel gets a second life as Google and Nvidia explore it as a TSMC backup for AI chips Most companies are flying blind on AI spending Frontier Radar #3: How agentic AI is turning tokens into a business metric Instagram AI chatbot breach may have affected over to 20,000 accounts, Meta discloses Microsoft tightens rules for conflict zones after investigation into Israel's military use of Azure Moonshot AI targets a $30 billion valuation, more than six times its late-2025 worth
Anthropic ships Claude Opus 4.8 as a "modest but tangible...
Matthias Bastian · 2026-05-29 · via The Decoder

Anthropic's latest flagship model, Claude Opus 4.8, leads most benchmarks and is designed to be more upfront about its own mistakes.

Anthropic says Opus 4.8 beats both its predecessor and OpenAI's GPT-5.5 and Google's Gemini 3.1 Pro across most tested categories. On agentic coding (SWE-Bench Pro), the model hits 69.2 percent, up from 64.3 percent for Opus 4.7 and 58.6 percent for GPT-5.5. For multidisciplinary reasoning (Humanity's Last Exam), Opus 4.8 scores 49.8 percent without tools and 57.9 percent with tools, the highest marks in the field.

Opus 4.8 stacked against Opus 4.7, GPT-5.5, and Gemini 3.1 Pro. | Image: Anthropic

Less fake progress, more honesty

Anthropic calls the model's improved honesty one of its most noticeable upgrades. AI models have a habit of jumping to conclusions and claiming progress that falls apart on closer look. It's a widespread problem.

"Early testers report that Opus 4.8 is more likely to flag uncertainties about its work and less likely to make unsupported claims," Anthropic says. The company backs that up with its own coding evaluations, where the model lets bugs slip through without comment about four times less often than Opus 4.7.

The model also sets new highs on prosocial traits like supporting user autonomy. Deception attempts and other unaligned behavior are said to be at Claude Mythos levels. Details are in the Claude Opus 4.8 System Card. The first Mythos-class models are expected to roll out to all customers in the coming weeks, once all safety measures are in place, the company says.

Dynamic workflows and effort controls steal the show

The new features Anthropic shipped alongside the model may matter more than the model update itself, which the company calls "modest but tangible."

The biggest is "dynamic workflows." The model can plan a task and then spin up hundreds of parallel sub-agents in a single session. Anthropic says Claude Code with Opus 4.8 can now handle codebase-wide migrations across hundreds of thousands of lines, from planning all the way to merge. The feature is available on Enterprise, Team, and Max plans.

On claude.ai and in Cowork, there's now an effort control next to the model picker. It lets you decide how hard Claude works on a given response. Crank it up for deeper thinking and better results. Turn it down for faster answers that use less of your rate limit.

Opus 4.8 defaults to "high." For tough tasks, Anthropic recommends "extra" (called "xhigh" in Claude Code) or "max." These modes burn more tokens, but Anthropic says higher rate limits for Claude Code users help offset that. Anthropic's advice is to just pick whatever level feels right for the task.

API prices stay the same, fast mode gets cheaper

Fast Mode, which runs Opus 4.8 at 2.5x speed, now costs a third of what it did for earlier models. Pricing sits at $10 per million input tokens and $50 per million output tokens.

Standard prices are unchanged from Opus 4.7: $5 per million input tokens and $25 per million output tokens. But 4.7 was already about 30 to 40 percent pricier in practice than its predecessor, 4.6, because it chewed through more tokens without delivering noticeable gains on many everyday tasks.

Opus 4.8 might actually cost less to run

According to Artificial Analysis, Opus 4.8 could ease that 4.7 price bump. On the GDPval-AA benchmark, which tests real-world knowledge work tasks, the model needs 15 percent fewer passes per task and 35 percent fewer output tokens than Opus 4.7.

In practice, that could mean noticeably lower costs. But Opus 4.8 still uses roughly 30 percent more passes than OpenAI's GPT-5.5, the second-place model.

At the "max" effort level, Opus 4.8 scored 1,890 points on GDÜVvall-AA, 137 points above Opus 4.7 and 121 points ahead of GPT-5.5, a win rate of about 67 percent head-to-head against GPT-5.5.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now