惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
Jina AI
Jina AI
雷峰网
雷峰网
有赞技术团队
有赞技术团队
WordPress大学
WordPress大学
美团技术团队
V
V2EX
酷 壳 – CoolShell
酷 壳 – CoolShell
小众软件
小众软件
博客园 - Franky
博客园 - 三生石上(FineUI控件)
月光博客
月光博客
博客园 - 叶小钗
大猫的无限游戏
大猫的无限游戏
爱范儿
爱范儿
Hugging Face - Blog
Hugging Face - Blog
宝玉的分享
宝玉的分享
Last Week in AI
Last Week in AI
Apple Machine Learning Research
Apple Machine Learning Research
量子位
IT之家
IT之家
人人都是产品经理
人人都是产品经理
博客园_首页
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com

Forbes - Consumer Tech

This Unhackable Quantum Navigation System Is The Size Of A Loaf Of Bread Apple At 50 — A Leadership Shift And An AR Future We Are Under-Investing In Robotics ... 90% Of Humanoid Robots Are Made In China Ditch The Apple White: Beats Expands Colorful Cable Line-Up With New 10-Foot Option Satechi’s New ChargeView 140W Desktop GaN Charger With Real-Time Display The Hasselblad In Your Pocket: Oppo’s Find X9 Ultra Challenges The Galaxy S26 Ultra There's No Such Thing As Brain Honey How AI Agents Could Rebuild Fashion’s Visual Production Layer QClaw Goes Global. The Agent Built Itself In 5 Days Apple’s Tim Cook Exit Hides A $4 Trillion Agentic AI Power Move EZQuest Reveals A New Line Of Pro Series USB-C Hubs For MacBook Neo Samsung Galaxy Z TriFold 2 Already In The Works, Report Claims Apple Revealed New Siri Release Date For iPhone, Latest Report Claims How Arcani’s HARK Is Designed For Modern Battlefield Acoustics The Newest Trend In Tech Embraces Femininity And Fun Samsung’s 75R95H Ushers In A New World Of LCD TVs New Apple iPhone Fold Design Pushes Smartphone Rivals To Go Wider And Taller iPhone 18 Pro Report: Four New Colors Leak As Apple Cancels Popular Shade Nothing’s Design-Led Strategy: Carl Pei Reveals The Tech Brand’s Philosophy iOS 26.5 Release Date: When To Expect Your iPhone Messaging Upgrade Google Pixel And Highsnobiety Build A Talent Pipeline For Fashion Android Circuit: Samsung Raises Galaxy Prices, Oppo Pad Mini Teased, Microsoft Closing Outlook App Apple Loop: iPhone Fold Launch Dates, iPad Air Upgrade, iPhone 18 Pro Specs Comcast $117.5 Million Breach Settlement — Are You Eligible? Amazfit Cheetah 2 Pro Takes Aim At The Garmin Audience Disney’s Launches ‘Infinity Vision’ Certification For Premium Theaters SoundPeats Reveals New Air6 HS Semi-Open Wireless Earbuds Amazon’s $11.57 Billion Leap Into Space: A Challenge To Starlink Meta Quest 3 Hit With $100 Price Increase Backblaze Stops Backing Up Dropbox And Others—Calls It An Improvement
AI’s Performance Gap Between Tests And Real Use Cases
Tim Bajarin · 2026-06-16 · via Forbes - Consumer Tech
AI sign

AI’s biggest risk isn’t future autonomy. It is its present unreliability.

getty

Last week, Anthropic released a white paper titled "When AI Builds Itself." As headlines went, it was bound to attract attention with the implication that AI would soon start building its own successors.

Anthropic’s call for coordinated consideration of how to prepare for pausing development before humans are no longer able to meaningfully guide the process has to be taken seriously. That warning comes in good faith from those who have a privileged view of what lies ahead.

I applaud this recommendation; it should have come much earlier than it did, given where we are in AI development. Even before ChatGPT launched in 2022, AI researchers already understood many of the pitfalls of AI without strong regulatory guidelines and rails.

The Immediate Problem Is Not Safety—It’s Reliability

However, the core economic issue in current AI systems is reliability. It affects every company, every developer, and everyone who pays for these tools. And it costs far more than it should.

Because a sizable portion of each dollar invested in current AI technologies buys you very little. Or worse, it prioritizes AI self-improvement over reliable problem-solving.

We’ve Seen This Pattern Before

In my decades in technology, I have watched many promises break after companies discovered reality. PCs promised to bring the power of computing to all and democratize innovation in the process. The Internet era offered instant connectivity anywhere. Mobile devices were supposed to give users true freedom and unleash creativity.

AI is now entering its discovery phase, when people are finding that the new tool does not quite live up to expectations. The promise is reliability: the ability to solve real problems when people rely on it in the real world.

Benchmarks Don’t Reflect Real-World Performance

Frontier models achieve incredible benchmark percentages of 80%, 90% and above on standardized tests. But put them up against real users with real, often challenging tasks, and everything changes.

And that shouldn't surprise you. Frontier models' inability to consistently perform is now well-documented. They may get 90% on benchmark testing, but provide consistent output less than a quarter of the time when run in production with the same task. Part of this phenomenon is easy to explain.

Many of the questions used in benchmark tests have been circulating on the web for years, and frontier models have effectively learned to recognize them. They know answers, not the solutions to problems you face.

The Hidden Cost: Wasted Work And Silent Errors

There are many names for this problem, from annoying to painful. Whatever term you choose, the point is the same: a sizable amount of effort is wasted because it produces nothing.

Not infrequently, not occasionally, but every day, across different industries, on various tasks, on a regular basis. And it shows up as wasted effort. It means retrying, reprompting, and saying "No, that's not what I asked." Even worse, it might mean accepting an answer that looks confident but turns out to be completely wrong, creating yet another problem for someone else weeks later.

Enterprise buyers are famous for being very patient. But patience, as all other virtues, has its limits. The day the return-on-investment calculation takes place always comes, and at that moment, reliability is the metric that counts, not the maximum benchmark percentage, but the minimum reliable floor.

The Problem Is Solvable—But Not Yet Solved

And the good news is that this problem is engineering-solvable. The industry already recognizes it, and large investments in reasoning, verification and reliability layers are underway. My concern, however, is that the safety discussion focuses on AI gaining too much power and autonomy, while the reliability issue is that AI keeps delivering an expensive solution that is confidently wrong.

The Industry Is Solving The Wrong Problem

Anthropic’s considerations need to be heard. But the industry should also confront the problem it already has: a widening reliability gap that undermines real-world value. The competition for ever more capable AI systems will not be decided by benchmark gains alone. If today’s systems cannot reliably complete a single task, the race for more powerful models will slow. Not because of regulation, but because the economics will no longer justify the effort.