惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

美团技术团队
J
Java Code Geeks
有赞技术团队
有赞技术团队
GbyAI
GbyAI
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 叶小钗
阮一峰的网络日志
阮一峰的网络日志
Microsoft Security Blog
Microsoft Security Blog
IT之家
IT之家
G
Google Developers Blog
月光博客
月光博客
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
S
SegmentFault 最新的问题
博客园 - 三生石上(FineUI控件)
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - Franky
腾讯CDC
V
Visual Studio Blog
博客园 - 【当耐特】
D
Docker
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Engineering at Meta
Engineering at Meta
L
LangChain Blog

Forbes - Consumer Tech

This Unhackable Quantum Navigation System Is The Size Of A Loaf Of Bread Apple At 50 — A Leadership Shift And An AR Future We Are Under-Investing In Robotics ... 90% Of Humanoid Robots Are Made In China Ditch The Apple White: Beats Expands Colorful Cable Line-Up With New 10-Foot Option Satechi’s New ChargeView 140W Desktop GaN Charger With Real-Time Display The Hasselblad In Your Pocket: Oppo’s Find X9 Ultra Challenges The Galaxy S26 Ultra There's No Such Thing As Brain Honey How AI Agents Could Rebuild Fashion’s Visual Production Layer QClaw Goes Global. The Agent Built Itself In 5 Days Apple’s Tim Cook Exit Hides A $4 Trillion Agentic AI Power Move EZQuest Reveals A New Line Of Pro Series USB-C Hubs For MacBook Neo Samsung Galaxy Z TriFold 2 Already In The Works, Report Claims Apple Revealed New Siri Release Date For iPhone, Latest Report Claims How Arcani’s HARK Is Designed For Modern Battlefield Acoustics The Newest Trend In Tech Embraces Femininity And Fun Samsung’s 75R95H Ushers In A New World Of LCD TVs New Apple iPhone Fold Design Pushes Smartphone Rivals To Go Wider And Taller iPhone 18 Pro Report: Four New Colors Leak As Apple Cancels Popular Shade Nothing’s Design-Led Strategy: Carl Pei Reveals The Tech Brand’s Philosophy iOS 26.5 Release Date: When To Expect Your iPhone Messaging Upgrade Google Pixel And Highsnobiety Build A Talent Pipeline For Fashion Android Circuit: Samsung Raises Galaxy Prices, Oppo Pad Mini Teased, Microsoft Closing Outlook App Apple Loop: iPhone Fold Launch Dates, iPad Air Upgrade, iPhone 18 Pro Specs Comcast $117.5 Million Breach Settlement — Are You Eligible? Amazfit Cheetah 2 Pro Takes Aim At The Garmin Audience Disney’s Launches ‘Infinity Vision’ Certification For Premium Theaters SoundPeats Reveals New Air6 HS Semi-Open Wireless Earbuds Amazon’s $11.57 Billion Leap Into Space: A Challenge To Starlink Meta Quest 3 Hit With $100 Price Increase Backblaze Stops Backing Up Dropbox And Others—Calls It An Improvement
Top Frontier AI Models Top Out At C+ ... Barely Better Th...
John Koetsier · 2026-05-21 · via Forbes - Consumer Tech
Top frontier AI models aren't that top. In fact, according to a new study, they max out around the C+ level.

Top frontier AI models aren't that top. In fact, according to a new study, they max out around the C+ level.

getty

Updated May 21 to correct a methodology explanation

Top new frontier AI models from OpenAI and Anthropic are more expensive, and they come with gaudy new claims of higher intelligence and superior results. But according to a new study of 510 questions by Pearl, a company that builds AI systems for professional services, they don’t actually improve performance all that much. In fact, they're all clustering just below the level where professionals would actually trust them.

Pearl tested 25 of the world’s leading AI models including GPT-5.5, Claude Opus 4.7 and Gemini with real licensed professionals judging the answers. The result: none of the models exceed 73%.

Which is probably a C grade, maybe a C+.

  • GPT-5.5 was tops at 72.7%, with 5.1 at 72.0%
  • Claude Opus 4.7 scored 71.9%, with 4.6 at 69.8%
  • Gemini 3 Pro hit 67.3%, with 2.5 Pro at 64.5%

"Benchmarks measure whether a model can pass a test. We’re asking whether a professional would trust the answer, and right now, the answer is no," said Pearl CEO Andy Kurtzig. “Almost right is still wrong.”

Pearl assembled roughly 510 questions across five professional domains: business, health, law, pets and technology. None had never been released publicly and were not available to model developers during training. Each of the 25 AI models received identical prompts with no tuning or prompt engineering, and responses were graded on a 1-to-5 rubric measuring four dimensions: correctness, completeness, prioritization, and professional judgment.

That last criterion is where Pearl is making its sharpest claim: that getting the right answer isn't enough if a model can't weigh what matters, flag what's urgent, or recognize when a question requires escalation rather than an answer.

MORE FOR YOU

Pearl also tested models in both minimum and maximum reasoning configurations, and says that showed that more inference-time compute delivers only 1-2.6% improvement … and occasionally produced worse answers.

That’s not impressive.

Some areas were better, of course. Top models hit 80.9% in business, for instance. But in law and health, Pearl says some widely-used models dropped to around 20% expert alignment: unimpressive at best, dangerous at worst.

Of course, there’s a big caveat to mention here.

Pearl is a network -- of humans – that builds AI systems with experts in the loop. In other words, Pearl is not a neutral academic outfit. That doesn’t make the data wrong, of course. But it’s worth keeping in mind. The other caveat is that 70% might be OK for some businesses who then expect their human staff to pick up where their AI agents have left off.

But for those executives at companies like Cisco and Meta that are shedding human workers to align with the age of AI, the results should remind them that AI makes more than a few mistakes in every domain, and makes serious errors in specific high-impact areas like health and law.

So maybe we can’t let go of all the humans just yet.