惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

大猫的无限游戏
大猫的无限游戏
U
Unit 42
T
Tailwind CSS Blog
罗磊的独立博客
WordPress大学
WordPress大学
小众软件
小众软件
Recent Announcements
Recent Announcements
博客园 - 聂微东
Jina AI
Jina AI
云风的 BLOG
云风的 BLOG
博客园 - 【当耐特】
爱范儿
爱范儿
Microsoft Azure Blog
Microsoft Azure Blog
GbyAI
GbyAI
V
V2EX
博客园 - 三生石上(FineUI控件)
I
InfoQ
雷峰网
雷峰网
G
Google Developers Blog
阮一峰的网络日志
阮一峰的网络日志
B
Blog
腾讯CDC
A
About on SuperTechFans
博客园 - 叶小钗

Futurism

Chinese Government Dismisses Tech Billionaire Calls for an AI Slowdown as "Fear-Mongering" Sam Altman Now Trying to Gain Control of Electric Grid Man Announces That He Has Synthesized a New Schizophrenia Treatment in His Garage, Based on a Formula Devised by ChatGPT Dyson's New AI Toothbrush Conducts Video Surveillance on the Inside of Your Mouth Man Pretends to Hallucinate in Job Interview With an AI Bot, Causing It to Go Haywire Even Babies Are Still Way Better at Learning Than AI Models Farmer Horrified as AI Gives Bad Advice That Kills 25 Acres of Crops Better Hope You Don't Need an Ambulance in New Orleans, Because the City's 911 Operators Are Now AI Amazon Is Gutting Its AI Division After Sustained Failure Suspicion Grows About OpenAI's Tale About Its Rogue Hacker AI A Chinese AI Model Just Shot to Number One on the Charts, Sending Shockwaves Through the American Tech Industry Tech Bros Puzzled by Why AI Hasn't "Massively Disrupted" Books Yet AI Companies Are Learning an Ironic Lesson as the People They Pay to Improve Their Chatbots Are Just Feeding AI Slop Into Them Sports Journalists Asked Microsoft’s Copilot to Predict World Cup Matches, and the Results May Surprise You Researchers Put AI Models in Charge of Analyzing Sports, and They Choked Spectacularly Fans Aghast as New York Jets Say They’re Switching to AI Companies That Adopted AI Agents Alarmed to Discover They’re Botching Incredibly Important Tasks Hackers Find That Inaudible Sounds Hidden in Podcasts or Random Videos Can Hijack Your AI Voice Chatbot Democrats’ 2024 Election Autopsy Shows Signs of Sloppy AI Generation Being a Crappy Boss to AI Chatbots Pushes Them Toward Spouting Marxist Rhetoric and Organizing With Their Compatriots, Researchers Find Amazon Employees Forced to Hit Quotas on AI Use, Immediately Start Using it for Everything Except Work These Smart Glasses That Show Captions of What Everyone’s Saying Without a Creepy Spy Camera Actually Seem Pretty Awesome AI Appears to Be Trapping Certain Job Applicants in a Limbo Where They Never Get an Interview for “Reasons” That Are Completely Unfair The AI Industry Is Secretly Powered by Homeless People ChatGPT Is Saying VWeird Things in Chinese America Trembles as Transportation Secretary Announces Plans for Air Traffic Controllers to Lean on AI Tools Today Is the Day Anthropic Promised That Fully Autonomous Employees Would Be Tearing Through the Business World China Is Starting to Pull Ahead of US in AI Race Berklee College of Music Students Furious That It’s Offering an AI “Songwriting” Class Usually, Young People Embrace New Technology. Gen Z’s Attitude Toward AI Should Worry the Entire Tech Industry
Frontier AI Is Faceplanting at Real-World Workplace Tasks
Joe Wilkins · 2026-07-22 · via Futurism

A stylized photo illustration of a robotic hand touching a computer keyboard.

Illustration by Tag Hartman-Simkins / Futurism. Source: Shutterstock

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

To date, AI industry spending has topped $1.6 trillion, and shows no sign of slowing anytime soon.

So what do we actually have to show for it? Historically, it’s been a whole lot of nothing: as numerous studies have shown us, tools like AI chatbots and autonomous agents have been ineffective at completing real world tasks in a competent way.

The tech industry insists that’s all about to change within the next few years, as AI’s capabilities grow by leaps and bounds, enabling economic growth the likes of which the world has never seen. But is it really?

Not necessarily. A new study out of the University of California Berkeley’s Center for Responsible, Decentralized Intelligence — flagged by the College Fix — shows that frontier AI tools of all makes and models are still incapable of completing the vast majority of workplace tasks at an acceptable level, throwing a major wrench in the tech industry’s assertions that the AI revolution is imminent.

To come to that conclusion, the UC researchers designed a rigorous assessment they call the “Agents’ Last Exam,” developed to test “job-readiness” across numerous state-of-the-art AI models. Basically, the ALE — an impish riff on “Humanity’s Last Exam” — is designed to put an AI system through its paces, covering “more than 1,500 expert-sourced tasks spanning 55 occupations,” the researchers wrote in apress release.

Those test spans the typical line-up of AI-exposed jobs like software engineering and graphic design, but also a substantial number of jobs whose fates remain less certain, such as maritime engineering, agriculture, audio production, and public health operations.

Using the ALE benchmark, researchers took a hard look at advanced “closed” models — proprietary AI systems developed by private companies — like Anthropic’s Fable 5, OpenAI’s GPT-5.5, Cursor’s Composer 2.5, and Google’s Gemini 3.1 Pro. (For good measure, they also looked at two open-source models by Chinese developers.)

As cutting-edge as these AI models are, the research found that they’re far from ready for the complex needs of the modern workplace. Out of all of the models put through the gauntlet, each of them failed spectacularly. OpenAI’s GPT-5.5 came in with the highest score: a passing rate of just 24 percent overall.

“Today’s agents can solve a meaningful fraction of professional tasks,” the researchers wrote. “However, when we look at the hardest tasks that require sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance.”

And as tasks became more complicated, even those meager aggregate scores fell off fast.

“On ALE’s hardest tier, every frontier agent we tested, including Fable 5, achieved a 0 percent success rate,” the presser explains.

The researchers also break down some cost considerations. The cutting edge Fable 5, they note, delivers “similar performance” to models like GPT-5.5 and Composer 2.5, “while costing roughly 4-12× more per completed task.”

Despite the horrible test results, researchers caution that the technology could still upend the job market for more AI-exposed — as plenty of corporate executives have shown us, the tech doesn’t need to work particularly well to keep workers on their back heels.

“Even if current pass rates remain relatively low, occupations dominated by routine and well-defined procedures are likely to experience disruption first, while decision-intensive roles will remain more resilient for longer,” Berkeley computer science researcher and study co-author Dawn Song told College Fix.

“The key factor,” Song added, “is not the industry itself, but the nature of the work.”

More on AI: OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin