惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

A
Arctic Wolf
博客园 - 聂微东
F
Fortinet All Blogs
云风的 BLOG
云风的 BLOG
小众软件
小众软件
V
Visual Studio Blog
博客园 - 三生石上(FineUI控件)
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Apple Machine Learning Research
Apple Machine Learning Research
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
The Cloudflare Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园_首页
L
LangChain Blog
A
About on SuperTechFans
阮一峰的网络日志
阮一峰的网络日志
I
Intezer
T
The Blog of Author Tim Ferriss
Security Latest
Security Latest
C
CXSECURITY Database RSS Feed - CXSecurity.com
Know Your Adversary
Know Your Adversary
Simon Willison's Weblog
Simon Willison's Weblog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
P
Palo Alto Networks Blog
Scott Helme
Scott Helme
S
Secure Thoughts
Spread Privacy
Spread Privacy
T
Threat Research - Cisco Blogs
Attack and Defense Labs
Attack and Defense Labs
P
Privacy & Cybersecurity Law Blog
O
OpenAI News
H
Heimdal Security Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Help Net Security
Help Net Security
C
Cyber Attacks, Cyber Crime and Cyber Security
Blog — PlanetScale
Blog — PlanetScale
GbyAI
GbyAI
G
Google Developers Blog
博客园 - Franky
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
K
Kaspersky official blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
T
Tor Project blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Tenable Blog
Google Online Security Blog
Google Online Security Blog
PCI Perspectives
PCI Perspectives

Futurism

A Chinese AI Model Just Shot to Number One on the Charts, Sending Shockwaves Through the American Tech Industry Tech Bros Puzzled by Why AI Hasn't "Massively Disrupted" Books Yet AI Companies Are Learning an Ironic Lesson as the People They Pay to Improve Their Chatbots Are Just Feeding AI Slop Into Them Sports Journalists Asked Microsoft’s Copilot to Predict World Cup Matches, and the Results May Surprise You Researchers Put AI Models in Charge of Analyzing Sports, and They Choked Spectacularly Fans Aghast as New York Jets Say They’re Switching to AI Companies That Adopted AI Agents Alarmed to Discover They’re Botching Incredibly Important Tasks Hackers Find That Inaudible Sounds Hidden in Podcasts or Random Videos Can Hijack Your AI Voice Chatbot Democrats’ 2024 Election Autopsy Shows Signs of Sloppy AI Generation Being a Crappy Boss to AI Chatbots Pushes Them Toward Spouting Marxist Rhetoric and Organizing With Their Compatriots, Researchers Find Amazon Employees Forced to Hit Quotas on AI Use, Immediately Start Using it for Everything Except Work These Smart Glasses That Show Captions of What Everyone’s Saying Without a Creepy Spy Camera Actually Seem Pretty Awesome AI Appears to Be Trapping Certain Job Applicants in a Limbo Where They Never Get an Interview for “Reasons” That Are Completely Unfair The AI Industry Is Secretly Powered by Homeless People ChatGPT Is Saying VWeird Things in Chinese America Trembles as Transportation Secretary Announces Plans for Air Traffic Controllers to Lean on AI Tools Today Is the Day Anthropic Promised That Fully Autonomous Employees Would Be Tearing Through the Business World China Is Starting to Pull Ahead of US in AI Race Berklee College of Music Students Furious That It’s Offering an AI “Songwriting” Class Usually, Young People Embrace New Technology. Gen Z’s Attitude Toward AI Should Worry the Entire Tech Industry Sam Altman’s Coworkers Say He Can Barely Code and Misunderstands Basic Machine Learning Concepts
Frontier AI Is Faceplanting at Real-World Workplace Tasks
Joe Wilkins · 2026-07-22 · via Futurism

A stylized photo illustration of a robotic hand touching a computer keyboard.

Illustration by Tag Hartman-Simkins / Futurism. Source: Shutterstock

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

To date, AI industry spending has topped $1.6 trillion, and shows no sign of slowing anytime soon.

So what do we actually have to show for it? Historically, it’s been a whole lot of nothing: as numerous studies have shown us, tools like AI chatbots and autonomous agents have been ineffective at completing real world tasks in a competent way.

The tech industry insists that’s all about to change within the next few years, as AI’s capabilities grow by leaps and bounds, enabling economic growth the likes of which the world has never seen. But is it really?

Not necessarily. A new study out of the University of California Berkeley’s Center for Responsible, Decentralized Intelligence — flagged by the College Fix — shows that frontier AI tools of all makes and models are still incapable of completing the vast majority of workplace tasks at an acceptable level, throwing a major wrench in the tech industry’s assertions that the AI revolution is imminent.

To come to that conclusion, the UC researchers designed a rigorous assessment they call the “Agents’ Last Exam,” developed to test “job-readiness” across numerous state-of-the-art AI models. Basically, the ALE — an impish riff on “Humanity’s Last Exam” — is designed to put an AI system through its paces, covering “more than 1,500 expert-sourced tasks spanning 55 occupations,” the researchers wrote in apress release.

Those test spans the typical line-up of AI-exposed jobs like software engineering and graphic design, but also a substantial number of jobs whose fates remain less certain, such as maritime engineering, agriculture, audio production, and public health operations.

Using the ALE benchmark, researchers took a hard look at advanced “closed” models — proprietary AI systems developed by private companies — like Anthropic’s Fable 5, OpenAI’s GPT-5.5, Cursor’s Composer 2.5, and Google’s Gemini 3.1 Pro. (For good measure, they also looked at two open-source models by Chinese developers.)

As cutting-edge as these AI models are, the research found that they’re far from ready for the complex needs of the modern workplace. Out of all of the models put through the gauntlet, each of them failed spectacularly. OpenAI’s GPT-5.5 came in with the highest score: a passing rate of just 24 percent overall.

“Today’s agents can solve a meaningful fraction of professional tasks,” the researchers wrote. “However, when we look at the hardest tasks that require sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance.”

And as tasks became more complicated, even those meager aggregate scores fell off fast.

“On ALE’s hardest tier, every frontier agent we tested, including Fable 5, achieved a 0 percent success rate,” the presser explains.

The researchers also break down some cost considerations. The cutting edge Fable 5, they note, delivers “similar performance” to models like GPT-5.5 and Composer 2.5, “while costing roughly 4-12× more per completed task.”

Despite the horrible test results, researchers caution that the technology could still upend the job market for more AI-exposed — as plenty of corporate executives have shown us, the tech doesn’t need to work particularly well to keep workers on their back heels.

“Even if current pass rates remain relatively low, occupations dominated by routine and well-defined procedures are likely to experience disruption first, while decision-intensive roles will remain more resilient for longer,” Berkeley computer science researcher and study co-author Dawn Song told College Fix.

“The key factor,” Song added, “is not the industry itself, but the nature of the work.”

More on AI: OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin