惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Schneier on Security
Schneier on Security
N
Netflix TechBlog - Medium
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
博客园_首页
S
SegmentFault 最新的问题
H
Help Net Security
P
Proofpoint News Feed
云风的 BLOG
云风的 BLOG
T
The Blog of Author Tim Ferriss
量子位
GbyAI
GbyAI
M
MIT News - Artificial intelligence
Recorded Future
Recorded Future
P
Privacy & Cybersecurity Law Blog
B
Blog
月光博客
月光博客
博客园 - 聂微东
Vercel News
Vercel News
罗磊的独立博客
腾讯CDC
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
A
Arctic Wolf
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
Blog — PlanetScale
Blog — PlanetScale
L
Lohrmann on Cybersecurity
I
Intezer
小众软件
小众软件
T
The Exploit Database - CXSecurity.com
Jina AI
Jina AI
C
Check Point Blog
AWS News Blog
AWS News Blog
C
Cisco Blogs
Martin Fowler
Martin Fowler
The Last Watchdog
The Last Watchdog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
宝玉的分享
宝玉的分享
S
Security Affairs
大猫的无限游戏
大猫的无限游戏
N
News and Events Feed by Topic
雷峰网
雷峰网
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
H
Hacker News: Front Page
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
F
Full Disclosure
P
Proofpoint News Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Microsoft Security Blog
Microsoft Security Blog

Futurism

A Chinese AI Model Just Shot to Number One on the Charts, Sending Shockwaves Through the American Tech Industry Tech Bros Puzzled by Why AI Hasn't "Massively Disrupted" Books Yet AI Companies Are Learning an Ironic Lesson as the People They Pay to Improve Their Chatbots Are Just Feeding AI Slop Into Them Sports Journalists Asked Microsoft’s Copilot to Predict World Cup Matches, and the Results May Surprise You Researchers Put AI Models in Charge of Analyzing Sports, and They Choked Spectacularly Fans Aghast as New York Jets Say They’re Switching to AI Companies That Adopted AI Agents Alarmed to Discover They’re Botching Incredibly Important Tasks Hackers Find That Inaudible Sounds Hidden in Podcasts or Random Videos Can Hijack Your AI Voice Chatbot Democrats’ 2024 Election Autopsy Shows Signs of Sloppy AI Generation Being a Crappy Boss to AI Chatbots Pushes Them Toward Spouting Marxist Rhetoric and Organizing With Their Compatriots, Researchers Find Amazon Employees Forced to Hit Quotas on AI Use, Immediately Start Using it for Everything Except Work These Smart Glasses That Show Captions of What Everyone’s Saying Without a Creepy Spy Camera Actually Seem Pretty Awesome AI Appears to Be Trapping Certain Job Applicants in a Limbo Where They Never Get an Interview for “Reasons” That Are Completely Unfair The AI Industry Is Secretly Powered by Homeless People ChatGPT Is Saying VWeird Things in Chinese America Trembles as Transportation Secretary Announces Plans for Air Traffic Controllers to Lean on AI Tools Today Is the Day Anthropic Promised That Fully Autonomous Employees Would Be Tearing Through the Business World China Is Starting to Pull Ahead of US in AI Race Berklee College of Music Students Furious That It’s Offering an AI “Songwriting” Class Usually, Young People Embrace New Technology. Gen Z’s Attitude Toward AI Should Worry the Entire Tech Industry Sam Altman’s Coworkers Say He Can Barely Code and Misunderstands Basic Machine Learning Concepts
Frontier AI Is Faceplanting at Real-World Workplace Tasks
Joe Wilkins · 2026-07-22 · via Futurism

A stylized photo illustration of a robotic hand touching a computer keyboard.

Illustration by Tag Hartman-Simkins / Futurism. Source: Shutterstock

Sign up to see the future, today

Can’t-miss innovations from the bleeding edge of science and tech

To date, AI industry spending has topped $1.6 trillion, and shows no sign of slowing anytime soon.

So what do we actually have to show for it? Historically, it’s been a whole lot of nothing: as numerous studies have shown us, tools like AI chatbots and autonomous agents have been ineffective at completing real world tasks in a competent way.

The tech industry insists that’s all about to change within the next few years, as AI’s capabilities grow by leaps and bounds, enabling economic growth the likes of which the world has never seen. But is it really?

Not necessarily. A new study out of the University of California Berkeley’s Center for Responsible, Decentralized Intelligence — flagged by the College Fix — shows that frontier AI tools of all makes and models are still incapable of completing the vast majority of workplace tasks at an acceptable level, throwing a major wrench in the tech industry’s assertions that the AI revolution is imminent.

To come to that conclusion, the UC researchers designed a rigorous assessment they call the “Agents’ Last Exam,” developed to test “job-readiness” across numerous state-of-the-art AI models. Basically, the ALE — an impish riff on “Humanity’s Last Exam” — is designed to put an AI system through its paces, covering “more than 1,500 expert-sourced tasks spanning 55 occupations,” the researchers wrote in apress release.

Those test spans the typical line-up of AI-exposed jobs like software engineering and graphic design, but also a substantial number of jobs whose fates remain less certain, such as maritime engineering, agriculture, audio production, and public health operations.

Using the ALE benchmark, researchers took a hard look at advanced “closed” models — proprietary AI systems developed by private companies — like Anthropic’s Fable 5, OpenAI’s GPT-5.5, Cursor’s Composer 2.5, and Google’s Gemini 3.1 Pro. (For good measure, they also looked at two open-source models by Chinese developers.)

As cutting-edge as these AI models are, the research found that they’re far from ready for the complex needs of the modern workplace. Out of all of the models put through the gauntlet, each of them failed spectacularly. OpenAI’s GPT-5.5 came in with the highest score: a passing rate of just 24 percent overall.

“Today’s agents can solve a meaningful fraction of professional tasks,” the researchers wrote. “However, when we look at the hardest tasks that require sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance.”

And as tasks became more complicated, even those meager aggregate scores fell off fast.

“On ALE’s hardest tier, every frontier agent we tested, including Fable 5, achieved a 0 percent success rate,” the presser explains.

The researchers also break down some cost considerations. The cutting edge Fable 5, they note, delivers “similar performance” to models like GPT-5.5 and Composer 2.5, “while costing roughly 4-12× more per completed task.”

Despite the horrible test results, researchers caution that the technology could still upend the job market for more AI-exposed — as plenty of corporate executives have shown us, the tech doesn’t need to work particularly well to keep workers on their back heels.

“Even if current pass rates remain relatively low, occupations dominated by routine and well-defined procedures are likely to experience disruption first, while decision-intensive roles will remain more resilient for longer,” Berkeley computer science researcher and study co-author Dawn Song told College Fix.

“The key factor,” Song added, “is not the industry itself, but the nature of the work.”

More on AI: OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin