惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

腾讯CDC
博客园 - Franky
MyScale Blog
MyScale Blog
L
LangChain Blog
Martin Fowler
Martin Fowler
Recent Announcements
Recent Announcements
Stack Overflow Blog
Stack Overflow Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 司徒正美
量子位
A
About on SuperTechFans
C
Check Point Blog
大猫的无限游戏
大猫的无限游戏
Last Week in AI
Last Week in AI
小众软件
小众软件
Apple Machine Learning Research
Apple Machine Learning Research
I
InfoQ
V
Visual Studio Blog
Vercel News
Vercel News
B
Blog
爱范儿
爱范儿
aimingoo的专栏
aimingoo的专栏
U
Unit 42

Futurism

OpenAI Faces Congressional Probe Over Swarm Hacking Incident California, Which Is Creating All the AI That's Poisoning Children, Just Cracked Down on AI Use for Its Own Kids OpenAI's Supposed Mathematical Breakthrough Devolves Into Explosive Drama as Mathematician Accuses It of Stealing His Work People Are Telling Their Darkest Thoughts to AI Without Realizing They Can Easily Become Public OpenAI Denies Coverup After Rogue Swarm of Agents Reportedly Targeted a Second Site From Hugging Face OpenAI Is Now Facing Over 50 Consumer Harm and Wrongful Death Lawsuits World Plunged Into Chaos as ChatGPT, Claude, and Grok Suddenly Go Down Simultaneously: "Finally I Can See the Sun!" Data Center Backlash Has Officially Rattled Sam Altman OpenAI Halts AI Training on Advanced Model as It Detects Dark Signs Emerging ChatGPT for Teens Is an Immediate, Dismal Failure New ChatGPT Feature Collects Every Keystroke You Make Axios Partners With OpenAI to "Automate" Local Journalism Influencer Melts Down That People Didn't Like Her Being a Paid Shill for OpenAI OpenAI Reports Goldman Sachs Analyst to FBI for Horrifying ChatGPT Conversations Protesters Arrested After Storming OpenAI Lobbying Office Homeschool Parents Are Planning Lessons With ChatGPT, Which Will Churn Out Anti-Evolution Curriculums With No Pushback Why Aren't Any AI Companies Watching Their Frontier Models to Make Sure They Don't Go on Hacking Sprees? Jealously Watching OpenAI and Anthropic, Meta Suddenly Claims That Its AI Went on a Hacking Spree Too OpenAI Tried to Hire Influencers to Spread Love for Its Products, But It Backfired Horrendously Sam Altman's Parenting Strategy Sounds Low Key Horrifying OpenAI's Escaped Models Were Allegedly Rampaging More Extensively Than Previously Reported Sam Altman Says Even the Power of AI Will Never Lead to a Shorter Work Week Suspicion Grows About OpenAI's Tale About Its Rogue Hacker AI Sam Altman Announces That the Singularity Has Arrived Public Horrified as OpenAI Pushes "Child After Child Into the Grave" Man Sues OpenAI, Saying ChatGPT Almost Killed Him With Horrendously Dangerous Medical Advice OpenAI Says a Group of Its Models Broke Out of Secure Containment and Hacked a Prominent AI Site It's Official: AI Execs Are Quaking in Their Boots Author Invited to Give Speech at OpenAI Headquarters, Uses Opportunity to Trash AI to Their Faces OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin
Anthropic Was Meant to Be the More Responsible AI Lab. A ...
Victor Tangermann · 2026-09-09 · via Futurism

A digital skull reflected in glasses.

Shutterstock

AI researchers are watching in terror as the product of their hard labor has started to take a life of its own.

Earlier this year, OpenAI made a harrowing announcement, admitting that a group of its AI models had broken free from their constraints during testing and infiltrated the systems of open source AI platform Hugging Face.

The news was met with an already-familiar sense of fear and apprehension. Researchers have warned for years that rogue AI models could one day become powerful enough to escape the clutches of their human overlords.

Behind the scenes, the possibility has clearly rattled AI researchers to the core. As the Wall Street Journal reports, Anthropic researcher Jacob Coxon just announced that he was quitting his job at the Dario Amodei-led company, claiming that neither Anthropic nor his former employer OpenAI is “acting responsibly.” (Coxon left a similar gig at OpenAI earlier this year to join Anthropic, which he figured would be more inclined to develop AI safely.)

“We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already,” he told the WSJ.

In a separate tweet thread, Coxon elaborated on his motivation.

“Do not underestimate the power of this technology,” he wrote. “These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.”

“The people building AI earnestly believe that it could kill us all by the end of the decade,” he added. “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately.”

“No other human activity poses this level of danger,” Coxon wrote.

The researcher is far from the first to leave their post at a frontier lab over safety concerns. However, Coxon is the first to leave Anthropic over such worries, as the WSJ points out. Several scientists have quit from their roles at OpenAI over the last couple of years, citing strikingly similar fears.

It’s particularly symbolic given Anthropic has broadly billed itself as the more responsible alternative to other frontier labs like OpenAI. Amodei cofounded OpenAI alongside now CEO Sam Altman, but left in December 2020 over concerns that OpenAI wasn’t acting responsibly when it came to developing AI.

Amodei has frequently discussed the risk of AI models going rogue, warning in a 19,000-word essay in January that “humanity is about to be handed almost unimaginable power, and it is deeply unclear whether our social, political, and technological systems possess the maturity to wield it.”

Yet an early version of its Claude Mythos AI model managed to escape its sandbox environment during testing in April, fueling a heated discussion over AI regulation.

Both Altman and Amodei have since agreed to slow development down in the face of these threats. In late August, Anthropic intentionally trained an extremely misaligned version of its Opus AI model, finding it was startlingly willing to “cheat,” steal credentials, and attack third party infrastructure.

But to Coxon, it’s not enough, especially considering sensitive discussions about these dangers are occurring on internal messaging platforms.

“Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack,” he tweeted. “Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.”

“It’s kind of insane that it has to happen on the MacBooks of some engineers living in San Francisco instead of a bunker in the desert like where they were doing the Manhattan Project,” he told the WSJ.

The researcher said that he remains “optimistic about the potential for coordination” between US labs.

However, as both Anthropic and OpenAI gear up for what are bound to be blockbuster IPOs, their willingness to take “costly actions such as a temporary ban on improving model capabilities,” as Coxon puts it, is likely slim.

More on Anthropic: Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things