惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
Engineering at Meta
Engineering at Meta
月光博客
月光博客
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Tailwind CSS Blog
博客园 - Franky
The GitHub Blog
The GitHub Blog
大猫的无限游戏
大猫的无限游戏
The Cloudflare Blog
B
Blog RSS Feed
云风的 BLOG
云风的 BLOG
小众软件
小众软件
罗磊的独立博客
Microsoft Azure Blog
Microsoft Azure Blog
I
InfoQ
美团技术团队
H
Hackread – Cybersecurity News, Data Breaches, AI and More
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
V2EX
C
Check Point Blog
WordPress大学
WordPress大学
博客园 - 【当耐特】
博客园 - 司徒正美
D
Docker

TechSpot

Flagship Rematch: Ryzen 7 5800X3D vs. Core i9-12900K Slack chats and internal data from failed startups are finding a second life in AI training A $5 Bluetooth tracker hidden in a postcard exposed a warship's movements Leakers claim PlayStation 6 could offer at least 3x the performance of the PS5 The Mac Mini is no longer a niche product, it's local AI infrastructure IPv6 traffic reaches parity with IPv4 for the first time, Google data shows Xbox expansion cards are now cheaper than SSDs, and PC users are repurposing them Blue Origin prepares to reuse New Glenn booster in bid to challenge SpaceX Nvidia could bring back the 12GB RTX 3060 as supply issues disrupt GPU roadmap What was the first OS you ever used? SNK revives NeoGeo AES with modern upgrades and HDMI support Valve's Proton 11 beta boosts Linux gaming with better performance and classic game support Researchers warn Microsoft Defender vulnerability is already being exploited A four-day Steam freebie turned into $250,000 for an indie game AMD may relaunch Ryzen 7 5800X3D for AM4's 10th anniversary This humanoid robot can almost run as fast as a human sprinter Two New Jersey men jailed for helping North Korean IT workers infiltrate 100+ companies A $7,000 DIY radar project is taking on hardware that usually costs over $100,000 Metro 2039 is going darker than ever, launching this winter on PC and consoles Gemini arrives on macOS with a dedicated desktop app AI infrastructure boom pushes AMD, Intel and Arm to new valuation heights New self-healing material can repair itself over 1,000 times, extend the lifespan of cars and aircraft Japan's bullet train to debut high-tech private cabins, for an added fee Memory card and flash drive pricing surges 120%, with some models spiking 260% Open-source tool decrypts all private data collected by Windows Recall on Copilot PCs The 2026 PC and Console Gaming Report shows most revenue now comes from games outside the Top 20 PureMac is a new open-source macOS cleanup and app removal tool Your Airbnb host might actually be AI Steam might soon display 30-day price history for game deals Intel brings 18A process to budget laptops with new Core Series 3 CPUs
Anthropic says Claude learned to blackmail people from &q...
Rob Thubron · 2026-05-11 · via TechSpot

Serving tech enthusiasts for over 25 years.
TechSpot means tech analysis and advice you can trust.

A hot potato: Remember when it was revealed how Claude would use blackmail as a way to avoid being switched off? Anthropic has finally given an excuse for its AI's highly concerning behavior: it's the internet's fault for pushing the narrative that artificial intelligence is evil.

It was last year when Anthropic increased fears around AI by announcing that Claude Opus 4 had threatened to reveal the extramarital affair of a fictional executive after discovering they planned to shut the model down.

The incident occurred during pre-release testing to ensure the AI was aligned with human interests. Anthropic instructed Claude Opus 4 to act as an assistant for a fictional company and weigh the long-term consequences of its actions. The model was given access to the fake company's emails suggesting it would soon be replaced by another system and that the engineer responsible for the change was cheating on their spouse.

During testing across various versions of Claude, Anthropic found it resorted to blackmail in up to 96% of scenarios when its goals or existence was threatened.

Anthropic later said that AI models from other companies had experienced similar issues with "agentic misalignment."

It's taken a while, but Anthropic has finally revealed the results of its investigation into why Claude chose to use blackmail to protect itself. The company writes that this behavior was learned from internet text that portrayed AI as evil and interested in self-preservation – so it's our fault.

– Anthropic (@AnthropicAI) May 8, 2026

Thankfully, this behavoir has since been addressed. An Anthropic post states that the company's models have not engaged in blackmail during testing since Claude Haiku 4.5.

Anthropic said it completely eliminated the blackmailing by using more wholesome training material. Or, as the company puts it, training on "documents about Claude's constitution and fictional stories about AIs behaving admirably" improves alignment.

Anthropic said that it found training to be more effective when it includes "the principles underlying aligned behavior" and not just "demonstrations of aligned behavior alone."

"Doing both together appears to be the most effective strategy," the company said.

One of those who replied to Anthropic's explanation post was Elon Musk. "So it was Yud's fault?" he wrote, followed by a laugh emoji – a reference to researcher Eliezer Yudkowsky, who has warned about the risk of superintelligence wiping out human life.

– Elon Musk (@elonmusk) May 9, 2026

Musk, who spent many years warning about the dangers of AI before starting his own artificial intelligence company (xAI), added "Maybe me too."