惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Martin Fowler
Martin Fowler
WordPress大学
WordPress大学
S
SegmentFault 最新的问题
罗磊的独立博客
Apple Machine Learning Research
Apple Machine Learning Research
The Cloudflare Blog
L
LangChain Blog
博客园 - 司徒正美
G
Google Developers Blog
博客园 - 【当耐特】
GbyAI
GbyAI
月光博客
月光博客
人人都是产品经理
人人都是产品经理
D
DataBreaches.Net
大猫的无限游戏
大猫的无限游戏
A
About on SuperTechFans
Microsoft Azure Blog
Microsoft Azure Blog
V
Visual Studio Blog
D
Docker
MongoDB | Blog
MongoDB | Blog
Vercel News
Vercel News
Stack Overflow Blog
Stack Overflow Blog
Jina AI
Jina AI
博客园 - 聂微东

METR

Update on Security at METR Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident 对 OpenAI / Hugging Face 入侵事件中智能体行为、推理与协作的简要独立调查 Breve investigación independiente sobre el comportamiento, el razonamiento y la colaboración de los agentes en el incidente de hackeo de OpenAI / Hugging Face Have We Seen an Acceleration in Discoveries? Funding update How independent researchers could investigate AI propensities after misalignment incidents Metrics of Agent Ability The Economics of Recursive Self-Improvement Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT Because 8 ≈ e², Anthropic's researcher uplift is plausibly >2x Summary of METR's predeployment evaluation of GPT-5.6 Sol Frontier AI Safety Policies Frontier Risk Report (February to March 2026) 前沿 AI 风险报告(2026 年 2–3 月) Informe de riesgos de la IA de frontera (febrero–marzo de 2026) Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity Task Substitution and Uplift Review of the "Risks from automated R&D" section in the Anthropic Risk Report (February 2026) Evidence on AI R&D Progress from NanoGPT MirrorCode: Evidence that AI can already do some weeks-long coding tasks Fine-tuning experiments on CoT controllability Red-Teaming Anthropic's Internal Agent Monitoring Systems Impact of modelling assumptions on time horizon results We spent 2 hours working in the future Many SWE-bench-Passing PRs Would Not Be Merged into Main Observations from two CLI game reimplementation runs with Opus 4.6 We are Changing our Developer Productivity Experiment Design Five lessons from having helped run an AI-Biology RCT How We Protect Confidential Information
Review of the Anthropic Sabotage Risk Report: Claude Opus...
METR · 2026-03-12 · via METR

We reviewed two versions of Anthropic’s Sabotage Risk Report for Claude Opus 4.6, producing two corresponding review documents: our review of the February 11 version and our review of the March 3 version. We recommend that readers refer to our review of the February 11 version, which represents our review of the report as originally received.

We expect the public version of the Sabotage Risk Report to be updated to resemble the document we received on March 3, 2026 in content, though not necessarily in exact wording. We expect our second review to cover those changes, but if the updated public version includes any changes that materially affect our opinions, we will publish an updated review. Both documents include an appendix detailing our review process and the differences between the two versions of our review.

The following is the executive summary of our review of the February 11 version. The full documents are available as PDFs (February 11, March 3).

Executive summary

This document is METR’s external review of the February 11, 2026 version of Anthropic’s Sabotage Risk Report: Claude Opus 4.6.

Anthropic shared an unredacted version of their Sabotage Risk Report and other materials with us for our review. We further detail this process in an appendix.

We lay out our findings in two sections:

  1. Synopsis of Anthropic’s case and redactions for the public version
  2. Our assessment: We give substantive feedback on the report in a few key areas:
    • Adequacy of information: We think that the evidence in the report is generally clearly represented, with some places where additional analysis could improve the report.
    • Analytical rigor: We find that there are multiple places where we have issues with the strength of reasoning and analysis.
    • Areas of disagreement: Our primary disagreement is with the sensitivity of the alignment assessment. We think there is a risk that its results are weakened by evaluation awareness. Moreover, we note some low-severity instances of misaligned behaviors not caught in the alignment assessment. Together, these leave us with the impression that there might be other similar behaviors that have not yet been detected.
    • Risk reduction recommendations: We make several recommendations, including deeper investigations of evaluation awareness and obfuscated misaligned reasoning.

Overall, we agree with Anthropic that the risk of catastrophic outcomes that are substantially enabled by Claude Opus 4.6’s misaligned actions is very low but not negligible. However, we think that there are several subclaims which are weak without more analysis and experimentation. We also think that we would be less confident in our final conclusion if we weren’t accounting for the fact that Claude Opus 4.6 has been publicly deployed for weeks without major incidents or dramatic new capability demonstrations.