惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
爱范儿
爱范儿
Y
Y Combinator Blog
T
Tor Project blog
V
Visual Studio Blog
U
Unit 42
B
Blog RSS Feed
博客园 - 叶小钗
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
阮一峰的网络日志
阮一峰的网络日志
T
Tailwind CSS Blog
G
Google Developers Blog
I
InfoQ
Stack Overflow Blog
Stack Overflow Blog
IT之家
IT之家
Microsoft Azure Blog
Microsoft Azure Blog
T
The Blog of Author Tim Ferriss
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
The Cloudflare Blog
Google DeepMind News
Google DeepMind News
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
H
Hackread – Cybersecurity News, Data Breaches, AI and More
F
Fortinet All Blogs
人人都是产品经理
人人都是产品经理
Apple Machine Learning Research
Apple Machine Learning Research
The GitHub Blog
The GitHub Blog
Recorded Future
Recorded Future
博客园_首页
罗磊的独立博客
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
量子位
P
Proofpoint News Feed
Jina AI
Jina AI
博客园 - 【当耐特】
S
Security @ Cisco Blogs
I
Intezer
MyScale Blog
MyScale Blog
Simon Willison's Weblog
Simon Willison's Weblog
P
Privacy & Cybersecurity Law Blog
腾讯CDC
T
Tenable Blog
A
Arctic Wolf
T
Threat Research - Cisco Blogs
S
Securelist
Know Your Adversary
Know Your Adversary
Spread Privacy
Spread Privacy
C
Check Point Blog
NISL@THU
NISL@THU
Microsoft Security Blog
Microsoft Security Blog
V
Vulnerabilities – Threatpost

METR

Metrics of Agent Ability The Economics of Recursive Self-Improvement Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT Because 8 ≈ e², Anthropic's researcher uplift is plausibly >2x Summary of METR's predeployment evaluation of GPT-5.6 Sol Frontier AI Safety Policies Frontier Risk Report (February to March 2026) 前沿 AI 风险报告(2026 年 2–3 月) Informe de riesgos de la IA de frontera (febrero–marzo de 2026) Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity Task Substitution and Uplift Review of the "Risks from automated R&D" section in the Anthropic Risk Report (February 2026) Evidence on AI R&D Progress from NanoGPT MirrorCode: Evidence that AI can already do some weeks-long coding tasks Fine-tuning experiments on CoT controllability Red-Teaming Anthropic's Internal Agent Monitoring Systems Impact of modelling assumptions on time horizon results We spent 2 hours working in the future Review of the Anthropic Sabotage Risk Report: Claude Opus 4.6 Many SWE-bench-Passing PRs Would Not Be Merged into Main Observations from two CLI game reimplementation runs with Opus 4.6 We are Changing our Developer Productivity Experiment Design How We Protect Confidential Information Analyzing coding agent transcripts to upper bound productivity gains from AI agents Measuring Time Horizon using Claude Code and Codex A simpler AI timelines model predicts 99% AI R&D automation in ~2032 Frontier AI safety regulations: A reference for lab staff 前沿 AI 安全法规:AI 公司员工参考指南 Regulación de seguridad de IA de frontera: una referencia para el personal de laboratorios Time Horizon 1.1 Clarifying limitations of time horizon Early work on monitorability evaluations Common Elements of Frontier AI Safety Policies (December 2025 Update) Details about METR's evaluation of OpenAI GPT-5.1-Codex-Max Review of the Anthropic Summer 2025 Pilot Sabotage Risk Report Summary of our gpt-oss methodology review MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity Early Results on Monitorability in QA Settings Claude, GPT, and Gemini All Struggle to Evade Monitors Forecasting the Impacts of AI R&D Acceleration: Results of a Pilot Study Research Update: Algorithmic vs. Holistic Evaluation Notes on Scientific Communication at METR CoT May Be Highly Informative Despite “Unfaithfulness” Details about METR's evaluation of OpenAI GPT-5 How Does Time Horizon Vary Across Domains? Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity What should companies share about risks from frontier AI models? Details about METR's preliminary evaluation of DeepSeek and Qwen models Recent Frontier Models Are Reward Hacking Details about METR's preliminary evaluation of OpenAI's o3 and o4-mini Details about METR's preliminary evaluation of Claude 3.7 HCAST: Human-Calibrated Autonomy Software Tasks Measuring AI Ability to Complete Long Tasks Response to OSTP on AI Action Plan Why it’s good for AI reasoning to be legible and faithful 为什么 AI 推理应当可读,并如实反映模型的实际决策过程 Por qué conviene que el razonamiento de la IA sea comprensible y fiel Details about METR's preliminary evaluation of DeepSeek-R1 METR’s GPT-4.5 pre-deployment evaluations Measuring Automated Kernel Engineering Details about METR's preliminary evaluation of DeepSeek-V3 An update on our preliminary evaluations of Claude 3.5 Sonnet and o1 AI models can be dangerous before public deployment Evaluating frontier AI R&D capabilities of language model agents against human experts The Rogue Replication Threat Model Response to Bureau of Industry and Security’s proposed AI reporting requirements New Support Through The Audacious Project Details about METR's preliminary evaluation of OpenAI o1-preview Response to U.S. AISI Draft “Managing Misuse Risk for Dual-Use Foundation Models” Vivaria Details about METR's preliminary evaluation of GPT-4o An update on our general capability evaluations Response to NIST Draft Generative AI Profile ML Engineers Needed for New AI R&D Evals Project Emma Abele is METR’s new Executive Director Autonomy Evaluation Resources Example autonomy evaluation protocol Guidelines for capability elicitation Measuring the impact of post-training enhancements GitHub - METR/public-tasks Portable Evaluation Tasks via the METR Task Standard 2023 Year In Review Bounty: Diverse hard tasks for LLM agents ARC Evals is now METR Responsible Scaling Policies (RSPs) 负责任扩展政策(RSP) Políticas de escalamiento responsable (RSP) ARC Evals is spinning out from ARC New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks Response to RfC on AI Accountability Policy Update on ARC's recent eval efforts
Five lessons from having helped run an AI-Biology RCT
METR · 2026-02-19 · via METR

Evidence-based AI policy is important but hard. We need more in-depth studies – which often don’t fit into commercial release cycles.

NOTE: This post reflects my personal meta takeaways about the role of Randomized Controlled Trials (RCTs) in AI safety testing. If you have not yet read the Active Site RCT study itself, consider doing so first: see the main results and forecasts.

In early 2025, AI systems began outperforming biology experts on biology benchmarks – OpenAI’s o3 outperformed 94% of virology experts on troubleshooting questions in their own specialties. However, it remained unclear how much this translated to real-world novice “uplift”: Could a novice actually use AI to perform wet-lab tasks they could not otherwise perform?

Over the summer, I tested this question directly with Active Site (formerly called Panoplia Laboratories). We recruited 153 novices and randomly divided them into an LLM group and an Internet-only group. Over 8 weeks, participants performed fundamental wet-lab tasks involved in molecular biology workflows like reconstructing a virus from a genetic sequence.

We found that, while AI showed signs of helpfulness at individual steps, it did not produce a significant effect on end-to-end success across the three core tasks together – a result that surprised many experts.

The result provided a mid-2025 snapshot of how well AIs assist novices at molecular biology. I think there are at least two reasons why this result is very informative:

  1. It surprised most domain experts and superforecasters, causing several people to revise their current risk estimates somewhat downwards. That’s helpful for decision making.
  2. It helped establish a baseline for novice ability – previously a major point of disagreement. Now, if future RCTs find larger AI uplift – say, 15 percentage points – we will better understand the implications for biosecurity and policy decisions.

Charts of how much LLMs helped

In this post, I’ll share five lessons for evidence-based AI policy from running this RCT.

  1. Rigorous long-term studies don’t fit hectic commercial release schedules. We need an additional pipeline of RCTs to validate benchmarks, ideally run every six months
  2. The main barrier to more RCTs is talent—especially excellent ops—not cost. We should pool efforts to build dedicated teams. I hope Active Site can find great hires.
  3. Many critical threat models will require RCTs that will be substantially harder to design and execute. We should begin piloting new study designs for expert uplift now.
  4. RCTs are informative but have their own caveats. We don’t yet know if RCTs over- or underestimate ‘real-world’ AI-biology effects – and future studies should dig into this.
  5. AI firms should develop safeguards before RCTs find there is an urgent need to deploy them. Gaining experience is good; as is thinking about what results trigger this.

Lesson 1: Rigorous long-term studies don’t fit hectic commercial release schedules

A fundamental timing mismatch exists. Rigorous RCTs can require many weeks to run, while AI companies typically have days or few weeks between finalizing model training and planning deployment. Consider that, in the time we ran our RCT, we saw four new frontier models launch!

Until recently, companies could rely on biology benchmarks AIs performing worse than experts to help rule out risks. Since AIs now outperform experts on most of these benchmarks, benchmarks alone are insufficient for risk assessment. That means we now need rigorous RCTs—and so this mismatch in timing is a real issue.

I’m sympathetic to the view that running an RCT for every model release is infeasible – and testing multiple models together is more efficient. Instead, I propose an additional pipeline of in-depth RCTs – perhaps twice annually and more over time1 – whose results can help validate the benchmark-based testing that can still be run for each model between RCTs.

This proposal only applies where AI release decisions are, at least to a meaningful degree, ‘reversible’. If a closed-weight model is released and a subsequent RCT finds its safeguards need strengthening, those safeguards can still be tightened — operating at an elevated risk level for a few months may be tolerable if the likelihood of misuse is low. This assumption breaks down for AI releases that cannot be rolled back, or where the stakes between RCTs are intolerably high, or when the gap to the next RCT is too long.

Diagram of proposal

Lesson 2: The main barrier to more RCTs is talent—especially excellent ops—not cost

If the Active Site RCT is so informative, why didn’t it happen earlier? I think it should have. Having now helped set one up, I understand the obstacles better.

Setting up an RCT is time-intensive, especially when it is novel: designing and piloting laboratory tasks, securing IRB and ethics approval, recruiting 153 participants, and the study itself lasted several weeks. Although some AI-biology pilots happened earlier [1,2], a large-scale RCT required a team of 4 people’s dedicated focus for over a year.

I think these timelines can be shortened without sacrificing rigor, now that we have learned a lot about how to run these studies. But new challenges and applications will come up. How do we get novices to learn how to elicit new AI models? How do we go through hundreds of pages to notice where LLMs gave wrong advice? Answering questions like these and turning this into a well-oiled machine will still demand an excellent team.

Cost is less of a barrier than people might assume. The Active Site RCT cost $1.9M, which is comparable to leading bio benchmarks like VCT (est. $0.4M) and LAB-Bench ($2.9M).2 RCTs can also be funded by pooled AI company contributions or be run by AI Safety Institutes. Most costs are fixed; once participants have been recruited, adding another AI model for them to use is straightforward.

The challenge is building centers of excellence that are up for the task. I am very excited about Active Site and expanding RCT that can also have applications to other threat models and questions in AI-biology. They are currently hiring!

Active Site job listings

Lesson 3: Many critical threat models will require RCTs that will be substantially harder to design and execute

Part of what motivates me to encourage more talent into this field is that the Active Site RCT only directly addressed a single threat model:3 Could a novice actually use AI to perform wet-lab tasks they could not otherwise perform?

Many future studies – such as if AIs help teams of experts at novel CBRN designs – will be substantially harder to implement. Recruiting expert teams is harder than recruiting novices; defining success metrics for novel research is harder than assessing performance on established protocols; and studying dual-use research demands more responsible oversight.

I think these challenges are surmountable – and that such studies are both feasible and essential for studying low-probability high-consequence misuse risks. But we must begin now, especially if we are to iterate on our methodologies. Again, I see Active Site as well-positioned to lead this work, and I encourage others to contribute to its growth. I am also excited for AISIs to be empowered to facilitate more public-private partnerships for national security.

Lesson 4: RCTs are informative but have their own caveats

While I see our RCT as very informative, it would be a mistake to see it as replacing all other safety testing. As noted, it only directly covers one threat model. More fundamentally: every evaluation is one source of evidence to a broader picture of risk. RCT results may well be “first among equals”, but they also require careful interpretation.

There are many hypotheses to explain why an RCT and a benchmark may not line up. For the purpose of safety testing, one important hypothesis is that RCT test conditions might underestimate the model’s ‘real-world’ capabilities. RCTs like ours measure the combined effect of a model’s maximum capability and participants’ ability to use it. Many participants may be slow to learn how to best make use of AIs.4 Similarly, it could take months for third parties to develop tools and techniques built on top of the AI model to increase its usefulness.

Framework for uplift RCTs

To clarify, other explanations also exist, and the Active Site still suggests current AI-biology risk is lower than most experts had thought from looking at the benchmark results alone. But, if RCTs do underestimate capabilities relative to real-world scenarios they are being used to inform, they cannot be the sole basis for risk assessment. Policymakers should still look at the whole portfolio of evidence available to them.

Two things help: (1) RCT methodologies must be transparent so policymakers can accurately interpret results; (2) pre-registering predictions is a very useful reference point, as it was for us.

Lesson 5: AI firms should develop safeguards proactively, before RCTs identify the need for them

One thing that running this study with a goal of informing AI policy helped me clarify is that AI companies face two questions when anticipating risks: (1) When should they start developing safeguards to be ready in time? (2) When do safeguards now need to be implemented reliably?

Over the last year, we have also learned that safeguards take time to develop. Recent CBRN classifiers likely required months of development by teams of software engineers. By the time an RCT identifies a risk threshold being crossed, battle-tested safeguards should be ready to deploy – not only then scrambled upon (and potentially risking deployment delays altogether).

I suspect proactive investment in safeguards will pay off: early deployment lets developers refine them through stress-testing and customer feedback. This can help reduce costs and their ‘false positive’ rate – i.e. AI models refusing to answer scientific questions, which will only become more important as AIs will hopefully become more useful for science.

Diagram of safeguards stages

Conclusion

In theory, almost everyone agrees AI policy should be evidence-based [1,2,3,4,5]. But, in practice, science is often messy and full of caveats – things that sit uneasily with policymakers’ demands for clean “yes/no” answers.

In 2025, as many AI benchmarks saturated and could no longer give a clear “no” on risk, we saw some of the first large RCTs examining AI effects in more realistic settings – from AI R\&D to political persuasion. Many of these RCTs revealed messier pictures of AI capabilities and new caveats of their own.

I remain excited about evidence-based AI policy. Recent RCTs show how much we still have to learn. We’ve moved from a world where we had benchmarks that showed a clear “no”, to one where some hard calls about emerging AI risks will need to be made. I think the bar for establishing that evidence will be high. Again, we should spend less time proving that today’s AIs are safe and more time figuring out how to tell if tomorrow’s AIs are dangerous.

Diagram of difference between theory and practice in bio evals