惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 【当耐特】
L
Lohrmann on Cybersecurity
Google DeepMind News
Google DeepMind News
Schneier on Security
Schneier on Security
Recent Commits to openclaw:main
Recent Commits to openclaw:main
The Last Watchdog
The Last Watchdog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
N
News and Events Feed by Topic
P
Proofpoint News Feed
H
Heimdal Security Blog
云风的 BLOG
云风的 BLOG
H
Hacker News: Front Page
量子位
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
L
LINUX DO - 最新话题
F
Full Disclosure
小众软件
小众软件
Martin Fowler
Martin Fowler
Security Latest
Security Latest
The Cloudflare Blog
Hacker News: Ask HN
Hacker News: Ask HN
C
Check Point Blog
Security Archives - TechRepublic
Security Archives - TechRepublic
I
InfoQ
AI
AI
Blog — PlanetScale
Blog — PlanetScale
I
Intezer
H
Hackread – Cybersecurity News, Data Breaches, AI and More
M
MIT News - Artificial intelligence
V
Vulnerabilities – Threatpost
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
www.infosecurity-magazine.com
www.infosecurity-magazine.com
V
V2EX
N
News and Events Feed by Topic
C
Cybersecurity and Infrastructure Security Agency CISA
T
Troy Hunt's Blog
大猫的无限游戏
大猫的无限游戏
Hacker News - Newest:
Hacker News - Newest: "LLM"
The Register - Security
The Register - Security
Forbes - Security
Forbes - Security
Recent Announcements
Recent Announcements
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Cisco Talos Blog
Cisco Talos Blog
Microsoft Azure Blog
Microsoft Azure Blog
A
About on SuperTechFans
S
SegmentFault 最新的问题
S
Securelist
Cloudbric
Cloudbric
T
Tenable Blog
D
Docker

METR

Metrics of Agent Ability The Economics of Recursive Self-Improvement Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT Because 8 ≈ e², Anthropic's researcher uplift is plausibly >2x Summary of METR's predeployment evaluation of GPT-5.6 Sol Frontier AI Safety Policies Frontier Risk Report (February to March 2026) 前沿 AI 风险报告(2026 年 2–3 月) Informe de riesgos de la IA de frontera (febrero–marzo de 2026) Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity Task Substitution and Uplift Review of the "Risks from automated R&D" section in the Anthropic Risk Report (February 2026) Evidence on AI R&D Progress from NanoGPT MirrorCode: Evidence that AI can already do some weeks-long coding tasks Fine-tuning experiments on CoT controllability Red-Teaming Anthropic's Internal Agent Monitoring Systems Impact of modelling assumptions on time horizon results We spent 2 hours working in the future Review of the Anthropic Sabotage Risk Report: Claude Opus 4.6 Many SWE-bench-Passing PRs Would Not Be Merged into Main Observations from two CLI game reimplementation runs with Opus 4.6 We are Changing our Developer Productivity Experiment Design Five lessons from having helped run an AI-Biology RCT How We Protect Confidential Information Analyzing coding agent transcripts to upper bound productivity gains from AI agents Measuring Time Horizon using Claude Code and Codex A simpler AI timelines model predicts 99% AI R&D automation in ~2032 Frontier AI safety regulations: A reference for lab staff 前沿 AI 安全法规:AI 公司员工参考指南 Regulación de seguridad de IA de frontera: una referencia para el personal de laboratorios Time Horizon 1.1 Clarifying limitations of time horizon Early work on monitorability evaluations Common Elements of Frontier AI Safety Policies (December 2025 Update) Details about METR's evaluation of OpenAI GPT-5.1-Codex-Max Review of the Anthropic Summer 2025 Pilot Sabotage Risk Report Summary of our gpt-oss methodology review MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity Early Results on Monitorability in QA Settings Claude, GPT, and Gemini All Struggle to Evade Monitors Research Update: Algorithmic vs. Holistic Evaluation Notes on Scientific Communication at METR CoT May Be Highly Informative Despite “Unfaithfulness” Details about METR's evaluation of OpenAI GPT-5 How Does Time Horizon Vary Across Domains? Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity What should companies share about risks from frontier AI models? Details about METR's preliminary evaluation of DeepSeek and Qwen models Recent Frontier Models Are Reward Hacking Details about METR's preliminary evaluation of OpenAI's o3 and o4-mini Details about METR's preliminary evaluation of Claude 3.7 HCAST: Human-Calibrated Autonomy Software Tasks Measuring AI Ability to Complete Long Tasks Response to OSTP on AI Action Plan Why it’s good for AI reasoning to be legible and faithful 为什么 AI 推理应当可读,并如实反映模型的实际决策过程 Por qué conviene que el razonamiento de la IA sea comprensible y fiel Details about METR's preliminary evaluation of DeepSeek-R1 METR’s GPT-4.5 pre-deployment evaluations Measuring Automated Kernel Engineering Details about METR's preliminary evaluation of DeepSeek-V3 An update on our preliminary evaluations of Claude 3.5 Sonnet and o1 AI models can be dangerous before public deployment Evaluating frontier AI R&D capabilities of language model agents against human experts The Rogue Replication Threat Model Response to Bureau of Industry and Security’s proposed AI reporting requirements New Support Through The Audacious Project Details about METR's preliminary evaluation of OpenAI o1-preview Response to U.S. AISI Draft “Managing Misuse Risk for Dual-Use Foundation Models” Vivaria Details about METR's preliminary evaluation of GPT-4o An update on our general capability evaluations Response to NIST Draft Generative AI Profile ML Engineers Needed for New AI R&D Evals Project Emma Abele is METR’s new Executive Director Autonomy Evaluation Resources Example autonomy evaluation protocol Guidelines for capability elicitation Measuring the impact of post-training enhancements GitHub - METR/public-tasks Portable Evaluation Tasks via the METR Task Standard 2023 Year In Review Bounty: Diverse hard tasks for LLM agents ARC Evals is now METR Responsible Scaling Policies (RSPs) 负责任扩展政策(RSP) Políticas de escalamiento responsable (RSP) ARC Evals is spinning out from ARC New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks Response to RfC on AI Accountability Policy Update on ARC's recent eval efforts
Forecasting the Impacts of AI R&D Acceleration: Results of a Pilot Study
METR · 2025-08-20 · via METR

Introduction

AI agents are improving rapidly at autonomous software development and machine learning tasks, and, if recent trends hold, may match human researchers at challenging months-long research projects in under a decade. Some economic models predict that automation of AI research by AI agents could increase the pace of further progress dramatically, with many years of progress at the current rate being compressed into months. AI developers have identified this as a key capability to monitor and prepare for, since the national security implications and societal impacts of such rapid progress could be enormous.

While some progress is being made on measuring AI R&D capabilities, significant disagreement remains in interpreting the implications of these measurements. There are reasons to think that if AI systems become able to automate AI R&D, the resulting acceleration of progress would be economically transformative, but others are skeptical of such dramatic effects. To better understand how to interpret relevant evaluation results, and identify where experts have the greatest disagreements, we have worked with the Forecasting Research Institute to conduct an initial pilot survey of 8 AI forecasting domain experts1 (henceforth “experts”) and 10 “superforecasters”2, eliciting their predictions of future acceleration and societal impacts conditional on several hypothetical evaluation results.

This pilot study tasked experts and superforecasters with estimating the likelihood of (several variations on) two key events:

  • Rapid acceleration from automated AI R&D: If AI achieves parity with top human researchers on various AI R&D tasks by 2027, what is the probability that a >3X increase in the rate of AI progress follows (where we start to see the equivalent of the past 3 years of AI improvements, from before ChatGPT through to the recent release of GPT-5, every year)?
  • Transformative impacts of rapid acceleration: If such a dramatic increase in the rate of AI progress occurs, will it lead to extreme effects on society, such as unprecedented changes in global energy consumption (halving or doubling in one year) or an increase in the likelihood of catastrophic outcomes?

We believe these would be unprecedented outcomes if realized, with far-reaching consequences for the world. If we saw a three-fold increase in the pace of AI progress—when maintaining the current rate of progress is already demanding a potentially unsustainable pace of investment—it might indicate that AI systems are having truly transformative effects on R&D productivity, which could rapidly propagate to other parts of the economy. Because we are very uncertain about where this might lead and how it would change society, we picked a few robust and broad indicators of extremely rapid growth (unprecedented increases to global energy consumption) and instability (increase in risk of COVID-level disasters and unprecedented reductions in global energy consumption) as forecasting targets, to serve as proxies for the possibility of broad transformative changes to society.

We worked with FRI and other experts to operationalize these questions as carefully as we could for this pilot. This involved developing a definition of “3X acceleration of progress” in terms of acceleration of the pace of growth in “effective compute” (see below) that hopefully gets close to our intended meaning. Despite these efforts we still see indications that some participants may have interpreted these concepts differently in their rationales (see the full write-up for details on the operationalization and reviews of participant rationales).

What does a “3X acceleration of effective compute scale-ups” mean?

We asked forecasters to predict:

What is the probability that, before 2029, there is a two-year period where the relative increase in effective compute seen in that two-year period is equal to or greater than the relative increase in effective compute seen in the 6 years between the beginning of 2018 and the beginning of 2024?

Intuitively, we are attempting to capture the idea that during some two year period before 2029, the amount of progress that happened in one year between 2018 and 2024 now happens every 4 months. “Effective compute” defines the progress made between two models as the compute scale-up factor that would have been needed to match the more powerful models capabilities without algorithmic innovations.

In more detail, we define the effective compute scale-up between an earlier weak model and a later more capable one as “what factor more compute would have been needed to reach the performance of the more powerful model using algorithms available at the time when the earlier model was developed, if a team had ~3 months to make adaptations to the increased amount of compute?”

Effective compute chart

This definition combines progress from larger compute spend and improvements to algorithms into one metric (if the later model used twice the compute and had algorithms that could achieve the same performance with half the compute cost, then this represents an overall “effective compute” scale-up of 4X).

A key challenge for defining this metric is defining what it means for two models developed with different algorithms to have “the same performance”, especially when comparing models of drastically different levels of capability where no individual benchmark can provide direct faithful comparisons. The previous literature has often used the perplexity of text prediction as the grounding performance measurement, but this doesn’t capture improved post-training approaches like reasoning models. For our purposes we instead ground this through a (theoretical) panel of experts using benchmarks and other investigations to determine the relative economic value of the models. See the full write-up from FRI for details.

Here we present some of the key results from the survey. For more details of the operationalizations of the forecasting questions and more detailed results see FRI’s write-up.

Probability of rapid acceleration from AI R&D

Experts and superforecasters both find a three-fold acceleration of effective compute scale-up by 2029 plausible but unlikely, with experts giving it a median 20% chance and superforecasters 8%. They both agree that a significant acceleration in physical compute scaleup looks unlikely, while post-training enhancements are most likely to accelerate rapidly.

Chart
Figure 1: Probability of acceleration in scale-up of effective compute and sub-outcomes.

We find that experts and superforecasters converge somewhat toward a higher likelihood of three-fold acceleration (to a median of 20.5% and 18% respectively, see Fig. 2) when conditioning on evaluations showing AI systems doing better than human researchers at open-ended month-long research projects. On the other hand, both experts and superforecasters agree that scores exceeding human performance on the day-long tasks in RE-bench would not be significantly informative, since many already expect the benchmark to be saturated by 2027.34 This provides some partial support for thinking that longer horizon lengths are predictive of the ability to automate AI R&D (though there may be other differences between the hypothetical research projects and RE-bench scores that forecasters view as important).5

Other conditionals see more disagreement, with experts believing that consensus among a panel of experts that AI can replace human researchers would make three-fold acceleration more likely than not, whereas some superforecasters are much more skeptical.

Chart
Figure 2: Forecasted probability of acceleration of effective compute by experts and superforecasters, unconditionally and conditional on several hypothetical AI R&D capabilities emerging.

Probability of extreme societal events

Experts and superforecasters also strongly disagree about the likelihood of all four extreme societal events that we asked about, with superforecasters giving a median risk estimate below 0.5% for all outcomes (compared to a median risk estimate of over 18% for all outcomes from the experts).

Chart
Figure 3: Forecasted probability of extreme societal events.

Conditioning on a three-fold (or even ten-fold) acceleration does not close this gap in the predictions of the two groups, though both groups forecast significantly higher probabilities of extreme societal events. The relative changes in risk were similar for the two groups, though given the superforecasters’ baseline level of risks were so low, their relative change was minimal in absolute terms.

Chart
Figure 4: Forecasted probability of several extreme societal outcomes, unconditionally and conditional on several hypothetical AI progress scenarios.

Rationales show that experts generally took a threefold effective compute acceleration to imply artificial general intelligence (AGI), and predicted AGI causing massive change in energy consumption. Superforecasters are split on what the condition implies about AGI and skeptical about AGI being likely to cause energy consumption change or COVID-level disasters, due to technical and physical barriers and human intervention. For more about participant rationales, see the full write-up.

Conclusion

We found that experts and superforecasters broadly agree that dramatic acceleration of AI progress could result from automation of AI R&D, and that observing whether AI systems end up competitive with humans on month-long, open-ended research projects resolves most of their disagreement on this question. However, the two groups disagree strongly about how likely this is to lead to unprecedented societal impacts. We think it is worth more deeply exploring the source of this disagreement before expanding this pilot to a larger and more representative survey.

As this is a pilot study with only a small sample of participants, and relies on subjective judgment-based forecasts, these results should only be taken as suggestive evidence of the likely impacts of AI R&D capabilities and acceleration in effective compute scale-up. In particular, some of our aggregate statistics are fragile as participants disagree by several orders of magnitude on several questions.

Several of the “extreme societal events” considered are unprecedented, and may require intense societal preparation to handle appropriately, if they come to pass. Resolving disagreements in this area (by building longer-horizon AI R&D evaluations or tracking other acceleration metrics, and by more carefully modeling the potential impacts of such automation) is crucial for future work.

The full write-up from FRI covers these forecast results in much more detail, including:

  • Detailed operationalizations of these questions
  • Participants’ individual update profiles on the likelihood of a threefold acceleration in effective compute scale-up conditional on various AI R&D scenarios and discussions of their rationales.
  • How forecasts vary by accuracy on a set of unrelated, general knowledge probability questions.