惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

U
Unit 42
博客园 - 司徒正美
V
Visual Studio Blog
博客园 - 【当耐特】
T
Tailwind CSS Blog
美团技术团队
博客园 - 叶小钗
Jina AI
Jina AI
宝玉的分享
宝玉的分享
IT之家
IT之家
Hugging Face - Blog
Hugging Face - Blog
雷峰网
雷峰网
Stack Overflow Blog
Stack Overflow Blog
博客园_首页
人人都是产品经理
人人都是产品经理
T
The Blog of Author Tim Ferriss
P
Proofpoint News Feed
Microsoft Security Blog
Microsoft Security Blog
Y
Y Combinator Blog
GbyAI
GbyAI
大猫的无限游戏
大猫的无限游戏
Martin Fowler
Martin Fowler
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
腾讯CDC

Help Net Security

Police arrest 10 suspected members of Black Axe cybercrime gang ShinyHunters claims it stole 1.4 million records from Udemy Sevii unveils Cyber Swarm Defense Mode to stop AI-driven attacks at scale Alleged Chinese hacker extradited to US over cyberattacks targeting COVID-19 research Cequence Agent Personas bring granular control and governance to enterprise AI agents NowSecure MARI gives enterprises evidence-based visibility into third-party mobile app risk The metrics killing your SOC, and what to use instead US state privacy fines reached $3.425 billion in 2025 Canada’s first SMS blaster case leads to three arrests Linux storage management tool Stratis 3.9.0 adds online encryption and cache-less pool startup TLS Connect gives SMBs a right-sized automated tool to manage TLS certificates Aptori expands its platform with autonomous offensive testing to reduce security bottlenecks Your IAM was built for humans, AI agents don’t care The AI criminal mastermind is already hiring on gig platforms 25 open-source cybersecurity tools that don’t care about your budget Product showcase: LuLu reveals unauthorized outbound connections from Mac apps Week in review: Claude Mythos finds 271 Firefox flaws, Vercel breach Users advised to drop passwords and make room for passkeys - Help Net Security Indirect prompt injection is taking hold in the wild - Help Net Security Compromised everyday devices power Chinese cyber espionage operations - Help Net Security New Cisco firewall malware can only be killed by pulling the plug - Help Net Security Meta is overhauling how you sign in, manage settings, and protect your accounts - Help Net Security Ubuntu 26.04 LTS delivers memory-safe system tools and live patching for Arm servers - Help Net Security OpenAI’s GPT-5.5 is out with expanded cybersecurity safeguards - Help Net Security AI is speeding up nation-state cyber programs - Help Net Security A study of 1,000 Android apps finds a privacy policy logging gap - Help Net Security IT spending to hit $6.31 trillion record, thanks to AI - Help Net Security Where AI in CI/CD is working for engineering teams - Help Net Security With AI's help, North Korean hackers stumbled into a near-undetectable attack - Help Net Security Hacker with a special interest in breaching sports institutions ends behind bars - Help Net Security
AI cyber capability is speeding past earlier projections
Sinisa Marko · 2026-05-14 · via Help Net Security

AI cyber capability is improving faster than expected, with newer models surpassing earlier projections, according to the UK government’s AI Security Institute (AISI).

AI cyber capability

AISI measures AI cyber capability using “time horizon benchmarks”, which estimate how long AI systems can complete cybersecurity tasks autonomously compared to human experts.

“In February 2026, we estimated that frontier models’ 80%-reliability cyber time horizon had doubled every 4.7 months since reasoning models emerged in late 2024, given a 2.5M token limit. This was around half our November 2025 doubling time estimate, which was 8 months for both 50% and 80% reliability,” AISI wrote in a blog post.

“Claude Mythos Preview and GPT-5.5 have since significantly outperformed this trend,” researchers added.

According to the institute, it remains unclear whether this represents “an isolated break from existing rates of progress or part of a new, faster trend.”

Researchers also said the latest frontier models are beginning to exceed the limits of the current cyber evaluation framework.

Claude Mythos Preview and GPT-5.5 achieved near-100% success rates on the longest tasks in the limited cyber test suite, even with a 2.5 million token limit applied to each task.

The institute noted that the benchmark becomes harder to measure once models consistently complete the most difficult tasks.

Removing the token cap would push success rates high enough that “time horizon” estimates could no longer be calculated reliably, researchers added.

“No single benchmark result should be read as a precise measure of AI capability,” they cautioned.

AISI noted that its cyber capability estimates are consistent with findings from METR, a nonprofit research group tracking AI performance on software engineering tasks. According to METR, AI software-engineering capability has been doubling roughly every 4.2 months since late 2024.

Researchers also evaluated frontier models in simulated enterprise attack environments known as cyber ranges. The tests measure whether AI systems can carry out longer multi-step intrusion operations after gaining initial access to a target network.

In the latest testing, Claude Mythos Preview became the first model to complete the two evaluated cyber ranges. The model solved “The Last Ones,” a 32-step simulated corporate network attack, in 6 out of 10 attempts and the previously unsolved “Cooling Tower,” a 7-step industrial control system attack, in 3 out of 10 attempts. GPT-5.5 completed “The Last Ones” in 3 out of 10 attempts.

“Frontier AI’s autonomous cyber and software capability is advancing quickly: the length of cyber tasks that frontier models can complete autonomously has doubled on the order of months, not years. What this evidence does not tell us is how the rate of progress will evolve, when AI will reach specific capability thresholds, or how these capabilities will perform against defended enterprise systems,” AISI concluded.