惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Securelist
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
WordPress大学
WordPress大学
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
T
Tailwind CSS Blog
V
V2EX
小众软件
小众软件
博客园 - 聂微东
H
Help Net Security
阮一峰的网络日志
阮一峰的网络日志
云风的 BLOG
云风的 BLOG
Blog — PlanetScale
Blog — PlanetScale
M
MIT News - Artificial intelligence
人人都是产品经理
人人都是产品经理
F
Fortinet All Blogs
S
Schneier on Security
Martin Fowler
Martin Fowler
MyScale Blog
MyScale Blog
Vercel News
Vercel News
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Google DeepMind News
Google DeepMind News
Google Online Security Blog
Google Online Security Blog
Webroot Blog
Webroot Blog
A
Arctic Wolf
量子位
博客园 - 叶小钗
I
Intezer
C
Check Point Blog
Cloudbric
Cloudbric
IT之家
IT之家
Last Week in AI
Last Week in AI
GbyAI
GbyAI
Attack and Defense Labs
Attack and Defense Labs
T
The Blog of Author Tim Ferriss
Y
Y Combinator Blog
Jina AI
Jina AI
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
T
Threat Research - Cisco Blogs
C
CERT Recently Published Vulnerability Notes
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
Cisco Blogs
J
Java Code Geeks
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Engineering at Meta
Engineering at Meta
酷 壳 – CoolShell
酷 壳 – CoolShell
L
Lohrmann on Cybersecurity
有赞技术团队
有赞技术团队
Simon Willison's Weblog
Simon Willison's Weblog
The Register - Security
The Register - Security
T
Threatpost

Latest from TechRadar in Pro

VodafoneThree gets Ofcom approval to bring satellite connectivity to your smartphone Is this the tipping point for AI at work? New Gallup survey finds half of all US employees now use it in some way 'Every Apple user needs to know about this nasty scam': Fake warnings tell users their iCloud data will be… 'Makes it even more disappointing': Microsoft backs fossil fuel big time with $7 billion deal in race for AI… 'Maybe it’s not science fiction': Solar panels are causing rainwater to fall in one of the driest places… Maine becomes first US state to pass data centre construction ban Dozens of WordPress plugins hijacked to target thousands of sites Drone-killing laser weapons greenlit for use in US airspace – FAA and Defense Department say high-energy weapons are ‘ready to protect all air travelers from illicit drone use’ despite airspace restrictions and friendly-fire incidents 'We are currently being extorted' — crypto giant Kraken says it is facing extortion attack, here's… I tried 7 free MTD software – now I've ranked my top picks as a freelancer Jackery McGraw Hill becomes latest to see its Salesforce data hacked Looking for a new PC? Now might be great time to upgrade, as Gartner figures claim shipments are rising — while… The new engineering playbook: how AI design copilots are reshaping product development Farewell Surface Hub — Microsoft kills off its super-sized touchscreen displays, but you might still be able to get one if you act fast 'We have no interest in patient data in the UK': Palantir UK head defends record as criticisms rise Amazon’s new AI Bio Discovery tool can provide ‘every researcher’ with ‘lab-in-the-loop drug discovery’ – 40+ AI biology models can filter 300,000 novel antibody candidates down to the top results for testing in just weeks Over 100 Chrome Web Store extensions found stealing user data from thousands of accounts Europe wants tech sovereignty but is this realistic? Enterprise AI governance cannot live in a prompt. So where is the safety net? Why 2026 is the year of flexibility without friction: solving the multi-platform crisis OpenAI reveals its Mythos rival designed for cybersecurity pros When cyberattacks are inevitable, recovery becomes the strategy Closing the cloud complexity gap LaLiga uses AI to fight illegal streaming that costs its clubs $800m a year Intel and Google expand long-term chip partnership to power AI systems 'Chatbots respond not just to what you ask, but how you ask it': Report finds AI agents might be sucking up to… 'Smartphones have physical limitations': Report explains why AI is kickstarting a billion-dollar hardware arms… 'I’m pretty sure actually we really do not need to work for five days' Zoom CEO calls for end of traditional work schedules — says 3-day working week should become the norm 'It's more common than you think': Experts reveal how hackers are trying to hijack your inbox with these… 'This wasn’t just phishing — it was a full-service cybercrime platform': FBI reveals takedown of notorious W3LL phishing operation targeting thousands of victims From cloud to Agentic AI: Why security must evolve faster than innovation Basic-Fit gym group data breach exposes details of over 1 million members — here's what we know ‘Authorities can ask them to hand over data’: Report claims over 80% of Europeans don’t trust US and Chinese businesses to handle their data – Europe is desperate for homegrown AI, cloud, and telecoms as the rift with the US grows Booking.com confirms reservation data breach — tells customers hackers 'may have been able to access certain… Agility is the key to protecting against Malware-as-a-Service (MaaS) Rockstar hackers publish 78.6 million stolen records — but many of us will be disappointed Adobe issues emergency security patch — Reader and Acrobat users need to update now OpenAI flags third-party data issue — all macOS users should update now Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Hackers use Claude and ChatGPT in 'a significant evolution in offensive capability' to breach government agencies, leak hundreds of millions of citizen records ‘You’re effed’: Palantir CEO says AI ‘will destroy humanities jobs’ – but Gen Z workers are apparently deliberately sabotaging AI rollouts in an effort to fight back 'This is not your typical run-of-the-mill malware': CPUID download page hacked and tools replaced with links… Anthropic is bringing Claude's AI power to Microsoft Word How businesses can turn AI pilots into scalable solutions AI can transform customer experiences – when it lives up to its promise 'Regain control of our digital destiny': France to ditch Windows for Linux to reduce reliance on US tech How the memory crisis is strangling the UK's data center boom ‘No Decision’ is the new breach: Why inaction is becoming a career risk for CISOs in 2026 'That shouldn’t translate into investing in AI blindly, without a clear strategy': Experts warn UK firms want to keep spending big on AI - even if they can't prove it makes a difference How AI is rewriting the ERP investment playbook Rockstar confirms major third-party data breach: GTA VI maker says 'no impact on our organization or our… How to deploy physical AI effectively '71% of US households get routers from ISPs': Why new FCC rules could leave millions stuck with outdated,… 'The CPU is the system’s executive layer': Intel joins SambaNova as both face existential threat from… 'Just not sustainable': Why your monthly £25 broadband internet bill could soon hit £45 '$15K bill destroyed a solo developer’s startup': How hackers are using leaked Google API keys to… 'Today is the day you've been waiting for': eGPUs can now officially turn a humble Mac Mini into an AI… Linux pulls support for ancient CPU — unsurprisingly, Linus Torvald says there is 'zero real reason' to… 'AI is a once-in-a-lifetime opportunity': Amazon CEO Andy Jassy lays out his '6 truths' for the… 'A self-inflicted hit': Washington state just rolled back sales tax exemptions for AI data centers worth… 'There’s no one-size-fits-all office chair': Vari explains the design decisions behind its award-winning… 'Small business owners have significant creative control from start to finish' — VistaPrint reveals the… 'Experts' to rent for $1 per month: Hostinger debuts 7-person AI team to help SMBs save thousands on… Microsoft hands Linux Foundation key Surface data to help fix laptop battery life Adobe Reader users beware — experts flag months-old security flaw using booby-trapped PDFs to scope out victims 'Shockingly good value': New rugged Android tablet has a built-in 1080p projector, night-vision camera, and… Stop the presses — Microsoft is actually cutting cloud PC prices for SMBs, promises to make it 'more cost-effective for small and medium businesses' 'If one piece of your supply chain is delayed, then your whole project can't deliver': Nearly half of US data centers planned for 2026 canceled or delayed — and things could soon get much worse ChatGPT’s hidden backup model just got smarter — as OpenAI adds a cheaper Pro option 'The problem is not AI’s capability...what won’t improve on its own is the human side': Major study claims white-collar workers are fighting back against AI in the workplace Introducing Perspectives — the new home for premium contributed content on TechRadar Pro Introducing Perspectives — the new home for premium contributed content on TechRadar Pro The New Internet is Coming Lazarus and Kimsuky prove why infrastructure-level analysis is crucial for cybersecurity Claude Cowork is now available for enterprise use, adds analytics, access controls and more The internet has a trust problem - identity needs to travel OpenAI halts £31 billion Stargate UK project over rising energy costs and regulatory deadlock The 70% rule: Why your AI strategy is a people strategy Top WordPress Slider plugin hijacked to spread malware — here's what to look out for Why CIOs need a single source of truth for digital operations No, Elon Musk doesn't want to give you a $5,000 tax refund — it's a scam, here's what to look out… Intermedia Unite review 2026 Why enterprise AI will be defined by integration, not model aggregation ‘It’s a potential national security threat’: Proton study finds over 3,500 US legislators’ official emails leaked and exposed on the dark web Microsoft warns worrying security flaw exposed over 50 million Android users, says 'user credentials and financial… Google Chrome rolls out a new tool to try and stop infostealer malware in its tracks How to submit an article for TechRadar Pro Perspectives 'Orwellian Notion': Federal workers can access Claude AI again after judge ditches Trump's Anthropic ban 'Almost 100 TOPS': GMKTec debuts powerful AI Mini PC that supports three 8K screens and costs less than you… 'Remember BlackBerry?': Iconic phone maker’s patents used to hit Brother in a massive lawsuit that could… Breach exposes sensitive LAPD files stored in city attorney system ‘FlamingChina’ hacker claims to have stolen over 10 petabytes of advanced military data from China’s National Supercomputing Center in possibly the biggest hack of all time Mac users beware — experts say this attack 'stood out immediately' by making a major change to try… Could AMD's former foundry be quietly building up to become a major Arm — and AMD — rival? Now that's different - hackers use miniature SVG images to try and hide credit card stealer "A future-proof powerhouse for demanding tasks": MSI's RTX5090 creative laptop gets a $300 price cut… Closing the implementation gap in America's cyber strategy UK NHS chief champions Palantir’s 'outstanding results’ in England, pushes for deeper rollout despite… French email provider accidentally leaked 40 million records — L’Oreal, Renault, French government data…
What AI coding benchmarks still miss about software quality
Andrian Budantsov · 2026-05-21 · via Latest from TechRadar in Pro

Most AI coding benchmarks still ask the question: did the agent produce code that passes the current tests?

This is a useful question, but it is too narrow. Software development is iterative. Requirements change and edge cases appear. Old design decisions become constraints on new work. Code that passes today can still make the next change slower and more expensive, while also increasing risk.

The gap matters more as AI raises the volume of code change. When generation gets cheap, the real question shifts from ‘can the agent produce a working patch?’ to ‘what kind of codebase does repeated agent use create over time?’

CEO of Hypersequent.

A recent paper, SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks (Orlanski et al.), gets closer to that question than most benchmark work. Instead of scoring one-shot solutions, it makes agents extend their own prior code across 20 problems and 93 checkpoints.

Each checkpoint changes the specification. The agent does not start fresh and is not given an internal design to follow. It has to live with earlier choices.

This setup is closer to real development than most benchmark suites, because real teams inherit yesterday's shortcuts.

Green tests can hide a worse codebase

The paper tracks two quality signals alongside correctness. Verbosity measures redundant or duplicated code. Structural erosion measures how much of a codebase's complexity gets trapped inside functions that are already too complex.

Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!

Those are failure modes familiar for every engineering manager. A system can keep passing tests while more logic gets pushed into the same large functions and more special cases get bolted on. More files need to be touched for every feature. The software still works, but becomes more difficult to change.

The code-search example in the test is a good example of this issue. At first, the system only needs to find Python code using exact text or regular expressions. Later on, it needs to handle more languages, understand the code structure (AST matching), and even automatically fix problems.

If the initial design is too strict and makes early assumptions, it might pass the first tests but won't be able to handle the complex, later requirements easily.

The results are clear. None of the evaluated agents solved any problem end to end. The best strict solve rate was 17.2 percent, and by the final checkpoint strict solve rates fell to 0.5 percent. Across trajectories, verbosity rose in 89.8 percent of runs and structural erosion in 80 percent.

The comparison with human-maintained code is even more useful. Against 48 maintained Python repositories, agent-generated code was 2.2 times more verbose and more structurally eroded.

When the authors tracked 20 of those repositories over time, the human code was comparatively flat while the agent code kept worsening with each iteration.

A passing suite tells you the latest version satisfied known checks. It does not tell you whether the code is becoming more fragile or more expensive to extend.

Why this matters for QA

For QA leaders, there are two key takeaways. The first is obvious: AI-built product code can degrade under repeated change even while current tests stay green. Teams may read continued output as proof that the system is healthy. In reality, they may be accumulating future regression cost at higher speed.

The second is closer to home. QA teams are now using AI tools to write and maintain tests, especially functional UI automation in tools like Playwright. That work follows the same pattern as the paper: the product changes, the test has to change, the next feature adds another branch, another selector, another exception, another helper.

The paper is about coding broadly, not automation test suites specifically, but the mechanism carries over. A test suite can also become verbose and structurally weak under repeated AI-assisted edits.

A degraded test suite is harder to notice than degraded product code. The pipeline can still be green and the suite can still look larger on paper. Coverage can appear to improve.

Meanwhile, the core asset might be degrading. This could include bad selectors, weak checks, copied test steps, overly large helper functions, and UI tests that are hard to fix and easy to doubt. While test flakiness is obvious, problems like tests that don't do much or tests that run very slowly might not be noticed right away.

For QA leaders, that shifts the job. Quality assurance cannot stop at validating the latest output against today's requirements. It also has to watch whether repeated change is damaging both the product and the test system that is supposed to protect it.

The role of QA leadership is changing; quality assurance must now go beyond simply verifying the latest product output against current requirements. QA leaders must also monitor whether continuous change is negatively impacting both the product's quality and the integrity of the testing system designed to safeguard it.

Prompting will not solve this by itself

The paper also tested whether better prompts could control the drift. They helped at the start, but not for long. Quality-aware prompts lowered initial verbosity and erosion. One anti-slop prompt cut initial verbosity by about a third on GPT-5.4.

The change was minimal. Cleaner starting points still degraded at roughly the same rate, and the better-looking code did not reliably improve pass rates. In some cases, the prompts increased cost.

Many organizations treat prompting as a governance layer. While this helps, it is not enough. If the workflow keeps asking an agent to extend its own code under changing requirements, the organization still needs controls outside the prompt.

A better way to evaluate AI-assisted development

To manage AI-assisted development well, you need to look past quick wins. Check the code changes after a few adjustments, not just the first fix. Watch out for complex or repeated parts in the code.

Don't confuse success on the current feature with confidence in long-term stability. Consider how easy the code is to maintain as a release risk, especially for systems dealing with things like cost, user ID, access rights, money, or rules.

The same rule applies to tests. Review how AI-generated test code changes after several product iterations. Watch for suites that grow faster than their signal and UI tests that absorb behavior better covered at lower levels.

Also be aware of ‘self-healing’ maintenance that subtly lowers assertion strength. A larger suite doesn’t automatically mean better control.

Quality needs to move upstream. By the time a feature reaches final validation, some of the damage may already be baked into the path the system took to get there.

QA needs a voice earlier in the loop: in design constraints, review standards, regression strategy, and the definition of acceptable change quality for both product code and test code.

Ultimately, passing tests still matters, but as AI increases the volume of code change, the more useful question is whether each successful change leaves the codebase safer to extend or more dangerous to touch.

We've featured the best AI website builder.

This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today.

The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit