惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
The GitHub Blog
The GitHub Blog
B
Blog
小众软件
小众软件
Jina AI
Jina AI
WordPress大学
WordPress大学
V
V2EX
MongoDB | Blog
MongoDB | Blog
Blog — PlanetScale
Blog — PlanetScale
P
Proofpoint News Feed
Y
Y Combinator Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
人人都是产品经理
人人都是产品经理
Microsoft Azure Blog
Microsoft Azure Blog
aimingoo的专栏
aimingoo的专栏
C
CERT Recently Published Vulnerability Notes
C
Cisco Blogs
Project Zero
Project Zero
云风的 BLOG
云风的 BLOG
K
Kaspersky official blog
Google DeepMind News
Google DeepMind News
宝玉的分享
宝玉的分享
T
Threat Research - Cisco Blogs
S
Securelist
V
Vulnerabilities – Threatpost
雷峰网
雷峰网
F
Fortinet All Blogs
D
DataBreaches.Net
I
Intezer
D
Docker
The Hacker News
The Hacker News
The Last Watchdog
The Last Watchdog
SecWiki News
SecWiki News
MyScale Blog
MyScale Blog
腾讯CDC
博客园_首页
Martin Fowler
Martin Fowler
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Application and Cybersecurity Blog
Application and Cybersecurity Blog
H
Help Net Security
GbyAI
GbyAI
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
L
Lohrmann on Cybersecurity
I
InfoQ
H
Hacker News: Front Page
T
Threatpost
Stack Overflow Blog
Stack Overflow Blog
博客园 - 叶小钗
T
Troy Hunt's Blog
Microsoft Security Blog
Microsoft Security Blog

Latest from TechRadar

Quordle hints and answers for Monday, April 13 (game #1540) NYT Strands hints and answers for Monday, April 13 (game #771) NYT Connections hints and answers for Monday, April 13 (game #1037) Morbid Metal developer explains why he ditched an origami art direction in favor of gritty sci-fi — 'It worked, but it didn't really feel like me' '71% of US households get routers from ISPs': Why new FCC rules could leave millions stuck with outdated,… 'The CPU is the system’s executive layer': Intel joins SambaNova as both face existential threat from… ‘More bang for your buck’: 7 easy ways to boost your MacBook Neo’s performance for free DJI Romo P vs Roborock Saros 10R — which robot vacuum comes out on top when it comes to dodging obstacles? I put… I spent 6 hours with Genshin Impact on the Galaxy S26 Ultra, and I can't believe how far mobile gaming has come What is the release date for The Testaments episode 4 on Hulu and Disney+? I reviewed the LG G6 for 3 weeks, and it's a fantastic OLED TV that's the new best option for brighter rooms Is your bird feeder camera doing more harm than good? 3 tips for using it safely as RSPB issues urgent disease warning Chelsea vs Man City Live Streams: How to watch Premier League 2025/26 from anywhere in the world, team news How to watch Alcaraz vs Sinner for FREE: TV Channels for Monte-Carlo Masters Final Sunderland vs Tottenham Live Streams: How to watch Premier League 2025/26 from anywhere in the world, team news Are these the best-designed workout headphones ever? I used them for a month to find out How to watch Snooker 900 John Virgo online (it's free) – stream O'Sullivan vs Higgins anywhere I've only just discovered the Walk With Frodo app on Garmin's Connect IQ store — and as as a huge LOTR nerd, it's going to make the next 1,800 miles fly by 'Just not sustainable': Why your monthly £25 broadband internet bill could soon hit £45 How to watch Paris-Roubaix 2026: Free Streams & TV Info as Tadej Pogacar chases third Monument How to watch Euphoria season 3 online – stream Zendaya & Sydney Sweeney drama from anywhere today '$15K bill destroyed a solo developer’s startup': How hackers are using leaked Google API keys to… There's a sneaky way to watch UFC 327 really cheap... NYT Connections hints and answers for Sunday, April 12 (game #1036) NYT Strands hints and answers for Sunday, April 12 (game #770) Quordle hints and answers for Sunday, April 12 (game #1539) Amazon's Ring cameras are the perfect solution to secure your home on a budget — shop today's best deals… I've tested every iPhone since the iPhone 12, and Ceramic Shield 2 is the first iPhone glass I fully trust UFC 327 live stream: how to watch Procházka vs Ulberg, start time, preview, full card We're officially getting the DJI Pocket 4 on April 16, but here's how Insta360 could beat it 'Today is the day you've been waiting for': eGPUs can now officially turn a humble Mac Mini into an AI… Linux pulls support for ancient CPU — unsurprisingly, Linus Torvald says there is 'zero real reason' to… Keanu Reeves' new Apple TV movie Outcome has been slammed by critics — watch these 4 highly-rated films with the beloved actor instead 'AI is a once-in-a-lifetime opportunity': Amazon CEO Andy Jassy lays out his '6 truths' for the… How to watch Grand National 2026: Free Streams & TV Channels for Aintree National Hunt Race ‘I hadn’t verified a single thing’: Using ChatGPT for Iran war news changed how I trust information Want cafe-quality lattes at home without buying an expensive new coffee machine? Jura's new gadget upgrades your drinks with perfectly foamed milk every time 'A self-inflicted hit': Washington state just rolled back sales tax exemptions for AI data centers worth… Playing The Last of Us with friends made my favorite PlayStation game feel brand new again Mint Mobile's new Samsung Galaxy S26 series deal can save you up to $900 — enough to cover an entire device Not a squat, not a deadlift — the trap bar deadlift 'sits between' the two, builds muscle fast and is… Record Store Day 2026 starts soon! The date, the top vinyl drops, and everything else you need to know Women's Six Nations 2026 Free Streams: TV Channels, Preview, Table, Round 5 Fixtures, France vs England Time Beyond Paradise season 4 star would 'love' to do The Celebrity Traitors season 2 — and would be 'terrified' if one contestant came to Shipton Abbott 'There’s no one-size-fits-all office chair': Vari explains the design decisions behind its award-winning… I was a vacuum reviewer for two years — these are the 6 sub-£250 models I'd recommend in a heartbeat Save $200 and get the Samsung Galaxy S26 Ultra at its preorder price for a limited time at Amazon 'Small business owners have significant creative control from start to finish' — VistaPrint reveals the… TurboQuant isn't the RAM crisis savior you're hoping for, analysts say — as memory prices continue to… ICYMI: the 7 biggest tech stories of the week, from DJI's new robovac to Artemis II iPhone photos I matched the upgraded Meta AI against ChatGPT, and you can really tell which AI has social media roots Quordle hints and answers for Saturday, April 11 (game #1538) I created my dream coffee corner at IKEA for under $100 — and my mornings are about to get a lot cozier 'Experts' to rent for $1 per month: Hostinger debuts 7-person AI team to help SMBs save thousands on… The new MacBook Air has already dropped to a record-low price on Amazon I tested Turtle Beach's Mario-themed controller and headset for Nintendo Switch 2 — and they surprised me for… NYT Strands hints and answers for Saturday, April 11 (game #769) NYT Connections hints and answers for Saturday, April 11 (game #1035) After soaring 2,200%, DDR4 RAM prices finally fall — but don't get too excited It's "completely changed my home cleaning habits": The Dreame Z20 is a highly effective vacuum cleaner for even lsrger homes. Beyond no-log: Tor looks into seizure-proof servers that forget your data There's a sneaky way to watch IPL 2026 for FREE Microsoft hands Linux Foundation key Surface data to help fix laptop battery life Adobe Reader users beware — experts flag months-old security flaw using booby-trapped PDFs to scope out victims 'Shockingly good value': New rugged Android tablet has a built-in 1080p projector, night-vision camera, and… Stop the presses — Microsoft is actually cutting cloud PC prices for SMBs, promises to make it 'more cost-effective for small and medium businesses' Microsoft has begun stripping out AI from Windows 11 — but it's already being criticized for not going far… Euphoria season 3 episode 3 release date: when will it come out on HBO Max? 'If one piece of your supply chain is delayed, then your whole project can't deliver': Nearly half of US data centers planned for 2026 canceled or delayed — and things could soon get much worse ChatGPT’s hidden backup model just got smarter — as OpenAI adds a cheaper Pro option Forget Big Mistakes — new Netflix true crime series Trust Me: The False Prophet is the only TV show you need to… 'The problem is not AI’s capability...what won’t improve on its own is the human side': Major study claims white-collar workers are fighting back against AI in the workplace Introducing Perspectives — the new home for premium contributed content on TechRadar Pro ‘Computers are no longer a bicycle for the mind’: Frameworks founder says the Steve Jobs era is over and PCs are now a ‘self-driving car that takes you directly to the destination’ No, Elon Musk doesn't want to give you a $5,000 tax refund — it's a scam, here's what to look out… ‘It’s a potential national security threat’: Proton study finds over 3,500 US legislators’ official emails leaked and exposed on the dark web ‘I want to cancel’: YouTube Premium quietly hikes its US prices for the first time in three years, forcing… RTX 5090s and other high-powered graphics cards may carry risks of cable melting issues — but Asus thinks it has… Former Xbox exec thinks Naughty Dog's decision to cancel the 80% completed The Last of Us Online 'was the right call', but it shouldn't have greenlit it in the first place — 'The ambition was there, but the realistic upfront planning wasn't', she says West Ham vs Wolves Live Streams: How to watch Premier League 2025/26 from anywhere in the world Microsoft warns worrying security flaw exposed over 50 million Android users, says 'user credentials and financial… ‘Apple will grit its teeth and push through’ — new report suggests the iPhone Air 2 isn’t dead,… Google Chrome rolls out a new tool to try and stop infostealer malware in its tracks 'Two Hells collide' — Doom: The Dark Ages and Diablo Immortal unite in a limited-time crossover event,… Spotify is rolling out new video controls, and as someone who hates its in-app music videos, I know this will be a huge… 8 new movies and TV shows to watch on Netflix, Prime Video, HBO Max, and more this weekend (April 10) AdGuard VPN has a new app for iPhone — and you can try it out for 7 days for free Currys refuses to end its Easter sale — I've found the 21 best tech deals that are still available Amazon is slashing prices on Garmin watches — save up to $350 on best-rated models for running, biking and hiking Inspired to start running this summer? Here are 8 brilliant running shoes I'd recommend for beginners NASA used a 12-year-old GoPro to capture a sight called the ‘greatest gift’ by Artemis II pilot — and… iPhone owners urged to change this key privacy setting after FBI recovers suspect’s deleted Signal messages How to read Murder in Purple and Gold online from anywhere Garmin's cashing in on the screenless Whoop-style smart band trend with its upcoming CIRQA — here's the… YouTube insists that a 90-sec, unskippable ad format 'isn't something we are testing' — but furious… ‘Everything is magenta’: This wild hack got Mac OS X Cheetah working on a Nintendo Wii, and I can’t… A new free-to-play Borderlands game gets surprise drop on mobile, which Zynga says is part of a 'limited-time… The Xiaomi 17 outmuscles the iPhone 17 and Galaxy S26 in several key areas — read our full review In a sea of PlayStation Portal cases, the one I value the most has yet to be beaten How to submit an article for TechRadar Pro Perspectives
What AI coding benchmarks still miss about software quality
Andrian Budantsov · 2026-05-21 · via Latest from TechRadar

Most AI coding benchmarks still ask the question: did the agent produce code that passes the current tests?

This is a useful question, but it is too narrow. Software development is iterative. Requirements change and edge cases appear. Old design decisions become constraints on new work. Code that passes today can still make the next change slower and more expensive, while also increasing risk.

The gap matters more as AI raises the volume of code change. When generation gets cheap, the real question shifts from ‘can the agent produce a working patch?’ to ‘what kind of codebase does repeated agent use create over time?’

CEO of Hypersequent.

A recent paper, SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks (Orlanski et al.), gets closer to that question than most benchmark work. Instead of scoring one-shot solutions, it makes agents extend their own prior code across 20 problems and 93 checkpoints.

Each checkpoint changes the specification. The agent does not start fresh and is not given an internal design to follow. It has to live with earlier choices.

This setup is closer to real development than most benchmark suites, because real teams inherit yesterday's shortcuts.

Green tests can hide a worse codebase

The paper tracks two quality signals alongside correctness. Verbosity measures redundant or duplicated code. Structural erosion measures how much of a codebase's complexity gets trapped inside functions that are already too complex.

Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!

Those are failure modes familiar for every engineering manager. A system can keep passing tests while more logic gets pushed into the same large functions and more special cases get bolted on. More files need to be touched for every feature. The software still works, but becomes more difficult to change.

The code-search example in the test is a good example of this issue. At first, the system only needs to find Python code using exact text or regular expressions. Later on, it needs to handle more languages, understand the code structure (AST matching), and even automatically fix problems.

If the initial design is too strict and makes early assumptions, it might pass the first tests but won't be able to handle the complex, later requirements easily.

The results are clear. None of the evaluated agents solved any problem end to end. The best strict solve rate was 17.2 percent, and by the final checkpoint strict solve rates fell to 0.5 percent. Across trajectories, verbosity rose in 89.8 percent of runs and structural erosion in 80 percent.

The comparison with human-maintained code is even more useful. Against 48 maintained Python repositories, agent-generated code was 2.2 times more verbose and more structurally eroded.

When the authors tracked 20 of those repositories over time, the human code was comparatively flat while the agent code kept worsening with each iteration.

A passing suite tells you the latest version satisfied known checks. It does not tell you whether the code is becoming more fragile or more expensive to extend.

Why this matters for QA

For QA leaders, there are two key takeaways. The first is obvious: AI-built product code can degrade under repeated change even while current tests stay green. Teams may read continued output as proof that the system is healthy. In reality, they may be accumulating future regression cost at higher speed.

The second is closer to home. QA teams are now using AI tools to write and maintain tests, especially functional UI automation in tools like Playwright. That work follows the same pattern as the paper: the product changes, the test has to change, the next feature adds another branch, another selector, another exception, another helper.

The paper is about coding broadly, not automation test suites specifically, but the mechanism carries over. A test suite can also become verbose and structurally weak under repeated AI-assisted edits.

A degraded test suite is harder to notice than degraded product code. The pipeline can still be green and the suite can still look larger on paper. Coverage can appear to improve.

Meanwhile, the core asset might be degrading. This could include bad selectors, weak checks, copied test steps, overly large helper functions, and UI tests that are hard to fix and easy to doubt. While test flakiness is obvious, problems like tests that don't do much or tests that run very slowly might not be noticed right away.

For QA leaders, that shifts the job. Quality assurance cannot stop at validating the latest output against today's requirements. It also has to watch whether repeated change is damaging both the product and the test system that is supposed to protect it.

The role of QA leadership is changing; quality assurance must now go beyond simply verifying the latest product output against current requirements. QA leaders must also monitor whether continuous change is negatively impacting both the product's quality and the integrity of the testing system designed to safeguard it.

Prompting will not solve this by itself

The paper also tested whether better prompts could control the drift. They helped at the start, but not for long. Quality-aware prompts lowered initial verbosity and erosion. One anti-slop prompt cut initial verbosity by about a third on GPT-5.4.

The change was minimal. Cleaner starting points still degraded at roughly the same rate, and the better-looking code did not reliably improve pass rates. In some cases, the prompts increased cost.

Many organizations treat prompting as a governance layer. While this helps, it is not enough. If the workflow keeps asking an agent to extend its own code under changing requirements, the organization still needs controls outside the prompt.

A better way to evaluate AI-assisted development

To manage AI-assisted development well, you need to look past quick wins. Check the code changes after a few adjustments, not just the first fix. Watch out for complex or repeated parts in the code.

Don't confuse success on the current feature with confidence in long-term stability. Consider how easy the code is to maintain as a release risk, especially for systems dealing with things like cost, user ID, access rights, money, or rules.

The same rule applies to tests. Review how AI-generated test code changes after several product iterations. Watch for suites that grow faster than their signal and UI tests that absorb behavior better covered at lower levels.

Also be aware of ‘self-healing’ maintenance that subtly lowers assertion strength. A larger suite doesn’t automatically mean better control.

Quality needs to move upstream. By the time a feature reaches final validation, some of the damage may already be baked into the path the system took to get there.

QA needs a voice earlier in the loop: in design constraints, review standards, regression strategy, and the definition of acceptable change quality for both product code and test code.

Ultimately, passing tests still matters, but as AI increases the volume of code change, the more useful question is whether each successful change leaves the codebase safer to extend or more dangerous to touch.

We've featured the best AI website builder.

This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today.

The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit