惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
Martin Fowler
Martin Fowler
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
C
Cisco Blogs
Last Week in AI
Last Week in AI
T
The Blog of Author Tim Ferriss
The GitHub Blog
The GitHub Blog
T
Tenable Blog
A
Arctic Wolf
小众软件
小众软件
Google DeepMind News
Google DeepMind News
aimingoo的专栏
aimingoo的专栏
PCI Perspectives
PCI Perspectives
博客园 - 司徒正美
The Last Watchdog
The Last Watchdog
H
Hacker News: Front Page
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Stack Overflow Blog
Stack Overflow Blog
N
News and Events Feed by Topic
Security Archives - TechRepublic
Security Archives - TechRepublic
博客园 - 【当耐特】
S
Security @ Cisco Blogs
P
Proofpoint News Feed
Cloudbric
Cloudbric
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Jina AI
Jina AI
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
月光博客
月光博客
Schneier on Security
Schneier on Security
Hacker News: Ask HN
Hacker News: Ask HN
V
Visual Studio Blog
D
DataBreaches.Net
H
Help Net Security
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Project Zero
Project Zero
阮一峰的网络日志
阮一峰的网络日志
Cyberwarzone
Cyberwarzone
博客园 - Franky
Y
Y Combinator Blog
Spread Privacy
Spread Privacy
N
News and Events Feed by Topic
The Cloudflare Blog
Simon Willison's Weblog
Simon Willison's Weblog
S
SegmentFault 最新的问题
W
WeLiveSecurity
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
I
Intezer
Hugging Face - Blog
Hugging Face - Blog
Attack and Defense Labs
Attack and Defense Labs

Lobsters

Lunacy | Red Vice CIFSwitch: a non-universal Linux local root vulnerability RIPE NCC session fixation: poaching logins with an Atlas probe GNOME 2.20 but its Web Components Agentic Search for Context Engineering – Leonie Monigatti Garnix is shutting down [not OC] akashina.tngl.sh/jjc Concerning Emacs (and Jazz) Nitpicking the shell history scene in ‘Tron: Legacy’ What's cooking on SourceHut? Q2 2026 The tenth OpenPGP email summit Package managers that package package managers Clojure on Fennel part three: parsing WordPress at 23 GitHub - creusot-rs/creusot: Creusot helps you prove your Rust code is correct. Announcing Rust 1.96.0 | Rust Blog A Love Letter to Neovim sqlite AGENTS.md Am I a Bad Friend? CSS vs. JavaScript • Josh W. Comeau Erlang Ecosystem Foundation - Supporting the BEAM community A brief note about slot access cost in Common Lisp Keyboard latency probe Rethinking the GNOME clipboard issues Back to the Building Blocks’ Building Blocks Tech Notes: Theseus: translating win32 to wasm Fast is better than slow Content-addressed Rust builds (or, what kache actually caches) Intent to Prototype: Embedding API Canada’s Bill C-22 and the security cost of collecting more data 5 PostgreSQL locking behaviors that trip people up okmij.org Stop advertising in your commits! | AksDev GitHub - mplsllc/macsurf: A modern web browser for Classic Mac OS 9 PowerPC. Real CSS3, ES5 JavaScript, native HTTPS — built with CodeWarrior on the Carbon API. Introducing DoomBench - Can Your Data Stack Run DOOM? What are some of your favourite developer tools? Building a Scalable Ingestion Pipeline with Temporal (Part 1) Converting shallow Git bundles into normal repositories Are you a member of any professional associations? What is a harmonic? An interactive comic about additive synthesis How Virtual Tables Work in the Itanium C++ ABI Using SwiftUI to Build a Mac-assed App in 2026 Rust (and Slint) on a jailbroken Kindle. ~jack/lambda-on-lambda - Serverless Haskell on AWS - sourcehut git Human proof for FOSS contributions Extremely simple internet radio controlled via IRC Announcing BABLR Splitting Konsole views from Helix to run tools | AksDev GitHub - yugr/rust-slides Serving files over HTTP three ways: synchronous, epoll, and io_uring update docs with information about building with build.py (#979) · astral-sh/python-build-standalone@c9c40c5 A Simple Makefile Tutorial On C extensions, portability, and alternative compilers Switching to Colemak | Pedro Alves Just How Bad Was The Intel IAPX432? Nix's Substituter List Is Not a Routing Table Accelerating copy_if using SIMD Lambda on Lambda: Serverless Haskell on AWS | Blog Announcing feed-repeat v1.0 Scaling Akvorado BMP RIB with sharding EYG news: A host of CLI improvements, new guides and new effects The social contract of writing JS Crossword C array types are weird; and related topics Flatpak will depend on systemd – OSnews Migrating from Go to Rust | corrode Rust Consulting A portentous reunion Vivado Licensing Options How my minimal, memory-safe Go rsync steers clear of vulnerabilities the entropy layer of a wavelet codec, on its own GitHub - nferhat/fht-compositor: A dynamic tiling Wayland compositor. Debian SE Linux and PinTheft Does bulk memmove speed up std::remove_if? (No.) 声明式部分更新 | Blog | Chrome for Developers Fully in-browser container builds Dianne Skoll's Web Site - Remind The Architecture of Open Source Applications (Volume 1)Berkeley DB Pardon MIE? - ironPeak Blog “Long-Term Support” doesn’t mean what you think Jira IS Turing-Complete May I recommend thinking of Emacs as your Fortress of Solitude hershey Floodgap Gopher-HTTP gateway gopher://thelambdalab.xyz/1cuneiforth/ HP QuickWeb, Singular And Pointless That one time I used Go panics for flow control A new suite of modern tools coming for editing and publishing RFCs From the Tabletop… The Digital Antiquarian Building a Host-Tuned GCC to Make GCC Compile Faster Are we self-sovereign PKI yet? Claw Patrol: an open-source security firewall for agents | Deno Revised^7 Report on Scheme, Large: Procedural Fascicle Draft is now public A Network Allow-List Won't Stop Exfiltration — André Graf From AFSK to Goertzel – µArt.cz Software For My New Home Server Introducing Neptune: Direct3D virtualization for QEMU AI Agent Bankrupted Their Operator While Trying to Scan DN42 - Lan Tian @ Blog mimalloc: A new, high-performance, scalable memory allocator for the modern era Making wl_shm fast The Soul of Maintaining a New Machine - Third Draft | Books in Progress What is Git made of?
Finding Miscompiles for Fun, Not Profit
Justin Lebar · 2026-05-28 · via Lobsters

Update June 1: The day after we published, Anthropic released Opus 4.8 and “ultracode” mode in Claude Code. Our preliminary experiments indicate that together these are significantly better at filtering out low-severity bugs, and that the cost per medium-to-high severity bug found is maybe 1/5 (with very large error bars) that of the workflow described in this article.

I’ve worked on compilers for ML for the last decade across Google, Waymo, and OpenAI. This includes CUDA support in clang, XLA:GPU, Triton, and OpenAI’s custom hardware. I’ve seen stuff. But over the past week or so I had one of the most unsettling experiences of my career: In one afternoon, I spent more than $10,000 running AI agents over compiler code, finding hundreds of plausible bugs in LLVM, including many miscompiles and at least one that’s Quite Serious. This is the story of how I got here and where we might be going.

In January 2026, I decided to try to find some bugs in LLVM (the compiler behind clang, rustc, and AMD’s GPU compiler, among others), as a personal project. Codex and I collaboratively wrote a fuzzer. The basic idea is to generate a random program, run it through part of the compiler, and then check that the resulting program after compilation does the same thing as the original program (usually just by running the two programs). I spent a few weeks on it, and I found and fixed five bugs in instcombine, LLVM’s peephole optimization pass. After that, my fuzzer started taking longer to find bugs, and I lost interest.

Fast forward to mid-May 2026. I joined SemiAnalysis as a contractor, and I decided to try applying the same technique to NVIDIA’s low-level compiler, ptxas. I expected this to be less fruitful than fuzzing LLVM, for a few reasons.

  • In general, fuzzers can get “stuck”: Once they find a bug, they can keep finding new ways to trigger it. With an open-source compiler like LLVM, you can “just” fix the bug and then continue fuzzing. But with a closed-source compiler like ptxas, the best you can do is try to modify your fuzzer so it doesn’t generate inputs that trigger the same bug. Implementing this is tedious at best.

  • With LLVM I can run just one pass (e.g. instcombine), whereas with ptxas I have to run the whole compiler end-to-end. I worried that this would make some bugs require larger or more complicated reproducers, making them less likely to be found by a fuzzer.

  • I can build LLVM myself, so I can compile with flags that add instrumentation into the program under test to help the fuzzer choose “interesting” inputs that explore new parts of the program. Although AFL++ has modes for gathering instrumentation from precompiled binaries, they slow down the program under test, and in general I didn’t expect them to be as useful. (I didn’t end up actually using these modes when fuzzing ptxas; I just did purely undirected fuzzing.)

I expected maybe I’d find a handful of bugs after a few weeks of work, like before.

Instead, in three days, I had 40 programs that ptxas miscompiles. (A week later, this number is up to about 80.) Although some of these test cases probably reflect the same underlying bug in the compiler, I was still kind of astonished. Many of the reproducers reduce to fairly “normal-looking” instruction sequences.

If you want to have a look at the bugs I found (really, that Codex and Claude found), they live in the FuzzX repo on GitHub.

Why was my fuzzer so much easier to write this time? As far as I can tell, it was the difference between ChatGPT 5.2 and 5.5. This time, I vibe-coded this entire fuzzer, never looking at a line of code. The LLM did the tedious job of altering the fuzzer after every bug we found so that we wouldn’t get stuck finding the same bug over and over. It also minimized each test case it generated, often spending an hour or longer doing so. It independently decided which PTX instructions to fuzz, and which sequences of instructions were “safe” (i.e. didn’t trigger undefined behavior). I just put it in a loop using /goal and went to sleep.

To be clear, it’s not surprising that fuzzing found bugs. But it was surprising that I found this many bugs so quickly, with almost no manual effort.

Naturally, my next question was whether I could also find bugs in LLVM’s AMDGPU backend. You won’t be surprised to hear that I could, at roughly the same rate as I found bugs in ptxas. At some point, my personal ChatGPT Pro account ran out of credits and I switched to SemiAnalysis’s Claude account. I didn’t notice a difference in quality between Opus 4.7 and ChatGPT 5.5, they were both great.

The story could end here. I reported the bugs I’d found to AMD and NVIDIA. As of this writing, AMD has already fixed five of them, and because their compiler is open-source, I can apply their fixes immediately. If any of these bugs were critical to me, I could even fix them myself. Hooray for open-source toolchains.

At this point my ptxas and AMDGPU fuzzers started to slow down; they were running for increasingly long times without finding any new bugs. I almost stopped there, but I had a thought. What if I just asked Claude to read through LLVM and find bugs? In other words, what if I was building a fuzzer because I hadn’t absorbed the Bitter Lesson?

I asked Claude to spin up 50 subagents at a time looking for bugs. Boy did I get them. Claude found bugs at a rate of one every four minutes. (In contrast, at this point the fuzzer was taking hours to find new bugs.)

My initial reaction was, I don’t even think this means LLVM’s AMDGPU backend is particularly buggy. Indeed, a friend suggested I do the same to the x86 backend, and it produced bugs at a rate of almost two per minute.

My bug-finding agents didn’t show any signs of slowing down before I stopped them. I don’t know how many more issues are out there. Probably a lot?

The question with automated bug-finding usually is, “Do these bugs matter?” I’ve only gone through about 20% of the x86 bugs so far. The bugs found by agents are definitely less severe on average than those found by the fuzzer; every fuzz bug is a demonstrable miscompile, whereas agents are using their judgement to decide what is and isn’t a bug, and sometimes they’re wrong.

On the other hand, one of the 30 or so agent-discovered bugs I’ve examined includes this extremely frightening case where LLVM will turn an atomic store into two non-atomic stores. This bug would have been very difficult to find via fuzzing (fuzzing atomics is hard), and would likely be both quite bad and quite hard to root-cause if it happened in production (because downgrading an atomic to a non-atomic store will be fine 99% of the time, and 1% of the time will corrupt your data).

The other question you should ask is, how much did all this cost? The vibe-coded fuzzer was relatively cheap. I was using about double the weekly quota for my $200/month ChatGPT Pro account to write the AMD and NVIDIA fuzzers at the same time. It was more expensive with pay-per-token Opus 4.7, on the order of $1000 for a few days’ work. I don’t know if it would have been cheaper with a Claude Max account, and I don’t know if the work I was doing with Claude was more token-intensive than the work I’d been doing with Codex earlier.

On the other hand, having an army of subagents read the code was, um, not cheap. I spent more than $10,000 in a few hours, and I wasn’t even using fast mode. (Thanks for the tokens, Dylan!) Even though bugs found this way were on average lower-severity than fuzz bugs, the fact that code inspection can find whole classes of bugs that would be very challenging to find via fuzzing makes it valuable to me.

Frankly, just the one atomics bug mentioned could do far more than $10,000 in damage under the right (or perhaps wrong) conditions; silent data corruption like this will break just about any production system.  And even if your system had safeguards that allowed it to notice this kind of corruption, you’d still likely burn months of engineering time hunting for its source.  If I were doing this again, I wouldn’t bother with fuzzing, assuming I had the budget to unleash my herd of agents.

I’m still trying to make meaning out of this experience. I think it’s bigger than simply “With enough subagents, all bugs are shallow.” Maybe the lesson is this:

Things that were impossible five months ago are now “just” Very Expensive.

A corollary is, if you don’t have the budget, you’re operating in a smaller part of the possibility space than those who do. And I expect that gap to grow substantially over the coming months.

SemiAnalysis spent an order of magnitude more on my tokens than on paying me the day I ran the agents. If the value of something is how much someone is willing to pay for it, then this was the first time in my career that I delivered less value to my employer than my AI did.

It makes me wonder how much SemiAnalysis will be willing to pay for tokens in six months, and what life will be like for people and companies who can’t or won’t pay.

One last thought. I could do the “have Claude read the code” approach with the AMDGPU and x86 LLVM backends because I have the source code. I can’t do it with ptxas because I only have a binary. But…how sure am I that, given enough tokens, Opus 5.7 (or even 4.7) couldn’t find bugs just by reading the assembly?