惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

爱范儿
爱范儿
大猫的无限游戏
大猫的无限游戏
WordPress大学
WordPress大学
C
Cyber Attacks, Cyber Crime and Cyber Security
D
DataBreaches.Net
G
Google Developers Blog
博客园 - Franky
V
V2EX
博客园 - 叶小钗
D
Docker
The GitHub Blog
The GitHub Blog
Microsoft Security Blog
Microsoft Security Blog
博客园 - 【当耐特】
H
Hackread – Cybersecurity News, Data Breaches, AI and More
B
Blog RSS Feed
月光博客
月光博客
M
MIT News - Artificial intelligence
F
Fortinet All Blogs
Microsoft Azure Blog
Microsoft Azure Blog
人人都是产品经理
人人都是产品经理
IT之家
IT之家
Google DeepMind News
Google DeepMind News
Apple Machine Learning Research
Apple Machine Learning Research
V
Visual Studio Blog
博客园 - 司徒正美
Stack Overflow Blog
Stack Overflow Blog
罗磊的独立博客
J
Java Code Geeks
U
Unit 42
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 聂微东
T
Tailwind CSS Blog
T
The Blog of Author Tim Ferriss
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Blog — PlanetScale
Blog — PlanetScale
Jina AI
Jina AI
C
Check Point Blog
Y
Y Combinator Blog
MyScale Blog
MyScale Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
阮一峰的网络日志
阮一峰的网络日志
宝玉的分享
宝玉的分享
B
Blog
小众软件
小众软件
云风的 BLOG
云风的 BLOG
I
InfoQ
Recorded Future
Recorded Future
酷 壳 – CoolShell
酷 壳 – CoolShell
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
美团技术团队

Lobsters

CIFSwitch: a non-universal Linux local root vulnerability RIPE NCC session fixation: poaching logins with an Atlas probe GNOME 2.20 but its Web Components Agentic Search for Context Engineering – Leonie Monigatti Garnix is shutting down [not OC] akashina.tngl.sh/jjc Concerning Emacs (and Jazz) Nitpicking the shell history scene in ‘Tron: Legacy’ What's cooking on SourceHut? Q2 2026 The tenth OpenPGP email summit Package managers that package package managers Clojure on Fennel part three: parsing WordPress at 23 Finding Miscompiles for Fun, Not Profit GitHub - creusot-rs/creusot: Creusot helps you prove your Rust code is correct. Announcing Rust 1.96.0 | Rust Blog A Love Letter to Neovim sqlite AGENTS.md Am I a Bad Friend? CSS vs. JavaScript • Josh W. Comeau Erlang Ecosystem Foundation - Supporting the BEAM community A brief note about slot access cost in Common Lisp Keyboard latency probe Rethinking the GNOME clipboard issues Back to the Building Blocks’ Building Blocks Tech Notes: Theseus: translating win32 to wasm Fast is better than slow Content-addressed Rust builds (or, what kache actually caches) Intent to Prototype: Embedding API Canada’s Bill C-22 and the security cost of collecting more data 5 PostgreSQL locking behaviors that trip people up okmij.org Stop advertising in your commits! | AksDev GitHub - mplsllc/macsurf: A modern web browser for Classic Mac OS 9 PowerPC. Real CSS3, ES5 JavaScript, native HTTPS — built with CodeWarrior on the Carbon API. Introducing DoomBench - Can Your Data Stack Run DOOM? What are some of your favourite developer tools? Building a Scalable Ingestion Pipeline with Temporal (Part 1) Converting shallow Git bundles into normal repositories Are you a member of any professional associations? What is a harmonic? An interactive comic about additive synthesis How Virtual Tables Work in the Itanium C++ ABI Using SwiftUI to Build a Mac-assed App in 2026 Rust (and Slint) on a jailbroken Kindle. ~jack/lambda-on-lambda - Serverless Haskell on AWS - sourcehut git Human proof for FOSS contributions Extremely simple internet radio controlled via IRC Announcing BABLR Splitting Konsole views from Helix to run tools | AksDev GitHub - yugr/rust-slides Serving files over HTTP three ways: synchronous, epoll, and io_uring update docs with information about building with build.py (#979) · astral-sh/python-build-standalone@c9c40c5 A Simple Makefile Tutorial On C extensions, portability, and alternative compilers Switching to Colemak | Pedro Alves Just How Bad Was The Intel IAPX432? Nix's Substituter List Is Not a Routing Table Accelerating copy_if using SIMD Lambda on Lambda: Serverless Haskell on AWS | Blog Announcing feed-repeat v1.0 Scaling Akvorado BMP RIB with sharding EYG news: A host of CLI improvements, new guides and new effects The social contract of writing JS Crossword C array types are weird; and related topics Flatpak will depend on systemd – OSnews Migrating from Go to Rust | corrode Rust Consulting A portentous reunion Vivado Licensing Options How my minimal, memory-safe Go rsync steers clear of vulnerabilities the entropy layer of a wavelet codec, on its own GitHub - nferhat/fht-compositor: A dynamic tiling Wayland compositor. Debian SE Linux and PinTheft Does bulk memmove speed up std::remove_if? (No.) 声明式部分更新 | Blog | Chrome for Developers Fully in-browser container builds Dianne Skoll's Web Site - Remind The Architecture of Open Source Applications (Volume 1)Berkeley DB Pardon MIE? - ironPeak Blog “Long-Term Support” doesn’t mean what you think Jira IS Turing-Complete May I recommend thinking of Emacs as your Fortress of Solitude hershey Floodgap Gopher-HTTP gateway gopher://thelambdalab.xyz/1cuneiforth/ HP QuickWeb, Singular And Pointless That one time I used Go panics for flow control A new suite of modern tools coming for editing and publishing RFCs From the Tabletop… The Digital Antiquarian Building a Host-Tuned GCC to Make GCC Compile Faster Are we self-sovereign PKI yet? Claw Patrol: an open-source security firewall for agents | Deno Revised^7 Report on Scheme, Large: Procedural Fascicle Draft is now public A Network Allow-List Won't Stop Exfiltration — André Graf From AFSK to Goertzel – µArt.cz Software For My New Home Server Introducing Neptune: Direct3D virtualization for QEMU AI Agent Bankrupted Their Operator While Trying to Scan DN42 - Lan Tian @ Blog mimalloc: A new, high-performance, scalable memory allocator for the modern era Making wl_shm fast The Soul of Maintaining a New Machine - Third Draft | Books in Progress What is Git made of?
zlib-rs in Firefox - Trifecta Tech Foundation
trifectatech · 2026-06-16 · via Lobsters

As of 151.0.0, Firefox uses zlib-rs for gzip (de)compression. This is very exciting, and has both performance and safety advantages.

We first started talking to Mozilla engineers in summer 2024, and it took 2 years to actually get zlib-rs into production. What took us so long?

Integrating zlib-rs into the Firefox codebase

Switching to zlib-rs is not entirely trivial: we present zlib-rs as a drop-in compatible replacement, but there are some asterisks to this claim. We change the algorithms that are used at the different compression levels (in a way that is consistent with zlib-ng, but inconsistent with stock zlib), so the exact output bytes and output length can change slightly.

The Firefox test suite tested for the exact output bytes in some cases, and for the (rough) output length in more. This is a good fail safe against messing up the compression configuration, but now these tests all needed to be updated.

Firefox also adds a prefix to all symbols: instead of inflate it uses MOZ_Z_inflate to prevent symbol clashes. We've long supported prefixing the symbol name in various ways, so getting this to work was just a matter of configuration.

So some work was needed, but the changes were straightforward. All seemed well, until...

Intel CPU bug

We started seeing crashes. The logs showed that a bounds check had failed that logically couldn't fail. Of course, we're lucky that we even got a bounds check failure; in C you'd just get silent data corruption.

We could not reproduce the issue locally, and as more reports came in, a pattern started to emerge: our implementation triggered the infamous Intel Raptor Lake CPU bug.

This generation of CPUs is plagued by instability and degradation issues. Something in our code was prone to triggering these issues, but of course we had no idea what, or even how to track it down.

Eventually Fabian Giesen wrote "Oodle 2.9.14 and Intel 13th/14th gen CPUs", which identifies the problem as a particular instruction used in writing the result of Huffman coding to memory. Zlib also uses Huffman coding, and zlib-rs turned out to also use the offending instruction.

Still, finding and shipping the solution in Firefox is not a quick fix. This May, shortly after the 151 release, Mozilla engineers shipped the patch, "After a year, Firefox finally stops crashing on Intel's Raptor Lake CPUs — Mozilla releases new version patch critical flaw on Intel 13th-gen and 14th-gen CPUs".

Fixing the bug

Once you know what to look for, fixing the issue is reasonably straightforward. We had this function:

https://godbolt.org/z/GjfYdPe3x

pub fn push_dist(&mut self, dist: u16, len: u8) {
    let buf = &mut self.buf.as_mut_slice()[self.filled..][..3];
    let [dist1, dist2] = dist.to_le_bytes();

    buf[0] = dist1;
    buf[1] = dist2;
    buf[2] = len;

    self.filled += 3;
}

This code is dead simple: we assign three byte values to consecutive indices of an array. But the assembly for this function (with LLVM 22) has this move from ch to memory, which is bits 8-15 of the RCX register:

mov     byte ptr [rsi + rdi + 1], ch

Due to the hardware bug, occasionally this instruction will actually write bits 0-7 instead, causing the crashes we were seeing.

To work around LLVM emitting this particular instruction, we use a tiny bit of unsafe code (LLVM is clever, so this was the simplest way we've found to have it generate the right thing):

pub fn push_dist(&mut self, dist: u16, len: u8) {
    let buf = &mut self.buf.as_mut_slice()[self.filled..][..3];

    let bytes = dist.to_le_bytes();
    unsafe { buf.as_mut_ptr().cast::<[u8; 2]>().write_unaligned(bytes) }
    buf[2] = len;

    self.filled += 3;
}

The fix in Firefox by Mike Hommey is here. The patch has been upstreamed into zlib-rs and we will continue to carry that patch for the foreseeable future: it's a marginal amount of unsafe that is easily vetted. These are the sacrifices we make to run reliably on a variety of platforms.

It turns out that LLVM 23 no longer emits the offending instruction, although I believe that is serendipitous and not deliberate. When we bump our MSRV to a version that requires LLVM 23 (e.g. for custom allocators and c-variadic functions) we can drop this workaround.

Results

So why go through all of this trouble? Because zlib-rs is faster. Much faster. Especially on linux x86_64 the speedup is almost silly. These benchmarks from zlib-py compare stock zlib versus zlib-rs:

-------------------------------------------------------------------------
  ONE-SHOT DECOMPRESSION
-------------------------------------------------------------------------
Benchmark                   CPython zlib        zlib_py          Speedup
-------------------------------------------------------------------------
decompress   1 KB  level=1        7.1 us         1.3 us     5.66x faster
decompress   1 KB  level=6        7.0 us         2.1 us     3.34x faster
decompress   1 KB  level=9        7.0 us         2.1 us     3.33x faster
decompress  64 KB  level=1      219.4 us         6.8 us    32.50x faster
decompress  64 KB  level=6      218.6 us         7.6 us    28.70x faster
decompress  64 KB  level=9      217.9 us         7.9 us    27.53x faster
decompress   1 MB  level=1       3.41 ms       128.0 us    26.61x faster
decompress   1 MB  level=6       3.42 ms       125.2 us    27.30x faster
decompress   1 MB  level=9       3.33 ms       134.8 us    24.71x faster
decompress  10 MB  level=1      33.95 ms        1.74 ms    19.50x faster
decompress  10 MB  level=6      33.94 ms        1.68 ms    20.16x faster
decompress  10 MB  level=9      33.80 ms        1.74 ms    19.42x faster
-------------------------------------------------------------------------
  STREAMING DECOMPRESSION
-------------------------------------------------------------------------
Benchmark                    CPython zlib        zlib_py          Speedup
-------------------------------------------------------------------------
stream decompress   1 KB  L6       7.3 us         2.7 us     2.74x faster
stream decompress  64 KB  L6     221.3 us        22.7 us     9.75x faster
stream decompress   1 MB  L6      3.36 ms       309.0 us    10.86x faster
stream decompress  10 MB  L6     33.71 ms        3.79 ms     8.89x faster

Compression is also faster, but harder to compare because the difference in compression ratio.

Via these benchmarks we noticed that the speedup is smaller on aarch64 systems, especially those running macOS. It turns out that Apple provides a more optimized zlib dynamic library, which uses inline assembly for some of the most performance-sensitive parts. This made us realize that there are some optimizations that we missed before, and we're now in the process of integrating them.

Conclusion

Upgrading to zlib-rs should be straightforward, but in this case we encountered the toughest bug we've seen so far. With CPU bugs, there isn't much to go on, and our standard debugging tools are of little value. We spent months not really sure what to do, but now we have a workaround and can finally move forward.

We're very excited about zlib-rs now serving many more users. We want to thank Mozilla, and specifically Mike Hommey and Gabriele Svelto, for the integration work and tracking down and fixing the CPU bug.