惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Spread Privacy
Spread Privacy
S
Schneier on Security
G
GRAHAM CLULEY
AWS News Blog
AWS News Blog
Cisco Talos Blog
Cisco Talos Blog
The Hacker News
The Hacker News
T
The Exploit Database - CXSecurity.com
P
Proofpoint News Feed
L
LINUX DO - 热门话题
C
CXSECURITY Database RSS Feed - CXSecurity.com
Security Latest
Security Latest
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Google DeepMind News
Google DeepMind News
H
Hacker News: Front Page
PCI Perspectives
PCI Perspectives
T
Tenable Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
N
Netflix TechBlog - Medium
Application and Cybersecurity Blog
Application and Cybersecurity Blog
腾讯CDC
A
Arctic Wolf
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Project Zero
Project Zero
NISL@THU
NISL@THU
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
V
Vulnerabilities – Threatpost
Cyberwarzone
Cyberwarzone
I
Intezer
Apple Machine Learning Research
Apple Machine Learning Research
T
Threat Research - Cisco Blogs
爱范儿
爱范儿
Webroot Blog
Webroot Blog
Forbes - Security
Forbes - Security
The Cloudflare Blog
T
Tailwind CSS Blog
C
CERT Recently Published Vulnerability Notes
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
大猫的无限游戏
大猫的无限游戏
Microsoft Azure Blog
Microsoft Azure Blog
Know Your Adversary
Know Your Adversary
云风的 BLOG
云风的 BLOG
B
Blog
The Register - Security
The Register - Security
T
Threatpost
C
Cybersecurity and Infrastructure Security Agency CISA
V2EX - 技术
V2EX - 技术
P
Privacy & Cybersecurity Law Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org

Lobsters

CIFSwitch: a non-universal Linux local root vulnerability RIPE NCC session fixation: poaching logins with an Atlas probe GNOME 2.20 but its Web Components Agentic Search for Context Engineering – Leonie Monigatti Garnix is shutting down [not OC] akashina.tngl.sh/jjc Concerning Emacs (and Jazz) Nitpicking the shell history scene in ‘Tron: Legacy’ What's cooking on SourceHut? Q2 2026 The tenth OpenPGP email summit Package managers that package package managers Clojure on Fennel part three: parsing WordPress at 23 Finding Miscompiles for Fun, Not Profit GitHub - creusot-rs/creusot: Creusot helps you prove your Rust code is correct. Announcing Rust 1.96.0 | Rust Blog A Love Letter to Neovim sqlite AGENTS.md Am I a Bad Friend? CSS vs. JavaScript • Josh W. Comeau Erlang Ecosystem Foundation - Supporting the BEAM community A brief note about slot access cost in Common Lisp Keyboard latency probe Rethinking the GNOME clipboard issues Back to the Building Blocks’ Building Blocks Tech Notes: Theseus: translating win32 to wasm Fast is better than slow Content-addressed Rust builds (or, what kache actually caches) Intent to Prototype: Embedding API Canada’s Bill C-22 and the security cost of collecting more data 5 PostgreSQL locking behaviors that trip people up okmij.org Stop advertising in your commits! | AksDev GitHub - mplsllc/macsurf: A modern web browser for Classic Mac OS 9 PowerPC. Real CSS3, ES5 JavaScript, native HTTPS — built with CodeWarrior on the Carbon API. Introducing DoomBench - Can Your Data Stack Run DOOM? What are some of your favourite developer tools? Building a Scalable Ingestion Pipeline with Temporal (Part 1) Converting shallow Git bundles into normal repositories Are you a member of any professional associations? What is a harmonic? An interactive comic about additive synthesis How Virtual Tables Work in the Itanium C++ ABI Using SwiftUI to Build a Mac-assed App in 2026 Rust (and Slint) on a jailbroken Kindle. ~jack/lambda-on-lambda - Serverless Haskell on AWS - sourcehut git Human proof for FOSS contributions Extremely simple internet radio controlled via IRC Announcing BABLR Splitting Konsole views from Helix to run tools | AksDev GitHub - yugr/rust-slides Serving files over HTTP three ways: synchronous, epoll, and io_uring update docs with information about building with build.py (#979) · astral-sh/python-build-standalone@c9c40c5 A Simple Makefile Tutorial On C extensions, portability, and alternative compilers Switching to Colemak | Pedro Alves Just How Bad Was The Intel IAPX432? Nix's Substituter List Is Not a Routing Table Accelerating copy_if using SIMD Lambda on Lambda: Serverless Haskell on AWS | Blog Announcing feed-repeat v1.0 Scaling Akvorado BMP RIB with sharding EYG news: A host of CLI improvements, new guides and new effects The social contract of writing JS Crossword C array types are weird; and related topics Flatpak will depend on systemd – OSnews Migrating from Go to Rust | corrode Rust Consulting A portentous reunion Vivado Licensing Options How my minimal, memory-safe Go rsync steers clear of vulnerabilities the entropy layer of a wavelet codec, on its own GitHub - nferhat/fht-compositor: A dynamic tiling Wayland compositor. Debian SE Linux and PinTheft Does bulk memmove speed up std::remove_if? (No.) 声明式部分更新 | Blog | Chrome for Developers Fully in-browser container builds Dianne Skoll's Web Site - Remind The Architecture of Open Source Applications (Volume 1)Berkeley DB Pardon MIE? - ironPeak Blog “Long-Term Support” doesn’t mean what you think Jira IS Turing-Complete May I recommend thinking of Emacs as your Fortress of Solitude hershey Floodgap Gopher-HTTP gateway gopher://thelambdalab.xyz/1cuneiforth/ HP QuickWeb, Singular And Pointless That one time I used Go panics for flow control A new suite of modern tools coming for editing and publishing RFCs From the Tabletop… The Digital Antiquarian Building a Host-Tuned GCC to Make GCC Compile Faster Are we self-sovereign PKI yet? Claw Patrol: an open-source security firewall for agents | Deno Revised^7 Report on Scheme, Large: Procedural Fascicle Draft is now public A Network Allow-List Won't Stop Exfiltration — André Graf From AFSK to Goertzel – µArt.cz Software For My New Home Server Introducing Neptune: Direct3D virtualization for QEMU AI Agent Bankrupted Their Operator While Trying to Scan DN42 - Lan Tian @ Blog mimalloc: A new, high-performance, scalable memory allocator for the modern era Making wl_shm fast The Soul of Maintaining a New Machine - Third Draft | Books in Progress What is Git made of?
invlpg – Premature Optimization is Fun Sometimes
invlpg · 2026-06-08 · via Lobsters

A colleague of mine was recently discussing a connectivity monitoring system he is working on with me. It’s nothing fancy, just sending ICMP Echo Requests to a couple of different servers, and monitoring latency and dropped packet averages over 1-minute, 5-minute, and 15-minute periods. Up came the topic of how this data should be stored, the natural thought was a 512 entry ring buffer, containing entries like the following:

struct ping_timestamp {
    uint64_t sent_ns;       // when the packed was sent (unit: ns)
    uint64_t received_ns;   // when the packet was received (unit: ns)
    in_addr_t source_addr;  // source address
    uint16_t seq_no;        // echo request sequence number
    bool received;          // has the request been received?
};

And backed by the following array

struct ping_timestamp pings_rb[512];

// ...

printf("%zu\n", sizeof(pings_rb));

// 12288

12 KiB. Pretty wasteful, right? We can certainly do better. Do we need to keep fields for both sent and received? What we’re really interested is the latency. We need to know when a packet was sent, only up until we know when it was received, at that point, the data we want to keep is received - sent, so why don’t we make it a tagged union?

struct ping_timestamp_2 {
    union {
        uint64_t sent_ts;       // unit: 100μs
        uint64_t elapsed_ts;    // unit: 100μs
    };
    in_addr_t source_addr;
    uint16_t seq_no;
    bool received;
};

// ...

printf("%zu\n", sizeof(pings_rb));

// 8192

Not bad, we’ve shaved off an entire page. We can still do better.

Nanosecond precision? In our case, ping times are measured in the tens or hundreds, or even thousands of milliseconds. We don’t need to keep all of those extra bits around. If we change the unit from nanosecond to 100 microsecond increments (0.1ms), then 43-bits is sufficient for us to keep track of pings for up to 20-years1. 20-years is a bit excessive still, but it doesn’t hurt to be at least a little bit future proof.

And received? 8-bit for a true/false value seems altogether too much. The answer: bitfields.

struct ping_timestamp_3 {
    uint64_t sent_or_elapsed_ts: 43;
    uint64_t received: 1;
    uint64_t seq_no: 16;
    in_addr_t source_addr;
};

// ...

printf("%zu\n", sizeof(pings_rb));

// 8192

Wait, what? Why haven’t we saved any space?

The answer is struct padding. The layout of ping_timestamp_2 looks like this:

     0       8      16      24      32      40      48      56      64...
     +-------+-------+-------+-------+-------+-------+-------+-------+
     |                         sent/elapsed                          |
     +-------+-------+-------+-------+-------+-------+-------+-------+
 ...64      72      80      88      96      104     112     120     128
     +-------+-------+-------+-------+-------+-------+-------+-------+
     |         source_address        |    seq_no     | recv? |  pad  |
     +---------------------------------------------------------------+

Where the padding byte at the end is to ensure alignment requirements. ping_timestamp_3 on the other end, looks like this:

     0       8      16      24      32      40  R?  48      56      64...
     +-------+-------+-------+-------+-------+-------+-------+-------+
     |                 sent/elapsed            ||     seq_no     | P |
     +-------+-------+-------+-------+-------+-------+-------+-------+
 ...64      72      80      88      96      104     112     120     128
     +-------+-------+-------+-------+-------+-------+-------+-------+
     |         source_address        |            padding            |
     +---------------------------------------------------------------+

So our optimization there didn’t actually save any space. We’re wasting 36-bits of padding. Is there any way we can somehow do better?

We keep track of the source address due to frequent changes while our product is in operation (on a mobile data network). When the address changes, we also reset the sequence number for reasons that aren’t relevant to the current topic. We have seen, in the past, packets with differing source addreses but identical sequence numbers due to be processed by our application at the same time (the joys of asynchronous programming), so the source address serves to disambiguate these changes.

But there’s another way to disambiguate.

An ICMP echo request has a 16-bit long identifier field to allow applications to identify which echo request packets were sent by them. Its value is completely arbitray. On Linux iputils ping sets it to getpid() & 0xFFFF; on OpenBSD a random number is used instead.

Although it’s 16-bits long, we don’t actually need to use the full 16-bits. There’s 4 free bits left in the first 8-bytes of our ping_timestamp_3. Our thought was to use a rolling 4-bit counter, that is increased whenever our source address changes (this is monitored elsewhere in the application), allowing us to uniquely identify2 which source address the packet came from.

Our final struct looks like this:

struct ping_timestamp {
    uint64_t elapsed_or_sent_ts : 43;
    uint64_t received : 1;
    uint64_t counter: 4;
    uint64_t seq_no: 16;
};

// ...

printf("%zu\n", sizeof(pings_rb));

// 4096

Much better. A whole 8-kilobytes of savings, and down to a single page of data. You may have noticed that I have changed the order of the fields slightly. This is to line seq_no up on a 16-bit boundary, so that loading it is a single ldrh instruction rather than require a shift. Similarly, reading from elapsed_or_sent_ts only requires a mask.

In the end, this was a completely pointless exercise. Our application isn’t remotely memory constrained.

But it was fun.

Addendum 2025-06-21

I realized there’s a way to “optimize” this slightly further. By switching the order of the received/counter fields, accessing the received bit only requires a shift instruction rather than a shift and a mask:

struct ping_timestamp {
    uint64_t elapsed_or_sent_ts : 43;
    uint64_t counter: 4;
    uint64_t received : 1;
    uint64_t seq_no: 16;
};

Addendum 2025-06-22

There’s a slight “issue”3 with the above code: received is now much cheaper to access, at the expense of counter needing to mask out the received bit.

But we can fix this! We only ever read counter when received is true, i.e. 1. If received were zero, and we could tell the compiler to assume that it were zero, then no mask would be necessary.

The solution? Flip the meaning of the received bit.

struct ping_timestamp {
    uint64_t elapsed_or_sent_ts : 43;
    uint64_t counter: 4;
    uint64_t not_received : 1;
    uint64_t seq_no: 16;
};

Now, if the read of counter only ever happens inside of a conditional that checks if not_received is zero, then the compiler is able to completely elide the mask.


  1. The timestamps are taken from the linux kernel’s monotonic clock, which measures time elapsed since boot. If we were measuring from the Unix epoch, then we would need 51-bits.↩︎

  2. As long the IP address doesn’t change more than 16-times in the period we’re monitoring, which is not something we have ever seen.↩︎

  3. I’m using that term loosely, we haven’t even benchmarked anything.↩︎