惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Vulnerabilities – Threatpost
Know Your Adversary
Know Your Adversary
C
Cyber Attacks, Cyber Crime and Cyber Security
S
Secure Thoughts
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
宝玉的分享
宝玉的分享
Spread Privacy
Spread Privacy
AWS News Blog
AWS News Blog
D
Docker
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
TaoSecurity Blog
TaoSecurity Blog
博客园 - 三生石上(FineUI控件)
Apple Machine Learning Research
Apple Machine Learning Research
Cyberwarzone
Cyberwarzone
V
V2EX
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
WordPress大学
WordPress大学
P
Palo Alto Networks Blog
H
Heimdal Security Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 叶小钗
N
News and Events Feed by Topic
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Simon Willison's Weblog
Simon Willison's Weblog
Project Zero
Project Zero
Martin Fowler
Martin Fowler
大猫的无限游戏
大猫的无限游戏
D
DataBreaches.Net
Engineering at Meta
Engineering at Meta
S
Schneier on Security
Google DeepMind News
Google DeepMind News
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Hugging Face - Blog
Hugging Face - Blog
P
Proofpoint News Feed
S
SegmentFault 最新的问题
Hacker News: Ask HN
Hacker News: Ask HN
小众软件
小众软件
博客园 - 聂微东
S
Security Affairs
T
Tor Project blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
T
Threat Research - Cisco Blogs
T
Threatpost
博客园 - 【当耐特】
L
LINUX DO - 热门话题
G
Google Developers Blog
P
Privacy & Cybersecurity Law Blog
A
About on SuperTechFans
F
Fortinet All Blogs

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
C++ and Microarchitecture Nuances
Sami Al-Jamal · 2026-06-18 · via DEV Community

C++ source code is written in order. That does not mean the processor executes it in order.

This is the first correction. It is also the one many performance discussions manage to avoid.

Modern high-performance cores use out-of-order execution. They accept a sequential instruction stream, break it into internal operations, rename registers, place work into scheduling structures, execute ready operations early, and then retire the results in program order. The machine preserves the visible behavior of sequential execution. Internally, it is not taking attendance line by line.

For ordinary software, this is mostly invisible. For C++ intended to run in tens of nanoseconds, it is not invisible. At that scale, performance is not just about the number of instructions. It is about whether those instructions can be scheduled in parallel or whether the program quietly built a dependency chain and then acted surprised.

The processor is a dependency scheduler

Out-of-order execution exists because in-order pipelines waste time. If an older instruction stalls, an in-order processor must often wait even if later instructions are independent and ready. That is a poor use of hardware. The chip has execution units available. The instruction stream has more work.

Dynamic scheduling fixes part of this problem. The processor tracks which operations have their inputs ready. When an operation is ready and an execution unit is available, it can issue. Older operations may still be waiting. Later operations may run first. The final architectural state is still committed in order, so the program behaves correctly.

Tomasulo’s algorithm is the classic model for this idea. It used reservation stations and register renaming to allow instructions to execute when their operands became available rather than strictly when they appeared in the original program (Tomasulo, 1967). Later superscalar processors extended the same general approach with speculation and reorder buffers, but the central idea stayed the same: execute according to readiness, not textual order (Hennessy & Patterson, 2019).

That is the point relevant to C++: the processor is not primarily reading the program as a list. It is resolving a graph.

Nodes are operations. Edges are dependencies.

No edge, possible overlap.
Real edge, forced order.

True dependencies are the hard limit

Out-of-order execution cannot violate true data dependencies.

If instruction B needs a value produced by instruction A, B waits for A. No scheduler can change that. No amount of confidence in “modern CPUs are smart” changes that either.

Consider this reduction:

std::uint64_t s = 0;

for (std::size_t i = 0; i < n; ++i) {
    s += data[i];
}

The loop looks harmless. It also creates a loop-carried dependency. Each update to s depends on the previous value of s.

Conceptually:

s1 = s0 + data[0]
s2 = s1 + data[1]
s3 = s2 + data[2]
s4 = s3 + data[3]

The processor may overlap some surrounding work, and it may have several loads in flight, but the additions themselves form a chain. The next addition needs the previous result.

A better version exposes independent chains:

std::uint64_t s0 = 0;
std::uint64_t s1 = 0;
std::uint64_t s2 = 0;
std::uint64_t s3 = 0;

for (std::size_t i = 0; i < n; i += 4) {
    s0 += data[i + 0];
    s1 += data[i + 1];
    s2 += data[i + 2];
    s3 += data[i + 3];
}

std::uint64_t s = s0 + s1 + s2 + s3;

Now there are four shorter dependency chains instead of one long chain. The out-of-order core has more ready work. This does not make the program faster by aesthetic force. It makes it faster because the dependency graph improved.

This is the useful mental model. Optimizing for out-of-order execution means reducing critical-path length and increasing available independent work.

The processor cannot schedule independence that the program does not expose.

False dependencies are bookkeeping problems

Not all apparent dependencies are real.

A read-after-write dependency is real. If a later operation needs the result of an earlier operation, it must wait.

A write-after-read or write-after-write dependency can be false. These often come from reusing the same architectural register name, not from the actual values depending on each other.

Out-of-order processors handle this with register renaming. The instruction set exposes a limited set of architectural registers. Internally, the processor maps those architectural registers onto a larger pool of physical registers. This lets the machine separate unrelated values that happen to use the same architectural name.

For example, source code may produce machine instructions that appear to reuse a register. Internally, the processor can assign different physical registers to different live values. The name is reused. The storage is not necessarily reused.

That matters because it prevents false dependencies from blocking execution. The processor can see that two writes to the same architectural register do not need to serialize if no actual value relationship exists between them.

This is also why reading assembly is useful but incomplete. Assembly shows the architectural instruction stream. It does not show the physical register mappings, scheduling queues, issue timing, or reorder-buffer behavior. Assembly is closer to the machine than C++. It is still not the machine.

The reorder buffer keeps the lie consistent

If instructions can execute out of order, the processor needs a way to preserve precise program behavior. This is the job of the reorder buffer, or ROB.

Instructions may finish execution out of order. They do not usually retire out of order. The ROB holds results until each instruction is safe to commit in the original program order. If speculation was wrong, the processor can discard the speculative work and restore a correct state.

This separation matters:

execute: may happen out of order
retire: happens in program order

That is how the processor gets performance without giving up the language’s sequential behavior.

For a C++ programmer, this explains a common confusion. The processor may execute later independent operations before earlier blocked operations, but the program still appears to obey the abstract machine rules, subject to the usual caveat that undefined behavior gives the compiler and hardware no meaningful obligation. That caveat is not a footnote. It is where many “low-level tricks” go to die.

Out-of-order execution rewards boring independence

The easiest way to help out-of-order execution is to provide independent operations.

This often means splitting state.

Bad:

state = mix(state, x0);
state = mix(state, x1);
state = mix(state, x2);
state = mix(state, x3);

Every line depends on the previous value of state. The processor sees a chain.

Better, if the algorithm allows it:

s0 = mix(s0, x0);
s1 = mix(s1, x1);
s2 = mix(s2, x2);
s3 = mix(s3, x3);

state = combine(s0, s1, s2, s3);

Now the processor sees several independent chains. The final combination still has to happen, but much of the work can be scheduled earlier.

This pattern appears everywhere: checksums, reductions, counters, hash-like computations, parsing loops, pricing loops, and numeric kernels. The exact transformation depends on the algorithm. The principle does not.

Out-of-order execution does not reward clever-looking code. It rewards code that gives the scheduler options.

Pointer chasing defeats the scheduler

A dependent load chain is especially restrictive.

Node* p = head;

while (p != nullptr) {
    sum += p->value;
    p = p->next;
}

The next address depends on the current load. The processor cannot issue the load of p->next->next until it has loaded p->next. The chain is serial.

This is not mainly a lecture about cache locality. Cache locality matters, but the out-of-order point is narrower: the address of future work is not available yet. The scheduler cannot issue an operation whose address has not been computed.

That is why pointer-heavy structures often behave poorly in nanosecond-scale code. The processor has resources, but the program gives it one dependent step at a time. A large out-of-order window helps only when there is other independent work nearby. If the whole loop is a dependent chain, the window fills with waiting.

The code is then not CPU-bound in the useful sense. It is dependency-bound.

Function calls and abstraction can hide independence

Out-of-order execution operates on the instruction stream produced by the compiler. It does not see source-level intent. If useful independence is hidden behind function calls, aliasing, virtual dispatch, or opaque control flow, the compiler may fail to expose it in the emitted code.

Inlining can matter because it gives the optimizer more context. With more context, the compiler may remove redundant work, keep values in registers, reorder independent operations, or unroll a loop enough to expose multiple independent chains. Pikus emphasizes this interaction between C++ structure, compiler optimization, and CPU utilization: high-performance C++ depends on giving the compiler and hardware enough information to use available resources effectively (Pikus, 2021).

This does not mean every function call is bad. That would be a convenient rule, and therefore suspicious. The actual question is whether the final instruction stream exposes enough independent work for the core to schedule.

The source is not the object of measurement. The emitted machine code is.

Latency and throughput are different questions

Out-of-order execution also forces a distinction between latency and throughput.

An operation’s latency is how long its result takes to become available to a dependent operation.

An operation’s throughput is how many such operations can be started per cycle when enough independent work exists.

This distinction matters constantly.

If a multiply has a latency of several cycles but the processor can start one multiply every cycle, then one dependency chain pays the full latency repeatedly. Four independent chains can approach throughput limits instead. Same instruction. Different dependency graph. Different performance.

This is why the earlier accumulator example matters. It does not make addition itself faster. It changes the loop from latency-limited toward throughput-limited.

Low-latency C++ often consists of this kind of work: find the critical dependency chain, shorten it, split it, or move independent work across it. Not glamorous. Usually effective. A rare combination.

What to measure

A benchmark that reports only wall-clock time is often not enough. It may show that one version is faster, but not why.

For out-of-order behavior, useful questions include:

  • Is the loop limited by one dependency chain?
  • Did unrolling expose independent operations?
  • Did the compiler actually emit multiple accumulators?
  • Are instructions waiting on operands or execution resources?
  • Did a small source change alter the generated dependency graph?

Tools such as Compiler Explorer, objdump, LLVM-MCA, perf, and Intel VTune help answer these questions. They do not replace thinking. They reduce the amount of fiction in the thinking.

A simple timing result can say “faster.” The assembly and counters explain whether the speed came from less work, better scheduling, fewer stalls, or a shorter critical path.

For nanosecond-scale code, that difference matters.

The practical conclusion

Out-of-order execution is not a general blessing applied to slow code. It is a specific hardware strategy for finding ready work inside an instruction stream.

It can hide stalls when independent work exists.
It can remove false register dependencies through renaming.
It can execute operations before older stalled instructions.
It can preserve sequential behavior through in-order retirement.
It cannot violate true dependencies.
It cannot issue future work whose inputs are unknown.
It cannot rescue a program that presents one long chain and no alternatives.

The C++ programmer writing low-latency code should therefore ask a different question. Not “how many lines did I write?” Not even only “how many instructions did the compiler emit?”

The better question is: what dependency graph did I give the processor?

That graph is what the out-of-order core actually schedules. The source code is merely how the problem was submitted.

References

Hennessy, J. L., & Patterson, D. A. (2019). Computer Architecture: A Quantitative Approach (6th ed.). Morgan Kaufmann.

Kessler, R. E. (1999). The Alpha 21264 microprocessor. IEEE Micro, 19(2), 24–36.

Pikus, F. G. (2021). The Art of Writing Efficient Programs: An Advanced Programmer’s Guide to Efficient Hardware Utilization and Compiler Optimizations Using C++ Examples. Packt Publishing.

Smith, J. E., & Sohi, G. S. (1995). The microarchitecture of superscalar processors. Proceedings of the IEEE, 83(12), 1609–1624.

Tomasulo, R. M. (1967). An efficient algorithm for exploiting multiple arithmetic units. IBM Journal of Research and Development, 11(1), 25–33.