惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
Spread Privacy
Spread Privacy
D
Docker
Stack Overflow Blog
Stack Overflow Blog
Google DeepMind News
Google DeepMind News
F
Fortinet All Blogs
F
Full Disclosure
美团技术团队
Y
Y Combinator Blog
N
Netflix TechBlog - Medium
Security Latest
Security Latest
C
Check Point Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
罗磊的独立博客
A
Arctic Wolf
S
Schneier on Security
T
Threatpost
C
CERT Recently Published Vulnerability Notes
L
LangChain Blog
博客园 - 叶小钗
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
博客园 - 聂微东
T
The Exploit Database - CXSecurity.com
W
WeLiveSecurity
Engineering at Meta
Engineering at Meta
C
Cybersecurity and Infrastructure Security Agency CISA
GbyAI
GbyAI
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
H
Heimdal Security Blog
L
LINUX DO - 热门话题
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
O
OpenAI News
G
GRAHAM CLULEY
M
MIT News - Artificial intelligence
S
Security @ Cisco Blogs
博客园 - 司徒正美
N
News and Events Feed by Topic
Microsoft Azure Blog
Microsoft Azure Blog
Cisco Talos Blog
Cisco Talos Blog
P
Palo Alto Networks Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Schneier on Security
Schneier on Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
月光博客
月光博客
The Last Watchdog
The Last Watchdog
Apple Machine Learning Research
Apple Machine Learning Research
Microsoft Security Blog
Microsoft Security Blog
C
Cisco Blogs
雷峰网
雷峰网

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Linux Performance Tuning: CPU, Memory, I/O & Network
James Lee · 2026-05-17 · via DEV Community

Linux performance optimization covers four major subsystems:

Subsystem Typical Bottleneck Scenario
CPU Compute-intensive workloads (Nginx, Node.js, math/batch processing)
Memory Database workloads (MySQL) — heavy memory and storage consumption
I/O Disk-bound applications with heavy read/write
Network High-throughput web services

Section 1: CPU Performance

CPU is the most critical subsystem — responsible for all computation. Modern production servers use multi-core CPUs based on SMP (Symmetric Multiprocessing) architecture. In practice, CPU utilization is often below 5%, meaning significant resource waste.

CPU Cache Hierarchy

# lscpu
L1d cache:   32K    ← L1 data cache (static, per-core)
L1i cache:   32K    ← L1 instruction cache (static, per-core)
L2 cache:    256K   ← dynamic, shared
L3 cache:    8192K  ← dynamic, shared across cores

Enter fullscreen mode Exit fullscreen mode

  • L1 cache: static cache, split into data and instruction caches
  • L2 / L3 cache: dynamic cache; L2 is shared between cores

CPU Affinity

In SMP systems, the Linux scheduler may run the same thread on different cores across time slices. Since each core has its own memory space (not shared), this causes cache invalidation — the thread's data must be reloaded into the new core's cache, degrading performance.

CPU affinity pins a process to a specific core, maximizing cache hit rate:

# Pin process 73890 to CPU core 0
taskset -pc 0 73890

Enter fullscreen mode Exit fullscreen mode

NUMA (Non-Uniform Memory Access)

taskset alone doesn't guarantee local memory allocation. For NUMA architectures, use numactl:

NUMA topology:
┌──────────────┐     ┌──────────────┐
│  CPU Node 0  │     │  CPU Node 1  │
│  Local RAM   │     │  Local RAM   │
│  (fast)      │     │  (fast)      │
└──────┬───────┘     └──────┬───────┘
       │  remote access (slower)  │
       └──────────────────────────┘

Enter fullscreen mode Exit fullscreen mode

# View current NUMA configuration
numactl --show

# Bind program to specific NUMA node
numactl --cpunodebind=0 --membind=0 ./myapp

Enter fullscreen mode Exit fullscreen mode

⚠️ Database servers should NOT use NUMA by default. If required, start the DB with numactl --interleave=all to avoid memory hotspots.

CPU Scheduling Policies

Real-time scheduling (priority 1–99, higher = more urgent):

Policy Behavior
SCHED_FIFO Static priority; once running, holds CPU until higher-priority task arrives or it yields
SCHED_RR Round-robin with time slices; expired slice goes to end of queue — fair among equal-priority tasks

General scheduling (priority 100–139, lower number = higher priority):

Policy Behavior
SCHED_OTHER Default; priority determined by nice + counter values. Least recently scheduled gets priority.
SCHED_BATCH For batch processing
SCHED_IDLE For very low priority background tasks
# Adjust process priority with nice (-20 to 19, lower = higher priority)
renice 5 <pid>

# Modify real-time scheduling priority
chrt -r -p 50 <pid>

Enter fullscreen mode Exit fullscreen mode

Context Switches

The Linux kernel treats each core as an independent processor. Each core can run 50–50,000 processes. Each thread gets a time slice; when it expires or is preempted, a context switch occurs.

The more context switches, the heavier the kernel scheduling overhead.

Run Queue

Each CPU has a run queue. A thread is either sleeping (blocked on I/O) or runnable (waiting for CPU time).

load = currently running threads + threads in run queue

Example: 2 cores, 2 running + 4 queued → load = 6

Enter fullscreen mode Exit fullscreen mode

CPU Performance Targets

Healthy CPU metrics:
┌─────────────────────────────────────────┐
│  us (user)    60% – 70%                 │
│  sy (system)  30% – 35%                 │
│  id (idle)     0% –  5%                 │
│  run queue    ≤ 4 per core (ideal)      │
└─────────────────────────────────────────┘

Enter fullscreen mode Exit fullscreen mode

# Monitor with vmstat (1-second intervals, 5 samples)
vmstat 1 5

Enter fullscreen mode Exit fullscreen mode

procs -----------memory---------- ---swap-- -----io---- --system-- -----cpu------
 r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs  us  sy  id  wa
 3  0 1150840 271628 260684 5530984  0    0     2     1    0    0  22   4  73   0
 5  0 1150840 270264 260684 5531032  0    0     0     0  5873 6085 13  13  73   0

Enter fullscreen mode Exit fullscreen mode

High in (interrupts) and cs (context switches) indicate the kernel is constantly switching processes and servicing hardware requests.

Bind Interrupts to a Specific CPU

# Bind interrupt type 19 to CPU core 2
echo 03 > /proc/irq/19/smp_affinity

# Bind TCP interrupts to one CPU (reduces scheduler interference)

Enter fullscreen mode Exit fullscreen mode


Section 2: Memory Performance

Linux uses Virtual Memory Management (VMM) — writes go to filesystem cache in memory first, then flush to disk lazily. This is why available memory appears low after running Linux for a while: most is consumed by cache + buffer.

Optimization goal: reduce disk writes, improve write efficiency.

Dirty Data Flush Policy

# Trigger pdflush when dirty data exceeds 10% of physical memory
echo 10 > /proc/sys/vm/dirty_background_ratio

# Flush dirty data that has been in memory longer than 2000ms
echo 2000 > /proc/sys/vm/dirty_expire_centisecs

Enter fullscreen mode Exit fullscreen mode

⚠️ Tune carefully — these settings have a large impact on I/O performance.

Swap Tuning

When physical memory is insufficient, Linux uses LRU to swap out cold pages to disk, and swap in when needed.

# 0 = prefer physical memory; 100 = aggressively use swap
echo 10 > /proc/sys/vm/swappiness   # recommended for production

Enter fullscreen mode Exit fullscreen mode

Minimize swap usage in production. For Redis, disable overcommit:

echo 0 > /proc/sys/vm/overcommit_memory

Reclaiming Memory

sync
echo 3 > /proc/sys/vm/drop_caches
# 1 = drop page cache (buffers)
# 2 = drop slab cache (cached)
# 3 = drop both

Enter fullscreen mode Exit fullscreen mode

Huge Pages

Large page sizes reduce TLB misses and page table overhead:

cat /proc/meminfo | grep -i huge
# AnonHugePages: transparent huge pages (auto-managed)
# Hugepagesize:  2048 kB (standard huge page size)

# Manually set huge page count
sysctl vm.nr_hugepages=20

Enter fullscreen mode Exit fullscreen mode

32-bit: 4MB huge pages; 64-bit: 2MB huge pages.
Larger pages = less overhead but more internal fragmentation.

Page Faults

MPF (Major Page Fault):  data not in cache → read from disk (expensive)
MnPF (Minor Page Fault): data found in buffer cache → no disk I/O (cheap)

Enter fullscreen mode Exit fullscreen mode

# First run: mostly MPF (cold cache)
/usr/bin/time -v ./myapp

# Second run: mostly MnPF (warm cache)
/usr/bin/time -v ./myapp

Enter fullscreen mode Exit fullscreen mode

The File Buffer Cache continuously grows to reduce MPF and increase MnPF, until the kernel needs to reclaim memory for other processes. Low free memory ≠ memory pressure — Linux intentionally uses free memory for caching.


Section 3: Disk I/O Performance

The I/O subsystem is typically the slowest part of a Linux system — both due to physical distance from the CPU and mechanical/electrical constraints. Minimize disk I/O wherever possible.

I/O Scheduler

cat /sys/block/sda/queue/scheduler
# noop  anticipatory  deadline  [cfq]

Enter fullscreen mode Exit fullscreen mode

Scheduler Description Best For
CFQ (default) Completely Fair Queuing; up to 8 requests per time slice; idles waiting for more I/O from same process General workloads
Deadline Every request must be served before a deadline Databases, latency-sensitive
noop No scheduling; FIFO order SSDs, VMs
anticipatory Deprecated; good for write-heavy/read-light Legacy

I/O Priority

# Set process I/O priority (1=realtime, 2=best-effort, 3=idle)
ionice -c1 -p <pid>   # highest I/O priority for pid

Enter fullscreen mode Exit fullscreen mode

Page Size & Block Size

Linux kernel accesses disk I/O in pages (typically 4KB):

/usr/bin/time -v date   # shows page size info

Enter fullscreen mode Exit fullscreen mode

Tune page size and block size based on your application's I/O pattern (large sequential vs. small random).

DMA (Direct Memory Access)

DMA allows hardware to transfer data directly to/from memory without CPU involvement:

Without DMA: disk → CPU → memory  (CPU occupied during transfer)
With DMA:    disk ──────▶ memory  (CPU free during transfer)

Enter fullscreen mode Exit fullscreen mode

DMA transfer lifecycle: Request → Acknowledge → Transfer → Complete

Writing Data Back to Disk

# Force immediate flush
fsync()   # per-file
sync()    # system-wide

# pdflush runs periodically if not explicitly called

Enter fullscreen mode Exit fullscreen mode

Useful I/O Monitoring Commands

iotop      # per-process I/O usage
lsof       # list all open files and file descriptors
iostat     # disk I/O statistics
vmstat     # combined system stats including I/O

Enter fullscreen mode Exit fullscreen mode


Section 4: Network Performance

For web applications, network performance is critical. Potential bottlenecks include: application response time, Linux network subsystem, NIC, and bandwidth.

NIC Settings

# Check if NIC is in full-duplex mode
ethtool eth0

# Increase MTU for high-bandwidth (≥1Gbps) networks
ifconfig eth0 mtu 9000 up   # jumbo frames

Enter fullscreen mode Exit fullscreen mode

TCP Buffer Tuning

# TCP read buffer: min / initial / max (bytes)
sysctl -w net.ipv4.tcp_rmem="4096 87380 8388608"

# TCP write buffer: min / initial / max (bytes)
sysctl -w net.ipv4.tcp_wmem="4096 87380 8388608"

Enter fullscreen mode Exit fullscreen mode

TCP Window Scaling

# Disable window scaling and set fixed window size
# (similar to setting JVM -Xms = -Xmx for predictability)
sysctl -w net.ipv4.tcp_window_scaling=0

Enter fullscreen mode Exit fullscreen mode

TCP Connection Reuse

Reusing TIME_WAIT connections avoids the full 3-way handshake overhead — significant performance gain for web servers:

sysctl -w net.ipv4.tcp_tw_reuse=1
sysctl -w net.ipv4.tcp_tw_recycle=1

Enter fullscreen mode Exit fullscreen mode

Keepalive Timeout

# Release idle persistent connections sooner (default is much longer)
sysctl -w net.ipv4.tcp_keepalive_time=1800   # 1800 seconds

Enter fullscreen mode Exit fullscreen mode

SYN Backlog (DoS Protection)

# Max length of queue for TCP connections not yet ESTABLISHED
# Prevents server crash under SYN flood / DoS attacks
sysctl -w net.ipv4.tcp_max_syn_backlog=4096

Enter fullscreen mode Exit fullscreen mode

Disable Unnecessary Protocols

# Disable ICMP broadcast responses to reduce noise
sysctl -w net.ipv4.icmp_echo_ignore_broadcasts=1

Enter fullscreen mode Exit fullscreen mode

Bind Network Interrupts to One CPU

# Bind NIC interrupt to CPU core 1 (reduces scheduler interference)
echo 02 > /proc/irq/<nic_irq>/smp_affinity

Enter fullscreen mode Exit fullscreen mode


Recommended Monitoring Toolset

Tool Purpose
htop Interactive process and CPU monitor
vmstat CPU, memory, swap, I/O, context switches
iotop Per-process disk I/O
sar Historical system activity reports
strace Trace system calls for a process
iftop Real-time network bandwidth by connection
ss Socket statistics (faster replacement for netstat)
lsof List open files and file descriptors
ethtool NIC diagnostics and settings
mtr Combined traceroute + ping for network diagnosis