惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Schneier on Security
C
Cyber Attacks, Cyber Crime and Cyber Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Project Zero
Project Zero
T
The Exploit Database - CXSecurity.com
G
GRAHAM CLULEY
T
Threatpost
A
Arctic Wolf
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Scott Helme
Scott Helme
Simon Willison's Weblog
Simon Willison's Weblog
P
Proofpoint News Feed
C
Cisco Blogs
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
K
Kaspersky official blog
P
Palo Alto Networks Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
T
Threat Research - Cisco Blogs
The Hacker News
The Hacker News
T
Tor Project blog
NISL@THU
NISL@THU
The GitHub Blog
The GitHub Blog
Security Latest
Security Latest
aimingoo的专栏
aimingoo的专栏
C
CERT Recently Published Vulnerability Notes
Recorded Future
Recorded Future
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
Google DeepMind News
Google DeepMind News
Martin Fowler
Martin Fowler
N
News | PayPal Newsroom
P
Privacy & Cybersecurity Law Blog
MyScale Blog
MyScale Blog
G
Google Developers Blog
V
V2EX
V
Visual Studio Blog
P
Privacy International News Feed
Google Online Security Blog
Google Online Security Blog
Microsoft Azure Blog
Microsoft Azure Blog
宝玉的分享
宝玉的分享
博客园 - 【当耐特】
L
LINUX DO - 热门话题
MongoDB | Blog
MongoDB | Blog
腾讯CDC
J
Java Code Geeks
The Last Watchdog
The Last Watchdog
L
Lohrmann on Cybersecurity
Cyberwarzone
Cyberwarzone
博客园 - 聂微东
Webroot Blog
Webroot Blog
S
Secure Thoughts

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The Silent Death of the System: OOM Killer and My VPS Journey
Mustafa ERBA · 2026-05-14 · via DEV Community

Introduction: A VPS Experience in the Shadow of OOM Killer

For some time, I've been having some interesting and at times frustrating experiences with memory management for applications running on my own Virtual Private Server (VPS). Especially during peak periods or unexpected traffic surges, when the system's memory (RAM) consumption reached critical levels, I encountered the OOM Killer (Out-of-Memory Killer). This is a situation we could call the "silent death" of the system, because your applications can suddenly shut down, data loss can occur, or the entire system can become unresponsive. In this post, I will delve into what OOM Killer is, why it occurs, the specific cases I faced on my VPS, and the strategies I implemented to deal with this problem.

The purpose of this article is not just to explain how a problem was solved, but also to delve into the depths of operating system memory management mechanisms and offer a perspective rooted in direct field experience, far from a "corporate consultant" tone. This adventure I had with my own VPS taught me invaluable lessons about the complex inner workings of Linux and showed me how delicate the balance of system reliability is. Especially in the first half of 2026, increasing workloads and more complex service deployments have made memory management issues more visible.

What is OOM Killer and Why Does It Occur?

The Linux kernel is designed to efficiently manage system resources. Memory (RAM) is one of the most critical of these resources. All processes running on a system occupy space in memory. If the total memory demand on the system exceeds the available physical RAM and swap space, the kernel faces an "out of memory" situation. At this point, a mechanism kicks in to prevent the system from completely crashing: the OOM Killer. The OOM Killer attempts to free up memory by selecting and terminating one or more processes currently running on the system.

ℹ️ How OOM Killer Works

When deciding which process to terminate, OOM Killer typically uses an oom_score. This score is calculated based on various factors such as the amount of memory the process occupies, its runtime, and the time spent in kernel mode. Processes with a higher oom_score are considered more likely targets by the OOM Killer. The ability to adjust this score is one of the methods used to prevent certain critical processes from being terminated.

So, how does this "out of memory" situation arise? There can be multiple reasons:

  • Application Errors: Applications with memory leaks gradually consume more and more memory over time.
  • Unexpected Load Increase: A sudden surge in traffic or processing load can exceed available resources.
  • Misconfiguration: Incorrectly set memory limits for applications or system services.
  • Insufficient Resources: The server not having enough memory initially.
  • Too Many Services: Running more services than necessary on a single VPS.

In my case, several scenarios converged. One of the backend services of a production ERP system was consuming much more memory than expected, especially when certain reporting queries were run. In addition, a Time Series Database (TSDB) and several small helper services running on the same VPS also created a continuous memory load. This combination began to push the system to its limits.

A Concrete Case: Delayed Shipment Reports in Production ERP and Data Loss After OOM

A few months ago, while working on a large manufacturing company's ERP system, we noticed that shipment reports were consistently delayed. Reports typically arrived with a 2-hour delay, which made it difficult to make operational decisions. Initially, we thought the problem was with the reporting query itself. We added indexes to optimize the query, examined the query plan, but didn't get the expected performance. It took a full 3 days to find the root cause of the problem.

The gist of the incident was this: when a shipment report for a specific date range was run, the backend service was pulling an enormous amount of data and processing it in memory. The PostgreSQL database also used intensive memory for this query. After about 1.5 hours, both the application service and the database began to push their memory limits. The total memory of the VPS was around 32GB, and memory usage exceeded 95% while this query was running. At this critical moment, OOM Killer intervened. However, OOM Killer first terminated the database process, and then the main process of the backend service.

The result: The reporting process was interrupted, the database shut down unexpectedly, and the backend service crashed. Although the system tried to recover a few minutes later, the reporting process remained "incomplete" and data loss occurred. This situation not only caused reporting issues but also operational disruptions and potential financial losses. One of the most important lessons I learned from this case was that OOM Killer is not just a "bad thing," but also a recovery mechanism that prevents the system from completely crashing; however, the cost can be high.

⚠️ Post-OOM Killer Situation Assessment

When OOM Killer terminates a process, it usually leaves a relevant entry in log files such as /var/log/syslog or /var/log/messages. These logs contain clues about which process was terminated, how much memory it used, and why OOM Killer made that decision. For example, you might see a log line like this:

May 14 03:15:12 vps-host kernel: [12345.67890] Out of memory: Kill process 9876 (myapp) score 1234, message in ...

Regularly checking these logs is critical for understanding the source of OOM issues.

Analyzing Memory Usage: top, htop, and free

To understand and prevent the excessive memory usage that triggers OOM Killer, it's essential to regularly monitor memory usage on the system. There are several basic tools we can use for this:

  1. top Command: Shows real-time CPU and memory usage of running processes. The %MEM column indicates the percentage of total memory a process occupies. You can press Shift + M to sort processes by memory usage.

  2. htop Command: A more user-friendly and interactive version of the top command. It's more useful with features like colored output, process trees, and mouse support. Memory usage, CPU usage, and other information are also presented graphically.

  3. free Command: Summarizes the system's overall memory usage. It shows information such as used memory, free memory, buffers, and cache. The free -h command displays the output in a human-readable format (like MB, GB).

These tools became my closest companions during my VPS journey. The first thing I did every morning was to check the system's overall status with htop. Especially if memory usage was consistently in the 80-90% range, it was an alarm signal. The free -h command clearly showed the current total memory status. For example, once, the free -h output was like this:

              total        used        free      shared  buff/cache   available
Mem:           31Gi        29Gi       1.2Gi       100Mi       800Mi        1Gi
Swap:         2.0Gi       500Mi       1.5Gi

Enter fullscreen mode Exit fullscreen mode

This output shows that 29GB of 31GB memory is used, only 1.2GB is free, and half of the swap space is also full. This situation was a clear indication that OOM Killer could intervene shortly. The low available memory amount is particularly critical. This value indicates how "comfortable" the system is to start new processes or meet the memory demands of existing ones.

💡 Swap Space and Performance

Swap space is backup memory on disk used when physical RAM is full. Swap usage usually indicates a performance drop because disk access is much slower than RAM access. Using swap to avoid OOM Killer is a solution, but it reduces the overall speed of the system. The ideal is to keep swap usage to a minimum and provide sufficient RAM.

Strategies to Prevent OOM Killer

Several different strategies can be followed to prevent OOM Killer from being triggered. These may vary depending on the source of the problem and the system's architecture:

1. Detecting and Fixing Memory Leaks

This is the most ideal and permanent solution. You can use various tools to detect memory leaks in your applications:

  • Valgrind: A powerful tool for detecting memory errors and leaks. However, it can significantly slow down application execution, so it's typically used in development environments.
  • Application Performance Monitoring (APM) Tools: APM tools like Datadog, New Relic help you monitor application memory usage and detect anomalies.
  • Log Analysis: Looking for clues about memory usage in application logs.

In my ERP example, I analyzed query-based memory usage to identify the problem. PostgreSQL's own pg_stat_activity view and query logs provided clues about which queries consumed how much memory. Afterward, I realized that a function processing data on the Vue.js frontend was fetching too much data and holding it in memory. I optimized this function, rewriting it to fetch only necessary data and occupy less memory. With this change, the memory consumption of the reporting query decreased by 60%.

2. Setting Service-Based Memory Limits

Limiting how much memory each service can use prevents one service from overwhelming others. Container orchestration tools like Docker and Kubernetes are very effective in this regard. However, it's also possible to set limits via systemd service units in bare-metal or simple VPS setups.

For example, you can set memory limits by adding these lines to a systemd unit file for a service:

[Service]
LimitAS=2G     # Address space limit (including virtualized memory)
LimitRSS=1G    # Resident Set Size limit (physical memory)
MemoryHigh=1.5G # Soft memory limit, kernel becomes more cautious when exceeded
MemoryMax=2G   # Hard memory limit, OOM Killer is more likely to intervene when exceeded

Enter fullscreen mode Exit fullscreen mode

This configuration helps the system become more cautious when the relevant service starts using more than 1.5GB of memory. If it exceeds 2GB, the OOM Killer is more likely to target this service. This was a method I used to prevent critical services, especially those I developed myself and had direct control over, from harming the production ERP or database. For example, I set a limit like MemoryMax=512M for the service running the backend of my own financial calculator application, so my own application wouldn't affect the overall system.

3. Optimizing or Increasing Swap Space

If there are no memory leaks and your existing applications' memory needs are known, but sometimes exceed this need, increasing swap space can be a temporary solution. However, this is not a long-term solution as it causes performance degradation. It's also possible to optimize swap by adjusting the swappiness parameter.

The swappiness value ranges from 0 to 100. Lower values (e.g., 10) make the kernel avoid moving memory to swap space as much as possible. Higher values (e.g., 60, the default) encourage the kernel to aggressively move memory to swap space.

To change the swappiness value:

# Check current value

<figure>
  <Image src={cover} alt="Graph showing memory usage on VPS and OOM Killer logs" />
</figure>

cat /proc/sys/vm/swappiness

# Change temporarily (lost after reboot)
sudo sysctl vm.swappiness=10

# Change permanently (by editing /etc/sysctl.conf file)
# Add the following line to the file:
# vm.swappiness = 10
sudo sysctl -p

Enter fullscreen mode Exit fullscreen mode

Lowering the swappiness value to 10 on my VPS allowed the system to retain physical memory longer and reduced swap usage. This decreased the frequency of OOM Killer interventions but did not eliminate them entirely.

4. Adding More RAM

The simplest and most effective solution, if possible, is to add more RAM to your VPS. However, this increases costs and may not always be feasible. If cost and infrastructure allow, this is often the best and cleanest solution. In my situation, I already had the highest RAM capacity offered by my VPS provider, so this option was not applicable to me.

Configuring OOM Killer: oom_score_adj and oom_score

You can adjust the oom_score_adj value to control more precisely which processes OOM Killer terminates. This value is found in the /proc/<PID>/oom_score_adj file for each process. The value can range from -1000 to 1000.

  • Negative Values: Reduce the likelihood of the process being terminated by OOM Killer. Used for critical services.
  • Positive Values: Increase the likelihood of the process being terminated by OOM Killer. Can be used for background processes that are not prioritized.

For example, for a critical process like the ERP backend service, you can lower this value:

# Find the PID of the ERP backend process (e.g., 12345)
ps aux | grep my-erp-backend

# Set oom_score_adj to -500 for PID 12345
echo -500 | sudo tee /proc/12345/oom_score_adj

Enter fullscreen mode Exit fullscreen mode

This command applies to the currently running process. To make it permanent, you need to configure it in the systemd service unit or startup scripts. Similarly, I increased the oom_score_adj value for one of my helper services, ensuring that if a memory problem occurred, this service would be the first to shut down.

oom_score, on the other hand, shows the current oom_score of the process. By reading this, you can see which processes are more likely to be targeted:

# Read oom_score for PID 12345
cat /proc/12345/oom_score

Enter fullscreen mode Exit fullscreen mode

These tools are invaluable for understanding and intervening in OOM Killer's "decision-making" mechanism.

Conclusion: System Balance and Continuous Monitoring

My OOM Killer adventures on my VPS once again showed me how complex and delicately balanced systems are. OOM Killer is not a "debugging" tool, but a "recovery" mechanism, and its activation usually signals the existence of a deeper problem.

The most important lessons I learned during this process were:

  1. Understanding the Problem: Realizing that what triggers OOM Killer is not OOM Killer itself, but insufficient memory or a memory leak.
  2. Monitoring: Continuously monitoring system resources (especially memory) allows for intervention before problems escalate. htop, free, and log analysis play a critical role here.
  3. Optimization: Optimizing the memory usage of applications and databases is the most permanent solution.
  4. Configuration: If necessary, setting service-based memory limits or adjusting kernel parameters like oom_score_adj can be used as a protective measure for the system.
  5. Trade-offs: Every solution has a trade-off. Using swap degrades performance, and memory limits can prevent applications from reaching their full potential. Finding the best balance is important.

In the world of technology, "perfect" and "error-free" systems are rare. What matters is understanding how these systems work, being prepared for potential issues, and solving problems with a pragmatic approach. OOM Killer is a silent but powerful reminder from systems: use resources carefully and always have a Plan B. Such experiences have been a continuous driving force for me to become a better system architect and operator.

As I mentioned in my previous post [related: my VPS migration experience], there is always something new to learn. This OOM adventure also gave me a deep understanding of memory management.