惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
C
Cisco Blogs
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
IT之家
IT之家
博客园 - 【当耐特】
V
V2EX
博客园_首页
T
Tailwind CSS Blog
Last Week in AI
Last Week in AI
G
Google Developers Blog
The Last Watchdog
The Last Watchdog
C
CXSECURITY Database RSS Feed - CXSecurity.com
博客园 - 司徒正美
N
Netflix TechBlog - Medium
F
Fortinet All Blogs
Know Your Adversary
Know Your Adversary
S
Schneier on Security
V
Vulnerabilities – Threatpost
T
The Exploit Database - CXSecurity.com
Vercel News
Vercel News
量子位
G
GRAHAM CLULEY
T
Threatpost
D
Darknet – Hacking Tools, Hacker News & Cyber Security
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
C
Cybersecurity and Infrastructure Security Agency CISA
S
Security @ Cisco Blogs
B
Blog
Stack Overflow Blog
Stack Overflow Blog
T
Tor Project blog
A
About on SuperTechFans
博客园 - 叶小钗
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
月光博客
月光博客
S
Securelist
博客园 - 聂微东
Cloudbric
Cloudbric
N
News and Events Feed by Topic
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
H
Help Net Security
N
News | PayPal Newsroom
P
Privacy & Cybersecurity Law Blog
Schneier on Security
Schneier on Security
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
W
WeLiveSecurity
Martin Fowler
Martin Fowler
K
Kaspersky official blog
S
Security Affairs
TaoSecurity Blog
TaoSecurity Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Kubernetes VPA vs HPA vs KEDA: Which Autoscaler Actually Cuts Your Bill
Muskan · 2026-05-06 · via DEV Community

Muskan

The average Kubernetes cluster runs at 13% CPU utilization and 20% memory utilization. That means 87% of provisioned compute sits idle. Three autoscalers exist to close that gap — VPA, HPA, and KEDA — and each attacks the problem from a different angle. VPA shrinks oversized pods. HPA adds and removes pod replicas. KEDA extends HPA with event-driven triggers and the ability to scale to zero.

Choosing the wrong autoscaler does not just leave money on the table. It creates production incidents. VPA evicts pods to resize them, which restarts JVM applications cold. HPA thrashes replica counts without stabilization windows, causing 5xx errors during scale-down. KEDA's scale-to-zero adds cold start latency that breaks SLA commitments. The cost savings are real, but only if the autoscaler matches the workload.

Your Cluster Is Running at 13% CPU. The Autoscaler Choice Determines What Happens Next

The CNCF 2024 Kubernetes Benchmark Report analyzed 4,000 clusters. Average CPU utilization: 13%. Average memory utilization: 20%. Memory overprovisioning by cloud provider: Azure at 65%, AWS at 58%, GCP at 53%. Large clusters with 30,000+ CPUs reach 44% utilization — better, but still more than half idle.

Three autoscaler paths from an overprovisioned cluster

The waste is not hypothetical. An estimated 35% of Kubernetes spending goes to overprovisioned resources. A ZeonEdge case study documented a reduction from $50,000 to $22,000 per month — 56% — using VPA right-sizing as a core component. The autoscaler choice is the primary lever for closing the gap between what you provision and what you use.

How Each Autoscaler Works

VPA (Vertical Pod Autoscaler) adjusts CPU and memory requests on individual pods. The recommender component maintains a decaying histogram of actual usage over an 8-day window with a 24-hour half-life. When current requests diverge from recommendations, VPA evicts the pod and recreates it with updated resource values via the admission controller. VPA does not add pods. It makes existing pods the right size.

HPA (Horizontal Pod Autoscaler) scales the number of pod replicas. It polls metrics every 15 seconds, compares observed CPU or memory utilization against a target threshold, and adjusts the replica count. When utilization exceeds 70% (a common target), HPA adds replicas. When it drops, HPA removes them. The scaling loop takes 2-4 minutes end-to-end from metric observation to new pod serving traffic.

KEDA (Kubernetes Event-Driven Autoscaling) extends HPA with 70+ external event sources. Instead of scaling on CPU utilization, KEDA scales on queue depth (SQS, Kafka, RabbitMQ), custom Prometheus queries, cron schedules, or HTTP request rates. KEDA creates and manages an HPA resource internally. Its defining capability: scaling to zero replicas when no events are present, and back to one or more when events arrive.

VPA, HPA, and KEDA internal flows side by side

The Numbers: Cost Savings, Latency, and Reaction Speed

Dimension VPA HPA KEDA
What it scales Pod CPU/memory requests Pod replica count Replica count (including to zero)
Cost savings range 30-40% compute reduction 30-50% vs static provisioning 50-70% for idle workloads
Reaction speed 24-48 hours warmup for recommendations 2-4 minutes end-to-end 2-5 seconds cold start (simple apps)
Scale to zero No No (minimum 1 replica) Yes
Latency impact Pod eviction + restart P95 latency: 398ms to 750ms during scaling 15-60+ seconds cold start for heavy apps
Data source 8-day usage histogram CPU/memory metrics (15s poll) 70+ event sources
Best workload type Memory-bound, stable patterns CPU-bound web/API servers Event-driven, batch, intermittent

The numbers expose a fundamental tradeoff. VPA delivers the deepest per-pod savings but cannot react to real-time spikes — it is backward-looking by design. HPA responds within minutes but adds latency during scaling events. KEDA offers the most aggressive cost savings through scale-to-zero but introduces cold start latency that can break SLA commitments for latency-sensitive services.

A concrete example: a team running a queue processor that handles 10,000 messages during business hours and zero messages overnight. With HPA, the minimum replica stays running 24/7 — roughly 128 idle hours per week. With KEDA scaling to zero overnight, those 128 hours cost nothing. At $0.05 per pod-hour, that is $6.40 per week per service. Across 50 microservices, that is $16,640 per year from one configuration change.

The Compatibility Matrix: What You Can and Cannot Combine

Combination Works? Why
VPA + HPA on same metric (CPU) No Death spiral: VPA lowers requests, HPA sees higher utilization %, scales up, VPA lowers again
VPA + HPA on different metrics Yes HPA scales horizontally on CPU, VPA right-sizes memory with controlledResources: ["memory"]
KEDA + manual HPA No KEDA creates its own HPA internally — two HPAs on the same deployment causes erratic scaling
KEDA + VPA Yes KEDA handles replica count, VPA handles resource requests — different dimensions, no conflict
VPA + HPA + KEDA all three Partial KEDA replaces HPA; combine KEDA + VPA on different resources only

The death spiral between VPA and HPA on the same metric is the most common production mistake. VPA reduces CPU requests based on historical usage. This lowers the denominator in HPA's utilization calculation. HPA now sees artificially high utilization and scales up replicas. More replicas distribute load, reducing per-pod usage. VPA sees lower per-pod usage and reduces requests further. The cycle repeats until pods are tiny and replica count is enormous.

The safe combination: HPA (or KEDA) for horizontal scaling decisions, VPA for memory right-sizing only. Set VPA's controlledResources to ["memory"] so it ignores CPU entirely. This lets HPA own the CPU scaling decision without interference.

Production Gotchas That Will Bite You

Autoscaler Gotcha Symptom Fix
VPA Eviction cascades JVM loses JIT-compiled code, Redis drops cache, WebSocket sessions disconnect Use VPA in "Off" mode for recommendations only; apply during maintenance windows for stateful workloads
VPA Recommendations exceed node capacity Pods become permanently unschedulable Set maxAllowed in VPA policy to stay within node instance type limits
VPA JVM heap mismatch Container gets more memory but JVM heap flags unchanged — JVM does not use it Update JVM -Xmx flags in deployment spec when VPA increases memory
HPA Thrashing without stabilization 3-second traffic spike scales from 3 to 10 replicas and back, causing 5xx errors Set behavior.scaleDown.stabilizationWindowSeconds: 300 and scaleUp.stabilizationWindowSeconds: 60
HPA Slow response to traffic spikes 2-4 minutes from spike to new pod serving Pre-scale before known traffic events; use KEDA cron scaler for predictable patterns
KEDA Cold start on scale from zero First requests get 503 until pod passes readiness probe Set minReplicaCount: 1 for latency-sensitive services; accept scale-to-zero only for async workloads
KEDA Cooldown misconfiguration Pods scale to zero during brief pauses between events Increase cooldownPeriod beyond the maximum gap between events (default: 300 seconds)

The VPA eviction problem deserves emphasis. When VPA decides a pod needs different resources, it evicts the pod. The new pod starts fresh. For a JVM application, this means losing all JIT-compiled bytecode — the application runs interpreted for minutes until the JIT warms up again. For a Redis instance, the in-memory cache is gone, causing a thundering herd of cache misses. For a WebSocket server, every connected client disconnects.

The fix: use VPA in "Off" mode. It still generates recommendations. Apply them manually during maintenance windows or CI deploys rather than letting VPA evict pods autonomously. Kubernetes 1.35+ supports in-place pod resizing, which eliminates the eviction requirement — but adoption is still early.

The Decision Framework: Match Your Workload to Your Autoscaler

Decision tree: workload type to recommended autoscaler

CPU-bound workloads (web servers, API gateways, compute services): HPA with a CPU target of 60-70%. Add VPA in memory-only mode for right-sizing.

Memory-bound workloads (JVM applications, caches, data processing): VPA in "Off" or "Auto" mode. Avoid VPA for stateful workloads where eviction causes data loss.

Event-driven workloads (queue processors, batch jobs, webhooks, scheduled tasks): KEDA with the appropriate scaler. Scale-to-zero for intermittent workloads. KEDA cron scaler for predictable traffic patterns.

Mixed workloads: KEDA replaces HPA for the horizontal scaling decision. Add VPA with controlledResources: ["memory"] for vertical right-sizing. This combination covers both dimensions without conflict.


The 87% of idle CPU in your cluster is not a monitoring problem. It is an autoscaling problem. VPA closes the gap between what pods request and what they use. HPA matches replica count to actual traffic. KEDA eliminates cost for workloads that spend most of their time idle. Pick the autoscaler that matches your workload pattern, combine them safely where needed, and watch for the gotchas that turn cost savings into production incidents. The numbers are clear — 30-70% savings are achievable. The engineering is in choosing the right tool and configuring it correctly.