惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Tailwind CSS Blog
C
CERT Recently Published Vulnerability Notes
P
Proofpoint News Feed
Vercel News
Vercel News
博客园 - 三生石上(FineUI控件)
IT之家
IT之家
Help Net Security
Help Net Security
月光博客
月光博客
N
News and Events Feed by Topic
Cloudbric
Cloudbric
博客园 - 司徒正美
L
LangChain Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
Tenable Blog
The Register - Security
The Register - Security
The Hacker News
The Hacker News
I
InfoQ
The Last Watchdog
The Last Watchdog
MyScale Blog
MyScale Blog
Schneier on Security
Schneier on Security
WordPress大学
WordPress大学
小众软件
小众软件
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
宝玉的分享
宝玉的分享
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
K
Kaspersky official blog
L
LINUX DO - 热门话题
N
News | PayPal Newsroom
F
Fortinet All Blogs
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
S
Security @ Cisco Blogs
Recorded Future
Recorded Future
大猫的无限游戏
大猫的无限游戏
H
Help Net Security
Google Online Security Blog
Google Online Security Blog
S
Schneier on Security
C
Cisco Blogs
N
News and Events Feed by Topic
V2EX - 技术
V2EX - 技术
Latest news
Latest news
PCI Perspectives
PCI Perspectives
T
The Blog of Author Tim Ferriss
P
Palo Alto Networks Blog
T
Tor Project blog
Project Zero
Project Zero
云风的 BLOG
云风的 BLOG
Webroot Blog
Webroot Blog
Attack and Defense Labs
Attack and Defense Labs
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Autoscaling Is an Authority System, Not a Capacity System
NTCTech · 2026-06-25 · via DEV Community

NTCTech

Autoscaling authority is the condition most cloud operations teams have never formally defined. Every organization running Kubernetes has autoscaling configured. Almost none has treated those configurations as governance artifacts.

autoscaling authority — two-state flow diagram comparing explicit authority model versus scaling divergence

The Capacity Framing Is Wrong

Autoscaling authority governs what your execution plane is permitted to do under changing load conditions — and that framing is the one most engineering teams have never applied to it.

The standard framing is operational: autoscaling solves the overprovisioning versus underprovisioning tradeoff. Set a threshold, define bounds, let the system respond. That framing isn't wrong — it's just operating at the wrong level of analysis. Capacity is the visible output. Authority is the invisible decision structure underneath it.

Every autoscaling configuration answers a governance question before it answers a capacity question: what is this system permitted to do without asking a human? When an HPA scales a deployment from 3 to 30 replicas in response to a traffic event, no engineer approved that action. The execution plane acted under authority delegated at configuration time. That delegation is the architecture — and in most organizations, it's ungoverned.

Cloud architecture's central authority question is whether authority defined at the policy layer actually reaches the systems that execute work. Autoscaling is the most common place that question gets answered by accident.

What Autoscaling Actually Encodes

A scaling configuration is a policy artifact. Every value in it encodes a decision about what the system is authorized to do.

HPA target utilization: a decision. Min and max replica bounds: a decision. VPA update mode — Off, Initial, Recreate, Auto: a decision about whether the system may modify running pods without human approval. KEDA trigger sources and thresholds: decisions about what signals the execution plane is authorized to act on. Each of these encodes a judgment about permitted behavior. The problem isn't that the decisions are wrong. The problem is that they're almost never recognized as decisions at all.

They get written once, at initial deployment, calibrated to the workload as it existed at launch. Then they persist. Not because someone reviewed them and confirmed they were still valid. Because no one touched them, and defaults survive indefinitely unless deliberately revisited.

Silent Delegation

No architect signs a document stating: "The execution plane may increase application capacity by 500% without human review."

Yet that's exactly what many autoscaling policies authorize.

The delegation happened when the configuration was written. The authority persisted long after the original decision-maker stopped thinking about it. Teams don't experience this as a governance failure — they experience it as autoscaling working as designed.

Diagnostic: "Who authorized your execution plane to make this decision — and do they still work here?"

When Scaling Behavior Diverges From Intent

The named failure state for this pattern is Scaling Divergence: the condition where autoscaling behavior remains technically correct relative to configuration while becoming operationally incorrect relative to current intent.

Scaling Divergence doesn't announce itself. There's no alert for "autoscaling is no longer doing what you intended." The system continues functioning. Metrics look normal. Deployments scale. The divergence is only visible when you compare current scaling behavior against the workload model that originally justified the configuration — and most teams never make that comparison.

The clearest real-world scenario: a team sets HPA target utilization at 70% at application launch. At that point, the bottleneck is CPU — the application is compute-bound under normal load, and 70% is a reasonable ceiling before latency degrades. Nine months later, the application's bottleneck has shifted to database connection pool saturation. Response time is now dominated by wait time on external I/O, not CPU cycles.

The autoscaler continues making perfectly valid CPU-driven decisions. CPU utilization stays well below the threshold under rising load. The autoscaler holds replica count steady. Latency spikes. The operations team investigates CPU — the metric the autoscaler watches — and sees nothing wrong. The workload changed. The authority logic did not.

scaling divergence pattern — workload characteristics change while autoscaling authority remains static

Three Failure Modes

Scaling Divergence manifests in three distinct patterns, escalating in severity and how long they typically go undetected.

01 — CEILING BLINDNESS

Max replicas were set at deployment time, but weren't derived from actual infrastructure capacity headroom. During a traffic event, the autoscaler hits the ceiling before load abates. Latency spikes. No one can confirm whether the ceiling is a deliberate safety boundary or an artifact of the original spec. The ceiling is enforcing something — nobody knows what.

02 — FLOOR DRIFT

Min replicas were set conservatively at launch to reflect early-stage load. The application has since scaled to 3× its original steady-state size. The min floor still reflects day-one sizing. During an incident-driven scale-down, the autoscaler hits the floor and holds there — at a replica count that hasn't matched actual minimum viable capacity in months. The floor feels like a safety net. It's a liability from an expired sizing decision.

03 — MULTI-CONTROLLER CONFLICT

HPA and VPA are running on the same workload without a defined coordination mode. VPA adjusts resource requests. HPA recalculates utilization ratios against the new requests. Each controller is operating correctly within its own decision boundary. The authority conflict is at the policy layer — two systems making decisions neither was designed to coordinate. Scheduler behavior becomes unpredictable and the root cause is invisible until you trace back to the configuration layer.

Common mistake: Running HPA and VPA on the same workload without setting VPA to Off or Initial mode. The default assumption is that both systems will self-coordinate. They won't. VPA modifies resource requests; HPA recalculates utilization ratios against those modified requests. The interaction is defined, but it's not intuitive, and the failure mode only surfaces under load.

What Autoscaling Governance Looks Like

If autoscaling is an authority system, it requires four properties to function as one.

01 — DECLARED INTENT — Scaling boundaries documented as decisions, not just values. Not "HPA target: 70%" in a YAML file — but what load profile that threshold was calibrated for, what behavior is expected outside that profile, and what workload characteristic it assumes. Without declared intent, there's no basis for evaluating whether current configuration is still valid.

02 — AUTHORITY OWNERSHIP — Who owns this scaling policy? Platform team? Application team? SRE? Cloud operations? If the answer is unclear, nobody owns it — which means nobody is accountable when scaling behavior diverges from intent. The ownership question determines who reviews the configuration when workload characteristics change, and who gets paged when a scaling event produces unexpected behavior.

03 — CHANGE COUPLING — Scaling configuration treated as code: changes to workload characteristics trigger review of scaling policy. The artifact is an explicit coupling between workload change events and scaling policy review. When the application's performance profile changes, the scaling authority that governs it should be reconsidered as part of the same change process, not discovered six months later during an incident.

04 — BEHAVIORAL AUDIT — Periodic comparison of actual scaling events against expected behavior under the declared intent. This is distinct from monitoring, which is reactive. Behavioral audit is deliberate review — examining whether the autoscaler is making decisions that align with the workload model the configuration assumed. Not "did it scale?" but "did it scale in the way the policy intended?"

autoscaling governance framework — four properties of explicit autoscaling authority

The Operational Architecture Connection

Framework #152 Operational Authority Boundary defines the point at which authority must translate into executable operational behavior. Below the boundary: the execution plane acts as authority intends. Above it: execution proceeds independently.

Autoscaling sits directly at that boundary. Every scaling configuration is simultaneously a policy decision (above the boundary) and an executable instruction to the runtime (below it). The failure state — Authority Without Execution — typically manifests as systems that have policies with no enforcement. In autoscaling, the failure is the inverse: execution without current authority. The runtime is executing perfectly. It's just not executing what anyone still intends.

Autoscaling Authority Audit

Three questions that should be answerable for any autoscaling configuration in your estate:

1. What workload model does this configuration assume?
Not "what are the current values" — but what load profile, bottleneck type, and traffic pattern were used to derive those values. If nobody can answer this, the configuration is ungoverned regardless of whether it's functioning correctly.

2. When was this configuration last validated against actual workload behavior?
Not last modified. Last validated — meaning someone compared current scaling events against the intent the configuration was designed to express.

3. Who is accountable if this configuration produces scaling behavior that damages the application under load?
If that answer resolves to nobody, authority ownership is undefined. The execution plane is making decisions with no identified principal.

Architect's Verdict

Autoscaling is one of the most widely deployed authority systems in modern cloud infrastructure. Every scaling policy delegates operational decision-making to the execution plane. The question is whether the authority you intended actually reaches the runtime that makes decisions under load — and for autoscaling, it reaches the runtime once, at configuration time, and then persists until someone deliberately revisits it.

The operational failure most teams encounter isn't that autoscaling doesn't work. It's that it works exactly as configured, and the configuration no longer reflects anything anyone deliberately decided. That gap — between technical correctness and operational intent — is what Scaling Divergence names. It doesn't show up in scaling metrics. It shows up in incidents where the autoscaler performed as designed and the outcome was still wrong.

Every autoscaling system either has explicit authority or it has defaults. Defaults are simply authority decisions that survived long enough to become invisible.

Originally published at rack2cloud.com