惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
H
Help Net Security
博客园_首页
P
Privacy International News Feed
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Tenable Blog
Latest news
Latest news
D
Darknet – Hacking Tools, Hacker News & Cyber Security
爱范儿
爱范儿
Cyberwarzone
Cyberwarzone
P
Palo Alto Networks Blog
月光博客
月光博客
有赞技术团队
有赞技术团队
Know Your Adversary
Know Your Adversary
博客园 - 叶小钗
P
Proofpoint News Feed
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
S
Schneier on Security
A
About on SuperTechFans
F
Full Disclosure
The Cloudflare Blog
T
The Exploit Database - CXSecurity.com
S
SegmentFault 最新的问题
Apple Machine Learning Research
Apple Machine Learning Research
Microsoft Security Blog
Microsoft Security Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
N
News and Events Feed by Topic
www.infosecurity-magazine.com
www.infosecurity-magazine.com
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Webroot Blog
Webroot Blog
量子位
大猫的无限游戏
大猫的无限游戏
小众软件
小众软件
WordPress大学
WordPress大学
Last Week in AI
Last Week in AI
美团技术团队
Help Net Security
Help Net Security
Microsoft Azure Blog
Microsoft Azure Blog
GbyAI
GbyAI
云风的 BLOG
云风的 BLOG
Y
Y Combinator Blog
博客园 - 司徒正美
N
Netflix TechBlog - Medium
S
Security @ Cisco Blogs
B
Blog
P
Privacy & Cybersecurity Law Blog
Security Archives - TechRepublic
Security Archives - TechRepublic
Hacker News - Newest:
Hacker News - Newest: "LLM"
V
Visual Studio Blog
NISL@THU
NISL@THU

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
self-healing infrastructure 4 runbooks we deleted after automating them
Muskan · 2026-06-17 · via DEV Community

Every runbook your team executes manually is an open automation ticket that nobody filed. That is the central problem. The runbook library is not documentation. It is a backlog in

The Hidden Backlog Sitting in Your Runbook Library

Every runbook your team executes manually is an open automation ticket that nobody filed. That is the central problem. The runbook library is not documentation. It is a backlog in disguise, and most engineering teams never treat it that way.

Concept Runbook (Current State) Automation (Target State)
Classification Documentation / finished work Executable controller logic
Execution trigger On-call engineer at incident time Machine / controller automatically
Repeat toil signal Same runbook run 3+ times in a sprint Automation ticket filed and resolved
Deletion outcome 4 runbooks deleted after automation Conditional logic encoded in system
Audit action Pull runbooks executed in last 30 days Rank by frequency; top entry = first ticket

The mechanism is straightforward. When an engineer writes a runbook, they encode a decision tree: check this metric, restart that service, page this team if the threshold exceeds a value. That decision tree is executable logic. The moment it lives in a wiki instead of a controller, you have chosen human execution over machine execution.

Encoded logic belongs in controllers

That choice costs you every time an on-call engineer fires up the page at 2 a.m.

We built a self-healing infrastructure system and deleted 4 runbooks in the process (ZopDev, "Self-Healing Infrastructure: 4 Runbooks We Deleted After Automating Them"). Not archived. Deleted. The automation encoded the same conditional logic that the runbooks described, which made the documents redundant.

The deletion test in practice

That outcome is the proof point: if a runbook can be deleted after automation, it was always an automation candidate.

Runbooks as implicit tickets. Each operational procedure that requires a human to read, decide, and act represents a unit of toil that repeats on every incident. The repetition is the signal. If an engineer executed the same runbook three times in a sprint, that runbook belongs in a queue for automation, not a folder for reference.

The deletion test. The right question for any runbook is not "Is this documented?" but "Could a controller execute this without human input?" If the answer is yes, the runbook is blocking automation, not enabling operations. In our case, 4 runbooks passed that test completely. They are gone.

Runbooks as a hidden backlog

Backlog blindness. Teams fail to treat runbooks as automation candidates because runbooks feel like finished work. Writing the procedure feels like solving the problem. It is not. It is deferring the solution to the next on-call engineer.

The fix starts with a single audit: pull every runbook executed in the last 30 days, count the repeat executions, and rank them by frequency. The top entry on that list is your first automation ticket.

How Self-Healing Infrastructure Consumes Runbooks

Runbook automation works by transferring conditional logic from a human's working memory into a controller's execution loop, permanently. The human reads a runbook, evaluates state, and acts. The controller reads sensor data, evaluates state, and acts. The sequence is identical.

Aspect Human Executing Runbook Controller Executing Automation
Decision logic location Engineer's working memory Controller's execution loop
Latency & reliability Variable (human-dependent) Lower latency, higher reliability
Outcome logging Engineer fills in fields manually Controller logs same fields automatically
Document status after full automation Runbook deleted (not archived/deprecated) System encodes complete decision surface
Risk of keeping document alongside automation Two sources of truth diverge within weeks Eliminated by deletion
Suitable automation candidates Procedures requiring judgment/contextual deviation Deterministic procedures (same conditions, same choice every time)

From document to running system

The difference is in latency and reliability.

When we automated our first procedure, the runbook described a restart sequence triggered by a specific memory threshold. The controller we built watches the same threshold, executes the same restart, and logs the same outcome fields the engineer used to fill in manually. After 30 days of verified autonomous execution without a single human intervention, we deleted the runbook. Not as a symbolic gesture.

As a maintenance decision, a document that describes what a running system already does is a liability, not an asset. It will drift, contradict the implementation, and mislead the next engineer who reads it.

Three conditions for deletion

ZopDev's documented outcome is precise: 4 runbooks were deleted after automation, not retired to an archive, not marked deprecated (ZopDev, "Self-Healing Infrastructure: 4 Runbooks We Deleted After Automating Them"). Each deletion represents a complete transfer of decision logic from document to system. That number matters because it is concrete. Four procedures that previously required a human to wake up, read, and act now execute without human involvement.

Logic fidelity. The automation must encode the exact branching conditions the runbook described, including the edge cases buried in footnotes. Automation that covers only the happy path produces a controller that handles 80% of incidents and silently fails the other 20%. The fix is treating the runbook as a specification document during the build phase, not a reference document after deployment.

Deletion as validation. A runbook that cannot be deleted after automation was not fully automated. If the document still needs to exist for human reference, the controller is incomplete. It handles the common case but not the full decision surface. The 4 deletions at ZopDev confirm complete encoding, not partial coverage.

Selecting automation candidates

Drift prevention. Keeping a runbook alongside its automated equivalent creates two sources of truth. Within weeks, they diverge. An engineer following the runbook during a controller failure will execute steps the system no longer uses, against infrastructure the system has already changed. Deletion removes that failure mode entirely.

The selection criterion for automation candidates is execution frequency combined with decision determinism. A runbook executed repeatedly under the same conditions, where every engineer makes the same choice, is fully deterministic. Deterministic procedures automate completely. Procedures that require judgment, where experienced engineers sometimes deviate from the documented steps based on context, are not yet ready for full automation.

They need better specification first, then automation second. Start with the deterministic ones. The first deletion proves the model works.

Not Every Runbook Is Ready to Be Automated — How to Tell the Difference

The boundary between a runbook ready for full automation and one that still requires human judgment is not a matter of complexity. It is a matter of decision determinism.

The two-axis scoring framework

A deterministic runbook produces the same action every time the same conditions appear. No engineer deviates. No footnote says "use your judgment here." The conditional logic is complete, bounded, and reproducible. That is the automation signal.

A non-deterministic runbook contains at least one branch where experienced engineers sometimes choose differently based on context the runbook cannot fully capture. Automating that branch without resolving the ambiguity first produces a controller that acts confidently on an incomplete specification.

We built a scoring framework around two axes to make this evaluation concrete. The first axis is decision determinism: does every engineer who executes this runbook make the same choices under the same conditions? The second axis is blast radius. A runbook that restarts a single stateless service has a contained blast radius.

A runbook that modifies database connection pools or reroutes traffic across availability zones has a blast radius that crosses service boundaries and requires human accountability before action.

Axis Automation-Ready Requires Human Judgment
Decision determinism Same action every execution Engineers deviate based on context
Blast radius Single service, stateless Cross-service, stateful, or irreversible
Trigger clarity Metric threshold, no ambiguity Requires log interpretation or intuition
Rollback path Automated, tested Manual, partial, or undefined

The 4 runbooks deleted after automation at ZopDev ("Self-Healing Infrastructure: 4 Runbooks We Deleted After Automating Them") all cleared both axes. They encoded bounded conditional logic against measurable thresholds, and their remediation steps were reversible within the same execution context. That combination made deletion possible. A runbook that fails either axis produces a controller that handles the common case and creates a new failure mode for the exceptions.

Trigger, blast radius, rollback

Trigger clarity. Automation-ready runbooks fire on a specific, measurable condition: memory above a threshold, a pod restart count exceeding a limit, a health check returning a non-200 status. Runbooks that begin with "check the logs and determine if" are not yet ready. The determination step is human judgment encoded as prose, not as a sensor. The fix is extracting that determination into a concrete signal before writing any controller logic.

Blast radius scoring. Before automating any runbook, map every system it touches. A restart procedure that affects one stateless deployment scores low blast radius. A procedure that drains a node, reschedules workloads, and updates a load balancer target group scores high. High blast radius runbooks require a circuit breaker: the controller executes up to the point of irreversibility, then pages a human for the final confirmation.

That is not a failure of automation. It is the correct architecture for that risk profile.

Rollback completeness. A runbook is automation-ready only when its remediation steps are fully reversible by the same system that executes them. In our testing, procedures with undefined rollback paths produced controllers that could fix the immediate symptom and leave the system in a state no subsequent automation could safely interpret. We measured this failure mode in the first deployment week on two candidates we pulled back from automation. Both required rollback path specification before we rebuilt the controllers.

Running the automation audit

The specification-first rule is the most commonly skipped step. Teams see a frequently executed runbook and move directly to controller logic, treating the existing document as a complete specification. It is not. Runbooks written for human execution contain implicit knowledge: the engineer knows which log lines to ignore, which transient errors resolve without action, and which threshold spikes are artifacts of deployment pipelines.

None of that implicit knowledge appears in the document. The controller built from an incomplete specification will act on the artifacts and ignore the real signals.

The audit that surfaces automation candidates is straightforward. Pull every runbook executed in the last 90 days. For each one, ask three questions: Did every engineer who executed it make the same choices? Does its remediation stay within a single service boundary?

Does a tested rollback path exist? A runbook that answers yes to all three is ready for controller logic today. A runbook that answers no to any one of them needs specification work before a single line of automation is written.

The 4 deletions documented by ZopDev represent procedures that cleared all three criteria completely. The deletion was the outcome of that clarity, not the starting point. Start the audit this sprint. The runbooks that fail the determinism check are telling you exactly where your specification debt lives.

What You Should Measure Before and After Automating a Runbook

Measurement is what separates a successful automation from a successful-feeling one. Without baseline metrics captured before the controller goes live, you have no defensible answer when leadership asks whether the work was worth the engineering time. The three metrics that matter are mean time to recovery (MTTR), on-call page volume, and engineering hours consumed per incident class.

Page volume and engineering hours

MTTR is the primary signal. It measures the elapsed time from alert firing to system recovery. Before automation, that clock includes the time for an on-call engineer to wake up, read the procedure, evaluate conditions, and execute steps. After automation, the clock covers only detection latency plus execution time.

The human latency component, which includes context-switching overhead and the cognitive load of reading under pressure, disappears entirely. Record MTTR per incident type, not as a fleet average. A fleet average obscures which specific runbooks are delivering recovery gains.

On-call page volume measures whether automation is absorbing incidents or merely accelerating human response to them. A controller that remediates the condition before the alert threshold fires reduces page volume directly. A controller that remediates after the alert fires but before the engineer acts reduces MTTR without reducing pages. Both are wins, but they are different wins with different cost implications.

Tracking page volume separately from MTTR tells you which outcome you actually achieved.

Deletion count as governance signal

Engineering hours per incident class. This is the metric most teams skip because it requires pre-automation time logging. The mechanism is straightforward: each manual runbook execution consumes a measurable block of engineering time, including the interruption recovery cost after the incident closes. Without this baseline, you cannot calculate the labor cost recovered by automation. An on-call engineer interrupted at 2 a.m.

for a procedure that takes 20 minutes loses closer to 90 minutes of productive sleep and next-day focus.

Deletion count as a hard outcome. ZopDev tracked 4 runbooks deleted after automation (ZopDev, "Self-Healing Infrastructure: 4 Runbooks We Deleted After Automating Them"). Each deletion is a binary confirmation that the procedure no longer requires human execution under any documented condition. Deletion count is a governance metric, not a vanity metric. It confirms complete encoding rather than partial coverage.

Metric What It Confirms
MTTR per incident type Human latency removed from recovery path
On-call page volume Incidents absorbed before engineer wakes
Engineering hours per incident class Labor cost recovered per automation
Runbooks deleted Full decision transfer, no residual human dependency

30-day comparison cadence

Capture all four baselines before the controller deploys. The window for honest baseline data closes the moment the automation goes live and engineers stop executing the procedure manually. By sprint 3 of a typical automation initiative, teams that skipped pre-measurement are left reconstructing baselines from incident ticket timestamps, which are unreliable because engineers close tickets after the fact, not at the moment of resolution.

The measurement cadence matters as much as the metrics themselves. Run a 30-day comparison window: 30 days of pre-automation data against the first 30 days of controller operation. Shorter windows introduce noise from incident frequency variance. Longer windows allow the team to rationalize away regressions.

At 30 days, you have enough incident volume to distinguish signal from noise, and the comparison is still close enough in time that infrastructure conditions have not materially changed.

If MTTR drops but page volume holds flat, the controller is remediating after alert fire. The fix is adjusting the controller's trigger threshold to act before the alerting threshold is crossed. That is a tuning problem, not an architecture problem, and the measurement surface tells you exactly which knob to turn.

Building the Habit: Turning Runbook Reviews Into Automation Sprints

Runbook reviews become automation sprints only when the team treats the runbook backlog as a product backlog, with prioritization criteria, sprint commitments, and a definition of done that ends in deletion.

Practice Recommended Approach Failure Mode
Sprint commitment size 1–2 runbooks per two-week sprint 3+ runbooks causes context-switching overhead and degrades incident response
Review cadence 30-minute review at sprint start; any procedure run 3+ times is an immediate candidate Procedures run once per quarter recover less engineering time per controller
Retrospective metric Count runbooks deleted per quarter Counting controllers deployed instead misses parallel manual procedures still in use
Specification debt Assign owner + due date as a first-class ticket Without ownership, gaps accumulate and backlog stalls after early wins
Prioritization criterion Frequency over complexity; start with 3 highest-frequency if backlog exceeds 10 Optimizing for complexity rather than frequency reduces recovered engineering time
Deletion standard Runbook deleted = acceptance criterion; deployment is intermediate Stopping at deployment accumulates controllers alongside runbooks

Sprint structure and closure

The mechanism is straightforward. Every runbook executed in the last 90 days is a candidate. Each one enters a triage queue scored against the three-axis framework from the previous section. The highest-scoring candidates become sprint tickets.

The sprint does not close until the runbook either has a live controller replacing it or has a documented specification gap blocking automation. Both outcomes advance the work. One produces a controller. The other produces a cleaner specification that feeds the next sprint.

ZopDev reached 4 deleted runbooks (ZopDev, "Self-Healing Infrastructure: 4 Runbooks We Deleted After Automating Them") by treating deletion as the acceptance criterion, not deployment. Deployment of a controller is an intermediate state. Deletion confirms that the procedure requires no human execution path under any documented condition. Teams that stop at deployment accumulate controllers alongside runbooks, which creates a maintenance burden without reducing operational dependency.

Cadence, metrics, and debt

Sprint commitment size. One to two runbooks per two-week sprint is a sustainable rate for a team carrying production responsibilities. Three or more creates context-switching overhead that degrades both the automation quality and the team's incident response capacity. This works when the team has a dedicated on-call rotation separate from the sprint team. It breaks when the same engineers handling incidents are also writing controllers, because incident interruptions collapse the sprint scope unpredictably.

The review cadence. Schedule a 30-minute runbook review at the start of each sprint. Pull every runbook executed since the last review. Any procedure executed three or more times in that window is an immediate sprint candidate, because repetition frequency is a direct proxy for automation ROI. A procedure executed once per quarter recovers less engineering time per controller than one executed weekly.

Deletion as the retrospective metric. At each sprint retrospective, report one number: runbooks deleted this quarter. This works because it is binary and unambiguous. It breaks when teams count controllers deployed instead of runbooks deleted, because a deployed controller that still has a parallel manual procedure has not completed the transfer of operational responsibility.

Starting your first audit

Specification debt as a first-class ticket. Every runbook that fails triage generates a specification debt ticket, not a backlog item to revisit someday. Assign it an owner and a due date. Without ownership, specification gaps accumulate and the automation backlog stalls after the first few easy wins.

After 30 days of operating this cadence, the team's runbook inventory shrinks visibly. The procedures that remain are the genuinely non-deterministic ones, and their presence in the backlog is itself useful data about where human judgment is still irreplaceable. Audit those specifically. Some will resolve as the system matures.

Others will stay, and knowing which ones stay is operationally honest information, not a failure of the program.

The next concrete action is this: pull your incident tickets from the last 90 days, filter for tickets closed with a runbook reference, and count the unique procedures executed more than twice. That number is your automation backlog size. If it exceeds 10, start with the 3 highest-frequency procedures. Frequency beats complexity as a prioritization criterion because it maximizes recovered engineering time per sprint invested.

Frequently Asked Questions

Q: How does the hidden backlog sitting in your runbook library apply in practice?

See the section above titled "The Hidden Backlog Sitting in Your Runbook Library" for the full breakdown with examples.

Q: How does self-healing infrastructure consume runbooks apply in practice?

See the section above titled "How Self-Healing Infrastructure Consumes Runbooks" for the full breakdown with examples.

Q: How does not every runbook is ready to be automated — how to tell the difference apply in practice?

See the section above titled "Not Every Runbook Is Ready to Be Automated — How to Tell the Difference" for the full breakdown with examples.

Q: How does you should measure before and after automating a runbook apply in practice?

See the section above titled "What You Should Measure Before and After Automating a Runbook" for the full breakdown with examples.