惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
Stack Overflow Blog
Stack Overflow Blog
B
Blog RSS Feed
C
Check Point Blog
D
Docker
Y
Y Combinator Blog
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
MongoDB | Blog
MongoDB | Blog
博客园_首页
Apple Machine Learning Research
Apple Machine Learning Research
量子位
有赞技术团队
有赞技术团队
IT之家
IT之家
大猫的无限游戏
大猫的无限游戏
D
DataBreaches.Net
M
MIT News - Artificial intelligence
B
Blog
阮一峰的网络日志
阮一峰的网络日志
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
腾讯CDC
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
V
V2EX
月光博客
月光博客

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
Error Budgets in Practice: A No-BS Guide
Samson Tanimawo · 2026-06-26 · via DEV Community
Cover image for Error Budgets in Practice: A No-BS Guide

Samson Tanimawo

Everyone Talks About SLOs, Nobody Talks About Error Budgets

Every SRE conference has talks about SLOs. Set a target! 99.9%! Three nines! Standing ovation.

But the actually useful part — what you DO when you're burning your error budget — rarely gets discussed.

What's an Error Budget, Really?

If your SLO is 99.9% availability, your error budget is 0.1%. That's:

Monthly error budget at 99.9%:
  43.2 minutes of downtime

Or in request terms (1M requests/day):
  ~1,000 failed requests per day
  ~30,000 failed requests per month

The error budget isn't a target to hit. It's fuel you spend on shipping features.

The Error Budget Policy

This is the document that actually matters. Ours looks like this:

error_budget_policy:
  budget_remaining:
    above_50_percent:
      deploy_frequency: "unlimited"
      feature_vs_reliability: "80/20"
      review_cadence: "monthly"
    25_to_50_percent:
      deploy_frequency: "2x daily max"
      feature_vs_reliability: "50/50"
      review_cadence: "weekly"
    10_to_25_percent:
      deploy_frequency: "1x daily, with canary"
      feature_vs_reliability: "20/80"
      review_cadence: "daily standup"
    below_10_percent:
      deploy_frequency: "emergency only"
      feature_vs_reliability: "0/100"
      review_cadence: "war room until recovered"
    exhausted:
      action: "feature freeze until budget replenishes"
      escalation: "VP Engineering notified"

Making It Real: The Dashboard

We built a simple burn rate dashboard:

def calculate_burn_rate(slo_target, window_hours, error_count, total_count):
    """Calculate how fast we're burning error budget."""
    error_rate = error_count / total_count
    budget = 1 - slo_target  # e.g., 0.001 for 99.9%

    # Burn rate: how many times faster than allowed
    # burn_rate of 1.0 = exactly on budget
    # burn_rate of 2.0 = burning 2x too fast
    burn_rate = error_rate / budget

    # Time until budget exhausted at current rate
    budget_remaining = budget - error_rate
    hours_left = (budget_remaining / error_rate) * window_hours if error_rate > 0 else float('inf')

    return {
        'burn_rate': round(burn_rate, 2),
        'hours_until_exhausted': round(hours_left, 1),
        'budget_consumed_pct': round((error_rate / budget) * 100, 1)
    }

# Example
result = calculate_burn_rate(
    slo_target=0.999,
    window_hours=24,
    error_count=500,
    total_count=1_000_000
)
print(result)
# {'burn_rate': 0.5, 'hours_until_exhausted': 48.0, 'budget_consumed_pct': 50.0}

The Hardest Part: Enforcing the Freeze

When you tell a product manager "no more features this month because we burned our error budget," expect pushback. Here's how to handle it:

  1. Make it data-driven: Show the burn rate chart. Numbers don't argue.
  2. Connect to money: "Our SLO breach cost us $X in SLA credits last quarter."
  3. Make it a team agreement: The error budget policy should be signed off by engineering AND product leadership.
  4. Automate the gates: CI/CD should block non-critical deploys automatically.
# Example: GitLab CI gate
deploy_production:
  rules:
    - if: '$ERROR_BUDGET_REMAINING < 10'
      when: manual  # Require manual approval
      allow_failure: false
    - when: on_success  # Auto-deploy when budget healthy

The Cultural Shift

The real value of error budgets isn't technical. It reframes reliability from "SRE's job" to "everyone's job." When developers see that their buggy deploy consumed 30% of the monthly error budget, they start writing better tests.

It took us about two quarters to get this cultural shift, but now our product teams actively ask about error budget status before planning sprints.

If you're struggling to operationalize SLOs and error budgets, check out what we're building at Nova AI Ops.


Written by Dr. Samson Tanimawo
BSc · MSc · MBA · PhD
Founder & CEO, Nova AI Ops. https://novaaiops.com