惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
Martin Fowler
Martin Fowler
云风的 BLOG
云风的 BLOG
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
The Blog of Author Tim Ferriss
大猫的无限游戏
大猫的无限游戏
A
About on SuperTechFans
小众软件
小众软件
博客园_首页
博客园 - 聂微东
罗磊的独立博客
Recent Announcements
Recent Announcements
U
Unit 42
N
Netflix TechBlog - Medium
Blog — PlanetScale
Blog — PlanetScale
阮一峰的网络日志
阮一峰的网络日志
博客园 - 叶小钗
V
V2EX
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
IT之家
IT之家
Stack Overflow Blog
Stack Overflow Blog
博客园 - Franky
D
DataBreaches.Net
Last Week in AI
Last Week in AI

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
The Node.js Mistake That Cost My Client $3,000 in AWS Bills
Lolo · 2026-06-23 · via DEV Community

Lolo

Last year I was asked to investigate a startup's AWS bill.

It had jumped from roughly $200/month to over $3,000 in a few weeks.

Nobody knew why.

After digging through logs, metrics, and database traffic, I found the culprit: a polling loop with no backoff strategy.

The code looked harmless:

async function processQueue() {
  const jobs = await getJobs()

  for (const job of jobs) {
    await processFile(job)
  }

  processQueue()
}

processQueue()

At first glance, this seems reasonable. Process all available jobs, then check again.

The problem appears when the queue is empty.

When getJobs() returned no work, the loop immediately queried the database again. And again. And again.

There was no delay, no backoff, and no event-driven trigger.

As a result, the service continuously hammered the database looking for work that didn't exist.

Each iteration generated:

  • A database query
  • Network traffic
  • CPU usage
  • Logging overhead
  • Additional infrastructure load

Individually, each operation was cheap.

Executed hundreds of thousands of times per day, they became expensive.

The fix was simple:

async function processQueue() {
  while (true) {
    const jobs = await getJobs()

    for (const job of jobs) {
      await processFile(job)
    }

    await new Promise(resolve => setTimeout(resolve, 5000))
  }
}

Even better would have been replacing polling entirely with an event-driven design using a message queue.

What this incident taught me:

1. Empty queues are production workloads.

Many engineers optimize for peak traffic and forget about idle traffic. Systems often spend more time idle than busy.

2. Polling needs backoff.

If you're polling, always define what happens when no work is found.

3. Cost bugs rarely look like bugs.

Nothing crashed. No exceptions were thrown. The system was technically working exactly as written.

It was just doing useless work 24/7.

4. Always monitor cost alongside performance.

CPU, latency, and error rates looked normal.

The AWS bill was the first real alert.

One question I ask during reviews now:

"What does this code do when there's nothing to do?"

That single question has caught more production issues than many architecture discussions ever did.

What's the most expensive bug you've ever seen in production?