惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
Engineering at Meta
Engineering at Meta
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 聂微东
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Last Week in AI
Last Week in AI
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 三生石上(FineUI控件)
T
Tailwind CSS Blog
Apple Machine Learning Research
Apple Machine Learning Research
Hugging Face - Blog
Hugging Face - Blog
爱范儿
爱范儿
博客园 - 司徒正美
人人都是产品经理
人人都是产品经理
Jina AI
Jina AI
博客园 - 叶小钗
雷峰网
雷峰网
罗磊的独立博客
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - Franky
WordPress大学
WordPress大学
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
阮一峰的网络日志
阮一峰的网络日志
量子位

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
Output assertions: the cron job check most monitoring too...
Kriss · 2026-04-29 · via DEV Community

Output assertions: the cron job check most monitoring tools skip

A follow-up to A reader comment made me realise I'd only solved half the problem — this is a deeper reference guide on output assertions specifically.

"Did it run?" is the wrong question.

Every monitoring tool asks it. Heartbeat monitors, cron schedulers, even purpose-built tools like Cronitor and Healthchecks.io — they all fundamentally ask: did the job check in? If yes, green. If no, red.

It's a useful question. But it's not the useful question.

The failure mode that looks like success

Imagine a nightly job that syncs user records from your CRM into your database. It runs at midnight, takes about 90 seconds, and exits cleanly. Your heartbeat monitor sees the ping at 12:01:34am and marks it healthy.

What it doesn't see: the job synced 0 records. It has been syncing 0 records for eight days, since someone rotated the CRM API credentials and forgot to update the environment variable. The job connects, gets a 401, logs a warning, falls back to a no-op, and exits 0.

All monitoring: green. Business: broken for eight days.

This is not a hypothetical. Variants of this failure happen constantly. The job ran. That fact is true and also completely useless.

What "did it do anything?" looks like

Output assertions flip the question. Instead of only checking that the job pinged in, you also check what it reported.

A job that processes records should report how many it processed. A job that generates a file should report the file size. A job that sends emails should report how many it sent. You instrument the job to emit a count — one number representing meaningful work done — and your monitoring layer validates it falls within expected bounds.

The failure modes this catches:

  • Zero when non-zero expected: sync runs, processes nothing, exits clean
  • Suspiciously low counts: normally syncs 500 records, today synced 3
  • Count drift over time: weekly report used to include 10k rows, now consistently 200

None of these trip a heartbeat check. All of them are real problems.

Why most tools don't do this

Heartbeat monitoring is architecturally simple: job pings URL, URL records timestamp, alerting checks timestamp age. The data model is just "last seen at".

Output assertions require more: the job must emit structured data, the tool must store it, and the alerting logic must understand what "normal" looks like for that specific job. That's a significantly more complex product to build.

Most tools solve the simpler problem because it covers the obvious failure mode and is much easier to ship.

How to instrument your jobs

The instrumentation is lightweight. Pick a number that represents meaningful work and emit it at the end:

# Database backup — report dump file size
result = subprocess.run(["pg_dump", "-Fc", "mydb", "-f", "/backups/mydb.dump"])
dump_size = os.path.getsize("/backups/mydb.dump")
ping_monitor(count=dump_size)

# CRM sync — report records synced
synced = sync_from_crm()
ping_monitor(count=len(synced))

# Email campaign — report emails sent
sent = send_campaign(campaign_id)
ping_monitor(count=sent)

Enter fullscreen mode Exit fullscreen mode

Three extra lines per job. The return is knowing your job didn't just run — it did something. (ping_monitor is a wrapper around your monitoring call — implementation below.)

Sending the count to your monitor

DeadManCheck accepts a count parameter with each ping:

curl -fsS "https://deadmancheck.io/ping/YOUR-TOKEN?count=1547" > /dev/null

Enter fullscreen mode Exit fullscreen mode

You configure the assertion on the monitor: "alert if count is 0" or "alert if count drops below threshold". If the job checks in but reports zero records, you get alerted — even though the job technically ran fine.

It also does duration monitoring with rolling average anomaly detection. If your 90-second job starts taking 45 minutes, that gets flagged too. Jobs that hang are a separate silent failure mode that output counts don't catch on their own.

The right question

Monitoring that only asks "did it run?" will eventually lie to you at the worst possible moment.

The right question is "did it do anything useful?" Output assertions are how you ask that question automatically, at 2am, every night, without anyone having to check.

Start with your backup jobs. That's where the answer matters most.