惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
A
About on SuperTechFans
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 【当耐特】
W
WeLiveSecurity
博客园 - 三生石上(FineUI控件)
The Cloudflare Blog
I
InfoQ
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Application and Cybersecurity Blog
Application and Cybersecurity Blog
雷峰网
雷峰网
Hacker News - Newest:
Hacker News - Newest: "LLM"
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
T
Troy Hunt's Blog
S
SegmentFault 最新的问题
Help Net Security
Help Net Security
博客园_首页
博客园 - 叶小钗
O
OpenAI News
PCI Perspectives
PCI Perspectives
月光博客
月光博客
人人都是产品经理
人人都是产品经理
B
Blog RSS Feed
GbyAI
GbyAI
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
The Last Watchdog
The Last Watchdog
C
CXSECURITY Database RSS Feed - CXSecurity.com
有赞技术团队
有赞技术团队
D
Darknet – Hacking Tools, Hacker News & Cyber Security
腾讯CDC
Hacker News: Ask HN
Hacker News: Ask HN
I
Intezer
Y
Y Combinator Blog
阮一峰的网络日志
阮一峰的网络日志
Spread Privacy
Spread Privacy
T
Tailwind CSS Blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
量子位
Cyberwarzone
Cyberwarzone
The Hacker News
The Hacker News
N
News and Events Feed by Topic
P
Proofpoint News Feed
Scott Helme
Scott Helme
D
Docker
Know Your Adversary
Know Your Adversary
Recent Commits to openclaw:main
Recent Commits to openclaw:main
TaoSecurity Blog
TaoSecurity Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Tor Project blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
How Failing at Fantasy Baseball Made Me Fix My Cron Jobs with Temporal
Prithwish Na · 2026-05-06 · via DEV Community

Image

So I made a bad trade in my fantasy baseball league. Dropped Kaz Okamoto because — according to my data — he’d been cold for two weeks. In reality, he’s been on a tear for the last 9 days. 😅 This was a bad decision made because of bad data — my stats cron job had hit a rate limit, exited with no errors, and my FastAPI backend kept serving a stale JSON snapshot.

Well, I’d been meaning to fix that setup anyway. This time I did — and instead of patching the script, I tried out Temporal and…it worked embarrassingly well. Retries, backoff, execution history — things I’d normally bolt on manually were just… there. And if the network layer itself was flaky — rate limits, geo blocks — I could just add a proxy as a hardening layer.

This actually prompted me to go look at some of our production ingest jobs at work, and I thought: these are the same pattern, just with more surface area! I ended up swapping out one of them, tentatively, then another. Same pattern, just more scale.

This is my (admittedly very casual) write-up of what I learned. I hope it’s useful!

💡 I use Temporal’s Python SDK here, but they have one for TypeScript too — if that’s your thing.

How Cron Jobs Can Burn You

Here’s the brittle script I was using:

# One-shot MLB.com player fetch + write  

def main() -> None:
    player_url = os.getenv(
        "PLAYER_URL",
        "https://www.mlb.com/player/kazuma-okamoto-672960",
    )
    out_dir = Path(os.getenv("OUTPUT_DIR", "./data/runs"))
    out = out_dir / "latest.json"

    try:
        r = requests.get(player_url, timeout=60)
        if r.status_code != 200:
            print(f"WARN: HTTP {r.status_code}, leaving {out} unchanged")
            return
        stats = extract_stats_datatable(r.text)
        out_dir.mkdir(parents=True, exist_ok=True)
        out.write_text(json.dumps(stats), encoding="utf-8")
        print(f"Wrote stats to {out}")
    except Exception as e:
        print(f"WARN: ingest failed ({e!r}), leaving {out} unchanged")

if __name__ == "__main__":
    main()

Enter fullscreen mode Exit fullscreen mode

That script lived behind a super simple crontab line — once a night, fixed schedule, stdout/stderr logs:

# crontab -l (excerpt)
0 2 * * * cd /home/me/fantasy-stats && .venv/bin/python scripts/fetch_player_stats.py >> /var/log/mlb_fetch.log 2>&1

Enter fullscreen mode Exit fullscreen mode

The script is about ~20 lines of Python plus one line of schedule. How many potential failures can you spot here?

The first is the thing I thought was prudent: if the fetch looks wrong, don’t overwrite the snapshot. So any 429, timeout, or 200 with a layout that no longer contains the marker extract_stats_datatable expects becomes a printed WARN, a no-op, and main() returns — exit code 0. No raise_for_status(); no sys.exit(1). Cron is “happy”; the one-line warning vanishes in a log I wasn’t tailing; latest.json never updates. Nine “successful” runs later…I made a bad decision because I had bad data. (Flip it to raise_for_status() and you get the opposite smell: a non-zero exit, still no retry, still stale data until someone fixes the feed — pick your poison 😬)

The other two are subtler — and I didn’t personally run into them — but upon review, they were just as likely to have burnt me.

  • Fixed output path. Every run writes to latest.json. If two runs overlap — which happens the moment a run is slow and the next cron tick fires — it’s a race condition. One overwrites the other mid-write. You might read a corrupted file, or never know which run's data you actually have.
  • Non-atomic write. out.write_text() is not atomic. If the process dies mid-write — OOM, signal, anything — you get a partial JSON file. The next reader gets a parse error and now has to figure out if the file is corrupted or just empty. This is the exactly kind of bug that shows up at 2am on a production system.

The real problem isn’t any ONE of these — it’s really that cron gives you exactly one bit of feedback: exit zero or exit non-zero — and as this script shows, exit zero can lie. It can’t give you a retry policy, overlap protection, or artifact history. No way to answer “what did this job actually do at 3am last Tuesday?”

And yes, you can try to patch around that. Add a retry loop and exponential backoff, ship logs somewhere. But now your retry state lives in the process memory that disappears on crash, your backoff is hand-rolled, and your observability is still pretty much log spelunking. At that point you’re not using cron anymore. You’re rebuilding a tiny, worse workflow engine around cron.

What is Temporal?

Temporal is a “durable execution platform”. What that really means in practice is that you’ll write ordinary functions — a Workflow that orchestrates things, and Activities that do the actual work — and Temporal will:

  • Make the execution survive process crashes,
  • Retry failed steps with backoff,
  • Prevent overlapping runs, and
  • Record the full history of every execution.

💡 Durable execution is the simple idea that your code should keep running to completion even if the machine running it doesn’t. The mental model is that your workflow is a function call that cannot be interrupted, even if the worker reboots halfway through. State lives in Temporal’s history, not in the worker’s memory.

The architecture for our project will look like this (by default you get the Temporal Web UI on localhost:8233):

To grok Temporal properly, understand that the Workflow owns when things happen, while the Activity owns what happens — i.e the page fetch, the stats extraction, the file write. Workflows must stay deterministic; all side effects belong in Activities. Why does that matter? We’ll come back to that in a second.

Getting player data from MLB.com

Before any of the Temporal machinery, you need a reliable data ingestion. For us, that means loading an MLB.com player page and extracting the stats blob embedded in the initial HTML.

MLB.com currently renders a player page with a JavaScript object that starts with stats: {"statsDatatable"...}. That's convenient: no browser, no Playwright, no screenshot automation needed.

Critical to understand that this does not tell you that the stats blob is still there or that the row you care about was parsed correctly.

mlb_player_stats.py

"""  
MLB.com player page: HTTP fetch + embedded JSON extraction.  
"""  
from __future__ import annotations  
import os  
import re  
from json import JSONDecoder  
from typing import Any, Dict, List  
from urllib.parse import quote  
import requests  

def _strip_tags(s: Any) -> Any:  
    if not isinstance(s, str):  
        return s  
    s = re.sub(r"`<[^>`]+>", "", s)  
    return s.strip()  

def _sanitize_row(row: Dict[str, Any]) -> Dict[str, Any]:  
    out: Dict[str, Any] = {}  
    for k, v in row.items():  
        out[k] = _strip_tags(v)  
    return out  

def extract_stats_datatable(html: str) -> Dict[str, Any]:  
    needle = 'stats: {"statsDatatable"'  
    i = html.find(needle)  
    if i == -1:  
        raise ValueError(  
            "Could not find stats JSON marker (page layout may have changed)."  
        )  
    start = i + len("stats: ")  
    obj, _ = JSONDecoder().raw_decode(html[start:])  
    return obj  

def pick_current_season_row(rows: List[Dict[str, Any]]) -> Dict[str, Any] | None:  
    for row in rows:  
        h = row.get("header", "")  
        if isinstance(h, str) and "Regular Season" in h and "Career" not in h:  
            return row  
    return rows[0] if rows else None  

def build_requests_proxies() -> Dict[str, str] | None:  
    # Bright Data super proxy; returns None if credentials unset  
    explicit = os.getenv("BRIGHT_DATA_PROXY_URL", "").strip()  
    if explicit:  
        return {"http": explicit, "https": explicit}  
    host = os.getenv("BRIGHT_DATA_PROXY_HOST", "brd.superproxy.io").strip()  
    port = os.getenv("BRIGHT_DATA_PROXY_PORT", "33335").strip()  
    username = os.getenv("BRIGHT_DATA_PROXY_USERNAME", "").strip()  
    password = os.getenv("BRIGHT_DATA_PROXY_PASSWORD", "").strip()  
    if not username or not password:  
        return None  
    user_enc = quote(username, safe="")  
    pass_enc = quote(password, safe="")  
    proxy_url = f"http://{user_enc}:{pass_enc}@{host}:{port}"  
    return {"http": proxy_url, "https": proxy_url}  

def fetch_player_page(player_url: str, *, timeout: int = 60) -> requests.Response:  
    proxies = build_requests_proxies()  
    return requests.get(  
        player_url,  
        timeout=timeout,  
        proxies=proxies  
    )  

def build_stats_payload(player_url: str, *, timeout: int = 60) -> Dict[str, Any]:  
    # Fetch page, parse embedded hitting summary rows (sanitized)  
    r = fetch_player_page(player_url, timeout=timeout)  
    r.raise_for_status()  
    blob = extract_stats_datatable(r.text)  
    hitting_large = blob["statsDatatable"]["hitting"]["large"]  
    block = hitting_large[0] if isinstance(hitting_large, list) else hitting_large  
    filtered = block["filteredRows"]  
    current = pick_current_season_row(filtered)  
    career = next(  
        (row for row in filtered if row.get("header") == "Career Regular Season"),  
        None,  
    )  
    via_proxy = build_requests_proxies() is not None  
    return {  
        "source_url": player_url,  
        "http_status": r.status_code,  
        "via_bright_data_proxy": via_proxy,  
        "current_regular_season": _sanitize_row(current) if current else None,  
        "career_regular_season_row": _sanitize_row(career) if career else None,  
        "all_summary_rows": [_sanitize_row(row) for row in filtered],  
    }

Enter fullscreen mode Exit fullscreen mode

Keeping the fetch inside an activity means the execution model stays unchanged while the network path can evolve independently. The same fetch_player_page call can run directly or be routed through a proxy layer without touching the workflow logic.

Note that Temporal can only give you execution reliability: retries, timeouts, and visibility. It can not make a failing network succeed. If every attempt returns 429 from the same IP, Temporal will reliably retry a failing request until the policy is exhausted — you won’t know how valuable our proxy layer is until you really need it (if you’re following along, get it here.)

Data Extraction with Temporal.io

The Temporal Workflow

@workflow.defn  
class StatsCollectionWorkflow:  
    @workflow.run  
    async def run(self, job: StatsJob) -> CollectStatsResult:  
        info = workflow.info()  
        return await workflow.execute_activity(  
            collect_stats,  
            CollectStatsInput(  
                player_url=job.player_url,  
                workflow_id=info.workflow_id,  
                run_id=info.run_id,  
                output_dir=job.output_dir,  
            ),  
            start_to_close_timeout=timedelta(minutes=10),  
            retry_policy=RetryPolicy(  
                initial_interval=timedelta(seconds=3),  
                backoff_coefficient=2.0,  
                maximum_interval=timedelta(minutes=2),  
                maximum_attempts=8,  
            ),  
        )

Enter fullscreen mode Exit fullscreen mode

That RetryPolicy block is the part most of us have written manually at some point — a while loop, a try/except, a time.sleep, a counter, hopefully a max attempts check. Here it's declared once, lives outside the business logic, and survives worker crashes. If the worker process dies on attempt 3 of 8, the next worker that comes up picks up at attempt 4. The state is in Temporal, not in memory.

The start_to_close_timeout is the hard limit on how long a single activity attempt can run. Without it, a stalled HTTP request holds a worker slot indefinitely. Decidedly not what we want.

The Temporal Activity

Here’s activities.py:

"""Activities: MLB.com fetch + stats extraction + artifact write (all side effects here)."""  
from __future__ import annotations  
import json  
import os  
from dataclasses import dataclass  
from pathlib import Path  
from typing import Any, Dict, Union  
from temporalio import activity  
from temporal_cron.mlb_player_stats import build_stats_payload  

@dataclass  
class CollectStatsInput:  
    player_url: str  
    workflow_id: str  
    run_id: str  
    output_dir: str  

@dataclass  
class CollectStatsResult:  
    artifact_path: str  
    home_runs: Union[str, int]  
    player_url: str  

def _atomic_write_json(path: Path, data: Dict[str, Any]) -> None:  
    path.parent.mkdir(parents=True, exist_ok=True)  
    tmp = path.with_suffix(path.suffix + ".tmp")  
    tmp.write_text(json.dumps(data, indent=2), encoding="utf-8")  
    tmp.replace(path)  

@activity.defn  
def collect_stats(input: CollectStatsInput) -> CollectStatsResult:  
    # One HTTP try per activity attempt; workflow RetryPolicy owns backoff/attempts.  
    data = build_stats_payload(input.player_url, timeout=60)  
    current = data.get("current_regular_season") or {}  
    hr = current.get("homeRuns", 0)  
    base = Path(input.output_dir or os.getenv("OUTPUT_DIR", "./data/runs"))  
    safe_wid = input.workflow_id.replace(os.sep, "_").replace(":", "_")  
    safe_rid = input.run_id.replace(os.sep, "_").replace(":", "_")  
    out_path = base / f"{safe_wid}__{safe_rid}.json"  
    payload = {  
        "workflow_id": input.workflow_id,  
        "run_id": input.run_id,  
        "player_url": input.player_url,  
        "data": data,  
    }  
    _atomic_write_json(out_path, payload)  
    return CollectStatsResult(  
        artifact_path=str(out_path.resolve()),  
        home_runs=hr if hr is not None else 0,  
        player_url=input.player_url,  
    )

Enter fullscreen mode Exit fullscreen mode

The output path uses the run ID, not a fixed filename. Every execution gets its own artifact — stats-manual-abc123__run456.json. No races, no overwrites, and you have a full history of every run. You can diff two runs. You can see exactly what data you had on any given night. This alone would have saved me.

The file write is atomic: _atomic_write_json (in the excerpt above) writes to a .tmp file first, then replace() on the same filesystem. A reader either sees the old file or the new file — never a partial write. The brittle script calls write_text() directly; if the process died mid-write, you got corrupted JSON and a confusing parse error at the worst possible time.

💡 Note that collect_stats is a sync function, not async. That's intentional — sync activities run in a thread pool, so blocking I/O doesn't block the event loop. The Temporal SDK supports both; sync is the right call when your activity is mostly waiting on a network request.

The Temporal Worker

async def _main() -> None:  
    client = await Client.connect(host, namespace=namespace)  
    worker = Worker(  
        client,  
        task_queue=task_queue,  
        workflows=[StatsCollectionWorkflow],  
        activities=[collect_stats],  
    )  
    await worker.run()

Enter fullscreen mode Exit fullscreen mode

The Worker is the only process that ever touches MLB.com. The workflow and scheduler are pure orchestration — they tell Temporal what to do, they don’t do any work themselves. You can scale workers horizontally without touching the scheduling layer. You’ll want that when you take this pattern to production.

The worker view shows the stats-pipeline worker polling alongside Temporal's own system worker.

The Temporal Schedule

Putting this on a schedule is one function call:

schedule = Schedule(  
    action=ScheduleActionStartWorkflow(  
        StatsCollectionWorkflow.run,  
        job,  
        id=workflow_id,  
        task_queue=task_queue,  
    ),  
    spec=ScheduleSpec(cron_expressions=[cron]),  
    policy=SchedulePolicy(overlap=ScheduleOverlapPolicy.SKIP),  
)

Enter fullscreen mode Exit fullscreen mode

SKIP is the thing cron simply cannot do.

So, if MLB.com is slow today, or your fetches are getting rate-limited harder than usual, and your run takes 12 minutes. Your schedule fires every 15. Eventually the slow run bleeds into the next tick. With cron, you now have two instances running simultaneously, both writing to latest.json, racing each other. With SKIP, the new scheduled run sees the previous one is still active and does nothing. When the previous run finishes, the schedule resumes normally at the next tick. That's an entire class of bug you stop thinking about.

The schedule script also handles create-or-update correctly — describe() first, catch NOT_FOUND, then either create or update:

try:  
    await handle.describe()  
except RPCError as err:  
    if err.status != RPCStatusCode.NOT_FOUND:  
        raise  
    await client.create_schedule(schedule_id, schedule)  
else:  
    await handle.update(lambda _input: ScheduleUpdate(schedule=schedule))

Enter fullscreen mode Exit fullscreen mode

Run it once to create the schedule. Run it again to change the cron expression or job parameters. Same command either way.

What you actually see in the UI

Pop open http://localhost:8233 while a workflow is running. Every step is there — which activity ran, how many attempts it took, what the retry intervals were, what came back. If it failed on attempt 2 and succeeded on attempt 5, you can see that. You can see the exact input that went in and the exact output that came out. You can see how long each attempt took.

The workflow list gives you the first-level answer cron never gives cleanly: what ran, when, and whether it completed.

The timeline connects the workflow input, activity execution, and output artifact in one place. The event history is the audit trail: scheduled tasks, activity start/completion, workflow task transitions, and final result.

The activity details will also show the successful third attempt and preserve the previous failure: 429 Too Many Requests.

Temporal activity event showing attempt #3 with previous HTTP 429 failure details

Compare that to the cron version — just a process exit code and whatever you print() to stdout. If the job failed at 3am and you weren't tailing logs, that information is gone.

This observability is honestly why teams end up on Temporal even for jobs that aren’t that complicated. This removes entire categories of debugging work. No log spelunking to figure out whether a retry happened. No guessing how many attempts ran. No reconstructing a timeline from scattered stdout. You won’t have to infer from fragments; all this is something you can just look up.

Running it yourself

Commands below assume uv is available with a venv set up.

# Start the local Temporal server  
temporal server start-dev  
# In another terminal, start the worker  
uv run temporal-cron-worker  
# Trigger a single run  
uv run temporal-cron-start  
# Or put it on a schedule (defaults to every 15 minutes, set SCHEDULE_CRON in .env to change)  
uv run temporal-cron-schedule

Enter fullscreen mode Exit fullscreen mode

Open http://localhost:8233 and you'll see the workflow execution. Artifacts land under data/runs/, one file per run, named by workflow and run ID.

A Note on Production

This just runs against a local Temporal dev server with a single worker. Taking it to production means picking Temporal Cloud or running your own cluster, adding proper secrets management, structured logging, metrics, and deployment automation.

Cron is all you need for simple, isolated tasks. Probably not so much when you have to introduce retries, external dependencies, or jobs that can overlap or run longer than their schedule.

When you get into that zone, Temporal gives you a durable execution layer around ingest: retries, timeouts, overlap control, and a history you can audit. And then when that layer needs to scale up or if you start hitting IP-based (or geo-based, even) friction, Bright Data’s proxies give you the hardened network path you’ll need. It’s been a pretty natural pairing, in my experience. You can route the same requests.get() call through a proxy and let Temporal keep owning retries, timeouts, and audit history.

Regardless, this pattern is a cheat code right here for Temporal: Workflows own orchestration, and Activities own side effects.