惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

腾讯CDC
GbyAI
GbyAI
C
CERT Recently Published Vulnerability Notes
S
Security @ Cisco Blogs
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
The Hacker News
The Hacker News
D
Darknet – Hacking Tools, Hacker News & Cyber Security
P
Proofpoint News Feed
C
Cyber Attacks, Cyber Crime and Cyber Security
S
Securelist
Security Latest
Security Latest
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Simon Willison's Weblog
Simon Willison's Weblog
Latest news
Latest news
T
Tor Project blog
T
Threat Research - Cisco Blogs
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Spread Privacy
Spread Privacy
K
Kaspersky official blog
T
The Exploit Database - CXSecurity.com
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
V2EX - 技术
V2EX - 技术
L
Lohrmann on Cybersecurity
Google Online Security Blog
Google Online Security Blog
Cyberwarzone
Cyberwarzone
Help Net Security
Help Net Security
The Last Watchdog
The Last Watchdog
C
Cybersecurity and Infrastructure Security Agency CISA
Attack and Defense Labs
Attack and Defense Labs
大猫的无限游戏
大猫的无限游戏
Schneier on Security
Schneier on Security
H
Heimdal Security Blog
O
OpenAI News
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
Visual Studio Blog
AI
AI
H
Hacker News: Front Page
博客园_首页
博客园 - 【当耐特】
L
LINUX DO - 最新话题
MyScale Blog
MyScale Blog
量子位
Vercel News
Vercel News
C
Cisco Blogs
L
LINUX DO - 热门话题
Y
Y Combinator Blog
T
Threatpost
爱范儿
爱范儿
P
Privacy & Cybersecurity Law Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The Drift from Chat to Backlog: How My AI Task Planning Evolved Over Three Months
Nic Lydon · 2026-06-17 · via DEV Community

Three months ago, my entire task-management system was a chat window I'd lose when the tab closed. Today it's a Postgres backlog that three different coding agents — Claude Code, Codex, Grok — pull work off autonomously, stamp with attribution, and close against git history. I never decided to build a project-management system. I just kept hitting a wall, patching it, and hitting the next one the patch exposed.

There's a clean way to read the whole arc, though, and it comes down to a single variable: where the plan lives. Watch that, and every step makes sense — including why you probably want to stop well before the end.

The setup

I run a self-hosted personal data platform called Nexus, on a 128GB Strix Halo box named Furnace, surrounded by ~100 repos: MCP servers, ingestion pipelines, iOS apps, content tooling. My execution tools are Claude Code and Codex CLI. The work is bursty — during one 35-day stretch in spring I shipped roughly 557K lines across those repos — and that throughput is the pressure that broke each planning approach in turn. At a calmer pace you'll hit the same walls later, but you'll hit them.

Phase 1: The plan lives in the chat (mid-to-late March)

At the start there was no task system. Planning was the conversation. I'd open a chat, think out loud about an architecture problem, get to something coherent, and then go build it. The artifact of planning was a better mental model in my head, not a written thing.

This is, for the record, exactly what every best-practice guide tells you to do, and it's right. The math is brutal and well-known: if Claude makes the right call 80% of the time on any single decision, a feature with 20 decision points lands all 20 at 0.8^20 — about 1%. Planning collapses those 20 live decisions into a reviewed spec where each one is already made. I'd never give up the plan-first instinct; it's the one habit from this whole story that never changed.

The problem was narrower: the plan evaporated when the thread ended. A late-March session designing a multi-agent system for Nexus produced genuinely good architecture — deterministic behavior under load, agents that self-regulate instead of spiraling, adaptive thresholds. None of it was anywhere I could act on the next morning except my memory and a scrollback buffer. At one-feature-a-day that's survivable. At my pace it was lossy in a way that actively cost me work.

The wall: plans that exist only in chat history can't be acted on later, can't be prioritized against each other, and can't be handed to anything but the version of you that remembers the conversation.

Phase 2: The plan lives in a file (mid-to-late April)

The first durable fix was embarrassingly simple: a TODO.md in each repo. But the structure I landed on is the part worth stealing, because it wasn't a checklist. Each item was a small spec. Here's a real one, still in my broadside/TODO.md:

## Idempotency on publish operations

**Status:** captured — flagged during the 2026-04-26 reality-sync session.
**Trigger:** before letting an agent publish unsupervised at any volume.

Today, POST /api/posts/[id]/twitter (and the bluesky / linkedin / devto
siblings) don't refuse a re-publish. If a publish succeeds upstream but the
response gets lost in a network blip, an agent's retry would publish a
duplicate — visibly, to the audience.

**What the work will involve:**
1. Before calling upstream: SELECT posted_url, status FROM posts WHERE id=$1.
   If already posted, refuse with 409 + { already_posted, posted_url }.
2. Optional: accept an Idempotency-Key header, TTL'd table, ~10min window.
3. Update the broadside_publish_to_platform MCP tool description accordingly.

**Risks worth knowing:**
- The "republish" path is sometimes intentional (manual delete + re-fire).
  Mitigate with a posted_at recency check or a force=true param.

**Why it matters:** an agent that double-posts even once is visible to the
audience; the cost is reputational, not just internal.

Four things made this carry weight beyond a checkbox:

  • Status + date. Every item is stamped with when it was captured. Trivial-sounding; it's what lets you reason about staleness later.
  • Trigger, not priority. Instead of P1/P2/P3, each item records the condition under which it becomes urgent. "Before letting an agent publish unsupervised" tells future-me exactly when to pull this off the shelf; "high priority" tells me nothing.
  • The work is pre-decided. That numbered list is a plan captured at the moment I understood the problem best. Hand it to Claude Code three weeks later and the 20 decisions are already made — Phase 1's whole value, persisted.
  • Risks are written down. The "republish is sometimes intentional" note is exactly the edge case I'd have forgotten and an agent would have trampled.

Notice the recurring phrase: "reality-sync session." Concretely, that was a 20-minute pass, usually before a planning block: open each repo's TODO.md next to its recent git log, close anything the commits show as already shipped, and re-date anything still open so I could tell stale from live at a glance. Reconciling the plan against ground truth on a cadence — that habit turns out to be the seed of everything in Phase 3.

The wall: TODO.md is per-repo. With ~100 repos, I had no single surface that could tell me what to work on next across all of them, no way to prioritize globally, and nothing an agent could pull from as a queue. The plan was durable but fragmented.

Phase 3: The plan lives in a database with a gate (early-to-mid May onward)

This is the structural leap. Task state moved out of flat files and into Nexus's Postgres as the Operator Backlog (OB) — with a real intake-to-execution lifecycle instead of a list.

The shape:

  • Work enters as a candidate, in pending state — not yet a real task. A candidate won't become an OB item until I approve it in a #pmo-review Discord flow.
  • Approval mints an OB-##### row in operator_backlog_items.
  • Items move through status lanesrequires_triagerequires_decisionrequires_investigationrequires_clickopsautonomous_safe — and agents drain those lanes.
  • Every commit references its OB id, joining the backlog directly to git history. OB-27081 H2/M2 — close /register on RS is a real commit subject across several of my repos.

The lanes are the interesting part, and they map almost exactly onto a distinction Anthropic draws in Building Effective Agents: the difference between work an agent can drive autonomously and work that needs a human checkpoint "before irreversible actions." autonomous_safe is the lane an agent can just do. requires_decision is the lane that needs me. The backlog isn't just storage; it's a router that sorts work by how much human judgment it still needs.

The single most important addition in this phase wasn't the lanes, though. It was the approval gate. Phase 2 captured indiscriminately — anything I typed that looked like a task got written down. Phase 3 added a filter, and I know exactly why, because I watched it fail without one. From a real session log on June 3:

Drained requires_triage + requires_decision queues (19 items → autonomous_safe/closed/ignored); 8 decisions made; discovered auto-filer over-captures on soft prose in narration (e.g. "blocked on" → filed OB-4736 to tighten regex); priority-lane starvation (skillopt-train starving PM jobs) diagnosed in OB-4715.

An automated process was scraping my session narration for tasks and mistaking the phrase "blocked on" — used conversationally — for a real blocker. The system was filing garbage into its own backlog. The fix (OB-4736) was itself filed as an OB item, through the gate. The backlog had become self-correcting: its own intake bugs are tracked in the same substrate as everything else.

That same log entry shows the daily rhythm this phase settled into. "Draining queues" became a literal, recurring operation — pull the items in a lane, make the decisions, move them to autonomous_safe or close them. Eight decisions in one sweep. It reads like an on-call shift, because that's effectively what it is: I'm the operator, the backlog is the queue.

The wall: a database-backed queue with a governance gate is great, but it assumes disciplined intake and it assumes someone drains the lanes. As volume grew, I was the bottleneck. The backlog could hold more work than I could personally execute.

Phase 4: Many agents drain the backlog (late May → now)

The most recent shift isn't about how tasks are tracked — Phase 3 settled that. It's about who executes them, and how you stay accountable when it's not just you.

The OB backlog is now a shared work queue that multiple runners pull from: Claude Code, Codex, and Grok, each tagged so I can tell after the fact which agent did which item. The same status lanes from Phase 3 keep it safe — an agent only picks up work already in autonomous_safe, never anything still sitting in requires_decision. Conflict is handled by leases: a runner claims an item, a second runner sees it's taken and skips it, and if the first agent dies the lease expires and the item frees itself.

Attribution stamping is the load-bearing piece. Because each commit and OB resolution is tagged by runner, "who did this and why" stays answerable even with three agents touching ~100 repos — and the rule is that the human owns the commit while the runner tag lives in the backlog, never forged into git history. Each run executes in its own isolated git worktree, so parallel agents never touch the same files.

That's the whole machine. Here's the moment it stopped being theory for me. OB-1623 was "wire a model-provenance footer into report delivery." Claude Code claimed it on May 30, started working — and then refused to finish it. It had discovered the task's premise was wrong: the function the task named was a shared primitive with ten callers, and the files it pointed at didn't even hold the data the footer needed. Instead of forcing a fix that would have quietly broken nine other call sites, it blocked the item, filed a corrected prerequisite as a new OB, and released its claim with a note explaining exactly why. Two days later, after the prerequisite landed, Codex picked up the same OB cold — no shared memory with the Claude Code run, just the backlog row and the blocking note — and shipped it end to end: PR merged, Nexus deployed to both Furnace and Crucible at c584d2a8, the live footer renderer verified emitting the right markers.

Read that sequence again, because it's the whole point of Phase 4 in one item. Two different models, no human mediating the handoff, and the system self-corrected across the gap — one agent's refusal to do the wrong thing became another agent's clean win, because the reason for the refusal was written down in the one place both of them could see. That's not parallelism. That's the backlog doing the thing a good engineering team does: catching a bad assumption before it ships, and carrying the correction forward to whoever picks the work up next.

The pattern underneath

Here's the whole arc as a table. Read the "Fixed" and "New limit" columns as a chain — each phase's new limit is the next phase's reason to exist.

Phase Where the plan lives Fixed New limit
1. Conversational Chat history Decision quality (plan-first) Plans evaporate
2. TODO.md Per-repo files Persistence No global view or priority
3. OB / PMO Postgres + approval gate Global queue, governance, routing Needs disciplined intake; you drain it
4. Multi-agent OB Backlog + worktrees + attribution Parallel execution, accountability Coordination overhead

Two things are worth pulling out of that table.

First: you do not need to reach Phase 4 to get most of the value. Take the Phase 2 TODO.md format — status, trigger, pre-decided steps, risks — and nothing else. It's a text file; it costs nothing; and everything after it is just scaling that same captured-plan idea to more repos and more executors. If you steal one thing from this post, steal that.

Second, and this is the part I'd defend hardest: one habit spans all four phases and predates the tooling. I re-ground against real state before I plan. It's the April "reality-sync session." It's the Phase 3 gate reconciling candidates against truth. And it shows up, almost word for word, when I catch an analysis cutting corners. A workflow review I ran in June leaned on convenient pre-summarized views instead of the raw tables, and my response was blunt:

This doesn't seem like it's aware of any of the PMO processes or project initiation… did you look through all of the actual raw ingestion tables?

The tooling got more elaborate across these four months; the discipline never changed: plan only against verified current state. The backlog, in the end, is just the most durable place I've found to keep that state — so that planning, whether it's me or an agent doing it, starts from the ground and not from a guess. The plan-first math everyone quotes only holds if the plan rests on true premises. A perfectly-structured spec built on stale context fails all 20 decisions just as surely as no plan at all — it just fails them faster, and with more confidence. Every phase here was, underneath, a better answer to the same question: where do I keep the truth the plan depends on?


I build a self-hosted personal AI data platform in the open. The one design call I'm still least sure about: whether the human approval gate at Phase 3 is a permanent feature or just scaffolding I haven't automated away yet. If you've run agents against a shared queue, where did you draw that line — and did it hold?