惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
GbyAI
GbyAI
MongoDB | Blog
MongoDB | Blog
人人都是产品经理
人人都是产品经理
A
About on SuperTechFans
Microsoft Security Blog
Microsoft Security Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
量子位
腾讯CDC
Google DeepMind News
Google DeepMind News
Vercel News
Vercel News
Blog — PlanetScale
Blog — PlanetScale
The Register - Security
The Register - Security
博客园 - Franky
M
MIT News - Artificial intelligence
C
CERT Recently Published Vulnerability Notes
B
Blog RSS Feed
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Simon Willison's Weblog
Simon Willison's Weblog
Attack and Defense Labs
Attack and Defense Labs
L
Lohrmann on Cybersecurity
S
Schneier on Security
MyScale Blog
MyScale Blog
The Last Watchdog
The Last Watchdog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
PCI Perspectives
PCI Perspectives
博客园 - 聂微东
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
T
Troy Hunt's Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Threat Research - Cisco Blogs
K
Kaspersky official blog
F
Fortinet All Blogs
Application and Cybersecurity Blog
Application and Cybersecurity Blog
D
Docker
Security Latest
Security Latest
P
Privacy & Cybersecurity Law Blog
T
Tenable Blog
B
Blog
有赞技术团队
有赞技术团队
TaoSecurity Blog
TaoSecurity Blog
C
Check Point Blog
Latest news
Latest news
H
Hackread – Cybersecurity News, Data Breaches, AI and More
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Help Net Security
Help Net Security
D
DataBreaches.Net
Google DeepMind News
Google DeepMind News
N
News and Events Feed by Topic
J
Java Code Geeks

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
What Building a Geopolitical Simulation Taught Me About Claude Code
Vartika Tewa · 2026-04-22 · via DEV Community

The most valuable features weren't the ones that wrote code for me.


I was 45 minutes into a Sprint 2 session when the test runner caught something I'd completely missed. I'd just saved a React component file — nothing dramatic, a minor prop change — and two seconds later, a notification appeared in my terminal:

FAIL tests/components/game/EscalationLadder.test.tsx
  ✗ renders correct rung label for rung 4
    Expected: "Conventional Strike"
    Received: "Conventional strike"

Enter fullscreen mode Exit fullscreen mode

The PostToolUse hook had fired automatically. run-tests-on-save.sh had detected that the saved file path matched a test file pattern, run vitest on that specific test file, and surfaced the failure before I'd even moved on to the next task. I hadn't asked for a test run. I hadn't thought to run one. The hook just enforced it.

That moment crystallized something I'd been slow to internalize: the most powerful things about Claude Code aren't the AI suggestions. They're the mechanisms that enforce process when you're too deep in a problem to remember to enforce it yourself.


What We Built

GeoSim is an AI-powered geopolitical strategic simulation engine. Load an Iran crisis scenario, watch six actors — the US, Iran, Israel, Russia, China, and the Gulf States — simultaneously plan their moves using separate Claude Sonnet agent calls, then watch a resolution engine arbitrate the outcomes and a narrator synthesize them into intelligence reports.

The core mechanic is a git-like branching system: at any turn node, a player can fork the timeline, take control of an actor, and steer events down a different path. Every branch is immutable once committed. Every actor sees only their own intelligence picture — fog of war filtered at the database layer via Supabase RLS.

It's live at https://geosim-eight.vercel.app/.

The technical stack is Next.js 14 App Router, TypeScript, Supabase (Postgres + Auth + Realtime), Mapbox GL JS, and the Anthropic API. But the interesting story isn't the stack — it's how Claude Code's extensibility layer shaped the way we built it.


CLAUDE.md as Architecture Documentation

The first thing we did was write a serious CLAUDE.md. Not a "this project uses Next.js" stub — a 216-line living document with @imports to 15 reference files:

@docs/frontend-design.md         — Stitch visual identity, font rules
@docs/prompt-library.ts          — All AI agent system prompts
@docs/agent-architecture.ts      — Agent roles and game loop pseudocode
@docs/geosim-data-model.ts       — TypeScript types for every entity
@docs/testing-strategy.md        — Test priorities, TDD workflow, mocking strategy

Enter fullscreen mode Exit fullscreen mode

The pattern that made this work was modular @imports. Instead of cramming everything into one file, each domain got its own reference doc. Claude Code only loaded what was relevant to the current task. When we were writing a component, it read frontend-design.md. When we were working on AI agents, it read prompt-library.ts and agent-architecture.ts. The whole thing felt like progressive disclosure — the same principle we applied to the fog-of-war system.

CLAUDE.md evolved across 92 PRs. You can trace the project's maturation by reading the git history of that single file: Sprint 1 added TDD rules and branch conventions; Sprint 2 added Stitch design tokens; Sprint 3 added the node-centric branch architecture. It became the source of truth for every architectural decision, enforced through both documentation and Claude Code's ability to actually read and follow it.


Skills as Iterable Automation

We built 14 custom skills. The most instructive story is quality-gate.

Version 1 ran a comprehensive QA audit — tests, types, linting, security, CI verification. It was useful. It also silently failed on our development machine, which runs WSL2 on Windows. The problem: v1 called npm run test, but on our machine npm is a Windows binary that can't execute the Linux Vitest binary. The skill appeared to succeed (exit 0) while actually doing nothing.

Version 2 header: # quality-gate — v2 WSL2 + context-mode aware.

Two changes: replaced every npm call with bun, and piped all output through ctx_execute (the context-mode MCP sandbox) instead of letting large test output flood the context window. The fix was trivial once we understood the problem. The lesson was the process: skills are code, and code gets bugs, and bugs get fixed in v2.

The other skill that became indispensable was run-turn — a complete game simulation turn executed from the command line. It chains the actor agent calls, runs the resolution engine, invokes the judge, and synthesizes via the narrator. During AI pipeline development, being able to run a full turn cycle from a single skill invocation compressed a development loop that would otherwise involve navigating five separate API endpoints.


Hooks as Enforcement, Not Documentation

Five hooks in .claude/settings.json:

"hooks": {
  "PreToolUse": [{
    "matcher": "Edit|Write",
    "hooks": [{ "command": "bash .claude/hooks/protect-files.sh" }]
  }],
  "PostToolUse": [
    { "matcher": "Edit|Write", "hooks": [{ "command": "prettier --write $CLAUDE_FILE_PATH" }] },
    { "matcher": "Edit|Write", "hooks": [{ "command": "bash .claude/hooks/run-tests-on-save.sh" }] }
  ],
  "Stop": [{ "hooks": [{ "command": "git status --porcelain | grep -q . && echo '...uncommitted...'" }] }]
}

Enter fullscreen mode Exit fullscreen mode

The difference between a documented rule and a hook is enforcement. CLAUDE.md can say "never commit to .env.local" — but a PreToolUse hook that exits with code 2 when you try to write to .env* files actually stops it from happening. Documentation is advisory. Hooks are structural.

The run-tests-on-save.sh hook was the one we felt most during development. It runs vitest only on the specific test file that was saved — not the whole suite — which keeps the feedback loop under two seconds. During component development, we'd write a test, save it, watch it fail, write the implementation, save it, watch it go green. Red-green in the same terminal session, automatically.

The session lifecycle hooks — SessionStart injecting a system message to run /start-session, Stop warning about uncommitted changes — felt more ergonomic than impactful. But they're there, and they've caught things.


Worktrees for Parallel AI Development

One of Sprint 2's structural wins was worktree agents. Five PRs were merged from worktree-agent-* branches — each one a Claude Code agent session running in a fully isolated git worktree:

  • worktree-agent-a29a8263 — trivial bug fixes
  • worktree-agent-ae36607a — prompt caching for AI stable system prompts
  • worktree-agent-a27c703e — error boundaries and empty states
  • worktree-agent-a32132a5 — Israel decision catalog
  • worktree-agent-ab230cc1 — z-index fixes for actor panel / map controls

The pattern: a parent session dispatches a subagent into a worktree (superpowers:dispatching-parallel-agents skill), the subagent works in isolation with no shared file state, opens a PR when done, and the parent session reviews and merges. No merge conflicts with in-progress main branch work. No contamination of the parent session's context.

During Sprint 2, while one worktree session was implementing the prompt caching architecture for AI agents, the main session was building the Scenario Hub page. Two nontrivial features, developed in parallel, no coordination overhead.


TDD with Claude Code

The project's testable commit history shows five explicit red-before-green sequences. The clearest:

43e857b  test: add failing tests for node API routes (TDD)
d4f9c13  feat(#32): add generateDecisionOptions asset-aware, NEUTRALITY preamble

Enter fullscreen mode Exit fullscreen mode

The run-tests-on-save.sh hook makes the red-green loop visceral. You commit the failing test, watch it fail on every save until you wire the implementation, then watch it go green. The hook closes the feedback loop that TDD requires without requiring you to context-switch to run tests manually.

The test suite has 41 files now: 9 game logic, 2 AI, 3 API integration, 20 components, 3 library utilities, 4 E2E. The component tests are the weakest layer — most are smoke-level assertions (does the element render?) rather than behavioral contracts. The game logic and AI tests are where TDD was practiced most rigorously.


Honest Reflection

Three things we'd do differently:

C.L.E.A.R. reviews should be live, not retroactive. We had the review-pr.md skill configured from Sprint 1. We used it inconsistently. The PR reviews we posted retroactively to PRs #83, #88, and #89 are substantive — but they landed after merge, not before. In a future project, the review-pr skill gets wired to a PR creation checklist, not a sprint retrospective.

E2E tests belong in Sprint 1. Our Playwright config lived as a script stub with no actual test files until the project documentation phase. Four smoke tests against the deployed app would have caught the auth redirect regression in Sprint 2 that we found manually.

The .mcp.json file should ship with the repository from day one. We configured MCP servers via settings.json throughout the project — which works fine locally — but the rubric expects a shareable .mcp.json. It's a five-minute addition, but it signals to the next developer exactly how to reproduce the development environment.


Five Specific Takeaways

  1. Write hooks before you write features. A PostToolUse test runner costs 30 minutes to configure and saves hours of "wait, did I break something?" across a sprint.

  2. Version your skills. quality-gate v1 was wrong for our environment. quality-gate v2 was right. Treating skills as code — with versions, bug fixes, and iteration — is the mindset shift.

  3. CLAUDE.md @imports are a force multiplier. A single 200-line CLAUDE.md with @imports to domain-specific docs is more effective than one 800-line monolith that Claude Code has to parse in full every time.

  4. Worktree agents are underrated for parallel work. The agent isolation prevents the context contamination and merge conflicts that slow down parallel development. Use them aggressively.

  5. The neutrality principle requires explicit enforcement. Our NEUTRALITY_PREAMBLE (injected into every AI agent system prompt) exists because without it, AI agents drift toward protagonist bias. For a simulation that models six actors with equal rigor, this is a correctness requirement, not an ethical preference. Make your invariants explicit in CLAUDE.md or they won't hold.