惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Spread Privacy
Spread Privacy
T
Threatpost
L
LINUX DO - 热门话题
Google Online Security Blog
Google Online Security Blog
I
InfoQ
大猫的无限游戏
大猫的无限游戏
博客园_首页
爱范儿
爱范儿
有赞技术团队
有赞技术团队
V
Visual Studio Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
酷 壳 – CoolShell
酷 壳 – CoolShell
P
Privacy International News Feed
C
Cyber Attacks, Cyber Crime and Cyber Security
Jina AI
Jina AI
博客园 - 聂微东
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
C
CERT Recently Published Vulnerability Notes
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
aimingoo的专栏
aimingoo的专栏
P
Proofpoint News Feed
K
Kaspersky official blog
L
LangChain Blog
G
GRAHAM CLULEY
B
Blog RSS Feed
G
Google Developers Blog
Google DeepMind News
Google DeepMind News
The Cloudflare Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
A
About on SuperTechFans
阮一峰的网络日志
阮一峰的网络日志
Last Week in AI
Last Week in AI
T
Tailwind CSS Blog
Cyberwarzone
Cyberwarzone
C
Cybersecurity and Infrastructure Security Agency CISA
P
Proofpoint News Feed
Help Net Security
Help Net Security
S
Security @ Cisco Blogs
Cloudbric
Cloudbric
雷峰网
雷峰网
C
Check Point Blog
MongoDB | Blog
MongoDB | Blog
NISL@THU
NISL@THU
L
Lohrmann on Cybersecurity
Vercel News
Vercel News
T
Tor Project blog
T
The Exploit Database - CXSecurity.com
T
Troy Hunt's Blog
W
WeLiveSecurity
T
Threat Research - Cisco Blogs

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The x3.16 Developer | Part 1
Amir Arad · 2026-05-29 · via DEV Community

TL;DR

Taking the time to figure out what has value beyond this specific task and feeding it back into your setup is the single highest-leverage thing you can do.


1. The Promise and the Gap

Everyone's chasing the x10 developer these days. AI can make you ten times more productive, they say. Ship ten times faster. Do the work of a team.

When I started using AI for coding, the obvious move was to go faster: wire the best LLM to the codebase, give it a task suitable in size and complexity, and watch it change the code faster than I can keep track of. That didn't work. The model would produce things that looked right but I'd spend more time correcting and retracing than it would have taken me to build it myself. Every 6 months a new model or tool would come out that made things better, but nothing that really closes the gap. I think I got to x1.25, maybe.

Slowly, I fell back to old principles. TDD, acceptance tests, static analysis: structural constraints around the risky areas. Task after task I'd retrospect, fix, generalize and automate the things supporting the model's coding. And it started working. Not because the model got better, but because everything around it got better. Work quality went up, and speed followed without my optimizing for it directly. At some point the attention required per task dropped enough that I could run two at the same time. Then three. My current peak is three running coding tasks and a design brainstorm - all running at the same time.

Here's how I think about it. x10 breaks down into two axes: scale up, where each task gets done faster, and scale out, where you can run more tasks concurrently. If you can only reach x√10 on each axis, you get x10 total. √10 = 3.16 . That's where the title comes from.

Getting to x10 capacity by reaching x3.16 in both axes of scale

What follows is the method. How I think about the system, how I optimize it, and what the feedback loops look like. The specific tools and workflows get their own article.


2. The Engine

An Engine

So the first thing to understand about LLMs, they're basically trained to predict what comes next in a sentence, and then fine-tuned to be helpful. Which means they have this deep pull toward giving you something that looks like a good answer. Not necessarily is a good answer, looks like one. I think of that as a very powerful engine, without a sense of direction.

When you work with a person, they have domain knowledge, they have stakes in the outcome, they understand the problem. An LLM doesn't have any of that. It just has this pull toward whatever feels most helpful right now, message by message. And a lot of the time that looks like substance, but it's not always actually substance.

This is more or less how hallucinations happen. A hallucinated answer sounds knowledgeable, confident, relevant. It passes your immediate quality check, because that's what the model is designed for. It's not that the model is trying to deceive you, it's just how the system works.

If you're just chatting with it, asking questions, brainstorming, it works fine most of the time, because what sounds right and what is right are mostly aligned. The problems start when things get complex, and what sounds right and what is right drift further apart. That pull toward passing as helpful becomes counterproductive the more complex the domain gets.

This thing took me a while to internalize: the model, with all its intelligence, is the least trustworthy part in the system. Engineers instinctively treat LLM as the core and wrap uncertainty around it. But it's actually a statistical guessing box with some randomness on top. It is inherently not sensible. Everything around it has to compensate for that.


3. The Harness

A car with no engine

So if the LLM is the engine, there's all this stuff around it that makes it actually usable. System prompts, tools, verification steps, context management, guardrails. In agent engineering they call this the harness. It's basically everything between the model and reality.

When an agent fails, the reflex is to rewrite the prompt. Add more detail, be more specific. But agent behavior comes from the whole system, not just the prompt. Structural fixes at the harness level regularly outperform prompt tweaks by an order of magnitude.

Tejas Kumar recently took a 2023-era model, GPT-3.5 Turbo, and had it successfully complete a multi-step browser task through harness engineering alone (watch his excellent talk here). He then continued to explain that the harness has moving parts:

  • Tool registry. What the model can actually do, and how it perceives these capabilities.
  • Context management. The working memory of the conversation so far. This is the part that degrades over long sessions.
  • Guardrails. Fail-fast mechanisms. Max retries, permissions, resource caps.
  • The agent loop. Think, act, observe the result, decide what's next.
  • Verification. Checking whether what the model did actually worked, independent of what the model claims.

That last one closes the gap between "appeared successful" and "was successful": Without it, you're trusting the model to judge its own output. And the model is biased toward telling you everything went great.

While Tejas Kumar may talk from the context of fully autonomous agents, this is not an autonomous car. You're in the center of the operation. The engine does the heavy lifting, but you decide where it goes, when it stops, and whether what it produced is what you needed. The harness is set up around that. It doesn't need to handle everything on its own, it needs to make your interventions cheap and your oversight easy. Every automation starts with doing the task manually and learning the process, then automating piece by piece.


4. The Driver-Mechanic

A car driver+mechanic

So we have an engine and a car. But to get somewhere we need the driver.

The driver is an equal part of the system as the engine and the car. What you choose to check and where you save your energy, how you review the model's output, when you decide to interrupt, how you define tasks - these are skills and intuitions, different from person to person, and change over time. That's the driver part.

Right now general-purpose AI-coding feels a lot like the early days of cars. The technology works, but it's not something you can just use without thinking about it. If you owned a car in 1905 you carried a toolbox and you knew your way around the engine, because driving and maintenance weren't separate things yet.

That's roughly where we are with coding harnesses. You're not just using the tool, you're also the person who tweaks and maintains it. And the really good results aren't commoditized yet. You have to make them by yourself, around yourself.

So this is what we optimize. The engine is fixed. The harness is engineerable. The driver improves through practice and feedback, and also improves the harness. And the interfaces between layers, that's where most of the value lives and where most things go wrong.


5. The Optimization Process

So there's this idea from manufacturing that became very popular. In the 1950s Toyota couldn't compete with American manufacturers on volume or capital. So they focused on their process instead. They took ideas from people like Deming about quality and waste reduction and built them into how they worked, at every level, continuously. The approach is called Kaizen. The core of it is that you keep finding and removing waste from your process, over and over, and the improvements compound.

Some principles that translate directly:

  • Optimize throughput, not capacity. A powerful engine means nothing if you're losing half the power somewhere in the middle.
  • Remove waste. Anything that doesn't contribute to output, including time fighting your tools.
  • Shorten the feedback loop. The faster you learn per cycle, the faster you improve.
  • Small batches. Smaller iterations surface problems sooner.
  • Stop the line. Don't let defects propagate. Fix now, not later.

There's also something I know as the Cult of Done (by Bre Pettis and Kio Stark) which roughly says "Done is the engine of more. Ship the imperfect thing, learn from it, ship the next thing better". The temptation with AI tools is to grab more land every iteration, do more, extend further. That causes drift. Forcing small completions counters that.

Now bring it back to improving the harness:

Taking the time to figure out what has value beyond this specific task and feeding it back into your setup is the single highest-leverage thing you can do. It's not glamorous. It's you as a mechanic, working on the car. It's you as a driver, learning. and it has nothing to do with LLM.

It's what compounds.


6. The Interrupt Loop (Andon Cord)

Immediate feedback from Driver to Harness, during a task.

My default fix for tasks gone wrong is usually to restart the task. But I prefer to use local repair over global restart as much as possible, so for this I need early detection at an actionable point. A restart costs you everything, context, progress, warmup. A local repair costs almost nothing if you catch it early.

In Toyota's factories, any worker could pull the andon cord to stop the production line when they spotted a defect. Not a failure of the system, the system working as intended. The defect gets fixed at the source instead of propagating downstream.

Same thing when you're running a task. The agent is generating code and starts drifting, adding unnecessary abstraction, approaching something in a way that'll cost you later. You interrupt. Not to start over. To make a local repair.

Sometimes it's a prompt adjustment. Sometimes you realize the agent needs context it doesn't have, or a tool is leading it astray. Sometimes the fix is removing something, not adding something.

The instinct is to let it finish. But that's how defects propagate. It's always more expensive later. Pull the cord when you first feel something is off, even if you're not sure - just to ask it to explain why it is doing what you think may be wrong, or how it plans to handle some challenge you suspect is going to cause issues. This habit has two great side-effects, and you can always tell it to "resume" once you are satisfied that things are going well.

Side effect 1: you learn. Little by little, you get a sense for the red flags. This learning will slowly translate into increased speed and reduced attention waste, and later you could even submit some of that learned wisdom to the harness, as guardrails, system prompts, and validations.

Side effect 2: the LLM gets grounding. Like people, it is sometimes good for LLMs to reflect on why they are doing what they are doing. After providing reasoning or plans, the chances of drifting away from it in the near future reduce significantly.

But if the answers are not good enough, that's when this habit really pays off - it saves time that would have otherwise been wasted on letting it wander and bump around the problem, reading the garbage result, etc. that's the scale-up axis going from x1 to x1.05 right there.

The subtle part is learning to distinguish "wrong" from "different than I expected." Sometimes the agent takes a path you wouldn't have taken, and it works fine. The cord is for defects, not preferences. Learning that boundary is part of the driver skill.


7. The Synthesis Loop

Feedback from completed task back to Harness, between tasks

Task done. Feature shipped, bug fixed, thing works. There's a mess on the workbench. Failed approaches, temporary hooks, discovered patterns, workarounds, model behaviors that only surfaced under these conditions.

Most of this gets swept into the bin. Task done, next task. This is where most people leave compound improvement on the table.

After a task, I ask: does anything here have value beyond this specific task?

Sometimes no. Clean up, move on. But often: a hook you added, was the underlying problem specific to this task, or general? A prompting pattern that worked, can you encode it into the system prompt? A model behavior you discovered, does it change how you structure future tasks? A tool you built mid-task, is a minimal version worth keeping permanently?

Options from cheapest to most expensive: log it (just capture, zero processing), generalize it (make the specific solution apply to the class), or investigate it (spend time on something weird, was it noise or signal?).

Sometimes the right answer is "not worth any investment." That's fine. The Cult of Done applies to the feedback loop too. But don't skip it entirely.

Not all improvements come from the loop. Some arrive sideways. You're debugging something unrelated and notice a model behavior that changes how you structure tasks. You read about someone else's setup and realize it solves a problem you hadn't named yet. You can't schedule these, but you can avoid suppressing them. When something unexpected happens during a task, give it thirty seconds before you correct it.

What compounds: system prompts that evolved from empty to opinionated operating manuals. Not designed top-down, but built from dozens of small synthesis cycles, each one encoding a specific problem into a structural solution.


The full triad stack of engine, car, and driver

That's the first part, it's about how the solution works, in general. The next part will be more about the practical assets and habits I've collected for scale-up using the methods described here, and how I scale-out several tasks together. It builds on what's here: a harness you trust and a feedback loop that compounds. You can't run three cars at once if you don't trust the one you're driving.

I don't think the x10 developer is someone who found a better tool than everyone else. It's someone who keeps tuning their setup.

Try this once. After your next completed task, before you move to the next one, stop and ask whether anything you just did or learned has value beyond that specific task. You don't need a system for it, you don't need to do it every time. You can also use LLM for that. Just the question, once, and see what's there.