惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Secure Thoughts
博客园_首页
IT之家
IT之家
Engineering at Meta
Engineering at Meta
量子位
宝玉的分享
宝玉的分享
MyScale Blog
MyScale Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
L
LangChain Blog
爱范儿
爱范儿
WordPress大学
WordPress大学
F
Full Disclosure
T
Tailwind CSS Blog
GbyAI
GbyAI
Recorded Future
Recorded Future
美团技术团队
S
SegmentFault 最新的问题
A
About on SuperTechFans
小众软件
小众软件
云风的 BLOG
云风的 BLOG
人人都是产品经理
人人都是产品经理
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
Apple Machine Learning Research
Apple Machine Learning Research
D
DataBreaches.Net
J
Java Code Geeks
The Cloudflare Blog
The GitHub Blog
The GitHub Blog
Hugging Face - Blog
Hugging Face - Blog
D
Docker
Vercel News
Vercel News
H
Help Net Security
博客园 - 叶小钗
B
Blog
阮一峰的网络日志
阮一峰的网络日志
N
Netflix TechBlog - Medium
Blog — PlanetScale
Blog — PlanetScale
腾讯CDC
Microsoft Security Blog
Microsoft Security Blog
V
Visual Studio Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - 司徒正美
The Register - Security
The Register - Security
aimingoo的专栏
aimingoo的专栏
博客园 - 聂微东
月光博客
月光博客
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Last Week in AI
Last Week in AI
M
MIT News - Artificial intelligence
Jina AI
Jina AI

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
What Is Human-in-the-Loop (HITL) in AI? A Practical Guide
Brenn Hill · 2026-06-24 · via DEV Community

Human-in-the-loop (HITL) in AI means keeping a person involved in an automated system's decisions — approving, editing, or interrupting what an AI does — instead of letting it run fully on its own. For AI agents, human-in-the-loop is the practice of pausing the agent at chosen points so a human can review or steer an action before it takes effect. The hard part isn't adding a human; it's making sure that human can actually catch the mistakes that matter.

This guide explains what human-in-the-loop is, the three forms it takes in real AI agents, how it differs from full automation and human-on-the-loop, and — most importantly — why a review step is not the same as safety. Then it covers how to do HITL well, and when you should prevent a bad outcome instead of reviewing for it.

What human-in-the-loop actually means

The phrase comes from control systems and machine learning, where a "loop" is the cycle of an action, its result, and a correction. Putting a human in the loop means the cycle can't close without a person — the system stops and waits for input. Putting a human on the loop means the system runs autonomously while a person watches and can step in. Taking the human out of the loop means full automation.

In AI agents — code agents, computer-use agents, support bots, ops automations — human-in-the-loop is how teams try to keep oversight as autonomy grows. The agent proposes or starts an action; a human gets a say. That's the idea. Whether it works depends entirely on the details, which is where most implementations quietly fail.

The three modes of human-in-the-loop AI

HITL shows up in agents in three recognizable shapes. Most products use a mix.

1. Approve-before-act

The agent describes an action and waits for a yes before doing it: "Run this command?" "Send this email?" "Delete these rows?" This is the most common pattern and the most over-trusted. It feels safe because nothing happens without a click — but a click is not the same as understanding. (More on that below.)

2. Review-and-edit

The agent produces a draft — code, a message, a plan, a config change — and the human reviews and edits it before it ships. This is genuinely useful when the artifact is legible and the reviewer has time: a small diff, a short email, a single query. It degrades fast when the output is large or dense, because reviewers skim.

3. Interrupt-and-resume

The agent runs autonomously, but a human (or a monitor) can pause, redirect, or kill it mid-task. This is the human-on-the-loop end of the spectrum, and it's the right default for high-throughput work where stopping for every action would be absurd. It only counts as oversight if the interrupt is real — reachable, fast, and able to halt in-flight work.

HITL vs. full automation vs. human-on-the-loop

These aren't three boxes; they're points on an autonomy ladder, and the right point depends on the action.

  • Full automation — the agent acts, no human gate. Correct for trivial, reversible, contained actions where a human adds nothing.
  • Human-on-the-loop — the agent acts autonomously; a human supervises and can interrupt. Good for high-volume work with a healthy kill switch and monitoring.
  • Human-in-the-loop — the agent waits for a human before (or during) the action. Reserve it for the consequential, gateable middle.

The mistake is treating this as one global setting — "we have a human in the loop" — applied to everything an agent can do. That guarantees either fatigue (you gate trivia) or blind risk (you under-gate the dangerous stuff). The LoopRails framework argues you should pick the mode per action, not per product. For the full mapping of grades to autonomy modes, see the playbook.

The misconception: a review step is not safety

Here's the belief that breaks most HITL designs: if a human approves it, it's safe. It isn't, and the evidence is direct.

In research on AI coding agents (see the LoopRails codex), requiring plan-approval before the agent acted cut attack occurrence from roughly 90% down to 60–74%. That sounds like a win — until you look at whether the human actually caught the bad action when it slipped through. Human intervention success stayed at just 9–26% across every approval strategy tested. Gating reduced how often bad actions happened, but barely improved the human's ability to recognize and stop one. The gate was working as a speed bump, not as a detector.

Why? Automation bias. People over-trust system suggestions and approve them without real scrutiny — especially when the system has been right before, when the output looks confident, and when there's time pressure to keep moving. A confirmation prompt does not turn a person into a good error-catcher. It mostly turns them into a click.

Two failure modes follow from this:

  • The Rubber Stamp — approvals get clicked through reflexively, so the gate stops bad actions occasionally but rarely catches a targeted one.
  • The Moral Crumple Zone — when something goes wrong, the human who clicked "approve" gets the blame, even though they never had a realistic chance to catch the problem. The review existed to assign accountability, not to prevent harm.

If your oversight only proves that a review step exists, you have Phantom Oversight: a control that looks like safety on the org chart and does nothing in production.

The better question

Don't ask "should a human review this?" Ask: can a human realistically catch this mistake in time?

That reframes oversight as an engineering problem with a testable answer. If the reviewer can see the real action and its consequences, has the competence and the time to judge, and can actually stop or reverse it — then a gate can work. If they can't — if the consequence is high but their controllability is low — then review is a trap. You're staging a decision the human can't really make, and a confirmation prompt just launders the risk into their name.

This is the difference between oversight that prevents harm and oversight that exists to be pointed at after harm.

How to do human-in-the-loop well

LoopRails frames good HITL as four moves: Grade, Guard, Show, Prove.

Grade

Score every action an agent can take on three axes — reversibility, blast radius, and stakes — and let the highest axis set the grade, G0 to G3.

  • G0 — trivial: reversible, local, no stakes (read a file, run a read-only query). No gate; gating it just breeds fatigue.
  • G1 — low: at most one medium axis (edit a local file, run tests). Cheap undo beats a confirmation.
  • G2 — high: any one high axis (git push, spend within budget, send an internal message). Confirm-before with a real preview.
  • G3 — critical: irreversible and external or severe (deploy, pay, delete prod data, post publicly). Prevent, or escalate — review alone is not enough here.

Guard

Match the control to the grade. Don't spend attention on G0/G1; gate G2 with a preview; for G3, lean on prevention patterns over approval prompts — Sandbox-First (contain blast radius in the environment), Blast-Radius Cap (limit any single action's magnitude), Capability Lock (make the bad action impossible, not discouraged), Runtime Shield, Kill Switch, Circuit Breaker, and Maker-Checker (the proposer is never the approver).

Show

When you do pull a human in, design the moment. Show them the real action and its consequences — a diff, a preview, the side effects, whether it can be undone — not a bare "Approve?" Surface the agent's uncertainty and provenance so they can check rather than trust. And spend attention sparingly: interrupt rarely and at meaningful breakpoints, because over-prompting trains people to dismiss prompts.

Prove

Treat "a human reviews it" as a claim to validate, not a checkbox. Seed known errors and prompt-injection attempts into your pipeline and measure whether the human (or monitor) actually catches them. The number that matters is intervention-success rate, not approval rate. Untested oversight is unvalidated oversight.

Underneath all four moves, keep every governed action on the RAIL: Reversible, Authorized, Interruptible, and Logged. If an action satisfies those four, even a missed review is recoverable, scoped, stoppable, and accountable.

When HITL is the wrong tool — prevent instead

Sometimes the honest answer to "can a human catch this in time?" is no. The action is too fast, too opaque, or too irreversible, and no realistic prompt would let a person intervene effectively. In that case, don't add a review. Adding one creates a Rubber Stamp and a Moral Crumple Zone at once. Change the action instead so the bad outcome can't happen or can be undone.

The clearest example is the lethal trifecta. An agent that has (1) access to private data, (2) exposure to untrusted content, and (3) a way to send data externally can be tricked by prompt injection into exfiltrating that data. No "are you sure?" prompt reliably catches this — the malicious instruction is buried in content the human won't read, and the agent looks like it's doing its job. The fix isn't review; it's prevention. Remove any one leg — cut external send, isolate the private data, or sanitize the untrusted input — and the attack can't complete. That's a Capability Lock, not a gate.

When consequence is high and controllability is low, prevention beats review every time.

Key takeaways

  • Human-in-the-loop means a person can approve, edit, or interrupt an AI's action before it takes effect — the opposite of full automation.
  • It shows up in three modes: approve-before-act, review-and-edit, and interrupt-and-resume.
  • Adding a review step is not the same as safety: gates cut how often bad actions occur but barely improve a human's ability to catch one (9–26% intervention success), and automation bias makes approvals reflexive.
  • Ask "can a human realistically catch this in time?" — not "should a human review this?"
  • Do HITL well with Grade, Guard, Show, Prove, and keep every action Reversible, Authorized, Interruptible, Logged.
  • When a human can't catch the mistake in time, prevent the bad outcome instead of staging a review.

Get started

Stop asking whether you have a human in the loop and start grading your agent's actions. Run your riskiest actions through the interactive grader to see their G0–G3 grade and the controls that match, then work the four moves with the practitioner playbook. Keep the cheatsheet next to your next agent review — and the next time someone proposes "just add an approval step," ask whether the human can actually catch the mistake in time.


Originally published at looprails.dev/article-what-is-human-in-the-loop.html. LoopRails is a free, sourced framework for designing human-in-the-loop oversight of AI agents.