惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LangChain Blog
Engineering at Meta
Engineering at Meta
S
Securelist
M
MIT News - Artificial intelligence
GbyAI
GbyAI
O
OpenAI News
W
WeLiveSecurity
T
Troy Hunt's Blog
L
LINUX DO - 最新话题
博客园_首页
C
Check Point Blog
Martin Fowler
Martin Fowler
The Last Watchdog
The Last Watchdog
量子位
Cloudbric
Cloudbric
S
SegmentFault 最新的问题
Recent Commits to openclaw:main
Recent Commits to openclaw:main
SecWiki News
SecWiki News
L
Lohrmann on Cybersecurity
Forbes - Security
Forbes - Security
雷峰网
雷峰网
H
Heimdal Security Blog
P
Palo Alto Networks Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Y
Y Combinator Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
博客园 - 聂微东
Hacker News: Ask HN
Hacker News: Ask HN
Hacker News - Newest:
Hacker News - Newest: "LLM"
Help Net Security
Help Net Security
U
Unit 42
N
News and Events Feed by Topic
Hugging Face - Blog
Hugging Face - Blog
A
About on SuperTechFans
Stack Overflow Blog
Stack Overflow Blog
TaoSecurity Blog
TaoSecurity Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - Franky
Jina AI
Jina AI
美团技术团队
L
LINUX DO - 热门话题
F
Fortinet All Blogs
The GitHub Blog
The GitHub Blog
Security Latest
Security Latest
S
Secure Thoughts
Microsoft Azure Blog
Microsoft Azure Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
IT之家
IT之家
P
Privacy International News Feed
博客园 - 司徒正美

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
Build A Basic AI Agent From Scratch: Human in the Loop & Security
ruxudev · 2026-06-17 · via Hacker News - Newest: "AI"

40 minute read · Artificial Intelligence

Previous parts of Build a Basic AI Agent From Scratch:

You can find and clone this code in this blog series' Github repo.

In the previous part of the Build A Basic AI Agent From Scratch series, we gave our agent the ability to plan and work on long tasks. We added a scratchpad, a to-do list and a system prompt that explains to the model how to break work down, recover from failures and keep going until the task is actually done.

That made the agent much more useful, but it also made it more dangerous. Running commands and editing files indiscriminately can have bad consequences that cannot be undone. We want our agent to be able to work autonomously but at the same time check with you before running potentially harmful tools.

In this part of the series we will add human in the loop controls to our agent. The agent will still be autonomous, but it will have to stop and ask for permission before doing potentially risky actions. It will also get a new tool that lets it ask the user a question when it does not have enough information to proceed.

Human in the Loop

In AI Agents, the term human in the loop means that some decisions require the manual action by a human before they run. This ensures that some sensitive actions are not performed without passing the test of the criterion of a human.

What Should Require Permission?

Not every tool call needs the same level of scrutiny. If the agent asks the user for permission on every single tool call, it becomes annoying and slow. On the other hand, if the agent never asks for permission, it becomes unsafe.

So we will classify tools by risk:

  1. Read tools can inspect the filesystem but do not change it.
  2. Planning tools only update the agent's internal state.
  3. Interaction tools ask the user for clarification.
  4. Write tools modify files.
  5. Other action tools can have broader side effects, like running shell commands or fetching from the network.

For this version of the agent, the safe default is:

  • Reading files is allowed.
  • Planning is allowed.
  • Asking the user a question is allowed.
  • Writing files requires permission unless we explicitly start the agent in a mode that accepts edits inside the current project.
  • Running bash commands requires permission.
  • Fetching web pages requires permission.

Permission Modes

We will add three permission modes to the agent:

class PermissionMode(Enum):
    DEFAULT = "default"
    ACCEPT_EDITS = "acceptEdits"
    DANGEROUSLY_SKIP_PERMISSIONS = "dangerouslySkipPermissions"

The modes work like this:

  • default: read tools and planning tools are allowed, everything else asks for permission.
  • acceptEdits: read tools, planning tools and writes inside the current working directory are allowed, everything else asks for permission.
  • dangerouslySkipPermissions: all tools run without asking.

The last mode is intentionally named in a scary way. Running without any safeguards is the kind of mode you might use in a throwaway sandbox or a trusted automation environment. It shouldn't be the default for an agent running on your machine with precious files and credentials.

We can expose the permissions mode as a command line flag:

parser = argparse.ArgumentParser(
    description="Coding agent with configurable tool permission gating."
)
parser.add_argument(
    "--mode",
    choices=["default", "acceptEdits", "dangerouslySkipPermissions"],
    default="default",
    help=(
        "Permission mode for tool execution. "
        "'default': read tools are free, everything else requires approval. "
        "'acceptEdits': read + write tools are free when inside the working directory, "
        "everything else requires approval. "
        "'dangerouslySkipPermissions': all tools run without any prompt."
    ),
)

Then we capture the current working directory when the agent starts, which we will use as the trust boundary for the acceptEdits mode. The agent can edit files inside the project, but writing outside the project still requires permission.:

mode = PermissionMode(cli_args.mode)
working_dir = Path.cwd()

print(f"Agent started in '{mode.value}' mode  (working dir: {working_dir})")

client = get_llm_client()
agent_loop(client, mode, working_dir)

Next, we will group the tools in three groups. Tools that can only read files or be used for planning will always be allowed because they are safe. Write tools will be more limited:

# Always allowed: read-only filesystem tools
READ_TOOLS = {"read_file", "glob_files", "grep"}

# Always allowed: internal planning/bookkeeping and user-interaction tools
PLANNING_TOOLS = {
    "todo_append",
    "todo_list",
    "todo_update",
    "read_scratchpad",
    "write_scratchpad",
    "ask_question",
}

# Conditionally allowed in acceptEdits mode when target is within working dir
WRITE_TOOLS = {"write_file", "edit_file"}

Checking the Write Path

If the agent is in acceptEdits mode, we want to allow writes inside the project and block writes outside the project unless the user approves them.

That means we need to resolve the path and check whether it is inside the working directory:

def _resolve_tool_path(tool_name: str, args: dict) -> str | None:
    """Return the file-path argument for write tools, or None if not applicable."""
    if tool_name in WRITE_TOOLS:
        return args.get("path")
    return None


def _is_within_working_dir(path: str, working_dir: Path) -> bool:
    """Return True if *path* resolves to somewhere inside *working_dir*."""
    try:
        target = Path(path)
        if not target.is_absolute():
            target = working_dir / target
        target.resolve().relative_to(working_dir.resolve())
        return True
    except ValueError:
        return False

Asking for Permission

When the agent wants to run a tool that is not automatically allowed, we ask the user:

def _ask_permission(tool_name: str, args: dict) -> bool:
    """Interactively ask the user whether to allow a tool call.

    Returns True if the user grants permission, False otherwise.
    """
    print(f"\n  [permission required] {tool_name}")
    print(f"  Arguments: {json.dumps(args, ensure_ascii=False)}")
    while True:
        try:
            answer = input("  Allow this action? [y/n]: ").strip().lower()
        except EOFError:
            print("  (EOF - denying permission)")
            return False
        if answer in ("y", "yes"):
            return True
        if answer in ("n", "no"):
            return False
        print("  Please enter 'y' or 'n'.")

We make it easy to see for the user which action the agent is trying to perform so they can understand what's going on. Before a risky tool runs, the user sees the tool name and the exact arguments the model requested. The user can approve or deny it.

Now we can put all the rules together:

def check_permission(
    tool_name: str,
    args: dict,
    mode: PermissionMode,
    working_dir: Path,
) -> bool:
    """Decide whether a tool call is permitted under the current mode."""
    if tool_name in READ_TOOLS or tool_name in PLANNING_TOOLS:
        return True

    if mode == PermissionMode.DANGEROUSLY_SKIP_PERMISSIONS:
        return True

    if mode == PermissionMode.ACCEPT_EDITS and tool_name in WRITE_TOOLS:
        path = _resolve_tool_path(tool_name, args)
        if path and _is_within_working_dir(path, working_dir):
            return True

    return _ask_permission(tool_name, args)

The function returns a boolean that represents whether the harness allows the agent to proceed with the tool call.

Now we need to integrate check_permission into the tool execution path. This is the part of the agent loop that receives tool calls from the LLM and decides what to do with them:

def handle_tool_calls(
    tool_calls,
    messages,
    mode: PermissionMode,
    working_dir: Path,
):
    """Execute each tool the LLM requested and append the results to messages."""
    for tool_call in tool_calls:
        name = tool_call.function.name
        args = json.loads(tool_call.function.arguments)

        print(f"  [tool] {name}({args})")

        if name not in TOOL_REGISTRY:
            result = (
                f"Error: unknown tool '{name}'. "
                f"Available tools: {list(TOOL_REGISTRY.keys())}"
            )
        elif not check_permission(name, args, mode, working_dir):
            result = (
                f"Permission denied: the user did not allow '{name}' to run. "
                "Do not retry this tool call without asking the user first."
            )
        else:
            try:
                result = TOOL_REGISTRY[name](**args)
            except TypeError as e:
                result = (
                    f"Error: invalid arguments for tool '{name}': {e}. "
                    "Check the tool schema and retry with the correct arguments."
                )

        print(f"  [tool result] {result[:200]}{'...' if len(result) > 200 else ''}")

        messages.append({
            "role": "tool",
            "tool_call_id": tool_call.id,
            "content": result,
        })

If the permission is denied by the user, we return a tool result back to the model saying that the permission was denied and that it should not retry the same tool call.

This also keeps what happened clear for the agent. The model learns that its requested action did not happen, and it has to adapt.

Letting the Agent Ask Questions

Permission prompts are initiated by the harness. They happen when the model tries to do something risky.

But there is another kind of human in the loop interaction: the agent itself might realize that it is missing information. Maybe the user asked it to update "the config" but there are multiple config files. Maybe it needs to know which deployment target to use. Maybe it found two possible interpretations of the task and choosing wrong could cause damage.

For that, we add a new tool called ask_question:

def ask_question(question: str) -> str:
    """Ask the user a clarifying question and return their answer."""
    print(f"\n  [agent] {question}")
    try:
        answer = input("  Your answer: ").strip()
    except EOFError:
        return "(no answer - EOF)"
    return answer if answer else "(no answer provided)"

This tool is very small, but it changes the behavior of the agent. The agent no longer has to guess when guessing would be unsafe. It can stop, ask one focused question, and continue with the user's answer in context.

Then we register it in the tool registry:

from tools.interaction import ask_question


def get_tool_registry():
    return {
        "run_bash":          run_bash,
        "read_file":         read_file,
        "glob_files":        glob_files,
        "grep":              grep,
        "write_file":        write_file,
        "edit_file":         edit_file,
        "webfetch":          webfetch,
        "todo_append":       todo_append,
        "todo_list":         todo_list,
        "todo_update":       todo_update,
        "read_scratchpad":   read_scratchpad,
        "write_scratchpad":  write_scratchpad,
        "ask_question":      ask_question,
    }

And we expose it to the model with a schema:

{
    "type": "function",
    "function": {
        "name": "ask_question",
        "description": (
            "Ask the user a clarifying question and wait for their answer. "
            "Use this when you are missing information required to complete the task "
            "and cannot reasonably infer it from context. "
            "Ask one focused question at a time. "
            "Do not use this for progress updates or to confirm actions you can already "
            "take - only ask when you are genuinely blocked."
        ),
        "parameters": {
            "type": "object",
            "properties": {
                "question": {
                    "type": "string",
                    "description": "The question to ask the user.",
                },
            },
            "required": ["question"],
        },
    },
}

Updating the System Prompt

We also need to mention the new tool in the system prompt:

"- Clarification (ask_question): ask the user a single focused question when you "
"are genuinely blocked and cannot reasonably infer the missing information from "
"context. Do not use it for progress updates or to confirm actions you can already "
"take - only ask when it is strictly necessary to proceed.\n\n"

Let's test it!

Now we can run the agent in default mode:

$ python agent.py
Agent started in 'default' mode  (working dir: /Users/roger/project)
You: Create a markdown summary of this project
  [tool] glob_files({'pattern': '*', 'path': '.'})
  [tool result] ./README.md
./src/main.py
./pyproject.toml
  [tool] read_file({'path': './README.md'})
  [tool result] 1: # Example Project
2: This project is a small CLI...
  [tool] write_file({'path': 'summary.md', 'content': '# Project Summary\n\nThis project is a small CLI...'})

  [permission required] write_file
  Arguments: {"path": "summary.md", "content": "# Project Summary\n\nThis project is a small CLI..."}
  Allow this action? [y/n]: y
  [tool result] Wrote 61 bytes to summary.md
Assistant: Done. I created `summary.md` with a short summary of the project.

The agent could explore and read files without bothering us. But when it wanted to write a new file, it had to ask for permission.

If we deny the action, the tool does not run:

  [tool] run_bash({'command': 'rm -rf build'})

  [permission required] run_bash
  Arguments: {"command": "rm -rf build"}
  Allow this action? [y/n]: n
  [tool result] Permission denied: the user did not allow 'run_bash' to run. Do not retry this tool call without asking the user first.
Assistant: I did not remove the build directory because permission was denied.

This is the safety behavior we want. The agent can propose an action, but the user remains in control.

Now let's test the clarification tool:

$ python agent.py --mode acceptEdits
Agent started in 'acceptEdits' mode  (working dir: /Users/roger/project)
You: Update the config to use production
  [tool] glob_files({'pattern': '*config*', 'path': '.'})
  [tool result] ./config.dev.json
./config.prod.json
./src/config.py
  [tool] ask_question({'question': 'I found multiple config files. Which one should I update: config.dev.json, config.prod.json, or src/config.py?'})

  [agent] I found multiple config files. Which one should I update: config.dev.json, config.prod.json, or src/config.py?
  Your answer: src/config.py
  [tool result] src/config.py
  [tool] edit_file({'path': 'src/config.py', 'old_string': 'ENV = "development"', 'new_string': 'ENV = "production"'})
  [tool result] Edited src/config.py
Assistant: Updated `src/config.py` to use production.

In acceptEdits mode, the edit inside the working directory was allowed automatically. But the agent still asked a question first because it was not sure which file the user meant.

What You've Built

We now have an agent that is not only capable of using tools and planning long tasks, but also has a basic safety model around tool execution.

The agent can:

  • Read and explore without unnecessary interruption.
  • Track its work with planning tools.
  • Ask the user clarifying questions when it is genuinely blocked.
  • Prompt for permission before risky tool calls.
  • Automatically allow project-local edits in acceptEdits mode.
  • Refuse to execute a denied tool call and report that denial back to the model.

This is a big step toward building agents that can work autonomously without giving free reign to do whatever they want in your machine.

What's next?

Human in the loop is the first step towards security in AI Agents. But its not everything you should consider if you worry about agent security. In the next part, we will complete the missing parts of security by taking a look at sandboxing and audit logs.