惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
U
Unit 42
V
Visual Studio Blog
D
DataBreaches.Net
F
Fortinet All Blogs
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
月光博客
月光博客
大猫的无限游戏
大猫的无限游戏
T
The Blog of Author Tim Ferriss
GbyAI
GbyAI
博客园 - 叶小钗
Blog — PlanetScale
Blog — PlanetScale
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
MongoDB | Blog
MongoDB | Blog
The Cloudflare Blog
云风的 BLOG
云风的 BLOG
D
Docker
G
Google Developers Blog
罗磊的独立博客
博客园 - 三生石上(FineUI控件)
小众软件
小众软件
S
SegmentFault 最新的问题

Hacker News

GitHub - SeanFDZ/macmind: Single-layer transformer in HyperTalk for the classic Macintosh Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis Bonsai 1-bit WebGPU - a Hugging Face Space by webml-community Moving a large-scale metrics pipeline from StatsD to OpenTelemetry / Prometheus GitHub - Nightmare-Eclipse/RedSun: The Red Sun vulnerability repository GitHub - SethPyle376/hiraeth: Local AWS emulator focused on fast integration testing, with SQS support, SQLite-backed state, and a debug-friendly web UI. GitHub - macOS26/Agent: Any AI, replaces Claude Code, Cursor, OpenClaw. Over 18 LLM providers (Claude, OpenAI, Gemini, Ollama, Zai, HF, Qwen) wired into a native Mac app that writes code, builds Xcode projects, bumps versions, manages git, automates Safari, use AppleScript, JS or Accessibility, extend Agent! w/ MCP Servers, run tasks from your iPhone via Messages. YouTube now lets you turn off Shorts I Made a Terminal Pager Burgers | マクドナルド公式 Commands — HackerNews CLI documentation ChatGPT for Excel PiCore - Raspberry Pi Port of Tiny Core Linux Live Nation illegally monopolized ticketing market, jury finds Google Broke Its Promise to Me. Now ICE Has My Data. Founding Engineer at Adaptional | Y Combinator CRISPR takes important step toward silencing Down syndrome’s extra chromosome GitHub - saffron-health/libretto: The AI toolkit for building reliable browser automations US v. Heppner (S.D.N.Y. 2026) no attorney-client privilege for AI chats [pdf] Retrofitting JIT Compilers into C Interpreters IPv6 – Google The Accursèd Alphabetical Clock Cybersecurity Looks Like Proof of Work Now Fragments: April 14 Cal.com Goes Closed Source: Why AI Security Is Forcing Our Decision | Cal.com - Scheduling Software for Online Bookings Laravel raised money and now injects ads directly into your agent When moving fast, talking is the first thing to break Too much Discussion of the XOR swap trick – Heather Cafe Introduction to Spherical Harmonics for Graphics Programmers The Grand Line
GitHub - context-labs/HALO: Hierarchal Agent Loop Optimizer
mikepollard_ · 2026-06-24 · via Hacker News

✨ RLM-based agent optimizer using production traces✨

X (formerly Twitter) License GitHub

Quickstart • What is this? • Benchmarks • Development • Contributing

Quickstart

Install the HALO desktop app with:

curl -fsSL https://inference.net/halo/install.sh | sh

Read HALO reports

The installer downloads the latest release for your platform and sets up the desktop app. macOS uses a signed, notarized DMG. You can also install directly from the GitHub releases page.

If you're looking for a hosted, plug-and-play version of HALO, please sign up for inference.net and follow the instructions here.

What is this?

HALO is a methodology for building recursively self-improving agent harnesses using RLMs. This repository contains:

  • The HALO Desktop App for running HALO locally on your machine.
  • Information on HALO methodology.
  • A Python package that implements the core HALO-RLM engine. View on PyPI
  • A demo project that shows how to build HALO loops for your agents using the Python package. View demo
  • Benchmarking examples applying HALO to popular agent benchmarks. (View AppWorld).

HALO Loop

The core HALO loop is surprisingly simple:

  1. Collect execution traces from your agent harness. HALO uses OpenTelemetry-compatible tracing.
  2. Feed traces into HALO-RLM engine.
  3. The engine decomposes the traces to understand common failure modes across harness executions and produces a report with its findings.
  4. This report is fed into a coding agent like Cursor or Claude Code to generate and apply a set of changes to your harness.
  5. The harness is then re-deployed, more traces are gathered, and the cycle repeats.

HALO is great at finding issues in production agent deployments. We find high-traffic environments tend to generate more data with higher variance across executions, creating the type of issues that HALO is great at identifying.

Why an RLM?

A general-purpose harness like Claude Code is the wrong tool for trace analysis. This isn’t because the model isn’t smart, but because traces can get extremely long, and you need a specialized toolkit in order to make observations about systemic agentic behavior. We noticed in our testing that harnesses like CC would often overfit to an error present in a single/few traces rather than generalize to harness-level problems. This led us to creating a specialized form of a RLM.

rlm

Get Started

Install

Install the HALO engine + CLI from PyPI:

pip install halo-engine

# Verify installation
halo --help

Usage

  1. Integrate Tracing
  2. Collect traces by running your agent
  3. Run the HALO engine
export OPENAI_API_KEY=...
# Optional: point HALO at another OpenAI-compatible provider.
export OPENAI_BASE_URL=https://openrouter.ai/api/v1

halo path_to_your_traces.jsonl -p "Diagnose errors you find and suggest fixes"

HALO uses the canonical OpenAI env vars: OPENAI_API_KEY for credentials and OPENAI_BASE_URL for OpenAI-compatible providers. If OPENAI_BASE_URL is unset, HALO uses https://api.openai.com/v1. Run halo --help to see all CLI options. The CLI mirrors the model/provider settings exposed by the Python SDK's ModelConfig and ModelProviderConfig.

CLI options

Flag Default Description
TRACE_PATH required JSONL trace file
--prompt, -p required User prompt sent to the root agent
--model, -m gpt-5.4-mini Model name for root and subagent calls; also the fallback for synthesis and compaction
--synthesis-model --model Model for synthesis calls (trace summarization). A small, cheap model (e.g. gpt-4.1-nano) is recommended
--compaction-model --model Model for compaction calls (context summarization) — the biggest token consumer in large runs. A small, cheap model (e.g. gpt-4.1-nano) is recommended
--max-depth 2 Max subagent recursion depth
--max-turns 20 Max turns per agent
--max-parallel 10 Max concurrent subagents
--base-url OPENAI_BASE_URL / https://api.openai.com/v1 OpenAI-compatible API base URL
--api-key OPENAI_API_KEY Provider API key
--header, -H unset Provider header as NAME: VALUE. Repeat for multiple headers, matching curl's -H convention
--temperature provider default Sampling temperature forwarded to the model
--max-output-tokens provider default Maximum output tokens forwarded to the model
--parallel-tool-calls / --no-parallel-tool-calls enabled Allow models to issue parallel tool calls
--refusal-retries 0 Retry an agent model request this many times when the model refuses
--reasoning-effort model/provider default Reasoning effort for root and subagent calls.
--telemetry off Emit OpenInference traces of HALO's own LLM, tool, and agent activity

For example:

halo path_to_your_traces.jsonl \
  -p "Diagnose errors you find and suggest fixes" \
  --base-url https://openrouter.ai/api/v1 \
  -H "HTTP-Referer: https://example.com"

Telemetry

HALO can emit OpenInference-shaped traces of its own LLM, tool, and agent activity. It is off by default; nothing is emitted unless you pass --telemetry.

halo TRACE_PATH --prompt "..." --telemetry

When telemetry is enabled, CATALYST_OTLP_TOKEN uploads spans to inference.net Catalyst over OTLP. If it is unset, spans are written to a local JSONL file at ./halo-telemetry-{run_id}.jsonl in the current working directory.

Var Default Purpose
CATALYST_OTLP_TOKEN unset If set, uploads to Catalyst over OTLP. If unset, writes JSONL locally
CATALYST_OTLP_ENDPOINT catalyst-tracing default OTLP endpoint base URL, for example https://telemetry.inference.net
CATALYST_DEBUG unset Set to 1 to surface OTLP export errors
CATALYST_TRACING_RUN_ID unset Uses this HALO run id instead of a generated uuid
CATALYST_TRACING_* unset Generic catalyst-tracing passthrough
HALO_TELEMETRY_PATH ./halo-telemetry-{run_id}.jsonl Local fallback file path. Only used when CATALYST_OTLP_TOKEN is unset

We have provided a simple demo and an AppWorld demo.

Python API

The engine exposes four entry points from engine.main. Use whichever matches the trade-off you want between observability and code simplicity. The yielded types (AgentOutputItem and AgentTextDelta) are defined in engine/models/engine_output.py:

Function Sync / async Returns When to use
stream_engine_async async AsyncIterator[AgentOutputItem | AgentTextDelta] You want every event including streaming-token deltas (live UI, custom rendering).
stream_engine_output_async async AsyncIterator[AgentOutputItem] You want to log / persist each completed step (assistant message, tool call, tool result) as it lands.
run_engine_async async list[AgentOutputItem] You want the final list at the end and don't care about per-step observability.
stream_engine sync Iterator[AgentOutputItem | AgentTextDelta] Sync generator; yields every event including deltas. Drives the async iterator on a private event loop.
stream_engine_output sync Iterator[AgentOutputItem] Sync generator; yields completed items only. Same shape as the async variant for sync callers.
run_engine sync list[AgentOutputItem] Sync, collects to a list. Pure convenience over asyncio.run(run_engine_async(...)).
from engine.main import stream_engine_output_async

async for item in stream_engine_output_async(messages, cfg, trace_path):
    logger.info("step", extra={"sequence": item.sequence, "agent": item.agent_name})
    # item.item is an AgentMessage (assistant / tool / etc.)

Benchmarks

HALO is consistently capable of driving improvements on benchmarks, solely by optimizing the harness.

AppWorld

We applied HALO to the AppWorld benchmark, a set of agentic tasks that assess the LLM’s ability to use multi-app services like Spotify, Venmo, file systems, and phone contacts. We tested HALO’s ability to improve harnesses for both Gemini 3 Flash and Sonnet 4.6. We iterated on the harness using the dev split, and then used the test_normal split as a proxy to verify that improvements did not come from overfitting.

The feedback from HALO Engine surfaced failures in the harnesses such as hallucinated tool calls, redundant arguments in tools, refusal loops, and semantic correctness issues. Each issue mapped cleanly to a direct prompt edit. HALO’s claims were independently verified from the source trace files with the findings holding up under scrutiny.

app-world-sgc

The peak improvements over baseline were substantial for both models. For Gemini 3 Flash, dev SGC went from 36.8% to 52.6% (+15.8 points) and test_normal SGC went from 37.5% to 48.2% (+10.7 points). For Sonnet 4.6, dev SGC went from 73.7% to 89.5% (+15.8 points) and test_normal SGC went from 62.5% to 73.2% (+10.7 points).

Development

Local development against this repo uses uv for dependency management and go-task as the task runner.

Setup

git clone https://github.com/context-labs/HALO
cd HALO
task env:setup

task env:setup installs uv (if missing), syncs the venv from uv.lock, and configures the repo's git hooks. After that, the halo CLI is available via uv run halo ... (or activate .venv/).

Common tasks

Run task --list for the full list. The ones you'll use most:

Task What it does
task check Run all pre-commit checks: pinned-versions, lint, format, typecheck, unit tests
task check:fix Same, but auto-fix lint/format issues
task test:unit Unit tests under tests/unit/
task test:integration Integration tests under tests/integration/

License

MIT

Contributing

Contributions are welcome! Please feel free to submit a pull request.