惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
美团技术团队
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
月光博客
月光博客
J
Java Code Geeks
Jina AI
Jina AI
罗磊的独立博客
宝玉的分享
宝玉的分享
S
SegmentFault 最新的问题
D
DataBreaches.Net
博客园 - 叶小钗
腾讯CDC
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Last Week in AI
Last Week in AI
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Google DeepMind News
Google DeepMind News
阮一峰的网络日志
阮一峰的网络日志
B
Blog
V
Visual Studio Blog
雷峰网
雷峰网
博客园 - 【当耐特】
Apple Machine Learning Research
Apple Machine Learning Research
Engineering at Meta
Engineering at Meta
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报

Show HN

Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code.
GitHub - intent-bench/intent-bench: Intent fulfillment be...
ryan4rtmx · 2026-05-29 · via Show HN

An open-source benchmark measuring whether providing structured intent to coding agents improves implementation effectiveness.

What This Measures

Existing agent benchmarks (SWE-bench, HumanEval, Aider Polyglot) test single-requirement tasks or bug fixes. Real engineering involves multi-requirement features with dependency ordering, cross-cutting concerns, and state machines. intent-bench fills this gap.

Research question: Does providing structured requirements (with dependency ordering, acceptance criteria, and verification hooks) to a coding agent reduce token waste and increase task completion rate on complex greenfield implementations?

Design

intent-bench uses a controlled A/B design:

  • Control: Agent receives only a task prompt in an empty directory
  • Treatment: Agent receives the same prompt, plus structured intent artifacts seeded into the working directory

Both conditions use identical prompts. The treatment condition additionally receives intent via a pluggable treatment layer (e.g., RTMX with MCP server, or a plain markdown specification file).

Treatments

Treatment Description Delivery
rtmx RTMX requirements database + MCP server .rtmx/ dir + MCP config
manual-spec Plain markdown requirements file REQUIREMENTS.md in workdir

Experiments

Experiment Complexity Requirements Depth
url-shortener Baseline 10 2
task-manager Standard 13 5

Quick Start

# Prerequisites: bash 4+, python3, Claude Code CLI
make setup

# Validate configuration
make validate

# Run a single experiment
bash bench.sh run url-shortener --condition control --runs 5
bash bench.sh run url-shortener --condition treatment --runs 5

# Analyze results
make analyze
make charts

Results

Results are recorded in results/summary.csv (the data ledger). Each row captures a complete session: token counts, outcome, wall clock time, and knowledge entropy score.

Statistical analysis uses Mann-Whitney U for token comparison and Fisher exact test for completion rate comparison, with bootstrap confidence intervals.

Metrics

Metric Description
Completion rate Fraction of runs where all tests pass
Total tokens Input + output tokens consumed
Token efficiency ratio Control tokens / Treatment tokens
Knowledge entropy 0-10 score measuring agent process quality
Variance (CV) Coefficient of variation across runs
Backtrack rate Fraction of turns re-reading previously read files

Adding Treatments

Create treatments/<name>.sh with this interface:

# treatments/<name>.sh validate
#   Exit 0 if dependencies are met
# treatments/<name>.sh setup <workdir> <experiment> <fixture_dir>
#   Seed intent artifacts into workdir. Exit 0 = ready.

Adding Agents

Create agents/<name>.sh with this interface:

# agents/<name>.sh <workdir> <model> <prompt_file> <result_dir> <max_budget>
#   Must produce: $result_dir/transcript.jsonl, $result_dir/stderr.log
#   Exit 0 = completed, non-zero = crashed

Subject Selection Criteria

See docs/subject-criteria.md for what qualifies as a valid experiment, complexity classifications, and minimum sample size requirements.

Reproducing Results

See REPRODUCING.md for exact commands, cost estimates, and how to submit community results.

Related Work

  • SWE-bench: Single bug fixes in existing repos
  • HumanEval: Single-function code generation
  • Aider Polyglot: Multi-language refactoring
  • ProjDevBench: Multi-file project development
  • FeatureBench: Feature implementation
  • SWE-EVO: Evolving complexity

intent-bench uniquely combines: greenfield multi-requirement tasks, explicit dependency ordering, causal A/B design with blinding, and pluggable intent delivery mechanisms.

License

Apache 2.0