惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Privacy International News Feed
The Register - Security
The Register - Security
Microsoft Azure Blog
Microsoft Azure Blog
P
Proofpoint News Feed
M
MIT News - Artificial intelligence
Recorded Future
Recorded Future
H
Hackread – Cybersecurity News, Data Breaches, AI and More
F
Fortinet All Blogs
G
Google Developers Blog
Engineering at Meta
Engineering at Meta
B
Blog
aimingoo的专栏
aimingoo的专栏
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
N
Netflix TechBlog - Medium
Martin Fowler
Martin Fowler
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
MyScale Blog
MyScale Blog
L
LangChain Blog
T
The Blog of Author Tim Ferriss
U
Unit 42
Blog — PlanetScale
Blog — PlanetScale
C
Check Point Blog
Vercel News
Vercel News
Microsoft Security Blog
Microsoft Security Blog
D
DataBreaches.Net
Recent Announcements
Recent Announcements
云风的 BLOG
云风的 BLOG
Stack Overflow Blog
Stack Overflow Blog
博客园 - 聂微东
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 司徒正美
月光博客
月光博客
Jina AI
Jina AI
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
WordPress大学
WordPress大学
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - Franky
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Hugging Face - Blog
Hugging Face - Blog
Last Week in AI
Last Week in AI
The Last Watchdog
The Last Watchdog
P
Privacy & Cybersecurity Law Blog
有赞技术团队
有赞技术团队
G
GRAHAM CLULEY
腾讯CDC
Cyberwarzone
Cyberwarzone
爱范儿
爱范儿
I
Intezer
SecWiki News
SecWiki News

Show HN

GitHub - steveking-gh/firmion: Firmion is DSL and engine for firmware image generation. GitHub - villagesql/villagesql-skills: Agent skills for VillageSQL - gemini-cli-extension; claude-code-plugin GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
GitHub - shahar-dagan/openfusion: Combine the results from a panel of models into an enhanced response
shadag · 2026-06-18 · via Show HN

CI License: MIT Python 3.11+

An open-source, drop-in compound-model proxy. Point any OpenAI-compatible tool at it, set model: "openfusion", and your prompt is fanned out to a panel of LLMs in parallel — then a judge model reads every response (consensus, contradictions, blind spots) and streams back a single synthesized answer that aims to beat any one of them.

It's the open version of the mixture-of-agents idea behind OpenRouter's Fusion: better answers from models you already pay for, as a tunable, forkable recipe instead of a black box.

openfusion playground — panel fan-out to judge synthesis

Quick start · How it works · Playground · Routing & strategies · vs. OpenRouter Fusion · Benchmarks · Contributing

Project layout

New here? You only need the first two to run it; the rest is for tuning and contributing.

Path What it is
openfusion/ The proxy (FastAPI). Start with server.py; see docs/ARCHITECTURE.md for the module map.
web/ The playground UI source (React + shadcn). Built assets ship in openfusion/static/.
examples/ Copy-paste config recipes (preset, dev, panel, bench…). You don't need a config to start.
bench/ Reproducible head-to-head harness; bench/FINDINGS.md is where fusion does and doesn't pay off.
DESIGN.md · docs/ Design rationale, architecture, and security notes.

Status

Beta — panel fan-out, judge synthesis, SSE streaming, web-tool fusion, an Auto Router, debate/ vote/ranked aggregators, production limits, and an interactive playground. See DESIGN.md and docs/ARCHITECTURE.md for architecture and security notes.

Quick start

openfusion has two front ends — an interactive terminal chat and a web playground. No clone, no config, no env vars needed to start.

Chat in your terminal

uvx --from git+https://github.com/shahar-dagan/openfusion openfusion   # ephemeral, needs uv
# …or: pip install git+https://github.com/shahar-dagan/openfusion && openfusion

Bare openfusion drops you into a Rich-rendered chat with the model panel — a banner, a live panel-progress spinner, Markdown answers with syntax-highlighted code, and slash commands (/preset, /tokens, /models, /key, /clear). On first run it asks for your OpenRouter key and saves it (~/.config/openfusion/credentials), so later runs don't re-prompt; use /key to change it. Pipe for one-shots: echo "…" | openfusion.

Web playground

openfusion web                                  # opens the playground in your browser
# …or: docker run -p 8000:8000 ghcr.io/shahar-dagan/openfusion

openfusion web pops the playground open at http://localhost:8000 once the server is ready (pass --no-open, or it's skipped automatically in non-interactive/headless/Docker contexts). Paste your key (kept only in server memory) and fuse. With nothing configured it boots the Budget preset (a diverse panel + judge with web search) so the first run lands where fusion actually wins.

Install the command everywhere (no venv to activate)

uv tool install .     # from a clone — or: pipx install . && pipx ensurepath

For active development, pip install -e . inside an activated venv (the command then works only while that venv is active). A bare pip install -e . does not put openfusion on your global PATH — see Troubleshooting.

For a fixed recipe, write an openfusion.yaml (start from examples/preset.yaml.examplepreset: quality | budget, or examples/default.yaml.example for a fully spelled-out panel/judge). A preset expands to a diverse OpenRouter panel + judge with web tools on, mirroring OpenRouter Fusion's Quality/Budget switch:

Preset Panel Judge Tools
quality Claude Sonnet 4 · Gemini 3 Pro · DeepSeek V4 Pro Claude Sonnet 4 web search + fetch
budget GPT-4o-mini · DeepSeek V4 Pro · Kimi K2.6 DeepSeek V4 Pro web search + fetch

Use as a drop-in API from the OpenAI SDK (with openfusion web running):

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="local-dev")
stream = client.chat.completions.create(
    model="openfusion",
    messages=[{"role": "user", "content": "Explain mixture-of-agents in one paragraph."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Or straight from the terminal, no server needed:

openfusion ask "Compare Postgres and SQLite for a small SaaS." --max-tokens 800

ask runs one fusion against your configured panel and streams the synthesized answer to stdout (panel progress goes to stderr). --max-tokens caps every call — lower is faster and cheaper.

Speed & length. Fusion runs N panel calls plus a judge, so it's slower than one model — the panel runs in parallel and the judge streams as soon as the panel finishes. The judge is prompted to stay concise, and you cap length with --max-tokens (CLI), max_tokens (API), the response- length control in the playground Settings, or cost_controls in config.

Routing & strategies

Three knobs control whether and how a prompt is fused. All are optional and off/default.

  • Auto Router (router.enabled: true) — a per-prompt gate that answers simple prompts with a single pass-through call and reserves the panel for prompts that look like they benefit (long, analytical, or containing code). Default is a cheap heuristic (no extra model call); mode: model uses a small classifier model and falls back to the heuristic if it errors:

    router:
      enabled: true
      mode: heuristic     # heuristic | model | always | never
      min_chars: 280      # prompts at/over this length fuse
      # classifier:       # required for mode: model
      #   base_url: https://openrouter.ai/api/v1
      #   api_key: ${OPENROUTER_API_KEY}
      #   model: openai/gpt-4o-mini
  • Strategy (strategy:) — how the panel is produced: self_fusion (one model sampled N times), panel (a fixed diverse panel), or debate (a diverse panel where each member revises after seeing the others' answers, then the judge synthesizes). Debate trades extra cost/latency for cross-examination:

    strategy: debate
    debate:
      rounds: 1           # revision rounds before the judge
  • Aggregator (aggregator:) — how answers become one: judge (synthesis, default), vote (majority vote, cheaper, best for verifiable short-answer tasks), or ranked (one short judge call picks the single best answer — cheaper than synthesis, uses model judgment unlike vote).

  • Analysis transparency (analysis.emit: true) — surface the judge's structured reasoning (consensus / contradictions / partial coverage / unique insights / blind spots) as a separate SSE event: analysis (and an analysis field on non-streaming responses), without polluting the answer body.

  • Prompt caching (cache.enabled: true) — mark the shared prefix so self-fusion's N samples reuse a cached prompt on providers that support it (a no-op elsewhere).

Production limits

For public deployments, bound load and spend (both default to 0 = unlimited):

limits:
  max_in_flight: 64           # cap concurrent requests; over-limit returns 503
  rate_limit_per_minute: 60   # per gateway key (or per client when unauthenticated); over-limit returns 429

These are best-effort, single-process guards — pair them with provider-side budgets and, for multi-replica deployments, an edge rate limiter.

How it works

A request to model: "openfusion" is fanned out to a panel of models in parallel (each optionally doing its own web research), then a judge model reads every answer and synthesizes one — streamed back over SSE, with the structured analysis and cost alongside.

flowchart LR
    C["Client<br/>(Cursor · OpenAI SDK · anything)"] -->|"POST /v1/chat/completions<br/>model=openfusion"| R{"Router<br/><i>(optional)</i>"}
    R -->|simple prompt| S["Single model"] --> OUT
    R -->|worth fusing| P

    subgraph P ["Panel · parallel fan-out"]
        direction TB
        A["Model A 🔍"]
        B["Model B 🔍"]
        D["Model C 🔍"]
    end

    P --> J["Judge<br/>consensus · contradictions · blind spots"]
    J --> OUT["Streamed answer (SSE)<br/>+ analysis + token/cost"]
    C -.->|other model / client tools| S

    classDef accent fill:#eef2ff,stroke:#4f46e5,color:#3730a3;
    class J,R accent;
Loading
  • Drop-in. OpenAI-compatible POST /v1/chat/completions + /v1/models, real SSE streaming.
  • No lock-in. Each panel member + judge is {base_url, api_key, model}. OpenRouter is the default upstream; OpenAI, Together, local vLLM/Ollama all work.
  • Config-driven. Panel, judge, strategy, aggregator, router, and limits live in openfusion.yaml — or a one-word preset, or nothing at all (zero-config quick start).

openfusion vs. OpenRouter Fusion

openfusion is the open implementation of the same idea. The core mechanism is at parity; the differences are scale and a per-prompt router.

OpenRouter Fusion openfusion
Parallel panel → judge synthesis
Synthesis dimensions consensus · contradictions · partial coverage · unique insights · blind spots same
Web search + fetch on the panel ✅ (default) ✅ (on by default with preset:)
Quality / Budget presets ✅ (preset: quality | budget)
Override panel + judge ✅ (plugin fields) ✅ (any {base_url, api_key, model} in YAML)
Per-call cost breakdown ✅ (Activity) ✅ (SSE usage event + /metrics)
Self-hostable / forkable ❌ closed API ✅ MIT, any OpenAI-compatible provider
Per-prompt Auto Router ✅ heuristic or model classifier (router.enabled)
Structured analysis surfaced analysis.emit (SSE analysis event)
Multi-round debate strategy: debate
Concurrency cap + rate limiting limits (best-effort, single-process)
Interactive web playground ✅ embedded at /playground (zero-build)
Headline benchmark full DRACO (100 tasks) DRACO subset (10 tasks) — see bench/FINDINGS.md

Parameter precedence

Parameter Applies to Notes
temperature (client) Judge only indirectly via recipe Self-fusion varies panel temps from config, not client
max_tokens, stop, response_format Judge (visible output) Panel members use recipe defaults
stream, stream_options Judge path Panel always runs non-streamed internally
tools / tool_calls Fusion or pass-through Server-executable web tools (openrouter:web_search/web_fetch) are fused; client-side function tools and mid-conversation tool turns pass through

Environment variables

Variable Purpose
OPENROUTER_API_KEY Default upstream key (via ${OPENROUTER_API_KEY} in config)
OPENFUSION_CONFIG Path to config file (default: openfusion.yaml)
OPENFUSION_API_KEYS Comma-separated gateway allowlist (optional)
OPENFUSION_HOST / OPENFUSION_PORT Server bind address

Cost safety and live smoke tests

cost_controls in config caps max_tokens for pass-through, panel, and judge calls. Missing max_tokens values are filled from the configured ceiling; over-limit pass-through and judge requests return 400, while internal panel calls clamp to their ceiling.

Run the opt-in live OpenRouter smoke test only when you intend to spend a small number of credits:

export OPENROUTER_API_KEY=your-key
python scripts/openrouter_smoke.py --config examples/dev.yaml.example --yes-spend-credits

Benchmarks

Run the head-to-head benchmark (self-fusion vs solo model):

pip install -e ".[dev]"
python bench/run.py --config examples/default.yaml.example --tasks bench/tasks/sample.jsonl

Use --tasks bench/tasks/smoke.jsonl --max-tokens 32 before larger benchmark runs.

Each run reports accuracy plus the spend it took to get there — total_tokens and total_cost_usd per mode — so you can weigh any accuracy change against the extra cost of fanning out to a panel.

What we measure today

The bundled bench/tasks/sample.jsonl (20 short Q&A tasks) is saturated for a capable model — the solo baseline already scores ~100%, so there is no headroom for fusion to add accuracy. On a recent run with openai/gpt-4o-mini (self-fusion N=2, max_tokens=32):

Mode Accuracy Avg latency Tokens Cost
Solo 100% (20/20) 0.55s 536 $0.0001
Self-fusion 95% (19/20) 1.40s 4,669 $0.0008

So on easy tasks fusion does not beat a single call — it costs more (here ~9× the tokens) and can even regress, because the judge only has trivially-correct answers to choose between. This is expected: mixture-of-agents helps where a single model is unreliable, not where it is already right.

openfusion makes no "beats frontier" claim. Demonstrating where fusion earns its cost needs a harder eval (one the solo baseline does not already ace) scored on quality per dollar, not accuracy alone. That eval is in progress; this table will be updated to show where fusion does and doesn't pay off. Claim only what your own bench/run.py run proves on your model and tasks.

Observability

The proxy exposes Prometheus metrics at GET /metrics (no auth; scrape-only, bind accordingly):

  • openfusion_requests_total{route,outcome} — client-facing requests (fusion / pass_through).
  • openfusion_upstream_requests_total{phase,outcome} — upstream calls by panel / judge / pass_through.
  • openfusion_panel_members_total{outcome} — per-member success vs. degraded failures.
  • openfusion_tokens_total{phase,kind} and openfusion_cost_usd_total{phase} — token and cost spend.
  • openfusion_request_latency_ms / openfusion_upstream_latency_ms — latency summaries (_count + _sum).

Cost (usage.cost, when the upstream reports it) is also rolled into the per-request SSE event: usage payload and the non-streaming usage field, so a single fusion call shows what it spent across the panel and judge. Per-call structured logs remain on the openfusion.upstream logger.

Playground

The server hosts an interactive playground at GET /playground (and GET / redirects there). It's a React + Tailwind + shadcn UI whose built assets ship in the package (no Node needed to run); it talks only to the local /v1 API, so provider keys never reach the browser. You can:

  • paste your OpenRouter API key on first run (held only in server memory; enabled by allow_ui_api_key, on for the zero-config quick start),
  • pick a Quality / Budget / Custom panel and a "Fuse with" judge model,
  • toggle web search, send a prompt, and watch the panel → synthesis progress,
  • read the streamed answer plus the judge's structured analysis (consensus / contradictions / blind spots) and the token + cost breakdown.

The model selectors are editable when the server sets allow_request_overrides: true (on for the quick start), which enables the per-request openfusion: { preset | panel | judge | tools } field (mirroring OpenRouter Fusion's analysis_models/model plugin fields). Overrides reuse the server's upstream credentials — clients choose model ids, never keys — and stay bounded by gateway auth, cost ceilings, and rate limits. Read GET /v1/config for the active panel/judge and flags.

Developing the UI

The UI source lives in web/ (Vite + React + TypeScript + Tailwind v4 + shadcn-style components):

cd web
npm install
npm run dev      # dev server (proxy /v1 to a running openfusion on :8000)
npm run build    # writes built assets into openfusion/static/playground/ (commit them)

Troubleshooting

openfusion: command not found — the console script lives in the environment you installed it into. Either install it as a tool so it's always on PATH (uv tool install . or pipx install .), or activate the venv you used (source .venv/bin/activate). A bare pip install -e . does not put openfusion on your global PATH.

Playground says "Couldn't reach the server" — open the page at the URL the running server prints (default http://localhost:8000), not a dev-server port or a standalone file.

No upstream API key — set OPENROUTER_API_KEY, run openfusion setup, or paste your key into the playground.

Stack

Backend: Python 3.11+ / FastAPI / httpx / uvicorn. Frontend: React / Vite / Tailwind / shadcn.

Contributing

Contributions are welcome — openfusion is meant to be forked and tuned. See CONTRIBUTING.md for dev setup and the PR checklist, and CODE_OF_CONDUCT.md. Please report security issues privately per SECURITY.md rather than as a public issue.

License

MIT.