惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
The GitHub Blog
The GitHub Blog
Vercel News
Vercel News
D
DataBreaches.Net
MongoDB | Blog
MongoDB | Blog
H
Help Net Security
小众软件
小众软件
美团技术团队
T
The Blog of Author Tim Ferriss
爱范儿
爱范儿
D
Docker
Martin Fowler
Martin Fowler
大猫的无限游戏
大猫的无限游戏
博客园 - 聂微东
Blog — PlanetScale
Blog — PlanetScale
H
Hackread – Cybersecurity News, Data Breaches, AI and More
罗磊的独立博客
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
V2EX
S
SegmentFault 最新的问题
云风的 BLOG
云风的 BLOG
B
Blog
雷峰网
雷峰网
The Cloudflare Blog

Show HN

GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal).
Test LLMs Side-by-Side
dhavalt · 2026-06-17 · via Show HN

Iterate. Compare. Benchmark.

A local-first desktop client designed to test, grade, and benchmark prompts across major LLMs. Stop guessing how a model will perform and prove it against your datasets.

Parallel Model Testing

Send a single prompt template to GPT-4, Claude 3, and Gemini simultaneously. Instantly compare raw JSON outputs, latency metrics, and exact token consumption side-by-side without managing multiple browser tabs.

Local-First Privacy

Your API keys and prompt history are stored in a local SQLite database. Nothing touches our servers.

Automatic Prompt Checkpointing

Every iteration is automatically saved to your local database. Fork a prompt to test a new variable, track the exact changes that improved the output, and easily revert to past configurations.

Benchmark & Evaluate

Inject test data into your prompt templates to establish a baseline. When a new LLM drops, benchmark it against your historical data before trusting it in production.

Model Benchmarking

Run your prompt against a full test dataset across multiple models at once. Review the batch outputs side-by-side and assign pass/fail grades to see exactly which model handles your edge cases.

Version Control for Your Prompts.

Keep a clean history of your iterations. Fork a prompt to test a new variable, track the changes, and easily switch back to past versions.


Request-Level Debugging.

Chat interfaces hide the details. Inspect raw API responses, latency stats, and exact token usage for every single request.

Model Benchmarking

Run your prompt against a full test dataset across multiple models at once. Review the batch outputs side-by-side and assign pass/fail grades to see exactly which model handles your edge cases.

Bring Your Own Keys.

Keep your credentials on your machine. Your keys are encrypted via your OS keyring, saved to your local database, and sent strictly to the providers. We track nothing.


Credentials Vault

1. Provider Setup

Bring your own keys. Connect OpenAI, Anthropic, Mistral, Gemini and XAI in seconds. Toggle models on/off to keep your workspace clean.

2. Inference Settings

Adjust temperature, top_p, and frequency penalties to observe how different constraints impact your prompt results.

Under the Hood

We chose Electron for cross-platform support, but kept the stack as simple as possible.

Native Web Components

No heavy frameworks overhead. We built the interface using standard HTML, CSS, and vanilla JavaScript.

Local SQLite Database

Your data lives in a standard SQLite file on your disk. Backup, version control, or delete it whenever you want.

{
  "runtime": "Electron",
  "security": "Context Isolated",
  "frontend": "Vanilla JS + Web Components",
  "database": "SQLite3 (Local-only)"
}

Ship AI Features With Certainty.

Batch-test your datasets and prove model reliability before hitting production.

Download for Mac Download for Windows Download for Linux

Learn about our permanent licensing and early-adopter pricing:
View License Details →