惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

U
Unit 42
Vercel News
Vercel News
博客园 - 叶小钗
大猫的无限游戏
大猫的无限游戏
MyScale Blog
MyScale Blog
P
Proofpoint News Feed
量子位
Engineering at Meta
Engineering at Meta
B
Blog RSS Feed
博客园 - 【当耐特】
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
D
DataBreaches.Net
Stack Overflow Blog
Stack Overflow Blog
博客园 - 聂微东
小众软件
小众软件
Hugging Face - Blog
Hugging Face - Blog
人人都是产品经理
人人都是产品经理
IT之家
IT之家
T
The Blog of Author Tim Ferriss
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Jina AI
Jina AI
博客园 - 三生石上(FineUI控件)

Show HN

GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal).
Test LLMs Side-by-Side
dhavalt · 2026-06-17 · via Show HN

Iterate. Compare. Benchmark.

A local-first desktop client designed to test, grade, and benchmark prompts across major LLMs. Stop guessing how a model will perform and prove it against your datasets.

Parallel Model Testing

Send a single prompt template to GPT-4, Claude 3, and Gemini simultaneously. Instantly compare raw JSON outputs, latency metrics, and exact token consumption side-by-side without managing multiple browser tabs.

Local-First Privacy

Your API keys and prompt history are stored in a local SQLite database. Nothing touches our servers.

Automatic Prompt Checkpointing

Every iteration is automatically saved to your local database. Fork a prompt to test a new variable, track the exact changes that improved the output, and easily revert to past configurations.

Benchmark & Evaluate

Inject test data into your prompt templates to establish a baseline. When a new LLM drops, benchmark it against your historical data before trusting it in production.

Model Benchmarking

Run your prompt against a full test dataset across multiple models at once. Review the batch outputs side-by-side and assign pass/fail grades to see exactly which model handles your edge cases.

Version Control for Your Prompts.

Keep a clean history of your iterations. Fork a prompt to test a new variable, track the changes, and easily switch back to past versions.


Request-Level Debugging.

Chat interfaces hide the details. Inspect raw API responses, latency stats, and exact token usage for every single request.

Model Benchmarking

Run your prompt against a full test dataset across multiple models at once. Review the batch outputs side-by-side and assign pass/fail grades to see exactly which model handles your edge cases.

Bring Your Own Keys.

Keep your credentials on your machine. Your keys are encrypted via your OS keyring, saved to your local database, and sent strictly to the providers. We track nothing.


Credentials Vault

1. Provider Setup

Bring your own keys. Connect OpenAI, Anthropic, Mistral, Gemini and XAI in seconds. Toggle models on/off to keep your workspace clean.

2. Inference Settings

Adjust temperature, top_p, and frequency penalties to observe how different constraints impact your prompt results.

Under the Hood

We chose Electron for cross-platform support, but kept the stack as simple as possible.

Native Web Components

No heavy frameworks overhead. We built the interface using standard HTML, CSS, and vanilla JavaScript.

Local SQLite Database

Your data lives in a standard SQLite file on your disk. Backup, version control, or delete it whenever you want.

{
  "runtime": "Electron",
  "security": "Context Isolated",
  "frontend": "Vanilla JS + Web Components",
  "database": "SQLite3 (Local-only)"
}

Ship AI Features With Certainty.

Batch-test your datasets and prove model reliability before hitting production.

Download for Mac Download for Windows Download for Linux

Learn about our permanent licensing and early-adopter pricing:
View License Details →