惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

人人都是产品经理
人人都是产品经理
有赞技术团队
有赞技术团队
WordPress大学
WordPress大学
月光博客
月光博客
T
Tailwind CSS Blog
阮一峰的网络日志
阮一峰的网络日志
小众软件
小众软件
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Last Week in AI
Last Week in AI
大猫的无限游戏
大猫的无限游戏
S
SegmentFault 最新的问题
罗磊的独立博客
Jina AI
Jina AI
酷 壳 – CoolShell
酷 壳 – CoolShell
宝玉的分享
宝玉的分享
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 三生石上(FineUI控件)
量子位
雷峰网
雷峰网
Apple Machine Learning Research
Apple Machine Learning Research
美团技术团队
博客园 - 聂微东
V
V2EX

Hacker News: Show HN

PurrrrrFocus: Pomodoro Timer App - App Store Workflow Engine — Multi-Step Orchestration for Bun RapidPhoto: Pro Photo Editor App - App Store GitHub - think41/extrasuite: Token-efficient pull/edit/push workflow for AI agents editing Google Workspace files (Sheets, Docs, Slides, Forms) GitHub - DheerG/swarms: Achieve extraordinary results with claude code across a variety of tasks SPICE simulation → oscilloscope → verification with Claude Code — Lucas Gerads Show HN: VCoding – A 5 MB native Windows IDE with no dynamic dependencies Show HN: LLMs don't hallucinate because they're bad at math, it's the format GitHub - Agent-FM/agentfm-core: AgentFM is a peer-to-peer network that turns everyday computers into a decentralized AI supercomputer. AgentFM lets you run massive AI workloads directly across a global mesh of idle CPUs and GPUs. Show HN: Tracking Top US Science Olympiad Alumni over Last 25 Years GitHub - Potarix/agent-hub: One place to talk to all your agents Show HN: Runtime security for AI agents(injection,tool abuse, data exfiltration) GitHub - dubeyKartikay/lazyspotify: Terminal Spotify client for macOS and Linux GitHub - the-banana-tool/king-louie: Easy to use GUI Personal AI Assistant. Win/Linux/Mac. Show HN I made my vacation rental bookable by AI agents–no Airbnb, 0% commission GitHub - basteez/jsf-autoreload: maven plugin to enable hot reload on jsf projects uvm32/hosts/host-gdbstub at main · ringtailsoftware/uvm32 GitHub - labsai/EDDI: Config-driven engine that turns JSON into production-grade AI agents. Multi-agent orchestration, 12+ LLM providers, MCP/A2A protocols, RAG, persistent memory, and enterprise compliance (EU AI Act, GDPR, HIPAA). Built on Quarkus. GitHub - glitchnsec/fortyone-oss: AI Executive Assistant Platform Quickstart | Alien GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. GitHub - ocrbase-hq/ocrbase: 📄 PDF/IMG ->.MD/JSON Document OCR API for PaddleOCR and GLMOCR. Self-hostable. GitHub - impactjo/home-memory: MCP server that lets your AI assistant remember everything about your home. GitHub - Sets88/dbcls: DbCls is a powerful terminal database client that supports various databases GitHub - neptun2000/heor-agent-mcp GitHub - SeanFDZ/macmind: Single-layer transformer in HyperTalk for the classic Macintosh RollQuation: Math Puzzles - Apps on Google Play GitHub - dropbox/witchcraft Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis GitHub - opentalon/opentalon: OpenTalon is an open-source platform built from the ground up in Go as a robust alternative to OpenClaw
Whisper API: Self-Hostable Speech to Text Transcription
Ved Gupta Portfolio · 2026-04-11 · via Hacker News: Show HN

April 11, 2026

Whisper API is an open-source, self-hosted service for speech-to-text. It runs whisper.cpp under the hood and exposes a Deepgram-compatible surface (/v1/listen over REST and WebSocket), so full control of audio and data is maintained while reusing familiar integration patterns.

Screenshot_2026-04-11_at_10.52.37_PM.png

Key features

  • Familiar API — Drop-in style compatibility with /v1/listen for uploads, JSON bodies (including transcribe-from-URL), and streaming.
  • Rich output — JSON with timings, plus SRT and VTT export options.
  • Live streaming — Real-time transcription over WebSockets (16 kHz PCM).
  • Advanced options — Custom vocabulary-style prompting, optional audio windowing (start / duration) where supported.
  • Secure operations — API keys via a small CLI; URL ingest includes SSRF-oriented limits (size caps, host controls; redirects off by default).
  • Documentation — Browse the live whisper.api docs (installation, auth, REST, WebSocket streaming, models, Docker, examples). The same site is built from the docs/ folder (Astro Starlight); run it locally with Bun if you are developing the docs.

Stack (high level)

FastAPI, SQLAlchemy (SQLite by default; PostgreSQL-friendly), async HTTP and WebSockets, and the whisper-cli binary from whisper.cpp. See the repo requirements.txt and setup_whisper.sh for the full setup path.

Quick start

Install dependencies, configure the environment, and pull/build whisper assets:

pip install -r requirements.txt
cp .env.example .env
chmod +x setup_whisper.sh
./setup_whisper.sh

Initialize storage and create an API key:

python -m app.cli init
python -m app.cli create --name "MyAdminKey"

Start the server (default dev port 7860):

uvicorn app.main:app --host 0.0.0.0 --port 7860

For local-only testing, Swagger can expose POST /v1/auth/test-token when ENABLE_TEST_TOKEN_ENDPOINT=true in .env. That flag defaults to off and should stay off in production.

Example: transcribe a file

curl -X POST 'http://localhost:7860/v1/listen' \
  -H "Authorization: Token <YOUR_KEY>" \
  -H "Content-Type: audio/wav" \
  --data-binary @audio.wav

14993b62-9ffb-464a-8e9e-b4a605e00958.png

Example: transcribe from URL

The server fetches the audio, with configurable download limits and safety defaults:

curl -X POST 'http://localhost:7860/v1/listen' \
  -H "Authorization: Token <YOUR_KEY>" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/audio.mp3"}'

Documentation

Online: https://whisper.vedgupta.in/docs/ — architecture overview, setup, authentication, REST and WebSocket API reference, models, Docker deployment, code examples, and contributing.

Local (from the repo):

cd docs && bun install && bun run dev

Reference and credits

Loading Mascot