惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
Project Zero
Project Zero
阮一峰的网络日志
阮一峰的网络日志
博客园 - 【当耐特】
大猫的无限游戏
大猫的无限游戏
The Register - Security
The Register - Security
C
Check Point Blog
Attack and Defense Labs
Attack and Defense Labs
L
LangChain Blog
Simon Willison's Weblog
Simon Willison's Weblog
S
Schneier on Security
Recorded Future
Recorded Future
GbyAI
GbyAI
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Y
Y Combinator Blog
量子位
A
About on SuperTechFans
I
Intezer
T
Threat Research - Cisco Blogs
MongoDB | Blog
MongoDB | Blog
U
Unit 42
C
CERT Recently Published Vulnerability Notes
Scott Helme
Scott Helme
Cisco Talos Blog
Cisco Talos Blog
P
Palo Alto Networks Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Spread Privacy
Spread Privacy
M
MIT News - Artificial intelligence
雷峰网
雷峰网
博客园 - 聂微东
NISL@THU
NISL@THU
The Hacker News
The Hacker News
G
Google Developers Blog
F
Full Disclosure
博客园 - Franky
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
P
Privacy & Cybersecurity Law Blog
博客园 - 叶小钗
酷 壳 – CoolShell
酷 壳 – CoolShell
T
The Blog of Author Tim Ferriss
Security Latest
Security Latest
T
Tenable Blog
Know Your Adversary
Know Your Adversary
Stack Overflow Blog
Stack Overflow Blog
K
Kaspersky official blog
Blog — PlanetScale
Blog — PlanetScale
博客园 - 司徒正美
C
Cybersecurity and Infrastructure Security Agency CISA
Martin Fowler
Martin Fowler
Schneier on Security
Schneier on Security

Show HN

GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - pileax-ai/pileax: PileaX is an all-in-one AI knowledge base system. 🍀 GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
Whissle — Personal AI for Research, Voice, and Everyday Tasks
ksingla025 · 2026-06-13 · via Show HN

Run VoiceAI locally

ASR, TTS, voice calling, diarization, metadata, AI coaching — one Docker command.
Models download automatically. No cloud dependency.

Quick start

$ docker run -d --name whissle \
  -p 9000:9000 -p 8001:8001 -p 8003:8003 \
  -v whissle-models:/models -v whissle-data:/data \
  -e VARIANT=en-full \
  -e ANTHROPIC_API_KEY=your-key \
  whissleasr/whissle-gateway:latest

VARIANT=

DEVICE=

en-full · Downloads ~2 GB on first run (cached after)

What happens when you run it:

═══════════════════════════════════════════════
  Whissle Gateway — en-full
═══════════════════════════════════════════════
No GPU detected → using CPU

Shared models:
  ✓ speaker encoder + VAD           26 MB
  ✓ punctuation                    254 MB
  ✓ ITN (English + Hinglish)       1.5 MB

Variant: en-full
  ✓ en-in-tech-misc (485 MB)
  ✓ KenLM ENGLISH (1.5 GB)

Auth:
  Mode:    local
  Token:   wh_a1b2c3d4e5f6... (admin)
  Manage:  curl -H 'Authorization: Bearer ...' localhost:9000/auth/tokens

Starting services...
  PostgreSQL: :5432  ●
  ASR:        :8001  ●
  Video:      :8002  ●
  TTS:        :8003  ●
  Agent:      :8765  ●
  Pipecat:    :8000  ●
  Gateway:    :9000  ●

API

Six interfaces — batch REST, streaming WebSocket, text-to-speech, video intelligence, voice calling, and an intelligent agent.

POST localhost:8001/transcribe

$ curl -X POST http://localhost:8001/transcribe \
    -F "file=@call.mp3" \
    -F "diarize=true" \
    -F "num_speakers=2" \
    -F "punctuation=true" \
    -F "metadata_prob=true" \
    -F "summarize=sales_coaching" \
    -o result.json

Response — transcript + metadata per segment + AI analysis

{
  "segments": [
    {
      "speaker":  "SPEAKER_00",
      "text":     "Hello, good morning.",
      "start":    1.0,  "end": 1.9,
      "metadata": {
        "emotion":  "EMOTION_NEUTRAL",
        "behavior": "BEHAVIOR_DIRECT",
        "role":     "ROLE_INTERVIEWER",
        "age":      "AGE_30_45",
        "gender":   "GENDER_MALE"
      },
      "words": [{"word": "Hello", "start": 1.0, "end": 1.3}]
    }
  ],
  "analysis": {
    "overall_score": 78,
    "buyer_outcome": "Converted",
    "practices":     { "followed": 6, "total": 8 },
    "highlights":    [...]
  }
}

Parameters

All parameters for POST /transcribe.

ParameterTypeDefaultDescription
filefilerequiredAudio file (MP3, WAV, FLAC, OGG, M4A)
languagestringautoLanguage hint: en, hi, zh
diarizeboolfalseSpeaker diarization
num_speakersintautoExact speaker count (if known)
punctuationbooltrueRestore punctuation and capitalization
itnbooltrueInverse text normalization (numbers, currency)
use_lmbooltrueKenLM language model beam search
metadata_probboolfalseProbability distributions for metadata
word_timestampsboolfalsePer-word start/end timestamps
speech_analysisboolfalseSpeech patterns (pace, fillers, fluency)
summarizestringAI analysis: true, sales_coaching, collections, or custom prompt
hotwordsstringComma-separated hotwords for boosting

AI analysis modes

Add -F "summarize=mode" to any transcription. The diarized transcript + metadata is sent to your configured LLM for analysis.

sales_coaching

Sales Coaching

8 best practices scored. Rep/buyer identification. Highlights with timestamps. Behavior labels per segment. Overall score 0–100.

collections

Collections Compliance

Identity verification, reason stated, amount mentioned, no harassment. Call outcome (Promise to Pay / Dispute / Hardship). Next action.

true

General Summary

Overview, participants, key topics, emotional dynamics, entities, outcome. Markdown format.

your prompt here

Custom Prompt

Pass any prompt string. The LLM receives your instructions + full transcript with per-segment metadata.

Models

Each model extracts different metadata in a single ASR forward pass — no separate models or API calls.

BEHAVIOREMOTIONEVALROLEAGEGENDERENTITY

120M params, 26 Behavioral codes for coaching, therapy, interviews. 8 evaluation labels.

English · 6 heads, 51 classes

INTENTEMOTIONROLEAGEGENDERENTITY

115M params, Debt collection intents — pay-back, disputes, hardship. Agent/Customer role detection.

Hindi-English · 5 heads, 26 classes

DIALECTAGEGENDERENTITY

160M params, Mandarin with North/South dialect detection.

Mandarin · 3 heads, 12 classes

INTENTEMOTIONAGEGENDERENTITY

600M params, inline action tokens. 31 intent groups, 18K vocabulary.

23 languages · 5,500+ action tokens

55 voices

Non-autoregressive text-to-speech. Sub-200ms TTFB on CPU. Always included.

10 languages · Baked in

CapitalizationNumbers

Punctuation restoration and inverse text normalization.

EN + Hinglish · Auto-downloaded

Metadata per segment

Every segment includes these tags. Common tags appear on all models. Additional tags depend on the model.

TagValuesModels
emotionEMOTION_NEUTRAL, EMOTION_HAPPY, EMOTION_SAD, EMOTION_ANGRY, EMOTION_FEAR, EMOTION_SURPRISEAll
ageAGE_0_18, AGE_18_30, AGE_30_45, AGE_45_60, AGE_60+All
genderGENDER_MALE, GENDER_FEMALEAll
behavior26 types (BEHAVIOR_EXPLAIN, BEHAVIOR_QUESTION, BEHAVIOR_ACKNOWLEDGE, ...)en-in-tech-misc
evalEVAL_CORRECT, EVAL_PROBE, EVAL_PARTIAL, EVAL_INCORRECT, EVAL_HINT, EVAL_SKIPen-in-tech-misc
roleROLE_INTERVIEWER / ROLE_INTERVIEWEE or ROLE_AGENT / ROLE_CUSTOMERen-in-tech-misc, hinglish-loans
intent13 collections intents or 31 general intents (INTENT_GREETING, INTENT_QUESTION, ...)hinglish-loans, whissle-large
dialectDIALECT_NORTH, DIALECT_SOUTH, DIALECT_OTHERSzh

Variants

Choose your variant based on language and quality needs. Switch by changing VARIANT= and restarting. Cached models are reused.

VariantLanguagesDownloadBest for
hinglishHindi-English~515 MBDebt collections, Hindi-English call centers
en-liteEnglish~500 MBQuick testing, development
en-fullEnglish~2 GBSales coaching, interviews, therapy
multi-full23 languages~4 GBMultilingual, highest quality
multi-zh23 langs + Mandarin~5 GBMultilingual + dialect detection
allAll~6 GBMaximum flexibility

Runs everywhere

From your laptop (CPU) to data center GPUs. Same Docker, same API. Auto-detects GPU.

HardwareVRAMVariantConcurrent
MacBook / LaptopCPUAny1–3
Mac Mini M4 Pro24 GB unifieden-full3–8
NVIDIA T416 GBen-lite5–10
RTX 409024 GBen-full20–50
A100 40GB40 GBmulti-full50–80
RTX 6000 Ada48 GBall50–100
H10080 GBall150–300
DGX Spark128 GB unifiedall30–60
H200141 GBall250–500
Docker TagArchRuntime
whissleasr/whissle-gateway:latestamd64CPU — Mac (Rosetta), Linux, Windows
whissleasr/whissle-gateway:gpuamd64NVIDIA CUDA 12.4 + onnxruntime-gpu

Architecture

┌──────────────────────────────────────────────────────────────┐
│                     Docker Container                        │
│                                                             │
│  ┌────────┐ ┌────────┐ ┌────────┐ ┌────────┐ ┌──────────┐  │
│  │ ASR    │ │ Video  │ │ TTS    │ │Pipecat │ │ Agent    │  │
│  │ :8001  │ │ :8002  │ │ :8003  │ │ :8000  │ │ :8765    │  │
│  │        │ │        │ │ Kokoro │ │        │ │          │  │
│  │ ONNX   │ │MediaPip│ │ 82M    │ │ WebRTC │ │ Any LLM  │  │
│  │ +KenLM │ │+Vision │ │55 voice│ │ Twilio │ │ Cloud or │  │
│  │ +ECAPA │ │  LLM   │ │        │ │Voice AI│ │  Local   │  │
│  │ +VAD   │ │        │ │        │ │        │ │          │  │
│  │ +Punct │ │Face    │ │        │ │ Auth   │ │Summarize │  │
│  │ +ITN   │ │Gesture │ │        │ │Multiorg│ │ Coach    │  │
│  └────────┘ └────────┘ └────────┘ └────────┘ └──────────┘  │
│                     │                                       │
│              ┌──────────────┐                               │
│              │ PostgreSQL   │                               │
│              │ :5432        │                               │
│              └──────────────┘                               │
│                                                             │
│  /models  (Docker volume — cached ASR models)               │
│  /data    (Docker volume — PostgreSQL, auth, conversations) │
└──────────────────────────────────────────────────────────────┘

whissle-models volume

ASR models, KenLM, punctuation, ITN. Downloaded on first run, cached forever. Survives container restarts.

whissle-data volume

Conversations, analytics, agent configs, auth tokens. Persists across restarts. Only deleted by docker volume rm.

Get started

One command. Models download automatically. Ready in 2 minutes.
Built for contact centers, sales intelligence, behavioral AI, and more.