惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
Cyberwarzone
Cyberwarzone
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
W
WeLiveSecurity
博客园 - 叶小钗
Hugging Face - Blog
Hugging Face - Blog
Security Latest
Security Latest
Scott Helme
Scott Helme
TaoSecurity Blog
TaoSecurity Blog
N
Netflix TechBlog - Medium
爱范儿
爱范儿
Application and Cybersecurity Blog
Application and Cybersecurity Blog
G
Google Developers Blog
F
Fortinet All Blogs
N
News and Events Feed by Topic
V2EX - 技术
V2EX - 技术
Google Online Security Blog
Google Online Security Blog
L
LINUX DO - 热门话题
NISL@THU
NISL@THU
The GitHub Blog
The GitHub Blog
Spread Privacy
Spread Privacy
S
Secure Thoughts
T
Tailwind CSS Blog
Google DeepMind News
Google DeepMind News
Recorded Future
Recorded Future
N
News and Events Feed by Topic
SecWiki News
SecWiki News
S
Security @ Cisco Blogs
A
About on SuperTechFans
云风的 BLOG
云风的 BLOG
L
Lohrmann on Cybersecurity
P
Palo Alto Networks Blog
Know Your Adversary
Know Your Adversary
IT之家
IT之家
人人都是产品经理
人人都是产品经理
Attack and Defense Labs
Attack and Defense Labs
Hacker News - Newest:
Hacker News - Newest: "LLM"
MyScale Blog
MyScale Blog
宝玉的分享
宝玉的分享
T
The Blog of Author Tim Ferriss
H
Hacker News: Front Page
T
Tenable Blog
C
CERT Recently Published Vulnerability Notes
D
DataBreaches.Net
阮一峰的网络日志
阮一峰的网络日志
Help Net Security
Help Net Security
博客园_首页
S
Securelist
罗磊的独立博客

Show HN

GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - pileax-ai/pileax: PileaX is an all-in-one AI knowledge base system. 🍀 GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
GitHub - patrickxia/entropic: Entropic: information-driven variable-rate media playback
pxx · 2026-05-31 · via Show HN

Variable-rate audio/video playback that slows down for unfamiliar or high-information words and accelerates through predictable speech, using token-level surprisal from a fast local LLM.

How it works

  1. Transcribe with WhisperX (word-level timestamps via forced alignment)
  2. Score each word's surprisal: unigram (word frequency) or contextual (causal LM like distilgpt2/gpt2)
  3. Estimate each word's confidence using whisper token-level probabilities, optionally from a separate weaker model (--uncertainty-model)
  4. Assign speeds per one of two modes:
    • VBR (default): speed ∝ 1/surprisal
    • Skiplow: words above an info-rate threshold stay at 1x; words below get sped up proportionally; silences are fast-forwarded
  5. Time-stretch via librubberband.so (ctypes, real-time mode [to avoid clipping between samples and to obtain output timestamps])
  6. Retime video (if applicable) via mkvmerge timecodes; optionally burn-in debug subtitles (requires re-encode)

Surprisal model

  • Unigram (default): word rarity via wordfreq
  • Contextual (--model gpt2): causal LM surprisal with sliding window
  • Rarity blend (--rarity 1.5, default when using --model): adds rarity × unigram_surprisal, see Development Notes below.

Requirements

System packages

# Debian/Ubuntu
sudo apt install ffmpeg mkvtoolnix librubberband-dev
  • ffmpeg: audio extraction, muxing, subtitle burn-in
  • mkvtoolnix (mkvmerge): video frame retiming
  • librubberband (>= 3.0): time-stretching via C API

Python

Requires Python 3.11+.

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

Note on WhisperX: installed from git since it's not on PyPI. It pulls in faster-whisper, pyannote-audio, and torch as transitive dependencies.

Note on torch: the default pip install pulls CUDA 12.x wheels (~2 GB). For CPU-only:

pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt

Usage

Audio

# Basic — 1.5x average speed, unigram surprisal (VBR mode)
python entropic.py podcast.mp3

# 2x with contextual surprisal
python entropic.py podcast.mp3 -s 2 --model distilgpt2

# Custom speed bounds
python entropic.py lecture.mp3 -s 1.8 --min-speed 0.8 --max-speed 3.0 --model gpt2

# Skiplow mode — keep speech at 1x, compress silences and low-info words
python entropic.py podcast.mp3 --mode skiplow

# Skiplow targeting 2x overall speed (auto-computes threshold)
python entropic.py podcast.mp3 --mode skiplow -s 2

# Skiplow with explicit threshold (bits/s)
python entropic.py podcast.mp3 --mode skiplow --threshold 80

Video

# Retime video (no re-encode, copies video stream)
python entropic.py lecture.mp4 -s 2 --model distilgpt2

# With burned-in subtitles (re-encodes video with SVT-AV1)
python entropic.py lecture.mp4 -s 2 --subtitles --model distilgpt2

Options

Flag Default Description
--mode vbr vbr (target average speed) or skiplow (enhanced skip-silence)
-s, --speed 1.5 Target average speed. In skiplow, auto-computes threshold to achieve this
--threshold - Min info rate (bits/s) before speedup (skiplow only, overrides --speed)
--min-speed 1.0 Minimum playback speed
--max-speed 3.5 Maximum playback speed
--silence-speed max-speed Speed for gaps between words
--model gpt2 Causal LM for contextual surprisal (gpt2, distilgpt2, none for unigram-only)
--rarity 0.1 Unigram rarity weight when using --model
--clarity 1.5 Clarity penalty strength (0 = disabled)
--whisper-model turbo WhisperX model for transcription (tiny.en, base.en, turbo, etc.)
--uncertainty-model same as whisper-model Separate model for word confidence. A weaker model (tiny.en) gives more honest uncertainty on mumbled speech
--language en Language code for transcription, alignment, and word frequency
--subtitles off Generate word/speed/surprisal overlay. Video: burned in. Audio-only: saved as .ass file
--device cpu cpu or cuda for WhisperX and LM
--transcript auto Path to transcript cache JSON
--no-cache off Force re-transcription

Transcript caching

Transcripts are cached to <input>.transcript.json. Re-running with different speed/spread/model parameters reuses the cached word timestamps and only recomputes surprisal and speeds. Use --no-cache to force re-transcription (needed when changing --whisper-model).

Development notes

Maybe more suitable for a blog post, but we might as well write the beginnings of it.

Constant information rate

Originally the design was to target a constant output information rate, so that each word's output duration depends only on its surprisal. This ends up ignoring the input duration entirely, which isn't entirely information-free (speakers might slow down naturally to indicate emphasis). Constant output information rate also causes significant distortion; already-short words were made even shorter in duration due to their low information content. The current implementation has speed inversely proportional to information only, preserving a bit more of the original speech rhythm.

City of Manaus problem

While testing initial prototypes, the city "Manaus" was well-predicted and accelerated within context. Even a weak LLM like GPT-2 can predict reasonably well from "the Brazilian city of ", but for listeners who are not familiar with Brazilian geography, accelerating this unfamiliar word is not productive. In some ways, this indicates how even weak LLMs are "too intelligent" for this process. To compensate for this, --rarity was added, but the weight is a heuristic. Perhaps a pure unigram model is better for most usecases, but it still does stand to reason that rarer words, when encountered multiple times, should be sped up on repetition.

Implementation bugs introduced by Claude Code (with model Opus 4.7)

Opus 4.7 is a very powerful model but it made some very bizarre choices, causing the introduction of many hard-to-debug bugs that had to be resolved in parallel to the design questions above. Those included:

  • Transcription uncertainty was taken from the whisperx alignment model, not the transcription step, entirely corrupting the confidence metric.
  • Surprisal for multiple-token words was averaged across tokens, including punctuation, which is also essentially noise.
  • A subtrahend of 12 was used to convert from wordfreq.zipf_frequency which is occurrences-per-billion (1e9). It's unclear why we're even using zipf_frequency to begin with; using the raw frequency wouldn't have this detuning knob.
  • Binary search for the speedup with clamping was implemented as a bizarre greedy floodfill that would return the wrong answer frequently but not always (of note, it would never get the right answer if the naive solution of no speedup was already within the constraints). The binary search is still implemented "competition programming" style, but at least it works now.
  • min-speed and max-speed kept going back to this weird implementation of multiplier applied to the target speed.
  • Transcription uncertainty was multiplied into the surprisal via hacky parameters despite the fact that it can be natively described as "bits of entropy". The weight still is heuristic, so the digression here is a little more understandable.

It's very unclear if development was truly sped up by using an LLM. There are likely more bugs. Everything always looked like it was working, but things would feel "off," and the debug process would take quite some time. I ran out of Opus tokens and I migrated to deepseek-v4-pro midway through, which was able to debug and fix these issues when directly pointing them out (though I'm sure Opus would have too).

Automatically debugging A/V sync issues with just the prompt "the A/V sync is broken" worked quite well though. Those are normally a nightmare to fix.

Clarity penalty for multiple speakers

Some podcasts feature one host who is very clear and one host who mumbles all of their words (e.g. Money Stuff). The goal of the clarity penalty was to make the less comprehensible host more comprehensible. Whisper is "too good" at this -- even Matt Levine can be predicted with reasonably high accuracy. Fixing this is a WIP, but --uncertainty-model is maybe the first step. tiny.en sometimes has better average log-probability of understanding Katie Greifeld than Matt Levine but they rarely differ by much, which makes this harder. Perhaps we should inject additional noise? Was whisper trained predominantly on male voices?

Disclaimer

This is a personal project. The views, code, and opinions expressed here do not represent those of my current or past employers.