惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

B
Blog
D
Docker
J
Java Code Geeks
腾讯CDC
Blog — PlanetScale
Blog — PlanetScale
G
Google Developers Blog
M
MIT News - Artificial intelligence
L
LangChain Blog
T
The Blog of Author Tim Ferriss
P
Proofpoint News Feed
MyScale Blog
MyScale Blog
博客园 - Franky
GbyAI
GbyAI
Hugging Face - Blog
Hugging Face - Blog
aimingoo的专栏
aimingoo的专栏
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 聂微东
N
Netflix TechBlog - Medium
B
Blog RSS Feed
Y
Y Combinator Blog
阮一峰的网络日志
阮一峰的网络日志
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Google DeepMind News
Google DeepMind News

Show HN

GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal).
GitHub - infiniteregrets/kv-psi
infiniteregr · 2026-06-28 · via Show HN

PSI KV Governor

PSI KV Governor is a small reference implementation for using Linux Pressure Stall Information to trim an LLM KV cache when the system is under memory pressure.

Requirements

  • Linux with PSI enabled: cgroup memory.pressure or /proc/pressure/memory
  • Python 3.10+
  • llama.cpp build dependencies for the runner
  • a GGUF model, for example models/SmolLM2-135M-Instruct-Q2_K.gguf

Check PSI:

cat /proc/pressure/memory
PYTHONPATH=src python benchmarks/pressure_bench.py --preflight-only

Basic Usage

Run the reference simulator:

PYTHONPATH=src python -m psi_kv_governor.cli simulate

Build the llama.cpp runner:

scripts/build_llama_runner.sh

Download the small benchmark model if needed:

python scripts/download_demo_model.py

PSI Benchmark

Run both variant orders. This matters because PSI avg10, cache, and zram/swap state can carry over from the first pressure run into the second.

PYTHONPATH=src python benchmarks/pressure_bench.py \
  -c 2048 \
  -n 1536 \
  --keep 64 \
  --tail 256 \
  --min-prune 64 \
  --pressure-mib 6000 \
  --pressure-step-mib 1024 \
  --pressure-warmup-s 10 \
  --variant-cooldown-s 45 \
  --out-dir data/bench-pressure/fixed-first

PYTHONPATH=src python benchmarks/pressure_bench.py \
  --variant-order psi-first \
  -c 2048 \
  -n 1536 \
  --keep 64 \
  --tail 256 \
  --min-prune 64 \
  --pressure-mib 6000 \
  --pressure-step-mib 1024 \
  --pressure-warmup-s 10 \
  --variant-cooldown-s 45 \
  --out-dir data/bench-pressure/psi-first

Recent Jetson result:

run variant decoded tok/s prunes final KV external PSI some/full
fixed-first fixed 1536 94.00 0 1547 1.61/1.61
fixed-first PSI 1536 88.80 4 1291 4.14/3.94
psi-first PSI 1536 96.16 2 1004 2.46/2.33
psi-first fixed 1536 89.76 0 1547 5.56/5.56

Result directories:

  • data/bench-pressure/real-psi-6000m-1536tok-cooldown
  • data/bench-pressure/real-psi-6000m-1536tok-cooldown-psi-first