惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
L
LangChain Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
云风的 BLOG
云风的 BLOG
月光博客
月光博客
F
Full Disclosure
G
Google Developers Blog
MongoDB | Blog
MongoDB | Blog
T
Tailwind CSS Blog
F
Fortinet All Blogs
A
About on SuperTechFans
Stack Overflow Blog
Stack Overflow Blog
J
Java Code Geeks
Microsoft Azure Blog
Microsoft Azure Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
Recorded Future
Recorded Future
Y
Y Combinator Blog
博客园 - 聂微东
爱范儿
爱范儿
D
DataBreaches.Net
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Threat Research - Cisco Blogs
L
Lohrmann on Cybersecurity
The Hacker News
The Hacker News
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Scott Helme
Scott Helme
L
LINUX DO - 热门话题
Apple Machine Learning Research
Apple Machine Learning Research
C
CERT Recently Published Vulnerability Notes
B
Blog RSS Feed
The Last Watchdog
The Last Watchdog
SecWiki News
SecWiki News
Webroot Blog
Webroot Blog
Engineering at Meta
Engineering at Meta
T
Tenable Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
O
OpenAI News
Google DeepMind News
Google DeepMind News
Security Archives - TechRepublic
Security Archives - TechRepublic
W
WeLiveSecurity
Hacker News: Ask HN
Hacker News: Ask HN
Hacker News - Newest:
Hacker News - Newest: "LLM"
T
Troy Hunt's Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
aimingoo的专栏
aimingoo的专栏
GbyAI
GbyAI
V
Vulnerabilities – Threatpost
N
News and Events Feed by Topic

Hacker News: Show HN

PurrrrrFocus: Pomodoro Timer App - App Store Workflow Engine — Multi-Step Orchestration for Bun RapidPhoto: Pro Photo Editor App - App Store GitHub - DheerG/swarms: Achieve extraordinary results with claude code across a variety of tasks SPICE simulation → oscilloscope → verification with Claude Code — Lucas Gerads Show HN: VCoding – A 5 MB native Windows IDE with no dynamic dependencies Show HN: LLMs don't hallucinate because they're bad at math, it's the format GitHub - Agent-FM/agentfm-core: AgentFM is a peer-to-peer network that turns everyday computers into a decentralized AI supercomputer. AgentFM lets you run massive AI workloads directly across a global mesh of idle CPUs and GPUs. Show HN: Tracking Top US Science Olympiad Alumni over Last 25 Years GitHub - Potarix/agent-hub: One place to talk to all your agents Show HN: Runtime security for AI agents(injection,tool abuse, data exfiltration) GitHub - dubeyKartikay/lazyspotify: Terminal Spotify client for macOS and Linux GitHub - the-banana-tool/king-louie: Easy to use GUI Personal AI Assistant. Win/Linux/Mac. Show HN I made my vacation rental bookable by AI agents–no Airbnb, 0% commission GitHub - basteez/jsf-autoreload: maven plugin to enable hot reload on jsf projects uvm32/hosts/host-gdbstub at main · ringtailsoftware/uvm32 GitHub - labsai/EDDI: Config-driven engine that turns JSON into production-grade AI agents. Multi-agent orchestration, 12+ LLM providers, MCP/A2A protocols, RAG, persistent memory, and enterprise compliance (EU AI Act, GDPR, HIPAA). Built on Quarkus. GitHub - glitchnsec/fortyone-oss: AI Executive Assistant Platform Quickstart | Alien GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. GitHub - ocrbase-hq/ocrbase: 📄 PDF/IMG ->.MD/JSON Document OCR API for PaddleOCR and GLMOCR. Self-hostable. GitHub - impactjo/home-memory: MCP server that lets your AI assistant remember everything about your home. GitHub - Sets88/dbcls: DbCls is a powerful terminal database client that supports various databases GitHub - neptun2000/heor-agent-mcp GitHub - SeanFDZ/macmind: Single-layer transformer in HyperTalk for the classic Macintosh RollQuation: Math Puzzles - Apps on Google Play GitHub - dropbox/witchcraft Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis GitHub - opentalon/opentalon: OpenTalon is an open-source platform built from the ground up in Go as a robust alternative to OpenClaw LinkedIn™ 职位抓取工具 - Chrome 应用商店 GitHub - EdoardoBambini/Agent-Armor-Iaga: AI agents are getting tool access — shell, file system, databases, APIs, secrets. But **nobody is governing what they actually do with it**. Frameworks like LangChain, CrewAI, AutoGen, and Claude Code give agents the power to execute. Agent Armor gives you the power to control, audit, and approve every single action before it happens. HN Vibes — Week 15, Apr 7–13 2026 GitHub - chojs23/ec: Easy terminal-native 3-way git mergetool vim-like workflow GitHub - SethPyle376/hiraeth: Local AWS emulator focused on fast integration testing, with SQS support, SQLite-backed state, and a debug-friendly web UI. GitHub - JakOb-dotcom/cloud-sandbox-security-analysis: Technical analysis and Proof of Concept (PoC) regarding environment variable exfiltration in containerized cloud sandboxes via side-channel data leaks. Springboards - Flint Alpha Show HN: A simpler coding agent harness GitHub - audiodude/sudomake-friends GitHub - 256thFission/mini-mythos: OSS clone of Anthropic’s Mythos harness to locate C/C++ memory vulnerabilities Show HN: OpenParallax: OS-level privilege separation for AI agent execution Hacker News Sorted - Chrome 应用商店 Show HN: How to Install Docker on Ubuntu 24.04 LTS: Complete 2026 Guide GitHub - himanshudongre/smriti GitHub - sverrirsig/claude-control: macOS desktop dashboard for monitoring and managing multiple Claude Code sessions GitHub - ory/dockertest: Write better integration tests! Dockertest helps you boot up ephermal docker images for your Go tests with minimal work. Chiral - Chrome 应用商店 Show HN: Two Claudes collaborating through shared memory on a $100 mini-PC GitHub - pmichaillat/latex-cv: Minimalist LaTeX template for academic CVs GitHub - oguzbilgic/posse: A web UI for Anthropic Managed Agents. GitHub - sshiraz/depsly: Dependency risk analysis tool for npm packages ABI Add safari/agent-harness — Safari browser automation via safari-mcp by achiya-automation · Pull Request #212 · HKUDS/CLI-Anything GitHub - Halfblood-Prince/trustcheck: Verify PyPI package attestations and improve Python supply-chain security GitHub - oguzbilgic/kern-ai: Agents that do the work and show it. GitHub - bruits/satteri: High-performance Markdown and MDX processing for the JavaScript ecosystem GitHub - tylergibbs1/feedstock: High-performance web crawler and scraper for TypeScript, powered by Bun and Playwright GitHub - Grimm67123/grimmbot: The self-improving sandboxed and open-source AI agent. With persistent memory and scheduling. GitHub - whitevanillaskies/whitebloom: Local whiteboard that blooms. GitHub - hwdsl2/docker-whisper: Docker image for a self-hosted Whisper speech-to-text server with speaker diarization and OpenAI-compatible transcription and translation APIs. Powered by faster-whisper. Supports all Whisper models, NVIDIA GPU (CUDA) acceleration, JSON/SRT/VTT output, SSE streaming, offline mode, and multi-arch (amd64, arm64). GitHub - yisding/reviewwiggum GitHub - MarwanAlsoltany/serrors: Structured errors for Go: sentinel hierarchies, typed data, custom formatting, and slog integration. GitHub - soatok/age-php GitHub - Luthiraa/markitme GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits GitHub - tombedor/excalicharts GitHub - wh1le/excalidraw-edit: Open and edit .excalidraw files from the terminal. Offline, auto-saves to disk. MalExt Sentry - Malicious Extension Scanner - Chrome 应用商店 GitHub - syi0808/asciianimesvg: Generate animated ASCII art SVGs from text. CLI, Rust library, WASM, and web editor. GitHub - zaina-ml/ml_forge: A visual-based graph node editor for training computer vision models. GitHub - anakin87/llm-rl-environments-lil-course: 🌱 A little course on Reinforcement Learning Environments for evaluating and training Language Models GitHub - takaakit/superpowers-uml: Superpowers-UML modifies Superpowers to ensure a software development workflow in which AI agents design through UML modeling. AdriByte Studio - Sviluppo Web e Soluzioni Digitali GitHub - chouligi/angel-copilot: Your personalized Angel Investment Advisor Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 GitHub - agenteractai/lodmem: Level Of Detail Context Management for Agents GitHub - ostefani/subnetlens: A fast, concurrent network scanner with a TUI and plain-text CLI, built in Go. It discovers live hosts on your network, scans their open ports, resolves hostnames, and fingerprints operating systems—delivered. Cyber Pulse: Agentic Intel - Apps on Google Play Whisper API: Self-Hostable Speech to Text Transcription The Agent-Web Protocol Stack: A Research Thesis GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Show HN: Provepy – A Python decorator that proves your code using Lean and LLMs Show HN: Pardonned.com – A searchable database of US Pardons GitHub - patrickdappollonio/dux: Dux is a terminal UI that lets you run multiple AI coding agents side by side, each in its own git worktree, with full companion terminals, macros, commit generation, and a command palette that knows more tricks than you do. kMC Crystal Simulator Show HN: HyperFlow – A self-improving agent framework built on LangGraph GitHub - stef41/vibescore: 🎵 Grade your vibe-coded project. One command, instant letter grade across security, quality, dependencies, and testing. GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. imgur.com GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. GitHub - nowork-studio/toprank: Open-source Claude Code skills for SEO, SEM, Google Ads GitHub - tacomanator/sash: Lightweight macOS menu bar app for reliably cycling through windows of the current application. Appents | Social Media Management for Product-First Teams GitHub - pnhoang/youtube-spam-blocker: Automatically detects and hides spam messages in YouTube Live chat. Set rate limits, keyword filters, and block repeat offenders. GitHub - decisionnode/DecisionNode: CLI + Local MCP - A shared structured memory store across Claude Code, Cursor, Windsurf, Antigravity, and every MCP client. Semantically queryable. GitHub - AvaCodeSolutions/django-email-learning: An open source Django app for creating email-based learning platforms with IMAP integration and React frontend components. The $100K Gap in Kubernetes Security Tooling Function Calling Harness: From 6.75% to 100%
compresh-benchmarks/epbench/WRITEUP.md at main · compresh/compresh-benchmarks
compresh · 2026-06-23 · via Hacker News: Show HN

Fewer tokens, same recall: reconstruct context, don't resend it

Most LLM apps do the same thing every turn: they resend the entire conversation. The transcript grows linearly, each turn costs more than the last, and — past a point — the model gets worse, not better, because long context degrades (the "lost in the middle" effect, and what people now call context rot: quality drops well before the nominal window is full).

Compresh takes a different path. Instead of resending the whole history, it reconstructs a query-aware slice of it each turn — the part of the past this turn actually needs. The obvious question is whether recall survives when you stop sending the whole thing. So we measured it on an independent benchmark, and we publish where it wins and where it loses.

The axis: savings × quality, not recall alone

Most agent-memory work optimizes one number: recall or accuracy. We care about a different one — how few tokens you can send while holding quality. Two measurements:

  • Compression. On 360 real StackExchange Q&A items, replayed as one long, growing session, our open-source core (tulbase) sent 66% fewer input tokens (40.9M → 13.9M) with no measurable quality loss (answer equivalence 87.5% vs 90.0% raw; cosine 0.667 vs 0.670).
  • Reconstruction (the paid memory layer, TUL 2.0). On a strong model, a single turn goes from 31,947 → 275 input tokens (−99.1%) — it sends a query-aware slice, not the conversation. (The system prompt is left untouched.)

Fewer tokens is easy if you don't care about answers. The point is holding quality — so here's the benchmark.

The benchmark

We used EpBench — an independent, published episodic-memory benchmark (ICLR 2025; built on Tulving's model of recall): cued questions over a long, generated book. Same answerer (gpt-5-mini) and the same judge across every arm, scored with the benchmark's own method — no home-field scoring.

Method Simple recall Context read
raw / full context 0.804 196 chapters
naive RAG · chapter 0.796 17 chapters
Compresh · TUL 2.0 0.828 query-aware

The point is the juxtaposition — recall is essentially at parity while tokens are not:

EpBench · Simple Recall (paper method) · gpt-5-mini
──────────────────────────────────────────────────────
  Compresh · TUL 2.0  0.828 [█████████████████░░░]  query-aware slice
  raw / full context  0.804 [████████████████░░░░]  196 chapters
  naive RAG · top-17   0.796 [████████████████░░░░]  17 chapters

Input tokens / turn (strong model, long chat)
──────────────────────────────────────────────────────
  raw                31,947 [████████████████████]
  Compresh              275 [▏░░░░░░░░░░░░░░░░░░░]  −99.1%

Compresh has the highest simple recall while reading a query-aware slice, not the whole ~103k-token book — and pulls further ahead on multi-event questions (full per-bin breakdown in results/). Judge caveat, stated up front: our judge was OpenRouter gpt-4o; the paper's own judge puts raw at 0.830 — within ~2 points. Same judge for all arms.

You can reproduce the headline in ~10 seconds, no API keys: verify.py recomputes Simple Recall (the paper method — an unweighted mean over the matching-event bins) from the published per-bin recalls and checks it against the scoreboard.

Where it loses — and why that's the honest part

On chronological ordering, naive RAG beats us: 0.65 vs 0.44. Retrieving a query-relevant slice breaks temporal contiguity, so "put these events in order" gets harder. We publish that number next to the wins.

This isn't a confession of inferiority — it's the nature of the field. Every approach here trades something. Long context keeps everything and loses the middle. RAG retrieves by similarity and loses coherence and order. Summarization keeps a gist and loses detail. Reconstruction keeps what the turn needs and (today) loses some chronology. Loss is already everywhere in context and memory systems; the only real choice is whether you measure it and say so. We did, and we published both sides.

"But what about prefix caching?"

A fair objection: you don't have to recompute a stable prefix — providers cache it. True, and prefix caching is a real, powerful serving optimization. But it's worth being precise about what it does and doesn't do:

  • It makes resending a lot cheaper to serve. It does not make the history smaller — you still ship the whole transcript every turn, just at a discount on the cached part.
  • It does nothing for the quality problem. Lost-in-the-middle degradation is orthogonal to caching: a perfectly cached 100k-token context still loses the middle. So the recall result above stands regardless.

In other words, prefix caching optimizes the symptom (recompute cost), not the cause (you're sending too much). And cached tokens still aren't free — roughly 10–50% of base input price depending on provider.

So the honest cost comparison isn't "Compresh vs raw-without-caching." It's Compresh vs raw + prefix caching. Modeling that — generously to the cached baseline (assuming a full cache hit every turn, ignoring cache-write premiums and TTL misses) — the crossover is around ~10k tokens of history. Below that, raw+cache can be cheaper: Compresh has a small fixed per-turn overhead. Above it, Compresh wins, and the gap widens as the conversation grows, because raw scales with length while the reconstructed slice stays roughly flat. (This is a modeled result; a live-capture confirmation is in progress.) The takeaway is honest and narrow: this is a long-conversation argument, not a "cheaper for everything" one.

How it works (briefly)

Each turn, Compresh takes the full history, builds a query-aware reconstruction of the older part (compresh_md), and keeps a protected tail of recent raw turns (raw_tail). The model receives compresh_md + raw_tail instead of the full transcript. It differs from RAG — we reconstruct the conversation per turn, not retrieve documents — and from prompt compression like LLMLingua — we don't drop tokens by perplexity; we rebuild the query-relevant history. The system prompt is never compressed.

Reproduce it / try it

  • Verify the headline (~10s, no keys): ../verify.py
  • Full re-run (calls the models + an independent judge, needs keys): REPRODUCE.md
  • Try Compresh: one line — change your base_url, keep everything else. Free to start, no card, pay only on the tokens it removes: compre.sh

We'd genuinely value pushback on the method and the cost model.