惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
A
About on SuperTechFans
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 【当耐特】
W
WeLiveSecurity
博客园 - 三生石上(FineUI控件)
The Cloudflare Blog
I
InfoQ
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Application and Cybersecurity Blog
Application and Cybersecurity Blog
雷峰网
雷峰网
Hacker News - Newest:
Hacker News - Newest: "LLM"
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
T
Troy Hunt's Blog
S
SegmentFault 最新的问题
Help Net Security
Help Net Security
博客园_首页
博客园 - 叶小钗
O
OpenAI News
PCI Perspectives
PCI Perspectives
月光博客
月光博客
人人都是产品经理
人人都是产品经理
B
Blog RSS Feed
GbyAI
GbyAI
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
The Last Watchdog
The Last Watchdog
C
CXSECURITY Database RSS Feed - CXSecurity.com
有赞技术团队
有赞技术团队
D
Darknet – Hacking Tools, Hacker News & Cyber Security
腾讯CDC
Hacker News: Ask HN
Hacker News: Ask HN
I
Intezer
Y
Y Combinator Blog
阮一峰的网络日志
阮一峰的网络日志
Spread Privacy
Spread Privacy
T
Tailwind CSS Blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
量子位
Cyberwarzone
Cyberwarzone
The Hacker News
The Hacker News
N
News and Events Feed by Topic
P
Proofpoint News Feed
Scott Helme
Scott Helme
D
Docker
Know Your Adversary
Know Your Adversary
Recent Commits to openclaw:main
Recent Commits to openclaw:main
TaoSecurity Blog
TaoSecurity Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Tor Project blog

Hacker News: Show HN

PurrrrrFocus: Pomodoro Timer App - App Store Workflow Engine — Multi-Step Orchestration for Bun RapidPhoto: Pro Photo Editor App - App Store GitHub - DheerG/swarms: Achieve extraordinary results with claude code across a variety of tasks SPICE simulation → oscilloscope → verification with Claude Code — Lucas Gerads Show HN: VCoding – A 5 MB native Windows IDE with no dynamic dependencies Show HN: LLMs don't hallucinate because they're bad at math, it's the format GitHub - Agent-FM/agentfm-core: AgentFM is a peer-to-peer network that turns everyday computers into a decentralized AI supercomputer. AgentFM lets you run massive AI workloads directly across a global mesh of idle CPUs and GPUs. Show HN: Tracking Top US Science Olympiad Alumni over Last 25 Years GitHub - Potarix/agent-hub: One place to talk to all your agents Show HN: Runtime security for AI agents(injection,tool abuse, data exfiltration) GitHub - dubeyKartikay/lazyspotify: Terminal Spotify client for macOS and Linux GitHub - the-banana-tool/king-louie: Easy to use GUI Personal AI Assistant. Win/Linux/Mac. Show HN I made my vacation rental bookable by AI agents–no Airbnb, 0% commission GitHub - basteez/jsf-autoreload: maven plugin to enable hot reload on jsf projects uvm32/hosts/host-gdbstub at main · ringtailsoftware/uvm32 GitHub - labsai/EDDI: Config-driven engine that turns JSON into production-grade AI agents. Multi-agent orchestration, 12+ LLM providers, MCP/A2A protocols, RAG, persistent memory, and enterprise compliance (EU AI Act, GDPR, HIPAA). Built on Quarkus. GitHub - glitchnsec/fortyone-oss: AI Executive Assistant Platform Quickstart | Alien GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. GitHub - ocrbase-hq/ocrbase: 📄 PDF/IMG ->.MD/JSON Document OCR API for PaddleOCR and GLMOCR. Self-hostable. GitHub - impactjo/home-memory: MCP server that lets your AI assistant remember everything about your home. GitHub - Sets88/dbcls: DbCls is a powerful terminal database client that supports various databases GitHub - neptun2000/heor-agent-mcp GitHub - SeanFDZ/macmind: Single-layer transformer in HyperTalk for the classic Macintosh RollQuation: Math Puzzles - Apps on Google Play GitHub - dropbox/witchcraft Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis GitHub - opentalon/opentalon: OpenTalon is an open-source platform built from the ground up in Go as a robust alternative to OpenClaw LinkedIn™ 职位抓取工具 - Chrome 应用商店 GitHub - EdoardoBambini/Agent-Armor-Iaga: AI agents are getting tool access — shell, file system, databases, APIs, secrets. But **nobody is governing what they actually do with it**. Frameworks like LangChain, CrewAI, AutoGen, and Claude Code give agents the power to execute. Agent Armor gives you the power to control, audit, and approve every single action before it happens. HN Vibes — Week 15, Apr 7–13 2026 GitHub - chojs23/ec: Easy terminal-native 3-way git mergetool vim-like workflow GitHub - SethPyle376/hiraeth: Local AWS emulator focused on fast integration testing, with SQS support, SQLite-backed state, and a debug-friendly web UI. GitHub - JakOb-dotcom/cloud-sandbox-security-analysis: Technical analysis and Proof of Concept (PoC) regarding environment variable exfiltration in containerized cloud sandboxes via side-channel data leaks. Springboards - Flint Alpha Show HN: A simpler coding agent harness GitHub - audiodude/sudomake-friends GitHub - 256thFission/mini-mythos: OSS clone of Anthropic’s Mythos harness to locate C/C++ memory vulnerabilities Show HN: OpenParallax: OS-level privilege separation for AI agent execution Hacker News Sorted - Chrome 应用商店 Show HN: How to Install Docker on Ubuntu 24.04 LTS: Complete 2026 Guide GitHub - himanshudongre/smriti GitHub - sverrirsig/claude-control: macOS desktop dashboard for monitoring and managing multiple Claude Code sessions GitHub - ory/dockertest: Write better integration tests! Dockertest helps you boot up ephermal docker images for your Go tests with minimal work. Chiral - Chrome 应用商店 Show HN: Two Claudes collaborating through shared memory on a $100 mini-PC GitHub - pmichaillat/latex-cv: Minimalist LaTeX template for academic CVs GitHub - oguzbilgic/posse: A web UI for Anthropic Managed Agents. GitHub - sshiraz/depsly: Dependency risk analysis tool for npm packages ABI Add safari/agent-harness — Safari browser automation via safari-mcp by achiya-automation · Pull Request #212 · HKUDS/CLI-Anything GitHub - Halfblood-Prince/trustcheck: Verify PyPI package attestations and improve Python supply-chain security GitHub - oguzbilgic/kern-ai: Agents that do the work and show it. GitHub - bruits/satteri: High-performance Markdown and MDX processing for the JavaScript ecosystem GitHub - tylergibbs1/feedstock: High-performance web crawler and scraper for TypeScript, powered by Bun and Playwright GitHub - Grimm67123/grimmbot: The self-improving sandboxed and open-source AI agent. With persistent memory and scheduling. GitHub - whitevanillaskies/whitebloom: Local whiteboard that blooms. GitHub - hwdsl2/docker-whisper: Docker image for a self-hosted Whisper speech-to-text server with speaker diarization and OpenAI-compatible transcription and translation APIs. Powered by faster-whisper. Supports all Whisper models, NVIDIA GPU (CUDA) acceleration, JSON/SRT/VTT output, SSE streaming, offline mode, and multi-arch (amd64, arm64). GitHub - yisding/reviewwiggum GitHub - MarwanAlsoltany/serrors: Structured errors for Go: sentinel hierarchies, typed data, custom formatting, and slog integration. GitHub - soatok/age-php GitHub - Luthiraa/markitme GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits GitHub - tombedor/excalicharts GitHub - wh1le/excalidraw-edit: Open and edit .excalidraw files from the terminal. Offline, auto-saves to disk. MalExt Sentry - Malicious Extension Scanner - Chrome 应用商店 GitHub - syi0808/asciianimesvg: Generate animated ASCII art SVGs from text. CLI, Rust library, WASM, and web editor. GitHub - zaina-ml/ml_forge: A visual-based graph node editor for training computer vision models. GitHub - anakin87/llm-rl-environments-lil-course: 🌱 A little course on Reinforcement Learning Environments for evaluating and training Language Models GitHub - takaakit/superpowers-uml: Superpowers-UML modifies Superpowers to ensure a software development workflow in which AI agents design through UML modeling. AdriByte Studio - Sviluppo Web e Soluzioni Digitali GitHub - chouligi/angel-copilot: Your personalized Angel Investment Advisor Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 GitHub - agenteractai/lodmem: Level Of Detail Context Management for Agents GitHub - ostefani/subnetlens: A fast, concurrent network scanner with a TUI and plain-text CLI, built in Go. It discovers live hosts on your network, scans their open ports, resolves hostnames, and fingerprints operating systems—delivered. Cyber Pulse: Agentic Intel - Apps on Google Play Whisper API: Self-Hostable Speech to Text Transcription The Agent-Web Protocol Stack: A Research Thesis GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Show HN: Provepy – A Python decorator that proves your code using Lean and LLMs Show HN: Pardonned.com – A searchable database of US Pardons GitHub - patrickdappollonio/dux: Dux is a terminal UI that lets you run multiple AI coding agents side by side, each in its own git worktree, with full companion terminals, macros, commit generation, and a command palette that knows more tricks than you do. kMC Crystal Simulator Show HN: HyperFlow – A self-improving agent framework built on LangGraph GitHub - stef41/vibescore: 🎵 Grade your vibe-coded project. One command, instant letter grade across security, quality, dependencies, and testing. GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. imgur.com GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. GitHub - nowork-studio/toprank: Open-source Claude Code skills for SEO, SEM, Google Ads GitHub - tacomanator/sash: Lightweight macOS menu bar app for reliably cycling through windows of the current application. Appents | Social Media Management for Product-First Teams GitHub - pnhoang/youtube-spam-blocker: Automatically detects and hides spam messages in YouTube Live chat. Set rate limits, keyword filters, and block repeat offenders. GitHub - decisionnode/DecisionNode: CLI + Local MCP - A shared structured memory store across Claude Code, Cursor, Windsurf, Antigravity, and every MCP client. Semantically queryable. GitHub - AvaCodeSolutions/django-email-learning: An open source Django app for creating email-based learning platforms with IMAP integration and React frontend components. The $100K Gap in Kubernetes Security Tooling Function Calling Harness: From 6.75% to 100%
sparsemap
gregburd · 2026-05-19 · via Hacker News: Show HN

This is a C99 implementation of a sparse, compressed bitmap index. In the best case, it can store 2048 bits in just 8 bytes. In the worst case, it stores the 2048 bits uncompressed and requires an additional 8 bytes of overhead.

CI Pages License: MIT

A sparse, compressed bitmap library for C. Optimized for workloads with long runs of consecutive set or unset bits.

Why sparsemap

Bitmaps are great when bits are dense and the universe is small. They get expensive when either assumption fails — a 32-bit universe needs 512 MB just to hold one bit per integer, even if only a dozen are set.

Sparsemap stores only the chunks that contain actual data. In each chunk it picks one of two encodings depending on the local pattern:

  • Sparse encoding stores a 64-bit descriptor and only the bit vectors that contain a mix of set and unset bits. Uniform vectors (all-zero or all-one) take zero payload.
  • RLE encoding stores a single 64-bit descriptor for a contiguous run of set bits. A 2-billion-bit run takes 8 bytes.

Best case: 16 KB of consecutive set bits in 8 bytes. Worst case (random bits): identical to a raw bitmap plus 8 bytes of overhead.

When to use sparsemap

Good fit:

  • PostgreSQL extensions tracking TID sets, bitmap heap scans, posting lists.
  • Trigram / n-gram indexes where document identifiers cluster.
  • Allocation bitmaps for storage engines (free-list tracking).
  • Anywhere you'd reach for CRoaring but want a smaller, simpler library and don't need 32-bit integer support out of the box.

Not a fit:

  • Multi-threaded workloads: sparsemap is not thread-safe. A lock-free / wait-free variant is in design (see experiment/thread-safe).
  • 32-bit integer universes: sparsemap uses 64-bit indices.

Quick start

nix develop                              # optional dev shell
meson setup builddir
ninja -C builddir
ninja -C builddir test

Or with the Makefile wrapper:

make build
make test

Use it from C:

#include <sparsemap/sm.h>

sm_t *map = sm_create(4096);
sm_add(map, 42);
sm_add(map, 1024);
assert(sm_contains(map, 42));
assert(sm_cardinality(map) == 2);
sm_free(map);

Documentation

API docs (Doxygen) are published to gregburd.codeberg.page/sparsemap/.

Build options

meson setup builddir -Ddiagnostic=true   # enable __sm_assert + invariant checks
meson setup builddir -Db_sanitize=address  # ASan
meson setup builddir -Dbuildtype=release   # production: no asserts, max optimization

See meson_options.txt for the full list.

Consumers

Sparsemap is vendored by:

  • pg_tre — PostgreSQL trigram search extension.
  • postgres/undo — EnterpriseDB's PostgreSQL undo-log fork.

contrib/pg_tre_sync.sh and contrib/postgres_undo_sync.sh keep the vendored copies in sync with upstream.

Vendoring and symbol prefixing

The library is exactly two files, sm.h and sm.c; vendoring is a two-file copy. If you need two independently-vendored copies of sparsemap to coexist in one binary, rename every public symbol by defining SPARSEMAP_PREFIX before including the header:

#define SPARSEMAP_PREFIX myapp_
#include <sparsemap/sm.h>

myapp_sm_t *m = myapp_sm_create(4096);   /* renamed */
myapp_sm_add(m, 42);

Every public function and type picks up the prefix at both declaration and call sites (Berkeley DB --with-uniquename style). Compile-time macros (SM_IDX_MAX, the SM_VERSION_* values, enum constants) and the serialized wire format are unaffected.

Versioning and history

Releases follow SemVer. 3.0.0 is the first formal public release. The pre-3.0 development history (the library grew up vendored inside other projects) is preserved on the archive/v2.3.0 tag for archaeology; the published history starts clean at 3.0.0.

API stability vs ABI stability

Sparsemap promises source-level API stability within a major version: function signatures, macro names, and behavior of public sm_* symbols do not change in a way that breaks compiling consumer code.

Sparsemap does not promise ABI stability of the struct sparsemap layout. sizeof(sm_t) and the offsets of its fields may change in any minor release. Consumers must:

  • Always allocate sm_t via sm_create(), sm_create_with_allocator(), or sm_wrap() -- never embed it inline in another struct, never sizeof(sm_t) for an on-disk format, never memcpy(struct, ...) it.
  • Treat the type as opaque: access only via sm_* accessors.
  • Recompile (not just relink) after upgrading sparsemap.

The wire format produced by sm_serialize and consumed by sm_open/sm_deserialize is stable and is preserved across the 3.x series. This is the contract that matters for on-disk consumers.

Migrating from a pre-3.0 vendored copy

3.0.0 makes two source-level breaks, both mechanical:

  • The opaque type is now sm_t, not sparsemap_t. Migrate with sed -i 's/\\bsparsemap_t\\b/sm_t/g' your_files.c.
  • The vendoring prefix macro is SPARSEMAP_PREFIX, not SM_PREFIX. Rename it if you set it.

Everything else -- the sm_* function names, their signatures and behavior, and the serialized wire format -- is unchanged from the latest pre-3.0 vendored copies. See docs/MIGRATION.md for the full checklist.

Future work: SIMD

Sparsemap is scalar-only by design. No __builtin_popcount chains, no AVX intrinsics, no NEON — nothing target-specific. The same source compiles unchanged on x86_64, ARM, RISC-V, and anything else with a C99 compiler. This is deliberate: single-file vendoring and cross-platform reproducibility outrank per-architecture peak performance for our consumer profile (PostgreSQL extensions, embedded indexers, undo logs).

The aligned_alloc / aligned_free slots in sm_allocator_t exist so that adding SIMD later doesn't force another API break. Two tiers of work are plausible if a real workload ever justifies it. Both are deferred until a downstream consumer profiles a hotspot in a sparsemap operation.

Tier 1 — vectorize the inner loops without changing the wire format

  • sm_cardinality over MIXED runs. Walk chunks scalar-style to identify contiguous runs of MIXED bitvecs of length ≥ K (~4), gather them into an aligned scratch buffer, run AVX2/AVX-512 (or NEON) popcount, accumulate. Falls back to the current scalar loop for short runs and unsupported platforms.
  • Set ops on MIXED-MIXED chunk-pair runs. Same idea applied to sm_union / sm_intersection / sm_xor / sm_difference: when both inputs have aligned MIXED runs, dispatch to a vectorized vpand / vpor / vpxor loop.
  • Roughly 500 LOC of intrinsics, runtime CPU dispatch via __attribute__((target("avx2"))) plus a cpuid probe, and one aligned scratch buffer per inner-loop call (uses sm_allocator_t::aligned_alloc).
  • Realistic gain: 1.5–3× on dense (mostly-MIXED) maps; near zero on sparse maps because the gather overhead eats the win.

Tier 2 — wire-format extension for native SIMD layout

  • Add a fifth chunk payload type (e.g. SM_PAYLOAD_DENSE_RUN) that stores N contiguous bitvecs aligned on a 32-byte boundary, with a length prefix. The encoder switches to dense-run mode when emitting a long MIXED run.
  • The 2-bit flag space is full (00/01/10/11 all assigned), so the new mode requires an escape encoding via the chunk header.
  • Removes the gather step entirely; SIMD ops run directly on the serialized bytes.
  • Roughly 1500 LOC, codec rewrite, deserialize-backward-compat work, consumer wire format changes.
  • Realistic gain: 4–6× on dense maps.

Why neither is shipped today

Sparsemap's value proposition is "small wire format, single-file vendoring, no SIMD assumptions". Adding SIMD splits the code (scalar fallback + vector fast path), introduces runtime CPU dispatch, and forces every consumer's build system to handle target-feature flags. We will not pay that cost speculatively.

If and when a real workload pins sm_cardinality or set-op throughput as a measured bottleneck, Tier 1 is the right answer (small, contained, no wire-format change). Tier 2 is a CRoaring-shaped rewrite and probably the wrong tool for sparsemap's niche. Open an issue with profile data if you hit such a workload.

License

MIT. See LICENSE.