惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

PCI Perspectives
PCI Perspectives
P
Proofpoint News Feed
G
GRAHAM CLULEY
Know Your Adversary
Know Your Adversary
T
Tenable Blog
I
Intezer
Scott Helme
Scott Helme
Hacker News - Newest:
Hacker News - Newest: "LLM"
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Security Latest
Security Latest
SecWiki News
SecWiki News
Schneier on Security
Schneier on Security
T
The Exploit Database - CXSecurity.com
The Last Watchdog
The Last Watchdog
N
News and Events Feed by Topic
Google DeepMind News
Google DeepMind News
S
Secure Thoughts
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
H
Hacker News: Front Page
L
LINUX DO - 最新话题
U
Unit 42
宝玉的分享
宝玉的分享
Stack Overflow Blog
Stack Overflow Blog
L
LINUX DO - 热门话题
IT之家
IT之家
C
Cybersecurity and Infrastructure Security Agency CISA
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
量子位
博客园_首页
爱范儿
爱范儿
T
Threatpost
小众软件
小众软件
A
Arctic Wolf
Jina AI
Jina AI
H
Help Net Security
Webroot Blog
Webroot Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Cyberwarzone
Cyberwarzone
博客园 - 聂微东
月光博客
月光博客
The Register - Security
The Register - Security
Martin Fowler
Martin Fowler
博客园 - 司徒正美
V
V2EX
博客园 - Franky
B
Blog
B
Blog RSS Feed
H
Heimdal Security Blog
WordPress大学
WordPress大学
A
About on SuperTechFans

Show HN

GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - pileax-ai/pileax: PileaX is an all-in-one AI knowledge base system. 🍀 GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
GitHub - dannypesic/ril: A CLI tool that turns your Python scripts into a parallel data processing engine
dpesic · 2026-06-17 · via Show HN

Most of your CPU sits idle while one Python script grinds through a large file on a single core. ril spreads that work across all of them.

load data.csv | clean.py | featurize.py | save output.csv

ril is for slow, CPU bound Python that iterates through a dataset. It spawns worker processes that each run a Python interpreter and transform Arrow RecordBatches while streaming to each other. It automatically divides batches between workers and reassembles them in order. Since data is streamed, peak memory usage stays constant even as dataset size rises. It works on arbitrary Python scripts that take in a pyarrow.RecordBatch and return the processed one.

It is signficiantly faster than single core pandas and matches a multiprocessing.Pool (see Benchmarks), but memory streams by default and it doesn't require rewriting for a Pool or managing IO.

ril works well on large datasets and complex transforms. It fits work where fast and efficient output is prioritized over setting up infrastructure, such as simulations, research code, scientific computing, and one-off data jobs.

A rilfile:

load data.csv +2000
| clean.py
| tee cleaned_data.csv
| featurize.py
| score.py
| save output.csv

A stage:

from ril import rilfn
import pyarrow as pa

@rilfn
def process(batch):
    batch = pa.record_batch(batch)
    data = batch.to_pydict()
    data["score"] = [x * 2.0 for x in data["value"]]
    return pa.RecordBatch.from_pydict(data)

It works natively with pandas, numpy, and any other module that supports PyArrow.

Benchmarks

Applying a CPU bound Python function to every row of a one million row CSV:

variant time
single-core (pandas) 47.5 s
multiprocessing.Pool 11.8 s
ril (auto workers) 16.1 s
ril (x8, pinned) 11.9 s

With workers pinned, ril matches a manually tuned multiprocessing.Pool to within a few percent, and works without configuring a custom Pool. It works out to nearly 4x over single-core pandas on this machine, reaching the same ceiling as Pool. Peak memory also stays flat as the input grows, holding around 160MB at every size in this test while the pandas approaches climb past 700MB.

The gap between the two ril rows comes down to how you configure the pipeline; see Tuning. The full harness and the memory-scaling numbers are in benchmarks/.

Versus Other Frameworks

For DataFrame transforms that fit Polars or DuckDB's expression API, those engines will be faster and simpler. ril is effective for arbitrary Python that can't be purely experessed with relation logic. They aren't mutually exclusive: Polars is a natural choice inside a ril stage.

multiprocessing.Pool gives the same multi-core speedup (see Benchmarks), but requires more technical overhead and doesn't stream data. ril handles the parallelism and streams data through rather than buffering it.

Ray, Dask, and Spark are built for clusters. ril gives speedups with a single binary and no infrastructure.


Documentation

How it works

Each stage runs as its own process. Stages receive Arrow RecordBatches from the previous stage over a Unix pipe and forward results to the next. Stages run concurrently, and backpressure is automatic via the pipe buffer.

The @rilfn function is called once per chunk. Operations that require the full dataset (e.g. a global sort) do not belong inside a single stage.

Free-threaded Python

Free-threaded builds (3.13t, 3.14t) remove the GIL, but threads still share memory. Most of the data stack (numpy, pandas, and most ML libraries) isn't fully thread-safe, so you need to audit every dependency before using a threaded pipeline.

Since ril uses separate processes, each worker has its own interpreter and memory space, so none of that applies: your existing scripts work as they are. ril still works on free-threaded builds, giving extra headroom within each worker on top of the parallelism across them.

Writing a stage

Create a .py file with a function decorated with @rilfn:

from ril import rilfn
import pyarrow.compute as pc

@rilfn
def process(batch):
    # batch is a pyarrow RecordBatch
    mask = pc.greater(batch.column("value"), 0)
    return batch.filter(mask)

Place the file in your project directory and reference it by name in the rilfile. ril calls the decorated function once per chunk with a pyarrow.RecordBatch and expects a pyarrow.RecordBatch back.

Binary stages

Any executable that reads Arrow IPC from stdin and writes Arrow IPC to stdout works as a stage. Reference it by path:

load data.csv | ./transform | save output.csv

The binary receives batches one at a time and must flush stdout after writing each result. Worker count tags work the same as for Python stages.

Built-in stages

Stage Example Notes
load load data.csv streams in batches of 1000 rows by default
load load data.csv +500 custom batch size (rows per chunk)
save save output.csv terminal stage, writes final output
tee tee checkpoint.csv writes to file and passes batches through

Worker count

ril allocates workers automatically. It profiles the first few batches of each untagged script stage (dropping the slowest and fastest then taking the mean), then divides your CPU cores among the stages in proportion to how long each one takes, so a slower stage gets more workers. load, save, and tee are I/O-bound and always run as a single process.

To pin a stage to a fixed worker count, tag it with xN. A pinned stage gets exactly N workers and is left out of the automatic split; the remaining cores are shared among the untagged stages.

load data.csv | clean.py | featurize.py    | save output.csv   # both auto-allocated
load data.csv | clean.py | featurize.py x4 | save output.csv   # featurize pinned to 4, clean auto

Pinning all stages skips the profiling step, which removes the allocation startup on small datasets.

Tuning

Two settings decide how well a pipeline parallelizes.

Batch size is set on load, for example load data.csv +20000. ril splits each batch across the workers in a stage, so large batches parallelize coarsely while small batches spend more time on per-batch overhead rather than on your code. A few thousand to a few tens of thousands of rows is a reasonable place to start, which can be adjusted for your row width and per-row cost. This is the same kind of choice you would make for a multiprocessing.Pool chunksize.

The other setting is worker placement. Auto mode profiles the first five batches on each untagged stage on one worker, then scales up and runs as fast as a pinned pipeline from there. The profile is a one time cost that becomes negligible on larger datasets with reasonable batch sizes, and mostly becomes faster given the improved efficiency. It is recommended to use auto mode except on a dataset with few batches, as the overhead can cause a significant slowdown.

Failure semantics

ril is fail fast: an error in any stage aborts the whole pipeline and exits non-zero. It does not retry, skip, or resume for data integrity.

When your @rilfn raises, ril prints the stage and the batch index along with the full Python traceback:

ril error: stage 1 (featurize.py)

batch 3:
Traceback (most recent call last):
  File "featurize.py", line 10, in process
    raise ValueError("kaboom")
ValueError: kaboom

A few things to know before trying:

  • Output is not transactional. Stages stream, so by the time a later batch fails, save and tee may have already written every batch that came before it. ril does not roll those back, so treat an output file as valid only when the run exits 0.
  • The schema must be identical across batches. The output schema is fixed by the first batch a stage emits, and ril checks every later batch against it. A batch that returns different columns or types stops the stage with an error naming the batch and both schemas. Return the same columns and types from every call.
  • Only ordinary exceptions carry a traceback. A hard crash such as a segfault, the OOM killer, or os._exit can't report one, so you get a worker failed and a non-zero exit instead.
  • There is no timeout by default. A stage that hangs, whether on an infinite loop or a blocking call, hangs the pipeline.

Compatibility

ril connects to pip or uv in your project to manage the .venv:

  • pip: default
  • uv: used automatically if a uv.lock file is detected

On startup, ril automatically detects the newest compatible Python interpreter available on your PATH, trying python3.14, python3.13, python3.12, python3.11, then python3 as a fallback.

Feedback

Bug reports and feature requests are welcome. Open an issue on GitHub.

Building

Requires Rust and Python 3.11-3.14. On first run, ril automatically creates a .venv and ril.py in your project directory and installs pyarrow and arro3-core.

To build against a specific Python version, set PYO3_PYTHON before running cargo:

PYO3_PYTHON=python3.12 cargo build --release

PYO3_PYTHON controls which Python interpreter the binary links against at compile time. Pre-built release binaries are compiled per Python version, found in the Github releases page. Download the asset matching your Python install.