惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
D
DataBreaches.Net
Microsoft Azure Blog
Microsoft Azure Blog
大猫的无限游戏
大猫的无限游戏
雷峰网
雷峰网
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
P
Privacy & Cybersecurity Law Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
L
LINUX DO - 最新话题
L
LangChain Blog
量子位
P
Palo Alto Networks Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
L
Lohrmann on Cybersecurity
博客园 - 聂微东
人人都是产品经理
人人都是产品经理
Y
Y Combinator Blog
MongoDB | Blog
MongoDB | Blog
PCI Perspectives
PCI Perspectives
S
SegmentFault 最新的问题
O
OpenAI News
S
Securelist
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
C
Cybersecurity and Infrastructure Security Agency CISA
AWS News Blog
AWS News Blog
G
Google Developers Blog
博客园 - 叶小钗
C
Check Point Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Cisco Talos Blog
Cisco Talos Blog
罗磊的独立博客
V2EX - 技术
V2EX - 技术
小众软件
小众软件
IT之家
IT之家
Engineering at Meta
Engineering at Meta
Hacker News - Newest:
Hacker News - Newest: "LLM"
M
MIT News - Artificial intelligence
T
Threat Research - Cisco Blogs
Vercel News
Vercel News
酷 壳 – CoolShell
酷 壳 – CoolShell
A
About on SuperTechFans
Recorded Future
Recorded Future
N
News and Events Feed by Topic
Cloudbric
Cloudbric
W
WeLiveSecurity
T
The Exploit Database - CXSecurity.com
Martin Fowler
Martin Fowler
C
CXSECURITY Database RSS Feed - CXSecurity.com
The Cloudflare Blog
宝玉的分享
宝玉的分享

Show HN

GitHub - villagesql/villagesql-skills: Agent skills for VillageSQL - gemini-cli-extension; claude-code-plugin GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
Open-Source Low-Overhead NVIDIA CUDA PC Sampling | Polar Signals
Tommy Reilly · 2026-06-15 · via Show HN

One of the more powerful features of the CUDA Profiling Tools Interface (CUPTI) is support for Program Counter (PC) sampling. This lets developers of CUDA programs see where their code is spending time down to the instruction level. Building upon our support for kernel timing information, we've added the ability for our low-overhead continuous profiler to send PC sample data to our backend where it can be analyzed using the Polar Signals UI or run through your favorite LLM model using our MCP support. PC sampling is typically used in developer-oriented workflows with tools like NVidia NSight and Triton's Proton profiler, but with our approach to minimizing overhead it's actually possible to use it in production! But first, what is PC sampling and how does it work?

If you're just interested in trying this, this work is available in the v0.48.0 release of the open source Parca Agent. Check out this blog post on how to try it out on Kubernetes.

PC Sampling was introduced in the Maxwell architecture and had a simple API that piggy-backed on the CUPTI activity API. When the Volta architecture came out a new dedicated API for PC sampling was introduced. PC sampling works via dedicated hardware per-warp, on every sampling tick the state of each warp is recorded. The sampling interval is based on a power of 2 sampling factor where the hardware samples every 2^SAMPLING_FACTOR GPU cycles. The sampling factor is constrained to the range 5 to 31, so we could be taking a sample every 32 cycles (2^5) or only once every couple billion (2^31), tens of millions of samples a second at one extreme, barely one a second at the other. Quite the range! For our purposes we found 20 to be a good default sampling factor but of course it's configurable. That's a raw hardware rate of over 2k samples/s which may seem high for a "low-overhead" continuous profiler but depending on your application and GPU utilization the amount of data can be much less. That's because PC sampling records a PC offset/stall reason "bucket" and every time a sample is taken it's just incrementing the counter on that bucket. So no call stacks, no timestamps, just a pc offset / stall-reason pair.

Basically the pc/stall-reason information is collected in hardware, counters go up, and then periodically this information is flushed from the hardware buffers out to software buffers (all low-level stuff handled by CUPTI and the driver). We can configure this buffer size but the real trick is getting this information out of the buffer and into our CUPTI shim library. If you didn't commit the information from our prior blog posts to memory the way it works is your CUDA application is run with an env var (CUDA_INJECTION64_PATH) set to our shim library and it handles initializing CUPTI and listening for information about GPU happenings.

For PC sampling it works like this:

PC Sampling Overview

The real power of PC sampling is that it's not just recording the PC, it's the stall reason that comes with it, meaning that if the instruction at that cycle is (or is not) being issued the reason will be recorded. It's like if you had a CPU profiler that told you, down to the instruction level, if the instruction was retired normally or it had to wait on a pipeline stall, a cache miss or coherency delays. In GPU land there are a multitude of stall reasons but the main ones are "long scoreboard" memory latency dependencies (waiting on loads) or "short scoreboard" latency waiting on shared memory or specialized functional unit results. But there can also be queuing stalls (waiting for busy functional units to open up), synchronization barriers/memory fences etc. Similar to a CPU just a bigger menu of options.

The Polar Signals profiler takes the guess work out of understanding these stall reasons by including a brief explanation and linking to NVIDIA documentation for deeper understanding (see screenshot below).

A GB10 chip like the DGX Spark has 48 streaming multiprocessors (SMs) with 48 warps/SM which means it's sampling 2304 warps in parallel! Many levers are needed to deal with that much information! We've already talked about the sampling factor and how data is rolled up into PC/stall-reason pairs. Another lever we have is the PC sampling collection mode. PC sampling can be done in a "continuous" mode or a "kernel-serialized" mode. You'd think a continuous profiler would want to use continuous mode, but continuous mode makes attribution of samples to a particular kernel launch impossible: a GPU handles multiple kernel launches in parallel and schedules them as aggressively as it can, so you could end up with two kernel launches that share some of the same underlying CUDA binaries (cubins) both contributing to the same pc/stall-reason pairs. To take the guess work out of this we use kernel-serialized mode, which as you probably surmised, kills performance hard. So how do we make a continuous production mode profiler when we're killing performance? By sampling the samples!

Turns out you can enable/disable PC sampling pretty quickly so we have a dynamic algorithm that periodically turns on PC sampling for short intervals (~50ms) and then turns it off again where the delay between is tuned to get a target number of PC/stall reason pairs per second. By default we target 100/s (again also configurable) which seems to work well in practice. For simple GPU workloads there may be very little time between intervals, and for intense PyTorch training workloads there may be many seconds between sampling intervals.

So we know how to get the data off the hardware and into our shim library, but how do we efficiently get it over the network back to the collection service? Simple, we piggy back on our existing work and utilize USDT probes to allow our agent to place hooks into the shim library and extract all the goods.

Here are the new probes we added to support PC sampling:

And here's how everything is wired together:

There's a subtle problem hiding in that probe table. A pc_sample_batch record is useless on its own, to make sense of it you need the stall_reason_map (to decode the stall indices), the cubin_loaded bytes (to turn a PC offset back into source), and the gpu_config (to convert samples to time). But those are one-shot events. The stall reason map is emitted once at startup, and cubins fire as they're loaded, often long before anyone is watching.

Remember there's no coordination between the shim and the agent. The shim just fires these probes and does very little processing on the information other than to batch things up conveniently (you wouldn't want a BPF probe to fire on every single PC sample). This is great from a division-of-labor perspective, but it makes things tricky. What if the agent attaches mid-workload?

Typically the Polar Signals agent will be running 24/7 like a fly on the wall but customers will sometimes change label configs and restart for upgrades, and it's also the case that some of these PyTorch training workloads can run for a very long time (as we saw in my last blog). So we think it's important to handle this case.

Without going into excruciating detail basically what we do is listen and record the USDT semaphore counts on our probes which allows us to know when clients attach/detach from the probes and when this occurs we re-send cached stall reason maps and CUBIN information.

The agent installs BPF programs on all of these probes, feeding the stream of events into a BPF ring buffer for processing. The one real trick it has to play is caching kernel launches: PC samples show up after the fact tagged only with a correlation ID, so the agent keeps a cache of recent launches and matches each batch of samples back to the application stack that launched that kernel. We take the PC and stall reason and attach them as "labels" on the stack sample that can be grouped or filtered:

screen shot of PC sampling flamegraph

From there the samples are packed into Apache Arrow records for efficient transmission to the server. One nice property that falls out of how PC sampling works: because the data is already reported as counts, each pc/stall-reason bucket carries the number of times it was sampled, meaning there's nothing to deduplicate. The agent and backend can simply add buckets together as they arrive.

The last piece is symbolization. Rather than resolve instructions to source on the host, which would mean burning cycles inside the profiled process, the agent uploads each cubin to our debuginfo service and we symbolize on the backend. The wrinkle is that cubins don't ship with standard DWARF debug info mapping SASS instructions back to source lines, so we can't lean on the usual symbolization tooling. Instead we crack open the cubin, disassemble it, and build our own address-to-source tables to turn a pc offset back into a function, file, and line. Just be sure to include the -lineinfo flag on your nvcc command line (seen screenshot above for an example).

PC sampling has traditionally been confined to interactive, developer-time tools like NSight and Proton because of its overhead. By sampling the samples, replaying metadata so a late-attaching agent never misses the context it needs, and pushing symbolization to the backend, we've gotten it down to something you can leave running in production. The result is instruction-level GPU insight, right down to why a warp stalled, alongside the call stacks and kernel timings you already get from the Polar Signals continuous profiler.

You can get started with our free 14-Day trial today. If you're on Kubernetes, check out how to get started within just a few minutes without modifying your workload.