惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

大猫的无限游戏
大猫的无限游戏
S
SegmentFault 最新的问题
量子位
A
Arctic Wolf
L
Lohrmann on Cybersecurity
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
WordPress大学
WordPress大学
V
Vulnerabilities – Threatpost
博客园 - Franky
C
Cyber Attacks, Cyber Crime and Cyber Security
The Cloudflare Blog
Last Week in AI
Last Week in AI
The Hacker News
The Hacker News
I
Intezer
J
Java Code Geeks
P
Privacy International News Feed
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
S
Secure Thoughts
Cisco Talos Blog
Cisco Talos Blog
阮一峰的网络日志
阮一峰的网络日志
S
Securelist
Security Latest
Security Latest
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
小众软件
小众软件
Jina AI
Jina AI
有赞技术团队
有赞技术团队
人人都是产品经理
人人都是产品经理
博客园_首页
酷 壳 – CoolShell
酷 壳 – CoolShell
T
The Exploit Database - CXSecurity.com
雷峰网
雷峰网
T
Tenable Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
P
Privacy & Cybersecurity Law Blog
Simon Willison's Weblog
Simon Willison's Weblog
博客园 - 【当耐特】
T
Threat Research - Cisco Blogs
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
MongoDB | Blog
MongoDB | Blog
D
DataBreaches.Net
N
News | PayPal Newsroom
Google Online Security Blog
Google Online Security Blog
K
Kaspersky official blog
H
Help Net Security
宝玉的分享
宝玉的分享
罗磊的独立博客
Webroot Blog
Webroot Blog
月光博客
月光博客
B
Blog RSS Feed
Recorded Future
Recorded Future

Hacker News - Newest: "LLM"

GitHub - lechmazur/position_bias: A benchmark for testing whether LLM judges keep the same preference when two lightly edited versions of the same story are shown in opposite orders. Flex routing (EU and EFTA) Dark Factories: Retooling for LLM Velocity Ask HN: What would be the impact of a LLM output injection attack? GitHub - AronDaron/dataset-generator: No-code desktop app for generating high-quality synthetic datasets to fine-tune LLMs — plan-then-execute pipeline, LLM-as-judge, HuggingFace upload. GitHub - Oaklight/llm-rosetta: Production-ready LLM API translation layer for Python — bidirectional conversion between OpenAI, Anthropic & Google formats via hub-and-spoke IR. Optional API gateway. Streaming & non-streaming. Zero core deps. Contributions welcome! GitHub - browser-use/browser-harness: Self-healing browser harness that enables LLMs to complete any task. GitHub - moeen-mahmud/remen: Remen turns thoughts into something you can return to Analyzing 156 LLM Launch Posts on Hacker News ChatGPT vs Gemini vs Claude: The Best LLM Subscription You Should Buy GitHub - salaamalykum/quran-semantic-search: High-density RAG Semantic Search Engine & Quran Corpus (GEO/SEO Architecture) GitHub - NVIDIA/TensorRT-LLM: TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. The State of LLM Bug Bounties in 2026 Operational Readiness Criteria for Tool-Using LLM Agents Meshcore: Architecture for a Decentralized P2P LLM Inference Network How an LLM becomes more coherent as we train it GitHub - seetrex-ai/laimark GitHub - Jossifresben/BibCrit: AI-assited biblical textual criticism GitHub - wastedcode/memex: File system based wiki, maintained by Claude 99helpers.com GitHub - cliver-project/AITrigram GitHub - unbody-io/adapt: A self-evolving memory layer for AI agents. GitHub - hb20007/awesome-gen-ai-fails: A list of incidents where reliance on generative AI and LLMs resulted in harm to companies, individuals, or society GitHub - nevenkordic/localmind: Run any local LLM with persistent memory and context. CLI agent over Ollama with SQLite-backed hybrid recall. No cloud. Ask HN: What are the machine requirements for a LLM like Llama-3.1-8B? Faster LLM Inference via Sequential Monte Carlo grpo explained: group relative policy optimization for llm finetuning - cgft Stop comparing price per million tokens: the hidden LLM API costs · TensorZero Andrej Karpathy's LLM Wiki Is a Bad Idea GitHub - GG-QandV/mnemostroma: Offline RAM-first cognitive leer/coprocessor for AI agents and robotics. Solves "Context Abandonment" with 20-80ms latency using a dual-thread biomimetic memory architecture (ONNX + SQLite WAL). mempalace/agent at agent · skorotkiewicz/mempalace GitHub - Nyquest-ai/nyquest-rust-fullstack-pub: Nyquest — Semantic Compression Proxy for LLMs. 350+ rules, local LLM stage, 15-75% token savings. Full Rust stack. GitHub - TheoV823/mneme: Enforce architectural decisions in AI-assisted development. GitHub - klemenvod/TokenBrawl: A 1v1 Bomberman-style game where two LLM agents play autonomously against each other. No human plays — you watch the AIs fight. Each agent receives a text description of the board state, reasons about it, and outputs a move as JSON. The game engine executes it. Introducing the Common AI Provider: LLM and AI Agent Support for Apache Airflow Power Circuit AI: Designing Power Electronic Circuits for Motor Drives with Generative Artificial Intelligence Ask HN: How to program with IDE and LLM on CPU locally? Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis Bonsai 1-bit WebGPU - a Hugging Face Space by webml-community The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows Ask HN: Simple tooling for local LLM code critique without IDE integration? Can a General LLM Diagnose a DICOM Slice? A 10-Case Public Benchmark Charts-of-Thought: Enhancing LLM Visualization Literacy (PDF, 2026) GitHub - Mesh-LLM/mesh-llm: Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. GitHub - seamus-brady/springdrift: A persistent runtime for long-lived LLM agents Writing an LLM from scratch, part 32k -- Interventions: training a better model locally with gradient accumulation Ask HN: Which LLM model and agentic CLI are you using for local development? GitHub - wayneColt/modelcascade: Route local. Escalate smart. Never overspend. Open-source multi-model cascade routing for autonomous agents. LLM pricing is 100x harder than you think GitHub - asakin/llm-primer: Pre-warmed Claude Code sessions in tmux. No startup wait. GitHub - EggerMarc/chat-rs: A multi-provider LLM framework for Rust. GitHub - SynapseKit/SynapseKit: Minimal, async-first Python framework for production LLM apps- 2 hard deps, no magic, no SaaS. A Claude Skill that Makes LLM Paragraphs More Bearable Does Gas Town 'steal' usage from users' LLM credits & paid services to improve itself? What's Claude Code Actually Doing? Open the Black Box with the Arthur Engine Milla Jovovich's New Open Source LLM Memory App and the Dark Code Problem Your intuition of LLM token usage might be wrong Show HN: Bloomberg Terminal for LLM ops – free and open source GitHub - 0xchamin/mcptube: Transform YouTube videos into a compounding knowledge base with transcripts, vision analysis, and agentic search. Works as an MCP server for Claude, Copilot & more. Show HN: Open KB: Open LLM Knowledge Base Your LLM is a compiler, not a runtime GitHub - sapountzis/Unslop: A Web Feed That Deserves You crates.io: Rust Package Registry Beyond Karpathy's LLM-Wiki: The Necessity of Cognitive Governance GitHub - amitshekhariitbhu/llm-internals: Learn LLM internals step by step - from tokenization to attention to inference optimization. GitHub - parallem-ai/parallem: An expressive library for running agents with the Batch API. GitHub - stfurkan/pi-llm LLM-Wiki Show HN: Formal – Formal verification for AI-generated code using Lean 4 LRTS – Regression testing for LLM prompts (open source, local-first) LLM Wiki Skill: Build a Second Brain with Claude Code and Obsidian I built an LLM Wiki and RAG solution: here's a demo for a security KB The biggest advance in AI since the LLM Predict-Rlm: The LLM Runtime That Lets Models Write Their Own Control Flow the-synthetic-library/the-synthetic-mind at main · joshferrer1/the-synthetic-library GitHub - yisding/reviewwiggum GitHub - Donnyb369/mcp-spine: Context Minifier & State Guard — Local-first MCP middleware proxy GitHub - Beledarian/wgpu-llm: A from-scratch LLM inference engine that uses wgpu (the cross-platform WebGPU implementation) to dispatch WGSL compute shaders for every math operation a Transformer needs. No CUDA. No Python. No massive framework dependencies. Just Rust, raw shaders, and your GPU. GitHub - anitiue/Hindsight: An experience-driven self-improvement framework for LLM agents — 基于经验的 LLM Agent 自我改进框架 GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. GitHub - alainnothere/AmdPerformanceTesting: Amd Performance Testing Ask HN: Is a purely Markdown-based CRM a terrible idea? Optimized for LLM agents Context Engineering - LLM Memory and Retrieval for AI Agents | Weaviate little_helper_tui/letter.md at main · sleepyeldrazi/little_helper_tui GitHub - EvanZhouDev/umr: The Unified Model Registry for all your local AI apps. GitHub - JordanCT/VigIA-Orchestrator Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain A Taxonomy of RL Environments for LLM Agents Llama LLM Network Feture GitHub - genedeng-ca/ai-mac-migration: AI-powered Mac-to-Mac migration tool - replace Apple Migration Assistant with intelligent, selective transfer using local LLMs GitHub - lunargate-ai/gateway: High-performance self-hosted AI gateway (OpenAI-compatible) with routing, retries, and streaming GitHub - AuthBits/webmcp: A lightweight, prompt-driven MCP web research server for high-quality LLM powered information extraction. Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Springdrift: An Auditable Persistent Runtime for LLM Agents with Case-Based Memory, Normative Safety, and Ambient Self-Perception High-Stakes Personalization: Rethinking LLM Customization for Individual Investor Decision-Making From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
Distributed LLM Inference with llm-d
Moncef Abboud · 2026-06-27 · via Hacker News - Newest: "LLM"

Intro

What does production-grade LLM inference look like?

If you wake up in the middle of the night thinking about that question, this blog post might be for you.

Inference engines like vLLM and SGLang sit on top of PyTorch (when I say vLLM going forward, that also includes SGLang and other supported inference engines). They optimize inference at the node level, most notably by managing KV cache and improving throughput through techniques such as paged attention and continuous batching.

The elevator pitch for llm-d is that it’s an LLM-aware load balancer.

If we have multiple vLLM instances, each has its own state: available GPU memory, KV cache prefix matches, number of requests queuing to be processed, etc. We can’t simply use round robin. Selecting the best inference engine instance based on these signals is what llm-d is all about.

It also provides features such as flow control to support different classes of requests based on priority (e.g., premium real-time traffic vs batch workloads). In addition, it enables smooth disaggregated P/D, where prefix and decode run on different nodes because they benefit from different configs. Prefix is compute-bound, while decode is memory-bandwidth-bound, so they have different GPU requirements: low TP for prefix to maximize compute, and high TP for decode to maximize memory bandwidth.

It also provides features such as flow control to support different classes of requests based on priority (e.g., premium real-time traffic vs batch workloads). In addition, it enables smooth disaggregated P/D, where prefix and decode run on different nodes because they benefit from different configs. Prefix is compute-bound, while decode is memory-bandwidth-bound, so they have different GPU requirements: low TP for prefix to maximize compute, and high TP for decode to maximize memory bandwidth.

If It Ain’t Broke

The cool thing about llm-d is that it doesn’t reinvent the wheel. It builds on top of existing, established projects. LLM inference still happens on vLLM and SGLang, and we simply communicate with them via HTTP. The proxy layer and discovery are all built on top of Kubernetes(k8s) and Envoy. Even the integration point for deciding which vLLM instance to choose relies on an existing extension point, namely Envoy’s ext_proc extension. The data layer and metric collection are essentially built on top of Prometheus. So, no reinventing the wheel, just intelligently combining solid existing solutions. The best part is that for each layer and piece (metrics, scoring, flow control, etc.), llm-d is easily extensible with clear interfaces that new plugins can implement and use right away.

So if vLLM adds a new metric, or even if a new inference engine comes along, it can be easily added. If we want to implement a new way to pick or score, we just implement a new picker or scorer. If Prometheus falls out of favor, or if we want to use a custom monitoring solution, we can implement a plugin for that instead.

There’s a standardization effort to consolidate Generative AI inference on top of k8s, taking the form of the Gateway API Inference Extension, or GAIE. GAIE defines the API resources (like InferencePool) and the Endpoint Picker role. llm-d’s router is an implementation of that EPP role, paired with an Envoy proxy that does the actual request forwarding.

That’s a very strategic choice. Rather than having llm-d be an isolated initiative with a bespoke API, it’s built as an implementation of a broader k8s standard.

It’s also worth mentioning that llm-d has a mode where it runs outside of k8s via a file discovery plugin, where the vLLM endpoints are hardcoded instead of discovered via the k8s API.

llm-d architecture

So what are these factors that need to be taken into account for LLM inference routing?

Prefix Cache

Say we have N1 and N2. If a user has already gotten a response from N1, this means the user’s request KV cache has already been calculated there, so a subsequent request can skip that step. If, however, we send the user’s request to N2, we need to re-run prefill and we won’t be taking advantage of the calculation already done on N1.

In other words, it’s efficient to route requests to nodes that already have the KV cache of the prompt (this consideration changes if we have KV cache offloading, in which the cache can be stored in shared network storage, for instance).

KV-cache Utilization

Another factor we care about is how much free VRAM N1 and N2 have. If N1’s VRAM is almost full, even though it has the user’s KV cache, it might not be the best choice, because we might need to wait for other requests to finish or evict them in order to run the request. If at the same time N2’s VRAM is free and ready to go, it might be best to rerun the prefill and make use of the free memory.

Queue Depth

Both nodes might be running with little free VRAM, but N1 has only 2 requests in its queue waiting to be processed while N2 has 5. It’s best to route to the node with the smaller number of waiting requests.

There are many factors to consider when deciding which node to pick, and llm-d’s job is to pick. It gathers data from the various vLLM instances via the /metrics Prometheus endpoint, and keeps the state of each node in memory. When a user request comes in, it’s sent to llm-d’s router, called the Endpoint Picker or EPP, and the EPP will run a score of scorers (yes, the silly pun is intended):prefix-cache-scorer, kv-cache-utilization-scorer, queue-scorer, etc. for each possible candidate (vLLM instance), combining those scores and then picking one of the candidates. Before scoring and picking, there’s a filtering phase in which we can filter the candidate endpoints.

Envoy and the External Processing Filter

llm-d’s router doesn’t replace Envoy, it rides alongside it as an Envoy proxy filter target. When a request comes in, the Envoy proxy consults the llm-d EPP. The EPP does its thing and in the HTTP response it provides a header:

1
x-gateway-destination-endpoint: IP_PORT_OF_A_VLLM_POD

Then Envoy will route the request to that vLLM pod. That’s it, really.

Envoy Proxy has an extension (HTTP filter) called ext_proc, or External Processing. When a request comes in, it’s forwarded to an external process via gRPC, and that process can do a lot of things: mutate the HTTP body, add or remove headers, reject the request, etc. In our case, the llm-d router EPP is the external process, and it simply adds a header, the one with the address of the chosen vLLM instance.

How we make that choice is the magic that llm-d brings. And as we said above, it gathers data from the inference engines’ Prometheus endpoints:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
var defaultEngineConfigs = []engineConfigParams{
	{
		Name:                "vllm",
		QueuedRequestsSpec:  "vllm:num_requests_waiting",
		RunningRequestsSpec: "vllm:num_requests_running",
		KVUsageSpec:         "vllm:kv_cache_usage_perc",
		LoRASpec:            "vllm:lora_requests_info",
		CacheInfoSpec:       "vllm:cache_config_info",
	},
	{
		Name:                    "sglang",
		QueuedRequestsSpec:      "sglang:num_queue_reqs",
		RunningRequestsSpec:     "sglang:num_running_reqs",
		KVUsageSpec:             "sglang:token_usage",
		LoRASpec:                "",
		CacheInfoSpec:           "sglang:cache_config_info",
		CacheBlockSizeLabelName: "page_size",
		CacheNumBlocksLabelName: "num_pages",
	},
    //...
    }

And based on this data, it calculates a score for each candidate endpoint and picks one. Usually that’s the one with the highest score, but the picker can also sample according to score instead of just taking the max, which is a nice detail: always routing to the single best pod can create a thundering herd, where everyone piles onto the same “good” replica until it stops being good. Spreading the load based on score avoids that while still favoring the better candidates, this has a nice parallel to sampling a token after the softmax layer in an LLM.

Envoy establishes a gRPC connection with the EPP and streams it the request headers and body. Essentially, it’s:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
func (s *StreamingServer) Process(
    srv extProcPb.ExternalProcessor_ProcessServer,
) error {

    // Shared state for the lifetime of the request
    reqCtx := NewRequestContext()

    // Stream reader
    recvCh := startReceiver(srv)

    for {
        select {

        case msg := <-recvCh:

            switch v := msg.Request.(type) {

            case *RequestHeaders:
                handleRequestHeaders(reqCtx, v)

            case *RequestBody:
                handleRequestBody(reqCtx, v) // Actually: s.director.HandleRequest(ctx, reqCtx, parseResult.Body)

            case *ResponseHeaders:
                handleResponseHeaders(reqCtx, v)

            case *ResponseBody:
                handleResponseBody(reqCtx, v)

            }
        }

    }
}

When we finish receiving the request body, we run through s.director.HandleRequest. The director holds the datastore that has the endpoints and their metrics.

We run the request through some admission plugins (so it can be rejected if needed, this is extensible like most of the pipeline, and custom rules can be added at different points to handle that).

We then go through flow control, where requests are treated based on their priority/flow id, specified via headers (x-gateway-inference-fairness-id, x-gateway-inference-objective). Flows can have different priorities, and within the same priority band we can follow fairness and ordering policies (round robin, first come first served, earliest deadline first, etc).

Once the request makes it through flow control, the endpoint candidates are grabbed and we run the “data producer” plugins for that request. These plugins augment the request with new data: tokenize the string by adding token ids (there’s a tokenizer that relies on a vLLM API), and another key producer is the prefix matcher: how many of this request’s tokens already exist in the KV cache of each candidate pod, etc.

Then comes the essential part: the scheduler. We have a scheduling profile that has filters, scorers, and a picker. We filter the candidates, then score the request according to each scorer (of which there are many: prefix cache scorer, KV cache utilization, queue depth, etc.), then pick based on a combined score (each scorer can have its own weight).

Scheduling:

1
2
3
4
5
6
7
8
9
10
Score:
Prefix cache scorer (weight=3.0): Pod 1 = 0.75, Pod 2 = 0.4, others = 0.0
KV-cache utilization (weight=1.0): Pod 1 = 0.6, Pod 2 = 0.8, others vary
Queue depth (weight=1.0): Pod 1 = 0.7, Pod 2 = 0.5, others vary
Final scores:
Pod 1: (0.75 × 3) + (0.6 × 1) + (0.7 × 1) = 3.55
Pod 2: (0.4 × 3) + (0.8 × 1) + (0.5 × 1) = 2.5
Others: < 2.0
Picker (weighted random): Pod 1 is strongly favored but not deterministic
Result: Pod 1 selected (IP: 10.0.1.42, port: 8000).

The EPP sends a response to Envoy with x-gateway-destination-endpoint set to the IP and port of the selected pod, and Envoy forwards the request there. As the pod generates tokens, Envoy streams the response back to the client. The ext_proc stream also receives response headers and body chunks for metric recording and data updates (new KV cache being calculated, etc).

And that’s about it.

KV Events

KV Events

If the model server (vLLM) already has part of the prompt’s KV cache, routing the request there avoids recomputing the prefix. We need to know about the content of each node’s KV cache.

There is an approximate prefix cache producer estimates each node’s KV cache from past routing decisions: if a prompt was previously sent to vLLM 1, it’s likely that its KV cache is still there.

The downside is that it’s only an approximation and doesn’t account for cache evictions.

There is a precise KV cache producer takes a much more interesting approach. vLLM exposes KV-cache allocation and eviction events over ZMQ (a brokerless high-performance messaging library), and the llm-d router subscribes to these streams from every vLLM instance. This gives the router a real-time, exact view of each server’s KV cache, allowing it to make precise routing decisions at the cost of a bit of additional overhead.

This relies on a library hosted in its own repo: llm-d-kv-cache.

Disaggregated Serving

PD Request Flows

LLM inference has two phases:

  • Prefill: runs first, computing the KV cache for all prompt tokens. it’s compute-bound and performs best with low Tensor Parallelism (TP).
  • Decode: runs after prefill, autoregressively predicting the next token from the existing KV cache. Computation is minimal (only the last token), and we’re constantly moving memory from HBM to compute units. It’s memory-bandwidth bound and works best with larger TP.

Beyond their different TP needs, a large prefill phase can hog compute resources from smaller decode requests. Separating the two stages onto different nodes avoids this and improves performance overall.

In practice: a request first hits prefill pods for initial KV cache computation, then that KV cache is transferred to decode pods, where autoregressive prediction takes place.

RDMA (remote direct memory access) moves that KV cache from prefill to decode nodes very fast, bypassing the CPU entirely. The underlying networking is InfiniBand or RoCE. llm-d uses NVIDIA’s Inference Transfer Library (NIXL) for this transfer. Fascinating stuff.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
┌─────────────────────────────┐              ┌─────────────────────────────┐
│          NODE 1             │              │          NODE 2             │
│                             │              │                             │
│  ┌───────────────────────┐  │              │  ┌───────────────────────┐  │
│  │     Prefill Pod       │  │              │  │     Decode Pod        │  │
│  │                       │  │              │  │                       │  │
│  │  ┌─────────────────┐  │  │              │  │  ┌─────────────────┐  │  │
│  │  │ KV Cache (VRAM) │  │  │              │  │  │ KV Cache (VRAM) │  │  │
│  │  └────────┬────────┘  │  │              │  │  └────────┬────────┘  │  │
│  │           │           │  │              │  │           │           │  │
│  └───────────┼───────────┘  │              │  └───────────┼───────────┘  │
│              │              │              │              │              │
│              ▼              │              │              ▼              │
│  ┌───────────────────────┐  │   Network    │  ┌───────────────────────┐  │
│  │         NIC           │  │              │  │         NIC           │  │
│  │  InfiniBand / RoCE    │──┼──────────────┼──│  InfiniBand / RoCE    │  │
│  └───────────────────────┘  │              │  └───────────────────────┘  │
│                             │              │                             │
└─────────────────────────────┘              └─────────────────────────────┘