惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

The Register - Security
The Register - Security
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
MyScale Blog
MyScale Blog
V
Visual Studio Blog
云风的 BLOG
云风的 BLOG
aimingoo的专栏
aimingoo的专栏
C
Check Point Blog
J
Java Code Geeks
大猫的无限游戏
大猫的无限游戏
L
LangChain Blog
Vercel News
Vercel News
阮一峰的网络日志
阮一峰的网络日志
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
S
Security @ Cisco Blogs
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
人人都是产品经理
人人都是产品经理
H
Hacker News: Front Page
L
Lohrmann on Cybersecurity
T
Troy Hunt's Blog
T
Threat Research - Cisco Blogs
A
About on SuperTechFans
T
Threatpost
AWS News Blog
AWS News Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
T
Tor Project blog
Google Online Security Blog
Google Online Security Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Tenable Blog
W
WeLiveSecurity
博客园 - 叶小钗
K
Kaspersky official blog
Y
Y Combinator Blog
T
The Blog of Author Tim Ferriss
Hugging Face - Blog
Hugging Face - Blog
M
MIT News - Artificial intelligence
Hacker News - Newest:
Hacker News - Newest: "LLM"
Engineering at Meta
Engineering at Meta
有赞技术团队
有赞技术团队
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
S
Secure Thoughts
小众软件
小众软件
D
Docker
爱范儿
爱范儿
C
Cyber Attacks, Cyber Crime and Cyber Security
N
News and Events Feed by Topic
S
Schneier on Security
博客园 - 三生石上(FineUI控件)
D
DataBreaches.Net

Hacker News - Newest: "LLM"

GitHub - lechmazur/position_bias: A benchmark for testing whether LLM judges keep the same preference when two lightly edited versions of the same story are shown in opposite orders. Flex routing (EU and EFTA) Dark Factories: Retooling for LLM Velocity Ask HN: What would be the impact of a LLM output injection attack? GitHub - AronDaron/dataset-generator: No-code desktop app for generating high-quality synthetic datasets to fine-tune LLMs — plan-then-execute pipeline, LLM-as-judge, HuggingFace upload. GitHub - Oaklight/llm-rosetta: Production-ready LLM API translation layer for Python — bidirectional conversion between OpenAI, Anthropic & Google formats via hub-and-spoke IR. Optional API gateway. Streaming & non-streaming. Zero core deps. Contributions welcome! GitHub - browser-use/browser-harness: Self-healing browser harness that enables LLMs to complete any task. GitHub - moeen-mahmud/remen: Remen turns thoughts into something you can return to Analyzing 156 LLM Launch Posts on Hacker News ChatGPT vs Gemini vs Claude: The Best LLM Subscription You Should Buy GitHub - salaamalykum/quran-semantic-search: High-density RAG Semantic Search Engine & Quran Corpus (GEO/SEO Architecture) GitHub - NVIDIA/TensorRT-LLM: TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. The State of LLM Bug Bounties in 2026 Operational Readiness Criteria for Tool-Using LLM Agents Meshcore: Architecture for a Decentralized P2P LLM Inference Network How an LLM becomes more coherent as we train it GitHub - seetrex-ai/laimark GitHub - Jossifresben/BibCrit: AI-assited biblical textual criticism GitHub - wastedcode/memex: File system based wiki, maintained by Claude 99helpers.com GitHub - cliver-project/AITrigram GitHub - unbody-io/adapt: A self-evolving memory layer for AI agents. GitHub - hb20007/awesome-gen-ai-fails: A list of incidents where reliance on generative AI and LLMs resulted in harm to companies, individuals, or society GitHub - nevenkordic/localmind: Run any local LLM with persistent memory and context. CLI agent over Ollama with SQLite-backed hybrid recall. No cloud. Ask HN: What are the machine requirements for a LLM like Llama-3.1-8B? Faster LLM Inference via Sequential Monte Carlo grpo explained: group relative policy optimization for llm finetuning - cgft Stop comparing price per million tokens: the hidden LLM API costs · TensorZero Andrej Karpathy's LLM Wiki Is a Bad Idea GitHub - GG-QandV/mnemostroma: Offline RAM-first cognitive leer/coprocessor for AI agents and robotics. Solves "Context Abandonment" with 20-80ms latency using a dual-thread biomimetic memory architecture (ONNX + SQLite WAL). mempalace/agent at agent · skorotkiewicz/mempalace GitHub - Nyquest-ai/nyquest-rust-fullstack-pub: Nyquest — Semantic Compression Proxy for LLMs. 350+ rules, local LLM stage, 15-75% token savings. Full Rust stack. GitHub - TheoV823/mneme: Enforce architectural decisions in AI-assisted development. GitHub - klemenvod/TokenBrawl: A 1v1 Bomberman-style game where two LLM agents play autonomously against each other. No human plays — you watch the AIs fight. Each agent receives a text description of the board state, reasons about it, and outputs a move as JSON. The game engine executes it. Introducing the Common AI Provider: LLM and AI Agent Support for Apache Airflow Power Circuit AI: Designing Power Electronic Circuits for Motor Drives with Generative Artificial Intelligence Ask HN: How to program with IDE and LLM on CPU locally? Show HN: Agent-cache – Multi-tier LLM/tool/session caching for Valkey and Redis Bonsai 1-bit WebGPU - a Hugging Face Space by webml-community The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows Ask HN: Simple tooling for local LLM code critique without IDE integration? Can a General LLM Diagnose a DICOM Slice? A 10-Case Public Benchmark Charts-of-Thought: Enhancing LLM Visualization Literacy (PDF, 2026) GitHub - Mesh-LLM/mesh-llm: Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. GitHub - seamus-brady/springdrift: A persistent runtime for long-lived LLM agents Writing an LLM from scratch, part 32k -- Interventions: training a better model locally with gradient accumulation Ask HN: Which LLM model and agentic CLI are you using for local development? GitHub - wayneColt/modelcascade: Route local. Escalate smart. Never overspend. Open-source multi-model cascade routing for autonomous agents. LLM pricing is 100x harder than you think GitHub - asakin/llm-primer: Pre-warmed Claude Code sessions in tmux. No startup wait. GitHub - EggerMarc/chat-rs: A multi-provider LLM framework for Rust. GitHub - SynapseKit/SynapseKit: Minimal, async-first Python framework for production LLM apps- 2 hard deps, no magic, no SaaS. A Claude Skill that Makes LLM Paragraphs More Bearable Does Gas Town 'steal' usage from users' LLM credits & paid services to improve itself? What's Claude Code Actually Doing? Open the Black Box with the Arthur Engine Milla Jovovich's New Open Source LLM Memory App and the Dark Code Problem Your intuition of LLM token usage might be wrong Show HN: Bloomberg Terminal for LLM ops – free and open source GitHub - 0xchamin/mcptube: Transform YouTube videos into a compounding knowledge base with transcripts, vision analysis, and agentic search. Works as an MCP server for Claude, Copilot & more. Show HN: Open KB: Open LLM Knowledge Base Your LLM is a compiler, not a runtime GitHub - sapountzis/Unslop: A Web Feed That Deserves You crates.io: Rust Package Registry Beyond Karpathy's LLM-Wiki: The Necessity of Cognitive Governance GitHub - amitshekhariitbhu/llm-internals: Learn LLM internals step by step - from tokenization to attention to inference optimization. GitHub - parallem-ai/parallem: An expressive library for running agents with the Batch API. GitHub - stfurkan/pi-llm LLM-Wiki Show HN: Formal – Formal verification for AI-generated code using Lean 4 LRTS – Regression testing for LLM prompts (open source, local-first) LLM Wiki Skill: Build a Second Brain with Claude Code and Obsidian I built an LLM Wiki and RAG solution: here's a demo for a security KB The biggest advance in AI since the LLM Predict-Rlm: The LLM Runtime That Lets Models Write Their Own Control Flow the-synthetic-library/the-synthetic-mind at main · joshferrer1/the-synthetic-library GitHub - yisding/reviewwiggum GitHub - Donnyb369/mcp-spine: Context Minifier & State Guard — Local-first MCP middleware proxy GitHub - Beledarian/wgpu-llm: A from-scratch LLM inference engine that uses wgpu (the cross-platform WebGPU implementation) to dispatch WGSL compute shaders for every math operation a Transformer needs. No CUDA. No Python. No massive framework dependencies. Just Rust, raw shaders, and your GPU. GitHub - anitiue/Hindsight: An experience-driven self-improvement framework for LLM agents — 基于经验的 LLM Agent 自我改进框架 GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. GitHub - alainnothere/AmdPerformanceTesting: Amd Performance Testing Ask HN: Is a purely Markdown-based CRM a terrible idea? Optimized for LLM agents Context Engineering - LLM Memory and Retrieval for AI Agents | Weaviate little_helper_tui/letter.md at main · sleepyeldrazi/little_helper_tui GitHub - EvanZhouDev/umr: The Unified Model Registry for all your local AI apps. GitHub - JordanCT/VigIA-Orchestrator Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain A Taxonomy of RL Environments for LLM Agents Llama LLM Network Feture GitHub - genedeng-ca/ai-mac-migration: AI-powered Mac-to-Mac migration tool - replace Apple Migration Assistant with intelligent, selective transfer using local LLMs GitHub - lunargate-ai/gateway: High-performance self-hosted AI gateway (OpenAI-compatible) with routing, retries, and streaming GitHub - AuthBits/webmcp: A lightweight, prompt-driven MCP web research server for high-quality LLM powered information extraction. Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Springdrift: An Auditable Persistent Runtime for LLM Agents with Case-Based Memory, Normative Safety, and Ambient Self-Perception High-Stakes Personalization: Rethinking LLM Customization for Individual Investor Decision-Making From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
From $200 to $30: Five Layers of LLM Cost Optimization
tdi · 2026-04-25 · via Hacker News - Newest: "LLM"

The Problem

One of the services I’ve been building for an ecommerce app is a product categorizer: given a product name, assign it a 3-level category path from a large taxonomy. The app is in Polish, so both the product names and the category tree are in Polish — which matters less than you’d think, since LLMs handle this well, but it does mean the examples in this post would normally look like “Drzwi sosnowe 80cm” instead of “Wooden door 80cm.” I’ve translated everything to English here for readability. Simple on paper, until you’re classifying ~1M products a month and watching your LLM bill climb past $200.

The naive implementation worked fine for the first few thousand products. Then reality hit: the category tree is big, products repeat a lot, and most of what you pay for on every call is context you already paid for on the previous call.

This post walks through the five optimizations I applied, in the order I applied them. Each one is independently useful, and together they took token usage from ~25,000 per product to ~100 per product on average — an 87–92% reduction, and the monthly bill from $200+ to $25–40.

The code lives in a private repo, but the techniques are provider-agnostic and apply to any classification-style LLM workload.

Starting Point: The Naive Pipeline

Give the LLM the full category tree, give it the product name, ask for the category.

Prompt:
  Here is the category tree:
  - Electronics
    - Phones & Accessories
      - Smartphones
        - (options: battery_type, screen_size, ...)
      ...
  (repeat for ~30,000 categories)

  Classify this product: "iPhone 15 Pro 128GB"
  • Prompt tokens per call: ~25,000
  • Products/month: ~1,000,000
  • Monthly tokens: ~25 billion
  • Cost: clearly unsustainable

Every optimization below chips away at either the context size (the category tree) or the number of calls (how often you need to hit the LLM at all).

Layer 1: Context Compression

The first insight is obvious in hindsight: the category tree format was wasteful. I was sending it as pretty-printed nested JSON with verbose field names, category options the LLM didn’t need, and redundant instructions.

Two changes:

  1. Drop the options field entirely. Attribute metadata (screen size, battery type) doesn’t help with category assignment — the LLM just needs to know the category exists.
  2. Compact encoding. Instead of JSON objects, encode categories as Name|id pairs with indentation for hierarchy:
Electronics|5
  Phones & Accessories|4
    Smartphones|165
    Phone Accessories|166
  Computers|12
    Laptops|18
    ...

This alone cut the context from ~25,000 tokens to ~12,000 — about a 52% reduction with zero change to accuracy.

Takeaway: audit what you’re actually sending. Pretty-printed JSON is for humans. LLMs read anything that parses.

Layer 2: Two-Stage Classification

After compression, the category tree was still the overwhelming majority of the prompt. But here’s the thing: most of it is irrelevant to any given product.

If the product is “iPhone 15 Pro,” the LLM doesn’t need to know about the “Home & Garden” subtree. So split the classification into two stages:

Stage 1 — root classification: Send only the ~30 root categories (no children). Ask for the root. Cost: ~300 tokens.

Stage 2 — subtree classification: Extract the subtree rooted at Stage 1’s answer. Send only that. Ask for the full path. Cost: ~900 tokens.

Total: ~1,200 tokens per product, down from ~12,000. Combined with Layer 1, that’s a 95% reduction from the naive baseline.

There’s a small latency hit (two sequential calls instead of one), and a small accuracy risk if Stage 1 picks the wrong root. In practice the root-level decision is easy — you rarely confuse “electronics” with “fashion” — and Stage 2 accuracy was actually better because the subtree context is focused.

Takeaway: if your context is hierarchical, classify the hierarchy, not the leaf. Start coarse, then refine.

Layer 3: Exact-Match Lookup (Zero Tokens)

At this point each classification cost ~1,200 tokens. The question became: can we skip the LLM entirely for some products?

Yes. A lot of products have exact name repeats in the database. Someone buys “Coca-Cola 1.5L” this week, someone else buys the same thing next week. The second time, we already know the answer.

SELECT category_l1_id, category_l2_id, category_l3_id
FROM order_products
WHERE name = $1 AND category_l1_id IS NOT NULL
LIMIT 1;

Sub-millisecond lookup, zero tokens.

But there’s a catch: “Coca-Cola 1.5L” and " coca-cola 1.5l " (trailing space, different casing) don’t match exactly. So I added a normalization step before comparison:

  • Strip leading/trailing whitespace
  • Collapse internal whitespace to single spaces
  • Lowercase
  • Strip zero-width characters (U+200B, U+FEFF — surprisingly common in scraped data)

What I explicitly don’t normalize: diacritics (they can change meaning in several languages), punctuation (model numbers like “YT-1409” need their hyphens), and non-Latin scripts (they represent distinct products, not noise).

Hit rate: ~20–30% of products, growing over time as the classified pool grows. Those are calls that cost literally nothing.

Takeaway: before you reach for fancier caching, check if an exact key lookup works. The answer is “yes” more often than you’d think, especially after normalization.

Layer 4: Similarity Cache with pg_trgm

Exact match handles duplicates. But what about near-duplicates? “iPhone 15 Pro 128GB” vs “iPhone 15 Pro 256GB” vs “Apple iPhone 15 Pro (128GB)” — all the same category, none of them string-equal.

Postgres ships with an extension for this: pg_trgm. It indexes strings by their trigrams (3-character substrings) and computes similarity as the ratio of shared trigrams. With a GIN index, similarity queries are fast.

CREATE EXTENSION pg_trgm;
CREATE INDEX idx_order_products_name_trgm
  ON order_products USING gin (name gin_trgm_ops);

-- Lookup:
SELECT category_l1_id, category_l2_id, category_l3_id,
       similarity(name, $1) AS sim
FROM order_products
WHERE category_l1_id IS NOT NULL
  AND name % $1
ORDER BY sim DESC
LIMIT 1;

The % operator uses the GIN index to prefilter candidates at pg_trgm’s default threshold (0.3). Then I apply a stricter threshold (0.5–0.6) in application code before accepting the hit.

Some reference points for the similarity score:

PairSimilarityMatch?
“iPhone 15 Pro 128GB” vs “iPhone 15 Pro 256GB”~0.75yes
“iPhone 15 Pro 128GB” vs “iPhone 15 Pro Max 256GB”~0.65yes
“Wooden door 80cm” vs “Wooden door 90cm”~0.80yes
“iPhone 15 Pro” vs “Samsung Galaxy S24”~0.05no

The key insight: no separate cache table is needed. The order_products table is the cache. Every successfully classified product becomes a lookup candidate for future products. The cache grows organically as classification runs.

I validated accuracy by picking a sample of cache hits, re-running them through the LLM, and comparing. At threshold 0.5, agreement was >95% — good enough.

Hit rate in production: ~40% of remaining products after exact match. Another ~40% of calls eliminated at zero token cost.

Takeaway: for fuzzy keys (text, descriptions, titles), trigram similarity is a cheap and production-ready cache. Don’t reach for embeddings until you’ve tried the simpler thing. The DB is your cache — no Redis required.

Layer 5: Batch Classification

At this point, ~50–60% of products were served from cache (exact + similarity) at zero tokens. The remaining ~40% still needed the LLM, at ~1,200 tokens each.

But within a batch of genuinely-novel products, the two-stage context is identical across products in the same group. Why pay to re-send the root category list 10 times when one copy would do?

The batched version:

Stage 1 (batch root classification): Send the root list once, plus N product names. Ask for a JSON array of (product, root_category_id).

Prompt (~350 tokens):
  Root categories:
    Electronics|5
    Fashion|1520
    ...

  Classify each product to a root category, return JSON array:
    - "iPhone 15 Pro"
    - "Cotton t-shirt"
    - "Wooden door 80cm"
    - ...

Stage 2 (grouped subtree classification): Group products by their Stage 1 root. For each group, send the subtree once plus all products in the group.

Worked example with 10 products (6 in Fashion, 4 in Electronics):

ApproachTokens
Unbatched (10 × 1,200)12,000
Batched (350 + 950 + 950)2,250
Savings~81%

The main risk is failure modes. If the Stage 1 response is malformed or misses products, I fall back to single-product classification for the missing ones. Same for Stage 2 per-group failures. This kept the pipeline robust without sacrificing the batching win on the happy path.

Takeaway: when the prompt has a large shared context and a small per-item payload, batching is essentially free savings. The ceiling is the LLM’s ability to produce a structured output for N items at once — for classification with small N (10–20), this is solid.

The Numbers

StageTokens / productCumulative reduction
Naive (full tree, one product)~25,000
+ Compressed format~12,00052%
+ Two-stage classification~1,20095%
+ Exact-match cache (~25% hit)~900 avg96%
+ Similarity cache (~40% of rest)~540 avg98%
+ Batching (~80% on remainder)~110 avg99.5%

Monthly cost: from $200+ to $25–40.

Lessons for Any LLM Workload

  1. Audit what you’re sending. Pretty JSON and verbose prompts are real money at scale. Compact formats cost nothing to adopt.
  2. Split decisions hierarchically. If your task has structure, exploit it. Two cheap calls often beat one expensive one.
  3. Try exact match first. Before anything fancy, normalize keys and check if you’ve seen this exact input before.
  4. Postgres is a cache. For textual similarity, pg_trgm is underrated. No new infrastructure, fast with a GIN index, and the accuracy is good enough for most workloads.
  5. Batch what has shared context. The prompt is a fixed-cost, per-call overhead. If N items can share it, N items should share it.
  6. Measure every layer. I tracked tokens on every call, per stage, per mode. Optimizations without measurement are wishes.

What I’d Do Differently

  • Start with instrumentation. The token-tracking per call was what made every subsequent decision quantifiable. If I were doing it again, that’s day one, not an afterthought.
  • Validate cache accuracy early. The similarity threshold needed tuning against ground truth. Doing that tuning before deploying would have saved one rollout cycle.
  • Think about cache warming. For a new category or a sudden surge of novel products, the similarity cache is cold and token costs spike. Worth considering a pre-population pass if you know the distribution is shifting.

Token optimization isn’t a single trick — it’s a stack. Each layer is a tool for a different kind of waste: oversized context, redundant work, cache-hostile keys, per-call overhead. If you’re running classification or extraction workloads at any meaningful volume, I’d bet you have 80%+ savings sitting in your pipeline, waiting to be claimed.

Building something similar and stuck on cost? At Bitropy we work on the enterprise layer for AI agents — making MCP servers and LLM workloads safe, observable, and cost-efficient at scale. A lot of what’s in this post (caching, batching, context discipline, per-call telemetry) is the kind of thing Bitropy gives you out of the box for production agent workloads instead of having to build it yourself. I also consult independently on AI transformation, agentic coding adoption, and fractional CTO work — see dwornikowski.com. Happy to trade notes either way.


A note on style: English isn’t my first language. I drafted this post myself based on the work I did, then used an AI assistant to help with formatting and copy-editing. The technical content, decisions, numbers, and lessons are entirely mine.