惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
Check Point Blog
罗磊的独立博客
博客园 - 叶小钗
Google DeepMind News
Google DeepMind News
Hugging Face - Blog
Hugging Face - Blog
人人都是产品经理
人人都是产品经理
J
Java Code Geeks
WordPress大学
WordPress大学
大猫的无限游戏
大猫的无限游戏
Blog — PlanetScale
Blog — PlanetScale
F
Fortinet All Blogs
小众软件
小众软件
M
MIT News - Artificial intelligence
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
酷 壳 – CoolShell
酷 壳 – CoolShell
The GitHub Blog
The GitHub Blog
A
About on SuperTechFans
Y
Y Combinator Blog
Recorded Future
Recorded Future
量子位
美团技术团队
S
Security @ Cisco Blogs
G
Google Developers Blog
Cyberwarzone
Cyberwarzone
C
Cybersecurity and Infrastructure Security Agency CISA
博客园 - 三生石上(FineUI控件)
博客园 - 司徒正美
D
Docker
S
Schneier on Security
T
Tor Project blog
阮一峰的网络日志
阮一峰的网络日志
T
Threatpost
P
Privacy & Cybersecurity Law Blog
C
Cisco Blogs
L
Lohrmann on Cybersecurity
NISL@THU
NISL@THU
Microsoft Azure Blog
Microsoft Azure Blog
博客园 - 聂微东
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
The Exploit Database - CXSecurity.com
A
Arctic Wolf
I
Intezer
Latest news
Latest news
Martin Fowler
Martin Fowler
G
GRAHAM CLULEY
B
Blog
V
Vulnerabilities – Threatpost
The Register - Security
The Register - Security
S
Securelist
T
Tenable Blog

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Amazon AI Cancelling Webcomics Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI
Chris Cowart · 2026-04-12 · via Hacker News - Newest: "AI"

Twelfth post in a series on building business process automation at scale. Infrastructure. Automation. Where automation fails. Statistical validation. Predictive models. When models disagree. A working system. Framework vs. architecture. MCP server. Making small LLMs safe. Quantization benchmark. This time: the full experiment — quantization, a smaller model, self-training data, and a fine-tuning failure that taught more than success would have.


The short version: I benchmarked 7 variants of Llama 3 on a real production verification pipeline — not a synthetic benchmark. Three quantization levels of the 8B model all hit 92% accuracy. A 3B challenger got close at 84% but left an 8% gap from a different class of error. I built a self-training pipeline that generates labeled data from the cascade's own high-confidence outputs — 542 examples, zero manual labeling. Then I fine-tuned a QLoRA adapter on the 3B model to close the gap. It scored 12% — worse than random. The model said NONE to everything. Catastrophic forgetting from overaggressive LoRA rank on a small model. The failure was more instructive than success would have been. Here's the data and the debugging story.


The Task

Blueprint's KYB (Know Your Business) verification engine discovers company career pages through a 4-layer agentic cascade. Layer 2 sends a list of DOM elements scraped from a company's website to a local Llama 3 model via Ollama and asks: which element most likely leads to the target page? The model responds with a number or NONE.

Structured input. Bounded output space. Deterministic evaluation criteria. A constrained classification task — and that distinction matters for interpreting everything that follows.

The model runs on a single Hetzner dedicated server with an NVIDIA RTX 4060 (8GB VRAM). No cloud GPU spend. Everything local, everything sovereign.

This post covers the full experiment: can we cut cost through quantization, shrink the model, generate our own training data, and fine-tune our way to higher accuracy? The answer to the first three is yes. The last one broke.

Phase 1: Quantization Is Free

The first question was simple: does reducing numerical precision degrade accuracy on this task?

I built a benchmark harness that tests three quantization levels of Llama 3 8B against 25 golden test cases (15 careers, 10 contact) across 50 runs each. Same code path as production, just with instrumented timing and a fixed test set. The full deep-dive is here. The headline:

VariantQuantAccuracyP50 (ms)P95 (ms)MemoryTok/sCost/1K
Llama 3 8BQ4_092.0%2192945.0 GB8.9$0.05
Llama 3 8BQ8_092.0%4591,0878.8 GB3.9$0.12
Llama 3 8BFP1692.0%1,0913,64515.5 GB1.5$0.30

Identical accuracy across all three levels. Not "within margin of error" — the same precision, recall, and F1 scores per task type. The same confusion matrix. The same mistakes on the same cases.

Q4 is the Pareto champion: 5x faster, 3x less memory, 6x cheaper than FP16. The 8% error rate (false picks on ambiguous NONE cases) is identical across all three — it's a prompt/task design issue, not a precision issue. Quantization is a free lunch for this class of task.

But 92% isn't 100%. And the errors follow a pattern: the model picks an element when the correct answer is NONE. The 8B model's decision boundary for "good enough match vs. no match" is fuzzy. Could a different approach close the gap?

Phase 2: The 3B Challenger

If quantization doesn't degrade quality, what about a smaller model entirely? Llama 3.2 3B at Q4_K_M — less than half the parameters, half the VRAM.

The Numbers

MetricLlama 3 8B Q4Llama 3.2 3B Q4
Accuracy92.0%84.0%
P50 latency219 ms197 ms
P95 latency294 ms289 ms
Memory5.0 GB2.6 GB
Tok/s8.98.4
Cost/1K$0.050$0.053

The 3B is 10% faster and uses half the VRAM. It also drops 8 percentage points on accuracy.

Where the 8% Gap Lives

The error profiles are different in a way that matters.

The 8B model's errors are exclusively false picks — it selects an element when the correct answer is NONE. It never picks the wrong element; when there's a correct answer, it finds it. 100% recall.

The 3B model has a different problem. It has false picks too (20), but it also makes wrong picks (20) — selecting the wrong element from the list entirely. It picks "About Us" when the answer is "Careers." It picks "Contact Sales" when the answer is "Contact Us." The smaller model lacks the capacity to distinguish between superficially similar DOM elements.

This is the "8% problem" from the title. The gap between 84% and 92% isn't just a NONE-boundary issue — it's a reasoning quality issue. The 3B model makes confident but incorrect selections that the 8B model never makes.

The Cost Paradox

Here's the counterintuitive result: the 3B model is not cheaper. $0.053 vs $0.050 per thousand queries. The latency advantage is marginal (197ms vs 219ms). The only dimension where the 3B wins is VRAM — 2.6 GB vs 5.0 GB. If you're running multiple models on the same GPU, that headroom matters. For a single-model deployment, the 8B Q4 dominates on every metric that matters.

Accuracy vs cost for all model variants — the 8B Q4 and 3B Q4 are both Pareto-optimal on different tradeoff frontiers, while the LoRA models cluster at 12% accuracy with high cost

Memory footprint comparison — the 3B model at 2.6 GB is half the 8B Q4 at 5.0 GB, while FP16 variants exceed the GPU's 8 GB VRAM

Both models sit on the Pareto frontier — the 8B optimizes for accuracy, the 3B optimizes for memory. But for this pipeline, accuracy is the binding constraint.

Phase 3: Building a Self-Training Pipeline

The 8% gap suggested fine-tuning. But where do you get training data for a domain-specific classification task where the inputs are scraped DOM elements from arbitrary company websites?

The answer: the cascade itself.

Confidence-Gated Training Data

The KYB verification cascade is 4 layers deep. Each layer produces signals that, taken together, tell you how confident you should be in the result:

  • DOM scoring picks an element AND downstream ATS detection confirms a careers page exists → high-confidence correct pick. That (prompt, response) pair is a training example.
  • All 4 layers fail to find a target page → high-confidence NONE. The company genuinely doesn't have a public careers page.
  • The LLM says NONE, but the probe layer (Layer 4) finds a careers page anyway → the LLM was wrong. This is the most valuable training signal because it directly addresses the false-NONE error mode.
  • The LLM picks an element, but ATS detection doesn't confirm → low confidence. Hold it out for edge case evaluation, not training.

The insight is architectural: the production pipeline generates its own labeled training data. The downstream verification layers provide the labels. No manual annotation required.

The Numbers

I ran 1,000 companies through the full cascade and applied confidence gating:

  • 542 training examples from 446 high-confidence outputs
    • 328 positive picks (322 from DOM scoring, 6 from LLM layer)
    • 110 high-confidence NONEs
    • 8 probe-corrected examples (LLM said pick, probe said NONE)
    • 2x NONE oversampling to combat the false-pick bias → effective ratio of 328 picks to 220 NONEs
  • 117 edge cases from low-confidence outputs, held out for evaluation

One critical design choice: the training data formatter uses the exact same prompt-building functions as the production inference path. build_llm_prompt() and prepare_elements() produce identical prompts whether the model is being trained or doing real inference. No prompt template mismatch between training and serving — a common failure mode in fine-tuning pipelines.

Phase 4: QLoRA Fine-Tuning

This is where the experiment went sideways.

The Setup

QLoRA on Llama 3.2 3B Instruct via Unsloth. The 8B model didn't fit for training on 8GB VRAM even in 4-bit — Unsloth's fused cross-entropy loss kernels need more memory than inference alone.

Training config:

  • LoRA: rank=16, alpha=32, dropout=0.05
  • Target modules: all linear layers — q, k, v, o projections plus gate, up, down MLP layers
  • Training: 3 epochs, batch size 1, gradient accumulation 8 (effective batch 8)
  • Optimizer: AdamW 8-bit, learning rate 2e-4, warmup 10%, weight decay 0.01
  • Precision: bf16, 4-bit NF4 base model quantization
  • Duration: 204 steps, 32 minutes on the RTX 4060

The training loss curve looked perfect. Smooth descent from 2.0 to 0.57 over 204 steps. No spikes, no plateaus, no signs of instability.

I exported the merged model to GGUF at three quantization levels (FP16, Q8, Q4), registered all three with Ollama, and ran the benchmark.

The Results

VariantQuantLoRAAccuracyP50 (ms)MemoryCost/1K
Llama 3 8BQ4_092.0%2195.0 GB$0.050
Llama 3 8BQ8_092.0%4598.8 GB$0.115
Llama 3 8BFP1692.0%1,09115.5 GB$0.297
Llama 3.2 3BQ4_K_M84.0%1972.6 GB$0.053
Llama 3.2 3BQ4_K_MLoRA12.0%6972.6 GB$0.168
Llama 3.2 3BQ8_0LoRA12.0%9824.0 GB$0.224
Llama 3.2 3BFP16LoRA12.0%3,7177.1 GB$0.832

12% accuracy. Across all three quantization levels. The LoRA adapter didn't just fail to close the 8% gap — it destroyed the model's ability to do the task at all.

The confusion matrix tells the story instantly: 100% NONE rate. The fine-tuned model says NONE to every single query. 220 false NONEs (cases where there was a correct element to pick), 30 true NONEs (cases where NONE was actually right). The only "correct" answers are the NONE cases it got right by always saying NONE.

Zero correct picks. Zero wrong picks. Zero precision. Zero recall. The model learned one thing: say NONE.

Confusion matrices — the top row (8B variants) shows healthy distributions with high correct-pick rates. The 3B base model (middle row, left) shows the false-pick problem. The three LoRA variants (middle row right and bottom) show near-total false-NONE dominance

Latency distributions — the LoRA FP16 variant is dramatically slower than all others, with the base 8B Q4 and 3B models showing tight, fast distributions

The Debugging Journey

The first instinct was wrong. And so was the second.

Hypothesis 1: Chat template mismatch. The GGUF export uses a custom Modelfile with the Llama 3.2 Instruct chat template — start/end header IDs, role formatting, stop tokens. If the template was wrong, the model would see garbled input and produce garbage. I verified every token: <|start_header_id|>, <|end_header_id|>, <|eot_id|>, role tags. All correct. The template matched the base model's exactly. Not the cause.

Hypothesis 2: Merge corruption. Unsloth's save_pretrained_gguf() merges LoRA weights into the base model before GGUF conversion. If the merge produced NaN weights or numerical instability, all outputs would be garbage. But the model wasn't producing garbage — it was producing "NONE" in the correct format, consistently, coherently. It understood the task format. It just always chose the same answer. Corrupted weights produce gibberish, not consistent single-token responses. Not corruption.

Hypothesis 3: Catastrophic forgetting. I tested the model on basic questions outside the training domain. "What is 2+2?" The response: EVEN. Reply only (e.g., , ) or ONLY... — fragments of the training prompt format, not an answer. The model wasn't just broken on the KYB task. It had forgotten how to be a language model.

This was the root cause. Rank=16 LoRA modifying all seven linear layer types in every transformer block is too aggressive for a 3B parameter model. The ratio of modified parameters to total parameters was too high. The adapter didn't learn to supplement the base model's knowledge — it overwrote it.

The training loss was misleading. The smooth descent from 2.0 to 0.57 showed the model memorizing the format tokens of the training examples — learning to emit the right special tokens and response structure — not generalizing from the training distribution to unseen inputs. With only 542 examples and 3 epochs, there wasn't enough data diversity to prevent the model from collapsing to the most "safe" response: NONE is never a wrong pick, it's only a missed pick. The model learned that saying nothing is better than saying something wrong.

Why the Failure Is Instructive

This failure clarified three things that success would have obscured:

Loss curves lie. A smooth loss descent is necessary but not sufficient. The curve showed memorization of format tokens, not generalization to the task. Always benchmark the merged model on held-out data before declaring victory. If I'd only looked at the training metrics, I'd have shipped a model that says NONE to everything.

Small models need gentle fine-tuning. Rank=16 on a 3B model is too aggressive. The parameter budget of LoRA (rank x 2 x hidden_dim x num_layers x num_modules) as a fraction of total model parameters matters. For a 3B model, rank=4-8 with a lower learning rate (5e-5 instead of 2e-4) and fewer target modules would preserve more of the base model's capability. The same config that works on a 7B or 13B model can destroy a 3B model.

The cascade architecture is more robust than model improvement. The 4-layer verification cascade catches errors at layers above the model. The defensive AI patterns treat every model output as untrusted until verified. Trying to make the model "smarter" via LoRA is fragile. Making the system around the model smarter is durable. The architecture is the product, not the model.

What We Learned

  1. Quantization is free for classification tasks. Q4_0 matches FP16 exactly on this element-picking task. Same accuracy, same error profile, 6x cheaper. Don't pay for precision you don't need.

  2. Loss curves lie. A smooth descent from 2.0 to 0.57 looked like successful training. The model had memorized the format, not the task. Always benchmark the merged model on held-out data before declaring victory.

  3. Small models need gentle fine-tuning. Rank=16 on a 3B model is too aggressive. The ratio of modified parameters to total parameters matters. For small models, use rank=4-8, lower learning rate (5e-5), and fewer target modules.

  4. Self-training pipelines are the real asset. The 542-example dataset generated from production cascade outputs is reusable. When we fine-tune with better hyperparameters — or on a larger model with more VRAM — the data infrastructure is already built. Zero manual labeling.

  5. The Pareto frontier is clear without fine-tuning. Llama 3 8B at Q4 delivers 92% accuracy at 219ms for $0.05 per thousand queries on a consumer GPU. The 8% gap to perfect isn't worth chasing with the current hardware constraints — the cascade's downstream layers catch the errors the model misses.

Where This Goes Next

Two directions.

Better training strategy. The same 542 examples with gentler hyperparameters — rank=4, learning rate 5e-5, target only attention layers, single epoch. Or DPO (Direct Preference Optimization) instead of SFT, which teaches the model to prefer correct picks over NONE without overwriting its base capabilities. DPO is better suited for tasks where the model already mostly works and you're sharpening a decision boundary.

Cascade improvement over model improvement. Instead of making the LLM smarter, make the layers around it smarter. Better confidence thresholds for when to escalate. Better probe strategies for catching false NONEs. Better downstream verification for catching false picks. The architecture is designed to compensate for model limitations — investing in the compensation layers may be more durable than investing in the model itself.

The quantization benchmark proved we're not leaving performance on the table. The self-training pipeline proved we can generate our own training data at zero marginal cost. The LoRA failure proved that the next improvement isn't a model problem — it's a systems problem. Each experiment narrowed the search space for where the next 8% of accuracy actually lives.


The full source for the KYB pipeline, benchmark harness, and LoRA training code is on GitHub: avatar296/blueprint.


Have questions about this topic?

We love talking tech. Reach out and let's discuss how this applies to your business.

Get in Touch