惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
T
The Exploit Database - CXSecurity.com
C
CERT Recently Published Vulnerability Notes
Simon Willison's Weblog
Simon Willison's Weblog
T
Tor Project blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
D
DataBreaches.Net
The Hacker News
The Hacker News
有赞技术团队
有赞技术团队
Latest news
Latest news
T
Tailwind CSS Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
H
Help Net Security
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
Threat Research - Cisco Blogs
G
GRAHAM CLULEY
G
Google Developers Blog
W
WeLiveSecurity
Project Zero
Project Zero
WordPress大学
WordPress大学
人人都是产品经理
人人都是产品经理
博客园 - 司徒正美
博客园 - 三生石上(FineUI控件)
MyScale Blog
MyScale Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
F
Full Disclosure
The Last Watchdog
The Last Watchdog
Security Archives - TechRepublic
Security Archives - TechRepublic
Attack and Defense Labs
Attack and Defense Labs
N
News and Events Feed by Topic
博客园 - 【当耐特】
Google DeepMind News
Google DeepMind News
V
Visual Studio Blog
Blog — PlanetScale
Blog — PlanetScale
F
Fortinet All Blogs
PCI Perspectives
PCI Perspectives
小众软件
小众软件
N
News | PayPal Newsroom
罗磊的独立博客
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
AI
AI
T
Tenable Blog
S
Schneier on Security
O
OpenAI News
The Register - Security
The Register - Security
Google DeepMind News
Google DeepMind News
Engineering at Meta
Engineering at Meta
T
Threatpost
Hacker News: Ask HN
Hacker News: Ask HN

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Your RAG System Is Broken. Your Chunks Are Why.
Arnav Sharma · 2026-06-15 · via DEV Community

Arnav Sharma

80% of RAG failures trace back to one decision made before the first vector is ever stored. Most teams never look at it.

The Wrong Thing to Fix
Your RAG system is giving bad answers. You swap the LLM for a bigger one. Still bad. You rewrite the prompt. Marginally better. You switch embedding models. Barely moves the needle.
Meanwhile, nobody has looked at how the documents were chunked.
This is the most common failure pattern in production RAG systems in 2026, and it is almost entirely invisible during development. The system produces answers. The answers look reasonable in testing. And then users ask real questions and something is quietly, consistently wrong.
80% of RAG failures trace back to the ingestion and chunking layer, not the LLM. Most teams discover this after spending weeks tuning prompts and swapping models while their retrieval quietly returns the wrong context every third query.

What Chunking Is and Why It Matters So Much
When you build a RAG system, you cannot feed an entire document library into a vector database at once. You break documents into chunks — smaller pieces that get individually embedded and stored. When a query arrives, the system retrieves the most relevant chunks, not the most relevant documents.
This means the chunk is the atomic unit of your retrieval system. Everything depends on whether the right chunk surfaces for the right query.
If the chunk is too large, it contains multiple topics and the embedding becomes diluted — the vector represents a mixture of concepts rather than a single coherent idea. Retrieval suffers because nothing matches anything cleanly.
If the chunk is too small, it lacks the surrounding context that gives it meaning. The chunk surfaces correctly but the LLM cannot generate a useful answer from it because critical context was in the adjacent chunk that did not get retrieved.
If the chunks cut across the wrong boundaries — splitting a table halfway, breaking a paragraph mid-sentence, separating a question from its answer — the retrieved content is technically present but practically useless.
The largest controlled comparison of chunking strategies to date tested 36 methods, 6 domains, 5 embedding models, and 1,080 total configurations (Shaukat et al., arXiv:2603.06976, March 2026). It confirmed that content-aware chunking significantly outperforms naive fixed-length splitting, and the gap is not marginal.

The Default Is Wrong
Most teams start with fixed-size chunking. You pick a token count — say, 512 tokens — and every document gets cut into pieces of exactly that size, with or without overlap. It is easy to implement, it is the default in most frameworks, and it produces reliably mediocre retrieval.
Weaviate's September 2025 guide puts a number on the gap: the wrong chunking approach can open a difference of up to 9% in recall between the best and worst methods on the same corpus, with the same retriever.
9% recall sounds small. In a system answering 10,000 queries per day, a 9% recall gap means 900 queries per day where the LLM was missing information it should have had. Some of those will produce noticeably wrong answers. Most will produce subtly incomplete ones — answers that are close enough to pass casual review but wrong enough to matter when someone acts on them.
The January 2026 systematic analysis on arXiv produced a finding that upends conventional wisdom: chunk overlap, the near-universal default of adding 10% to 20% overlap between adjacent chunks to preserve context, provides no measurable benefit in retrieval quality. Teams are adding complexity and storage costs to their chunking pipelines for a technique that the most rigorous analysis to date found does not help.

The Hierarchy That Actually Works
The chunking approach with the strongest evidence behind it in 2026 is hierarchical chunking — sometimes called parent-child chunking.
The idea is straightforward. Documents are indexed at two levels. Large parent chunks — full sections, full paragraphs — capture context. Small child chunks capture specific claims, facts, or data points. When a query arrives, the system retrieves based on the small child chunks (which match more precisely) but returns the surrounding parent chunk (which provides the context the LLM needs to answer usefully).
NVIDIA's internal testing on university presentation decks found that hierarchical chunking improves answer accuracy from 61% with fixed-size chunks to 89%. That is a 28 percentage point improvement from a chunking decision alone — with the same model, the same embedding, and the same vector database.
A 28 point accuracy improvement is not what teams expect to find in their chunking layer. It is what they find when they finally look.

Re-Ranking: The Second Fix Nobody Uses
Even with good chunking, approximate nearest-neighbor search introduces noise. The retrieval step optimizes for speed and will include semantically adjacent chunks that are not actually relevant to the query. This is a property of vector similarity search — it finds things that are conceptually close, not things that are definitively correct.
Re-ranking addresses this. A cross-encoder re-ranker takes the retrieved chunks and scores them again, more carefully, against the actual query. It acts as a quality filter between retrieval and generation.
Cross-encoder re-ranking boosts precision by 18% to 42% compared to retrieval without re-ranking, according to multiple production evaluations. Re-rankers add 50 to 200ms of latency and compute cost — but they reduce LLM token consumption by passing fewer, more relevant chunks. At scale, the LLM cost savings frequently outweigh the re-ranker cost.
Most RAG systems deployed in 2024 and early 2025 do not have a re-ranking step. It was considered an optional optimization rather than a core component. By 2026, re-ranking has moved from optional to expected in production-grade RAG pipelines. Teams running systems without it are leaving significant accuracy on the table.

The Silent Decay Problem
There is one more dimension to the chunking problem that is rarely discussed: RAG systems degrade over time without changing.
A v1 RAG that scored 90 on launch can easily score 60 a year later without a single line of code changing. The world moves, the system does not.
Embedding models improve. The model you chose at launch is likely not the best available option twelve months later. Upgrading embedding models requires re-chunking and re-indexing everything — which most teams plan to do but few actually execute on schedule.
Source documents change. If your knowledge base is built on documents that get updated — policy documents, product documentation, regulatory filings — but your index is not refreshed at the same cadence, you are answering questions from stale context. The system looks like it is working. It is working from outdated information.
Evaluation coverage drifts. The questions your evaluation set was designed around are not necessarily the questions real users are asking six months after launch. A system optimized for the original test questions but misses the evolved user intent will show good numbers on internal benchmarks and bad results in production.

What Good Retrieval Infrastructure Makes Possible
The chunking decisions, the re-ranking layer, the index refresh cadence — all of these matter, but they all rest on the same foundation: a vector database that retrieves accurately and efficiently at the scale your system actually reaches.
Good chunking on a database with poor recall still misses results. The best re-ranking layer cannot recover from retrieved chunks that do not contain the right information to begin with. The architectural layers depend on each other, and the retrieval infrastructure is the layer everything else sits on.
This is why the retrieval database is not a commodity choice. High recall is not a nice-to-have. It is the baseline requirement that makes everything else in the pipeline work as designed.
The teams that get this right build systems that improve over time — better chunking, better re-ranking, better evaluation, all producing measurably better answers. The teams that get it wrong keep swapping models and rewriting prompts while the actual problem sits quietly in their chunking configuration.
Endee is an open-source vector database (Apache 2.0) that delivers the highest recall of any independently benchmarked database — the retrieval foundation that makes everything else in your RAG pipeline work correctly. Free to start at endee.io.