惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
N
Netflix TechBlog - Medium
博客园 - 叶小钗
大猫的无限游戏
大猫的无限游戏
S
SegmentFault 最新的问题
V
V2EX
IT之家
IT之家
J
Java Code Geeks
Hacker News - Newest:
Hacker News - Newest: "LLM"
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
GbyAI
GbyAI
D
Docker
S
Secure Thoughts
Recent Announcements
Recent Announcements
Webroot Blog
Webroot Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
云风的 BLOG
云风的 BLOG
博客园_首页
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Security Archives - TechRepublic
Security Archives - TechRepublic
酷 壳 – CoolShell
酷 壳 – CoolShell
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
N
News | PayPal Newsroom
S
Security @ Cisco Blogs
I
InfoQ
Last Week in AI
Last Week in AI
SecWiki News
SecWiki News
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
W
WeLiveSecurity
T
Troy Hunt's Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Attack and Defense Labs
Attack and Defense Labs
美团技术团队
T
The Blog of Author Tim Ferriss
Google DeepMind News
Google DeepMind News
Martin Fowler
Martin Fowler
B
Blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
Scott Helme
Scott Helme
T
Tor Project blog
Know Your Adversary
Know Your Adversary
有赞技术团队
有赞技术团队
Hugging Face - Blog
Hugging Face - Blog
Recorded Future
Recorded Future
C
Cyber Attacks, Cyber Crime and Cyber Security
AI
AI
G
Google Developers Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
I Spent 6 Months Fixing RAG. Here's What I Found (And Built)
vigneshwar · 2026-05-19 · via DEV Community

This is the story of a debugging session that turned into a research paper.

The Bug That Started Everything
I was building a document Q&A system — nothing exotic. Standard RAG setup. FAISS index, SBERT embeddings, GPT as the reader. Classic.

It worked fine on simple questions. "What is the refund policy?" → correct answer.

Then I tested it on a multi-hop question: "What are the environmental compliance requirements for facilities that process the chemicals used in the manufacturing process described in section 4.2?"

The model's answer was confident. Detailed. And completely wrong.

The retrieved documents were all there. Every piece of information needed to answer correctly was in the context window. But the model still hallucinated a number that appeared nowhere in any document.

I started logging. What I found was two distinct failure modes happening simultaneously:

Failure Mode 1: Semantic Drift. By the time the query had been reformulated for multi-hop retrieval, the embedding had drifted so far from the original intent that we were retrieving the wrong documents. Not slightly wrong. Documents from a completely different section of the corpus.

Failure Mode 2: Context Poisoning. Even when we retrieved mostly correct documents, 1–2 tangentially related but contradictory chunks were slipping through. And those poison chunks were enough to derail the model.

Standard RAG has no defense against either of these. The pipeline is essentially: embed → retrieve → stuff into context → hope.

I needed something better.

Six Months Later: VORTEXRAG
I'm releasing the full framework today. 7 layers, each targeting a specific failure mode. Here's what I built and why each layer exists.

Layer 1: Tri-Vector Encoding (TVE)
The problem: Single-vector embeddings collapse too much information. "The bank charged a fee" and "The river bank was steep" share a close embedding in SBERT space even though they're semantically unrelated in most retrieval contexts.

The solution: Three encoding arms:

Semantic arm: standard SBERT 768d

s = sbert_model.encode(chunk) # shape: (768,)

Syntactic arm: POS + dependency structure

t = syntactic_encoder(chunk) # shape: (64,)

Causal arm: verb-argument chains

c = causal_encoder(chunk) # shape: (32,)

Fused vector

v = concat([α·s, β·t, γ·c]) # shape: (864,)
The causal arm is the key innovation. It captures "X causes Y" relationships that pure semantic similarity misses entirely. This is what allows the pipeline to distinguish between a document that mentions a concept and a document that explains the causal mechanism behind it.

Layer 2: Vortex Retrieval Cone (VRC)
The problem: Flat cosine similarity treats all high-similarity documents equally. But document #1 and document #47 in your ranked list shouldn't have equal weight — there's a natural falloff in relevance.

The solution: Spiral ranking inspired by vortex dynamics:

spiral_rank = TVE · e^(−λr) · cos(nθ)
Where r is the radial distance (rank position) and θ is an angular phase that encodes causal depth. Documents with high causal relevance get a phase advantage that can overcome a slightly lower semantic similarity score.

In practice: the top-k documents returned by VRC are causally denser than those returned by a flat cosine search on the same index.

Layer 3: Semantic Drift Corrector (SDC)
The problem: Multi-hop queries reformulate themselves at each hop. Each reformulation can drift slightly from the original intent. Over 3–4 hops, this compounds into a completely different query.

The solution: Track the embedding trajectory. At each hop:

drift = query_embedding - anchor_embedding
SDS = 1 - tanh(np.linalg.norm(drift) / τ)

if SDS < 0.72:
query_embedding = re_anchor(query_embedding, anchor_embedding)
The SDS score (Semantic Drift Score) measures how far we've drifted. Below 0.72, we re-anchor to the original query intent. This single intervention eliminated most of our multi-hop hallucinations.

Layer 4: Context Poison Guard (CPG)
The problem: 1–2 contradictory or off-topic chunks in the context window is enough to poison the model's answer. We needed a way to identify and remove these before they reach the LLM.

The solution: Entity-salience ratio per chunk:

ESR = sum(SDS_i * w_i for each entity i) / (num_propositions + ε)

if ESR < 3.5:
flag_for_purging(chunk)
The purging algorithm is provably greedy-optimal (I include the formal proof in the paper). It maximizes total ESR across the retained context while respecting the token budget.

In ablation studies, removing CPG alone dropped faithfulness from 0.94 to 0.75 — an 0.19 point drop from a single layer.

Layer 5: Rank Fusion Gate (RFG)
The problem: Most rank fusion methods are additive. A single terrible signal gets diluted but doesn't eliminate a bad document.

The solution: Multiplicative fusion:

Φ = TVE^α × SDS^β × ESR^γ
If any of the three signals is near zero, Φ collapses toward zero. A document that scores 0.9 on semantic similarity but 0.1 on CPG gets Φ ≈ 0.09 — effectively eliminated.

This was a deliberate design choice. In high-stakes retrieval (medical, legal, financial), you want a veto mechanism, not a popularity contest.

Layer 6: Causal Context Builder (CCB)
The problem: LLMs have dramatically higher attention to the beginning and end of their context window. Documents buried in the middle get "lost" — the Lost-in-the-Middle problem (Liu et al., 2023).

The solution: Reorder chunks by causal depth:

position = rank * causal_depth

High causal depth → lower position number → placed at context start

Chunks that are causally central to answering the question get placed at the beginning of the context window, where attention weights are highest. Peripheral context gets pushed to the middle where it does less damage if ignored.

Layer 7: Faithfulness Verifier (FV)
The problem: Even with all 6 preceding layers, some hallucinations slip through. We need a final gate.

The solution: Score the candidate answer against the retrieved context:

ΔR = 1 - (ROUGE_L * NLI_entailment_score)

if ΔR > 0.15:
regenerate_answer()
If the answer diverges more than 15% from the source documents (measured by ROUGE-L weighted by NLI entailment), it gets thrown out and regenerated. This catches the subtle hallucinations — cases where the model paraphrases correctly but changes a critical number or name.

Results
Tested on 4 standard benchmarks (NQ, HotpotQA, MuSiQue, 2WikiMultiHopQA):

System EM F1 Faithfulness
Naive RAG 61.2 68.4 0.71
HyDE 63.8 71.2 0.74
Self-RAG 65.4 73.1 0.79
FLARE 64.9 72.8 0.77
VORTEXRAG 74.8 82.6 0.94
+13.6 EM. +14.2 F1. +0.23 Faithfulness over naive RAG.

The ablation shows all 7 layers contribute. The two biggest individual contributors are CPG (+0.19 faithfulness) and SDC (+0.08 EM on multi-hop benchmarks specifically).

Quick Start

pip install vortexrag

from vortexrag import VortexRAG, VortexConfig

config = VortexConfig(
sdc_threshold=0.72,
cpg_esr_threshold=3.5,
fv_delta_r_threshold=0.15,
)

rag = VortexRAG(config)
rag.index(your_documents)

answer = rag.query("Your complex multi-hop question here")
What's Next
The framework is MIT licensed. The full research paper (with formal proofs) is on Zenodo.

If you're building RAG systems and hitting hallucination walls — especially on multi-hop or domain-specific queries — this framework is designed for exactly that problem.

GitHub: https://github.com/vignesh2027/VORTEXRAG

Paper: https://doi.org/10.5281/zenodo.20285144

Live demo docs: https://vignesh2027.github.io/VORTEXRAG

Questions? Drop them in the comments — happy to go deep on any of the layers.

Post these tonight at 9:00 PM IST. Reddit first (r/MachineLearning, then r/LocalLLaMA 30 minutes apart), then Dev.to. Cross-link the Dev.to article in your Reddit comments with "I wrote a deeper walkthrough here" — that drives traffic both ways.