惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

MongoDB | Blog
MongoDB | Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
小众软件
小众软件
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - Franky
博客园 - 聂微东
V
Visual Studio Blog
I
InfoQ
罗磊的独立博客
Security Latest
Security Latest
G
Google Developers Blog
博客园_首页
P
Proofpoint News Feed
T
Threat Research - Cisco Blogs
D
DataBreaches.Net
PCI Perspectives
PCI Perspectives
Forbes - Security
Forbes - Security
P
Proofpoint News Feed
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Know Your Adversary
Know Your Adversary
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Recent Commits to openclaw:main
Recent Commits to openclaw:main
The Cloudflare Blog
AWS News Blog
AWS News Blog
Latest news
Latest news
T
Tailwind CSS Blog
P
Palo Alto Networks Blog
Hugging Face - Blog
Hugging Face - Blog
云风的 BLOG
云风的 BLOG
Cyberwarzone
Cyberwarzone
T
The Exploit Database - CXSecurity.com
WordPress大学
WordPress大学
Recorded Future
Recorded Future
A
Arctic Wolf
V
Vulnerabilities – Threatpost
Security Archives - TechRepublic
Security Archives - TechRepublic
宝玉的分享
宝玉的分享
人人都是产品经理
人人都是产品经理
月光博客
月光博客
有赞技术团队
有赞技术团队
P
Privacy & Cybersecurity Law Blog
Scott Helme
Scott Helme
美团技术团队
Hacker News - Newest:
Hacker News - Newest: "LLM"
Jina AI
Jina AI
N
News and Events Feed by Topic
Attack and Defense Labs
Attack and Defense Labs

Snorkel AI

Building AI-Native Systems for Federal Infrastructure: A Conversation with Rezaur Rahman Code World Models and AutoHarness for LLM Agents Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory Building FinQA: An Open RL Environment for Financial Reasoning Agents How Tool Discipline Let a 4B Model Outsmart a 235B Giant on Financial Tasks Coding agents don’t need to be perfect, they need to recover Closing the Evaluation Gap in Agentic AI SlopCodeBench: Measuring Code Erosion as Agents Iterate Introducing the Snorkel Agentic Coding Benchmark 2026: The year of environments Part V: Future Direction and Emerging Trends in Rubric-Based AI Evaluation The self-critique paradox: Why AI verification fails where it’s needed most Chat With the Terminal-Bench Team | Snorkel AI Intelligence per watt: A new metric for AI’s future Terminal-Bench 2.0: Raising the bar for AI agent evaluation Snorkeling in RL environments Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning Scaling Trust: Rubrics in Snorkel's Quality Process Evaluating Multi-Agent Systems in Enterprise Tool Use Evaluating Coding Agents with Terminal-Bench 2.0 Parsing isn’t neutral: why evaluation choices matter The science of rubric design The right tool for the job: An A-Z of rubrics Data quality and rubrics: how to build trust in your models Building the benchmark: inside our agentic insurance underwriting dataset Evaluating AI agents for insurance underwriting LLM observability: key practices, tools, and challenges Anthropic Claude + AWS: revolutionizing pharma data analytics with Snorkel AI Data-centric development of an enterprise AI agent with Snorkel Building the data development platform for specialized AI LLM-as-a-judge for enterprises: evaluate model alignment at scale Why GenAI evaluation requires SME-in-the-loop for validation and trust Research spotlight: is long chain-of-thought structure all that matters when it comes to LLM reasoning distillation? Why enterprise GenAI evaluation requires fine-grained metrics to be insightful What is specialized GenAI evaluation, and why is it so critical to enterprise AI? LLM alignment techniques: 4 post-training approaches Research spotlight: Is intent analysis the key to unlocking more accurate LLM question answering? Why enterprises should embrace LLM distillation What is large language model (LLM) alignment? Databricks + Snorkel Flow: integrated, streamlined AI development How LLM evaluation drives better models in Snorkel Flow Unlock proprietary data with Snorkel Flow and Amazon SageMaker LLM evaluation in enterprise applications: a new era in ML Snorkel AI joins the AWS ISV Accelerate Program and launches Snorkel Flow Availability in AWS Marketplace AI data development: a guide for data science projects SnorkelCon 2024: Inaugural Snorkel AI user conference gathers leaders from 30+ Fortune 500 companies Snorkel Flow 2024.R3: Supercharge your AI development with enhanced data-centric workflows Explore the new GenAI Evaluation Suite: Snorkel 2024.R3 New NLP features in Snorkel Flow 2024.R3 Enterprise data compliance and security review: Snorkel Flow 2024.R3 How a global financial services company built a specialized AI copilot accurate enough for production Task Me Anything: innovating multimodal model benchmarks Alfred: Data labeling with foundation models and weak supervision RAG: LLM performance boost with retrieval-augmented generation Call center AI for customer experience management: a case study New GenAI features, data annotation: Snorkel Flow 2024.R2 How data slices transform enterprise LLM evaluation Meta’s Llama 3.1 405B is the new Mr. Miyagi, now what? Meta’s new Llama 3.1 models are here! Are you ready for it? Data-centric AI with Snorkel and MinIO Weak supervision for non-categorical applications + superalignment Snorkel AI signs strategic collaboration agreement with AWS to help enterprises cross the demo-to-production chasm AI alignment made simple: innovative solutions for businesses How does the Snorkel Flow label model work? Vision language models: how LLMs boost image classification Long context models in the enterprise: benchmarks and beyond How to build production-grade RAG retrieval with Snorkel Flow How Bonito helps fine-tune specialized LLMs faster than ever Walking safely before building flying saucer seatbelts: introducing Enterprise Alignment Role-based access controls in Snorkel Flow secure enterprise data Accelerating AI development in manufacturing with Snorkel Flow and AWS SageMaker How ROBOSHOT boosts zero-shot foundation model performance Discover what’s new in Snorkel Flow: Flexible data and LLM connectivity, secure data controls, and more! Faster than ever document intelligence with new Snorkel Flow FM-first workflow The art of data development for Enterprise LLMs Crossing the demo-to-production chasm with Snorkel Custom How Snorkel topped the AlpacaEval leaderboard (and why we're not there anymore) CRFM's HELM and enterprise LLM evaluation beyond accuracy How we achieved 89% accuracy on contract question answering Five sessions not to miss at Google Cloud Next 24 Content filtering breakthrough: Snorkel client reaches 96% recall in 3 days Here's how Snorkel Flow + Google AI built an enterprise-ready model in a day Snorkel teams with Microsoft to showcase new AI research at NVIDIA GTC How Skill-it! enables faster, better LLM training Fine-tuned representation models boost LLM systems. Here's how Enterprise GenAI to surge in 2024: survey results Large language model training: how three training phases shape LLMs LoRA: Low-Rank Adaptation for LLMs LLM distillation demystified: a complete guide Enterprises must shift their focus from models to data in AI development Insurance’s GenAI revolution: a business perspective Scaling human preferences in AI: Snorkel's programmatic approach Building better enterprise AI: incorporating expert feedback in system development “Fall in love with your data”—Snorkel AI’s Enterprise LLM Summit Why QBE Ventures invested in Snorkel AI New benchmark results demonstrate value of Snorkel AI approach to LLM alignment Retrieval augmented generation (RAG): a conversation with its creator Snorkel Flow 2023.R4: enhanced UI + PDF and Databricks tools How Snorkel Flow users can register custom models to Databricks Stanford professor discusses exciting advances in foundation model evaluation
Retrieval-augmented generation (RAG) failure modes and how to fix them
Matthew Casey · 2025-02-06 · via Snorkel AI

Retrieval-augmented generation (RAG) represents a leap forward in natural language processing. Well-crafted RAG systems deliver meaningful business value in a user-friendly form factor. However, these systems contain multiple complex components. RAG failure modes can inhibit a system’s value and cause significant headaches.

In this article, we will explore some of the most common pitfalls encountered in RAG pipelines and provide actionable solutions to address them.

What is retrieval-augmented generation (RAG)?

RAG systems combine the strengths of reliable source documents with the generative capability of large language models (LLMs).

The simplest RAG system consists of a vector database, an LLM, a user interface, and an orchestrator such as LlamaIndex or LangChain. After a user enters their query, the system retrieves relevant documents or document chunks from the vector database and adds them to the initial request as context. This final prompt gives the LLM more context with which to answer the user’s question.

Learn more about retrieval-augmented generation in our guide.

Prompt engineering: crafting effective prompt templates

Poorly designed prompt templates can lead to off-target responses or outputs that lack the desired specificity. This is akin to giving a colleague or subordinate incomplete or poorly-formed instructions; if they don’t understand the task, they can’t complete it.

Solving challenges with prompt templates

Begin by clearly defining the prompt’s objective and the desired characteristics of the output. Experiment with different prompt structures, starting with simple instructions and iteratively incorporating more complex directives as needed. Consider using prompt engineering techniques such as few-shot learning, where relevant examples are included to guide the model’s response.

You may also need to add dynamic variables to your prompt template. For example, an LLM does not know the current date. If any intended task requires this knowledge, you should include a step in your pipeline that injects the current date and time into the prompt.

Whatever approach you take, iterate. Adjust your prompt, re-run it on a consistent set of representative queries, and inspect the results. You should soon find a working template.

Data coverage and quality

When data scientists build RAG systems, they do so with the intent of enriching users’ requests with high-quality, relevant context. The systems’ users assume that it’s drawing on a comprehensive set of high-quality documents, but that’s not always the case.

A RAG system that lacks complete coverage could leave the LLM with no context to draw on, increasing the likelihood that it generates a “hallucination.” Additionally, a system that includes poorly curated documents could give the end user misleading, incorrect, or outdated information.

Solving challenges with data coverage and quality

If your RAG system’s responses seem unanchored in the right underlying information, check the context included in the prompts. If the prompts lack appropriate context, thoroughly search your vector database and then fill in any gaps you find.

Teams managing RAG applications should regularly audit and expand the system’s dataset to ensure comprehensive coverage. They should also check older documents and ensure that they remain accurate; organizational policies and product offerings will change over time. Automated tools can help identify gaps or biases in the data, enabling proactive adjustments.

Document chunking: striking the balance

Ineffective document chunking can lead to information loss, noisy context, or irrelevant retrievals, hampering the performance of RAG pipelines. Chunking documents into sections that are too large may result in the system overlooking pertinent information. Overly granular chunks can separate pertinent information from important surrounding context.

If the chunks included in your RAG prompts are too long, too short, or cut off in the middle of vital information, you may have a chunking problem.

Improper document chunking is a common rag failure mode.
Chunk size 512, overlap 0.2. Notice how a single section is spread across multiple chunks.
Snorkel AI's dynamic chunker overcomes the common RAG failure mode of improper document chunking
The SnorkelDynamicChunker automatically chunks documents by section!

Improving RAG outcomes with better document chunking

While off-the-shelf RAG orchestration tools typically chunk documents according to a set number of tokens by default, they include other chunking options. Experimenting with different token windows and overlaps may yield usable results. If they don’t, switching to paragraph, page, or semantic segmentation might.

In Snorkel Flow, we use a proprietary algorithm that leverages semantic similarity and document structure to approximate how a human might separate document sections. This yields chunks of different sizes that typically include enough information to feel self-contained. In practice, we have found that this algorithm yields production-grade results.

Embedding models: capturing semantics accurately

Challenges with embedding models initially appear similar to challenges with document coverage; the prompts lack the appropriate context.

If you have this problem and verify that the appropriate information exists in your vector database, the challenge likely springs from your embedding model. Generalist embedding models usually won’t capture the semantic nuances of domain-specific data, causing the system to struggle to prioritize chunks for retrieval at inference time. Some retrieved chunks may be relevant. Others may not be. An off-the-shelf embedding model likely can’t tell the difference.

Improving retrieval through custom embedding models

Selecting the right pre-trained embedding model for your domain can significantly improve the relevance of retrieved chunks. However, optimal embedding model performance requires custom fine-tuning on your proprietary corpus.

Fine-tuning an embedding model calls for three-part training examples, which consist of:

  1. A query.
  2. A chunk of context that is highly relevant to that question.
  3. A chunk of context that is irrelevant to that question.

This process works best if the “wrong” example is similar to the “right” example. These “near misses” (also known as “hard negatives”) highlight document chunks that look or behave similarly to the query but are not truly related. These examples can be hard to find, and sometimes the process is referred to as “hard negative mining.”

Data scientists can achieve results faster by randomly pairing correct chunks with unrelated ones. This may yield an embedding model less sensitive to important nuances than one trained with hard negatives, but the result may satisfy deployment needs for far less effort.

Chunk separation on a on off-the-shelf (left) and customized (right) embedding model in a real-world, domain-specific use-case.

Chunk enrichment: enhancing contextual relevance

Sometimes, vector-based retrieval will fall short of enabling your system to prioritize all of the correct information at all times—regardless of how well you cusomize your embedding model. Some tasks demand prioritizing specific information that may not appear semantically similar to a user’s query.

For example, an application built to find payment dates may need to surface all passages that include references to dates, which an embedding model may struggle to prioritize on its own.

Enriching chunks with metada enables hybrid approaches that leverage categorical information as well as vector embeddings.

Better results through chunk tagging

Teams in charge of RAG systems can enrich document chunks with additional metadata or annotations to aid in retrieval and generation. This could include tagging chunks with key entities, summarizing sections, or linking related documents.

Data teams can add tags to chunks as a one-time exercise or build it into the pipeline. Supporting tools such as named entity recognition (NER) or topic models can automatically enrich incoming chunks with relevant contextual information.

Once they’ve added tags to the chunks, data scientists can use their orchestration framework to filter results accordingly. They can do this separately and in addition to finding chunks with high relevance scores. This process might, as a result, add two separate sets of chunks to the context—one optimized for relevancy and one optimized for tags—ensuring that the LLM gets everything it needs to answer the user’s question.

LLM fine-tuning: aligning with task-specific needs

Off-the-shelf LLMs can produce generic or contextually inappropriate outputs, undermining the effectiveness of the RAG pipeline. If your RAG system retrieves the correct context, but the model returns a result that is out of line with your expectations—tonally, factually, or format-wise—you likely need to fine-tune your LLM.

Improving RAG outcomes with customized LLMs

Fine-tuning LLMs on domain- and task-specific datasets aligns their generative capabilities with the desired output. You may want to begin by finding an open source LLM already fine-tuned to your domain. That may be enough to sharpen responses to meet your standards. If it isn’t—or you can’t find an LLM for your domain—the next step is to fine-tune your chosen open-source model with appropriate prompts and high-quality responses.

The team’s data scientists should work closely with subject matter experts to label a corpus of model responses (either historical or freshly generated) as high or low quality. We recommend that teams do this scalably with programmatic labeling, but other methods can also work.

Once curated, use the prompt/response pairs to fine-tune the model and align its outputs with what you want them to look like. This may require several iterations.

Debug your RAG system and enhance its enterprise value

RAG pipelines hold immense potential to transform how enterprises interact with information by providing contextually enriched, accurate responses. However, their success hinges on careful attention to the nuances of their implementation.

Organizations can harness the full power of RAG systems by understanding and addressing common failure modes—such as document chunking, embedding models, LLM fine-tuning, chunk enrichment, and prompt template engineering. Emphasizing a data-centric approach, continuous iteration, and user feedback will ensure these pipelines not only meet but exceed expectations, driving innovation and problem-solving to new heights.

Learn More

Follow Snorkel AI on LinkedInTwitter, and YouTube to be the first to see new posts and videos!

Featured image by Jessica Ruscello on Unsplash