惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Scott Helme
Scott Helme
有赞技术团队
有赞技术团队
阮一峰的网络日志
阮一峰的网络日志
雷峰网
雷峰网
D
Docker
Stack Overflow Blog
Stack Overflow Blog
Hugging Face - Blog
Hugging Face - Blog
爱范儿
爱范儿
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
MyScale Blog
MyScale Blog
A
About on SuperTechFans
博客园 - 【当耐特】
U
Unit 42
H
Help Net Security
博客园 - 三生石上(FineUI控件)
V2EX - 技术
V2EX - 技术
T
Tor Project blog
博客园 - 叶小钗
G
Google Developers Blog
S
Securelist
Security Latest
Security Latest
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
T
Threat Research - Cisco Blogs
aimingoo的专栏
aimingoo的专栏
C
Cybersecurity and Infrastructure Security Agency CISA
博客园_首页
V
Vulnerabilities – Threatpost
P
Palo Alto Networks Blog
T
The Exploit Database - CXSecurity.com
The Register - Security
The Register - Security
Recorded Future
Recorded Future
NISL@THU
NISL@THU
量子位
L
LangChain Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
C
Cyber Attacks, Cyber Crime and Cyber Security
C
CERT Recently Published Vulnerability Notes
The Hacker News
The Hacker News
D
DataBreaches.Net
小众软件
小众软件
罗磊的独立博客
Forbes - Security
Forbes - Security
The Last Watchdog
The Last Watchdog
Jina AI
Jina AI
I
InfoQ
S
Schneier on Security
Recent Announcements
Recent Announcements
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
S
Secure Thoughts

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis Monitor Aruba Central in Datadog How we centralize and remediate risks with Datadog Case Management Accelerate incident response with Datadog and ServiceNow Monitor your application and network load balancer logs Understanding Karpenter architecture for Kubernetes autoscaling Tools for collecting metrics and logs from Karpenter Monitor Karpenter with Datadog What your product data is actually saying Key metrics for monitoring Karpenter Securing Datadog’s platform in the AI age: The role of observability data Four ways engineering teams use the Datadog MCP Server to power AI agents Approaching your observability migration with the right mindset Meet the new Bits AI SRE: Deeper reasoning, twice as fast Key learnings from the 2026 State of DevSecOps study Use plain English to query your multi-cloud infrastructure in Resource Catalog Simplifying troubleshooting across the user journey with Datadog Synthetic Monitoring Protect your OCI resources with Datadog Cloud Security This Month in Datadog - February 2026 Amazon EC2 security: How misconfigured and public AMIs expand your cloud attack surface Enable end-to-end visibility into your Java apps with a single command Measure and improve mobile app startup performance with Datadog RUM Evaluating our AI Guard application to improve quality and control cost Identify untested code across every level of your codebase Make use of guardrail metrics and stop babysitting your releases Monitor Versa Networks SD-WAN performance in Datadog Improve performance and reliability with APM Recommendations Remediate transitive vulnerabilities faster with Datadog Software Composition Analysis Generate audit-ready vulnerability and compliance reports with Datadog Sheets Monitor Fortinet FortiManager performance in Datadog Improve test coverage across codebases with Datadog Code Coverage Move fast, don’t break things: Consistent testing standards at scale Enrich logs with ServiceNow CMDB context before routing to any SIEM or logging tool Monitor Lustre with Datadog Make faster, better product decisions with Datadog Product Analytics Surface and remediate runtime posture issues with Workload Protection Findings Protect agentic AI applications with Datadog AI Guard How to optimize JavaScript code with CSS Trace Google Pub/Sub workloads in Cloud Run with Datadog Detect human names in logs with ML in Sensitive Data Scanner How we cut our NLQ agent debugging time from hours to minutes with LLM Observability Debug PostgreSQL query latency faster with EXPLAIN ANALYZE in Datadog Database Monitoring Datadog acquires Propolis Unify and correlate frontend and backend data with retention filters Scale compliance across global frameworks with Datadog Cloud Security Monitor Arista VeloCloud SD-WAN performance with Datadog Building reliable dashboard agents with Datadog LLM Observability Simplify log collection and aggregation for MSSPs with Datadog Observability Pipelines Mitigation for Node.js denial-of-service vulnerability affecting Datadog APM Automate flaky test fixes with the Bits AI Dev Agent and Test Optimization How we built an AI SRE agent that investigates like a team of engineers Datadog integrations 2025 recap: Observability for AI, security, and hybrid cloud Design effective executive dashboards with Datadog Implement dbt data quality checks with dbt-expectations Bring faster visibility into AWS Lambda functions with remote instrumentation Troubleshoot faster with the GitLab Source Code integration in Datadog How Cambia Health Solutions saved $30,000 monthly with Cloud Cost Management and the Datadog Resource Catalog Normalize any logs for Cloud SIEM with Datadog's OCSF processor Optimizing Datadog at scale: Cost-efficient observability at Zendesk Detect, diagnose, and resolve network issues easily with CNM Network Health Connect engineering errors to user impact in early-stage products Cilium configuration for Kubernetes operations at scale Designing feedback loops for progressive delivery Ship features faster and safer with Datadog Feature Flags Choosing the right OpenTelemetry Collector distribution Route your monitor alerts with Datadog monitor notification rules Automate Cloud SIEM investigations with Bits AI Security Analyst Cloud threat detection: How to identify risky activity across control and data planes Collecting Kafka performance metrics Monitoring Kafka with Datadog Monitoring Kafka performance metrics
Troubleshooting RAG-based LLM applications
2024-11-08 · via Datadog | The Monitor blog
Jordan Obey

Jordan Obey

LLMs like GPT-4, Claude, and Llama are behind popular tools like intelligent assistants, customer service chatbots, natural language query interfaces, and many more. These solutions are incredibly useful, but they are often constrained by the information they were trained on. This often means that LLM applications are limited to providing generic responses that lack proprietary or context-specific knowledge, reducing their usefulness in specialized settings. For example, a customer interacting with a financial service company’s chatbot to ask about the status of their recent loan application or for specific details about their account may only receive generic responses about loan processing timelines or account management. That’s likely because the LLM lacks access to customer-specific information, recent transaction details, or any real-time updates from the company’s systems.

To address these limitations, organizations often integrate retrieval-augmented generation (RAG) into their LLM applications. RAG enhances the accuracy and relevance of responses by retrieving information from external datasets beyond an LLM’s preexisting knowledge base, such as websites, financial databases, company policy guidelines, and handbooks. By incorporating that data into the prompt-response cycle, models can generate responses enriched with context from unique data sources. For instance, by implementing RAG, our earlier chatbot example can now access the most recent policy documents from the company’s database in real time, providing up-to-date and accurate responses instead of relying on the old information it was originally trained on.

While RAG enhances LLMs, developers and AI engineers still face certain challenges when building these systems. Integrating all of the retrieval and generation components of RAG-based systems at scale introduces complexity in managing latency, ensuring the relevance of retrieved data, and maintaining model accuracy, all while handling large volumes of diverse information in real time. With so many moving pieces, it can be difficult to identify where and when something has gone wrong.

In this post, we’ll explore how to mitigate some of the common challenges you may face by:

  • Chunking and choosing the right embedding model to reduce latency
  • Implementing hybrid search to limit irrelevant or inaccurate responses
  • Using vector databases to exclude outdated information
  • Scanning prompts and responses to prevent accidental exposure of sensitive data

Chunking and choosing the right embedding model to reduce latency

RAG-based systems retrieve relevant information from a knowledge source or database at inference time, which means that it’s absolutely vital for retrieval to be fast and with minimal delay. High retrieval latencies increase the overall latency of your application, resulting in slow responses, degraded performance, and a negative user experience.

You may, for example, receive complaints that your application is slow at responding to prompts and frustrating your users, who wind up closing the session and leaving the application.

rag-01

There are a few steps you can take to mitigate the issue. For instance, you can improve latency when choosing an embedding model. Embeddings are numerical representations of text and other data types that LLMs use to understand the semantic and contextual relationship between different pieces of information they may encounter. In an embedding model, values representing words such as “cat” and “dog” might be closer together than the value for “rocketship” because cats and dogs are both animals.

When choosing embedding models, engineers are often tempted to select those with the highest dimensionality, which captures more context-rich and highly granular data. However, as dimensionality increases, so too does the amount of data that needs to be processed, which increases latency and drives up costs. By choosing lower-dimensional embeddings, you can reduce the computational load and speed up retrieval without significantly sacrificing accuracy, leading to faster and improved performance.

You should also consider chunking large documents or inputs into smaller, manageable parts that fit within token limits. By processing and retrieving smaller, contextually meaningful chunks, overall response times improve as models can retrieve the most relevant information faster. You might consider selecting models with smaller token limits, which can help improve latency because it means models will have less data to process.

Implementing hybrid search to limit irrelevant or inaccurate responses

The quality of responses your RAG system generates relies heavily on the relevance of the documents retrieved when it receives a query. If, for example, a customer service chatbot for a clothing brand’s e-commerce site is asked about sales on jackets and then erroneously retrieves documents about sales on jackets and shirts, it will return irrelevant information outside of the scope of the original query. Alternatively, your system might retrieve the correct documents but still generate factually inaccurate responses.

rag-02

You can improve your system’s document retrieval mechanism by implementing hybrid search. Hybrid search combines term-based retrieval, which looks for exact keyword matches, with vector-based retrieval, which uses embeddings to capture the meaning or semantics of both queries and documents. By using both retrieval methods, your system can filter the list of possible documents down to a manageable volume and then rerank the remaining documents to reorder search results based on semantic similarity. Breaking the list down into a smaller set with term-based retrieval improves the efficiency of vector-based retrieval, allowing the system to perform detailed similarity comparisons without overwhelming resources.

Responses include outdated information

It’s crucial that the data your RAG system retrieves stays fresh and up-to-date. This is particularly important in cases where systems are expected to deliver data that may be updated in real time, such as breaking news and stock market reports. If a system relies on outdated information, it not only loses its key advantage but can also lose the trust of its users. For example, if users rely on a RAG system to deliver breaking news but consistently receive stale information, they will leave for more reliable news sources.

You can mitigate the risk of retrieving irrelevant information by implementing effective metadata filtering, which uses structured data descriptors—such as tags, timestamps, categories, and authors—to refine and target searches. Metadata filtering is crucial for retrieving contextually relevant information at scale, as it narrows results to meet specific criteria (like recent dates or particular topics) to enhance response quality. You can also lower the risk of retrieving irrelevant data through regular updates and proper versioning. Establishing data pipelines for periodic updates and proper versioning ensures that data stays current and reliable.

Application retrieves sensitive data that shouldn’t be exposed

Without the proper guardrails, RAG systems can unintentionally retrieve and expose sensitive data that users should not be able to access—creating a serious security and privacy issue. This exposure can be due to a number of factors, including lack of strong access control permissions, insufficient query filtering, and the inclusion of sensitive information within training data.

Your first line of defense against security vulnerabilities is to make sure you have tight permission and access controls in place to restrict accessibility to sensitive information. For instance, let’s say you use Elasticsearch as the retrieval engine of your RAG system. With Elasticsearch you can define roles to specify permissions and ensure your system only retrieves documents that users have read access to.

“Security & Safety” detections in Datadog LLM Observability will scan LLM prompts and responses for signs of data leakage and attempted prompt injections. In the screenshot below, a user triggered the Prompt Injection Scanner by asking a chatbot to return an admin password. You can identify the bad actor who attempted this breach and take appropriate security measures, such as logging the attempt, blocking the user, or escalating the issue to the security team for further investigation and mitigation.

rag-03

Start troubleshooting your RAG-based LLMs today

In this post, we looked at some of the challenges AI developers and engineers face when building out RAG-based LLMs along with some steps you can take to overcome those challenges. And with Datadog LLM Observability, you can identify when your RAG system experiences high latencies, fails to respond accurately, and encounters other issues connected to these challenges so that you can quickly begin troubleshooting. To learn more about LLM Observability, check out our documentation.

And if you aren’t already using Datadog, sign up today for a 14-day free trial.