惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

月光博客
月光博客
Martin Fowler
Martin Fowler
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
The Last Watchdog
The Last Watchdog
S
Schneier on Security
C
Cisco Blogs
P
Privacy International News Feed
T
Tenable Blog
Spread Privacy
Spread Privacy
Recent Commits to openclaw:main
Recent Commits to openclaw:main
N
News and Events Feed by Topic
Security Archives - TechRepublic
Security Archives - TechRepublic
阮一峰的网络日志
阮一峰的网络日志
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
大猫的无限游戏
大猫的无限游戏
Project Zero
Project Zero
GbyAI
GbyAI
N
Netflix TechBlog - Medium
T
Tor Project blog
雷峰网
雷峰网
Y
Y Combinator Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
T
Threat Research - Cisco Blogs
Cyberwarzone
Cyberwarzone
L
LangChain Blog
MyScale Blog
MyScale Blog
C
CERT Recently Published Vulnerability Notes
C
Check Point Blog
G
Google Developers Blog
T
Tailwind CSS Blog
L
LINUX DO - 热门话题
宝玉的分享
宝玉的分享
IT之家
IT之家
F
Fortinet All Blogs
TaoSecurity Blog
TaoSecurity Blog
Recent Announcements
Recent Announcements
T
The Exploit Database - CXSecurity.com
Hacker News: Ask HN
Hacker News: Ask HN
aimingoo的专栏
aimingoo的专栏
云风的 BLOG
云风的 BLOG
Engineering at Meta
Engineering at Meta
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Google Online Security Blog
Google Online Security Blog
Help Net Security
Help Net Security
H
Hacker News: Front Page
小众软件
小众软件
U
Unit 42
Apple Machine Learning Research
Apple Machine Learning Research
P
Privacy & Cybersecurity Law Blog
T
Threatpost

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis Monitor Aruba Central in Datadog How we centralize and remediate risks with Datadog Case Management Accelerate incident response with Datadog and ServiceNow Monitor your application and network load balancer logs Understanding Karpenter architecture for Kubernetes autoscaling Tools for collecting metrics and logs from Karpenter Monitor Karpenter with Datadog What your product data is actually saying Key metrics for monitoring Karpenter Securing Datadog’s platform in the AI age: The role of observability data Four ways engineering teams use the Datadog MCP Server to power AI agents Approaching your observability migration with the right mindset Meet the new Bits AI SRE: Deeper reasoning, twice as fast Key learnings from the 2026 State of DevSecOps study Use plain English to query your multi-cloud infrastructure in Resource Catalog Simplifying troubleshooting across the user journey with Datadog Synthetic Monitoring Protect your OCI resources with Datadog Cloud Security This Month in Datadog - February 2026 Amazon EC2 security: How misconfigured and public AMIs expand your cloud attack surface Enable end-to-end visibility into your Java apps with a single command Measure and improve mobile app startup performance with Datadog RUM Evaluating our AI Guard application to improve quality and control cost Identify untested code across every level of your codebase Make use of guardrail metrics and stop babysitting your releases Monitor Versa Networks SD-WAN performance in Datadog Improve performance and reliability with APM Recommendations Remediate transitive vulnerabilities faster with Datadog Software Composition Analysis Generate audit-ready vulnerability and compliance reports with Datadog Sheets Monitor Fortinet FortiManager performance in Datadog Improve test coverage across codebases with Datadog Code Coverage Move fast, don’t break things: Consistent testing standards at scale Enrich logs with ServiceNow CMDB context before routing to any SIEM or logging tool Monitor Lustre with Datadog Make faster, better product decisions with Datadog Product Analytics Surface and remediate runtime posture issues with Workload Protection Findings Protect agentic AI applications with Datadog AI Guard How to optimize JavaScript code with CSS Trace Google Pub/Sub workloads in Cloud Run with Datadog Detect human names in logs with ML in Sensitive Data Scanner How we cut our NLQ agent debugging time from hours to minutes with LLM Observability Debug PostgreSQL query latency faster with EXPLAIN ANALYZE in Datadog Database Monitoring Datadog acquires Propolis Unify and correlate frontend and backend data with retention filters Scale compliance across global frameworks with Datadog Cloud Security Monitor Arista VeloCloud SD-WAN performance with Datadog Building reliable dashboard agents with Datadog LLM Observability Simplify log collection and aggregation for MSSPs with Datadog Observability Pipelines Mitigation for Node.js denial-of-service vulnerability affecting Datadog APM Automate flaky test fixes with the Bits AI Dev Agent and Test Optimization How we built an AI SRE agent that investigates like a team of engineers Datadog integrations 2025 recap: Observability for AI, security, and hybrid cloud Design effective executive dashboards with Datadog Implement dbt data quality checks with dbt-expectations Bring faster visibility into AWS Lambda functions with remote instrumentation Troubleshoot faster with the GitLab Source Code integration in Datadog How Cambia Health Solutions saved $30,000 monthly with Cloud Cost Management and the Datadog Resource Catalog Normalize any logs for Cloud SIEM with Datadog's OCSF processor Optimizing Datadog at scale: Cost-efficient observability at Zendesk Detect, diagnose, and resolve network issues easily with CNM Network Health Connect engineering errors to user impact in early-stage products Cilium configuration for Kubernetes operations at scale Designing feedback loops for progressive delivery Ship features faster and safer with Datadog Feature Flags Choosing the right OpenTelemetry Collector distribution Route your monitor alerts with Datadog monitor notification rules Automate Cloud SIEM investigations with Bits AI Security Analyst Cloud threat detection: How to identify risky activity across control and data planes Collecting Kafka performance metrics Monitoring Kafka with Datadog Monitoring Kafka performance metrics
Diagnose and resolve database performance issues faster with Database Investigator
Ethan Perez, Zhengda Lu, Joel Marcotte · 2026-05-08 · via Datadog | The Monitor blog

When your database performance degrades, diagnosing the root cause is rarely quick or straightforward. Your existing tools might surface metrics like CPU utilization, wait events, and query duration, but then leave you to correlate the data and identify what went wrong. Worse, what first appears to be the root cause can often just be a downstream effect of multiple interrelated issues. Cutting through all that complexity to get to an actionable fix requires deep database expertise, application knowledge, and institutional context. 

Database Investigator takes an agentic approach to solving your database performance issues. It draws on Datadog’s context about your database and application, and layers in experience from real-world incidents, to handle the diagnostic heavy lifting of root cause analysis. Engineers can ask questions and get answers in plain language about what broke, why it happened, and how to resolve it. For DBAs, platform teams, and application developers without deep database expertise, this means faster mean time to resolution (MTTR), fewer escalations, and the ability to find and fix performance issues.

In this blog post, we will cover how Database Investigator makes it easy for teams to:

Diagnose database issues without deep expertise

With Database Investigator, any engineer can diagnose and resolve database performance issues. It independently examines workload metrics, query samples, execution plans, and logs across your stack, then points to a root cause along with concrete remediation steps. Each suggested step includes links to the relevant queries, services, and database instances, with live graphs displayed to confirm symptoms or verify fixes. And after reviewing the results of an investigation, engineers can refine the analysis by asking follow-up questions or adding context.

Database Monitoring overview page with the Database Investigator panel open on the right.

Trace a latency spike back to its source

When a deployment causes a performance regression, identifying whether the database is involved can be surprisingly difficult. Traditional tooling forces you to bounce between Application Performance Monitoring (APM) traces, deployment logs, service health dashboards, and execution plans, leaving you to stitch that data together yourself. Database Investigator does that work for you by correlating distributed traces, query metrics, and node-level execution plans in a single view. With all this information at its disposal, Database Investigator can quickly tell you which query regressed and on which instance.

Here’s an example of how this works in practice: Imagine an on-call engineer is paged because the p95 latency of a service endpoint has just tripled. The engineer follows the traces through APM to Database Monitoring and launches a Database Investigator investigation. More than 15 health checks run immediately. The health checks reveal that query latency has jumped from 15 ms to 447 ms, that 770 MB of shared blocks are read with each query, and that cache hit ratio has dropped from 99.5% to 71.8%. Database Investigator uses this information to identify the latency spike as a query-level regression, not instance saturation.

Database Investigator panel showing root cause analysis of a query regression, including the offending query and remediation steps.

Pulling sampled execution plans, Database Investigator then identifies the cause: An index scan has flipped to a sequential scan on a large table, and the sequential scan is reading the entire table from disk. Cross-referencing schema and plan data, Database Investigator determines that the WHERE predicate in the updated query is not covered by an index. APM correlation ties the deploy directly to the latency spike and scan flip. The engineer adds a composite index and validates the fix by asking Database Investigator to re-check performance metrics. Logical reads are back to baseline, latency is at 16 ms, and execution plans are back to index scans.

Detect connection pool exhaustion

Connection pool exhaustion is notoriously difficult to identify as the root cause of poor database performance. When this is the case, the application might be throwing errors, but CPU utilization is often low, disk space is ample, and no individual query is failing. Without the right tooling, the root cause is effectively invisible.

Database Investigator can detect this type of problem because it can see the low-level interactions between application behavior and connection state. It also analyzes connection state breakdowns, transaction durations, and wait events together to surface what individual metrics cannot.

Consider a scenario where a service team sees “too many clients” errors growing, but database CPU utilization is remaining stable under 20%. Running the Database Investigator immediately surfaces non-obvious signals, such as multiple transactions running longer than five minutes and high transaction age. The team also sees that the database is waiting on the application and that latency has increased across more than half of the instance queries. Information about connection state exposes the core issue: 87 connections are stuck in idle, up from three at baseline, and 18 connections are blocked waiting for a slot. This is the signature of connection pool exhaustion: idle-in-transaction sessions holding slots that other connections are waiting for.

Database Investigator panel showing root cause analysis of connection pool exhaustion, including idle-in-transaction sessions and remediation steps.

Catch replication lag before it affects your data

Replication lag is another database issue that is difficult to diagnose, specifically because the symptoms and their root cause can live in different parts of the cluster. Stale data returned from replicas can point to the primary’s write throughput or the replica’s own I/O as the culprit, but often neither is the main problem. The real issue is that write-ahead log (WAL) replay has stalled. Database Investigator helps you diagnose replication lag by reasoning about replication internals across your entire cluster. It can trace a lag spiral to the specific query and service that are blocking WAL replay, giving you specific steps to address the root cause.

As an example, consider a scenario where your analytics reports are displaying data that is growing increasingly stale. You know the reports read from a replica, so you perform a quick check on the primary. Everything looks fine, though WALs are starting to accumulate on disk. Your team then starts a Database Investigator investigation on the replica. Initial health checks reveal a long-running transaction and a high replication transaction ID age. Pulling instance-specific telemetry data, Database Investigator finds that replication replay has climbed exponentially and is still growing. WAL write and flush lag are both normal on the primary, which rules out the primary as the source of the lag.

Database Investigator panel showing root cause analysis of replication lag, including a long-running transaction and remediation steps.

Database Investigator then builds on this information to identify a root cause: a runaway single session, open for almost an hour, that is pinning WAL replay. In this situation, Database Investigator recommends cancelling the query to allow WAL replay to catch up—and correctly notes that this is a short-term fix. For a permanent fix, the investigator recommends ensuring the transaction is committed properly in all cases. Finally, it provides optimizations to improve query performance for the analytics reports. In a scenario that could easily consume many hours of manual investigation across multiple instances, Database Investigator has given the team a clear path to resolution in minutes.

Resolve database performance issues faster with Database Investigator

Database Investigator gives DBAs, platform teams, and application developers a fast and accessible way to resolve database performance issues. It examines the evidence across your entire stack and delivers a root cause together with concrete remediation steps, enabling engineers at all levels to resolve issues with confidence.

To read more about Database Investigator, visit the documentation. To learn more about how Bits AI is bringing agentic AI across the Datadog platform, see our Bits AI documentation. If you’re new to Datadog, you can sign up for a 14-day free trial.