惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - Franky
有赞技术团队
有赞技术团队
宝玉的分享
宝玉的分享
雷峰网
雷峰网
Hugging Face - Blog
Hugging Face - Blog
V
V2EX
大猫的无限游戏
大猫的无限游戏
博客园 - 司徒正美
D
Docker
T
The Blog of Author Tim Ferriss
罗磊的独立博客
博客园 - 叶小钗
酷 壳 – CoolShell
酷 壳 – CoolShell
Blog — PlanetScale
Blog — PlanetScale
月光博客
月光博客
J
Java Code Geeks
Jina AI
Jina AI
博客园 - 【当耐特】
C
Check Point Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
腾讯CDC
Last Week in AI
Last Week in AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
V
Visual Studio Blog

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis
Understand the scope of user impact with Watchdog Impact ...
Nicholas Thomson · 2021-12-01 · via Datadog | The Monitor blog
Nicholas Thomson

Nicholas Thomson

Technical Content Writer

Watchdog is Datadog’s machine learning and AI engine, which leverages algorithms like anomaly detection to automatically surface performance issues in your infrastructure and applications. Without any manual setup or configuration, Watchdog generates a feed of Alerts—on anomalies such as latency spikes, elevated error rates, and network issues in cloud providers—to help you reduce your mean time to detection.

Watchdog now includes Impact Analysis, which shows you how performance issues in your application may be impacting users. Whenever Watchdog finds a new APM anomaly, it will automatically assess if the anomaly is adversely affecting any web or mobile pages that are instrumented with RUM by analyzing a variety of latency and error metrics submitted from the RUM SDKs. If Watchdog finds that the issue is impacting end users, it will provide information about which view paths and users were affected.

see the user impact of Watchdog alerts

At a glance, Watchdog Impact Analysis helps you determine the scope of a performance issue in a service, including which part(s) of your application are impacted and who the issue is affecting. In this post, we’ll show you how you can leverage Impact Analysis to:

  • Quickly prioritize which issues to troubleshoot

  • Reduce business impact by understanding which users were affected

  • Prevent similar issues from happening in the future by creating detailed postmortems

Prioritize troubleshooting problems that impact the greatest number of users

While Watchdog often surfaces actionable issues within your applications, it may also find anomalies that don’t require immediate intervention. It can be difficult to distinguish between these two cases, especially if you aren’t familiar with the particular services involved.

Watchdog Impact Analysis addresses this challenge by clearly showing you which application performance anomalies are impacting your users. Whenever Watchdog determines that an issue is associated with user impact, it will prominently display details in the “Impacts” section.

For example, let’s say that you see two Watchdog Alerts at the top of your feed. The first tells you that latency is up on the product recommendation service, but the Impact Analysis indicates that this issue is only affecting a few customers.

Alerts affecting a small number of users are low priority

The second Alert informs you of a faulty deployment of your address service, and Impact Analysis shows you that this issue is affecting significantly more customers and thus should be dealt with first.

Alerts affecting a large number of users are high priority

Once you’ve pinpointed the most important issue, you can easily troubleshoot by clicking on the Alert to see RUM data and APM traces. Whenever a Watchdog APM Alert involves multiple services, the Alert will include a dependency map that illustrates the spread of the performance issue through your application. If Watchdog is able to determine the root cause of the issue, it will be highlighted on the dependency map.

See exactly which users are impacted

In addition to helping you troubleshoot issues more effectively, Impact Analysis also provides you with a list of users that were potentially affected by the disruption.

View a list of users affected by a performance issue

The dropdown list above shows you that out of 1.22k total users who visited the impacted pages during the anomaly window, 183 experienced degraded performance. You can click to see the affected users’ contact information if you need to reach out to them. Additionally, you can click on the view paths pill to see Session Replays that show the impact from the perspective of users who actually experienced the disruption. Seeing what impacted users saw makes it even easier to assess whether the issue is critical or not.

See impacted users’ view paths

Create postmortems to document user impact and other findings

All postmortems should have an “impacts” section, where you can document an incident’s impact on users. You can easily export Impact Analysis data to a Datadog Notebook and supplement it with live graphs and other visualizations. This allows you to quickly create postmortems and follow best practices by centralizing all your data in a single place as you investigate user-facing incidents. Notebooks are also fully editable, so everyone on your team can leave comments and contribute additional data throughout the incident response process.

Streamline writing postmortems

Start using Watchdog Impact Analysis

Watchdog Impact Analysis is now automatically enabled for all Datadog APM and RUM users, so you can immediately get deeper visibility into the real-world impact of service performance issues, easily prioritize troubleshooting efforts, and streamline the postmortem process. See our documentation to learn more. If you’re new to Datadog, sign up for a 14-day free trial.