惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 司徒正美
大猫的无限游戏
大猫的无限游戏
腾讯CDC
J
Java Code Geeks
博客园 - 【当耐特】
Microsoft Azure Blog
Microsoft Azure Blog
V
Visual Studio Blog
人人都是产品经理
人人都是产品经理
博客园 - Franky
博客园 - 聂微东
阮一峰的网络日志
阮一峰的网络日志
美团技术团队
云风的 BLOG
云风的 BLOG
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
U
Unit 42
雷峰网
雷峰网
B
Blog RSS Feed
博客园_首页
量子位
F
Fortinet All Blogs
罗磊的独立博客
H
Hackread – Cybersecurity News, Data Breaches, AI and More
酷 壳 – CoolShell
酷 壳 – CoolShell
C
Check Point Blog

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis
Decrypt Devops alerts with contextual graphs, runbooks an...
Alexis Lê-Quôc · 2013-05-28 · via Datadog | The Monitor blog

It’s 3AM and you have just been woken up by the nagging ringtone of your phone. The PagerDuty service is contacting you - as your name was on the on-call support list for today. This late, it can’t be good. Sure enough, despite the bleary eyes and the utter lack of direction, you see a cryptic message: “service templeton is lagging”. You had heard that this new service being rolled out earlier this week but had not had a chance to ask anyone how this was supposed to run. By then you quickly run through your options:

(a) go back to bed and hope that it magically goes away

(b) get up, find your laptop, get some coffee started, find some runbooks and hope they are up-to-date.

You go for option (b) and you spend the next few hours looking for clues on what that alert means in the first place and how to fix the issue. With some luck you may even go back to bed before your morning alarm rings.

Sound familiar? Are you tired of spending precious time figuring out why you were alerted? Do you wish your alerts were more than just a whodunit in 140 characters or less?

You’re not alone: a better design of alerts and how to turn them into useful messages has been at the center of the community’s focus this year at Monitorama. The writing was on the wall, and we decided to do something about it.

At Datadog, just like you, we sometimes get alerts in the middle of the night, and we got tired of sifting through enigmatic text, stats and process names trying to make sense of what exactly was wrong. So we sat down to redesign alerting in a way that any alert would make immediate sense to the recipient. We’ve accomplished this by packing as much useful context as possible with graphs, up-to-date runbooks and routing.

Visualizing Alert Data with Graphs

For instance each alert comes with the graph that immediately shows what the data really looks like so that you can rule out temporary blips and go back to sleep right away.

Visualize Devops Alerts

Integrating Runbooks for Immediate Information

Each alert also comes with an integrated runbook. So the person who authors the alert can easily add simple diagnostic and remediation steps right in the alert. No more endless searches in the corporate wiki, no more out-of-date runbooks.

Runbook Devops Alert Documentation

Routing Alerts to the Right Recipients

Finally, each alert can be routed precisely to the right person, group or service so you don’t get bombarded with a barrage of alerts that are not relevant to you.

Devops alert routing

Adding context to your alerts is quick, easy and available to try for free. Signing up for Datadog takes just a few minutes, and these context-adding features for your devops alerts are available immediately.