惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
IT之家
IT之家
F
Fortinet All Blogs
腾讯CDC
酷 壳 – CoolShell
酷 壳 – CoolShell
月光博客
月光博客
GbyAI
GbyAI
Recent Announcements
Recent Announcements
博客园 - 叶小钗
H
Hackread – Cybersecurity News, Data Breaches, AI and More
博客园 - 司徒正美
Y
Y Combinator Blog
人人都是产品经理
人人都是产品经理
N
Netflix TechBlog - Medium
博客园_首页
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
爱范儿
爱范儿
F
Full Disclosure
博客园 - 【当耐特】
V
V2EX
M
MIT News - Artificial intelligence
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
G
Google Developers Blog
Stack Overflow Blog
Stack Overflow Blog
D
DataBreaches.Net
Martin Fowler
Martin Fowler
Cisco Talos Blog
Cisco Talos Blog
大猫的无限游戏
大猫的无限游戏
T
The Blog of Author Tim Ferriss
S
Schneier on Security
Latest news
Latest news
S
Secure Thoughts
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Hugging Face - Blog
Hugging Face - Blog
P
Proofpoint News Feed
Apple Machine Learning Research
Apple Machine Learning Research
T
Tenable Blog
Blog — PlanetScale
Blog — PlanetScale
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Cyberwarzone
Cyberwarzone
Spread Privacy
Spread Privacy
K
Kaspersky official blog
L
Lohrmann on Cybersecurity
宝玉的分享
宝玉的分享
A
About on SuperTechFans
Jina AI
Jina AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
P
Palo Alto Networks Blog
Microsoft Azure Blog
Microsoft Azure Blog

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis Monitor Aruba Central in Datadog How we centralize and remediate risks with Datadog Case Management Accelerate incident response with Datadog and ServiceNow Monitor your application and network load balancer logs Understanding Karpenter architecture for Kubernetes autoscaling Tools for collecting metrics and logs from Karpenter Monitor Karpenter with Datadog What your product data is actually saying Key metrics for monitoring Karpenter Securing Datadog’s platform in the AI age: The role of observability data Four ways engineering teams use the Datadog MCP Server to power AI agents Approaching your observability migration with the right mindset Meet the new Bits AI SRE: Deeper reasoning, twice as fast Key learnings from the 2026 State of DevSecOps study Use plain English to query your multi-cloud infrastructure in Resource Catalog Simplifying troubleshooting across the user journey with Datadog Synthetic Monitoring Protect your OCI resources with Datadog Cloud Security This Month in Datadog - February 2026 Amazon EC2 security: How misconfigured and public AMIs expand your cloud attack surface Enable end-to-end visibility into your Java apps with a single command Measure and improve mobile app startup performance with Datadog RUM Evaluating our AI Guard application to improve quality and control cost Identify untested code across every level of your codebase Make use of guardrail metrics and stop babysitting your releases Monitor Versa Networks SD-WAN performance in Datadog Improve performance and reliability with APM Recommendations Remediate transitive vulnerabilities faster with Datadog Software Composition Analysis Generate audit-ready vulnerability and compliance reports with Datadog Sheets Monitor Fortinet FortiManager performance in Datadog Improve test coverage across codebases with Datadog Code Coverage Move fast, don’t break things: Consistent testing standards at scale Enrich logs with ServiceNow CMDB context before routing to any SIEM or logging tool Monitor Lustre with Datadog Make faster, better product decisions with Datadog Product Analytics Surface and remediate runtime posture issues with Workload Protection Findings Protect agentic AI applications with Datadog AI Guard How to optimize JavaScript code with CSS Trace Google Pub/Sub workloads in Cloud Run with Datadog Detect human names in logs with ML in Sensitive Data Scanner How we cut our NLQ agent debugging time from hours to minutes with LLM Observability Debug PostgreSQL query latency faster with EXPLAIN ANALYZE in Datadog Database Monitoring Datadog acquires Propolis Unify and correlate frontend and backend data with retention filters Scale compliance across global frameworks with Datadog Cloud Security Monitor Arista VeloCloud SD-WAN performance with Datadog Building reliable dashboard agents with Datadog LLM Observability Simplify log collection and aggregation for MSSPs with Datadog Observability Pipelines Mitigation for Node.js denial-of-service vulnerability affecting Datadog APM Automate flaky test fixes with the Bits AI Dev Agent and Test Optimization How we built an AI SRE agent that investigates like a team of engineers Datadog integrations 2025 recap: Observability for AI, security, and hybrid cloud Design effective executive dashboards with Datadog Implement dbt data quality checks with dbt-expectations Bring faster visibility into AWS Lambda functions with remote instrumentation Troubleshoot faster with the GitLab Source Code integration in Datadog How Cambia Health Solutions saved $30,000 monthly with Cloud Cost Management and the Datadog Resource Catalog Normalize any logs for Cloud SIEM with Datadog's OCSF processor Optimizing Datadog at scale: Cost-efficient observability at Zendesk Detect, diagnose, and resolve network issues easily with CNM Network Health Connect engineering errors to user impact in early-stage products Cilium configuration for Kubernetes operations at scale Designing feedback loops for progressive delivery Ship features faster and safer with Datadog Feature Flags Choosing the right OpenTelemetry Collector distribution Route your monitor alerts with Datadog monitor notification rules Automate Cloud SIEM investigations with Bits AI Security Analyst Cloud threat detection: How to identify risky activity across control and data planes Collecting Kafka performance metrics Monitoring Kafka with Datadog Monitoring Kafka performance metrics
Remediate faster with apps built using Datadog App Builder
2024-06-17 · via Datadog | The Monitor blog
Alex Flinois

Alex Flinois

When troubleshooting an issue or remediating an outage, engineers need tools that are accessible and easy to use, closely integrated with their services, and tailored to their teams’ specific requirements. But the reality is that they often have to jump between their monitoring platform and other tools their team uses—such as internal full-stack tools, the CLI, and cloud provider consoles—to get a clear picture of the status and scope of the problem while taking the necessary action to communicate internally, inform customers, and remediate the issue. This creates bottlenecks and slows down mean time to resolution (MTTR).

Datadog App Builder—now generally available—is a low-code solution that enables you to create apps in Datadog that facilitate collaboration and direct action by teams throughout your organization. These apps integrate natively into Datadog’s monitoring platform, providing interactive visibility into your telemetry data. Apps are easy to build using drag-and-drop UI components, Datadog monitoring tools and data sources, custom HTTP requests and JavaScript code, and 550+ out-of-the-box actions for popular platforms and a host of AWS, Azure, and Google Cloud services.

In this post, we’ll show how you can use a custom runbook app built with App Builder to accelerate remediation by centralizing context and action into one unified view in Datadog.

The Checkout Service Runbook app in the App Builder app editor.
The Checkout Service Runbook app in the App Builder app editor.
The Checkout Service Runbook app in the App Builder app editor.

App Builder not only enables you to create apps that identify issues in your systems but also to take direct, immediate action to remediate those issues from the same view within Datadog. Let’s look at an example of a runbook app that enables you to troubleshoot and take action to remediate an incident.

Let’s say your team owns the checkout service for an ecommerce website, and uses Datadog, GitLab, Atlassian Opsgenie, and Atlassian Statuspage. A monitor notifies you that the service is experiencing an outage. By reviewing telemetry data on the service health dashboard, you discover that the root cause is a bug introduced by a broken commit. To remediate, you decide to rollback the commit and retry the GitLab deployment. To start this process, you open the Checkout Service Runbook app, which was linked directly as a runbook on the service page of Service Catalog for easy access. The app enables you to trigger an incident, update the public status page, roll back the commit, and kick off a new build in the pipeline—all without leaving the Datadog platform.

Your runbook app is organized into four sections, with each section focused on a specific task. In the first section, a custom snapshot of your service appears with the most relevant monitors to get a quick summary of its health.

The first section of the runbook app showing the checkout service health status.
The first section of the runbook app showing the checkout service health status.
The first section of the runbook app showing the checkout service health status.

Scrolling down to the second section, you can immediately create an incident directly from within the app by clicking the “Create New Incident” button. You’re also given useful context from Opsgenie, including the current on-call engineer, the number of active incidents on this service, and any past incidents associated with it. If you need to page the on-call engineer, you can click the “Page on-call” button and do that directly from the app as well.

Now that you’ve created the incident, your next priority is to inform your customers about the outage. To update the public status page on Atlassian Statuspage, you simply click “Modify status,” which provides a modal where you can adjust the incident impact level, status, name, and description before posting the update.

With your customers informed, you turn your attention to fixing the issue. The runbook app pulls data directly from GitLab, so you can easily review the pipelines and commits associated with each of the most recent deployments, then rollback the problematic commit with just a few clicks.

Finally, it’s time to update your customers and team about the remediation steps taken. You scroll back up to the incident management and status page sections of the app and change the statuses to “In Progress” and “Identified” respectively, provide an explanation of your solution, and indicate the expected turnaround time for the systems to go back to normal.

How we built this app

To build this app, we added actions from the Actions Catalog that communicate with third-party services. In this case, we used App Builder’s Opsgenie and GitLab actions, as well as custom HTTP actions to interact with public APIs for services that don’t have out-of-the-box App Builder integrations (including listing Statuspage updates or retrying a deployment in GitLab). We also used some additional JavaScript in the post-query transform to format data, such as dates, easily.

This is just one example of the many types of apps that you can build with App Builder. Others include database consoles to empower customer support teams to solve a wider range of issues on their own, custom analyses to extract net new insights, custom visualizations of your company’s heath for leadership teams, and portals for developers to easily spin up new GitHub projects on their own.

Build self-service tools to streamline DevOps processes

By helping DevOps teams not only collect and analyze monitoring data but also perform remediations, Datadog App Builder expands the scope of what Datadog can do. By enabling you to pull in data from and interact with key Datadog products as well as many integrated services and platforms, App Builder can help personnel across your organization take action on their monitoring insights. App Builder can improve teams’ productivity, facilitate smoother collaboration, and limit the cost of outages.

App Builder is now generally available for all Datadog customers. For more information, see our App Builder documentation. If you’re brand new to Datadog, sign up for a free trial.