惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

大猫的无限游戏
大猫的无限游戏
H
Hackread – Cybersecurity News, Data Breaches, AI and More
博客园_首页
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
D
Docker
酷 壳 – CoolShell
酷 壳 – CoolShell
宝玉的分享
宝玉的分享
Martin Fowler
Martin Fowler
美团技术团队
量子位
M
MIT News - Artificial intelligence
Apple Machine Learning Research
Apple Machine Learning Research
阮一峰的网络日志
阮一峰的网络日志
博客园 - 叶小钗
博客园 - 三生石上(FineUI控件)
腾讯CDC
Hugging Face - Blog
Hugging Face - Blog
博客园 - 【当耐特】
小众软件
小众软件
博客园 - 司徒正美
罗磊的独立博客
云风的 BLOG
云风的 BLOG
B
Blog RSS Feed
博客园 - 聂微东

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis
Monitor model performance with Superwise’s offering in th...
Bowen Chen · 2022-05-09 · via Datadog | The Monitor blog
Bowen Chen

Bowen Chen

Superwise is a monitoring platform that provides model observability for high-scale machine learning (ML) operations. Superwise provides teams with out-of-the-box (OOTB) metrics on their models’ production behavior, so they can effectively address drift, data quality issues, and other problems before they negatively impact the business.

Datadog’s Superwise integration automatically generates metrics based on the data entities that are specific to your model. Once you’ve set up the integration, your metrics will begin flowing into an OOTB dashboard in Datadog, so you can immediately visualize trends without any manual configuration. In this post, we’ll cover how to gain visibility into your model activity and drift with dashboards and how to configure incident monitoring with Superwise’s policies.

Use the out-of-the-box Superwise dashboard to monitor model metrics and incidents.

Monitor model drift and other issues

A model is only as accurate as the data set it was trained on, which makes it essential for your production data to closely mirror your model baseline. If your data distribution begins to drift from your baseline, your model may reflect these changes in its prediction accuracy and recall. Superwise’s flexible policy builder offers a wide range of templates that can help you monitor model drift and other issues. You can create customized policies based on a wide range of Superwise metrics, such as changes in your model’s performance, data quality, activity, drift, and your own performance metrics and business KPIs.

Configure policies using Superwise’s flexible policy template builder.

Once you’ve configured a policy, Superwise will scan for anomalies within the chosen logic (policies allow for any and/or segment, metric, and feature combination), allowing you to dynamically monitor your model without having to manually configure thresholds. Performance degradation and prediction shift policies allow you to detect when your model falls below production-standard levels. You can also configure missing value and outlier policies to monitor the quality of your distribution data.

The OOTB Superwise dashboard provides a quick overview of the number of active models, their drift, and total predictions over time. With these metrics, you can determine when high-impact models need to be retrained to improve their accuracy.

Visualize metrics such as model drift using dashboard widgets.

For a more flexible view, you can clone and customize the dashboard by adding or removing widgets to showcase metrics of your choice. For example, you can customize your dashboard to show overall input drift from a specific model or use it to monitor your models alongside other integrated services your infrastructure depends on.

Investigate Superwise incidents with Datadog

If Superwise detects a violation of any of your configured policies, it will automatically create an incident. When correlated policy violations are detected, Superwise will aggregate them into a single incident to reduce noise and provide a focused view into your model’s issues. You have full control over which incidents get sent to Datadog to ensure that each incident reaches the appropriate channel. To keep track of Superwise incidents in Datadog, the OOTB dashboard widgets display the number of models with ongoing incidents, their incident distribution, and allow you to home in on individual incidents for more details.

Once you’ve configured Datadog as a notification channel for Superwise incidents, they will begin to appear in Datadog Incident Management. With cross-platform visibility, your team can triage and analyze the downstream impact of model issues by correlating Superwise metrics with other data from your environment.

Configure Superwise to send incidents to Datadog Incident Management.

Integrated model visibility and insights

With Datadog’s Superwise integration, it is easier than ever to monitor your ML models at enterprise scale. You can now track trends in your models’ performance, data quality, and drift straight from Datadog, alongside the other services your infrastructure depends on. To get started, install the Superwise integration and sign up for a Superwise subscription from the Datadog Marketplace. If you’re not already a Datadog customer, sign up today with a free 14-day trial.

The ability to promote branded marketing tools is a membership benefit offered through the Datadog Partner Network. You can learn more about the Datadog Marketplace in this blog post. If you’re interested in developing an integration or application that you’d like to promote, you can contact us at marketplace@datadog.com.