惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

aimingoo的专栏
aimingoo的专栏
Y
Y Combinator Blog
云风的 BLOG
云风的 BLOG
Microsoft Azure Blog
Microsoft Azure Blog
腾讯CDC
T
The Blog of Author Tim Ferriss
P
Proofpoint News Feed
Hugging Face - Blog
Hugging Face - Blog
博客园_首页
小众软件
小众软件
美团技术团队
Martin Fowler
Martin Fowler
爱范儿
爱范儿
有赞技术团队
有赞技术团队
博客园 - 【当耐特】
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Microsoft Security Blog
Microsoft Security Blog
宝玉的分享
宝玉的分享
J
Java Code Geeks
B
Blog
V
V2EX
Stack Overflow Blog
Stack Overflow Blog
B
Blog RSS Feed
博客园 - Franky

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis
Monitoring multi-cloud container storage with Portworx an...
2018-08-13 · via Datadog | The Monitor blog
Prashant Rathi

Prashant Rathi

This is a guest post by Prashant Rathi, Director of Product Management at Portworx.

About Portworx

Portworx provides solutions for Kubernetes storage as well as other leading container schedulers, dramatically reducing storage, compute, and infrastructure costs for running mission-critical, multi-cloud applications with zero downtime or data loss. With Portworx, you can manage any database or stateful service on any infrastructure using any container scheduler. Portworx is trusted by many of the world’s most sophisticated IT organizations including Comcast, GE, Lufthansa Systems, the U.S. Department of Homeland Security, and Verizon.

Portworx has CLI and UI tools to manage and monitor the health of a running cluster. For production use cases, Portworx also provides timeseries monitoring and log analytics that can be integrated with dedicated monitoring services for historical and infrastructure-wide context. We are pleased to announce a new integration between Portworx and Datadog, so you can correlate performance, throughput, and latency metrics from Portworx with data from infrastructure and application components to help pinpoint performance bottlenecks and provision resources appropriately.

Monitor the health of Portworx clusters and nodes

A custom dashboard for tracking the health of a Portworx cluster.
A Datadog dashboard tracking the health of a Portworx cluster
A custom dashboard for tracking the health of a Portworx cluster.

Customers deploy Portworx on hundreds of nodes and across multiple clusters. In these distributed environments, you need a traffic-light dashboard that helps you monitor each cluster’s health and resource usage. In Datadog, you can build a single dashboard to monitor all your clusters, and then use template variables to drill down to a specific cluster using built-in tags. For more focused troubleshooting, you can drill down to metrics from individual nodes in seconds.

With the new integration, Datadog collects cluster-level metrics from Portworx such as capacity usage, pending I/O, and more. You can use that data to set Datadog alerts for indicators like quorum or capacity used, enabling you to proactively prepare for maintenance events. With machine learning features like outlier detection, you can be notified automatically if a single node is behaving different than others.

Understand usage for capacity planning

As more workloads and users onboard, capacity planning becomes crucial for continued operations. Measuring overall usage against available resources and ranking nodes by usage are simple ways to track capacity. With the host map, Datadog provides a quick and intuitive way to segment usage and identify heavily utilized resources. In a cluster dashboard like the one pictured in the section above, you can use the overall size of the hexagons in a host map to denote the node’s capacity, and the color-coding to indicate the ratio of usage to capacity.

And as shown below, Datadog forecasts can help extrapolate usage trends into the future to trigger an alert when the usage is predicted to cross a predefined threshold (1.5 TB in this case) within a given interval.

Forecasting the resource usage of a Portworx cluster in Datadog

Monitor usage, latency, and I/O performance metrics in context

Cluster-wide monitoring is important for daily operations, but when it comes to troubleshooting performance issues, you need to go deeper. Application developers often seek to answer questions such as: Why is my application slow? What changed between yesterday and today? At this point, metrics from individual data paths become invaluable for connecting application performance to the underlying storage layer.

With Portworx volume metrics in Datadog, it is easy for developers to understand per-volume I/O, throughput, and latency. By visualizing these metrics in conjunction with other infrastructure and application performance data, you can create custom dashboards tailored to a specific scenario. For example, you can build a dashboard tracking service-level performance alongside metrics from the volumes used by the application, so you can see at a glance whether any slowdowns are due to capacity issues, or abnormally high CPU utilization or I/O on the app hosts.

An anomaly detection graph in Datadog tracks the p95 latency for a storage volume, with the expected bounds based on past performance overlaid as a gray band on the graph.
Anomaly detection in Datadog analyzes the volume latency of a Portworx node
An anomaly detection graph in Datadog tracks the p95 latency for a storage volume, with the expected bounds based on past performance overlaid as a gray band on the graph.

Additionally, anomaly detection in Datadog can compare performance data against expectations based on recurring patterns and trends, which is not possible with static threshold–based monitoring. By defining the evaluation window, the alerting and recovery threshold, and the allowable deviations from the prediction, you can monitor for anomalies in the 95th-percentile latency on any given volume, as shown above.

Get started

To start monitoring your Portworx clusters and nodes alongside the rest of your infrastructure and applications, check out the documentation on the new integration.