惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

有赞技术团队
有赞技术团队
M
MIT News - Artificial intelligence
Hugging Face - Blog
Hugging Face - Blog
博客园 - 聂微东
量子位
S
SegmentFault 最新的问题
V
Visual Studio Blog
博客园 - 【当耐特】
Apple Machine Learning Research
Apple Machine Learning Research
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
小众软件
小众软件
Stack Overflow Blog
Stack Overflow Blog
Vercel News
Vercel News
D
Docker
J
Java Code Geeks
博客园 - 三生石上(FineUI控件)
博客园 - Franky
Recent Announcements
Recent Announcements
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
MongoDB | Blog
MongoDB | Blog
D
DataBreaches.Net
Y
Y Combinator Blog
云风的 BLOG
云风的 BLOG
V
V2EX

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis
Monitor your NVIDIA Jetson IoT devices with Datadog
2021-01-14 · via Datadog | The Monitor blog

NVIDIA Jetson is a family of embedded, low-power computing boards designed to support machine learning and AI applications at the edge. Organizations use Jetson boards for complex video and image processing and analysis, automating build processes in factories, and improving city infrastructures. For example, Jetson-based devices enable cities to analyze traffic patterns with their existing traffic cameras in order to find ways to improve their most congested intersections.

To help you monitor your fleet of Jetson devices, the Datadog IoT Agent now supports the current portfolio of Jetson boards, giving you even more visibility into your IoT environments. Datadog captures critical performance metrics from your Jetson hardware, including GPU utilization and frequency (i.e., GR3D), the amount of memory dedicated to the GPU (i.e., IRAM utilization), and external memory controller utilization (i.e., EMC). In addition to Jetson metrics, the IoT Agent automatically collects standard system metrics for CPU, memory, and network I/O, giving you deeper insight into what is happening on each of your devices. You can view all of these metrics in Datadog’s out-of-the-box Jetson dashboard, so you can get a high-level overview of your fleet.

Visualize NVIDIA Jetson metrics with a built-in dashboard

End-to-end visibility into your device network

IoT networks can be a large and complex web of hundreds or thousands of devices, making it difficult to see how they connect to and support your services. Visibility into your entire network is important for quickly pinpointing issues such as a device that is performing poorly or unexpectedly goes offline, which can cause disruptions for your teams, their services, and your customers.

Datadog provides full visibility into your IoT network, so you can make informed decisions on how to maintain all of your devices. You can tag your devices with identifiers such as their geographic location to easily compare their performance using Datadog’s Host Map. For example, you can visualize your fleet’s GPU utilization across multiple locations in order to identify which devices need an upgrade in order to keep up with a service’s processing demand. This information can be invaluable for machine learning and computer vision use cases, where developers need to know how much their models are taxing the device.

You can also use Datadog to proactively monitor your network with alerts that automatically notify you when a device goes offline, or when there are unusual drops in a device’s GPU utilization.

NVIDIA Jetson outlier alert

As seen in the example above, alert notifications can be customized to include device-specific tags, so you know exactly which devices in your fleet were affected and how to fix them.

Monitor the performance of your resource-intensive workflows

Jetson devices are ideal for processing video and image data. Since these types of processing jobs are resource intensive, it’s important that you have visibility into the health and performance of each of your devices to ensure they continue supporting your overall workflows.

With Datadog, you can monitor critical resource metrics for your devices, such as how much memory is allocated to a device and how much it is utilizing. This can help you determine if a device is reaching its limits for executing a complex processing job and needs to be upgraded. Datadog can also help you monitor the state of your devices after you’ve deployed an update to their software (e.g., video analytics or automation software).

Overlay events with NVIDIA Jetson metrics

As seen in the example above, you can overlay events on a graph in order to track how specific events like a software update might have affected key device metrics (e.g., IRAM, EMC). A spike in IRAM metrics after an update, for example, could be an indicator that the update is consuming too much memory and needs to be rolled back.

As with tracking a device’s memory usage, monitoring its power consumption can also ensure that a fleet is optimized to support your services, and that its power usage stays within your team’s energy budget. With Datadog, you can quickly identify the source of unusual changes in a device’s energy usage.

Power usage by NVIDIA Jetson device

In the example above, you can see spikes in average power usage for several devices, which could be due to a resource-intensive processing job or inefficient hardware. You can quickly pivot to related logs to troubleshoot further and determine if you need to improve a processing workflow or schedule hardware upgrades (e.g., a new battery or radio) for the affected devices, ensuring that your fleet is working optimally.

Meet the Jetsons

The NVIDIA Jetson family of devices powers applications and robotics critical for improving manufacturing and shipping workflows and city infrastructures, use cases that require end-to-end visibility into device performance. With Datadog, you can monitor all of your Jetson-powered IoT devices and seamlessly correlate hardware metrics with other infrastructure metrics to ensure your devices and the systems they support are performing optimally. Check out our documentation to learn more. If you don’t already have a Datadog account, you can get started with a free 14-day trial today.