惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
N
Netflix TechBlog - Medium
博客园 - 叶小钗
大猫的无限游戏
大猫的无限游戏
S
SegmentFault 最新的问题
V
V2EX
IT之家
IT之家
J
Java Code Geeks
Hacker News - Newest:
Hacker News - Newest: "LLM"
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
GbyAI
GbyAI
D
Docker
S
Secure Thoughts
Recent Announcements
Recent Announcements
Webroot Blog
Webroot Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
云风的 BLOG
云风的 BLOG
博客园_首页
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Security Archives - TechRepublic
Security Archives - TechRepublic
酷 壳 – CoolShell
酷 壳 – CoolShell
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
N
News | PayPal Newsroom
S
Security @ Cisco Blogs
I
InfoQ
Last Week in AI
Last Week in AI
SecWiki News
SecWiki News
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
W
WeLiveSecurity
T
Troy Hunt's Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Attack and Defense Labs
Attack and Defense Labs
美团技术团队
T
The Blog of Author Tim Ferriss
Google DeepMind News
Google DeepMind News
Martin Fowler
Martin Fowler
B
Blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
Scott Helme
Scott Helme
T
Tor Project blog
Know Your Adversary
Know Your Adversary
有赞技术团队
有赞技术团队
Hugging Face - Blog
Hugging Face - Blog
Recorded Future
Recorded Future
C
Cyber Attacks, Cyber Crime and Cyber Security
AI
AI
G
Google Developers Blog

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis Monitor Aruba Central in Datadog How we centralize and remediate risks with Datadog Case Management Accelerate incident response with Datadog and ServiceNow Monitor your application and network load balancer logs Understanding Karpenter architecture for Kubernetes autoscaling Tools for collecting metrics and logs from Karpenter Monitor Karpenter with Datadog What your product data is actually saying Key metrics for monitoring Karpenter Securing Datadog’s platform in the AI age: The role of observability data Four ways engineering teams use the Datadog MCP Server to power AI agents Approaching your observability migration with the right mindset Meet the new Bits AI SRE: Deeper reasoning, twice as fast Key learnings from the 2026 State of DevSecOps study Use plain English to query your multi-cloud infrastructure in Resource Catalog Simplifying troubleshooting across the user journey with Datadog Synthetic Monitoring Protect your OCI resources with Datadog Cloud Security This Month in Datadog - February 2026 Amazon EC2 security: How misconfigured and public AMIs expand your cloud attack surface Enable end-to-end visibility into your Java apps with a single command Measure and improve mobile app startup performance with Datadog RUM Evaluating our AI Guard application to improve quality and control cost Identify untested code across every level of your codebase Make use of guardrail metrics and stop babysitting your releases Monitor Versa Networks SD-WAN performance in Datadog Improve performance and reliability with APM Recommendations Remediate transitive vulnerabilities faster with Datadog Software Composition Analysis Generate audit-ready vulnerability and compliance reports with Datadog Sheets Monitor Fortinet FortiManager performance in Datadog Improve test coverage across codebases with Datadog Code Coverage Move fast, don’t break things: Consistent testing standards at scale Enrich logs with ServiceNow CMDB context before routing to any SIEM or logging tool Monitor Lustre with Datadog Make faster, better product decisions with Datadog Product Analytics Surface and remediate runtime posture issues with Workload Protection Findings Protect agentic AI applications with Datadog AI Guard How to optimize JavaScript code with CSS Trace Google Pub/Sub workloads in Cloud Run with Datadog Detect human names in logs with ML in Sensitive Data Scanner How we cut our NLQ agent debugging time from hours to minutes with LLM Observability Debug PostgreSQL query latency faster with EXPLAIN ANALYZE in Datadog Database Monitoring Datadog acquires Propolis Unify and correlate frontend and backend data with retention filters Scale compliance across global frameworks with Datadog Cloud Security Monitor Arista VeloCloud SD-WAN performance with Datadog Building reliable dashboard agents with Datadog LLM Observability Simplify log collection and aggregation for MSSPs with Datadog Observability Pipelines Mitigation for Node.js denial-of-service vulnerability affecting Datadog APM Automate flaky test fixes with the Bits AI Dev Agent and Test Optimization How we built an AI SRE agent that investigates like a team of engineers Datadog integrations 2025 recap: Observability for AI, security, and hybrid cloud Design effective executive dashboards with Datadog Implement dbt data quality checks with dbt-expectations Bring faster visibility into AWS Lambda functions with remote instrumentation Troubleshoot faster with the GitLab Source Code integration in Datadog How Cambia Health Solutions saved $30,000 monthly with Cloud Cost Management and the Datadog Resource Catalog Normalize any logs for Cloud SIEM with Datadog's OCSF processor Optimizing Datadog at scale: Cost-efficient observability at Zendesk Detect, diagnose, and resolve network issues easily with CNM Network Health Connect engineering errors to user impact in early-stage products Cilium configuration for Kubernetes operations at scale Designing feedback loops for progressive delivery Ship features faster and safer with Datadog Feature Flags Choosing the right OpenTelemetry Collector distribution Route your monitor alerts with Datadog monitor notification rules Automate Cloud SIEM investigations with Bits AI Security Analyst Cloud threat detection: How to identify risky activity across control and data planes Collecting Kafka performance metrics Monitoring Kafka with Datadog Monitoring Kafka performance metrics
Increase visibility into your infrastructure processes with Process Tag Rules
2025-03-27 · via Datadog | The Monitor blog

Monitoring the health of your infrastructure and services requires you to understand the performance of fundamental system processes. But particularly in large environments, the sheer volume of processes can make their performance and resource usage difficult to track, let alone troubleshoot. To further complicate matters, it can be challenging to gather the necessary context on your processes to effectively categorize and monitor them—historically, it’s only been possible to identify processes using tags that come from the host they’re running on, either directly or through elements of the associated cloud infrastructure, such as containers, pods, and other resources. An individual process’s command line contains information on its associated groups, services, and configuration, but this metadata is not always parsed as facets or metrics in your observability platform.

Datadog Live Processes now includes Process Tag Rules, which enable you to enrich your system’s end-to-end visibility by deriving tags from a process’s command line, and to create performance metrics and alerts specific to your processes’ tags.

In this blog, we’ll explore how Processes Tag Rules help enhance troubleshooting and monitoring your host processes by enabling you to:

Quickly find common issues with groups of processes

A process or a group of processes can represent vital application- and system-level operations. However, it can be difficult to obtain visibility into the data that will tell you how to group your processes for more effective monitoring.

Process Tag Rules enable you to discover what processes are related to each other using information embedded in the command line configuration of your processes. By creating these tags, you can monitor performance and resource utilization for groups of related processes, giving you visibility into system-wide performance to help you troubleshoot issues.

You can create a new Process Tag Rule from the Manage Process Tags tab in Datadog Live Processes.

Manage Process Tags tab in Datadog Live Processes

After you click on New Process Tag Rule, you can configure your rule, starting with name, scope, and expression. Rule expressions are where you define your tags based on the command line configuration of your processes, using the standard Grok language syntax.

For example, you can configure a Process Tag Rule that matches the command line for tini processes in a particular environment and extracts the command line flags as tags. You can filter your processes down to command:tini and a specific cluster_name to see how many processes your rule will apply to. The next step is to create an expression for the rule—in the screenshot below, this is defined as tini -- %{notSpace:tini_start_process_name} %{notSpace:tini_sub_process_name}. This rule will capture the two tokens following tini -- as new tags that define the facets tini_start_process_name and tini_sub_process_name.

Example process tag rule in Datadog Live Processes

You can test your new rule by providing a sample of a command line, such as tini -- bash startstuff.sh. If the rule is validated, you’ll see example tags from your command sample, so you can make sure these new tags meet your expectations.

Test your process tag rule expression in the Datadog UI

After you click Create Rule to save the new Process Tag Rule, it will be applied to all matching processes, and your new tags will be available to search from the Processes view. To view process groups with these newly generated tags, you can group your processes by the tag group you just created—in our case, tini_start_process_name.

List of processes based on a newly created process tag rule

You can also use Process Tag Rules to group processes by service in certain scenarios. For example, with the shell svchost Windows system process, the -s command line option will load the specific service that the process is running on. With a Process Tag Rule, you can define this value to be extracted as a tag.

Monitor processes running on specific services with Process Tag Rules

Track system health by creating metrics and monitors on your custom process groups

Creating new tags also enables you to create metrics based on the process groups you define. You can create a new metric right from the Manage Metrics tab by clicking New Metric.

Create a new metric in Datadog Live Processes

From there, you can provide the metric dimensions you want to collect for the process groups you’re interested in.

Customize the dimensions for new process metrics

For more information on configuring Process Metrics, check out our documentation.

You can also create Live Process Monitors scoped to your rule-based tags—simply choose the New Monitor option from the New Metric dropdown in your Processes view.

Create a new process monitor in Datadog Live Processes

From here, you can create a Live Process Monitor that will show you whether processes tagged based on your custom rules are running. For example, using the svchost_service tagging rule we created early, you can create a monitor on the svchost_service:WinRM process group that ensures that Windows Remote Management is running across your fleet of Windows hosts.

Customize the dimensions of your new process monitor using process tag rules

You can also set metric monitors based on the metrics you create for your custom process groups. For example, you may want to be alerted when certain groups of processes reach specific thresholds for CPU and RSS memory usage, as high values for these metrics might signal that adjustments are required for process replicas running for a particular service.

Scope metric monitors to focus on processes grouped by your custom tags

Enhance your process-level observability

Datadog’s Process Tag Rules expand your organization’s ability to troubleshoot and monitor your process-level data. By proactively monitoring and addressing these issues early, you can safeguard mission-critical services, optimize resource allocation, and build resilience into your systems.

Process Tag Rules are available if you are a Datadog customer. If you’re not already using Datadog, you can start today with a 14-day free trial.