惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
酷 壳 – CoolShell
酷 壳 – CoolShell
宝玉的分享
宝玉的分享
IT之家
IT之家
Recent Announcements
Recent Announcements
T
The Blog of Author Tim Ferriss
雷峰网
雷峰网
阮一峰的网络日志
阮一峰的网络日志
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
M
MIT News - Artificial intelligence
D
Docker
C
CERT Recently Published Vulnerability Notes
月光博客
月光博客
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Recorded Future
Recorded Future
博客园 - 司徒正美
D
DataBreaches.Net
Last Week in AI
Last Week in AI
U
Unit 42
人人都是产品经理
人人都是产品经理
博客园_首页
Blog — PlanetScale
Blog — PlanetScale
量子位
大猫的无限游戏
大猫的无限游戏
博客园 - Franky
T
Tailwind CSS Blog
小众软件
小众软件
Y
Y Combinator Blog
WordPress大学
WordPress大学
B
Blog RSS Feed
C
Check Point Blog
H
Help Net Security
The Last Watchdog
The Last Watchdog
F
Full Disclosure
腾讯CDC
V
Visual Studio Blog
Google Online Security Blog
Google Online Security Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Troy Hunt's Blog
N
News and Events Feed by Topic
F
Fortinet All Blogs
B
Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
J
Java Code Geeks
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
有赞技术团队
有赞技术团队
博客园 - 三生石上(FineUI控件)
TaoSecurity Blog
TaoSecurity Blog
I
InfoQ
V
Vulnerabilities – Threatpost

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis Monitor Aruba Central in Datadog How we centralize and remediate risks with Datadog Case Management Accelerate incident response with Datadog and ServiceNow Monitor your application and network load balancer logs Understanding Karpenter architecture for Kubernetes autoscaling Tools for collecting metrics and logs from Karpenter Monitor Karpenter with Datadog What your product data is actually saying Key metrics for monitoring Karpenter Securing Datadog’s platform in the AI age: The role of observability data Four ways engineering teams use the Datadog MCP Server to power AI agents Approaching your observability migration with the right mindset Meet the new Bits AI SRE: Deeper reasoning, twice as fast Key learnings from the 2026 State of DevSecOps study Use plain English to query your multi-cloud infrastructure in Resource Catalog Simplifying troubleshooting across the user journey with Datadog Synthetic Monitoring Protect your OCI resources with Datadog Cloud Security This Month in Datadog - February 2026 Amazon EC2 security: How misconfigured and public AMIs expand your cloud attack surface Enable end-to-end visibility into your Java apps with a single command Measure and improve mobile app startup performance with Datadog RUM Evaluating our AI Guard application to improve quality and control cost Identify untested code across every level of your codebase Make use of guardrail metrics and stop babysitting your releases Monitor Versa Networks SD-WAN performance in Datadog Improve performance and reliability with APM Recommendations Remediate transitive vulnerabilities faster with Datadog Software Composition Analysis Generate audit-ready vulnerability and compliance reports with Datadog Sheets Monitor Fortinet FortiManager performance in Datadog Improve test coverage across codebases with Datadog Code Coverage Move fast, don’t break things: Consistent testing standards at scale Enrich logs with ServiceNow CMDB context before routing to any SIEM or logging tool Monitor Lustre with Datadog Make faster, better product decisions with Datadog Product Analytics Surface and remediate runtime posture issues with Workload Protection Findings Protect agentic AI applications with Datadog AI Guard How to optimize JavaScript code with CSS Trace Google Pub/Sub workloads in Cloud Run with Datadog Detect human names in logs with ML in Sensitive Data Scanner How we cut our NLQ agent debugging time from hours to minutes with LLM Observability Debug PostgreSQL query latency faster with EXPLAIN ANALYZE in Datadog Database Monitoring Datadog acquires Propolis Unify and correlate frontend and backend data with retention filters Scale compliance across global frameworks with Datadog Cloud Security Monitor Arista VeloCloud SD-WAN performance with Datadog Building reliable dashboard agents with Datadog LLM Observability Simplify log collection and aggregation for MSSPs with Datadog Observability Pipelines Mitigation for Node.js denial-of-service vulnerability affecting Datadog APM Automate flaky test fixes with the Bits AI Dev Agent and Test Optimization How we built an AI SRE agent that investigates like a team of engineers Datadog integrations 2025 recap: Observability for AI, security, and hybrid cloud Design effective executive dashboards with Datadog Implement dbt data quality checks with dbt-expectations Bring faster visibility into AWS Lambda functions with remote instrumentation Troubleshoot faster with the GitLab Source Code integration in Datadog How Cambia Health Solutions saved $30,000 monthly with Cloud Cost Management and the Datadog Resource Catalog Normalize any logs for Cloud SIEM with Datadog's OCSF processor Optimizing Datadog at scale: Cost-efficient observability at Zendesk Detect, diagnose, and resolve network issues easily with CNM Network Health Connect engineering errors to user impact in early-stage products Cilium configuration for Kubernetes operations at scale Designing feedback loops for progressive delivery Ship features faster and safer with Datadog Feature Flags Choosing the right OpenTelemetry Collector distribution Route your monitor alerts with Datadog monitor notification rules Automate Cloud SIEM investigations with Bits AI Security Analyst Cloud threat detection: How to identify risky activity across control and data planes Collecting Kafka performance metrics Monitoring Kafka with Datadog Monitoring Kafka performance metrics
Monitor your Argo CD clusters with Datadog
Addie Beach, Shri Subramanian · 2023-02-10 · via Datadog | The Monitor blog
Addie Beach

Addie Beach

Technical Content Writer

Shri Subramanian

Shri Subramanian

Argo CD is a declarative continuous delivery tool for Kubernetes developed by the Cloud Native Computing Foundation (CNCF). Argo CD automates your application deployment by continuously monitoring the live state of your containers and comparing it against the desired state in your Kubernetes manifest files, then pulling changes into your Kubernetes clusters as needed. Because your cluster setup is sourced directly from the manifest files in your application’s repository, you can track any changes made to your container infrastructure via Git, making Argo CD GitOps-compatible. And with an easy-to-use interface, you can quickly check whether all your containers are in sync with your repository at any time.

The Datadog Argo CD integration helps you ensure that your Kubernetes cluster is up to date with your latest manifest files via metrics from every Argo CD component. By using the Argo CD out-of-the-box (OOTB) dashboard, you can monitor how quickly and accurately your infrastructure changes are being applied to your cluster. Plus, you can easily leverage preconfigured monitors for key Argo CD metrics to notify you of any sync issues. In this post, we’ll explain how the Argo CD integration can help you:

  • Visualize activity across your Argo CD clusters

  • Troubleshoot with metrics from every Argo CD component

  • Quickly detect application sync issues with Argo CD monitors

The out-of-the-box dashboard for Argo CD, with an overview of hosts running Argo CD and metrics for each of the components.

Visualize activity across your Argo CD clusters

There are three main components to Argo CD: the repository server, application controller, and API server. Each component handles a different part of the sync process: the repository server maintains a local cache of your manifest file, the application controller compares the manifest file against the current state of your cluster to look for changes, and the API server provides endpoints for the deployments and rollbacks needed to reconcile these changes. To ensure your Argo CD clusters are able to stay in sync with your manifest file, you need to monitor all three components.

The Argo CD OOTB dashboard helps you visualize metrics for each component, allowing you to see how well Argo CD is managing your application deployments and syncs. You can access information such as how many app syncs are occurring and how many are successful or failing, giving you granular visibility into your container configuration. With this data, you can quickly spot any deployment issues that could be affecting your Kubernetes clusters and syncs. Some metrics on the dashboard that can give you insight into the state of your cluster include:

  • argocd.app_controller.app.sync.count: the total number of application syncs

  • argocd.app_controller.app.info: how many applications are not in sync, categorized by host and status

  • argocd.api_server.grpc.server.handled.count: the total number of requests for each service

  • argocd.repo_service.git.request.duration.seconds.bucket: the performance of Git fetch requests

In addition to performance data from the main three components, you can also access Argo CD logs, utilization metrics, and Kubernetes cluster stats to help you with troubleshooting. Additionally, the dashboard comes with template variables that help you easily drill down into specific cluster attributes, such as namespace, health status, and repository. You can use these attributes to narrow the scope of your investigation based on the hosts you want to target. For example, health status enables you to target hosts that are up to date, currently progressing through a sync, or whose status is missing or unknown.

Troubleshoot with metrics from every Argo CD component

You can also use the Argo CD dashboard to help you dig deeper into your incident investigations. Let’s say you receive an error while attempting to deploy a change to your Kubernetes configuration. You can pivot to the Argo CD dashboard to see whether there seem to be any sync issues in your cluster. Via the argocd.app_controller.app.info metric, you’re able to see that a high number of applications are not in sync.

Graph of applications not in sync by status and host, with the number of applications increasing over time.

This metric also shows you the specific hosts experiencing this issue. By jumping to the logs widget on the dashboard, you see a high number of error messages from these hosts with additional details about the issue, such as the relevant failure codes and the time that the hosts started failing. You can then view these hosts in Datadog Container Monitoring to pinpoint the problem, whether it’s syntax errors in the YAML manifest file or an overloaded resource.

Quickly detect application sync issues with Argo CD monitors

In order to catch issues in your Kubernetes deployments even faster, you can set up Argo CD monitors to notify you of sync issues. The Argo CD integration includes a recommended, preconfigured monitor that alerts you to any app sync failures by filtering the argocd.app_controller.app.info metric to unsuccessful syncs. It also includes a default sync failure threshold and 30-minute interval for sync status checks, so you can be notified as soon as it’s clear that there’s a meaningful issue without getting distracted by false alarms.

The default recommended monitor for the Argo CD integration.

Let’s say your monitor alerts you to an out-of-sync cluster. By pivoting to the Argo CD dashboard, you can immediately see that the Argo CD API server has exceeded the amount of available CPU on its host, making it difficult to process new requests. You can then use the dashboard to investigate whether there’s an unusual spike in requests due to a one-off event or whether this is an ongoing issue indicating that you may need to allocate more resources to your servers.

Start monitoring Argo CD performance with Datadog today

Argo CD can help you easily manage your Kubernetes clusters while maintaining a single source of truth via Git. With the Datadog integration, you can quickly detect application sync failures that could cause your cluster configuration to drift from your application deployments. The OOTB dashboard enables you to visualize activity throughout your entire cluster, while the recommended Argo CD monitor helps you catch meaningful sync issues fast.

To get started, you can set up the Argo CD integration using our documentation. Or, if you’re not yet a customer, you can sign up for a 14-day free trial today.