惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Security Archives - TechRepublic
Security Archives - TechRepublic
S
Secure Thoughts
V2EX - 技术
V2EX - 技术
Schneier on Security
Schneier on Security
Application and Cybersecurity Blog
Application and Cybersecurity Blog
L
LangChain Blog
博客园_首页
Jina AI
Jina AI
IT之家
IT之家
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
T
Tenable Blog
量子位
V
V2EX
酷 壳 – CoolShell
酷 壳 – CoolShell
S
Security Affairs
Last Week in AI
Last Week in AI
Scott Helme
Scott Helme
月光博客
月光博客
D
Darknet – Hacking Tools, Hacker News & Cyber Security
博客园 - 叶小钗
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
T
Tailwind CSS Blog
Simon Willison's Weblog
Simon Willison's Weblog
PCI Perspectives
PCI Perspectives
人人都是产品经理
人人都是产品经理
N
News and Events Feed by Topic
腾讯CDC
P
Proofpoint News Feed
T
The Exploit Database - CXSecurity.com
J
Java Code Geeks
博客园 - 司徒正美
博客园 - Franky
Latest news
Latest news
S
SegmentFault 最新的问题
小众软件
小众软件
博客园 - 三生石上(FineUI控件)
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
I
Intezer
Attack and Defense Labs
Attack and Defense Labs
H
Heimdal Security Blog
H
Hacker News: Front Page
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 【当耐特】
T
Troy Hunt's Blog
N
News | PayPal Newsroom
P
Palo Alto Networks Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Recent Commits to openclaw:main
Recent Commits to openclaw:main
雷峰网
雷峰网

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis Monitor Aruba Central in Datadog How we centralize and remediate risks with Datadog Case Management Accelerate incident response with Datadog and ServiceNow Monitor your application and network load balancer logs Understanding Karpenter architecture for Kubernetes autoscaling Tools for collecting metrics and logs from Karpenter Monitor Karpenter with Datadog What your product data is actually saying Key metrics for monitoring Karpenter Securing Datadog’s platform in the AI age: The role of observability data Four ways engineering teams use the Datadog MCP Server to power AI agents Approaching your observability migration with the right mindset Meet the new Bits AI SRE: Deeper reasoning, twice as fast Key learnings from the 2026 State of DevSecOps study Use plain English to query your multi-cloud infrastructure in Resource Catalog Simplifying troubleshooting across the user journey with Datadog Synthetic Monitoring Protect your OCI resources with Datadog Cloud Security This Month in Datadog - February 2026 Amazon EC2 security: How misconfigured and public AMIs expand your cloud attack surface Enable end-to-end visibility into your Java apps with a single command Measure and improve mobile app startup performance with Datadog RUM Evaluating our AI Guard application to improve quality and control cost Identify untested code across every level of your codebase Make use of guardrail metrics and stop babysitting your releases Monitor Versa Networks SD-WAN performance in Datadog Improve performance and reliability with APM Recommendations Remediate transitive vulnerabilities faster with Datadog Software Composition Analysis Generate audit-ready vulnerability and compliance reports with Datadog Sheets Monitor Fortinet FortiManager performance in Datadog Improve test coverage across codebases with Datadog Code Coverage Move fast, don’t break things: Consistent testing standards at scale Enrich logs with ServiceNow CMDB context before routing to any SIEM or logging tool Monitor Lustre with Datadog Make faster, better product decisions with Datadog Product Analytics Surface and remediate runtime posture issues with Workload Protection Findings Protect agentic AI applications with Datadog AI Guard How to optimize JavaScript code with CSS Trace Google Pub/Sub workloads in Cloud Run with Datadog Detect human names in logs with ML in Sensitive Data Scanner How we cut our NLQ agent debugging time from hours to minutes with LLM Observability Debug PostgreSQL query latency faster with EXPLAIN ANALYZE in Datadog Database Monitoring Datadog acquires Propolis Unify and correlate frontend and backend data with retention filters Scale compliance across global frameworks with Datadog Cloud Security Monitor Arista VeloCloud SD-WAN performance with Datadog Building reliable dashboard agents with Datadog LLM Observability Simplify log collection and aggregation for MSSPs with Datadog Observability Pipelines Mitigation for Node.js denial-of-service vulnerability affecting Datadog APM Automate flaky test fixes with the Bits AI Dev Agent and Test Optimization How we built an AI SRE agent that investigates like a team of engineers Datadog integrations 2025 recap: Observability for AI, security, and hybrid cloud Design effective executive dashboards with Datadog Implement dbt data quality checks with dbt-expectations Bring faster visibility into AWS Lambda functions with remote instrumentation Troubleshoot faster with the GitLab Source Code integration in Datadog How Cambia Health Solutions saved $30,000 monthly with Cloud Cost Management and the Datadog Resource Catalog Normalize any logs for Cloud SIEM with Datadog's OCSF processor Optimizing Datadog at scale: Cost-efficient observability at Zendesk Detect, diagnose, and resolve network issues easily with CNM Network Health Connect engineering errors to user impact in early-stage products Cilium configuration for Kubernetes operations at scale Designing feedback loops for progressive delivery Ship features faster and safer with Datadog Feature Flags Choosing the right OpenTelemetry Collector distribution Route your monitor alerts with Datadog monitor notification rules Automate Cloud SIEM investigations with Bits AI Security Analyst Cloud threat detection: How to identify risky activity across control and data planes Collecting Kafka performance metrics Monitoring Kafka with Datadog Monitoring Kafka performance metrics
A guide on scaling out your Kubernetes pods with the Watermark Pod Autoscaler
Alex Preston · 2024-11-12 · via Datadog | The Monitor blog

While overprovisioning Kubernetes workloads can provide stability during the launch of new products, it’s often only sustainable because large companies have substantial budgets and favorable deals with cloud providers. As highlighted in Datadog’s State of Cloud Costs report, cloud spending continues to grow, but a significant portion of that cost is often due to inefficiencies like overprovisioning. That’s why for most companies, proper resource provisioning and autoscaling play crucial roles in keeping costs down and helping the product respond to traffic spikes.

The Kubernetes built-in Horizontal Pod Autoscaler (HPA) scales based on a single target threshold. At Datadog, we built the open source Watermark Pod Autoscaler (WPA) controller to extend the HPA, allowing for more flexibility by introducing high and low watermarks for scaling decisions and more fine-grained control over autoscaling behavior.

In this post, we’ll take a look at how the WPA helps you keep your cloud spend under control by managing scaling policies and simulating actions before deployment. For a more detailed guide on scaling your Kubernetes workloads and clusters, check out our blog post.

But first, we’ll cover some important provisioning principles we learned while implementing the WPA.

Sizing your pods correctly

Vertical provisioning

High resource utilization is essential prior to setting up autoscaling. As we were setting up the WPA for one of our own internal query engines, we realized we had overprovisioned worker pods. Our investigation revealed these pods were only seeing 10 percent CPU utilization during peak hours, making it difficult to determine when the pods were actually under load.

In order to achieve high utilization, we gradually decreased the CPU cores per pod in 10 percent increments, closely monitoring performance to maintain high utilization without sacrificing latency. This process allowed us to reduce CPU cores per pod by 40 percent, leading to higher utilization and a savings of ~$30k per month prior to setting up the autoscaler.

Horizontal provisioning

Similar to vertical provisioning, horizontal provisioning should come before autoscaling and can be done incrementally. By gradually adjusting the number of pods and evaluating workload distribution, you can find the optimal balance between resource allocation and performance.

We ran numerous A/B tests in our data centers to determine how scaling down vertically or horizontally would affect our cluster performance. As an example, running one zonal cluster with 15 pods, one at 12, and one at 10 allowed us to compare latency changes. These A/B tests are invaluable for determining the optimal minimum and maximum replica counts that your autoscaler should adjust to. Typically, services reach a point of diminishing returns, where adding more pods no longer improves performance—such as reducing latency—but only increases costs. Similarly, these tests help identify the minimum number of replicas required before performance degradation occurs.

Redrive step functions directly from Datadog

This latency comparison offers two points to consider during the provisioning stage.

Redrive step functions directly from Datadog

The first is to identify your service level objectives (SLOs). For example, you may want the p99 latency to be less than five seconds or each pod to handle 600 requests per second. This can help you identify your range for a steady state and will guide your autoscaling configuration decisions for scaling delays and velocity (which we’ll talk about later).

If your service is very expensive and you want to squeeze out as much efficiency as possible, the next step may be to load test your service. Load testing shows you how quickly performance falls off during times of high stress, which is useful for calculating an efficiency ratio. In other words, how much traffic does an extra CPU core or gigabyte of memory buy you? This can also help determine a minimum and maximum bound for the number of pods in your cluster.

Lastly, prior to autoscaling it’s important to figure out your cloud spend to establish a baseline. We used internal metrics to determine our hourly price per pod, but Datadog Cloud Cost Management provides out-of-the-box tooling to determine any service’s costs.

Is autoscaling right for your workload?

The principles of vertical and horizontal provisioning apply to scaling as well. Vertical scaling increases resources (CPU, memory, etc.) for a single pod, while horizontal scaling adds more pods to distribute the workload across nodes. In this post, we’ll primarily focus on horizontal scaling, as it was the key requirement for our service. Not every workload benefits from autoscaling, so it’s important to determine whether it’s the right choice for your specific needs.

Before implementing autoscaling, ask the following:

  • Does your workload experience fluctuating or unpredictable traffic patterns? If traffic surges are frequent but short-lived, autoscaling can prevent overprovisioning and wasted resources.

  • Is your application sensitive to delays caused by scaling up? If immediate response times are critical, autoscaling may need to be paired with other optimizations such as preloaded caches, rate limiting, or scheduled scaling.

  • Are you consistently seeing high resource utilization (CPU, memory)? Low utilization may indicate overprovisioning, signaling that autoscaling could help reduce costs.

Below is an example chart that we used to determine autoscaling was the right choice for our workload. Notice the traffic surges during peak hours, followed by lower demand periods; this pattern made static provisioning inefficient and wasteful. This chart matches expected behavior, as customers use the product mostly during business hours and less so at nights and on weekends.

Redrive step functions directly from Datadog

Setting up the autoscaler

Choosing a metric

Selecting the right metric is fundamental for effective autoscaling. There are two main types of metrics to consider: container metrics, such as CPU, memory usage, and network load that directly reflect workload resource needs, and custom metrics, like queue length, request latency, or database transaction time for more precise control over specialized applications.

We found that some metrics took longer to show workload stress than others. For example, CPU utilization might show a traffic spike 15 seconds later than request queue size. We opted for using the length of our request queue as our scaling metric. It’s important to remember that scaling decisions involve cascading steps—the Datadog Agent has to pick up the new metric, intake it, and then spin up new pods—which can take minutes, so finding a proactive metric is ideal.

The WPA can bake in your scaling metric during deployment or reference an externally deployed metric. We referenced an external metric for more granular configurations—check out this post on autoscaling workloads to learn more about how to set up external metrics. While it’s possible to use multiple metrics for scaling, it is not recommended because it can introduce complexity and inefficiency. For example, relying on different metrics for scaling up and down can lead to conflicting scaling decisions, and external dependencies can distort metrics, making scaling less predictable. It’s generally more efficient to scale based on a single, well-defined metric, like CPU utilization, and adjust scaling parameters to control behavior.

From Terraform to production

Once you understand your metrics, it’s time to set up the autoscaler itself. The first key decisions to make are what kind of autoscaler you want to use—WPA or HPA—and how you want to scale—by cluster or data center. In general, the HPA is well-suited for predictable scaling needs while the WPA is better for handling erratic traffic patterns. If you need a wider range of thresholds and more granular control over your scaling configurations, the WPA is the recommended choice.

Depending on your infrastructure, you will need to decide whether to scale individual clusters or scale workloads across multiple data centers. The latter is beneficial for high availability and fault tolerance but can introduce more complexity in scaling decisions.

Next, you can jump into writing Terraform. Implementing the WPA in Terraform is easy and allows you to manage scaling policies in a consistent, templated manner. Terraform simplifies scaling configurations and enables you to manage autoscalers across environments.

apiVersion: datadoghq.com/v1alpha1

kind: WatermarkPodAutoscaler

metadata:

name: {{ template "workload.name" . }}

namespace: {{ $.Release.Namespace }}

spec:

scaleTargetRef:

apiVersion: apps/v1

kind: Deployment

name: {{ template "workload.name" . }}

minReplicas: {{ .Values.server.autoscaling.minReplicas }}

maxReplicas: {{ .Values.server.autoscaling.maxReplicas }}

tolerateZero: true # Tolerate our Datadog Metric being zero (queue length is zero)

downscaleForbiddenWindowSeconds: 1200 # 20min

upscaleForbiddenWindowSeconds: 60 # 1min

scaleDownLimitFactor: 10

scaleUpLimitFactor: 100

dryRun: {{ $.Values.server.autoscaling.dryRun }}

replicaScalingAbsoluteModulo: 1

metrics:

- external:

highWatermark: "10"

lowWatermark: "5"

metricName: "datadogmetric@{{ $.Release.Namespace }}:{{ printf "%s-queue-length" $.Values.trino_cluster }}"

metricSelector:

matchLabels:

app: {{ template "workload.name" . }}

datacenter: {{ $.Values.global.datacenter.datacenter }}

It’s recommended to start the WPA in dryRun mode, which lets you simulate scaling actions based on metrics without actually modifying your infrastructure. This is a useful feature for testing how your autoscaler will react to different conditions such as sudden traffic spikes, before rolling it out in production. This also provides adequate data to build out and tune dashboards, monitors, and other alerting systems. The WPA in dryRun mode emits wpa_controller_replicas_scaling_proposal and wpa_controller_replicas_scaling_effective metrics, which are useful for determining scaling velocity and how your metric affects the autoscaler. A sample dashboard shown below includes widgets we found useful:

Redrive step functions directly from Datadog

Prior to letting your autoscaler loose in production, it’s important to ensure that monitoring and alerting are in place. It’s essential to monitor your autoscaler’s performance to detect potential issues early, which includes setting up alerts for abnormal scaling behaviors or failures. You should also know when metrics stop emitting, as autoscalers might stop working as expected. In such cases, fallback configurations should be in place, such as relying on a minimum replica count.

Additionally, we found it useful to build out runbooks for turning on and off the autoscaler. This can be done with the kubectl CLI tool by turning dryRun back on. These commands can be strung together in a bash script to disable or enable autoscalers across all your clusters at once:

kubectl patch wpa <WPA_NAME> --type='json' -p='[{"op": "replace", "path": "/spec/dryRun", "value":true}]'

As a note, the WPA will not perform any scaling actions when metrics aren’t received. It is often recommended to build uptime monitors for whichever metric the autoscaler is using. Datadog Workflow Automation can also be a great remediation tool if the WPA or metrics start to fail. As an example, if your metric stops reporting data you can prompt an engineer and scale all pods up to the maxReplica until the metric starts reporting again.

While using dryRun metrics provides valuable insights into how the system would behave under load, these metrics alone aren’t a perfect substitute for having the WPA fully enabled. As long as you start with conservative configurations, it’s best to enable WPA and start tuning rather than using dryRun for weeks on end.

Tuning the autoscaler

After setting up your autoscaler, continuous tuning is necessary to optimize it. Tuning the autoscaler is an iterative process aimed at optimizing both cost efficiency and cluster performance. To achieve this, several key areas must be considered.

The first is scaling velocity, which refers to how quickly your system scales up or down. It’s important to maintain balance, as scaling too quickly may result in overprovisioning, while scaling too slowly can hinder performance. This is controlled by the scaleUpLimitFactor and scaleDownLimitFactor parameters. For more detailed information on how these parameters work in the scaling algorithm, refer to the official WPA GitHub page.

The second consideration is scaling speed. A common best practice is to scale up fast and scale down slowly, allowing the system to handle sudden load spikes quickly while scaling down conservatively to avoid oscillation (frequent scaling up and down) that can lead to instability and increased costs. Another crucial factor is tuning cooldown periods, which helps prevent overreacting to short-lived spikes in demand. After weeks of iterative tuning, we found the following configurations gave us a well-tuned autoscaler for our clusters. As a note, these numbers are specific to our service but can be a good starting point:

  • downscaleForbiddenWindowSeconds: 1200 # 20min

  • upscaleForbiddenWindowSeconds: 60 # 1min

  • scaleDownLimitFactor: 10

  • scaleUpLimitFactor: 100

A useful heuristic for assessing whether your autoscaler is functioning efficiently is to monitor CPU and memory utilization. The chosen metric should remain stable on average across your cluster or datacenter. For example, if your workload is optimal at 75 percent CPU utilization, the autoscaler should keep this value at a steady state without significant fluctuations.

Lastly, when it comes to tuning the autoscaler, changes may take significantly longer than you initially expect. We iterated through numerous configurations and tested different scaling velocities, cooldown periods, and resource limits to prevent performance degradation while still ensuring the system could handle traffic spikes. Each adjustment required extensive monitoring for days at a time to observe the effects over various traffic patterns and workloads, which prolonged the tuning process.

Test out the Watermark Pod Autoscaler in your workloads today

Autoscaling offers a powerful and flexible approach to managing Kubernetes workloads by dynamically scaling on real-time metrics. However, regardless of which horizontal autoscaling framework you choose (WPA or HPA), it is not a “set it and forget it” solution; it requires continuous monitoring and adjustments. As your product grows or your autoscaler becomes less responsive, revisiting and fine-tuning configurations—whether monthly or quarterly—is essential to ensure optimal performance.

For a fully managed experience, Datadog Kubernetes Autoscaling combines watermark-based horizontal scaling, continuous vertical scaling, and built-in monitoring to simplify your scaling strategies and setup time. Try it out with a 14-day free trial.