惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
V
V2EX
博客园_首页
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Recent Announcements
Recent Announcements
博客园 - 司徒正美
Microsoft Security Blog
Microsoft Security Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Latest news
Latest news
Vercel News
Vercel News
The Register - Security
The Register - Security
T
The Exploit Database - CXSecurity.com
S
Schneier on Security
N
Netflix TechBlog - Medium
WordPress大学
WordPress大学
小众软件
小众软件
L
Lohrmann on Cybersecurity
GbyAI
GbyAI
P
Privacy & Cybersecurity Law Blog
T
Tor Project blog
AWS News Blog
AWS News Blog
美团技术团队
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
K
Kaspersky official blog
B
Blog RSS Feed
G
Google Developers Blog
量子位
大猫的无限游戏
大猫的无限游戏
Google DeepMind News
Google DeepMind News
Scott Helme
Scott Helme
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
I
Intezer
雷峰网
雷峰网
Martin Fowler
Martin Fowler
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Blog — PlanetScale
Blog — PlanetScale
IT之家
IT之家
F
Full Disclosure
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - 【当耐特】
The Hacker News
The Hacker News
U
Unit 42
S
SegmentFault 最新的问题
I
InfoQ
aimingoo的专栏
aimingoo的专栏
Y
Y Combinator Blog
宝玉的分享
宝玉的分享
罗磊的独立博客
Spread Privacy
Spread Privacy
C
CERT Recently Published Vulnerability Notes

Datadog | The Monitor blog

Introducing our open source AI-native SAST Instrument and monitor Boomi integration flows with OpenTelemetry and Datadog Not all index scans are equal: How we cut query latency by over 99% Platform engineering metrics: What to measure and what to ignore Integrate Recorded Future threat intelligence with Datadog Cloud SIEM CI/CD security: threat modeling using a MITRE-style threat matrix CI/CD security: How to secure your GitHub ecosystem Ingress NGINX is EOL: A practical guide for migrating to Kubernetes Gateway API Operating agentic AI with Amazon Bedrock AgentCore and Datadog LLM Observability: Lessons from NTT DATA Introducing the Datadog Code Security MCP Capture and analyze custom heatmaps in Session Replay Understand session replays faster with AI summaries and smart chapters Monitor ClickHouse query performance with Datadog Database Monitoring How we designed empathetic alert sounds for on-call engineers Search and act across Datadog to resolve issues faster with Bits Assistant Measure the business impact of every product change with Datadog Experiments Analyzing round trip query latency Configuring JavaScript caches for better performance Introducing Bits AI Dev Agent for Code Security Datadog achieves ISO 42001 certification for responsible AI Monitor Nutanix clusters, hosts, and VMs with Datadog Monitor Juniper Mist in Datadog A new Host Map for modern infrastructure Annotate traces to improve LLM quality with Datadog LLM Observability What’s new in Cloud SIEM: AI-powered investigations, enhanced threat intelligence, and scalable security operations Explore Kubernetes with native OpenTelemetry data Monitor Oracle Fusion Cloud Applications with Datadog Announcing the Datadog Terraform provider v4.0.0 Scaling Kubernetes workloads on custom metrics How to design cloud environments for AI-powered threat analysis Monitor Aruba Central in Datadog How we centralize and remediate risks with Datadog Case Management Accelerate incident response with Datadog and ServiceNow Monitor your application and network load balancer logs Understanding Karpenter architecture for Kubernetes autoscaling Tools for collecting metrics and logs from Karpenter Monitor Karpenter with Datadog What your product data is actually saying Key metrics for monitoring Karpenter Securing Datadog’s platform in the AI age: The role of observability data Four ways engineering teams use the Datadog MCP Server to power AI agents Approaching your observability migration with the right mindset Meet the new Bits AI SRE: Deeper reasoning, twice as fast Key learnings from the 2026 State of DevSecOps study Use plain English to query your multi-cloud infrastructure in Resource Catalog Simplifying troubleshooting across the user journey with Datadog Synthetic Monitoring Protect your OCI resources with Datadog Cloud Security This Month in Datadog - February 2026 Amazon EC2 security: How misconfigured and public AMIs expand your cloud attack surface Enable end-to-end visibility into your Java apps with a single command Measure and improve mobile app startup performance with Datadog RUM Evaluating our AI Guard application to improve quality and control cost Identify untested code across every level of your codebase Make use of guardrail metrics and stop babysitting your releases Monitor Versa Networks SD-WAN performance in Datadog Improve performance and reliability with APM Recommendations Remediate transitive vulnerabilities faster with Datadog Software Composition Analysis Generate audit-ready vulnerability and compliance reports with Datadog Sheets Monitor Fortinet FortiManager performance in Datadog Improve test coverage across codebases with Datadog Code Coverage Move fast, don’t break things: Consistent testing standards at scale Enrich logs with ServiceNow CMDB context before routing to any SIEM or logging tool Monitor Lustre with Datadog Make faster, better product decisions with Datadog Product Analytics Surface and remediate runtime posture issues with Workload Protection Findings Protect agentic AI applications with Datadog AI Guard How to optimize JavaScript code with CSS Trace Google Pub/Sub workloads in Cloud Run with Datadog Detect human names in logs with ML in Sensitive Data Scanner How we cut our NLQ agent debugging time from hours to minutes with LLM Observability Debug PostgreSQL query latency faster with EXPLAIN ANALYZE in Datadog Database Monitoring Datadog acquires Propolis Unify and correlate frontend and backend data with retention filters Scale compliance across global frameworks with Datadog Cloud Security Monitor Arista VeloCloud SD-WAN performance with Datadog Building reliable dashboard agents with Datadog LLM Observability Simplify log collection and aggregation for MSSPs with Datadog Observability Pipelines Mitigation for Node.js denial-of-service vulnerability affecting Datadog APM Automate flaky test fixes with the Bits AI Dev Agent and Test Optimization How we built an AI SRE agent that investigates like a team of engineers Datadog integrations 2025 recap: Observability for AI, security, and hybrid cloud Design effective executive dashboards with Datadog Implement dbt data quality checks with dbt-expectations Bring faster visibility into AWS Lambda functions with remote instrumentation Troubleshoot faster with the GitLab Source Code integration in Datadog How Cambia Health Solutions saved $30,000 monthly with Cloud Cost Management and the Datadog Resource Catalog Normalize any logs for Cloud SIEM with Datadog's OCSF processor Optimizing Datadog at scale: Cost-efficient observability at Zendesk Detect, diagnose, and resolve network issues easily with CNM Network Health Connect engineering errors to user impact in early-stage products Cilium configuration for Kubernetes operations at scale Designing feedback loops for progressive delivery Ship features faster and safer with Datadog Feature Flags Choosing the right OpenTelemetry Collector distribution Route your monitor alerts with Datadog monitor notification rules Automate Cloud SIEM investigations with Bits AI Security Analyst Cloud threat detection: How to identify risky activity across control and data planes Collecting Kafka performance metrics Monitoring Kafka with Datadog Monitoring Kafka performance metrics
How we cut Spark compute costs by 44% with agentic AI and Datadog Jobs Monitoring
Charles Yu, Meghna Banerjee, Eddie Cai · 2026-06-01 · via Datadog | The Monitor blog

Spark jobs only get more expensive and harder to debug as they scale. It’s a problem we’ve run into ourselves. Our Referential Data Platform team builds and maintains the knowledge graph that maps relationships between customers’ observability entities. ServiceQueryEdge is at the center of that graph, mapping service entities to their associated metric and log queries. It runs daily across seven datacenters, with individual partitions processing up to 27 TB of input and 16 billion records. At that scale, we were averaging $1.5k of infrastructure costs daily, with each run taking over 17 hours.

AI agents seemed like a natural fit for this problem. They’re good at reasoning over code, connecting symptoms to root causes, and generating hypotheses quickly. But an agent working from code alone is still guessing. It needs to know what’s actually slow.

In this post, we’ll walk through how we used Datadog’s Data Observability Jobs Monitoring and an AI agent built on Claude to debug and optimize ServiceQueryEdge. We’ll cover what worked, what didn’t, and the specific changes that cut our daily compute costs by 44% and reduced run duration by 60% in US1, our largest data center.

Closing the gap between Jobs Monitoring and the codebase

To understand where inefficiencies are, we rely on Jobs Monitoring with the Spark SQL Plan to get a visual, interactive representation of the full execution plan. However, even with that visibility, correlating a slow operator in the SQL Plan back to the relevant section of application code can still take time, particularly for a large, complex job like ServiceQueryEdge.

The full Spark SQL Plan for our ServiceQueryEdge job showing 76.9% of executor time on shuffle and 65.5% skew ratio across a deep execution graph.

To speed up debugging, we built an AI agent to surface any bottlenecks across the execution graph and suggest fixes. We created a custom prompt structure that ingests the same data shown in Jobs Monitoring, such as stage metrics, the SQL execution plan, and telemetry data, alongside the source code. This allows the agent to perform correlation work that would usually fall on one of the team’s engineers, saving up to hours of manual investigation. For every issue the agent flags, the engineer lands directly at the relevant node with context on why it matters.

Getting signal from noise: scoping data for AI-assisted debugging

At first, we ran into problems with Claude depleting its context while making Model Context Protocol (MCP) calls through our Datadog MCP Server to collect Spark data from Jobs Monitoring. The agent pulled job run telemetry data, represented as traces, using the get_datadog_trace, apm_search_spans, and apm_explore_trace tools. Multiple runs made the problem worse. The agent exhausted its context window before completing meaningful analysis. Suggestions became incomplete or incoherent. 

We alleviated this by using subagents that delegated the acquisition of specific information into targeted tasks, preserving context for the analysis work that actually mattered. Agent output quality depended less on data volume than on how precisely that data was scoped.

However, the agent’s initial suggestions didn’t work. Many recommendations were either off target or addressed symptoms rather than root causes. For example, the agent suggested pruning column reads to reduce data read in, which was redundant because Spark had already handled that optimization. 

Our first instinct was to try reducing the noise by feeding deeper embeddings of Spark runtime info and metrics from Jobs Monitoring into the agent. The additional context helped the agent identify more issues and generate more recommendations. But false positives also increased significantly. 

Our next step was to add another subagent that filtered issues we deemed irrelevant or perfunctory. The validator works as a grader rather than a generator. It pulls the same context as the main agent, job health and the SQL plan, but its job is to find reasons a proposed fix won’t work.

Optimization #2: Apply spam filter before the main join (line 766)

Contraindication Checks:

1. Does the spam filter actually reduce significant rows? The LeftAnti join outputs 3,406,939,340 rows same as the left input (3,406,939,340 from the

Exchange). This means the 628 spammy metrics contribute near-zero rows to the 3.4B total. The spam filter removes metrics by (org_id, metric_id) but the main join key

is (org_id, metric_name). Even if these 628 metrics were removed earlier, the 3.4B metrics-side of the main join would barely decrease.

2. Is the filter already applied before the main join? Looking at the data flow: metricsWithTagsRaw spam filter metricsWithTags salt

metricsWithTagsSalted main join at line 766. The spam filter IS already applied before the main join in the code.

3. Would earlier filtering help? The spam filter requires a groupBy($"org_id", $"metric_id").agg(countDistinct(...)) over the full metricsWithTagsRaw. This

aggregation itself is expensive. Moving it "earlier" doesn't change the fact that you need the full dataset to compute the counts.

SQL Plan Evidence:

- The SortMergeJoin (LeftAnti) outputs 3,406,939,340 rows — virtually identical to its 7,431,377,983 / 2 input (accounting for the join filtering ~half). The 628

metrics removed have negligible impact on downstream volume.

Impact Assessment:

- The spam filter already runs before the main join in the current code flow (lines 737-754 before line 766)

- The 628 removed metrics represent a tiny fraction of the 3.4B rows

Recommendation: ❌ Do not proceed — The spam filter already runs before the main join. The 628 removed metrics have negligible impact on the downstream 3.4B row

count. This optimization is based on an incorrect assumption about the ordering.

For each optimization type, the validator checks a specific set of contraindications. Some checks are about whether the fix actually addresses the measured bottleneck. Others look for cases where Spark is already handling the issue automatically or where the fix could introduce new problems downstream. It also checks whether the root cause originates upstream of the flagged stage rather than in the stage itself.

The validator also estimates the potential impact of each fix, using stage CPU contribution and bottleneck type to rate proposals as high, medium, or low priority, so the main agent can rank what to tackle first. The contraindication list is updated to encode team knowledge directly into the validation step. When the same poor recommendation kept surfacing, we added a rule to catch it.

This pattern of using a second agent consistently helps identify gaps the original agent missed and refines suggestions to get to a fix that works.

Agent workflow diagram showing data sources feeding into automated triage and remediation, with human review before implementation.

From here, the agent acted as an effective debugging companion. It drew connections that would have taken hours to surface manually. For example, it mapped a slow shuffle operation in the execution plan back to a specific transformation in the code and flagged a join pattern worth trying as a broadcast hint.

Spark optimization findings and implementation

Once we worked with the subagents to output the relevant context, the agent yielded three primary recommendations for improvement:

  • Removing redundant aggregation

  • Replacing a sorted join with a broadcast join

  • Fixing improper salting

To validate our changes efficiently, we built a dashboard using Jobs Monitoring data that compared key performance metrics, such as executor time, skew, shuffle, memory, and overall duration. This helped us summarize and evaluate results.

Removing redundant aggregation

Jobs Monitoring indicated that Spark was performing a HashAggregate operation when joining queryVertices and metricsWithTag, despite no explicit aggregate instruction. The agent traced this back to a distinct() call, which was redundant given the structure of the upstream data. 

The agent also identified that every downstream stage performed the same aggregation on the joined data. Those rows were being shuffled unnecessarily before reaching an identical aggregation step.

Spark SQL Plan showing 3.12 TB peak memory and 165k hash probes/key from a redundant distinct() call.

We decided to aggregate the data early to reduce the number of rows shuffled, saving CPU and network I/O. This reduced the number of rows being sent to the subsequent stages to less than a tenth of the original size on some shards.

Spark SQL Plan after removing distinct() showing peak memory down to 2.19 TB and shuffle data reduced to 1.63 TB.

Replacing a sorted join with a broadcast join

Next, the Spark SQL plan indicated a left_anti join between a table of around 500 rows and another of over seven billion rows. Spark defaulted to a SortMergeJoin strategy, forcing a sort of the entire multi-billion-row table. We attempted to change this to a broadcast join to avoid that sort entirely.

Spark SQL Plan showing SortMergeJoin between a 500-row table and a multi-billion-row table, with full-table sort operations.
Spark SQL Plan showing SortMergeJoin between a 500-row table and a multi-billion-row table, with full-table sort operations.

🥈 #2 Bottleneck: Context Spam Filter / LeftAnti Join (Stage 2137998469030960929)

Duration: 16 min | CPU: 24.4 hours | Spill: 14.3 TB memory, 3.6 TB disk

This stage runs the spam filtering at ServiceQueryEdge.scala:737-754:

val metricsToSkipDf = metricsWithTagsRaw

.groupBy($"org_id", $"metric_id")

.agg(countDistinct($"bhandle0", $"bhandle1") as "contexts_count")

.filter($"contexts_count" > contextSpamThreshold) // threshold = 2,097,152

The Sort node before the LeftAnti SortMergeJoin spills 10.4 TB sorting 7.4B rows. The right side (628 spammy metrics) is tiny but the left side is massive.

Recommendations:

1. Use a broadcast for the spam filter Only 628 metrics exceed the threshold. Broadcast the spam list and use a BroadcastHashJoin instead of SortMergeJoin to avoid the 10.4 TB sort

spill.

2. Apply spam filter earlier in the pipeline Filter before the expensive salt/join rather than after, to reduce downstream row counts.

The initial impact of this change appeared marginal in isolation. We suspected this was because the persistent data skew in the worst stage was limiting its effectiveness. Once we addressed the skew in the following step, the broadcast join contributed meaningfully to the overall gains. This is worth noting as a reminder that optimizations don’t always surface their full value immediately. Improvements that appear incremental on their own may unlock headroom that a later fix will realize. When isolating and quantifying the impact of individual changes, the order in which they are applied matters.

Fixing improper salting

As mentioned earlier, the Spark Plan UI and agent recognized a large data skew and suggested increasing salting values, which did not directly resolve the issue. The agent had enough signal from the stage metrics to identify that skew was the problem, but lacked the context to understand the salting implementation that caused it. The fix only surfaced when an engineer manually added context about the implementation. This showed that agents reason best when given complete context. 

While dropping the bhandle* columns, we identified a flaw in our salting logic. 

Our salting logic was creating salts by applying modulus to a skewed column, which meant the resulting salt values inherited the same non-uniform distribution.

Jobs Monitoring histogram before salting fix showing p99 tasks at 2h 55m and a long tail extending past 7 hours.

We switched to using rand() instead, which breaks the dependency on the skewed column and distributes rows uniformly across partitions. Once this change was in place, the earlier suggestion of increasing salting values worked as expected. Duration skew on the worst task significantly improved from a max-p50 of 25 minutes to around 10 minutes, and the overall duration fell over 57%.

Jobs Monitoring histogram after switching to rand() salting showing worst-case task duration reduced from 2h 55m to 1h 13m.

Results and impact

In the week following deployment, cloud cost monitoring confirmed the impact directly. Daily compute costs for ServiceQueryEdge dropped from an average of $1.5k per day to approximately $830 per day, a 44% reduction. In our US1 datacenter specifically, run duration fell by 60% and allocated executor time dropped by nearly 50%. 

At a pre-optimization annual cost of approximately $600k in US1, these gains represent an estimated $250k in annual savings from that datacenter alone. Across all other datacenters, which account for roughly $120k per year in compute costs, the same optimizations project an additional $50k per year in savings. Further cost reductions are expected as cluster pods are right-sized to reflect the lower peak execution memory requirements, though that impact has not yet been quantified.

Cost Management chart showing daily costs drop from $1.89k to $1.17k, confirming 44% reduction after deploying optimizations.

Getting started with agentic Spark optimization and the Datadog MCP Server

From our experiment, we found that treating AI agents as collaborative partners rather than autonomous problem solvers produced better outcomes. Deeper optimizations emerged that neither the team nor the agent would have reached independently. The agent’s value came not from producing ready-made answers but from surfacing connections across the execution plan that pointed engineers toward the right questions. Structuring the agent’s access to data deliberately, starting broad and retrieving fine-grained detail only where needed, was as important as the analysis itself. 

This principle directly informed how we built the latest Spark tools for the Datadog MCP Server. Rather than exposing everything at once, we built two purpose-scoped tools. get_spark_health returns the overall health of a Spark job and its worst-performing stages, giving the agent a ranked starting point without overwhelming its context. get_spark_sql_plan retrieves the low-level execution plan and stage metrics for a specific trace, allowing the agent to go deep once the right stage has been identified. A sample prompt may look like:

“Help me optimize the <job name> Spark job. Evaluate its performance over the past day using get_spark_health. Look at the worst stages, correlate them with the code at <repo path>, and retrieve the associated execution plans using get_spark_sql_plan. For the worst stage, create hypotheses for the root cause and discuss them with me.”

You can also include a Markdown file of your team’s known Spark patterns so the agent validates any proposed fixes against existing standards.

The Datadog MCP Server is generally available. Its data-observability toolset is in Public Preview and includes the Spark tools used in this post. Configure the toolset to bring Data Observability Jobs Monitoring data into your own agentic debugging workflow.

To get started with Jobs Monitoring, use our documentation or sign up for a 14-day free trial.