惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
SegmentFault 最新的问题
The Last Watchdog
The Last Watchdog
P
Proofpoint News Feed
C
Cybersecurity and Infrastructure Security Agency CISA
L
LINUX DO - 热门话题
Cyberwarzone
Cyberwarzone
S
Schneier on Security
C
CERT Recently Published Vulnerability Notes
Latest news
Latest news
I
Intezer
A
Arctic Wolf
IT之家
IT之家
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
Cisco Blogs
AWS News Blog
AWS News Blog
博客园 - 三生石上(FineUI控件)
C
CXSECURITY Database RSS Feed - CXSecurity.com
F
Fortinet All Blogs
Microsoft Azure Blog
Microsoft Azure Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
T
The Exploit Database - CXSecurity.com
Google DeepMind News
Google DeepMind News
M
MIT News - Artificial intelligence
D
Docker
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
C
Cyber Attacks, Cyber Crime and Cyber Security
MongoDB | Blog
MongoDB | Blog
B
Blog
博客园 - 叶小钗
V2EX - 技术
V2EX - 技术
Simon Willison's Weblog
Simon Willison's Weblog
MyScale Blog
MyScale Blog
Hugging Face - Blog
Hugging Face - Blog
Engineering at Meta
Engineering at Meta
NISL@THU
NISL@THU
WordPress大学
WordPress大学
人人都是产品经理
人人都是产品经理
Stack Overflow Blog
Stack Overflow Blog
N
Netflix TechBlog - Medium
The GitHub Blog
The GitHub Blog
V
V2EX
PCI Perspectives
PCI Perspectives
N
News | PayPal Newsroom
V
Visual Studio Blog
Vercel News
Vercel News
P
Proofpoint News Feed
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
J
Java Code Geeks
O
OpenAI News
爱范儿
爱范儿

OneUptime Blog

How to Monitor Azure App Services (PaaS) with OpenTelemetry Grafana Stack vs OneUptime: DIY Observability or Unified Platform? The Great Observability Consolidation Is Here How to Write Custom Object Classes for Ceph How to Write Custom Ceph Manager Modules How to Write a ceph.conf Configuration File How to Use Rook-Ceph with OpenShift How to Use Rook-Ceph with Longhorn for Comparison How to Configure Volume Snapshot Class for RBD in Rook How to Configure VolumeReplicationClass Scheduling Intervals in Rook How to Set Up Volume Replication with Rook-Ceph How to Create Volume Group Snapshots with Rook CSI How to Visualize Ceph Network Performance in Grafana How to Enable Virtual Host-Style Bucket Access in Rook How to View Runtime Configuration via Admin Socket How to View Quota Settings and Update Stats in Ceph RGW How to View PG Scaling Recommendations with autoscale-status How to View PG Distribution via Admin Socket How to View Performance Metrics in the Ceph Dashboard How to View OSD Performance Counters in Ceph How to View Connection Status via Admin Socket How to View Ceph Cluster Summary Dashboard via CLI How to Version Control Rook-Ceph Configuration How to Version Control Ceph Infrastructure with Terraform How to Verify Kubernetes Node Requirements for Rook-Ceph Deployment How to Verify Health Before and After Rook Upgrades How to Verify Data Integrity with Deep Scrubbing How to Verify Complete Rook-Ceph Cleanup How to Verify Backup Integrity from Ceph Snapshots How to Use Rook-Ceph with Velero for Kubernetes Backup How to Integrate HashiCorp Vault with Rook-Ceph (Token Auth) How to Configure TLS for Vault Integration in Rook How to Integrate HashiCorp Vault with Rook-Ceph (Kubernetes Auth) How to Validate Ceph Cluster Configuration After Deployment How to Understand User Type and ID Notation (TYPE.ID) in Ceph How to Configure User Management in the Ceph Dashboard How to Use Rook-Ceph with Kubernetes Operators How to Use Rook-Ceph with Helm Chart Deployments How to Use the Swift API with Ceph RGW How to Use SQLite Databases Stored on Ceph How to Use s3cmd with Ceph RGW How to Use the S3 API with Ceph RGW How to Use Red Hat Ceph with RHEL Virtualization How to Use RBD with QEMU How to Use RBD with Nomad How to Use RBD with CloudStack How to Use RBD Snapshot Rollback How to Use rados bench for Object Storage Benchmarking How to Secure Rook-Ceph with Pod Security Admission How to Use pg-upmap for PG Mapping in Ceph How to Use Multipath Devices with Ceph OSDs How to Use MinIO Client (mc) with Ceph RGW How to Use fs swap for CephFS How to Use fio for Ceph Block Storage Benchmarking How to Use the CephFS Shell How to Use Ceph RGW for Media Asset Management How to Use Ceph RGW for Log Storage and Archival How to Use Ceph RGW for Data Lake Storage How to Use Ceph RGW for Backup Repository Storage How to Use the ceph-authtool Utility How to Use boto3 (Python) with Ceph RGW S3 How to Use AWS CLI with Ceph RGW S3 How to Use the Admin Ops API with Ceph RGW How to Configure Usage Log Key Transition in Ceph RGW How to Handle Rook-Ceph Upgrades in GitOps Pipelines How to Upgrade Rook-Ceph with Zero Downtime How to Create a Ceph Upgrade Runbook How to Upgrade the Rook Operator from v1.18 to v1.19 How to Upgrade the Rook Operator on Kubernetes How to Upgrade External Cluster Connections in Rook How to Upgrade the Ceph Version in Rook How to Upgrade from Ceph Reef to Squid How to Upgrade from Ceph Quincy to Reef How to Upgrade Ceph Clusters in Stretch Mode How to Update Kernel for CephFS Feature Compatibility How to Update Ceph Configuration on a Running Rook Cluster How to Create Unique Kubernetes Services per NFS Server in Rook How to Understand When Compression Helps vs Hurts in Ceph How to Understand User Types (Individual vs System) in Ceph How to Understand the undersized PG State in Ceph How to Understand the stale PG State in Ceph How to Understand the repair PG State in Ceph How to Understand the remapped PG State in Ceph How to Understand Red Hat Ceph Storage vs Upstream Ceph How to Understand Placement Groups in Ceph How to Understand PG Splitting in Ceph How to Understand the peering PG State in Ceph How to Understand OSD Recovery Process in Ceph How to Understand the OSD Map in Ceph How to Understand New Features in Each Ceph Release How to Understand Monitor Leadership in Ceph How to Understand MDS States in CephFS How to Understand Deprecated Features in Ceph Reef How to Understand the degraded PG State in Ceph How to Understand D3N in Ceph How to Understand the creating PG State in Ceph How to Understand the clean PG State in Ceph How to Understand CephX Authentication Protocol How to Understand CephX Authentication Flow How to Understand What Data Ceph Telemetry Collects
Your AI Workloads Are About to Blow Up Your Observability Bill
Jamie Mallers · 2026-04-01 · via OneUptime Blog

Every company is shipping AI features right now. RAG pipelines, agent workflows, LLM-powered support bots, code assistants, content generators. The race to production is real.

But here is what nobody talks about at the architecture review: your observability bill is about to explode.

Traditional microservices generate predictable telemetry. A request comes in, hits a few services, returns a response. You trace it, you log it, you metric it. Done.

AI workloads are fundamentally different. And if you are monitoring them with the same tools and the same approach, you are about to get a very unpleasant invoice.

The telemetry explosion nobody planned for

A single LLM inference request can generate:

  • Prompt and completion tokens - often thousands per request, each needing tracking for cost attribution
  • Embedding vectors - high-dimensional data that traditional log systems were never designed to store
  • Chain-of-thought traces - multi-step reasoning that creates deep, branching trace trees
  • Evaluation metrics - accuracy, hallucination rates, relevance scores that do not map to HTTP status codes
  • Retry and fallback logs - model timeouts, rate limits, and graceful degradation events

A typical RAG pipeline hitting a vector database, retrieving context, calling an LLM, and post-processing the response generates 10-50x more telemetry data than an equivalent traditional API call.

Grafana's 2026 Observability Survey - the largest community-driven survey in the space, with over 1,300 respondents - found that 92% of practitioners see value in using AI within observability. But buried in that optimism is a cost problem that is already hitting teams hard.

Where traditional observability tools fail

Here is the core issue: tools like Datadog, New Relic, and Splunk price by data volume. Logs per GB. Custom metrics per host. Traces per span. APM per host.

That pricing model was designed for a world where telemetry volume scaled roughly linearly with traffic. More users meant proportionally more logs, more traces, more metrics.

AI workloads break that assumption completely.

The token economy creates a new cost dimension

Every LLM call needs token-level cost attribution. You need to know:

  • Cost per request, per user, per feature
  • Which prompts are burning budget (a poorly tuned system prompt can 10x your token usage)
  • Whether your caching layer is actually working
  • Cost-performance tradeoffs between models (is GPT-4 class accuracy worth 10x the cost of a smaller model for this use case?)

Traditional APM tools treat all this as "custom metrics" - and charge accordingly.

Evaluation metrics do not fit existing models

In traditional systems, "is it working?" means: is the HTTP status 200? Is latency under the SLO? Is the error rate below threshold?

For AI workloads, "is it working?" means something entirely different:

  • Is the model hallucinating more than last week?
  • Are responses relevant to user queries?
  • Is the model drifting from its training baseline?
  • Are embeddings still clustering correctly?

These semantic quality metrics generate continuous streams of evaluation data. Ship that to Datadog as custom metrics and watch your bill double overnight.

Traces become trees become forests

A simple chatbot interaction might look like this under the hood:

  1. User message received
  2. Intent classification (LLM call #1)
  3. Context retrieval from vector DB
  4. Document ranking and filtering
  5. Prompt assembly with retrieved context
  6. LLM completion (LLM call #2)
  7. Response validation (LLM call #3)
  8. Safety check (LLM call #4)
  9. Response delivered

That is 4 LLM calls, a vector DB query, and multiple processing steps - for a single user message. Each step generates spans, logs, and metrics. Your trace storage just went up 4-8x compared to a traditional request-response flow.

The real numbers

Let us do the math on a mid-sized AI deployment:

Scenario: Customer support bot handling 10,000 conversations per day, averaging 5 turns each.

  • 50,000 user messages per day
  • 4 LLM calls per message = 200,000 LLM invocations
  • Average 2,000 tokens per invocation = 400 million tokens per day
  • Each invocation generates ~5 trace spans = 1 million spans per day
  • Token tracking, evaluation, and cost metrics = ~20 custom metrics per invocation = 4 million metric data points per day
  • Logs for prompts, completions, and debugging = ~2KB per invocation = 400 MB of logs per day

On Datadog's pricing:

  • APM: $31/host/month (but you need dedicated GPU hosts)
  • Log Management: $0.10/GB ingested + $2.55/million log events
  • Custom Metrics: $0.05 per custom metric per month (4M data points gets expensive fast)

Teams report that adding AI workload monitoring to their existing Datadog setup has increased their observability bill by 40-200%, depending on the volume and how many custom metrics they instrument.

Some teams have responded by simply not monitoring their AI workloads properly - which is worse.

What actually works

The observability industry is waking up to this. Here is what forward-thinking teams are doing:

1. Separate your AI telemetry pipeline

Do not shove AI telemetry into the same pipeline as your application telemetry. The data shapes are different, the retention requirements are different, and the query patterns are different.

AI observability data includes large text payloads (prompts and completions), high-cardinality token counts, and evaluation scores that need statistical analysis - not just threshold alerting.

2. Use OpenTelemetry for AI instrumentation

OpenTelemetry's semantic conventions for GenAI are maturing fast. The gen_ai.* attribute namespace provides standardized ways to capture:

  • gen_ai.usage.input_tokens / gen_ai.usage.output_tokens
  • gen_ai.request.model
  • gen_ai.response.finish_reasons

This gives you vendor-neutral telemetry that does not lock you into any specific observability platform. If your bill spikes, you can switch backends without re-instrumenting.

3. Sample intelligently, not uniformly

Not every AI request needs full telemetry. A 10% sample of normal conversations gives you solid statistical coverage. But you want 100% capture of:

  • Requests that hit cost thresholds
  • Requests with low evaluation scores
  • Requests that trigger safety filters
  • Requests from high-value users or enterprise accounts

Head-based sampling does not work for AI workloads. You need tail-based sampling that makes decisions after seeing the full request lifecycle.

4. Consolidate your observability stack

This is the biggest lever. If you are paying Datadog for APM, a separate vendor for log management, another for traces, and now adding AI-specific tooling on top - you are burning money on integration overhead alone.

An open-source, self-hosted observability platform eliminates per-GB and per-metric pricing entirely. Your cost becomes compute and storage - which are predictable, controllable, and do not spike when you ship a new AI feature.

OneUptime handles monitoring, incident management, status pages, logs, APM, and traces in a single platform. Self-hosted, it is free. On SaaS, pricing is usage-based by GB ingested - not by the number of custom metrics, not by host count, not with surprise overage charges.

For AI workloads specifically, this means:

  • Log all your prompts and completions without worrying about per-event pricing
  • Track custom evaluation metrics without per-metric charges
  • Store deep traces for every AI pipeline step without span-based billing
  • Set up alerts on model drift, cost anomalies, and quality degradation using the same platform you already use for infrastructure monitoring

5. Build cost attribution into your pipeline from day one

Do not wait until the bill arrives to figure out where the money went. Instrument cost tracking at the application level:

# Tag every LLM call with cost context
span.set_attribute("ai.cost.input_tokens", input_tokens)
span.set_attribute("ai.cost.output_tokens", output_tokens)
span.set_attribute("ai.cost.model_price_per_1k", model_price)
span.set_attribute("ai.cost.estimated_usd", estimated_cost)
span.set_attribute("ai.feature", "customer_support_bot")
span.set_attribute("ai.customer_tier", customer.tier)

This lets you answer questions like "How much does our AI support bot cost per enterprise customer per month?" - which is the question your CFO will ask approximately three months after launch.

The bottom line

AI workloads are not just "more data." They are a fundamentally different telemetry profile that breaks the assumptions traditional observability pricing was built on.

The teams that figure this out early - who build cost-aware AI telemetry pipelines, who consolidate their observability stack, who use open standards and open-source tooling - are the ones who will actually be able to afford to run AI in production long-term.

The rest will discover that their monitoring bill has quietly become their second-largest infrastructure cost, right behind the GPU instances themselves.

Do not let your observability vendor be the biggest winner from your AI investment. Own your telemetry. Control your costs. Ship AI features without the bill shock.