惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
V2EX - 技术
V2EX - 技术
S
Secure Thoughts
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
L
LINUX DO - 最新话题
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Hacker News: Ask HN
Hacker News: Ask HN
T
Troy Hunt's Blog
Forbes - Security
Forbes - Security
Application and Cybersecurity Blog
Application and Cybersecurity Blog
P
Proofpoint News Feed
Know Your Adversary
Know Your Adversary
Schneier on Security
Schneier on Security
H
Heimdal Security Blog
C
Cybersecurity and Infrastructure Security Agency CISA
Simon Willison's Weblog
Simon Willison's Weblog
V
Vulnerabilities – Threatpost
月光博客
月光博客
罗磊的独立博客
Webroot Blog
Webroot Blog
博客园 - 【当耐特】
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
The Cloudflare Blog
爱范儿
爱范儿
Last Week in AI
Last Week in AI
博客园 - 聂微东
博客园 - 叶小钗
美团技术团队
A
Arctic Wolf
P
Palo Alto Networks Blog
T
Tailwind CSS Blog
Cyberwarzone
Cyberwarzone
雷峰网
雷峰网
Apple Machine Learning Research
Apple Machine Learning Research
人人都是产品经理
人人都是产品经理
宝玉的分享
宝玉的分享
H
Hacker News: Front Page
大猫的无限游戏
大猫的无限游戏
S
SegmentFault 最新的问题
Jina AI
Jina AI
C
Cyber Attacks, Cyber Crime and Cyber Security
The Last Watchdog
The Last Watchdog
IT之家
IT之家
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
酷 壳 – CoolShell
酷 壳 – CoolShell
阮一峰的网络日志
阮一峰的网络日志
J
Java Code Geeks
B
Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
P
Privacy & Cybersecurity Law Blog

OneUptime Blog

Grafana Stack vs OneUptime: DIY Observability or Unified Platform? Your AI Workloads Are About to Blow Up Your Observability Bill The Great Observability Consolidation Is Here How to Write Custom Object Classes for Ceph How to Write Custom Ceph Manager Modules How to Write a ceph.conf Configuration File How to Use Rook-Ceph with OpenShift How to Use Rook-Ceph with Longhorn for Comparison How to Configure Volume Snapshot Class for RBD in Rook How to Configure VolumeReplicationClass Scheduling Intervals in Rook How to Set Up Volume Replication with Rook-Ceph How to Create Volume Group Snapshots with Rook CSI How to Visualize Ceph Network Performance in Grafana How to Enable Virtual Host-Style Bucket Access in Rook How to View Runtime Configuration via Admin Socket How to View Quota Settings and Update Stats in Ceph RGW How to View PG Scaling Recommendations with autoscale-status How to View PG Distribution via Admin Socket How to View Performance Metrics in the Ceph Dashboard How to View OSD Performance Counters in Ceph How to View Connection Status via Admin Socket How to View Ceph Cluster Summary Dashboard via CLI How to Version Control Rook-Ceph Configuration How to Version Control Ceph Infrastructure with Terraform How to Verify Kubernetes Node Requirements for Rook-Ceph Deployment How to Verify Health Before and After Rook Upgrades How to Verify Data Integrity with Deep Scrubbing How to Verify Complete Rook-Ceph Cleanup How to Verify Backup Integrity from Ceph Snapshots How to Use Rook-Ceph with Velero for Kubernetes Backup How to Integrate HashiCorp Vault with Rook-Ceph (Token Auth) How to Configure TLS for Vault Integration in Rook How to Integrate HashiCorp Vault with Rook-Ceph (Kubernetes Auth) How to Validate Ceph Cluster Configuration After Deployment How to Understand User Type and ID Notation (TYPE.ID) in Ceph How to Configure User Management in the Ceph Dashboard How to Use Rook-Ceph with Kubernetes Operators How to Use Rook-Ceph with Helm Chart Deployments How to Use the Swift API with Ceph RGW How to Use SQLite Databases Stored on Ceph How to Use s3cmd with Ceph RGW How to Use the S3 API with Ceph RGW How to Use Red Hat Ceph with RHEL Virtualization How to Use RBD with QEMU How to Use RBD with Nomad How to Use RBD with CloudStack How to Use RBD Snapshot Rollback How to Use rados bench for Object Storage Benchmarking How to Secure Rook-Ceph with Pod Security Admission How to Use pg-upmap for PG Mapping in Ceph How to Use Multipath Devices with Ceph OSDs How to Use MinIO Client (mc) with Ceph RGW How to Use fs swap for CephFS How to Use fio for Ceph Block Storage Benchmarking How to Use the CephFS Shell How to Use Ceph RGW for Media Asset Management How to Use Ceph RGW for Log Storage and Archival How to Use Ceph RGW for Data Lake Storage How to Use Ceph RGW for Backup Repository Storage How to Use the ceph-authtool Utility How to Use boto3 (Python) with Ceph RGW S3 How to Use AWS CLI with Ceph RGW S3 How to Use the Admin Ops API with Ceph RGW How to Configure Usage Log Key Transition in Ceph RGW How to Handle Rook-Ceph Upgrades in GitOps Pipelines How to Upgrade Rook-Ceph with Zero Downtime How to Create a Ceph Upgrade Runbook How to Upgrade the Rook Operator from v1.18 to v1.19 How to Upgrade the Rook Operator on Kubernetes How to Upgrade External Cluster Connections in Rook How to Upgrade the Ceph Version in Rook How to Upgrade from Ceph Reef to Squid How to Upgrade from Ceph Quincy to Reef How to Upgrade Ceph Clusters in Stretch Mode How to Update Kernel for CephFS Feature Compatibility How to Update Ceph Configuration on a Running Rook Cluster How to Create Unique Kubernetes Services per NFS Server in Rook How to Understand When Compression Helps vs Hurts in Ceph How to Understand User Types (Individual vs System) in Ceph How to Understand the undersized PG State in Ceph How to Understand the stale PG State in Ceph How to Understand the repair PG State in Ceph How to Understand the remapped PG State in Ceph How to Understand Red Hat Ceph Storage vs Upstream Ceph How to Understand Placement Groups in Ceph How to Understand PG Splitting in Ceph How to Understand the peering PG State in Ceph How to Understand OSD Recovery Process in Ceph How to Understand the OSD Map in Ceph How to Understand New Features in Each Ceph Release How to Understand Monitor Leadership in Ceph How to Understand MDS States in CephFS How to Understand Deprecated Features in Ceph Reef How to Understand the degraded PG State in Ceph How to Understand D3N in Ceph How to Understand the creating PG State in Ceph How to Understand the clean PG State in Ceph How to Understand CephX Authentication Protocol How to Understand CephX Authentication Flow How to Understand What Data Ceph Telemetry Collects
How to Monitor Azure App Services (PaaS) with OpenTelemetry
Jamie Mallers · 2026-04-17 · via OneUptime Blog

Azure App Service is a Platform-as-a-Service offering. You don't get the host, you don't get root, and you can't run a daemon next to your app the way you would on a VM. That changes how you approach OpenTelemetry.

On IaaS you'd drop a Collector agent on the box, point your application at localhost:4318, and be done. On App Service you have to work around a locked-down sandbox - which means in-process instrumentation in the app itself, and a separate pipeline for the platform-level telemetry the SDK can't see. This post walks through a setup that covers both.


The architecture at a glance

There are two telemetry streams on App Service, and you need both:

  1. Application telemetry - traces, metrics, and logs emitted by your code. The OpenTelemetry SDK lives inside your app process and exports OTLP directly (or to a Collector).
  2. Platform telemetry - HTTP logs, console logs, platform/container events, and platform metrics such as CPU, memory, HTTP queue length, and response times. These come from the App Service runtime itself and are exposed through Diagnostic Settings. They need a separate path to OTLP.

Treat them as two pipelines that happen to land in the same backend. Teams that only do #1 are blind when the platform misbehaves (scale-in killing requests, cold starts, quota throttling). Teams that only do #2 have no trace data and can't correlate errors to requests.


Prerequisites

  • An Azure App Service running Node.js, .NET, Java, or Python. Node.js, .NET, and Java can run on Linux or Windows; Python's built-in App Service runtime is Linux-only unless you bring a custom Windows container.
  • Permission to edit Application settings and Diagnostic settings on the App Service
  • An OTLP-compatible backend (this guide uses OneUptime)
  • Optionally, an OpenTelemetry Collector deployed somewhere reachable (Container Apps, AKS, or a VM)

Step 1: Instrument the application

The fastest path is auto-instrumentation. You install the OpenTelemetry SDK for your runtime, set a handful of environment variables, and the SDK takes over - wrapping HTTP handlers, database clients, and outbound calls without code changes.

Node.js

Add the auto-instrumentation packages to your package.json:

npm install @opentelemetry/api \
            @opentelemetry/auto-instrumentations-node \
            @opentelemetry/exporter-trace-otlp-http \
            @opentelemetry/exporter-metrics-otlp-http \
            @opentelemetry/exporter-logs-otlp-http

Then in the Azure Portal, go to your App Service → Configuration → Application settings and add:

NODE_OPTIONS = --require @opentelemetry/auto-instrumentations-node/register
OTEL_SERVICE_NAME = my-node-app
OTEL_EXPORTER_OTLP_ENDPOINT = https://oneuptime.com/otlp
OTEL_EXPORTER_OTLP_HEADERS = x-oneuptime-token=YOUR_TOKEN
OTEL_RESOURCE_ATTRIBUTES = deployment.environment.name=prod,cloud.provider=azure,cloud.platform=azure.app_service
OTEL_TRACES_EXPORTER = otlp
OTEL_METRICS_EXPORTER = otlp
OTEL_LOGS_EXPORTER = otlp

Restart the app. Incoming HTTP requests, outbound fetch/http calls, and popular database clients will produce spans immediately.

.NET

For .NET 6+ the cleanest option is the OpenTelemetry Auto-Instrumentation module. Upload OpenTelemetry.AutoInstrumentation to /home/site/wwwroot/otel/ via Kudu (or bundle it in your deployment zip), then set:

OTEL_DOTNET_AUTO_HOME = /home/site/wwwroot/otel
CORECLR_ENABLE_PROFILING = 1
CORECLR_PROFILER = {918728DD-259F-4A6A-AC2B-B85E1B658318}
CORECLR_PROFILER_PATH = /home/site/wwwroot/otel/linux-x64/OpenTelemetry.AutoInstrumentation.Native.so
DOTNET_STARTUP_HOOKS = /home/site/wwwroot/otel/net/OpenTelemetry.AutoInstrumentation.StartupHook.dll
DOTNET_ADDITIONAL_DEPS = /home/site/wwwroot/otel/AdditionalDeps
DOTNET_SHARED_STORE = /home/site/wwwroot/otel/store
OTEL_SERVICE_NAME = my-dotnet-app
OTEL_EXPORTER_OTLP_ENDPOINT = https://oneuptime.com/otlp
OTEL_EXPORTER_OTLP_HEADERS = x-oneuptime-token=YOUR_TOKEN

On Windows App Service, use CORECLR_PROFILER_PATH_32 or CORECLR_PROFILER_PATH_64 for the Windows DLL path; the profiler CLSID stays the same. If you prefer explicit wiring, add the OpenTelemetry.Extensions.Hosting NuGet and configure the tracer provider in Program.cs - either approach works.

Java

Java has the smoothest ride on App Service because the Java agent attaches at JVM startup with a single flag. Upload opentelemetry-javaagent.jar to /home/site/wwwroot/ and set this for Java SE apps:

JAVA_OPTS = -javaagent:/home/site/wwwroot/opentelemetry-javaagent.jar
OTEL_SERVICE_NAME = my-java-app
OTEL_EXPORTER_OTLP_ENDPOINT = https://oneuptime.com/otlp
OTEL_EXPORTER_OTLP_HEADERS = x-oneuptime-token=YOUR_TOKEN
OTEL_EXPORTER_OTLP_PROTOCOL = http/protobuf

For Tomcat apps, put the same -javaagent:/home/site/wwwroot/opentelemetry-javaagent.jar value in CATALINA_OPTS instead of JAVA_OPTS.

Tomcat, Spring Boot, JDBC, Kafka, and the rest of the common stack are instrumented out of the box.

Python

Install the agent packages and run the bootstrap helper during your deployment:

pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install

Then change the startup command under Configuration → General settings → Startup Command to wrap your entrypoint:

opentelemetry-instrument gunicorn --bind=0.0.0.0 --workers=4 app:app

Set the usual OTEL_* variables as Application settings, same as the Node.js example.


Step 2: Set the right resource attributes

App Service exposes a handful of environment variables that uniquely identify a running instance. Map them into OpenTelemetry resource attributes so every span, metric, and log can be pinned to a specific slot, region, and instance. If you control a Linux startup script or container entrypoint, build the attribute string there so the shell expands the App Service variables:

export OTEL_RESOURCE_ATTRIBUTES="cloud.provider=azure,cloud.platform=azure.app_service,cloud.region=${REGION_NAME},service.instance.id=${WEBSITE_INSTANCE_ID:-$WEBSITE_ROLE_INSTANCE_ID},deployment.environment.name=${WEBSITE_SLOT_NAME:-Production}"

If you configure OTEL_RESOURCE_ATTRIBUTES directly in the Azure Portal as an Application setting, App Service passes the value as-is; it does not expand $REGION_NAME or $WEBSITE_INSTANCE_ID inside the setting. In that case, set concrete values or set the attribute string in startup code. The payoff shows up the first time you're debugging a bad deploy and can filter to a single instance or slot without writing a custom tag.


Step 3: Get platform telemetry out via Diagnostic Settings

The SDK inside your app can't see what happens before the request reaches your code - front-end HTTP logs, App Service platform events, resource metrics, container startup failures. For that, use Diagnostic Settings.

Go to App Service → Monitoring → Diagnostic settings → Add diagnostic setting and enable at least:

  • AppServiceHTTPLogs - the front-end HTTP request log
  • AppServiceConsoleLogs - stdout/stderr from your container
  • AppServicePlatformLogs - container lifecycle and platform events
  • AllMetrics - CPU, memory, HTTP queue length, response times

Route them to an Event Hub. If you choose a named Event Hub in the diagnostic setting, all selected categories can go through that hub. If you leave the Event Hub name blank, Azure Monitor writes category-specific hubs such as insights-logs-appservicehttplogs and insights-metrics-pt1m, so configure one receiver per hub.

Then run an OpenTelemetry Collector with the azure_event_hub receiver to consume the stream and translate it into OTLP:

receivers:
  azure_event_hub:
    connection: Endpoint=sb://...;SharedAccessKeyName=...;SharedAccessKey=...;EntityPath=appservice-diagnostics
    format: azure

processors:
  batch: {}
  resource:
    attributes:
      - key: cloud.provider
        value: azure
        action: upsert
      - key: cloud.platform
        value: azure.app_service
        action: upsert

exporters:
  otlphttp:
    endpoint: https://oneuptime.com/otlp
    headers:
      x-oneuptime-token: YOUR_TOKEN

service:
  pipelines:
    logs:
      receivers: [azure_event_hub]
      processors: [resource, batch]
      exporters: [otlphttp]
    metrics:
      receivers: [azure_event_hub]
      processors: [resource, batch]
      exporters: [otlphttp]

Run that Collector on Azure Container Apps or an AKS cluster in the same region as your App Service. Now platform logs land in the same backend as your application traces, and you can correlate a 502 in the front-end HTTP log to the exception that caused it in the app.


Step 4: Decide whether you need a Collector in front of the app

For a small app, exporting OTLP straight from the SDK to your backend works. For anything that matters in production, put a Collector between the two. Reasons:

  • Batching and retries. The Collector buffers during backend hiccups. The SDK dropping spans during a 30-second outage is a familiar pain.
  • Enrichment. Strip secrets, add tenant IDs, normalize attribute names once - not in every service.
  • Multi-backend fan-out. Send traces to OneUptime and metrics to a long-term metrics store without re-instrumenting.

On App Service you have two Collector placement options:

  1. Sidecar container on Web App for Containers. Add the otel/opentelemetry-collector-contrib image as a sidecar to your main app container. The SDK talks to the sidecar on localhost:4318. This is the path Microsoft now recommends - Docker Compose multi-container support is on retirement (ending March 31, 2027) in favor of sidecar containers, so start here even if you've used compose in the past.
  2. Dedicated Collector on Azure Container Apps or AKS. The SDK exports over HTTPS to the Collector's public endpoint (or private endpoint via VNet integration). Works for every App Service runtime without changing the deployment artifact.

App Service-specific gotchas

These are the things that consistently bite teams moving OpenTelemetry onto App Service:

  • Enable "Always On". Without it, the app unloads during idle periods and takes the SDK's background exporter with it - partial traces, missing metrics, and cold-start latency that looks like a bug in your code.
  • The writable path is /home on Linux and D:\home on Windows. If you want to write the Collector binary or a Java agent somewhere persistent on Linux, put it under /home/site/wwwroot/. For custom containers, make sure App Service storage is enabled if you expect /home to persist.
  • Slot swaps recycle the process. Keep OTEL_BSP_SCHEDULE_DELAY at its default 5000 ms, or lower it carefully, so in-flight spans flush before the old slot shuts down. Just don't bump it up.
  • Private networking. If the Collector is inside a VNet, enable VNet Integration on the App Service and verify the outbound egress rules allow traffic to the Collector's private endpoint.
  • Windows App Service sandbox restrictions. Some profilers and native hooks are blocked. If the .NET auto-instrumentation fails silently on Windows, switch to code-based instrumentation with OpenTelemetry.Extensions.Hosting.
  • Container startup timeout. App Service gives a Linux container 230 seconds by default to start responding on the configured port (tunable via WEBSITES_CONTAINER_START_TIME_LIMIT, 10–1800 seconds). A misconfigured Collector sidecar, or an app entrypoint that waits on it, can make the app fail to boot. Keep the sidecar's startup fast and don't make the app wait on it.
  • Resource CPU/memory limits. The Collector sidecar shares the plan's CPU and RAM. On a B1 plan you'll notice the overhead. Size the plan accordingly if you're running a Collector sidecar.

Verifying the pipeline

A quick end-to-end check before you declare victory:

  1. Hit a route on your app and look for a trace in your backend with service.name=my-app and cloud.platform=azure.app_service.
  2. Force a 502 by stopping the app mid-request, then check that the platform HTTP log arrived through the Event Hub → Collector path.
  3. Restart the app and confirm a span appears for the first request after cold start. If it doesn't, Always On is probably off.
  4. Filter by service.instance.id and confirm you can see each instance separately during a scale-out.

If all four pass, your two pipelines are healthy and you can start building SLOs against them.


Sending it all to OneUptime

Point OTEL_EXPORTER_OTLP_ENDPOINT at your OneUptime OTLP ingest URL, set the x-oneuptime-token header, and traces, metrics, and logs land in the corresponding telemetry services - no Azure Monitor workspace or vendor-specific agents. The Collector pipeline for platform logs sends to the same endpoint, so application and platform telemetry sit next to each other in the same UI.

The point of OpenTelemetry on a PaaS like App Service is to keep the observability story portable. You get the same SDK, the same wire format, and the same backend choices you'd have on Kubernetes or bare metal. App Service just adds a few constraints - a locked-down sandbox, a separate path for platform telemetry, and a handful of runtime quirks - that are worth knowing up front.