惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
The Blog of Author Tim Ferriss
S
Schneier on Security
博客园 - 聂微东
爱范儿
爱范儿
大猫的无限游戏
大猫的无限游戏
有赞技术团队
有赞技术团队
腾讯CDC
博客园 - 叶小钗
WordPress大学
WordPress大学
博客园_首页
J
Java Code Geeks
Last Week in AI
Last Week in AI
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
V
V2EX
Microsoft Azure Blog
Microsoft Azure Blog
The GitHub Blog
The GitHub Blog
N
Netflix TechBlog - Medium
Y
Y Combinator Blog
Schneier on Security
Schneier on Security
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
Recorded Future
Recorded Future
The Register - Security
The Register - Security
C
Cybersecurity and Infrastructure Security Agency CISA
P
Privacy & Cybersecurity Law Blog
P
Proofpoint News Feed
P
Privacy International News Feed
K
Kaspersky official blog
C
CERT Recently Published Vulnerability Notes
阮一峰的网络日志
阮一峰的网络日志
F
Full Disclosure
NISL@THU
NISL@THU
AWS News Blog
AWS News Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
U
Unit 42
MongoDB | Blog
MongoDB | Blog
A
Arctic Wolf
云风的 BLOG
云风的 BLOG
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Threatpost
D
Docker
人人都是产品经理
人人都是产品经理
T
Tailwind CSS Blog
V2EX - 技术
V2EX - 技术
G
GRAHAM CLULEY
M
MIT News - Artificial intelligence
H
Heimdal Security Blog
N
News and Events Feed by Topic
P
Proofpoint News Feed

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Why Your FSx for ONTAP Audit Logs Deserve Better Than EC2
Yoshiki Fuji · 2026-05-17 · via DEV Community

TL;DR

FSx for ONTAP file access audit logs are usually consumed through EC2-based patterns — mounted audit volumes and agent-based forwarders such as Splunk Universal Forwarder. This series explores an EC2-free alternative: configure ONTAP to write audit logs to an audit volume, expose that volume through an FSx for ONTAP S3 Access Point, use EventBridge Scheduler to invoke Lambda, and ship normalized events to observability platforms such as Datadog, Splunk, New Relic, Grafana Cloud, Elastic, and OpenTelemetry-compatible backends.

What This Post Covers

This post introduces the architecture and the open-source pattern library. It does not yet cover:

  • Full Datadog deployment walkthrough (Part 2)
  • Vendor-specific field mappings
  • Cost/performance benchmarking
  • ARP + EMS webhook + Datadog alerting (Part 3)
  • FPolicy binary protocol internals (future post)

The Problem Nobody Talks About

You're running Amazon FSx for NetApp ONTAP. You've enabled file access auditing because compliance requires it — or because you genuinely want to know who's accessing what on your file shares.

But where do those audit logs go?

If you followed the official AWS blog post, you likely ended up with EC2-based collectors: syslog-ng for cluster/admin audit forwarding, and a mounted audit volume plus Splunk Universal Forwarder for file access audit logs. It works. But now you have:

  • EC2 instances to patch and maintain
  • NFS mounts to the audit volume
  • syslog-ng configuration for admin audit forwarding
  • Splunk Universal Forwarder configuration for file access logs
  • A single point of failure unless you build your own HA pattern
  • Vendor lock-in to Splunk's agent-based model

What if you could replace that EC2-based collector pattern with managed services — Lambda reads audit logs via S3 APIs, no NFS mount required — and ship to any observability platform?

That's the goal of this project.

Important Distinction: Two Types of ONTAP Audit

Before diving in, a clarification. FSx for ONTAP has two distinct audit mechanisms:

  1. Cluster/admin activity audit logs — Administrative operations (CLI/API commands). These are forwarded via syslog to a log destination, as described in the AWS blog.

  2. File access audit logs — SMB/NFS file operations (open, read, write, delete, permission changes). These are recorded based on ONTAP audit policies and SACLs/NFSv4 ACLs, stored in EVTX or XML format, depending on your ONTAP audit configuration, on an audit volume inside the SVM.

In this series, "audit logs" refers to file access audit logs (type 2). The cluster admin audit forwarding via syslog is a separate concern.

The EC2-Free Alternative

I'm building an open-source pattern library that targets 9 observability vendors using Lambda, EventBridge Scheduler, and ECS Fargate — eliminating the need for self-managed EC2 instances.

This is EC2-free, not necessarily Lambda-only:

  • The audit-log and EMS paths are Lambda patterns (scheduled and event-driven respectively).
  • The FPolicy path uses ECS Fargate because ONTAP FPolicy requires a persistent TCP listener.

How Audit Logs Flow

ONTAP's file access auditing writes rotated audit log files to a configured destination path inside the SVM. In this project, that destination is an audit volume exposed through an FSx for ONTAP S3 Access Point. Lambda does not mount NFS or SMB; it reads the rotated audit log files through S3 APIs.

Because this pattern does not rely on S3 ObjectCreated events from FSx for ONTAP S3 Access Points, the audit processor is invoked on a schedule and uses checkpointing to process only newly rotated log files.

FSx for ONTAP audit configuration (`vserver audit`)
    │
    ▼ audit logs written to /audit volume
Audit volume exposed via FSx for ONTAP S3 Access Point
    │
    ▼ EventBridge Scheduler (periodic invocation)
Lambda audit processor (Python 3.12)
    │
    ▼ parse EVTX/XML → normalize → vendor API
Datadog / Splunk / New Relic / Grafana / Elastic / ...

Enter fullscreen mode Exit fullscreen mode

The key shift from the EC2 pattern: Lambda does not mount the audit volume over NFS or SMB. It reads rotated ONTAP audit log files through an FSx for ONTAP S3 Access Point using S3 APIs, while the data itself remains on the FSx for ONTAP file system.

A note on FSx for ONTAP S3 Access Points: FSx for ONTAP S3 Access Points let applications use S3 APIs to access data that still resides on FSx for ONTAP volumes. They are excellent as a serverless access boundary, but they are not the same as standard S3 buckets. In particular, you should not rely on S3 ObjectCreated notifications from an FSx for ONTAP S3 Access Point. Instead, this project uses EventBridge Scheduler plus checkpointing to discover and process newly rotated audit log files.

Three Event Sources, One Architecture

FSx for ONTAP generates observability data through three distinct channels:

1. File Access Audit Logs (FSx for ONTAP S3 AP)

Depending on your ONTAP audit configuration and SACL/NFSv4 ACL settings, file operations such as create, delete, read, write, and permission changes can be recorded as ONTAP audit logs in EVTX or XML format.

  • Delivery: ONTAP writes rotated audit log files to an audit volume inside the SVM
  • Access path: Lambda reads those files through an FSx for ONTAP S3 Access Point
  • Trigger: EventBridge Scheduler invokes Lambda periodically; Lambda uses checkpointing to process newly rotated files
  • Compute: Lambda (scheduled, pay-per-invocation)
  • Latency: Near-real-time rather than sub-second streaming. End-to-end latency depends on your ONTAP audit log rotation interval and the EventBridge Scheduler frequency.
  • Use case: Compliance auditing, access pattern analysis, data governance

2. EMS (Event Management System) Webhooks

ONTAP's built-in event system can push critical alerts via HTTP webhooks. This includes:

  • Autonomous Ransomware Protection (ARP) alerts — ONTAP detects encryption patterns and fires an event
  • Quota threshold violations
  • Hardware failures
  • Replication issues

  • Delivery: ONTAP pushes HTTPS webhook to API Gateway

  • Trigger: API Gateway invocation (event-driven)

  • Compute: Lambda (behind API Gateway)

  • Use case: Security alerting, operational monitoring

3. FPolicy (File Policy) Events

FPolicy intercepts file operations at the protocol level (CIFS/NFS) and forwards them in real-time via a proprietary TCP protocol. Unlike the other two sources, FPolicy requires a persistent TCP listener — which is why this path uses ECS Fargate rather than Lambda.

  • Delivery: ONTAP connects to Fargate task via TCP:9898
  • Trigger: Fargate receives FPolicy events → enqueues to SQS → Lambda processes
  • Compute: ECS Fargate (TCP listener) + Lambda (vendor shipping)
  • Use case: File activity monitoring, DLP, suspicious behavior detection

Note: The FPolicy path is the one exception to the "pure Lambda" model. ONTAP's FPolicy protocol is a proprietary binary format over TCP — it cannot be received by API Gateway or Lambda directly. Fargate handles the protocol translation, then hands off to Lambda via SQS for the vendor-specific shipping. It's still EC2-free, but not entirely serverless in the strictest sense.

The Architecture

Each event source feeds into the same delivery pattern:

┌─────────────────────────────────────────────────────────────────┐
│                    FSx for ONTAP                                │
├──────────────┬──────────────────────┬───────────────────────────┤
│ File Access  │   EMS Webhook        │   FPolicy (TCP:9898)      │
│ Audit Logs   │                      │                           │
└──────┬───────┴──────────┬───────────┴───────────┬───────────────┘
       │                  │                       │
       ▼                  ▼                       ▼
  FSx S3 AP +        API Gateway            ECS Fargate
  Scheduler               │                       │
       │                  ▼                       ▼
       ▼             Lambda (EMS)           SQS → Lambda
  Lambda (parser)         │                       │
       │                  │                       │
       └──────────────────┼───────────────────────┘
                          ▼
              Observability Vendor API
              (Datadog, Splunk, New Relic, ...)

Enter fullscreen mode Exit fullscreen mode

Each integration packages the parser and vendor shipper together in a single Lambda, but the pattern is the same: normalize ONTAP events, then send them to the vendor API. Swap the integration Lambda, and you switch vendors. Vendor-specific Lambdas are optimized for quick adoption and native API behavior, while the OpenTelemetry integration provides a vendor-neutral path for organizations standardizing on OTLP.

The Gotcha That Cost Me a Day

Here's something that isn't immediately obvious from the documentation:

In my validation, a Lambda function placed in a VPC with only an S3 Gateway Endpoint could not read from the FSx for ONTAP S3 Access Point and timed out. Adding NAT Gateway egress resolved the issue.

This gotcha matters because this project intentionally reads audit logs through FSx for ONTAP S3 Access Points rather than mounting the audit volume over NFS/SMB from an EC2 instance.

Tested with:

  • Lambda in private subnets (ap-northeast-1)
  • FSx for ONTAP S3 Access Point attached to an FSx volume
  • S3 Gateway VPC Endpoint only
  • No NAT Gateway
  • Failure mode: timeout (no response, not AccessDenied)

Your options:

Lambda Placement FSx for ONTAP S3 AP Access Recommendation
Outside VPC ✅ Works Simplest for read-only access
VPC + NAT Gateway ✅ Works Production recommended
VPC + S3 Gateway EP only ❌ Timeout Not recommended based on this validation

This is based on my validation environment (ap-northeast-1). Always test the network path in your own account and Region, as AWS may update this behavior.

Target Vendors

The project targets 9 observability platforms. Datadog is fully verified end-to-end (the subject of Parts 2 and 3 of this series). The remaining vendors have initial implementations that I'll be verifying and writing about in upcoming posts:

Vendor Delivery Method Status
Datadog Logs API v2 E2E verified
Splunk HEC (HTTP Event Collector) 🧪 Implementation ready, verification planned
New Relic Log API v1 🧪 Implementation ready, verification planned
Grafana Cloud Loki Push API 🧪 Implementation ready, verification planned
Elastic Bulk API 🧪 Implementation ready, verification planned
Dynatrace Log Ingest API v2 🧪 Implementation ready, verification planned
Sumo Logic HTTP Source 🧪 Implementation ready, verification planned
Honeycomb Events Batch API 🧪 Implementation ready, verification planned
OpenTelemetry OTLP/HTTP (vendor-neutral) 🧪 Implementation ready, verification planned

Status definitions:

  • E2E verified — Deployed and validated with real FSx for ONTAP audit logs
  • 🧪 Implementation ready — Code and CloudFormation available; E2E validation pending
  • 🚧 Planned — Design exists; implementation pending

Each vendor integration is designed as a self-contained CloudFormation stack with its own Lambda, IAM roles, DLQ, and CloudWatch alarms. As I verify each one, I'll publish a dedicated article with the results and any vendor-specific gotchas I encounter.

What's in the Repo

The project is structured for easy adoption:

fsxn-observability-integrations/
├── integrations/
│   ├── datadog/           # ✅ Verified: Lambda + CFn + tests + docs
│   ├── splunk-serverless/ # 🧪 Implementation ready
│   ├── new-relic/         # 🧪 Implementation ready
│   ├── grafana/           # 🧪 Implementation ready
│   ├── elastic/           # 🧪 Implementation ready
│   ├── dynatrace/         # 🧪 Implementation ready
│   ├── sumo-logic/        # 🧪 Implementation ready
│   ├── honeycomb/         # 🧪 Implementation ready
│   └── otel-collector/    # 🧪 Implementation ready
├── shared/
│   ├── lambda-layers/     # Reusable log parser (EVTX/XML) + S3 AP reader
│   ├── templates/         # Prerequisites CFn (EventBridge Scheduler, IAM)
│   └── scripts/           # Deploy + test utilities
└── docs/                  # Bilingual (EN/JA) documentation

Enter fullscreen mode Exit fullscreen mode

The shared infrastructure (EventBridge Scheduler, log parser layer, IAM roles) is vendor-agnostic and already proven through the Datadog verification. Each vendor directory follows the same structure, so once you understand one, you understand them all. Each stack is designed to include DLQ, CloudWatch alarms, and operational visibility out of the box; the Datadog stack also includes the verified CloudWatch operational dashboard used during E2E validation.

GitHub: github.com/Yoshiki0705/fsxn-observability-integrations

Related Posts

If you've been following my FSx for ONTAP S3 Access Points series, this project builds directly on those foundations:

This observability integrations project is the natural next step: taking those serverless patterns and applying them specifically to audit log shipping across multiple vendors.

Design Considerations

Based on early feedback, here are key points for different audiences:

Design philosophy: The goal is not just to remove EC2. The goal is to move undifferentiated collector operations into managed services, make failures observable and replayable, and keep the integration layer small enough for customers to operate themselves.

Where this pattern matters: This pattern is especially useful for enterprise file workloads where auditability matters but EC2-based collectors add operational overhead —
departmental file shares, enterprise application interface directories such as SAP, Oracle, or SQL Server adjacent file shares, VDI/EUC home directories, engineering and design repositories, regulated file repositories, and ransomware investigation workflows.

Non-intrusive by design: This pipeline observes audit logs after ONTAP records them; it does not sit in the application data path. NFS/SMB access patterns are unchanged. No application code changes are required.

Telemetry ownership: This pattern treats ONTAP as the authoritative source of file activity telemetry, while AWS managed services provide the event processing and delivery layer.

Compliance note: This pattern helps centralize and analyze audit events, but retention, immutability, and regulatory controls should be designed according to your organization's compliance requirements. This is an audit log delivery pattern, not a compliance certification. For audit evidence, consider separately how long raw EVTX/XML files should be retained on the audit volume or archived outside the observability pipeline.

Audit policy dependency: The quality and volume of events depend heavily on your ONTAP audit policy, SACLs, NFSv4 ACLs, and rotation interval. Enabling read auditing on high-traffic volumes can produce significant log volume — design your audit policy carefully.

Cost variables: The biggest cost factors are audit event volume, log rotation frequency, EventBridge Scheduler frequency, Lambda runtime, NAT Gateway usage (if Lambda is in VPC), and vendor ingest pricing. Compared to the EC2 pattern, you trade always-on instance cost for pay-per-invocation compute and vendor-ingest-driven cost.

Multi-account deployment: This pattern can be deployed per workload account or centralized into a logging/security account, depending on your organization's landing zone design.

Reliability: The stack includes DLQ for failed events, CloudWatch alarms for error/throttle detection, and checkpointing to avoid reprocessing already completed audit log files. Delivery to external vendor APIs should be treated as at-least-once; DLQ messages can be replayed after resolving the root cause.

What's Coming Next

This is Part 1 of a series. In the upcoming posts, I'll deep-dive into:

  • Part 2: Implementing the Datadog integration end-to-end — from CloudFormation to seeing logs in the Datadog Log Explorer
  • Part 3: Event-driven ransomware detection using ONTAP's Autonomous Ransomware Protection (ARP) + EMS webhooks + Datadog alerting

Beyond this Datadog series, I'll be verifying and writing about each vendor integration as I go:

  • Replacing the EC2-based Splunk pattern with Lambda + HEC
  • OpenTelemetry as the vendor-neutral escape hatch
  • Grafana Cloud + Loki for the open-source stack
  • And more — each with its own E2E verification and lessons learned

The goal is to build a comprehensive, battle-tested pattern library where you can pick your vendor and deploy with confidence. Follow along as I work through each one.

Try It Yourself

The Datadog integration is fully verified and ready to deploy. You'll need:

  • An FSx for ONTAP file system with audit logging enabled
  • An FSx for ONTAP S3 Access Point attached to the audit volume
git clone https://github.com/Yoshiki0705/fsxn-observability-integrations.git
cd fsxn-observability-integrations

# Deploy Datadog integration
# (FsxS3AccessPointArn = your FSx for ONTAP S3 Access Point ARN)
aws cloudformation deploy \
  --template-file integrations/datadog/template.yaml \
  --stack-name fsxn-datadog-integration \
  --parameter-overrides \
    FsxS3AccessPointArn=<your-fsx-s3-ap-arn> \
    DatadogApiKeySecretArn=<your-secret-arn> \
  --capabilities CAPABILITY_NAMED_IAM

Enter fullscreen mode Exit fullscreen mode

This stack deploys the scheduled Lambda processor, IAM permissions for reading from the FSx for ONTAP S3 Access Point, checkpoint storage, DLQ, CloudWatch alarms, and the Datadog shipping logic. The processor keeps track of already-processed audit log files so each scheduled invocation only ships newly rotated logs.

After deployment, you should see:

  • EventBridge Scheduler invoking the Lambda processor on your configured interval
  • Checkpoint storage updated after processing rotated audit logs
  • Parsed FSx for ONTAP audit events arriving in Datadog Logs (source:fsxn)
  • CloudWatch alarms and DLQ ready for operational visibility

Full setup guide in the repo's Prerequisites doc.


Have questions or want to see a specific vendor integration verified next? Drop a comment below — it'll help me prioritize the series.

Next up: Shipping FSx for ONTAP Logs to Datadog — The Serverless Way