惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
AWS News Blog
AWS News Blog
Y
Y Combinator Blog
云风的 BLOG
云风的 BLOG
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
F
Full Disclosure
H
Help Net Security
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
A
About on SuperTechFans
J
Java Code Geeks
Jina AI
Jina AI
GbyAI
GbyAI
酷 壳 – CoolShell
酷 壳 – CoolShell
爱范儿
爱范儿
美团技术团队
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Latest news
Latest news
Vercel News
Vercel News
博客园 - 【当耐特】
P
Privacy & Cybersecurity Law Blog
P
Proofpoint News Feed
阮一峰的网络日志
阮一峰的网络日志
V
Vulnerabilities – Threatpost
Stack Overflow Blog
Stack Overflow Blog
Hugging Face - Blog
Hugging Face - Blog
D
Docker
Microsoft Security Blog
Microsoft Security Blog
博客园_首页
S
Securelist
WordPress大学
WordPress大学
S
Secure Thoughts
博客园 - 聂微东
Cloudbric
Cloudbric
Help Net Security
Help Net Security
腾讯CDC
T
Threat Research - Cisco Blogs
T
Tor Project blog
L
LINUX DO - 热门话题
Project Zero
Project Zero
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - Franky
N
Netflix TechBlog - Medium
小众软件
小众软件
Cyberwarzone
Cyberwarzone
量子位
MyScale Blog
MyScale Blog
W
WeLiveSecurity
MongoDB | Blog
MongoDB | Blog
I
InfoQ
M
MIT News - Artificial intelligence

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
From YAML to AI Agents: Building Smarter DevOps Pipelines with MCP
Nimesh Kulka · 2026-05-23 · via DEV Community

From YAML to AI agents: building smarter DevOps pipelines with MCP

DevOps teams have spent years turning manual work into YAML.

That helped. CI runs on every pull request. Deployments can be triggered from a commit. Kubernetes can reconcile desired state. Terraform can plan infrastructure before it changes anything.

But a lot of DevOps work still sits outside the pipeline:

  • reading failed CI logs
  • checking whether a deployment is safe
  • connecting traces, alerts, recent commits, and infra changes
  • deciding whether to roll forward or roll back
  • writing the same runbook steps again and again
  • asking five tools for the same incident context

This is where AI automation gets interesting. Not as a magic replacement for DevOps engineers, but as a better interface for operational work.

The strongest version of this stack is not just "AI in CI/CD." It is an AI-native DevOps layer built around three pieces:

  1. MCP servers for tool access
  2. Skills for repeatable expert workflows
  3. Plugins for company-specific infrastructure actions

If you build it well, the pipeline gets faster because the boring glue work disappears. If you build it badly, you get an AI bot with production credentials and vague judgment. That is not automation. That is a future incident report.

Why MCP matters for DevOps

MCP, or Model Context Protocol, gives AI applications a standard way to connect to external systems.

The official MCP docs describe three main server-side primitives:

  • tools: functions an AI app can call, like file operations, API calls, database queries, or deployment actions
  • resources: context an AI app can read, like docs, schemas, logs, runbooks, or service metadata
  • prompts: reusable templates for structured workflows

That maps cleanly to DevOps.

A platform team could expose separate MCP servers for:

  • GitHub or GitLab
  • CI/CD logs
  • Kubernetes
  • Terraform or OpenTofu
  • Argo CD
  • Prometheus, Grafana, Datadog, or OpenTelemetry backends
  • cloud cost data
  • incident management
  • internal service catalog

The AI agent does not need to scrape random dashboards or guess from partial screenshots. It can ask real tools for real state.

For example:

User: Why did the production deploy fail?

Agent flow:
1. Read the failed GitHub Actions job logs.
2. Check the changed files in the pull request.
3. Query Argo CD for sync status.
4. Read Kubernetes events for the affected namespace.
5. Pull recent error traces from observability.
6. Summarize the likely failure and suggest the smallest safe fix.

Enter fullscreen mode Exit fullscreen mode

That is not replacing the DevOps engineer. It is removing the tab-hopping tax.

Skills are where the real expertise lives

MCP gives the agent access. Skills tell it how to work.

A skill is a reusable procedure for a specific job. In DevOps, that matters because production work has rules. You do not want an agent inventing a deployment strategy every time someone asks a question.

Good DevOps skills could look like this:

skill: debug_failed_ci
steps:
  - fetch failed jobs
  - group logs by failure type
  - check if failure is test, lint, dependency, infra, or runner-related
  - compare against recent commits
  - suggest the smallest code or config fix
  - never rerun expensive jobs more than once without approval

Enter fullscreen mode Exit fullscreen mode

skill: safe_kubernetes_rollout
steps:
  - check current deployment health
  - verify image tag and git SHA
  - check recent incidents for the service
  - confirm SLO status before rollout
  - deploy to one environment first
  - watch error rate, latency, and pod readiness
  - stop if guardrail thresholds fail

Enter fullscreen mode Exit fullscreen mode

skill: terraform_plan_review
steps:
  - read the Terraform plan
  - classify adds, changes, and destroys
  - flag IAM, networking, database, and public exposure changes
  - check cost-sensitive resources
  - summarize blast radius
  - require human approval for destructive or privilege-expanding changes

Enter fullscreen mode Exit fullscreen mode

This is the part I think most people miss. The value is not just that an AI can call tools. The value is that it can call tools through a workflow your team already trusts.

Plugins make it fit your company

Every company has weird infrastructure.

Maybe your deploys go through Argo CD, but production still needs a Slack approval. Maybe your Terraform state is split across workspaces. Maybe the service catalog is internal. Maybe your rollback process depends on a custom CLI that only three people understand.

Plugins are how you expose that reality safely.

A plugin can wrap a company-specific action like:

  • get_service_owner(service_name)
  • fetch_deploy_risk_score(pr_number)
  • create_change_request(environment, service, sha)
  • run_internal_canary(service, image_tag)
  • open_incident_with_context(summary, traces, logs)
  • estimate_cloud_cost_diff(terraform_plan_id)

The plugin should not give the agent unlimited shell access and vibes. It should expose narrow, typed actions with logs, permissions, and guardrails.

A good internal DevOps plugin feels boring:

{
  "name": "request_production_deploy",
  "input": {
    "service": "checkout-api",
    "image_tag": "2026.05.23.4",
    "change_summary": "Fix timeout handling in payment gateway client",
    "risk_level": "medium"
  },
  "requires_approval": true,
  "audit_log": true
}

Enter fullscreen mode Exit fullscreen mode

Boring is good here. Boring means it can survive production.

Who should use this?

This stack is useful for a bunch of specialists, but each one should use it differently.

DevOps engineers can use it to debug CI/CD failures faster, generate release notes, identify flaky jobs, and automate repetitive deployment checks.

Platform engineers can turn internal developer platforms into agent-accessible systems. Instead of making every developer learn five dashboards, they can expose safe workflows through MCP servers and skills.

SREs can use it for incident triage: correlate alerts, attach traces, find recent deployments, pull service ownership, and suggest runbooks.

Cloud infrastructure engineers can use it to review Terraform plans, detect risky IAM changes, estimate cost impact, and standardize provisioning workflows.

Release engineers can use it to decide whether a release is ready, what changed, what failed, what needs approval, and what rollback path exists.

DevSecOps engineers can connect security checks into the pipeline: secret scanning, policy checks, dependency review, artifact provenance, image scanning, and permission drift.

AI infrastructure engineers can use the same pattern to manage model-serving deployments, GPU capacity, eval gates, prompt/version rollouts, and inference observability.

The common thread is simple: if your job involves reading state from multiple systems and taking careful action, AI agents can help. But only if you give them structured tools and clear operating procedures.

A practical AI-native CI/CD pipeline

Here is a realistic pipeline architecture.

AI-native DevOps pipeline diagram

The shape is simple: pull request, CI checks, AI agent, MCP tool layer, guardrails, then GitOps/deploy automation. The agent speeds up context gathering. The pipeline still owns execution, approval, and audit history.

Pull request opened
        |
        v
CI runs tests, lint, security checks
        |
        v
Agent reads CI result through MCP
        |
        v
If failed:
  - summarize failure
  - identify likely owner
  - suggest fix
  - open comment with exact logs and files

If passed:
  - read diff
  - check Terraform or Kubernetes changes
  - classify deployment risk
  - verify service ownership and runbook
  - prepare deploy summary
        |
        v
Human approval for production
        |
        v
Argo CD / deploy tool syncs desired state
        |
        v
Agent watches rollout health
        |
        v
If healthy:
  - close deploy task
  - attach release summary

If unhealthy:
  - collect logs, traces, events
  - recommend rollback or roll-forward
  - require approval before mutation

Enter fullscreen mode Exit fullscreen mode

This is faster because the agent handles context collection. It is safer because the agent does not blindly mutate production.

That balance matters.

Where Kubernetes and GitOps fit

Kubernetes already works like an automation platform. Controllers watch desired state and reconcile actual state. The Kubernetes docs describe the controller pattern as programs that read an object's spec, act on it, and update status.

GitOps tools like Argo CD build on that idea. Argo CD treats Git as the source of truth, compares live cluster state against desired state, and syncs when needed.

AI should not replace that control loop.

It should sit above it.

The agent can explain what changed, detect risk, connect symptoms to recent deploys, and recommend action. Kubernetes and Argo CD should still do the actual reconciliation with clear audit history.

That gives you the best version of both worlds:

  • deterministic infrastructure control loops
  • human-readable operational reasoning
  • faster triage
  • safer approvals

Observability is the agent's fuel

An AI DevOps agent is only as good as the context it can retrieve.

OpenTelemetry matters here because it gives teams a common way to collect traces, metrics, and logs. Traces are especially useful because they show the path of a request across services.

For an agent, this context can answer questions like:

  • Did the error start after a deployment?
  • Which dependency is adding latency?
  • Is this one service failing or a full user journey?
  • Are we seeing infrastructure failure, code failure, or traffic shape change?
  • Did the rollback actually improve user-facing symptoms?

Without observability, the agent is just guessing politely.

Guardrails that should exist before production actions

If an AI agent can touch production, the guardrails need to be boring and strict.

Start with these:

  • read-only by default
  • separate permissions per environment
  • mandatory approval for production mutations
  • audit logs for every tool call
  • allowlists for safe actions
  • dry-run support for infrastructure changes
  • policy checks before apply
  • rollback plans attached to deploy actions
  • rate limits for repeated retries
  • no secret exposure in prompts or logs

Terraform run tasks are a good mental model. HCP Terraform can call external systems between plan and apply, show messages in the run pipeline, and block the apply phase when needed. That is exactly the kind of control point AI automation should respect.

Do not start by letting an agent run kubectl delete in prod.

Start by letting it explain what it would do, why it would do it, what it needs approval for, and how to undo it.

What to build first

If I were building this for a real DevOps team, I would not start with autonomous deployment.

I would start with three narrow workflows.

First: CI failure explanation.

The agent reads failed logs, groups the error, identifies the likely cause, links exact lines, and comments on the PR. Low risk, high value.

Second: Terraform plan review.

The agent summarizes infrastructure changes, flags destructive actions, points out IAM/network/database risk, and asks for review before apply.

Third: deployment health summary.

The agent watches rollout status, error rate, latency, pod readiness, recent traces, and recent alerts. It posts one clean summary instead of making someone manually check six tools.

Once those work, you can add more automation.

A good maturity path looks like this:

Animated AI DevOps pipeline flow

Level 1: Read-only assistant
Level 2: Suggests fixes and runbooks
Level 3: Opens tickets, comments, and summaries
Level 4: Runs approved low-risk actions
Level 5: Handles narrow autonomous remediation with hard guardrails

Enter fullscreen mode Exit fullscreen mode

Most teams should live at Level 2 or Level 3 for a while. That is not slow. That is how trust gets built.

The main takeaway

AI-native DevOps is not about replacing YAML with a chatbot.

It is about giving DevOps specialists a faster way to move through the work they already do: gather context, understand risk, apply a known workflow, and take the next safe action.

MCP gives the agent a standard way to reach tools. Skills give it repeatable expert behavior. Plugins make it fit the company's real infrastructure.

The result is a better pipeline:

  • faster CI/CD debugging
  • cleaner infrastructure reviews
  • safer releases
  • better incident context
  • less repetitive manual work

The best DevOps AI systems will not be the ones that act the most independently. They will be the ones that know when not to act.

Start with read-only context. Add skills. Wrap dangerous actions in plugins with approvals. Then automate the boring work first.

That is how AI makes DevOps faster without making production scarier.

References

  1. Model Context Protocol, Architecture overview https://modelcontextprotocol.io/docs/learn
  2. Model Context Protocol, Understanding MCP servers https://modelcontextprotocol.io/docs/learn/server-concepts
  3. GitHub Docs, GitHub Actions documentation https://docs.github.com/actions
  4. GitHub Docs, Workflows https://docs.github.com/en/actions/concepts/workflows-and-actions/workflows
  5. Argo CD Docs, Declarative GitOps CD for Kubernetes https://argo-cd.readthedocs.io/en/latest/
  6. Kubernetes Docs, Extending Kubernetes https://kubernetes.io/docs/concepts/extend-kubernetes/
  7. HashiCorp Developer, Set up HCP Terraform run task integrations https://developer.hashicorp.com/terraform/cloud-docs/integrations/run-tasks
  8. OpenTelemetry Docs, OpenTelemetry Concepts https://opentelemetry.io/docs/concepts/
  9. OpenTelemetry Docs, Traces https://opentelemetry.io/docs/concepts/signals/traces/