惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

量子位
aimingoo的专栏
aimingoo的专栏
C
CXSECURITY Database RSS Feed - CXSecurity.com
Stack Overflow Blog
Stack Overflow Blog
C
CERT Recently Published Vulnerability Notes
T
Tailwind CSS Blog
腾讯CDC
罗磊的独立博客
Security Latest
Security Latest
K
Kaspersky official blog
A
Arctic Wolf
博客园 - Franky
D
Docker
博客园 - 司徒正美
GbyAI
GbyAI
T
Tenable Blog
Engineering at Meta
Engineering at Meta
A
About on SuperTechFans
H
Help Net Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
Lohrmann on Cybersecurity
小众软件
小众软件
V
V2EX
T
Threatpost
T
Threat Research - Cisco Blogs
T
The Exploit Database - CXSecurity.com
P
Palo Alto Networks Blog
P
Privacy & Cybersecurity Law Blog
S
Securelist
Google DeepMind News
Google DeepMind News
I
Intezer
The Register - Security
The Register - Security
NISL@THU
NISL@THU
L
LINUX DO - 热门话题
C
Cisco Blogs
AWS News Blog
AWS News Blog
MyScale Blog
MyScale Blog
S
Schneier on Security
Scott Helme
Scott Helme
T
The Blog of Author Tim Ferriss
G
Google Developers Blog
Project Zero
Project Zero
Cyberwarzone
Cyberwarzone
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
I
InfoQ
Cisco Talos Blog
Cisco Talos Blog
Know Your Adversary
Know Your Adversary
L
LangChain Blog
P
Proofpoint News Feed

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Your AI Agent Knows What to Do But, Does It Know How?
Seenivasa Ra · 2026-05-07 · via DEV Community

The missing piece in most LLM applications, and how AgentSkills fix it. We've gotten pretty good at telling AI agents who they are.

"You are an expert software engineer." "You are a seasoned marketing strategist." We hand them a persona, dump in some context, maybe paste in a few examples and then we hit send and hope for the best.

And for simple tasks? That works fine.

But the moment you ask an agent to do something that involves multiple steps, decisions, and potential failure points things start to fall apart in ways that are hard to predict and even harder to debug.

The agent sounds confident. It just doesn't behave consistently.

Here's why and what to do about it.

The Gap Nobody Talks About

There's a meaningful difference between knowing what needs to be done and knowing how to do it reliably.

A new employee on their first day might understand the goal perfectly "onboard this customer" but still flounder without a clear process. Do they send the welcome email first or set up the account? What if the system throws an error? Who do they escalate to?

Without a procedure, they improvise. Sometimes that works. Often it doesn't.

LLM agents have the exact same problem.

You can give an agent all the context in the world about what it's supposed to accomplish, and it'll still invent its own process every single time it runs. Skipping steps. Hallucinating validations. Silently glossing over failures.

This is the gap and it's where most LLM applications quietly break down.

Enter AgentSkills (and Why They're a Big Deal)

AgentSkills also called Procedure Skills are exactly what they sound like: explicit, step-by-step instructions that teach an agent how to execute a task, not just what the task is.

Think of it less like a prompt and more like a standard operating procedure. A playbook. A binder on the shelf.

Industry leaders like Anthropic and Microsoft have both converged on this idea and formalized it around a portable format called SKILL.md. That's not a coincidence it signals that the field is maturing from "prompt engineering" toward something more rigorous: procedure engineering.

What a Skill Actually Looks Like

A skill isn't a single prompt tucked inside a system message. It's a structured, self-contained unit of procedural knowledge a directory that bundles everything an agent needs to execute a specific type of task.

Here's how it breaks down:

SKILL.md is the core the instruction manual. It contains YAML frontmatter that lets the agent automatically discover and select the right skill for the job, plus detailed step-by-step execution instructions.

scripts/ holds small, single purpose automation scripts (Python, Bash, Node.js) for the steps that LLMs consistently get wrong when left to their own devices. Repetitive operations, file handling, API calls these belong in code, not in natural language instructions.

resources/ contains domain specific knowledge company standards, data schemas, regulatory rules anything the agent needs to reference but shouldn't be expected to memorize.

assets/ stores output templates. JSON schemas, document layouts, checklists so the agent produces consistent, structured results every time.

Put it all together and you get a self contained playbook instructions, tools, references, and templates in one place.

The Three Layers Most Teams Confuse

Before you can appreciate why skills matter, it helps to get clear on what they're not:

Most teams have prompts. Many now have tools. Very few have skills.

A skill is where workflow intelligence lives. It's the layer that answers the questions nobody bothers to write down: What comes first? What needs to be validated before moving on? What happens if this step fails?

Why Embedding All of This in a System Prompt Fails

The intuitive response to all of this is: "Can't I just put the procedure in the system prompt?"

You can. And for a single, small workflow it might work okay. But it breaks down fast for a few predictable reasons.

Fragility. Large, instruction-heavy prompts are brittle. One small tweak to the wording can cascade into completely different agent behavior. There's no modularity, no separation of concerns.

Token waste. Every time the agent runs, it pays the full token cost of every procedure even the ones that are completely irrelevant to the current task. At scale, this adds up fast.

Inconsistency. Without explicit validation steps ("check whether the file exists before editing it"), agents will invent shortcuts. They'll confidently skip steps and never tell you they did it.

The result is the thing that makes AI in production so frustrating: agents that sound certain and behave unpredictably.

The Idea That Changes Everything: Progressive Disclosure

Here's the mental model that ties this all together and it's dead simple.

Imagine your new employee's first day. You have two options:

Bad approach: Pile every binder all 50 of them on their desk. Tell them to read all of it before they start. By 11am they're exhausted, overwhelmed, and can't remember a thing.

Good approach: Put the binders on a shelf with clear labels. They glance at the labels, grab the one they need, read it, and do the job. Tomorrow, they grab a different one.

That's Progressive Disclosure.

In practice, it works in two phases:

Discovery Phase The agent loads only skill names and short descriptions. A table of contents for procedural knowledge. Minimal tokens, maximum orientation.

Activation Phase When a user request matches a skill's description, the agent loads the full SKILL.md and supporting assets into active memory. Only what's needed, only when it's needed.

The payoff is real: fewer hallucinations, lower token costs, better decisions when many skills exist simultaneously.

How to Design Skills That Actually Work

If you're going to build skills, these principles are worth internalizing from day one:

Write in third person imperative. "Extract the text." Not "You should try to extract the text." Precision matters ambiguous instructions produce ambiguous behavior.

Define failure states explicitly. What should the agent do when a script errors? When a file is missing? When validation fails? If you don't specify, the agent will improvise and you won't like the improvisation.

Keep skills small and composable. A skill called "Marketing" is a red flag. A skill called "Ad Copy Generation" is useful. A skill called "SEO Analysis" is useful. Small, focused skills compose into larger workflows. Monolithic skills just become another fragile mega prompt in disguise.

When Does This Actually Matter?

Not every situation calls for this level of structure. If you have one skill and it's always needed, just hand it to the agent upfront. Progressive disclosure doesn't help when there's nothing to disclose progressively.

But as your agent grows more tasks, more workflows, more edge cases the calculus changes:

  • 10 skills, one needed at a time? Huge savings. Show only what's needed.
  • 50 skills? Progressive disclosure becomes essential. Otherwise the agent drowns.
  • Complex multi step workflows? Explicit failure states and validation steps stop being nice-to-have and become the difference between an agent that works and one that confidently fails.

The Shift Worth Making

AgentSkills represent a genuine change in how we think about building with LLMs.

We're moving from prompt engineering which is ultimately about describing what we want to procedure engineering, which is about encoding how to reliably do it.

From probabilistic answers to deterministic execution.

From agents that talk about the work to agents that actually do it.

The tools and the personas are important. But without skills, you've hired a brilliant employee who has no idea how your company actually operates. Give them the binders. Label them clearly. Put them on the shelf.

That's the whole idea.

The one takeaway: MCP gives the LLM the tools. Skills tell the LLM when to use them. Progressive disclosure means "show only what's needed, when it's needed."

Thanks
Sreeni Ramadorai