惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

爱范儿
爱范儿
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
G
GRAHAM CLULEY
www.infosecurity-magazine.com
www.infosecurity-magazine.com
V2EX - 技术
V2EX - 技术
The Last Watchdog
The Last Watchdog
S
Secure Thoughts
Webroot Blog
Webroot Blog
PCI Perspectives
PCI Perspectives
L
LINUX DO - 最新话题
Hacker News: Ask HN
Hacker News: Ask HN
N
News and Events Feed by Topic
H
Heimdal Security Blog
H
Help Net Security
T
The Blog of Author Tim Ferriss
P
Proofpoint News Feed
The GitHub Blog
The GitHub Blog
Jina AI
Jina AI
Recent Commits to openclaw:main
Recent Commits to openclaw:main
F
Full Disclosure
小众软件
小众软件
S
Securelist
罗磊的独立博客
NISL@THU
NISL@THU
D
Darknet – Hacking Tools, Hacker News & Cyber Security
C
Cisco Blogs
云风的 BLOG
云风的 BLOG
C
CERT Recently Published Vulnerability Notes
Cisco Talos Blog
Cisco Talos Blog
Know Your Adversary
Know Your Adversary
S
Schneier on Security
D
DataBreaches.Net
M
MIT News - Artificial intelligence
V
Vulnerabilities – Threatpost
N
News and Events Feed by Topic
有赞技术团队
有赞技术团队
F
Fortinet All Blogs
T
Tenable Blog
The Register - Security
The Register - Security
C
Check Point Blog
AWS News Blog
AWS News Blog
Cloudbric
Cloudbric
C
CXSECURITY Database RSS Feed - CXSecurity.com
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
C
Cyber Attacks, Cyber Crime and Cyber Security
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Google Online Security Blog
Google Online Security Blog
博客园 - 叶小钗
Hacker News - Newest:
Hacker News - Newest: "LLM"
博客园 - 司徒正美

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The End of "One-Shot AI": Why Context Engineering Is Replacing Prompt Engineering
Yao Xiao · 2026-06-29 · via DEV Community

Most "prompt engineering" advice circulating today is already obsolete for anyone building production-grade AI. Granular phrasing matters for simple, single-turn tasks — but the moment a system involves retrieval, memory, tool calls, or multi-step reasoning, the wording of your prompt becomes a second-order variable.

Drawing from years of building high-performance quantitative data pipelines, the principle is familiar: optimizing a model with corrupted or incomplete input data never works, regardless of how elegant the model itself is. Context engineering applies this same rigorous logic to LLMs. The industry's definitive shift toward context engineering isn't a rebranding — it's the structural and mathematical foundation required to scale AI reliability beyond the single-turn demo.

What Prompt Engineering Actually Is (and Where It Stops)

Prompt engineering is the practice of crafting and refining the specific text you send to a model. Given a fixed model and a fixed task, better phrasing produces better outputs. That is real and measurable.

The problem is its scope. A prompt is a single input to a stateless transaction. The model sees it, generates a response, and the interaction is over. For simple, one-turn tasks — generate a summary, classify this text, rewrite this paragraph — prompt engineering is genuinely sufficient.

But production AI systems are rarely single-turn. They involve multi-step reasoning, access to external knowledge, memory of prior interactions, tool calls, retrieved documents, and structured constraints. At that level, the wording of your prompt becomes a second-order concern. What matters is what the model has access to when it runs.

Why Production Models Fail: Context Rot, Hallucinations, and Buried Instructions

Most LLM failures in production are not model failures. They are context failures.

The model hallucinates because it doesn't have the right reference material available. It drifts off-topic because the system prompt is competing with a wall of unrelated chat history. It gives a generic answer because the user's specific constraints were never surfaced in the context window. The instruction was there — it just got buried.

This phenomenon has been documented in the research literature under terms like "lost in the middle" — the observation that language models systematically underweight information placed in the center of long contexts, even when that information is directly relevant. A 2026 paper on arXiv (2603.09619) formalizes this with the concept of context rot: as more information is pushed into the context window without curation, the model's effective attention on any individual piece degrades. More context, counterintuitively, can mean worse performance.

The fix isn't a better prompt. The fix is better context architecture.

What Context Engineering Actually Means

Context engineering is the discipline of designing the entire information environment a model operates in — not just the words you type, but the full stack of what it knows at inference time. Every component that touches the context window is in scope: inference time latency, token optimization, semantic search quality, retrieval ranking, and memory summarization.

That includes:

  • The system prompt: the standing instructions, persona, and constraints
  • Retrieved documents: chunks from a vector database, API results, knowledge base entries
  • Memory: summaries of prior sessions, user preferences, established facts
  • Tool outputs: the results of function calls, code execution, search results
  • Conversation history: selectively filtered to avoid diluting attention

A skilled context engineer treats all of these as a pipeline, not as afterthoughts appended to a clever instruction. The question isn't "how should I word this?" It's "what does the model need to know, in what order, at what granularity, to perform well on this task?" This is closer to software architecture than to copywriting — and in my experience, engineers with a background in strict data pipeline design adapt to it faster than anyone else.

The Same Question, Two Completely Different Systems

Abstract distinctions are easy to miss. Here's a concrete scenario.

The situation: A user types into an AI customer service assistant — "Can I still get a refund for my order?"

What the prompt engineer does: Refines the instruction. The system prompt becomes something like: "You are a helpful, empathetic customer service agent. Answer refund questions accurately and concisely." The model receives the user's message plus this instruction, and generates a response.

The result: the model responds in the right tone — polite, professional. But it doesn't know the user's order date, the company's actual return policy, or whether the specific product category is even refund-eligible. It either hallucinates a policy number, gives a generic "please contact support" deflection, or asks five clarifying questions that a real agent would have already known the answers to.

What the context engineer does: Before a single token of the prompt fires, the system executes a pipeline:

  1. Retrieves the user's order record from the database — order date: June 1, item: consumer electronics, delivery status: confirmed
  2. Fetches the current return policy from the knowledge base — electronics category: 15-day return window, no exceptions for opened items
  3. Computes the time delta — today is June 23, which is 22 days post-purchase, outside the return window
  4. Filters conversation history — drops the last four irrelevant exchanges about shipping, keeps only the one message referencing the order number
  5. Injects all of this as structured context above the system prompt, clearly labeled by source

The model now sees: the user's order date, the exact policy, the computed eligibility status, and a clean conversation history. It responds: "Your order from June 1 falls outside the 15-day return window for electronics. A standard refund isn't available at this point. If there are extenuating circumstances — a defective item, for example — I can escalate this to our support team directly."

Same model. Same base instruction. The output is unrecognizably better — not because the prompt was better, but because the context was engineered.

The key insight: the prompt engineer asked "how do I say this?" The context engineer asked "what does the model need to know before I say anything?" That is the entire difference.

The Five Properties of a Well-Engineered Context

The arXiv paper (2603.09619) proposes five production-grade criteria for evaluating context quality. Use these as a diagnostic checklist against any failing pipeline:

  1. 🎯 Relevance — Include only information that bears on the current task. Irrelevant content doesn't disappear from the model's attention; it competes with relevant content for it.
  2. Sufficiency — The model must have enough information to answer correctly without guessing. Insufficient context causes hallucination just as reliably as incorrect context does.
  3. 🔒 Isolation — Separate task-specific context from global state and prior conversation history. Mixing long cross-session history with an immediate instruction is one of the most common causes of degraded performance.
  4. 💰 Economy — Every unnecessary token carries a cost: money, inference time latency, and attention. A bloated context window is not a safety net; it's a liability.
  5. 🔍 Provenance — In high-stakes applications, the model must be able to trace where each piece of information came from. This matters for auditability and for calibrating source trust at inference time.

Author's note: When I audit failing AI pipelines, the root cause almost never turns out to be the prompt. It's context violating one of these five criteria — usually Relevance or Economy. Teams add retrieval, history, and tool outputs, then never prune any of it. The context window becomes a landfill, and the model's outputs reflect that exactly.

RAG Is Not a Feature — It's a Context Engineering Problem

Retrieval-Augmented Generation has become standard in enterprise AI. But most teams implement it as a plumbing problem: connect the database, retrieve the top-k chunks, append them to the prompt. Done.

The performance gap between teams that treat RAG as plumbing and teams that treat it as a context engineering challenge is substantial. The hard questions aren't about retrieval recall — they're about what to do with retrieved content once you have it.

How do you handle retrieval failures gracefully? How do you prevent retrieved text from contradicting the system prompt? How do you tell the model which document to trust when two retrieved chunks say different things? How do you maintain coherent reasoning across a multi-step chain where each step retrieves different context?

If you're working on RAG systems and running into unexplained accuracy regressions, the Advanced RAG Prompting Strategies piece breaks down exactly where most retrieval pipelines fail at the prompt-context interface — it's a useful companion to this topic.

Memory Systems: The Missing Layer in Most AI Architectures

Single-session AI interactions are increasingly the exception. Users return. They have preferences, prior work, established context. And yet most AI deployments reset completely with every new session.

This is a context engineering gap, not a model limitation. Modern systems handle this through persistent memory architectures: episodic memory (summaries of past sessions), semantic memory (long-term facts and preferences), and working memory (the current task state). The model doesn't inherently "remember" anything — but a well-engineered system can surface the right memories into its context at the right time, making it behave as if it does.

Memory systems turn one-shot interactions into cumulative relationships. This is where significant productivity gains live for power users and enterprise deployments alike — and it's entirely orthogonal to the quality of any individual prompt.

For a practical breakdown of how memory, planning, and tool access work together as a system, the Memory, Planning, Tools: The Three Pillars article maps the architecture clearly for anyone building or using agentic workflows.

Prompt Engineering as a Subset, Not a Replacement

Context engineering doesn't make prompt engineering obsolete. The instruction you put in the system prompt still matters. The phrasing of a few-shot example still matters. The role definition still matters.

What changes is the hierarchy. Prompt engineering becomes a component of context engineering — the layer that handles the instruction format, tone, and constraint specification within an already well-designed information environment.

Think of it like this: a skilled author chooses words carefully. But choosing words carefully inside a structurally broken outline still produces a bad piece of writing. The context is the outline. The prompt is the sentence-level craft. Both matter, but the outline comes first.

Practical Pitfall: The "Just Add More Context" Trap

One of the most common mistakes I see teams make when they first learn about context engineering is treating it as a license to stuff more information into the context window. More documents, more history, more examples — surely more is better?

It isn't. Context engineering is fundamentally about curation, not accumulation. The goal is the minimum sufficient context: exactly what the model needs, nothing it doesn't. Every redundant token increases cost, increases latency, and dilutes attention on what actually matters.

A useful mental model: treat your context window like a whiteboard in a focused meeting. A clean whiteboard with the right information drives good decisions. A whiteboard covered in every note from every meeting for the past six months drives confusion.

What This Means for How You Work

For engineers building production AI systems, the implication is architectural: context design needs to be a first-class concern from the start. Retrofitting context management onto a system that was built purely around prompt iteration is painful and usually incomplete.

For knowledge workers using AI tools, the implication is more immediate. You can start practicing context engineering right now by being intentional about what you surface to the model before asking your question: relevant documents, prior decisions, constraints, the specific sub-task at hand. This is what experienced AI users do intuitively — they prepare the context before firing the prompt.

To enforce this architectural discipline at the individual prompt level, I built Prompt Scaffold — a strictly local-first, zero-backend in-browser tool. By forcing you to define Role, Task, Context, Format, and Constraints before a single token ever leaves your device, it eliminates the underspecified instructions that silently break RAG pipelines and agentic workflows. Because it runs entirely in your browser with no server involved, you can engineer highly sensitive prompt context — proprietary business logic, internal data schemas, confidential constraints — with absolute data privacy. For teams operating in regulated environments or handling sensitive data, that's not a nice-to-have; it's a hard requirement.

Where the Field Is Going

The terminology is settling. "Context engineering" was informal jargon in early 2025; by mid-2025 it had been adopted by Anthropic, Google, and LangChain as the preferred frame for discussing production AI system design. The arXiv paper formalizing its criteria situates it within a four-level maturity model: Prompt Engineering → Context Engineering → Intent Engineering → Specification Engineering. Each level abstracts upward from the previous.

The direction is clear. As models become more capable and agent systems become more complex, the leverage in the stack shifts further from the individual prompt and further toward the systems that determine what the model knows when it runs.

Prompt engineering was always a workaround for the absence of better tooling. Context engineering is what fills that gap.

What You Can Do This Week

Start auditing your existing AI interactions or pipelines against the five context quality criteria: Relevance, Sufficiency, Isolation, Economy, Provenance. You don't need new tooling to do this — you need the right diagnostic frame.

For each failure case you're seeing, ask: is this a prompt problem (bad instruction) or a context problem (wrong information available)? The answer will tell you where to spend your improvement effort.Most of the time, it's the context.