惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Google DeepMind News
Google DeepMind News
MongoDB | Blog
MongoDB | Blog
有赞技术团队
有赞技术团队
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
人人都是产品经理
人人都是产品经理
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
B
Blog RSS Feed
T
Tor Project blog
T
Threat Research - Cisco Blogs
Microsoft Azure Blog
Microsoft Azure Blog
M
MIT News - Artificial intelligence
V
Vulnerabilities – Threatpost
Project Zero
Project Zero
C
CXSECURITY Database RSS Feed - CXSecurity.com
The Register - Security
The Register - Security
Latest news
Latest news
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
The Hacker News
The Hacker News
Google DeepMind News
Google DeepMind News
L
LINUX DO - 最新话题
U
Unit 42
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园 - 司徒正美
T
Tenable Blog
H
Hacker News: Front Page
B
Blog
宝玉的分享
宝玉的分享
C
Check Point Blog
美团技术团队
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
C
CERT Recently Published Vulnerability Notes
P
Proofpoint News Feed
The GitHub Blog
The GitHub Blog
G
GRAHAM CLULEY
Google Online Security Blog
Google Online Security Blog
Security Archives - TechRepublic
Security Archives - TechRepublic
P
Proofpoint News Feed
GbyAI
GbyAI
酷 壳 – CoolShell
酷 壳 – CoolShell
Hugging Face - Blog
Hugging Face - Blog
Y
Y Combinator Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Hacker News - Newest:
Hacker News - Newest: "LLM"
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Scott Helme
Scott Helme
L
Lohrmann on Cybersecurity
量子位
A
About on SuperTechFans
V2EX - 技术
V2EX - 技术
T
The Exploit Database - CXSecurity.com

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Inside Systems 01: Your Verification Process Did Not Break. It Was Replaced.
Mr Chandrava · 2026-05-16 · via DEV Community

Why fluent AI outputs quietly replace verification habits without announcing the change

The Rule Vikram Built

Vikram had a rule he followed for seven years.

Never forward a number you did not personally trace to its source.

Not because his manager told him. Because in 2017 he sent a procurement estimate to a client that was built on a number a colleague had provided without sourcing it, and the error cost him three weeks of repair work and one relationship that never fully recovered. The rule came from that. It was specific, operational, and his.

By late 2024, he had been using the tool for five months. He was faster. His first drafts were cleaner. The sourcing rule had quietly stopped applying to outputs the tool produced, not as a decision, not as an exception he consciously granted, but as a gap that opened without being named.

He did not notice the gap. The outputs were good.

The Substitution Happened Quietly

The illusion operating here is precise enough to name in one sentence.

Fluency looks like diligence, and diligence is what verification is supposed to confirm, so fluency quietly substitutes for the confirmation without either party agreeing to the substitution.

That sentence is not an accusation. It is a description of how calibration actually works under repeated exposure to high-quality outputs.

The tool produces language that has the texture of careful work. Complete sentences, confident framing, appropriate hedges in the right places, specific-sounding figures. Every surface signal that trained readers associate with reliable material is present. The signal that is absent, the one that says "this was independently checked against a primary source," does not have a surface form. It was never visible. You inferred it from the other signals.

For seven years, those signals correlated strongly enough with actual verification that treating them as proxies was reasonable. Then the tool arrived and produced the same surface signals from a completely different process.

The inference kept running. The correlation had broken.

What The Architecture Actually Learned

Gary Marcus has made this point repeatedly in technical contexts: the gap between a system producing correct-shaped output and a system that has any access to whether the output is correct is not a limitation to be patched in the next version.

It is the architecture.

These systems were trained on text. Text that was mostly written by people who had verified their claims before writing them. The training absorbed the surface form of verified writing, the vocabulary, the cadence, the citation-adjacent phrasing, without absorbing the verification process that produced it.

When you read the output, you receive all the surface markers of diligent work. The work behind those markers is not there. It was never part of what was learned.

This is not a criticism of the tool. It is a description of what it is. The problem is not using a tool that produces fluent unverified text. The problem is using a tool that produces fluent unverified text and having your judgment respond as though the verification happened.

Verification Became The Expensive Part

The mechanism that makes this persistent is not carelessness. It is something more structural.

Verification is expensive relative to reading. It requires locating primary sources, cross-referencing specific claims, catching the specific failure modes of specific subject areas. It is the slow part of intellectual work, the part that does not scale, the part that does not get faster with practice in the same way that writing does.

When a tool compresses the writing phase dramatically, the ratio changes. If writing a first draft took two hours and verification took one, the ratio was 2:1. When the tool compresses the writing phase to twenty minutes, the same absolute verification time is now 1:3 in the other direction. Verification becomes the dominant cost.

Under cost pressure, the behavior that gets reduced is the expensive one. This is not a moral failure. It is how systems respond to changed cost ratios. The behavior looks the same from outside, the output still arrives, still gets sent, still performs the function. What changed is invisible in the output itself.

What changed is how much of the process that used to precede the output is still happening.

What Expertise Is Actually Built From

Ted Chiang would locate the deeper problem one layer below the verification question.

Expertise in any analytical field is partly the accumulated history of catching your own errors before they leave your hands. The researcher who learned not to trust rounded numbers in secondary sources because one wrong rounded number sent her down a three-week dead end in 2019. The engineer who runs a specific sanity check on every cost estimate because he learned in 2021 that a particular class of error was invisible until the project was already committed.

That knowledge is not transferable through instruction. It is built through the specific experience of being wrong in recoverable situations, where the cost was high enough to register but not high enough to be catastrophic.

When the tool produces the first draft, the occasions for that experience change. You are no longer building from raw material and encountering your own errors on the way to a finished product. You are editing a finished-seeming product and encountering errors, when you encounter them, in a different cognitive position. The editor's relationship to a draft's errors is different from the author's. Both catch things. They do not build the same instincts.

The instinct Vikram built in 2017 required being the person who passed the unsourced number. The tool creates conditions where that experience is less likely to occur. Which means the conditions that would update or reinforce that instinct are also less likely to occur.

The instinct does not disappear. It atrophies through reduced exercise in the specific conditions that maintain it.

The Calibration Problem

The cost that is already running is not located in any single output.

It is located in the calibration itself.

Vikram's judgment about what needs verification is being trained on tool outputs. When an output is accurate, he files: this kind of output does not require checking. When an output is accurate again, he files the same. Over five months of mostly accurate outputs, the prior probability that any given output needs checking has shifted downward without a deliberate decision.

This is rational updating on available evidence. The problem is that the evidence set is missing the failure cases that would correct it. The tool's errors are not randomly distributed across outputs. They cluster in specific domains, specific types of claims, specific question shapes where the training produced confident-sounding text that does not correspond to verifiable fact. Those failure modes are not visible in the surface form of the output. You encounter them by checking, and the rate at which you check has been declining.

The calibration is updating on a sample that excludes the data most relevant to calibration.

The Identity Still Feels Intact

The identity this behavior protects is worth naming directly.

Vikram considers himself someone who does not cut corners on accuracy. This is not self-flattery in his case. He has a documented history of the opposite, built over seven years including one expensive correction that produced a lasting rule. The professional identity is real and earned.

The behavior of not applying the sourcing rule to tool outputs is not experienced as cutting a corner. It is experienced as appropriate adaptation to a new tool. A different kind of process, not a reduced one.

Both things are simultaneously true: the identity is accurate, and the behavior contradicts what the identity would require if applied consistently. They coexist because the new process feels substantively different from what the rule was designed to prevent, even when the failure mode it prevents is identical.

The rule was built to catch unsourced numbers passed to clients. The tool produces unsourced numbers with excellent surface form. The rule does not automatically transfer because the experience of using the tool does not feel like the experience that created the rule.

The Internal Argument

Here is the internal argument that does not resolve cleanly.

One side: the tool is a productivity tool, using it well requires calibrating trust to actual performance, the outputs are reliable enough that applying the full verification protocol to every output would eliminate the efficiency gain that justifies using it, treating the tool with blanket suspicion is not a practical working posture.

Other side: the failure modes are invisible in the surface form, the calibration is updating on a biased sample, the instincts being maintained through reduced exercise are the ones most needed for the cases the tool gets wrong, and those cases are exactly the ones where no surface signal flags the need for checking.

Both arguments are structurally sound. Neither defeats the other.

The working resolution most people land on is implicit rather than explicit: they verify when something feels uncertain and trust when it does not. The problem with that resolution is that the tool's errors frequently do not feel uncertain. They arrive feeling like the outputs that were correct, because they were produced by the same process.

Feeling uncertain is not a reliable trigger for verification when the signal is absent in the class of errors that matter most.

The Practical Adjustment

The practical adjustment is narrower than it sounds.

Not full verification of every output. That eliminates the tool's value and is not what careful work requires even without the tool.

One question, asked after reading any output that will be passed forward: what specifically would need to be true for this to be correct, and have I confirmed any of it independently?

The question does not require answering yes. It requires being asked.

Vikram's 2017 rule was not "verify everything." It was "never forward a number you did not personally trace." The precision of the rule is what made it operational. It covered the specific failure mode he had encountered. It did not require treating everything with equal suspicion.

The equivalent for tool outputs is a similarly narrow rule, derived from the actual failure modes of the actual tool you are using, applied to the specific class of claims where those failure modes cluster.

That rule does not exist in generic form. You have to build it from your own failure cases.

Which means you have to be present when the failure cases occur.

Vikram's March Incident

Vikram found his failure case in March, eight months into regular use.

A unit cost figure in a supplier analysis. Accurate-sounding, appropriately specific, surrounded by a well-constructed paragraph. Wrong by 23 percent. The paragraph was so well constructed that he had read it as confirmation of a figure he vaguely remembered rather than as a claim requiring independent verification.

He traced the error. He updated his practice. He now asks the question after every output that contains specific quantitative claims.

The adjustment took one incident to make and five minutes per output to run.

The incident cost two days of rework and one client conversation he would have preferred not to have.

The question he did not ask for eight months was: what would I need to confirm for this to be usable?

He knew how to ask it. The tool had trained him, through eight months of mostly accurate outputs, to ask it less often.

That is what the tool actually did to his process. Not through any design intent. Through the same mechanism that any reliable system uses to reduce the perceived need for verification: it was right often enough that checking felt redundant, until the specific case where it was not.