惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
Cyber Attacks, Cyber Crime and Cyber Security
Cisco Talos Blog
Cisco Talos Blog
Scott Helme
Scott Helme
The Last Watchdog
The Last Watchdog
G
GRAHAM CLULEY
T
Tenable Blog
PCI Perspectives
PCI Perspectives
Simon Willison's Weblog
Simon Willison's Weblog
N
News and Events Feed by Topic
Know Your Adversary
Know Your Adversary
S
Schneier on Security
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
P
Privacy International News Feed
C
CERT Recently Published Vulnerability Notes
NISL@THU
NISL@THU
SecWiki News
SecWiki News
S
Securelist
D
Docker
阮一峰的网络日志
阮一峰的网络日志
人人都是产品经理
人人都是产品经理
T
Tailwind CSS Blog
T
Troy Hunt's Blog
The Register - Security
The Register - Security
K
Kaspersky official blog
Blog — PlanetScale
Blog — PlanetScale
云风的 BLOG
云风的 BLOG
Hacker News: Ask HN
Hacker News: Ask HN
S
Secure Thoughts
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
博客园 - 司徒正美
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
F
Fortinet All Blogs
T
Threatpost
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
小众软件
小众软件
WordPress大学
WordPress大学
Security Archives - TechRepublic
Security Archives - TechRepublic
博客园 - 聂微东
Attack and Defense Labs
Attack and Defense Labs
B
Blog RSS Feed
Project Zero
Project Zero
Y
Y Combinator Blog
T
The Blog of Author Tim Ferriss
博客园 - 【当耐特】
V
V2EX
Help Net Security
Help Net Security
P
Proofpoint News Feed
A
Arctic Wolf

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Our VP's AI Wrote 3,000 Tests. Production Cost $700K. I Deleted Every Single One.
xulingfeng · 2026-06-05 · via DEV Community

Based on real industry trends. About an AI testing tool promising 300x efficiency, a VP who rebranded hand-written automation as "manual testing," and a $700K SLA bill nobody saw coming.


Act 1: The All-Hands

VP Harrison stood in front of the screen, the AI testing dashboard glowing behind him.

"Three days. Three thousand test cases. Zero human intervention."

He paused. Let his eyes sweep the room. They landed on me.

"And some people — six years. Maintained four hundred automated test cases. That's less than twenty per person per year."

A few people looked at their phones. Others studiously avoided my direction.

"I'm not here to debate efficiency. I'm here to ask — why does your team still exist?"

I opened my notebook.

"Mr. Harrison. What's the coverage on those three thousand AI-generated tests?"

"One hundred percent."

"And how many bugs did they find?"

A beat.

"The first phase is about regression coverage, not —"

"Zero." I cut him off. "Three thousand tests, zero bugs. You ran three thousand checks on 'what the code does' and not a single one on 'what the code should do.' That's coverage, not quality."

VP Harrison smiled.

"I understand your anxiety. When new technology threatens your domain, people always rationalize resistance. But the data doesn't lie — three hundred times the efficiency, zero incremental cost. What took you six years to prove — AI did in three days."

He didn't look at me again. Clicked to the next slide.


Act 2: The Sideline

That afternoon, HR notified me: my team was being reassigned from Quality Assurance to the AI Engineering Group. My reporting line now went through VP Harrison's deputy.

I walked to his office.

"Mr. Harrison. The AI testing tool needs a three-week trial run. I need to verify its behavior under production traffic patterns —"

"You don't need to verify it. I already did."

"What environment?"

"Staging. One hundred percent pass rate."

"Staging doesn't simulate real traffic shapes —"

"Are you telling me your four hundred manual test cases are more effective than three thousand AI-generated ones?" He leaned back. "Do you actually believe that?"

I stood at his office door. He didn't ask me to sit.

"Your new desk is on the third floor. AI Engineering Group. Report tomorrow."


Act 3: The Report

I spent three nights pulling and reviewing all 3,000 AI-generated test cases.

The AI tool itself wasn't bad. The problem wasn't the technology.

The problem was the configuration. VP Harrison's deputy had set the input boundary to "90th percentile of historical production data." The AI faithfully generated tests within that boundary. All three thousand cases lived inside the 90th percentile. Inside that range, the AI validated "what the code does according to the config." It couldn't — by design — validate "what the code should do at the boundaries." The config never asked it to look there. The AI didn't err. It flawlessly executed a flawed instruction set.

I wrote a full analysis report with configuration screenshots and comparative data.

Sent it to VP Harrison. No CC.

Twenty-three minutes later, his reply landed:

"Noted. The edge scenarios you identified have an estimated probability below 0.3%. Per our risk prioritization framework, we will not allocate resources to cover them. I suggest you focus on learning the new tools rather than finding reasons to reject them."

I read that line twice.

Then I filed the report into a folder called RCA_2026Q3 and went back to maintaining my test suite.


Act 4: The Rollout

Three weeks later. The AI tests went live on the main release pipeline.

VP Harrison published a piece in the company newsletter:
"Why We Retired Manual Testing — And Why Your Team Might Be Next"

One line found me in the company-wide email:

"Some people spent three weeks trying to prove AI wouldn't work. Two weeks in production — zero incidents. Sometimes, what you're resisting isn't the technology's flaws. It's your own insecurity."

The one remaining tester on my team walked over to my desk.

"Boss... that part about insecurity. Was that about you?"

I closed the email tab.

"Two weeks zero incidents. Let's see what week three brings."


Act 5: The $700K Breakdown

1:14 AM. PagerDuty lit up like a Christmas tree.

A module the AI tests had cleared — thanks to that "90th percentile" boundary — hit a data race condition under real traffic. Every AI-generated test ran inside "normal traffic" parameters. Not a single test covered "resource contention when call frequency exceeds threshold." Because the AI's configuration never told it to check.

Cascading failure. Core transaction pipeline down. Nine hours of data recovery.

Initial damage: $700K.

The CTO called an RCA. Meeting time: Monday, 9 AM. Attendees: VP level and above... and me.


Act 6: The Meeting Room

9 AM. The conference room.

The CEO walked in. Didn't sit. Stood at the head of the table and placed a printed report on the surface.

"Mr. Harrison. You go first."

VP Harrison cleared his throat.

"This was a tool-level edge case. The AI testing framework lacks built-in detection for this scenario. We've contacted the vendor — the next release will include a fix."

The CEO listened standing. He didn't interrupt.

When Harrison finished, he closed the folder and looked around the table.

"Before we went live — did anyone raise a concern like this?"

Silence. VP Harrison said nothing. The CTO studied his laptop.

Three seconds.

I opened my notebook.

"Yes. One month ago."

The CEO's eyes shifted to me.

"A report analyzing the AI testing tool's configuration — input boundary set at the 90th percentile, leaving twenty-three categories of low-probability, high-impact scenarios uncovered. Including the race condition that caused last night's outage."

CEO: "Who did you send it to?"

"Mr. Harrison."

I plugged my laptop into the conference room projector. The email screenshot filled the screen.

"To: Mr. Harrison. Sent: June 7, 11:23 PM."

Several people pulled out their phones to photograph it.

The CEO glanced at the screen, then back at VP Harrison.

"You received it?"

"I did."

"And?"

A pause.

"At the time, the assessment was —"

I flipped to the next page.

"Noted. The edge scenarios you identified have an estimated probability below 0.3%. Per our risk prioritization framework, we will not allocate resources to cover them. I suggest you focus on learning the new tools rather than finding reasons to reject them."

The CEO read it. Nodded once. No raised voice. No theatrics.

He placed the report back on the table and looked at VP Harrison.

"He sent it to you. Why didn't you escalate?"

VP Harrison had no answer.


Act 7: The Aftermath

VP Harrison submitted his resignation two days later.

Internal memo from the CTO: Quality Assurance restored as an independent division, reporting directly to the CTO. Budget doubled. I was appointed department head.

That afternoon, I walked back to my old desk on the fifth floor. Still empty. Everything where I'd left it.

A yellow sticky note on the corner of my monitor. Not mine — left by a former teammate who'd left the company months ago.

It read: "Don't let them touch your tests."

I never knew if he meant the VP, or the AI.

I peeled it off and tucked it into the first page of my notebook.

I opened the AI testing platform. Typed test_case list --all --source ai. 3,000 records.

Select all. Delete. Confirm.

"This action cannot be undone. Delete 3,000 test cases?"

Confirm.

Three thousand cases. Three seconds. Gone.

The one remaining tester on my team stood behind me.

"Boss... you just deleted everything?"

"The configuration was wrong. Every test was built on a broken foundation. If the foundation is crooked, no test on top of it will save you."

He didn't answer.

I opened our four hundred automated test cases — six years of writing, one line at a time. Four hundred cases on the legacy system. Zero production incidents in six years.

"What about the new module?" he asked.

"We write them. Starting today."

But I didn't close the AI testing platform.

I opened its test generator. Pasted the new module's API spec. Changed the boundary config from the default 90th percentile to — unlimited. Generate everything. I'll curate.

The AI generated 87 candidates.

I reviewed every single one. Kept 42. Deleted 45. Added 8 boundary scenarios the AI never considered.

Fifty cases merged into our test suite.

Four hundred human-written, plus fifty AI-assisted — running together.

The tester looked at the green checkmarks on the dashboard.

"Boss... you're using AI too?"

"The AI isn't the problem. The problem is who decides what it tests."

"The VP bought AI as a mask. I use AI as a microscope."


"AI makes mistakes. Humans make mistakes. But the worst part is — someone stacks both mistakes on top of each other, then blames it all on the AI."


AI-generated tests pass at 100%. They verify what the code does — not what it should do. When the code itself is wrong — AI will prove the wrong code is right.


Have you seen someone package a process failure as a technology failure? What happened next?

Follow for more stories about AI testing, quality engineering, and what happens when the tools are smarter than the process.


More stories like this:

If these stories made you think or saved you time, buy me a coffee ☕ — currently maintaining 400 automated test cases and manually reviewing every AI-generated one. Caffeine is the only config I trust at 💯.