惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
DataBreaches.Net
J
Java Code Geeks
有赞技术团队
有赞技术团队
H
Hackread – Cybersecurity News, Data Breaches, AI and More
The GitHub Blog
The GitHub Blog
T
The Blog of Author Tim Ferriss
Apple Machine Learning Research
Apple Machine Learning Research
量子位
D
Docker
V
Visual Studio Blog
博客园 - 司徒正美
Martin Fowler
Martin Fowler
人人都是产品经理
人人都是产品经理
WordPress大学
WordPress大学
U
Unit 42
M
MIT News - Artificial intelligence
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
G
Google Developers Blog
Engineering at Meta
Engineering at Meta
V
V2EX
大猫的无限游戏
大猫的无限游戏
雷峰网
雷峰网
Vercel News
Vercel News
C
Check Point Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
The 55.8 Percent Productivity Number From Doshi And Vaish...
A3E Ecosyste · 2026-05-20 · via DEV Community

When Doshi and Vaishnav published their controlled experiment on AI code completion in Science (2023), the headline that propagated everywhere was "55.8% faster." Repeat it enough and it becomes received wisdom.

The actual paper measured time-to-completion on a single well-defined HTTP server task. A problem with a known shape, a stable target, and a scoring function that rewarded a specific solution path. The 55.8% lift was real for that task. It is also the narrowest possible reading of what "AI productivity" means in software work.

A more careful follow-up at HICSS-59 (Stray et al., 2026) looked at sustained workflow integration over weeks instead of a single benchmarked task. Numbers compressed. Across mixed work (greenfield, debugging, refactoring, code review) aggregate time savings landed closer to 10-20%, with high variance across task class. Debugging and code review barely moved. Greenfield CRUD work moved the most.

That gap between single-task lab benchmark and integrated weekly workflow is where most engineering org AI productivity decisions are silently going wrong.

The mechanism gap

A code-completion model is doing one thing: predicting the next plausible token sequence given local context. Fantastic when the context is a half-finished function with a clear signature and the loss function would reward the standard completion. Much weaker when the work involves:

  • Tracing a bug through three repos and a queue
  • Deciding which refactor is worth doing
  • Reading existing code to understand intent before touching it
  • Negotiating a schema change with another team
  • Writing the test that catches the actual failure mode

None of those are next-token problems.

Where the gains actually compound

Builders shipping production AI workflows in 2025-26 are seeing real durable lift, but not by turning on Copilot and waiting. The compounding wins look like:

  1. Stack reduction. Skip a build step entirely. Replace a 4-step ETL with a single LLM-and-validator pass for cases where the validator can be trusted.

  2. Context elimination. Cut the time it takes to load a problem into working memory. Quick orientation queries on a strange codebase, API surface lookup, error message triage.

  3. Boilerplate elimination at the boundary. Form validators, type-mapping, mock data, fixture generation.

  4. Spec to first-draft compression. Get a structurally-correct first cut, then spend the saved time on the parts that need taste.

What this means for tooling decisions

Stop comparing AI tooling claims on single-task benchmarks. Ask vendors for sustained-workflow time-distribution data over weeks of real engineering work.

Measure your own lift the same way. Pick three task classes, instrument time-to-merge over a 4-week window, compare against baseline.

Hire for orchestration skill, not typing speed. The bottleneck moved.

Summary

The 55.8 percent number is not wrong, it is narrow. Sustained workflow integration data puts realistic aggregate productivity lift in the low double digits, concentrated in specific task classes.

Sources: Doshi and Vaishnav, Science 2023. Stray et al., HICSS-59 proceedings 2026.