惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

大猫的无限游戏
大猫的无限游戏
H
Hacker News: Front Page
T
The Blog of Author Tim Ferriss
WordPress大学
WordPress大学
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Blog — PlanetScale
Blog — PlanetScale
Stack Overflow Blog
Stack Overflow Blog
F
Fortinet All Blogs
H
Help Net Security
罗磊的独立博客
D
DataBreaches.Net
MyScale Blog
MyScale Blog
美团技术团队
人人都是产品经理
人人都是产品经理
L
LangChain Blog
M
MIT News - Artificial intelligence
C
Check Point Blog
GbyAI
GbyAI
B
Blog RSS Feed
Microsoft Azure Blog
Microsoft Azure Blog
Y
Y Combinator Blog
雷峰网
雷峰网
Last Week in AI
Last Week in AI
F
Full Disclosure
量子位
V
Visual Studio Blog
Google DeepMind News
Google DeepMind News
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
S
SegmentFault 最新的问题
云风的 BLOG
云风的 BLOG
H
Hackread – Cybersecurity News, Data Breaches, AI and More
P
Proofpoint News Feed
爱范儿
爱范儿
A
About on SuperTechFans
MongoDB | Blog
MongoDB | Blog
腾讯CDC
博客园 - 【当耐特】
U
Unit 42
Martin Fowler
Martin Fowler
NISL@THU
NISL@THU
B
Blog
T
The Exploit Database - CXSecurity.com
Apple Machine Learning Research
Apple Machine Learning Research
L
Lohrmann on Cybersecurity
P
Proofpoint News Feed
有赞技术团队
有赞技术团队
C
CERT Recently Published Vulnerability Notes
The GitHub Blog
The GitHub Blog
T
Threatpost

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
20 Years of GPUs in Numbers: How FLOPS & TDP Grew, and Who Led the NVIDIA vs AMD Race (open dataset, 13.5k GPUs)
Max Vyazniko · 2026-05-26 · via DEV Community

We run a GPU spec catalog, and over a couple of years it grew into a database of 13,566 GPUs — from the GeForce 256 (1999) all the way to Blackwell and the MI355X (2025). At some point the interesting question stopped being "which card is faster" and became "how did the whole industry move?" — how much did FLOPS actually grow, where did the thermal envelope hit a wall, and who really led the NVIDIA-vs-AMD race in different years.

Here's a breakdown from our data. I'll give you two things up front that usually get hidden: the methodology (what I measured, where the data is noisy) and, at the end, an open dataset — grab it and dig in yourself.

TL;DR

  • Peak FP32 of the flagship grew ~400× in 19 years: 0.3 TFLOPS (GeForce 8800 GTX, 2006) → 126 TFLOPS (Blackwell, 2025). On a semi-log axis it's almost a straight line.
  • TDP crept up slowly (155 → 300 W over 2006–2020), then exploded in the datacenter: 700 W (H100), 1000 W (MI325X / B200), 1400 W (MI355X, 2025).
  • Meanwhile performance per watt grew ~100× — they "draw more" but "do far more per watt." The main driver is the process node (90 nm → 3 nm) plus architecture.
  • The NVIDIA/AMD duel by peak FP32 moved in waves: AMD led in the early 2010s (GCN era) and again in 2023–24 (Instinct MI300/MI325), NVIDIA in 2016–2020 (the AI pivot) and in 2025 (Blackwell). But "raw FP32" is a misleading metric, and I'll explain why below.

Methodology (please read before yelling at the numbers)

What these TFLOPS are and why they're "theoretical." Every FP32 figure here is the theoretical peak that vendors compute as:

FP32 TFLOPS = (shader ALUs / CUDA cores) × boost clock (Hz) × 2 / 10^12

Enter fullscreen mode Exit fullscreen mode

The ×2 is because an FMA (fused multiply-add) does a multiply and an add in one cycle — two operations. This is a ceiling, not real-world throughput: in practice you hit noticeably less — typically 60–90% on well-optimized compute-bound kernels, and a fraction of that on memory-bound ones — because memory bandwidth, SM occupancy, instruction mix, and the fact that boost clocks don't hold under sustained load and thermal limits all get in the way. Theory diverging from practice is normal. The theoretical peak is valuable for a different reason: it's computed by one formula across every card and generation, so it's a fair comparable yardstick for a historical look — which is exactly what spec sheets list and what we use. Real performance is measured with benchmarks (those are a separate table in the dataset).

A few more honesty notes:

  • "Flagship of the year" = the card with the maximum fp32_performance released that year, tracked separately for NVIDIA and AMD.
  • For the TDP/efficiency curves I excluded dual-GPU cards (GTX 295, HD 6990, R9 295X2, etc.) — otherwise the power and FLOPS double up and wreck the trend.
  • vendor is filled in for ~2,360 of 13,566 cards (the rest are mostly OEM/partner board variants). Medians use the labeled subset; flagship peaks are fully labeled.

1. FLOPS: an almost perfectly straight exponential

FP32 of flagship GPUs by year, log scale

Peak FP32 of the single flagship by year (NVIDIA):

Year Flagship FP32, TFLOPS
2006 GeForce 8800 GTX 0.3
2010 GeForce GTX 580 1.6
2013 GeForce GTX 780 Ti 5.3
2016 Quadro P6000 12.6
2017 Tesla V100 15.7
2020 RTX A6000 38.7
2022 L40S 91.6
2025 RTX PRO 6000 Blackwell 126.0

≈400× in 19 years is a CAGR of about 37% per year. On a semi-log axis the line is almost perfectly straight — a classic exponential that only recently started bending on the desktop segment and moved into the datacenter.

2. TDP: a quiet climb, then a datacenter explosion

TDP of flagship GPUs by year

Year Card TDP, W
2006 GeForce 8800 GTX 155
2010 GTX 580 244
2017 Tesla V100 250
2020 RTX A6000 300
2022 H100 SXM 700
2024 MI325X / B200 1000
2025 MI355X 1400

For a decade and a half the flagship thermal envelope stayed in a 150–300 W band. The break comes after 2020, and it's entirely datacenter-driven: AI accelerators (SXM/OAM modules) jumped to 700–1400 W because they're cooled by liquid in a rack, not by a fan in a case. The desktop ceiling separately hit ~450–600 W (RTX 4090/5090).

There's a curious gap if you look only at NVIDIA's consumer flagships: the GeForce flagship sat at exactly 250 W for seven years (2013–2019) — GTX 780 Ti, Titan X, 1080 Ti, 2080 Ti — and only broke that ceiling with the RTX 3090 (350 W, 2020), then 4090 (450 W) and 5090 (575 W). Datacenter accelerators, by contrast, went to 700–1400 W almost immediately. It looks like what capped gaming TDP wasn't the silicon so much as the market — cases, PSUs, and buyer habits; in a rack none of those constraints apply, and watts grew without looking back. (That's interpretation, of course — the spec stores watts, not intentions — but a 250 W plateau across seven generations shows up clearly in the data.)

3. Performance per watt: this is where the real progress is

FP32 per watt by year

If you only look at TDP it feels like "everything's getting worse, cards guzzle power." But FP32 per watt tells the opposite story:

Year Flagship TFLOPS/W
2006 8800 GTX 0.002
2013 GTX 780 Ti 0.021
2016 Quadro P6000 0.051
2020 RTX A6000 0.129
2022 L40S 0.262
2025 RTX PRO 6000 Blackwell 0.21

~100× in efficiency. Peak "classic" efficiency lands in 2022 (Ada/L40S); the 2024–25 datacenter cards sometimes lose on TFLOPS/W because they deliberately trade efficiency for absolute compute density in the rack. The main drivers of efficiency are the process node (90 nm → 3 nm) and architecture, not clock speed.

4. The NVIDIA vs AMD duel

NVIDIA vs AMD FP32 leadership timeline

If you mark, year by year, whose single flagship had the higher FP32:

Period Leader Context
2007–2008 AMD FireStream 9170/9270
2010–2013 AMD GCN: HD 6970, HD 7970 GHz, R9 290X
2014 NVIDIA Titan Black (5.6) vs FirePro W9100 (5.2)
2015 AMD Fury X (8.6)
2016–2020 NVIDIA Pascal → Ampere, the AI pivot
2021 AMD Instinct MI250X (47.9)
2022 NVIDIA L40S / Hopper
2023–2024 AMD Instinct MI300A/MI325X (81.7)
2025 NVIDIA Blackwell (126)

The picture is wavy, and honestly I included it partly to give AMD its due — because on raw FP32, AMD led more often than people remember, both in the GCN era and again on recent Instinct parts. But raw FP32 is a deceptive metric for the modern world. The AI era isn't won on FP32; it's won on FP16/BF16/FP8 and on software.

And here's the catch that trips up most "raw FP16" comparisons:

FP16 vs FP32 by vendor, dense values

FP16/tensor numbers are not directly comparable between vendors because of structured sparsity. Starting with Ampere (A100), NVIDIA's spec sheets quote tensor FP16/BF16 figures with sparsity already applied — that's 2× the dense value (the feature processes sparse matrices twice as fast). AMD has no equivalent spec line — those are dense. So NVIDIA's raw FP16 numbers (A100+) need to be halved to compare fairly with AMD: A100 = 624 (sparse) → 312 dense, H100 = 1979 → ~990 dense. With tensor cores (since V100, 2017) and the CUDA ecosystem, NVIDIA built a moat that raw FP32 simply doesn't show.

H100 isn't one card — it's a family

A good example of why "an H100" isn't yet a spec: several different SKUs of the same chip live under that name in the database. TDP swings 2× — from 350 W on the PCIe 80 GB version to 700 W on the SXM5 modules; there's also an NVL at 94 GB and 400 W. Memory is 80 / 94 / 96 GB, and the PCIe-80 ships on slower HBM2e while SXM and the higher PCIe parts are on HBM3. Even peak FP32 "floats" by ~30%: 51 TFLOPS on the PCIe-80 vs 67 on any SXM5 (same silicon, different clocks and power). So when a price list or an email just says "H100," it's worth asking which one — otherwise you'll compare a 350 W PCIe card to a 700 W SXM module as if they're the same thing. In the dataset, each is its own row.

Variant TDP VRAM Memory FP32
H100 PCIe 80 GB 350 W 80 GB HBM2e 51.2
H100 NVL 94 GB 400 W 94 GB HBM3 60.3
H100 PCIe 96 GB 700 W 96 GB HBM3 62.1
H100 SXM5 (80/94/96 GB) 700 W 80–96 GB HBM3 66.9

Open dataset — take it

We published a cleaned dump of our GPU spec database for anyone who wants to dig in themselves — it's on the GPU Ark open dataset page (CSV, SQLite, and a single .tar.gz):

  • 13,566 GPUs (fields: vendor, manufacturer, release date, architecture, process node, transistors, clocks, memory size and type, bus, FP16/FP32/FP64/BF16/TF32/INT8, TDP, NVLink, CUDA SM, and more) + 993 third-party benchmark results (join on gpu_id).
  • Formats: CSV (Excel/pandas) and SQLite (ready-made SQL) — two tables, gpu_specs and benchmarks.
  • License: CC BY 4.0 (attribution to gpuark.com).
  • Only public specs — no competitor prices, no internal mappings, no editorial content.
  • A KNOWN_ISSUES.md sits next to it with the same caveats as above (sparse vendor, VRAM units on old cards, FP16 noise), plus a README.md with a five-line pandas example.

If you'd rather explore interactively before downloading, the same data powers the GPU comparison tool on the site.

Takeaways

  1. FLOPS grew as an almost perfect exponential (~37%/yr) — but the "free" growth is over; from here we pay with thermal envelope and a move into the rack.
  2. Real progress is measured not in watts and not in raw FP32, but in performance per watt (×100) — and that rides on the process node.
  3. AMD fought and led on "raw" numbers more often than the common narrative admits; but the AI era was defined by tensor + software, not FP32.

The data is open — if you find something in it we missed, tell me and I'll add it.