惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 叶小钗
D
Darknet – Hacking Tools, Hacker News & Cyber Security
S
SegmentFault 最新的问题
博客园 - 三生石上(FineUI控件)
雷峰网
雷峰网
WordPress大学
WordPress大学
有赞技术团队
有赞技术团队
博客园 - 【当耐特】
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
V
V2EX
V
Visual Studio Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 聂微东
P
Proofpoint News Feed
Last Week in AI
Last Week in AI
U
Unit 42
W
WeLiveSecurity
博客园 - Franky
Recent Announcements
Recent Announcements
Hacker News - Newest:
Hacker News - Newest: "LLM"
Attack and Defense Labs
Attack and Defense Labs
月光博客
月光博客
The Cloudflare Blog
Spread Privacy
Spread Privacy
腾讯CDC
P
Privacy International News Feed
N
News and Events Feed by Topic
AWS News Blog
AWS News Blog
NISL@THU
NISL@THU
T
Troy Hunt's Blog
小众软件
小众软件
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Microsoft Security Blog
Microsoft Security Blog
L
Lohrmann on Cybersecurity
Webroot Blog
Webroot Blog
Y
Y Combinator Blog
量子位
P
Palo Alto Networks Blog
N
News and Events Feed by Topic
V
Vulnerabilities – Threatpost
K
Kaspersky official blog
IT之家
IT之家
T
Threat Research - Cisco Blogs
Cloudbric
Cloudbric
云风的 BLOG
云风的 BLOG
C
Check Point Blog
Blog — PlanetScale
Blog — PlanetScale
爱范儿
爱范儿
G
Google Developers Blog
S
Secure Thoughts

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
750,000 Chips, 140 Trillion Tokens: The Math Behind DeepSeek's Permanent Price Cut
keeper · 2026-05-23 · via DEV Community

DeepSeek made its V4-Pro 75% price cut permanent on May 22. The conventional read: "they got cheaper hardware." The real story is more interesting — and it's about a gap that's not closing fast enough.


What Happened

On May 22, 2026, DeepSeek announced that the 75% discount on its V4-Pro API would become permanent. The new pricing:

Metric Before After Cut
Input (cache miss) ¥12 / 1M tokens ¥3 / 1M tokens 75%
Output ¥24 / 1M tokens ¥6 / 1M tokens 75%
Input (cache hit) ¥0.1 / 1M tokens ¥0.025 / 1M tokens 75%

At current exchange rates, that's roughly $0.44/M input and $0.87/M output — making V4-Pro one of the cheapest frontier-class models on the market, on par with DeepSeek's own V4-Flash but with significantly more capability.

The move came exactly four weeks after V4's launch on April 24, and coincided with growing user frustration over rate limits at Google Gemini and Anthropic Claude.


The Standard Narrative

The surface-level story has three parts:

1. Architectural efficiency. V4 uses a Mixture-of-Experts architecture with 1.6 trillion parameters, but only activates a fraction per token. This gives it a structural cost advantage over dense models of comparable capability — roughly 30% of the gap.

2. Supply chain scaling. Huawei's Ascend 950PR entered mass production in April 2026. Huawei plans to ship ~750,000 units through the year — a 2.5x increase over 2025's 910C output. DeepSeek specifically optimized V4 for the Ascend architecture. More chips → lower unit cost → lower API pricing.

3. Competitive positioning. Western AI providers (Google, Anthropic) have been quietly tightening rate limits as demand overwhelms their GPU supply. DeepSeek is exploiting the backlash, offering unlimited usage at a fraction of the cost to capture disgruntled developers.

All three are true. But none of them fully explains the magnitude of the cut — or why it's permanent rather than promotional.


The Math That Changes Everything

Let's check the numbers.

Demand Side

China's daily token consumption hit 140 trillion in March 2026, according to the National Data Administration. The growth trajectory:

  • Early 2024: 0.1 trillion/day
  • End of 2025: 100 trillion/day
  • March 2026: 140 trillion/day

That's a 1,000x increase in two years, and a 40% jump in just the last quarter — implying ~13% month-over-month growth.

Supply Side

Huawei's mass-produced chip for 2026 is the Ascend 950PR (Prefill-optimized, 1 PFLOPS FP8), with the higher-end 950DT (2 PFLOPS FP8) coming in Q4. The numbers:

Chip FP8 Memory Bandwidth Inference Throughput (est.)
950PR 1 PFLOPS 128GB HBM 1.6 TB/s ~1,200 tokens/sec
950DT 2 PFLOPS 144GB HBM 4 TB/s ~2,400 tokens/sec

(Throughput derived from Huawei's published Atlas 950 SuperNode benchmark: 19.6M tokens/sec across 8,192 cards.)

Now the arithmetic:

Item Value
Total chips (2026 target) 750,000 (70% PR + 30% DT)
Raw daily throughput 85.7 trillion tokens/day
Inference-allocated (60%) 51.4 trillion tokens/day
vs Current demand (140T) 37% coverage
vs Demand in 6 months (~291T) 18% coverage

Even in the most optimistic scenario — every single chip dedicated to inference at 100% utilization:

Scenario vs Current vs +6 months
100% inference, 100% utilization 61% coverage 29% coverage

The conclusion is stark: 750,000 Ascend 950 chips can't cover today's demand — let alone the demand in six months.


So Why Cut Prices?

If supply is still a fraction of demand, permanent price cuts don't make sense in a normal market. But this is not a normal market.

The Real Logic: Pre-Commitment, Not Surplus

DeepSeek is not cutting prices because it has spare compute. It's cutting prices to lock in routing commitments before the hardware arrives.

Here's the timeline:

April 24:  V4 launched, optimized for Ascend
April 24+: ByteDance orders 350,000 Ascend 950 chips (~¥40B)
May 4:     Ascend 950PR mass production confirmed
May 22:    DeepSeek makes V4-Pro 75% cut permanent

Enter fullscreen mode Exit fullscreen mode

The critical insight: DeepSeek's price cut is not a cost pass-through. It's a market share pre-commitment — using the promise of future Ascend supply to grab developer mindshare now, before Western competitors can resolve their own capacity issues.

The Numbers Behind the Strategy

Western providers are capacity-constrained:

Provider Constraint Signal
Google Gemini TSMC CoWoS capacity Rate limits tightened, user backlash
Anthropic Claude H100/B200 availability API throttling, compute-use monitoring
OpenAI Inference cluster rollout Delayed GPT-5 token limits

DeepSeek's bet: "Spend the next 6 months building developer dependency on V4-Pro's API — by the time Ascend supply catches up in H2 2026, those developers won't switch back."

This is AWS in 2006. AWS wasn't cheaper than running your own servers in 2006. But it would be once scale kicked in. AWS priced for the scale it planned to have, not the scale it had. DeepSeek is doing the same.


What 750,000 Chips Actually Buys

The popular framing in Chinese media is "75万颗昇腾950产能大爆发." But as the math shows, 750,000 chips isn't abundance — it's barely adequacy.

Think of it this way: China's token demand is growing at roughly 0.5 trillion tokens per day every single month (the monthly increment itself is larger than the entire market 18 months ago). By year-end, demand will be 300-400+ trillion. Against that, 750K chips at the 950PR/DT mix buy roughly 50-85T/day of inference capacity.

Timeframe Demand (est.) Inference Supply Gap
March 2026 140T ~50T 90T
June 2026 ~200T ~50T 150T
September 2026 ~290T ~55T (DT ramp) 235T
December 2026 ~420T ~65T 355T

The gap is growing, not shrinking. Even with 75万 chips fully deployed, the supply-demand deficit more than triples over nine months.

This means DeepSeek's price cut isn't a sign of market saturation. It's a sign of exactly the opposite: a market so unsaturated that the winner gets to define the default API for an entire generation of developers, if they can lock them in before the hardware arrives.


Three Counter-Arguments (And Why They're Weak)

"But cache hits reduce the effective compute needed"

True — cache-hit tokens cost ~1/100th of miss tokens. And DeepSeek's cache hit rates can be high for workloads with stable system prompts. But cache hits are mostly in the input direction. Output tokens — the expensive ones — still need full compute. And as agentic workloads grow (multi-turn, chain-of-thought), output-to-input ratios increase, making cache less effective.

"But not all 140T tokens need 950-class inference"

Also true. Many tokens are generated by smaller models (Flash variants, Qwen, etc.) that don't need 950-level compute. But the growth is in the frontier-class tokens — longer context, more complex reasoning, higher quality requirements. That's exactly where 950-class chips are needed.

"But they can still buy H20 / smuggled H100"

H20 is less capable than 950PR per chip (the US-designed it to be worse). And the CHIPS Act + export controls have made H100 procurement increasingly difficult. Relying on smuggled hardware is not a supply chain strategy.


What This Means

For Developers

Your inference costs are likely going down over the next 12 months, not up — even though demand is exploding. That's unprecedented in any computing market. The driver isn't efficiency gains or manufacturing scale. It's a strategic subsidy by Chinese AI firms betting that locking in your API calls today is worth negative margins for a year.

Take the subsidy. But don't assume today's prices reflect tomorrow's costs — they reflect tomorrow's hopes.

For the Industry

The AI API market has entered a phase that looks like price war but functions like infrastructure land-grab. The playbook is AWS 2006, DoorDash 2019, Uber 2015: lose money on every transaction to own the default routing.

When the hardware does catch up — when Ascend 960 (2027) or 970 (2028) ships with 3-5x the throughput — the providers with the largest captive developer bases will convert negative margins to positive ones. Everyone else will be competing on price against incumbents they can't dislodge.


The Bottom Line

DeepSeek's permanent price cut is not evidence that Chinese AI compute supply has caught up with demand. The math shows it hasn't — and won't for at least 12-18 months. It's evidence that DeepSeek is playing the long game: use today's negative margins to own tomorrow's default inference route, and trust that Huawei's future chips will eventually close a gap that's currently 3-5x wider than headlines suggest.

The 75% cut isn't a cost breakthrough. It's a bet that developer lock-in is worth more than current margins — and that the 75万 Ascend 950 chips shipping this year are just the beginning.


Numbers sourced from: National Data Administration (China daily token data, March 2026), Huawei Connect 2025 (Ascend 950 specs and roadmap), SCMP/DW (ByteDance order volume), DeepSeek official pricing page (May 2026). Throughput calculations based on published Atlas 950 SuperNode benchmarks. Growth projections assume continuation of 40%/quarter rate per published data.