惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

N
Netflix TechBlog - Medium
S
Secure Thoughts
aimingoo的专栏
aimingoo的专栏
量子位
F
Full Disclosure
云风的 BLOG
云风的 BLOG
D
Docker
Stack Overflow Blog
Stack Overflow Blog
Last Week in AI
Last Week in AI
Security Latest
Security Latest
Project Zero
Project Zero
Cyberwarzone
Cyberwarzone
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
AWS News Blog
AWS News Blog
G
GRAHAM CLULEY
T
Threat Research - Cisco Blogs
The Hacker News
The Hacker News
The Register - Security
The Register - Security
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
MyScale Blog
MyScale Blog
博客园_首页
阮一峰的网络日志
阮一峰的网络日志
P
Privacy & Cybersecurity Law Blog
V
Visual Studio Blog
A
Arctic Wolf
L
LINUX DO - 热门话题
W
WeLiveSecurity
Google DeepMind News
Google DeepMind News
TaoSecurity Blog
TaoSecurity Blog
Recorded Future
Recorded Future
K
Kaspersky official blog
L
Lohrmann on Cybersecurity
V
V2EX
Recent Announcements
Recent Announcements
N
News and Events Feed by Topic
MongoDB | Blog
MongoDB | Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
S
Securelist
S
Security Affairs
Forbes - Security
Forbes - Security
P
Proofpoint News Feed
Google DeepMind News
Google DeepMind News
C
Cisco Blogs
AI
AI
Vercel News
Vercel News
博客园 - 三生石上(FineUI控件)
Cisco Talos Blog
Cisco Talos Blog
T
Tor Project blog
Scott Helme
Scott Helme
Spread Privacy
Spread Privacy

Forbes - Innovation

Why Do Humans Have Fingerprints? Hint: It’s Not What You Think Booking.com Confirms Data Breach, Reservation PIN Codes Changed Why Major News Sites Are Blocking The Internet Archive’s Wayback Machine iPhone Fold Release Date: New Report Details Frustrating Apple News Comet Tracker: How To See Pan-STARRS And Three Planets On Wednesday NYT Mini Crossword Today: Tuesday, April 14 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Tuesday, April 14 (It’s A Little Unclear) Today’s Wordle #1760 Hints And Answer For Tuesday, April 14 Most Of The Microplastics In Urban Air Come From Tires Today’s Wordle #1759 Hints And Answer For Monday, April 13 NYT Mini Crossword Today: Monday, April 13 Hints And Answers NYT Pips Today: Hints, Answers And Walkthrough For Monday, April 13 The YC Chief Who Codes 10,000 Lines A Day Has A Simple Secret Samsung Expands One UI 8.5 Beta To More Galaxy Owners Why You Should Stop Using Your iPhone If It’s On This List Chamath Says Firms That Treat AI As A Strategy Hand Rivals Their Edge 3 Unexpected Habits Of Secure Couples, By A Psychologist The First Lamp That Folds Your Clothes Samsung’s Disappointing Price Update For Galaxy Phone Buyers 3 Subtle Signs Someone Is Falling In Love With You, By A Psychologist Do Mantis Shrimp See More Colors Than Humans? A Biologist Explains NYT Connections Answers Explained For Monday, April 13 (#1,037) NYT Connections Hints Today: Monday, April 13 Clues And Answers (#1,037) LEGO Luigi & Mach 8 (72050) Review: 2026’s Best Set Yet? Marc Andreessen Says AI Productivity Will Trigger A Hiring Boom 3D Printing Is The Ultimate Hack To Reduce Household Spending Apple iPhone Fold: Striking Design Revealed In Leaked Photos Apple Smart Glasses: New Leak Reveals A Major Design Twist To Beat Meta Tested: The AI Coming To The Rivian R2 Quordle Hints Today: Monday, April 13 Clues And Answers Companies And H-1B Employees Endure Immigration Waits At Consulates 3 Easy Ways To Turn Anxiety Into Sustained Focus, By A Psychologist Here’s The Most Affordable Humanoid Robot You Can Buy Now UFC 327 Results: 5 Biggest Takeaways From A Wild Night In Miami UFC 327 Results, Bonus Winners, Highlights And Reactions Dana White Announces Huge New Fight For UFC White House Today’s NYT Strands Hints, Spangram, Answers: Sunday, April 12 (Get Ready) Tesla ‘Model 2’ Rises From The Ashes Today’s Wordle #1758 Hints And Answer For Sunday, April 12 NYT Pips Today: Hints, Answers And Walkthrough For Sunday, April 12 Tyson Fury Vs. Arslanbek Mahkmudov Results: Highlights and Reaction NYT Mini Crossword Today: Sunday, April 12 Hints And Answers How Shadow AI Culture Is Destroying Your Business Venture Capital Funds That Market Like Startups Win More Deals Conor Benn Vs. Regis Prograis Results: Highlights and Reaction Samsung’s Disappointing Price Update For Galaxy Phone Buyers Artemis Reached The Moon. The Grid Can Reach The 21st Century A Biologist Explains How Archerfish Shoot Down Prey. Hint: Their Aim Rivals Human Throwing Is It Time For Apple To Forget About The MacBook Air NYT Connections Hints Today: Sunday, April 12 Clues And Answers (#1036) Trump’s 2027 Budget To Reshape U.S. Environmental And Energy Policy CDC Delays Reporting Of COVID-19 Vaccine Benefits—Here’s What To Know Oura Has Designed A Solution To A Big Smart Ring Problem Netflix’s Best New Show Has A Near-Perfect 95% Rotten Tomatoes Score Coachella 2026 Is Being Taken Over By Creator Streams Quordle Hints Today: Sunday, April 12 Clues And Answers This Startup Wants To Use AI To Help Digitize History How To Get The Best Shield In ‘Crimson Desert’ Microsoft Venom Attack Targets C-Suite Executives ‘Maul: Shadow Lord’ Sets Even More Star Wars Rotten Tomatoes Records 3 Ways Happy Couples Argue Differently, By A Psychologist Success For Leapmotor Might Have Negatives For Stellantis New Names Surface As Potential Rogue And Wonder Woman In The MCU And DCU 4 Reasons Artemis Mission Matters Even If You Think It Is Wasteful Fast ‘Crimson Desert’ Patch Adds New Moves, Shield Hiding And One Great Feature Why Do Humans Blush? An Evolutionary Biologist Explains The Signal We Can’t Control Apple iPhone Fold: Striking Design Revealed In Leaked Photos Adobe Attacks Underway—Windows And Mac Users Given 72 Hours To Update iOS 26.4.1 Release: Crucial iPhone Feature Update Arrives, But No Security Fix Fury vs. Makhmudov Full Card, Ring Walk Times and How to Watch Can’t Stand Liquid Glass? This New Hidden iPhone Setting Is A Game-Changer Test-Driving The 2026 Changan Deepal S05: Italian Style Made In China NSA Warning—Reboot Your Internet Router Now Ways That Human-AI Collaboration Slides People Into ‘AI Brain Fry’ And Cognitive Downturns Stop Using These Networks—Google, NSA And TSA Warn NASA Changes Moon Plan: Landing Now Depends On SpaceX Or Blue Origin Samsung Expands One UI 8.5 Beta To More Galaxy Owners The Evolution Of Programmable Hardware At Xilinx NYT Mini Today: Saturday, April 11 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Saturday, April 11 (You’re Putting Me On) Splashdown! NASA’s Artemis II Returns To Earth After Moon Mission Attention Is All You Need. The Human Kind Is Still The One That Counts Today’s Wordle #1757 Hints And Answer For Saturday, April 11 NYT Pips Today: Hints, Answers And Walkthrough For Saturday, April 11 Android Circuit: Galaxy S27 Pro Emerges, Honor 600 Pre-Order Offers, Pixel 11 Display Leaks Apple Loop: iPhone 18 Pro Leak, Urgent iOS Update, MacBook Neo Issues Morgan Stanley Has Mostly Positive Outlook On Tesla Robotaxi, FSD V15 Running Out Of AI Tokens Faster Than Ever? Here’s Why CoreWeave Shares Pop 13% After Anthropic Deal ‘Euphoria’ Season 3’s Rotten Tomatoes Score Crashes, Has Lost Key Player People Don’t Agree On What AI Can Do, But They Don’t Even Use The Same Product ‘Overwhelming’—Google Issues Gemini Update For Gmail Users NYT Connections Hints Today: Saturday, April 11 Clues And Answers (#1035) Quordle Hints Today: Saturday, April 11 Clues And Answers The Costly Dream Of Space-Based AI Infrastructure Can You See The Watcher In This ‘Daredevil: Born Again’ Shot? Adobe Attacks Underway—Windows And Mac Users Given 72 Hours To Update You Just Watched The Backdoor Pilot For ‘The Pitt: Night Shift’ Are Nicotine Pouches Like Zyn And VELO Safe To Use? A Doctor Answers Human Resources (HR) Is The Key To AI Success Per WalkMe ( SAP)
The GPU Boom Is Over—The Cloud Boom Has Just Begun
Vasu Raj Jain · 2026-06-18 · via Forbes - Innovation

Vasu Raj Jain is an Impact-Focused Engineering Lead at Amazon Ads, leading infrastructure powering ad serving systems at massive scale.

getty

Two years ago, the AI playbook was straightforward: Acquire GPUs. Companies competed for Nvidia H100 allocations like traders bidding on oil futures, accepting 12-month waitlists and signing data center leases before the concrete had dried. The assumption was simple: Whoever accumulated the most silicon would win the AI race.

That assumption is no longer sufficient on its own. AI infrastructure has shifted from a training-centric model to one increasingly defined by inference, and that transition changes what AI compute (and AI infrastructure itself) actually means.​

Training Is A Project, Inference Is A Product

This is the distinction most teams miss when planning infrastructure. Training has a start date, an end date, a fixed cluster and a known workload. You assemble GPUs, run the job and produce a model; done. You could own that hardware outright, and it made sense.

Inference runs 24/7, scales with your user base, spikes unpredictably and never ends. It has to happen everywhere, simultaneously and at variable scale. This is not only a hardware scaling problem; it is increasingly a distributed systems problem.​

In 2023, inference drove one-third of AI compute. By 2026, Deloitte projects it will be two-thirds. It's growing at a 79% CAGR compared to 25% for training. Teams that planned infrastructure like a training run are stuck with clusters that can't handle what inference demands.​

Why Training Infrastructure Breaks Down For Inference

I've watched teams repurpose training clusters for inference. It fails for reasons that aren't obvious until you're in production.

1. Distribution is nonnegotiable. When an AI agent is reasoning through a task, every 100 milliseconds of latency degrades the experience. An estimated 80% to 85% of inference workloads will need global distribution within two years. You can't serve that from one cluster.

2. Demand is impossible to predict. Training demand is deterministic. Inference follows your user base: launches, viral moments and Monday spikes. Fixed hardware means idle capacity most of the time or failure under load when it matters.

3. Heterogeneity is the norm. Production AI runs dozens of models with different precisions, latencies and scaling needs. You need intelligent routing, not a homogeneous GPU farm.​

Treating Inference As A First-Class Production Service

The biggest failure mode isn't picking the wrong hardware. It's treating model deployment like a batch job instead of a service release. Here's what this looks like in practice.​

Operational rigor matters.

You need service level objectives (SLOs) for latency and throughput, not just uptime. A model returning answers in 800 milliseconds when your user experience (UX) requires 200 milliseconds is a broken product. Cost-per-inference becomes a tracked metric at the team level, because without it nobody owns the bill. Model updates need canary deployments, not direct-to-production releases.

Architecture has to change.

​Separate inference compute from training clusters entirely. Rightsize instance fleets to each model’s serving profile: A 7B-parameter model and a 70B-parameter model have fundamentally different compute and memory requirements, yet teams routinely deploy both on identical infrastructure. Build request routing that accounts for model warmup behavior, because newly loaded models often produce higher latency until caches, weights and serving pipelines stabilize, putting your latency SLOs at risk.​

Organization has to change, too.

Inference operations need dedicated on-call ownership rather than sharing coverage with training. The failure modes are different: When a training job fails, the cost is typically wasted compute and delayed iteration; when inference fails, the impact is immediate through degraded user experience, interrupted workflows or lost revenue. Runbooks should distinguish model degradation from infrastructure failure because the remediation paths are entirely different.

This is where many teams get stuck. They optimize relentlessly for training throughput while underinvesting in serving reliability. They monitor CPU utilization and pod counts but lack observability into per-request inference performance and output quality. And they treat deployment as a milestone rather than an ongoing operational discipline.​​

Where Fixed Infrastructure Falls Short In The Inference Era

The instinct to own hardware comes from training-era thinking, where utilization is predictable and sustained. Inference breaks this structurally. Demand follows time zones, product cycles and viral unpredictability. A cluster sized for peak sits idle most of the day. A cluster sized for average fails during spikes. There is rarely a single "right size" for fixed inference hardware in fast-growing or consumer-scale products.​

The operational overhead compounds it. Every hour managing firmware, network topology and physical redundancy is an hour not spent on model quality. And when the next model generation arrives with different memory requirements, fixed hardware can become a liability you can't swap out with a configuration change.

Cloud changes the economics and operating model. It makes it easier to adjust architectures, scale capacity more dynamically and match hardware profiles to different model requirements without continuously reconfiguring physical infrastructure. ​

That flexibility is not universally required, but for many inference workloads, especially those with variable demand, geographic distribution and heterogeneous models, it becomes increasingly difficult to replicate efficiently with fixed infrastructure alone.​​

The Question You Should Actually Be Asking

Most teams still ask, "How do I get more GPUs?" The question I believe should be asked instead is "How do I build an inference platform that scales globally, handles heterogeneous models, adapts to bursty demand and doesn't require a dedicated hardware operations team?"

For many organizations, that question is evolving into a cloud-first infrastructure problem.​ The GPU boom built the models, but the cloud boom will deliver them to the world.​​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?