惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

N
Netflix TechBlog - Medium
I
InfoQ
Engineering at Meta
Engineering at Meta
Jina AI
Jina AI
Recent Announcements
Recent Announcements
T
The Blog of Author Tim Ferriss
P
Proofpoint News Feed
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
D
Docker
Microsoft Security Blog
Microsoft Security Blog
宝玉的分享
宝玉的分享
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
GbyAI
GbyAI
博客园 - Franky
博客园 - 聂微东
Microsoft Azure Blog
Microsoft Azure Blog
博客园 - 叶小钗
酷 壳 – CoolShell
酷 壳 – CoolShell
B
Blog RSS Feed
WordPress大学
WordPress大学
MyScale Blog
MyScale Blog
月光博客
月光博客
罗磊的独立博客

Forbes - Innovation

Why Do Humans Have Fingerprints? Hint: It’s Not What You Think Booking.com Confirms Data Breach, Reservation PIN Codes Changed Why Major News Sites Are Blocking The Internet Archive’s Wayback Machine iPhone Fold Release Date: New Report Details Frustrating Apple News Comet Tracker: How To See Pan-STARRS And Three Planets On Wednesday NYT Mini Crossword Today: Tuesday, April 14 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Tuesday, April 14 (It’s A Little Unclear) Today’s Wordle #1760 Hints And Answer For Tuesday, April 14 Most Of The Microplastics In Urban Air Come From Tires Today’s Wordle #1759 Hints And Answer For Monday, April 13 NYT Mini Crossword Today: Monday, April 13 Hints And Answers NYT Pips Today: Hints, Answers And Walkthrough For Monday, April 13 The YC Chief Who Codes 10,000 Lines A Day Has A Simple Secret Samsung Expands One UI 8.5 Beta To More Galaxy Owners Why You Should Stop Using Your iPhone If It’s On This List Chamath Says Firms That Treat AI As A Strategy Hand Rivals Their Edge 3 Unexpected Habits Of Secure Couples, By A Psychologist The First Lamp That Folds Your Clothes Samsung’s Disappointing Price Update For Galaxy Phone Buyers 3 Subtle Signs Someone Is Falling In Love With You, By A Psychologist Do Mantis Shrimp See More Colors Than Humans? A Biologist Explains NYT Connections Answers Explained For Monday, April 13 (#1,037) NYT Connections Hints Today: Monday, April 13 Clues And Answers (#1,037) LEGO Luigi & Mach 8 (72050) Review: 2026’s Best Set Yet? Marc Andreessen Says AI Productivity Will Trigger A Hiring Boom 3D Printing Is The Ultimate Hack To Reduce Household Spending Apple iPhone Fold: Striking Design Revealed In Leaked Photos Apple Smart Glasses: New Leak Reveals A Major Design Twist To Beat Meta Tested: The AI Coming To The Rivian R2 Quordle Hints Today: Monday, April 13 Clues And Answers
AI’s Memory Crisis Is Here: Don’t Hoard, Optimize
Liran Zvibel · 2026-05-13 · via Forbes - Innovation

Liran Zvibel, Cofounder & CEO, WEKA.

getty

​Training AI demands raw GPU compute. Inference demands something else entirely: memory. The GPUs powering today's models carry limited high-bandwidth memory (HBM) before external memory is required—that's the memory wall, and at inference scale, every model hits it. As the industry shifts from training to inference, memory has become the defining constraint in AI infrastructure.

DRAM supply remains tight amid strong demand, and I’m seeing enterprises paying 300-1100% more for memory chips since last year, a number expected to climb fast. Procurement timelines that once took days now stretch to months. Yet the most common response I’m seeing across the industry is the wrong one: hoard more hardware.​

I’ve spent the last year talking with customers across every segment of the AI market, neoclouds, large enterprises and AI model builders—the pattern is consistent. Organizations compensating for architectural inefficiency by buying more capacity are now exposed. With memory stockpiles constrained through late 2027, that strategy no longer works.

The shortage didn’t create the problem. It’s the forcing function that revealed what was always broken, and made it impossible to ignore.

AI-Optimized Chips, Crippling Memory Scarcity

When chip manufacturers in 2024 shifted wafer capacity away from Dynamic Random Access Memory (DRAM) and toward AI-friendly HBM, global DRAM supply contracted sharply. The inventory situation has deteriorated to two to four weeks of product on hand, and the shortfall is expected to persist through late 2027.

This crisis exposes a fundamental flaw that has been percolating for years: the AI industry has been papering over architectural inefficiency with raw capacity. With memory now genuinely scarce, that approach has run out of road. The organizations that win the AI race won’t be the ones who secured the most chips. They’ll be the ones who deployed software-defined architectures that transform underutilized resources, fresh NVMe and spare CPU capacity already sitting in GPU servers, into high-performance memory extensions.

How To Maximize Value From Your GPU Stacks

The memory shortage will only intensify over the next 18-24 months, and organizations can’t afford to wait for manufacturing to pick up as AI projects accelerate. Instead, savvy teams will deploy software that can transform underutilized hardware into high-performance AI systems.

Here are five strategies to help you not only survive the memory shortage, but thrive in spite of it:

1. Stop Wasting What You Have

Audit your current GPU utilization before placing any new orders. You’ll likely find rates well below capacity—not because you lack GPUs, but because your storage can’t feed them fast enough. I regularly see enterprise customers discover they’re leaving 50%–70% of available capacity on the table. Benchmark your data delivery rates against GPU consumption first. The answers are usually already in your infrastructure.

2. Extend GPU Memory By 1,000x Vs. Adding More Storage

The critical bottleneck isn’t storage capacity measured in terabytes—it’s GPU memory measured in gigabytes. When inference exhausts the limited HBM integrated into GPU packages, systems waste expensive compute cycles recomputing tokens they’ve already processed. You cannot buy more HBM; it’s fixed in the hardware and more constrained than any other memory type.​

Implement GPU memory extension using direct communication technologies that transform abundant NVMe into a functional extension of scarce HBM. Since the aim is to replace or extend memory that is incredibly fast, focus on picking an NVMe flash-based solution that optimizes for lowest latency. With near-perfect KV cache hit rates, organizations can achieve dramatically faster time-to-first-token and serve significantly more concurrent users per GPU—without waiting months for new GPUs that face the same memory constraints.

3. Deploy Software-Defined Architecture, Not More Hardware

The instinct during shortages is to secure more silicon at any price. But when new fabs won’t be operational until 2027, that strategy means your AI roadmap stalls while competitors move forward. The hardware-first instinct is deeply ingrained—I understand it. When you can see the constraint, buying your way out feels like control. But the math no longer works.

Shift to software-defined approaches that transform underutilized resources in existing GPU servers into high-performance, memory-class infrastructure. Co-located deployment eliminates separate storage procurement entirely, activating NVMe and CPU already in the servers you’re buying for your GPUs anyway. This isn’t theory; it’s deployable in weeks using capacity you already own.​

4. Prioritize Performance Density Over Raw Capacity

During shortages, “reclaimed” drives can seem like an attractive stopgap to increase raw capacity. Resist the temptation. These drives are approaching end-of-life and introduce unpredictable latency spikes that drag down GPU productivity.

Focus on workload density—the amount of consistent, reliable work you can extract per drive—rather than grabbing cheap gigabytes. Aging hardware masks underlying throughput problems that will eventually force you back into the strained memory market.

5. Let Software—Not Static Policies—Manage Your Storage Tiers

Most storage tiering strategies were configured once, years ago, and never revisited. The result: cold data clogs expensive flash while frequently accessed datasets get demoted to slow storage, forcing manual copies and redundant data to keep workflows moving. AI workloads make this exponentially worse, with checkpoints sitting idle until they’re suddenly needed at full speed.

Deploy automated, behavior-based tiering that continuously adapts to actual access patterns. Systems that transparently move data between tiers without blocking workloads can turn object storage into a viable overflow valve during shortages.​

Questions Worth Asking Now

The organizations that thrive through this shortage won’t be the ones who bought the most HBM. They’ll be the ones who asked harder questions while competitors were treading water.​

• Is your storage actually fast enough to feed your compute?

• Can you deploy on the infrastructure you have instead of waiting for what you can’t get?

• Are you optimizing for capacity or utilization?

• Does your architecture extend GPU memory, or does it simply manage storage?​

The memory crisis didn’t create a new set of problems. It revealed the one problem the industry has avoided: we’ve been buying our way around an architectural flaw rather than fixing it. The shortage is the forcing function. The solution has been available all along.


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?