惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - Franky
有赞技术团队
有赞技术团队
宝玉的分享
宝玉的分享
雷峰网
雷峰网
Hugging Face - Blog
Hugging Face - Blog
V
V2EX
大猫的无限游戏
大猫的无限游戏
博客园 - 司徒正美
D
Docker
T
The Blog of Author Tim Ferriss
罗磊的独立博客
博客园 - 叶小钗
酷 壳 – CoolShell
酷 壳 – CoolShell
Blog — PlanetScale
Blog — PlanetScale
月光博客
月光博客
J
Java Code Geeks
Jina AI
Jina AI
博客园 - 【当耐特】
C
Check Point Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
腾讯CDC
Last Week in AI
Last Week in AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
V
Visual Studio Blog

StorageNewsletter

Vast Data Valued at $30 Billion as AI Drives a New Infrastructure Stack Wasabi Technologies Closes $250M Credit Facility to Expand Cloud Storage Innovation NAB Show 2026: TVC Soho Selects EditShare High-Performance NVMe Storage to Support Resolve Finishing Workflows KIOXIA Unveils Value-Oriented QLC-based EG7 Series SSDs for PC OEMs NinjaOne Unified Backup Surpasses Fifteen Thousand Customers Portworx by Everpure is Redefining Modern Virtualization for Customers with Proven, Enterprise-Ready Solutions Peer Software Strengthens Global Partner Program to Unify Fragmented File Environments for the AI Era NetApp Collaborates with Google Cloud to Power Data Infrastructure for Distributed Cloud Sidus Space Expands Existing Agreement with Lonestar Data Holdings, Inc. to Support Additional StarVault Orbital Data Storage Payload From SNIA: SCSI Continues to Innovate Data Storage with SBC-5 Microsoft Technology Licensing Assigned Patent Linux Kernel 7.0 is Out NAB Show 2026: ATTO Technology Ignites Next Era of Media Connectivity NAB Show 2026: Promise Technology to Showcase Integrated Storage Plug-in for Video and Image Creative Workflows NAB Show 2026: EditShare Advances Analytical AI and NVMe Performance for Modern Broadcast and Post NAB Show 2026: Elements Introduces GRID, a New Node-Based Scale-Out NAS Platform NAB Show 2026: MASV Expands Global Partner Ecosystem to Accelerate End-to-End Media Workflows NAB Show 2026: UnifyDrive to Showcase Full NAS Lineup NAB Show 2026: Strada Releases Easiest Remote Editing Platform on the Market Synology: Three Security Advisories on Resolved Vulnerabilities Mastercard International Assigned Patent NAB Show 2026: Promise Technology to Unveil AI-Optimized Storage Solutions NAB Show 2026: QNAP Releases HDP Recovery Media Creator: Building Windows DR Media in USB and ISO Formats NAB Show 2026: SNS Unveils Three New Products, Expanded Ecosystem NAB Show 2026: Symply Unveils Centara Platform, World’s First Quad-Interface LTO with Thunderbolt 5, and Spark One Portable NVMe NAB Show 2026: Other World Computing Launches OWC Express 4M2 Ultra Thunderbolt 5 Four-Slot NVMe M.2 SSD Enclosure Panmnesia to Mass-Produce PCIe 6.4-CXL 3.2 Fusion Switch CIQ Delivers the First Enterprise Linux Compliance Platform for Federal Cryptographic Validation and Post-Quantum Readiness Adata Launches Urban Tapsafe Up to 2TB USB 3.2 Gen2 External SSD Raidon Technology Introduces 4-Bay STARDOM SR4-BA32 20Gb/s USB-C RAID-5 Desktop Storage System
Weka and Oracle Cloud Infrastructure Validate 10x Through...
Philippe Nicolas · 2026-06-11 · via StorageNewsletter

RAIDON

Joint benchmarks on OCI H100 infrastructure showed 10x more concurrent users, 10x higher token throughput, and 7x more tokens served without adding GPUs

This is a Press Release edited by StorageNewsletter.com on June 11, 2026 at 2:01 pm

Weka, a AI data and memory infrastructure company, announced production-scale benchmarks that show how organizations can improve the economics of long-context AI inference by serving more users and tokens on the same GPU footprint.Weka Expands India OperationsThe benchmarks show that Weka’s NeuralMesh platform with Augmented Memory Grid on Oracle Cloud Infrastructure (OCI) serves 10x more concurrent users, delivers 10x higher token throughput, and produces 7x more tokens per GPU than DRAM-only configurations without adding infrastructure. The results were validated on a nine-node OCI bare-metal H100 cluster with 100,000-token context windows.

“Enterprise AI workloads are pushing context windows and GPU utilization to new limits,” said Pablo Selem, senior director, software development, Oracle Cloud Infrastructure. “These benchmarks show how Weka’s NeuralMesh platform with Augmented Memory Grid on OCI helps remove memory bottlenecks so customers can support larger, more demanding inference workloads without simply adding more GPUs.”

Three Outcomes That Change the Math on Inference
Validated at production scale on a bare-metal H100 cluster (nine nodes, 72 GPUs, 100,000-token context windows, thousands of concurrent users), NeuralMesh with Augmented Memory Grid on OCI delivered:

  • 10x more concurrent users served, without adding infrastructure. NeuralMesh with Augmented Memory Grid scaled past 5,000 concurrent users vs. about 600 for DRAM-only configurations. This eliminates the failure cliff that hits when cache saturates by expanding the active cache working set from 8.64 TiB of DRAM to 287 TiB of usable NVMe. In addition, more users per GPU means the same investment stretches further
  • 10x higher token throughput. More output from every GPU in the cluster. On OCI, NeuralMesh with Augmented Memory Grid reached approx. two million tokens per second, compared to under 200,000 for the DRAM-only baseline. For product teams running real-time AI features, including search, summarization, code assist, and multi-turn agents, the throughput determines the ceiling for how many users can be served, how fast features respond, and how much revenue the infrastructure can support
  • 7x more tokens served. Lower cost per token at scale. NeuralMesh with Augmented Memory Grid served five billion tokens, compared to 700 million for the DRAM-only baseline, in a single one-hour, 2,400-user test. For organizations running agentic workflows, DRAM saturation quietly drains GPU capacity through constant recomputation, creating a direct hit on cost per token and ROI

“Inference is bottlenecked by how much effective memory is available to GPUs,” said Liran Zvibel, CEO, Weka. “These results prove that AI token economics aren’t solved by hardware alone; they’re solved by eliminating the memory wall that has been the real ceiling on what existing hardware can do. NeuralMesh with Augmented Memory Grid running on OCI brings orders of magnitude more tokens to customers in an extremely cost-efficient way.”

Transforming AI Economics with Context Memory Infrastructure
As inference demand grows, AI infrastructure inefficiencies compound. Every key-value (KV) cache eviction is a tax: on GPU cycles, latency, user experience, and the cost of every token served. For long-context and agentic workloads, where inputs routinely run to 100,000 tokens or more, that tax is not a rounding error. It is a direct hit on the unit economics of every organization running production AI.

Augmented Memory Grid, a capability of NeuralMesh, solves the problem at the architectural level by decoupling KV cache from local GPU memory and storing it in a high-performance token warehouse accessible across the cluster. Any host can serve any session with cache hits intact, eliminating rigid session stickiness while delivering superior performance to DRAM, improving load balancing, and enabling clean horizontal scaling as concurrency grows. The result is persistent context memory for AI agents and the cost lever that makes long-context inference economical to run at scale.

Production-Grade Proof
OCI published the full benchmark methodology, system configuration, and results on its AI & Data Science blog on May 13, 2026. The benchmarks, executed on a nine-node OCI bare-metal H100 cluster, move beyond the prior phase of validation, which demonstrated 1000x more KV cache capacity and up to 20x faster time to first token at 128,000 tokens. This latest phase tests the full economics of inference in production: concurrency density, sustained throughput, cache persistence, and service level objective (SLO) stability when demand spikes under high load.

Available on Oracle Marketplace
NeuralMesh with Augmented Memory Grid is available to Weka customers and on the Oracle Marketplace, with OCI as Weka’s exclusive cloud launch partner. Organizations running long-context inference on OCI can deploy a validated, production-ready architecture today.

Share this news : Twitter Facebook Linkedin email pdf

Articles_bottom

SNL Awards_2026

AIC