惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
腾讯CDC
Recent Announcements
Recent Announcements
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Hugging Face - Blog
Hugging Face - Blog
H
Help Net Security
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Last Week in AI
Last Week in AI
博客园_首页
D
DataBreaches.Net
P
Proofpoint News Feed
云风的 BLOG
云风的 BLOG
V
Visual Studio Blog
月光博客
月光博客
Jina AI
Jina AI
Stack Overflow Blog
Stack Overflow Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 【当耐特】
Vercel News
Vercel News
WordPress大学
WordPress大学
J
Java Code Geeks
博客园 - 聂微东
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
U
Unit 42

TechPowerUp

Microsoft Visual Studio Professional 2026 + 15 Coding Courses Just $50 No Surprise Microsoft Office Hikes—Own it for Life for $50 The Witcher 3: Wild Hunt Reaches 65 Million Copies Ahead of Songs of the Past Launch Valve Steam Deck Sells Out in 24 Hours Despite Price Hike 007 First Light Exceeds 1.5 Million Sales in 24 Hours Rockstar Workers Unionize Ahead of GTA VI Launch Following Dismissal of 31 Workers Unknown Worlds Earns $250M Performance Bonus After Stellar Subnautica 2 Launch Dell Technologies Delivers First Quarter Fiscal 2027 Financial Results Marathon Season 2 To Start With Free Week and Plenty of New Content ASRock iBox Fanless Mini PCs Get Intel Panther Lake Upgrade Acer Broadens Portfolio with Two New Laptops Powered by the Latest Snapdragon Processors Qualcomm Introduces Snapdragon C Entry‑Level Processors for Budget Laptops OneXPlayer 3 Gaming Handheld Emerges With Intel Arc G3 Extreme Intel Arc G3 CPU Family Officially Released for Handheld Gaming PCs Samsung Exynos 2600 SoC Annotated, Showcases 3-tier CPU, AMD RDNA 4 iGPU Acer Expands Gaming Portfolio with Predator Atlas 8 Handheld Powered by Intel Silicon Motion Introduces SM2524XT PCIe Gen 5 DRAMless SSD Controller PXN Launches the Vector X Professional-Grade Sim Racing Pedals TP-Link Introduces Archer 8, Its First Wi-Fi 8 Router Platform Synology Announces Availability of New FlashStation FS200T Philips Announces Evnia 32M2N8900P QD-OLED 4K 240 Hz Gaming Monitor Samsung Display Develops First 4K 360 Hz QD-OLED Panel for Monitors LG Display Begins Mass Production of World's First 240 Hz RGB Stripe OLED ZALMAN Intros ZM-STC11 Silicone-based Thermal Paste Stream Deck Becomes the Action Layer for AI, Starting with NVIDIA G-Assist GIGABYTE Debuts New BRIX Mini PC Powered by Panther Lake to Scale Enterprise AI Sharkoon Announces the S25 Series Cases Scythe Intros the Magoroku Dual Fin-stack Air CPU Cooler First Look at ZOTAC's GeForce RTX 50-series 20th Anniversary Edition Graphics Cards ADATA TRUSTA AI Scaler Extended Memory Solution Breaks GPU Limits
Lexar AI Storage Solution Helps Limited DRAM Run Larger L...
by AleksandarK · 2026-06-16 · via TechPowerUp

Lexar has been exploring various technologies to help consumers achieve faster data throughput and more reliable storage. The company now envisions a shift as the PC evolves from a traditional personal computer to a local AI-enhanced experience. We interviewed Lexar's Chief Technical Officer, Daniel Guo, about the technology Lexar is developing to help reduce some DRAM demand and shift the balance towards integrated hardware-software AI storage solution with NAND Flash. According to Guo, DRAM is about six times more expensive to manufacture than NAND Flash, and AI SSDs present opportunities to reduce DRAM requirements for running AI models on local hardware. This is where the Lexar AI Storage Solution comes into play, as the company is creating new storage solutions to support local AI deployments in constrained DRAM capacity by offloading some parts of large language models to SSDs. This approach allows larger and more powerful LLMs to fit into a PC build, reducing DRAM capacity requirements by up to 40% in specific scenarios.

Based on internal testing, Lexar managed to run the Qwen 3.5 122B AI model on a local PC. Traditionally, users would need to invest about $4,500 in a PC with a decent CPU and 128 GB of DRAM to run this model. Through hardware and software optimization, the Lexar AI suite with the Lexar AI SSD can reduce the DRAM requirement to 32 GB and run the model with 35 billion parameters at 15.6 tokens per second, compared to only 5.2 tokens per second using traditional frameworks. When attempting to load the 122B model on 32 GB of DRAM, the traditional Llama.cpp fails to load and crashes, while Lexar's SSD offloading provides about 4.4 tokens per second.

With a more robust configuration featuring 64 GB of DRAM, running the 122B model with a larger context window is possible only with SSD offloading. With about 4,000 tokens in context, both traditional configurations and the Lexar AI stack run at a slightly higher speed. However, for larger contexts, often needed at 256K tokens, only the Lexar AI suite can launch and manage to produce about 19.3 tokens per second. This setup isn't perfect, and not every model size can be offloaded to the SSD. With larger LLMs, system latency increases significantly, as the time between submitting a prompt and receiving a response grows exponentially.

The time to first token, often called TTFM, is about two seconds before the first token appears after the prompt is submitted with a 2K context window. When the context is larger at 4K, the delay increases to between 6 and 8 seconds. Technically, users could offload models about 400 billion parameters large, but the tokens per second and TTFM would be very slow. For some, this might be suitable, but for others, buying more DRAM is the better solution. Either way, this is an intriguing concept from Lexar.

At Computex 2026, the company developed a concept for Mini-PCs and desktops featuring an M.2 slot designed for multiple insertions. An M.2 SSD is encased in a metal jacket and inserted into a 25 mm-wide slot on the front panel of a mini PC, connecting directly to the M.2 slot wired to the processor or chipset. This design eliminates other overheads. The hot-swappable SSD, which offloads AI models onto NAND Flash, reduces dependency on DRAM and aids in running larger models. It is available in both PCIe Gen 5 and Gen 4 versions, with the Gen 5 version offering more bandwidth. This M.2 SSD uses Lexar's custom Storage Processing Unit (SPU) DRAM-less controller for complete control over data movement.