惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

量子位
Vercel News
Vercel News
Microsoft Azure Blog
Microsoft Azure Blog
爱范儿
爱范儿
N
Netflix TechBlog - Medium
Google DeepMind News
Google DeepMind News
H
Help Net Security
罗磊的独立博客
The Cloudflare Blog
J
Java Code Geeks
博客园 - 叶小钗
I
InfoQ
B
Blog
Blog — PlanetScale
Blog — PlanetScale
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
腾讯CDC
月光博客
月光博客
博客园_首页
雷峰网
雷峰网
M
MIT News - Artificial intelligence
博客园 - 【当耐特】
美团技术团队
T
The Blog of Author Tim Ferriss
博客园 - 司徒正美

Latest from TechRadar in News

VodafoneThree gets Ofcom approval to bring satellite connectivity to your smartphone NYT Connections today – my hints and answers for April 16 (#1040) Quordle hints and answers for Thursday, April 16 (game #1543) NYT Strands hints and answers for Thursday, April 16 (game #774) Is this the tipping point for AI at work? New Gallup survey finds half of all US employees now use it in some way Allbirds — the shoe viral company — just pivoted into AI, and I wish this were an Onion headline 'Every Apple user needs to know about this nasty scam': Fake warnings tell users their iCloud data will be… 'Makes it even more disappointing': Microsoft backs fossil fuel big time with $7 billion deal in race for AI… 'Maybe it’s not science fiction': Solar panels are causing rainwater to fall in one of the driest places… Maine becomes first US state to pass data centre construction ban Dozens of WordPress plugins hijacked to target thousands of sites Drone-killing laser weapons greenlit for use in US airspace – FAA and Defense Department say high-energy weapons are ‘ready to protect all air travelers from illicit drone use’ despite airspace restrictions and friendly-fire incidents 'We are currently being extorted' — crypto giant Kraken says it is facing extortion attack, here's… McGraw Hill becomes latest to see its Salesforce data hacked Looking for a new PC? Now might be great time to upgrade, as Gartner figures claim shipments are rising — while… Farewell Surface Hub — Microsoft kills off its super-sized touchscreen displays, but you might still be able to get one if you act fast 'We have no interest in patient data in the UK': Palantir UK head defends record as criticisms rise Amazon’s new AI Bio Discovery tool can provide ‘every researcher’ with ‘lab-in-the-loop drug discovery’ – 40+ AI biology models can filter 300,000 novel antibody candidates down to the top results for testing in just weeks Over 100 Chrome Web Store extensions found stealing user data from thousands of accounts OpenAI reveals its Mythos rival designed for cybersecurity pros NYT Connections hints and answers for Tuesday, April 14 (game #1038) Forget Dr Doolittle, study finds animals might not only want to use tech, but they also want to talk to us with it… 'The decision is deeply troubling': Tesla gets a green light for Full Self-Driving in Europe — but not… OpenAI flags third-party data issue — all macOS users should update now Microsoft says Copilot is for ‘entertainment' not work, Meta’s Muse Spark and 7 other AI stories you… Man Utd vs Leeds Live Streams: How to watch Premier League 2025/26 from anywhere in the world, team news What is the release date for Invincible season 4 episode 7 on Prime Video? Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… The Lenovo Legion Go 2 handheld costs more than two Nvidia RTX 5080 GPUs — and that's genuinely absurd Secretlab is launching its first Diablo desk, with a design that 'traces the infernal history' of the series
This tiny AMD PC just ran a massive 397B AI Model that re...
https://www.techradar.com/sg/author/rahim-amir · 2026-06-19 · via Latest from TechRadar in News
AMD Ryzen AI
(Image credit: AMD)

AMD's Ryzen AI Halo recently went on sale for $4,000, sparking an interesting debate about how it compares to Nvidia's slightly pricier DGX Spark offering.

The configuration that the Ryzen AI Halo offers, however, has been on the market for a few months now, and while most OEMs and enterprise providers are offering the same flavor and configuration, Shenzhen-based memory and storage company Longsys has taken things a step further.

The storage giant demonstrated a localized version of a 397B-parameter AI model running on its own version of the Ryzen AI Halo, featuring the same 16-core Ryzen AI Max+ 395 and 128GB of RAM configuration.

How was the Ryzen AI Max+ 395 able to run such a massive model with only 128GB of RAM?

While the model being run was not explicitly stated, it seems to be a customized version derived from Alibaba's Qwen 3.5 397B (A17B), a multimodal foundation model that leverages a Mixture-of-Experts (MoE) approach, which made the original DeepSeek such a potent challenger.

Even if it was leveraging INT4 quantization, the memory requirements far exceed the memory the device demonstrating the feat had on offer: only 96GB of VRAM is available to the GPU in a 128GB unified configuration, versus an estimated 200-250GB of VRAM the model needs to run.

The secret sauce is Longsys's recently unveiled custom SPU and iSA configuration that offers the ability to compress data in real time, a feat that the company says allows it to fit as much as twice the amount of data in storage drives of up to 128GB, leveraging a caching layer that reduces DRAM requirements considerably.

The approach involves offloading experts not in active use to a large, fast storage buffer that the AI chip can then reintroduce them from if needed.

Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!

In a press release, Longsys claimed its approach worked by targeting, "the pain points of MoE LLMs", such as large parameter counts, rapid KV Cache expansion, and I/O latency that hampers inference efficiency

"It leverages expert offloading, intelligent cache management, and predictive prefetch algorithms to efficiently resolve storage scheduling challenges and comprehensively improve local AI inference smoothness," the company added.

It is important to note that while the move itself is an impressive feat, Longsys did not provide specifics on compute power in terms of tokens per second, where the Ryzen AI chip is relatively limited compared to most modern AI GPU offerings.

Regardless, the approach that essentially treats storage as memory suggests that localized AI might be able to run considerably larger models, and that memory might not be as hard a constraint for certain approaches.

It signifies that memory constraints can be circumvented by leveraging fast storage and running a frontier-level model that would otherwise require tens of thousands of dollars in AI hardware, which is no small feat. It means that models that were previously constrained to datacenters only can now be run on a device that fits in the palm of your hand.


Google logo on a black background next to text reading 'Click to follow TechRadar'

Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.

Rahim Amir is a UAE-based tech writer who enjoys building PCs as much as he enjoys writing about them. He has been professionally writing about PC hardware since 2023, focusing on buyer’s guides, hardware reviews, and sponsored content and features related to tech.

Having built hundreds of gaming PCs and being an avid gamer in his spare time, Rahim tends to have stronger opinions about hardware than most. This is particularly on display when he gets his way with powerful, but minimalistic RGB builds even as Small Form Factor (SFF) PCs come a close second.