惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
腾讯CDC
M
MIT News - Artificial intelligence
Y
Y Combinator Blog
L
LangChain Blog
Vercel News
Vercel News
云风的 BLOG
云风的 BLOG
GbyAI
GbyAI
Stack Overflow Blog
Stack Overflow Blog
Microsoft Azure Blog
Microsoft Azure Blog
B
Blog RSS Feed
The GitHub Blog
The GitHub Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
B
Blog
P
Proofpoint News Feed
H
Hackread – Cybersecurity News, Data Breaches, AI and More
博客园_首页
Google DeepMind News
Google DeepMind News
WordPress大学
WordPress大学
aimingoo的专栏
aimingoo的专栏
小众软件
小众软件
IT之家
IT之家
A
About on SuperTechFans
H
Help Net Security

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
AMD Ryzen™ AI Max+ AI PCs Deliver Exceptional Intelligenc...
teleforce · 2026-04-25 · via Hacker News - Newest: "AI"

For the last few years, powerful AI has mostly lived somewhere else: behind an API or in a distant data center. With the AMD Ryzen™ AI Max+ Series, that center of computational gravity began to move back toward the user. Most researchers, innovators, businesses, and entrepreneurs desire cloud-level AI capabilities they can access from their own desks.

Cloud quality AI on your desk

A key question from customers is simple: can they get something close to ChatGPT level quality locally, without sending everything to the cloud?

Recent evaluations of the open GPT-OSS 120B model show around 80% on the GPQA Diamond benchmark, which focuses on PHD-level science questions, and roughly 90% on MMLU, a broad measure of reasoning across college level exams. Running that same model on AMD Ryzen™ AI Max+ processor means this capability is now accessible on an AMD Ryzen™ AI Max+ powered system (with 128GB memory) instead of a remote cluster.

Local inference changes how teams think about AI. They gain tighter control over sensitive data, more predictable performance regardless of network conditions, and the freedom to customize and fine tune models without worrying about API limits or variable per token pricing.

Slide showcasing a difference in intelligence between the cloud-based ChatGP o4 Mini and GPT-OSS 120B (local) models.

Large models without the usual slowdown

Historically, bigger models have also meant significantly slower models. That tradeoff is starting to shift thanks to breakthroughs in AI model development and LLM architectures.

On an AMD Ryzen™ AI Max+ 395+ processor, OpenAI’s GPT-OSS 120B (with 116.8 billion total parameters) runs about 10 times faster than Meta Llama 3 70B (with 70 billion total parameters) in our comparative measurements. This is primarily due to model architectures becoming far more efficient and having lower activated parameters. At the same time, GPT-OSS 120B also offers a dramatically longer context window of around 128K tokens compared to about 8K before.

In practical terms, a contract review tool can see an entire agreement set at once instead of working section by section. A coding assistant can keep more of your repository, logs, and documentation in view. Analysts can paste full reports and multiyear time series into a single prompt and keep the conversation coherent.

Workflows that once required intricate prompt strategies and manual chunking now simply fit.

Slide showcasing the performance difference between an older SOTA local model and GPT-OSS 120B.

Class leading tokens per dollar inference in LM Studio

Using llama.cpp-based application LM Studio (which is available on both Windows and Linux) to evaluate four large models, GPT-OSS 20B, GPT-OSS 120B, GLM 4.5 Air, and DeepSeek R1 Distill 70B, we measured an average 1.7 times more tokens per dollar on an AMD Ryzen™ AI Max+ processor compared to an NVIDIA DGX Spark configuration targeting the same workloads.

For enterprises and businesses budgeting for rollout of "available-on-your-desk" AI capability for their employees, tokens per second is one of the key planning metrics. Startups can run meaningful experiments on a few AMD Ryzen™ AI Max+ workstations instead of renting time on large clusters. Universities and labs can put advanced models into more hands without expanding budgets.

Slide showcasing tokens per second performance difference between the NVIDIA DGX Spark and AMD Ryzen AI Max

One platform for Windows and Linux

The AI ecosystem runs on a mix of Windows tools and Linux native ML frameworks. The AMD Ryzen™ AI Max+ platform is designed for both.

Creators can stay in familiar Windows creative suites while offloading AI workloads to local models. The entire CAD ecosystem is supported and is only available officially on Windows. These systems support thousands of existing Windows applications while also enabling rich Linux environments for development and deployment.

Slide showcasing the various form factors and OEM designs available for the AMD Ryzen AI Max+

A step toward the next AI era

The story of the AMD Ryzen™ AI Max+ is truly about where AI is going.

We see hybrid AI architectures that combine cloud services, private clusters, and powerful local machines. We see open and customizable models giving organizations more control over behavior and governance. And we see AI native PCs and workstations becoming the default tools for knowledge work, creativity, and engineering.

With cloud grade models running locally, large context reasoning at interactive speeds, and class leading efficiency measured in tokens per dollar in apps like LM Studio, the AMD Ryzen™ AI Max+ Series processors help bring that future into the present.

The next wave of AI breakthroughs will not come only from bigger data centers but also from what people build when powerful, efficient AI is running on the systems they use every day.

SHOP-26: Testing as of December 2025 by AMD. All tests conducted in LM Studio 0.3.35 (Build 1). Vulkan llama.cpp v 1.64.0 used with Ubuntu 24.03.3 and therock-gfx1151-7.9rc1 for AMD Ryzen™ AI Max+ 128GB. CUDA llama.cpp v1.64.0 used with DGX OS and Driver Version 580.95.05 and CUDA Toolkit Version 13.1.0-1 for NVIDIA DGX Spark. Flash Attention = ON in all cases. Token/s: sustained performance average of multiple runs with specimen prompt “How long would it take for a ball dropped from 10 meter height to hit the ground?”. Models tested: OpenAI GPT-OSS 120B, OpenAI GPT-OSS 20B, GLM 4.5 Air and DeepSeek R1 Distill Llama 70b. Tokens per second per dollar measured using market pricing of USD $2566 for the Framework Desktop 128GB and USD $4000 for the NVIDIA DGX Spark as of December 2025. AMD Ryzen™ AI Max+ 395 on a Framework Desktop with equivalent storage capacity and 128GB memory and NVIDIA DGX Spark with 128GB memory. Performance may vary. SHOP-26

SHOP-27: Testing as of November 2025 by AMD. All tests conducted in LM Studio 0.3.30 (Build 2). Vulkan llama.cpp v 1.57.1 used with Ubuntu 24.0.4.3 and therock-gfx1151-7.9rc1 for AMD Ryzen™ AI Max+ 128GB. Flash Attention = ON in all cases. MMLU and GPQA scores as-reported from research papers and github repos (gpt-oss-120b & gpt-oss-20b Model Card.& meta-llama/Meta-Llama-3-70B · Hugging Face). Cloud-quality statement from OpenAI “The gpt-oss-120b model achieves near-parity with OpenAI o4-mini on core reasoning benchmarks..” (Introducing gpt-oss) AMD Ryzen™ AI Max+ 395 PRO on an HP Z2 Mini G1a with 128GB memory.  200 billion parameters (in 4-bit quant) require 128GB of unified memory. The AMD Ryzen™ AI Max+ was the first x86 processor to launch with 128GB of unified memory. Performance may vary. SHOP-27