惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
aimingoo的专栏
aimingoo的专栏
IT之家
IT之家
N
Netflix TechBlog - Medium
MyScale Blog
MyScale Blog
雷峰网
雷峰网
T
Tailwind CSS Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
The Blog of Author Tim Ferriss
S
Schneier on Security
C
CERT Recently Published Vulnerability Notes
Help Net Security
Help Net Security
云风的 BLOG
云风的 BLOG
GbyAI
GbyAI
I
InfoQ
H
Help Net Security
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
酷 壳 – CoolShell
酷 壳 – CoolShell
G
GRAHAM CLULEY
Blog — PlanetScale
Blog — PlanetScale
G
Google Developers Blog
I
Intezer
大猫的无限游戏
大猫的无限游戏
AWS News Blog
AWS News Blog
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
Spread Privacy
Spread Privacy
博客园_首页
宝玉的分享
宝玉的分享
量子位
T
Threatpost
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Security Latest
Security Latest
C
Cybersecurity and Infrastructure Security Agency CISA
SecWiki News
SecWiki News
H
Hackread – Cybersecurity News, Data Breaches, AI and More
博客园 - Franky
C
CXSECURITY Database RSS Feed - CXSecurity.com
T
The Exploit Database - CXSecurity.com
T
Tenable Blog
Know Your Adversary
Know Your Adversary
P
Proofpoint News Feed
The Register - Security
The Register - Security
V2EX - 技术
V2EX - 技术
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Last Week in AI
Last Week in AI
L
LangChain Blog
T
Tor Project blog
Stack Overflow Blog
Stack Overflow Blog
月光博客
月光博客

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
Local AI Hardware: Break Even in 2.6 Years?
© 2026 SkepticCTO LLC · 2026-05-30 · via Hacker News - Newest: "AI"
Economics of local agentic AI

As you may have noticed, large Mac Mini M4 Pros have disappeared.

Apple’s cute little desktop has become impossible to find. First, shipping delays stretched to sixteen weeks. Then, Apple pulled entire configurations from its US store. First, the 64GB Mac Mini was gone, and the 128GB and larger (196GB, 256GB, and 512GB) Mac Studio models soon followed. On its 2026 Q2 earnings call, Tim Cook revealed why. “Both of these are amazing platforms for AI and agentic tools,” he told investors, “and the customer recognition of that is happening faster than what we had predicted.”

Autonomous AI agents on local hardware (specifically OpenClaw and later Hermes Agent) exploded onto the AI community. OpenClaw now has over 350,000 GitHub stars, overtaking React to become the most-starred software project. Hermes Agent, from Nous Research (and OpenClaw variants such as NVidia NemoClaw), follows a similar philosophy: give it a task through messaging apps like WhatsApp or Telegram, and it will independently work on your behalf.

These agentic frameworks can use local LLMs. Their rise has triggered a hardware buying spree. If you own the hardware, you can escape from your LLM API bill forever…

But being generous, it will take 2.6 years to recoup your investment! Let’s see why…

The Setup

You can’t buy a new Mac Studio with 128 GB of memory right now. Viable alternatives include the NVidia DGX spark (the cheapest being a 128 GB Asus at $3494) and the Ryzen AI Max+395 (the cheapest being a 128 GB GMKtec EVO-X2 at $3,299). The important aspect of these machines is that they use 128GB of unified LPDDR5X memory. “Unified” means that we can allocate memory for either the CPU or GPU, which at 128GB allows us to run very capable mid-sided LLMs with large contexts (such as 256K tokens).

Let’s start with GMKtec EVO-X2: $3,299.

For the model, let’s use Gemma 4 26B-A4B. This is a rather capable mixture-of-experts model with 25.2 billion parameters (3.8 billion active). It runs well on this hardware, benchmarks competitively with models several times its size, and represents the class of open-weight models people are actually deploying for agent workflows.

For the cloud comparison, we’ll use DeepInfra, a pretty cheap provider for this model: $0.07/M input, $0.34/M output (roughly $0.10/M overall).

The Setup

The (Generous) Math

We’ll apply a variant of the “Principle of Generosity”: when we make assumptions, we will choose numbers that favor buying the hardware. That way, if local inference still looks bad, it won’t be because of our assumptions.

Assumption 1: We’ll get our money’s worth and run the machine at maximum inference 24/7.

Assumption 2: We’ll focus on output tokens because they represent the best savings using local inference. Output tokens cost $0.34/M and the machine’s peak concurrent output rate is about 120 t/s (achievable at 5–8 concurrent requests). For comparison, at $0.07/M and 240t/s, input token savings $529.80/year, less than half of the savings for input tokens calculated below.

So:

120 tokens/sec × 31,536,000 seconds/year = 3,764,320,000 tokens/year
3,764,320,000 × $0.34/1,000,000 = $1,279.07/year in avoided API costs

Break-even: $3,299 ÷ $1,279/year ≈ 2.58 years

A local AI machine running full-bore, 24/7, will pay for itself in about two and a half years.

For a more casual single-user agentic workflow with10% utilization, it’s going to take 25 years to break even.

2.58 year wait

But Wait… There’s More

The break-even calculation above leaves out some additional costs:

Electricity. The EVO-X2 draws roughly 140W under sustained inference load, which at $0.16/kWh adds approximately $195/year (if we ran it 24/7).

Maintenance. The AMD ROCm/Vulkan/llama.cpp software stack is rapidly changing. Backend updates, driver regressions, and model compatibility issues can easily cost you hours mer month.

Depreciation. Amazon, OpenAI, and Anthropic have depreciation schedules of 5.5 to 6 years for their inference hardware. Consumer APU hardware typically turns over faster (potentially 3-5 years).

Where Local Inference Does Make Sense

There are great reasons to run local inference.

Privacy and compliance: Local inference is a simpler way to meet HIPAA, attorney-client privilege, classified research, and GDPR data residency requirements. If the data cannot leave your network, the cost comparison doesn’t matter.

Air-gapped environments:.The same logic applies to defense, certain financial institutions, secure R\&D.

Very high sustained volume: Against open-weight models on APIs like DeepInfra, Braincuber estimates you’d need roughly 500 million tokens per day of sustained output before self-hosting a 70B-class model beats the API on total cost of ownership.

Learning and experimentation: Understanding inference infrastructure, running fine-tuning experiments, gaming, or simply having the machine for development would justify these powerful machines. The AI capability then comes along for “free”.

However, It May Get Worse

Today’s $3,299 price for a 128GB EVO-X2 today is not normal. In September 2025, the same machine sold for $1,799 at MicroCenter.

DRAM prices surged 90% from Q4 2025 to Q1 2026. Data centers now consume an estimated 70% of all memory chips produced worldwide, and analysts say this will not be a simple temporary shortage. A leaked SK Hynix internal analysis indicates that LPDDR5X supply at consumer prices will not normalize until 2028 or 2029, because almost all new memory lines are going to AI data centers first.

In addition AMD officially announced the Ryzen AI Max+ 400 series (“Gorgon Halo”) on May 21, 2026. The PRO 495 will support up to 192GB of unified memory, with systems shipping from ASUS, HP, and Lenovo in Q3 2026. That may reduce prices for older models, but won’t help reduce the demand for memory.

There is irony here: the impressive AI advancements driving the demand for local inference hardware is also the reason that hardware costs are so high.

So if you are setting up your dream OpenClaw box with local inference, seriously consider if you really need to keep your data local. While it is tempting to view local LLMs as “free”, you’ll be waiting a while (at least 2.6 years) to break even.


Robert “Butch” Buccigrossi, Ph.D., is CTO of TCG, Inc. and founder of SkepticCTO. He writes about AI from a scientific skeptical evidence-based perspective.