惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
T
The Blog of Author Tim Ferriss
GbyAI
GbyAI
雷峰网
雷峰网
Last Week in AI
Last Week in AI
人人都是产品经理
人人都是产品经理
F
Fortinet All Blogs
酷 壳 – CoolShell
酷 壳 – CoolShell
T
Tailwind CSS Blog
Y
Y Combinator Blog
J
Java Code Geeks
S
SegmentFault 最新的问题
罗磊的独立博客
爱范儿
爱范儿
F
Full Disclosure
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
V
V2EX
G
Google Developers Blog
腾讯CDC
美团技术团队
Martin Fowler
Martin Fowler
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
D
DataBreaches.Net
大猫的无限游戏
大猫的无限游戏
博客园 - 【当耐特】
B
Blog
Recorded Future
Recorded Future
月光博客
月光博客
Blog — PlanetScale
Blog — PlanetScale
IT之家
IT之家
N
Netflix TechBlog - Medium
P
Proofpoint News Feed
云风的 BLOG
云风的 BLOG
博客园 - 聂微东
阮一峰的网络日志
阮一峰的网络日志
B
Blog RSS Feed
aimingoo的专栏
aimingoo的专栏
W
WeLiveSecurity
Recent Announcements
Recent Announcements
P
Palo Alto Networks Blog
Apple Machine Learning Research
Apple Machine Learning Research
MongoDB | Blog
MongoDB | Blog
G
GRAHAM CLULEY
A
Arctic Wolf
AWS News Blog
AWS News Blog
Project Zero
Project Zero
博客园 - Franky
V
Vulnerabilities – Threatpost

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
Local AI Is Not Ready for Coding. Yet? - Methodos Mechanicus
Markus Lachinger · 2026-06-16 · via Hacker News - Newest: "AI"

In the GasTown post I made a throwaway claim: your worker agents — the polecats that actually write the code — don’t need your strongest model. Run them on Sonnet, not Opus, and your API bill stops looking like a car payment.

That got me thinking about the obvious next step. If a polecat is just a scoped worker that picks up a task and implements it, why pay for an API at all? Why not run it on a model on my own machine? No per-token bill, no rate limits, no data leaving the building. For a fleet of workers that spin up dozens of times a day, “free and private” is a very loud pitch.

So I tested it. Properly — not “ask a local model to write fizzbuzz,” but drop a local model into our real production harness and tell it to go do a job, exactly like the cloud agents do.

The short version: the model that fits on your laptop can’t carry the job, and the model that can carry the job doesn’t fit on your laptop. The longer version is more interesting, because the local models were both more and less capable than I expected — sometimes in the same sentence.

The Setup#

I wanted a fair fight, so I gave the local models the exact same harness the cloud agents use.

  • Machine: a maxed-out MacBook Pro — M5 Max, 48-core GPU, 128 GB unified memory. This is as much model as consumer hardware will hold.
  • Serving: Ollama on the host, exposed over its OpenAI-compatible API, driven by the opencode harness so the agent gets real tools: bash, read, write, edit, grep.
  • The two contestants:
    • Devstral 24B — a dedicated local coding model (~25 GB at Q8). Small, fast, purpose-built.
    • Qwen3.5 122B-a10b — a 122-billion-parameter mixture-of-experts (~10B active), ~81 GB in memory at Q4. The most capable model I can physically load on this machine. When people say “run a frontier model locally,” this is roughly the ceiling of what that means on a laptop.
  • The harness: GasCity, the production successor to the GasTown setup I wrote about earlier. Same mental model — a Mayor that plans, polecats that implement, a Reviewer, a Refinery that merges. Work flows as beads (tasks) bundled into convoys, and a polecat runs a formula (a multi-step workflow) to take a task from “assigned” to “merged.”

The task itself was deliberately trivial: “Create a file with one line of specific text.” If a model can’t do that, nothing else matters. If it can, we learn where the wheels come off.

Before judging the local models, I ran the identical task through a Claude Sonnet polecat. It sailed through the whole pipeline — created the file, committed it, handed off to the Reviewer, got merged to main, closed the bead. No drama. That matters: it proves the harness, the dispatch, and the task are all sound. Everything that follows is the model’s contribution, not a broken setup.

Devstral (24B): Eager, Capable, and Can’t Spell#

Handed our full agent onboarding prompt — about 25 KB of role, rules, and protocol — Devstral did something I didn’t expect. It introduced itself.

“I’m here to assist you with software engineering tasks. To get started, please let me know what you need help with.”

Zero tool calls. It read 15,000 tokens of “you are an autonomous worker, here is how to find and execute your assigned task,” and concluded that the correct move was to wait politely for instructions. A frontier model treats that same prompt as a starting gun; Devstral treated it as a company handbook.

But here’s the twist that makes “local AI isn’t ready” too glib: when I stripped the ceremony and gave it a direct, explicit instruction — “create this file, use your write tool, then run ls — it just did it. Tool call, file written, verified. About three times out of four. The capability is real; it only shows up when the ask is concrete.

And then, on the trivial task, it did this:

The model that renamed the deliverable

I asked for a file named MINSTRAL.md. Devstral created minstrel.md — silently “corrected” my spelling and lowercased it — then explained, with total confidence, that this was fine “since filenames are case-sensitive on Linux.”

That justification is not just wrong, it’s backwards. Case-sensitivity is exactly why MINSTRAL.md and minstrel.md are two different files. The model produced a plausible-sounding sentence to rationalize an error it didn’t notice it was making. This is the local-model failure mode in miniature: confident, fluent, and subtly off — on a task with one requirement.

For glue code and scoped edits where you’re checking the output anyway, a 24B coding model is genuinely useful. For anything you’re not going to read line-by-line, that confidence-without-correctness is a tax you pay later.

Qwen3.5 (122B): Genuinely Agentic — and Genuinely Lost#

The 122B model is a different animal, and this is the result that surprised me most.

Handed the same full onboarding prompt that made Devstral freeze, Qwen self-started. It ran the work-discovery command on its own, found its assigned bead, claimed it atomically, checked its mailbox, and read the workflow definition — fourteen real, correctly-formed tool calls in sequence. This is not a model that “can’t use tools.” It understood it was an autonomous agent in a system and started operating the system.

Then it never wrote a single line of code.

It spent every remaining turn spelunking. Inspecting the convoy. Reading the convoy’s children. Dumping metadata. Checking dependencies on three different beads. Re-reading the workflow. Eighteen minutes of impeccable, purposeful-looking tool calls, and the actual task — write one line to a file — never happened. It got lost in the cartography of our system and forgot there was territory to cover.

I confirmed it wasn’t a context-window problem: it ingested the full ~15k-token prompt without truncation and chose to explore. The complexity didn’t overflow the model; it captured it.

Our harness wasn’t too complex for the model’s intelligence — Qwen is plenty intelligent. It was too complex for the model’s executive function. A frontier model holds “I am here to do ONE small thing” as an anchor while it navigates ceremony. The 122B local model let the ceremony become the task. That’s the line, and it’s not the line I expected to find.

The Fix: We Had to Strip the System to the Studs#

If complexity was the captor, the obvious experiment was to remove it. So I built a “local edition” of the worker: I cut the 25 KB onboarding prompt down to about 40 lines, and collapsed the seven-step workflow (load context → record pool → set up worktree → preflight → implement → self-review → submit) into two steps: do the thing, then hand it off.

Qwen, on the slim setup, reached the task and created the file with the exact right content. From “18 minutes, zero files” to “done” — purely by deleting instructions. That is a genuinely hopeful result.

But the tail was long. It first stopped the instant the file existed — declaring victory before committing or handing off — so I had to make the prompt scream that “done” means committed and submitted, not “file written.” On the next run it actually ran the full sequence… and then tripped over a configuration bug on my side (the test repo had a branch-name mismatch and no remote), which it handled by improvising a throwaway git repo and pushing into the void. Some of that was my fault, not the model’s. But the deeper point stands: a frontier agent would have noticed the git error and recovered. The local model barreled past it confidently — the same failure mode as minstrel.md, just with git init instead of a filename.

The lesson isn’t “it can’t be done.” It’s that every layer of abstraction you love about your agent platform is a layer the local model has to pay for in executive budget it doesn’t have. You can get there — by stripping the platform back down to almost nothing, which is most of what the platform was for.

The Hardware Wall#

Here’s the part that turns “interesting” into “not yet.”

Everything above was the most powerful model I can run on a $5,000 laptop, and it still couldn’t autonomously drive the harness. The natural response is “so use a bigger, smarter open model” — the genuinely frontier-competitive open weights like DeepSeek-V3 (671B) or GLM-4.6 (~355B). And you can! You just can’t do it on consumer hardware.

DeepSeek-V3 is roughly 400 GB quantized to 4-bit, ~750 GB at full precision. To serve it properly you want around 1 TB of GPU memory, which in practice means an 8× NVIDIA H200 node (1,152 GB of VRAM, comfortably enough).

What that costs, mid-2026:

  • To buy: an 8×H200 server / DGX-class box runs ~300,000to300,000 to 500,000+ once you add CPUs, networking, NVMe, and cooling. The GPUs alone are ~$35–40K each.
  • To rent: roughly 25–35perhour∗∗foran8−GPUnodeon−demand—callit∗∗ 25–35 per hour** for an 8-GPU node on-demand — call it **~20,000 a month if you keep it warm the way an always-on agent fleet needs.

So the bill to run a model that can actually drive an autonomous coding harness is somewhere between “a luxury car every month” and “a house.” Set against that, a hosted frontier API — billed per token, zero idle cost, someone else eating the DRAM shortage — stops looking like the expensive option and starts looking like the adult one.

The Apple-silicon footnote

Yes, you can load a 671B model on a single 512 GB M3 Ultra Mac Studio — it was the one consumer box that could hold it, at around $10,000. Two catches. First, Apple pulled the 512 GB tier in March 2026 amid the global DRAM squeeze, so you can’t even buy it now. Second, even when you could: ~6 tokens/second, and a ~14-minute wait just to ingest an 8,000-token prompt before it emits a single output token. Our agent’s onboarding prompt alone was ~15,000 tokens. An agentic loop is dozens of those round-trips. The math doesn’t close.

That’s the whole thesis in one frame: the model that fits on your desk can’t carry the work, and the model that can carry the work needs a data center. The middle ground — frontier capability at consumer cost and consumer latency — doesn’t exist today.

What I Believe Now#

After actually running this, here’s where I land — and it’s not “local AI is a toy,” because that’s not what I saw.

  • Local models are genuinely capable at scoped, well-framed, single-shot work. Devstral writes glue code; Qwen makes real, correct tool calls and self-starts. If your use case is “private autocomplete and bounded edits I’m going to review anyway,” local is already good enough, and the privacy and zero-marginal-cost story is real.
  • Local models are not ready to be autonomous agents in a complex system. They lack the executive function to stay anchored to a goal while navigating operational ceremony, and they fail confidently — the wrong filename, the wrong git recovery, the plausible-but-false justification. In an unattended loop, confident-and-wrong is the expensive kind of wrong.
  • The harness has to meet them halfway. The single biggest unlock wasn’t a better model — it was a simpler system. If you want local workers, design a deliberately dumb, flat, push-the-task-at-them workflow. Everything elegant about your orchestration layer is friction to a model running on executive fumes.
  • The hardware curve is the real gate. This is why it’s “yet.” Small models are getting smarter every quarter, harnesses can be simplified, and one day the 100B-class model on your laptop will have the executive function the 122B one is missing today. But right now, the capable open models need $300K of NVIDIA or a data-center lease, and that math beats “just call the API” for almost everyone.

When non-engineers can’t get good code out of a chatbot, it doesn’t mean the tools are bad — it means the tools amplify the operator. Local agentic coding is the same story, one level deeper: it amplifies not just your engineering skill but your systems design. Give a local model a clean, narrow, well-lit path and it’ll walk it. Give it your beautiful, abstract, production orchestration layer and it’ll admire the architecture until it runs out of time.

Not ready. But the “yet” is doing real work in that sentence — and I’d bet on it shrinking fast.

If you want to try this yourself

  • Start with a dedicated coding model (Devstral-class) for scoped edits before reaching for a giant generalist.
  • Give it explicit, concrete instructions — local models are far worse at inferring the task from context than at executing a stated one.
  • Radically simplify any agent workflow you hand it. Two steps, not seven. Push the task in; don’t make it go discover the task.
  • Keep a human (or a frontier model) on the review pass. The failure mode is confident-and-subtly-wrong, which is exactly what review catches.

Where do you draw the local-vs-cloud line?

Have you run local models as real agents — not just chat? Where did they surprise you, and where did they fall over? And is “free and private” worth the executive-function tax for your workload? Tell me in the comments.