惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
腾讯CDC
Recent Announcements
Recent Announcements
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Hugging Face - Blog
Hugging Face - Blog
H
Help Net Security
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Last Week in AI
Last Week in AI
博客园_首页
D
DataBreaches.Net
P
Proofpoint News Feed
云风的 BLOG
云风的 BLOG
V
Visual Studio Blog
月光博客
月光博客
Jina AI
Jina AI
Stack Overflow Blog
Stack Overflow Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 【当耐特】
Vercel News
Vercel News
WordPress大学
WordPress大学
J
Java Code Geeks
博客园 - 聂微东
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
U
Unit 42

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
Stop chasing parameter counts. Build the toolbelt instead...
Ángela López Mendoza · 2026-06-02 · via DEV Community

For the last few months I've been building Tlamatini, an open-source local-first AI developer assistant. Along the way I kept bumping into the same assumption — both in articles and in my own head — that to build something useful, you need the biggest model you can afford. GPT-4. Claude Opus. Llama 70B at minimum.

Then I started actually shipping with smaller local models, and I learned something that flipped my thinking.

The real lesson

A 20B-parameter LLM, given the right tools, the right agents, and skills fine-tuned to your operating procedures, is good enough to power most of your company's real workflows.

Parameter count is not the bottleneck. The bottleneck is whether the model can act — and that's a tools problem, not a parameters problem.

What "the right tools" actually means

In Tlamatini, we wired the LLM into 75 concrete capabilities:

  • Shell and Python execution
  • File operations
  • Browser automation (Playwright)
  • Screenshots and keyboard/mouse control
  • Email, Telegram, WhatsApp bridges
  • A hybrid RAG pipeline (FAISS + BM25) so the model sees the right code, not random chunks
  • Multi-agent orchestration via ACPX — the assistant can delegate sub-tasks to Claude Code, Cursor, Codex, or Gemini CLI and relay output between them

With this toolbelt, a 20B model running locally on Ollama can:

  • Read your codebase and answer accurate questions about it
  • Refactor a module, run the tests, and report back
  • Open a browser, fill a form, screenshot the result
  • Build and flash firmware to an STM32 microcontroller (yes, really)
  • Chain all of the above into a single conversation

A 200B cloud model with no tools cannot do any of those things.

Why this matters for companies

Most internal AI projects fail because teams reach for the biggest model and the smallest scope. They get an expensive chatbot that drafts emails.

Flip it: give a modest model a real toolbox and skills fine-tuned to your actual operating procedures (your CRM, your ticketing system, your build pipeline), and you get an operator — something that participates in the workflow instead of describing it.

Local 20B + tools > cloud 200B + chat box. Almost every time.

The practical takeaway

If you're thinking about adopting AI in your company and the budget conversation is stuck on which API to pay for, consider stepping back:

  1. What are your repeatable operating procedures?
  2. What tools would an agent need to actually execute them?
  3. Can you wrap those tools cleanly enough that a local 20B model can call them reliably?

If yes, you don't need to send anything to the cloud. You don't need to pay per token. You don't need permission from a vendor. You just need to build the toolbelt.

That's what Tlamatini is — an open-source toolbelt and orchestration layer for local LLMs. Built in Django, runs on Ollama, GPL-3.0.

I'd love to hear from other people who've shipped agent systems on smaller local models — what's working for you? What's still painful? What tools made the biggest difference?