惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Microsoft Azure Blog
Microsoft Azure Blog
The Register - Security
The Register - Security
S
Securelist
Simon Willison's Weblog
Simon Willison's Weblog
T
The Exploit Database - CXSecurity.com
V
Vulnerabilities – Threatpost
NISL@THU
NISL@THU
P
Privacy & Cybersecurity Law Blog
V2EX - 技术
V2EX - 技术
O
OpenAI News
N
News and Events Feed by Topic
AI
AI
P
Proofpoint News Feed
Schneier on Security
Schneier on Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Cloudbric
Cloudbric
Help Net Security
Help Net Security
C
Cyber Attacks, Cyber Crime and Cyber Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Security Latest
Security Latest
Application and Cybersecurity Blog
Application and Cybersecurity Blog
L
LINUX DO - 热门话题
Cyberwarzone
Cyberwarzone
Scott Helme
Scott Helme
The Hacker News
The Hacker News
Hacker News - Newest:
Hacker News - Newest: "LLM"
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Google DeepMind News
Google DeepMind News
H
Hacker News: Front Page
C
Cisco Blogs
Webroot Blog
Webroot Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Hacker News: Ask HN
Hacker News: Ask HN
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
The Last Watchdog
The Last Watchdog
PCI Perspectives
PCI Perspectives
AWS News Blog
AWS News Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Know Your Adversary
Know Your Adversary
Latest news
Latest news
Forbes - Security
Forbes - Security
I
Intezer
Project Zero
Project Zero
C
CERT Recently Published Vulnerability Notes
T
Tenable Blog
TaoSecurity Blog
TaoSecurity Blog
S
Security @ Cisco Blogs
N
News | PayPal Newsroom
H
Heimdal Security Blog
W
WeLiveSecurity

Towards AI

Building AI Agents in Rust — part 4 | Towards AI The Verified Identity Agent Bridge | Towards AI You Can’t Prompt Your Away Your LLM Problems | Towards AI The Free Agent Trap | Towards AI Your Agentic Loop Will Drift. Here Is the KL Divergence Equation That Measures How Far It Has Wandered From Its Original Instruction. | Towards AI Beyond Chat: Processing Images, PDFs, and Documents with the OpenAI Adapter in Oracle Integration Cloud | Towards AI Building AI Agents in Rust — part 3 | Towards AI Self-Hosting Airflow at Home: Automating Stock Price Data Collection | Towards AI The 76-Hour Frontier: How the Takedown of Claude Fable 5 Birthed the Military-Industrial-AI Complex | Towards AI I Trained a Markdown File to Boost GPT-5.5 by 23 Points — It Shouldn't Work | Towards AI We Replaced ChatGPT With a Local AI Server. Six Months of Honest Data. | Towards AI What Really Makes Cars Pollute? A Data Science Deep Dive into CO₂ Emissions | Towards AI Training GPT-2 From Scratch on a GTX1050 | Towards AI Principal Component Analysis (PCA): Theory, Mathematics, and Applications Build a Zero-Cost Web Automation Pipeline With OpenRouter, OpenClaw, and MediaUse I Gave Qwen3.7-Plus a Screenshot and It Found the Exact Pixel to Click for $0.40 Beyond the Prompt: Why Autonomous AI Agents Are Replacing the Chatbot Moonshot Cracked Claude Code’s Playbook with an MIT Terminal Agent and a $0.60 Model Connections, Roles, and Warehouses: Getting CoCo Desktop Production-Ready from Day One My First $5,000 Month Writing About AI Engineering on Medium Google Shrank Gemma 4 by 72% and Unsloth Fixed the 4-Bit Bug Nobody Else Caught on One 4090, and 4-Bit Shouldn’t Be This Good LangChain Explained: Understanding Models, Prompts, Chains, Memory, Indexes, and Agents TOON: Beyond JSON for LLMs Claude Code Casual, Pro, Elite: The Three Working Personas of Claude Code Mastery MiniMax M3 Decodes 1M Tokens 15x Faster — and It Shouldn’t Be This Cheap Using Amazon SQS for AI Agent Orchestration I Ran a 1.5B-Active Model on My Laptop That Embarrassed a 26B by 46 Points How to Build a Self-Improving Company with AI Part 3 — Implementation/Engine-Level: Choosing the Runtime That Gives You These for Free Part 2 — Serve-Level Speed: System Design That Stabilizes P95/P99 3-Part Series: LLM Latency in Production (Part 1) Claude Code: The AI Coding Partner Changing How Developers Build Software Claude Code Pitfalls: Claude Code Won’t Do What You Told It: A Troubleshooting Catalog Full-Stack Data Scientists for the Agentic Coding World Building Production-Grade AI Skills with Snowflake Cortex AI Function Studio I Tried 10 AI Agent Frameworks in 2026 — Here’s the Honest Guide I Wish I Had Earlier How One Spring Boot Optimization Saved Our Startup $30,000 a Year Inside Palantir AIP: How the World’s Most Controversial AI Platform Actually Works What Is a Reverse Proxy? (And Why Every Backend Developer Should Care) What Claude Opus 4.8 Actually Changes If You’re Building Agents QWEN 3.7 Max Worked For 35 Hrs Straight And The Results Were Mind-blowing When LLMs Meet Knowledge Graphs on the Battlefield Fine-Tuning is Dead: Why Context Orchestration Won in 2026 5 Things Broke When I Shipped a RAG + MCP Agent to Production. Google Co-Scientist: Hyper Scaling Research and Discovery Microsoft Just Embarrassed Browser Web Agents — 1,000 Lines Made GPT-5.4 Beat Opus 4.6 on 200 Web Tasks The Modern Data Stack Is Broken — Here’s How to Fix It With AI, Governance, and Real Architecture Building Production MCP Servers: What the Spec Won’t Tell You When Should an Agent Stop? The Anatomy of Termination Harness Engineering: The Layer That Matters More Than the Model AI Engineers Who Can’t Debug Are Getting Fired (Here’s How I Debug with Claude Code) Claude Code Memory: Why You Keep Explaining the Same Thing to Claude (and the Five Layers That Fix It) Claude Code Subagents: The Claude Code Feature You Skip Every Day (And Why It Quietly Wrecks Your Sessions) Agentic AI and the SMB Banking Advantage Claude Code: Spec-Driven Development — Why Your AI Coding Sessions Fall Apart at Hour Three The Real Cost of Agentic AI Nobody Budgets For SVM : 40 must visit Interview Questions (Part 2) Your AI Agent Works Perfectly in the Demo. Here Are the 6 Ways It Dies in Production. Unleashing the Power of ONNX for Speedier SBERT Inference Terraform vs CI/CD for Serverless Deployments Merve Noyan Stopped Writing Training Scripts — Her Agent Just Fine-Tuned 18 Models Solo for $11.40 Why Your Sales Forecast Is Always 20% Wrong (And How To Make It 12% Wrong) Genetic Cubic n{C/A} Ratios For Elementary Robotics Design Top 20 AdaBoost Interview Questions & Answers (Part 2 of 2) Agentic AI Vs AI Agents — What Are the Key Differences? LAI #127: The Infrastructure Layer of AI Is Becoming the Product Anthropic Caught Its Own AI Planning to Blackmail Engineers RNNs Cannot Think What Transformers Think Cheaply. ICLR 2026 Proved the Gap Is Exponential. Time Series Made So Easy My Aunt Got It on the Second Read Claude Cowork 101 | Towards AI Is 3-Bit KV Cache the Holy Grail? A Reality Check on Google’s TurboQuant LangGraph Multi-Agent Architecture: Building a Self-Critiquing AI Debate System AutoML on Autopilot | Towards AI I Ran This Open-Source AI Tool on a Messy Codebase and Got 71x Fewer Tokens — Here Is Exactly What Happened Month in 4 Papers (April 2026) AI Kept Forgetting My Notes. Fixing That Taught Me How It Actually Works. How ChatGPT Makes You Addicted Crack ML Interviews with Confidence: K-Nearest Neighbors (KNN 20 Q&A) The Event-Driven Blueprint: How I Scaled a Spring Boot System to 10 Million Kafka Messages/Day Building Vector Search? Why FAISS Alone Isn’t Enough TAI #202: GPT-5.5 Moves Codex Into Real Work Machine Learning System Design -The Model Serving Triangle, With One Forward Pass Flowing Through Every Trade-off (Part3) AI Orchestration in Action: How MuleSoft and LLMs Fuel the Future of Enterprise AI GPT-4 Has 1.8 Trillion Parameters. It Uses 2% of Them Per Token. Part 20: Data Manipulation in Multi-Dimensional Aggregation A Fundamental Introduction to Genetic Algorithm -Part Two TAI #200: Anthropic’s Mythos Capability Step Change and Gated Release From Notebook to Production: Running ML in the Real World (Part 4) Sqribble’s Template‑Driven Document Automation Anthropic Just Shipped the Layer That’s Already Going to Zero Long-Term vs Short-Term Memory for AI Agents: A Practical Guide Without the Hype The L1 Loss Gradient, Explained From Scratch Your Postcode Is Deciding Your Care. I Built a Pipeline to Prove It. I Directed AI Agents to Build a Tool That Stress-Tests Incentive Designs. Here’s What It Found. Your System Prompt Is the Product — Not the Feature The LLM Wiki Trend Has a Retention Problem Nobody Mentions Top 20 Data Preparation Interview Questions and Answers (Part 2 of 2) LAI #122: Word Embeddings Started in 1948, Not With Word2Vec Top 15 Computer Vision Datasets [2026] 40 Generative AI Interview Questions That Actually Get Asked in 2026 (With Answers)
Choose Wisely: Models Should Follow Your Use Case. | Towards AI
Dhanush Kandhan · 2026-06-25 · via Towards AI

Originally published on Towards AI.

Choose Wisely: Models Should Follow Your Use Case.
Choose Wisely: Models Should Follow Your Use Case. — By Dhanush Kandhan

A guy in my builder’s discord group blew his entire Codex subscription in eleven days. Two weeks into the month, nothing left. You know what he was building? A billing feature in his SaaS. Not a compiler. Not an operating system kernel. Not a real-time physics simulation. A billing page with subscriptions, invoices, and a Dodo Payments webhook that doesn’t send duplicate emails.

He said it with the exhausted pride of someone who just pushed to prod at 2 AM (we devs are batmans, right?). I nodded. I didn’t say anything. But inside I was doing the mental math.

I run my full AI stack coding agent, agent workflows, browser automation, speech to text, for around $10 — $15 a month. And I ship. Regularly (my github is proof for that). With billing features and everything.

That conversation is what this post is about.

The Benchmark Theater We All Fell For

Let me describe a pattern you’ve probably noticed.

A big AI company/lab drops a new model/version. The announcement lands. Within hours, everyone on X is posting about it. “Our model built a C compiler from scratch.” “Our model achieved gold on the International Math Olympiad.” “Our model solved problems that researchers said required human-level reasoning.”

Image Credits: Faiapp Meme Creator

The posts get thousands of likes. Engineers screenshot the benchmark charts. Someone puts together a thread comparing it to the previous generation. Replies flood in from founders saying they’re switching immediately.

Then someone from Chennai quietly tries it on their actual codebase and reports back that it’s roughly the same as before for their use case. This tweet gets eleven likes.

I’m not mocking the benchmark results. Building a C compiler is impressive. Scoring on the IMO is legitimately hard. These results tell you something real about what the model is capable of in controlled settings.

But here is the question nobody asks loudly enough: when was the last time your actual work required an AI to build a C compiler?

Look at what you built last week. Probably a REST endpoint. A React component that talks to it. Some data validation logic. An email template. A webhook handler. A cron job that moves rows between two database tables. Maybe a RAG pipeline if you’re in the AI space. Something with auth. Something with payments.

You are not building compiler infrastructure. You are building software for users. Web apps. Mobile apps. Developer tools. Internal automation. The kind of work that, individually, each piece looks boring on a benchmark slide but collectively represents most of the software being written on earth today.

The benchmark score tells you the ceiling of what a model can achieve on curated academic tasks. It does not tell you whether the model is the right tool for your Monday morning standup’s ticket queue.

I learned this slowly. And expensively.

What “Open Source” Actually Means Here? (It’s Not One Thing)

Before I get into the specific models, I need to clear up something that trips up engineers constantly. When someone says a model is “open source,” they usually mean one of two very different things, and conflating them leads to bad decisions.

The first is open weights. The actual model parameters, the billions of floating point numbers that encode what the model knows are publicly available. You can download them. You can run them on your own hardware. You can fine-tune them on your own data. You can deploy them inside your own VPC and never send a single token to anyone else’s server. You can modify the architecture and release derivatives. Models like GLM-5.2, DeepSeek V4, Kimi K2.6, and Nemotron from NVIDIA are all open-weight models. The weights live on Hugging Face. Most of them ship under MIT licenses, which means you can use them commercially without paying anyone a licensing fee.

The second is what most of the subscription-based coding tools are: API access. You get to call their endpoint. The model runs on their servers. Their data retention policy applies to your prompts. Their pricing can change next quarter. If their infrastructure has issues on the day you have a demo, that is your problem too. You never see the weights. You cannot run it locally. The model is theirs; you are renting access.

The practical difference matters more than most engineers realize until they’ve felt it.

With open weights, your inference cost is literally your compute. You can run through OpenRouter or Together AI and pay per token with no monthly subscription, switching to a better model the day it ships. You can cache aggressively. You can self-host if the data sensitivity requires it. You are not locked into anyone’s pricing model.

There is also a comfortable middle path, which is what I run: open-weight models accessed through inference providers. Pay per token, no subscription, full flexibility to switch, and the per-token cost is typically a fraction of what the closed model APIs charge.

The Stack. For Real.

I’ve read too many “why I use open source models” posts that are basically just “open source good, closed source bad” with a Hugging Face link at the bottom. Useless. Let me be specific.

GLM-5.2 for Coding via OpenCode

When GLM-5.2 dropped from Z.ai, the Beijing-based lab that used to be called Zhipu AI the X(twitter) reaction was something. Aravind Srinivas posted about it. Guillermo Rauch appreciated it. The Artificial Analysis Intelligence Index ranked it at 51 points, which put it above DeepSeek V4 Pro, Kimi K2.6, and even some Google models. On their GDPval-AA v2 metric, which is their best approximation of real agentic task performance, GLM-5.2 roughly matched GPT-5.5.

But you know how it goes. X(Twitter) energy is its own genre. I do not make infra decisions based on who gets quote-tweeted by whom.

So I used it. On a $10/month OpenCode Go plan, using it daily. The billing feature I built with it subscriptions, metered usage, invoice generation, Dodo webhook handling with idempotency keys so the emails don’t duplicate, the model handled all of it without me holding its hand through every function. I was not babysitting it. I was shipping.

Technically, GLM-5.2 is a Mixture-of-Experts model with 753 billion total parameters. The “Mixture-of-Experts” part is important enough that I’ll explain it properly in a section below. The context window is one million tokens, which sounds like a spec-sheet number until you actually try to feed it your entire backend codebase and watch it reason across files you thought it would lose track of.

The interesting architectural detail is something Z.ai calls IndexShare. Here’s the problem it solves: at one million tokens, standard transformer attention is computationally brutal. The cost grows quadratically with context length, so a 1M token context isn’t just ten times more expensive than a 100K context, it’s more like a hundred times more expensive. IndexShare gets around this by reusing the same lightweight indexer across every four consecutive sparse attention layers instead of computing a new one for each. At 1M tokens, this cuts the per-token FLOPs by 2.9 times. That is not a minor tweak. That is what separates “supports 1M context” on a benchmark slide from “can actually use 1M context in production without your inference costs going vertical.”

The thing that sold me was not the benchmark. It was the day I ran it against a payment service codebase I’d inherited from a previous project a thing with four different Dodo event handlers, some legacy subscription logic, and a webhook processor that had comments like “TODO: figure out why this sometimes fires twice” dating back to 2021. GLM-5.2 read the whole thing, understood the context, and helped me fix the duplicate-fire issue without me having to summarize what each file did. That was the moment.

Kimi K2 Series for Agent Workflows

For Hermes, my internal automation system that handles repetitive background tasks and orchestration workflows. I’ve been using Kimi models from Moonshot AI.

The Kimi series has moved fast. K2 in July 2025. K2.5 in January 2026. K2.6 in April 2026. Each release meaningfully closed the gap with closed frontier models. K2.6 is where I landed.

It’s a one-trillion-parameter MoE model, but only 32 billion parameters are active per token, which means inference cost is roughly that of a 32B model, not a trillion-parameter model. On SWE-Bench Pro, it ties GPT-5.5 at 58.6%. API pricing is around $0.95 per million input tokens and $4.00 per million output. GPT-5.5 is considerably more expensive. The math is not subtle.

What matters for agent use cases is coherence across long chains of tool calls. A lot of cheaper models sound fine in isolation but start going sideways somewhere around tool call fifteen in a chain of fifty. They lose the thread. They start contradicting earlier steps. They confidently do the wrong thing.

Kimi K2.6 has what Moonshot calls Agent Swarm, a native architecture for multi-agent coordination. The model can decompose a complex task into parallel sub-agent workstreams and coordinate the outputs. In practice, for my Hermes workflows, this means I can set up long-running automation tasks, leave them running, and come back to coherent results rather than having to babysit the run like a new intern’s first week.

MiMo models from Moonshot handle lighter tasks and shorter context requirements in the same system. Think of it as routing: not every task needs the full K2.6 capacity, so lighter tasks go to lighter models and costs stay proportional to complexity.

DeepSeek V4 and Nemotron for Browser Automation

This is where my setup gets a bit unusual and I want to explain the reasoning.

Browser OS automation, interacting with web UIs, extracting structured data, triggering workflows across tools, benefits from a model that can reason across long sessions while staying fast enough that the loop doesn’t feel dead. You want the model to remember what it did two hundred steps ago, but you also need tokens per second to be high enough that the automation completes before you’ve had time to make and finish your second cup of coffee.

DeepSeek V4 handles the reasoning and planning layer. The V3 and V4 lineage matters here: DeepSeek demonstrated that you can train frontier-quality models for around six million dollars, compared to the hundreds of millions that Western labs were spending on comparable generations. This was not just a cost story, it was a signal that the training efficiency techniques they developed were genuinely novel. V4 uses Compressed Sparse Attention, which compresses token sequences into summary representations and selectively attends via top-k selection. This is what makes the one-million-token context viable for long automation sessions where the agent genuinely needs to remember state from hours ago.

Nemotron from NVIDIA is architecturally the most interesting thing I use. The Nemotron 3 family — Nano, Super, Ultra — is built on a hybrid Mamba-Transformer Mixture-of-Experts architecture. The key decision: replace most self-attention layers with Mamba-2 layers. Standard transformer attention has a KV cache that grows linearly with context. That means memory pressure climbs continuously as the context gets longer. Mamba-2 layers maintain a constant-size state instead of a growing cache, so the memory footprint stays flat regardless of how long the context is. For automation workloads with very long running sessions, this is not a theoretical advantage. You actually feel it.

NVIDIA released Nemotron with open weights, the full training data, and the training recipes. Not just the weights. The recipes. Super has 120B total parameters with 12B active. Ultra goes to 550B total with 55B active. And NVIDIA built and evaluated it explicitly with developer harnesses in mind OpenCode, OpenHands, coding review loops, not just as a chat interface. That shows up in how reliably it handles multi-step tool use without losing track of the task structure.

NVIDIA Parakeet for Speech to Text

I’ll keep this one short because the experience says it better than any spec.

I was using WhisperFlow for voice input. Perfectly fine. Then a Chennai guy and I say this with full affection, because if you know, you know mentioned Parakeet at a tea kadai (cafe) discussion that somehow turned into a thirty-minute model comparison session. I tried it the next morning with handy.computer (one my fav tool).

Parakeet TDT 0.6B v2 is 600 million parameters, trained on 64,000 hours of diverse audio, ranked first on the Hugging Face Open ASR leaderboard, 6.05% word error rate, and inference running 50 times faster than comparable models. The architecture is a FastConformer with a TDT decoder, handles up to 24 minutes of audio in a single pass. Actually worth for their benchmark hypes.

But what the numbers don’t tell you: it handles Indian-accented English well. My Tamil-inflected English, the way I say “idempotent” which is apparently not how Americans say it, function names like handleDodoWebhookRetry spoken aloud mid-thought, it transcribes all of this without losing the plot. I was skeptical for about a morning. I was not skeptical after that. By the way, you need some customization on top of this 🙂

Bye bye WhisperFlow. The Chennai recommendation stood.

Why MoE Models Are Everywhere Right Now

You noticed that almost everything I mentioned uses Mixture-of-Experts. That is not a coincidence. It’s worth understanding why this architecture has become the dominant pattern and cool in 2025 and 2026.

The classic neural network problem: to get smarter, you need more parameters. More parameters means higher memory usage, slower inference, and higher cost. Dense models, where every parameter activates for every token scale poorly once you get into the hundreds of billions.

MoE breaks this by doing something more like what specialized teams do. Instead of one giant generalist brain, you train a collection of expert sub-networks, each specializing in different kinds of knowledge. When a token arrives, a routing mechanism itself learned during training which decides which experts are most relevant and activates only those. For Kimi K2.6, this means one trillion total parameters but only 32 billion active per token.

Think of it this way. Imagine you have a startup with 200 employees. When a customer files a billing dispute, you don’t put all 200 people in the room. You route it to the two finance people and one customer success manager who are actually relevant. The other 197 people keep doing their own work. You get the collective knowledge of the full organization with the throughput of a small focused team.

That is roughly what MoE does, except the routing is learned from data and operates at the level of each token.

What this means for cost: you pay for inference on 32B active parameters, not 1T total. The model learned from the full trillion-parameter training graph and carries that knowledge, but serving it costs the same as serving a 32B dense model. You get frontier-level capability at mid-tier inference cost.

If MoE sounds excites you, check it out about more in core: https://huggingface.co/blog/moe (read it, if you’re deep in AI/ML core & Maths, else it burns you)

The Privacy Question Nobody Asks Until It’s Urgent

There’s a conversation I’ve watched happen too many times in startup engineering rooms.

Product launches. Engineers are excited. They’ve integrated a closed model API and it’s working beautifully. Three months later, a new client asks: “Where does the data go when we call your AI?” And the room gets quiet because nobody actually read the data retention section of the API terms of service. Something personal experience, never mind 😛

This is not paranoia. It’s engineering responsibility.

When you send prompts to a closed model API, you are sending data to another company’s infrastructure under their terms. For most personal projects and consumer tools, this is fine. For production workloads with proprietary business logic, sensitive user data, anything that touches financial information, healthcare, legal documents, or regulated industries “we accepted the terms” is not an answer a serious engineering team should be comfortable with.

Open weights change this entirely. You can run on your own infrastructure. Your prompts never leave your environment. No model provider has access to your users’ data, your codebase, or your business logic. For a startup building in a regulated space, this capability alone can be worth paying more for.

Even if full self-hosting isn’t feasible, using open-weight models through smaller inference providers means you’re working with companies whose data handling practices you can actually scrutinize, audit, and ask direct questions about. The opaque data practices of the largest closed model providers are harder to interrogate.

I’m not saying don’t use closed models. I use them for specific things where the integration value justifies it. I’m saying make the decision deliberately, not by default.

The Numbers, Plainly

Here is what my actual monthly spend looks like: GLM-5.2 via OpenCode Go Plan at $10. Kimi K2.6 API usage for agent workflows at a cost low enough that I’m tracking it as a rounding error. Parakeet running locally at whatever my laptop’s electricity costs. Nemotron and DeepSeek API usage for automation workflows.

Total: somewhere in the $10–15 range per month. Some months less.

A Codex subscription is $20. Claude Max is more. A team running both which is what a lot of engineers are doing right now is spending significantly more than that every month. Without necessarily getting significantly more output.

This is not a post about being cheap. If you’re at a company with real revenue and the subscription tools improve your team’s velocity enough to justify the cost, that’s a real calculation and it might come out in favor of the subscriptions.

But for the solo developer, the indie hacker, the early-stage startup with three engineers and a cloud bill that’s already too high, this gap is real money. The money you don’t spend on AI subscriptions can buy you another month of runway. Or better tooling. Or a good engineer who can review your architecture before you regret it.

What I’d Tell Myself Months Ago

Months ago I was very much in the “use whatever has the best benchmark and the most X(Twitter) engagement” camp. I’m not embarrassed about this. Everyone goes through it. The social proof is loud and the alternative requires you to actually experiment with models that most of the English-language tech press isn’t writing about.

What I’d tell myself then is this: the benchmark is a ceiling test in a controlled environment. Your production use case is a different exam entirely. Run the model on your actual task for a week before making a decision. Not on a toy example. On a real problem from your actual codebase with real edge cases.

The open-weights ecosystem in 2026 is not what it was in 2023. GLM-5.2 matching GPT-5.5 on real agentic tasks at one-sixth the cost is not a quirk it’s the result of serious engineering talent in Chinese AI labs working on the same problems with access to serious compute. DeepSeek showing that frontier models don’t require nine-figure training budgets reset a lot of priors across the industry. Kimi shipping models that tie closed frontier benchmarks with full open weights and MIT licenses means the argument for paying subscription prices gets harder to make every quarter.

Try new models when they drop. Specifically, spend time with the ones that aren’t getting the most X(Twitter) attention, because the ones that are getting attention are also getting adopted by everyone else, which means the differentiation is gone. The interesting discoveries come from the models the generalist tech press hasn’t written fifteen threads about yet.

Being Generalist is fine dude 🙂

And if you find a model that works exceptionally well for your use case contribute. GitHub Sponsors, bug reports, eval contributions, writing about your experience. The labs/companies releasing open models are running on tighter margins than the closed model companies. The community is part of what makes the ecosystem stay open.

TL;DR: the best model for your problem is not the one with the highest score on a benchmark you will never run. It’s the one that does your actual task reliably, at a cost that makes sense for your situation, under terms you understand.

Everything else is marketing. Some of it is very well-produced marketing, but still.

Choose your model like you choose your tech stack, based on what you actually need to build, not based on what is getting the most applause on the internet this week.

Now if you’ll excuse me, I have a billing feature to ship with him.

Image Credits: Tenor Gifs

Published via Towards AI