惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
Engineering at Meta
Engineering at Meta
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Martin Fowler
Martin Fowler
雷峰网
雷峰网
Recent Announcements
Recent Announcements
博客园 - 叶小钗
F
Full Disclosure
C
Check Point Blog
A
About on SuperTechFans
L
LangChain Blog
Vercel News
Vercel News
T
The Blog of Author Tim Ferriss
博客园 - 司徒正美
C
Cybersecurity and Infrastructure Security Agency CISA
C
CXSECURITY Database RSS Feed - CXSecurity.com
Cisco Talos Blog
Cisco Talos Blog
Hugging Face - Blog
Hugging Face - Blog
Recorded Future
Recorded Future
MongoDB | Blog
MongoDB | Blog
Project Zero
Project Zero
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Know Your Adversary
Know Your Adversary
D
Docker
U
Unit 42
酷 壳 – CoolShell
酷 壳 – CoolShell
T
The Exploit Database - CXSecurity.com
L
Lohrmann on Cybersecurity
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Latest news
Latest news
V
Visual Studio Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
L
LINUX DO - 最新话题
Google DeepMind News
Google DeepMind News
小众软件
小众软件
NISL@THU
NISL@THU
GbyAI
GbyAI
H
Heimdal Security Blog
Jina AI
Jina AI
I
InfoQ
PCI Perspectives
PCI Perspectives
P
Privacy International News Feed
P
Proofpoint News Feed
N
News and Events Feed by Topic
C
CERT Recently Published Vulnerability Notes
aimingoo的专栏
aimingoo的专栏
SecWiki News
SecWiki News
B
Blog RSS Feed
阮一峰的网络日志
阮一峰的网络日志

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
I Put Gemma 4 Behind My Homelab AI Gateway. This Is the Beginning.
Nic Lydon · 2026-05-13 · via DEV Community

Most model experiments start with a notebook, a benchmark script, or a quick API call.

This one started with a production-shaped question:

Can I swap out an entire model family that is currently serviing the default paths through my actual local AI gateway?

Not a side demo. Not a one-off curl. Not "look, it runs."

I mean the real route: the gateway that agents, background jobs, app surfaces, benchmark harnesses, and my own tools already call.

That is the experiment I started with Gemma 4.

This post is the beginning of that story, not the final verdict. I am writing it while the platform is still in the trial window. The follow-up will be more interesting: what stayed stable, what broke under real load, what got rolled back, and what I would keep after a week or two of actual use.

For now, this is the setup: what I changed, why I changed it, and what failed immediately.

The Platform Before The Swap

My local AI stack is built around a gateway I call Forge.

Forge gives callers one OpenAI-ish API surface and handles the messy parts behind it:

  • which model should answer this kind of request
  • which machine is hosting it
  • whether the model is hot, cold, deprecated, or on-demand
  • whether a request is chat, vision, embedding, transcription, code, extraction, or something else
  • whether a backend is available or should be skipped

The machines behind it are consumer hardware, not datacenter gear:

Host Role
Furnace Primary inference box, AMD Strix Halo, 96 GB unified VRAM allocated to the iGPU
Crucible Secondary AMD box for creative workloads, permissive models, and burst/bulk work
Anvil M4 Mac mini, useful for MLX/Metal paths and lightweight resident services

Before this experiment, the default local text path was mostly Qwen-family. That was not an accident. Qwen had become the operating baseline because it was predictable enough for a platform, not just impressive in isolation.

I had also tested other models. Devstral2, for example, was interesting enough to onboard and benchmark seriously. The smaller 24B variant was competitive in code scenarios, but it did not become the default path. The 123B model was too slow for the role I needed. That distinction matters:

A model can be good and still not be a good platform default.

That is the bar Gemma 4 had to clear.

Why I Did An In-Place Swap

I could have added Gemma 4 as another optional model and called it a day.

That would have been safer. It also would have taught me much less.

Instead, I treated it like a real migration. For the trial window, Gemma 4 took over the canonical roles that real callers already use.

Role Previous route Trial route
default chat qwen3.6-chat-35b-a3b gemma-4-chat-31b
priority chat qwen3-8b gemma-4-chat-26b-a4b
vision / multimodal qwen3-vl-30b-a3b gemma-4-multimodal-8b-e4b
prompt enhancement qwen3-4b gemma-4-multimodal-2b-e2b

The old Qwen routes were not deleted. They were marked deprecated with a planned rollback window. That gives me a clean flip-back path if the experiment does not earn its keep.

This is the part I think model posts often skip. A real model migration is not just "can I run it?" It is:

  • do I have the right weights on disk?
  • does my serving stack understand the architecture?
  • can I fit the hot set in memory?
  • do my existing aliases and callers still work?
  • can I roll back without spelunking through five repos?
  • do I have telemetry that will tell me the difference between model failure, gateway failure, and benchmark nonsense?

That last one matters more than I expected.

The First Failure Was Not The Model

The first deploy crashed.

Forge restarted cleanly. The model catalog showed the new Gemma 4 ids. The first smoke request hit the gateway, routed to llama-swap, and came back as a 502.

The useful error was one layer lower:

unknown model architecture: 'gemma4'

Enter fullscreen mode Exit fullscreen mode

The problem was not Gemma 4 quality. The problem was my serving binary.

My llama.cpp build was from April 1. It was 466 commits behind the branch I needed. The GGUF files declared general.architecture: gemma4, and the old build simply did not know what that meant.

So the first chapter of the Gemma 4 experiment was not prompting. It was infrastructure:

  • back up the existing build tree
  • rebuild llama.cpp with ROCm/HIP for Strix Halo
  • verify the new binary recognizes the Gemma 4 architecture
  • regression-check the existing Qwen route
  • restart the serving layer
  • smoke test through the actual gateway, not a side process

Only after that did the model start answering.

That is a useful reminder: "model support" is not a binary property. A model can be downloadable, quantized, and present on disk, and still be unusable because the serving stack is one architecture handler behind.

The Second Failure Was More Interesting

Once Gemma 4 loaded, the first real chat benchmark looked bad.

Not "a little worse than Qwen" bad. Broken bad.

On the initial chat-bench run, gemma-4-chat-31b failed the structured extraction and format-compliance scenarios. It was also slow enough that something was clearly wrong. These were not hard prompts. These were the boring, throughput-oriented tasks that agents and background workers need to complete cleanly.

A direct request showed the issue immediately:

<think>
The user is asking a basic arithmetic question...
</think>

2 + 2 = 4

Enter fullscreen mode Exit fullscreen mode

The model was spending the answer budget on a reasoning block.

For a human chat UI, visible thinking might be useful. For a benchmark expecting JSON, or an agent expecting a short answer, it is poison. The model can know the right answer and still fail the task because the caller never receives the shape it asked for.

This was familiar. Forge had already solved the same class of problem for Qwen3.

The fix was to make "thinking mode" a gateway policy, not a model identity.

Programmatic callers get:

{
  "chat_template_kwargs": {
    "enable_thinking": false
  }
}

Enter fullscreen mode Exit fullscreen mode

injected by default when the model family needs it.

Chat UIs can opt back in explicitly:

{
  "chat_template_kwargs": {
    "enable_thinking": true
  }
}

Enter fullscreen mode Exit fullscreen mode

That is the right abstraction for my platform. I do not want every agent, benchmark, worker, and internal tool to remember which local model family wraps output in thinking tokens this week. I want the gateway to know that once.

After that change, the chat benchmarks became meaningful. The three relevant routes - Gemma 4 31B, Gemma 4 26B-A4B, and the displaced Qwen3.6 35B-A3B baseline - reached the same pass-rate shape across the default chat scenarios.

The interesting result was latency. The 26B-A4B route was materially faster than both the dense 31B and the Qwen3.6 baseline on several workloads, while keeping the same pass rate in the corrected harness.

That is the kind of result I care about. Not "model X wins," but "model X belongs in this role."

Vision Exposed A Different Problem

The multimodal side taught a separate lesson.

I added a new VLM benchmark harness and ran the obvious first test. The initial scenario was too weak. It was good enough as a smoke test, but not good enough to tell me which model or host was better.

So I built a more discriminating scenario with three generated fixtures:

  • a bar chart that required OCR and chart reasoning
  • a code screenshot that required reading a function name and language
  • a homelab topology diagram that required identifying the hub and connected nodes

Then another problem appeared: concurrency.

Promptfoo's default concurrency sent multiple image-bearing requests at once during a backend startup window, which produced misleading 502s. Some errors appeared to implicate the wrong backend because parallel requests were failing around the same time.

Sequential runs told the truth.

With concurrency set to 1 and the output budget raised, the final VLM run passed cleanly across the tested local routes. The surprising part was not that Gemma 4 could read the images. The surprising part was that an M4 Mac mini running an MLX path was effectively tied with the AMD inference box on this small, practical vision benchmark.

That is not a leaderboard result. It is a routing result.

It tells me Anvil is not just a dev box. For some multimodal work, it is a useful inference target.

Audio Needed A Sidecar

Gemma 4's multimodal story is not just text and images. Audio is part of the interesting surface.

But my normal GGUF plus llama.cpp path did not support Gemma 4 audio input yet. The text and vision path worked through llama-swap. The audio conformer path did not.

So I built it as a sidecar:

  • a separate FastAPI worker
  • safetensors weights
  • HuggingFace Transformers
  • ROCm-specific PyTorch wheels
  • a Forge route at /v1/audio_qa

That route is intentionally not a replacement for Whisper.

Whisper remains the right tool for long-form transcription. Gemma 4 audio is more interesting for short audio understanding, audio Q&A, emotion or intent questions, and cross-modal prompts that combine audio and an image.

The first useful test was simple: the JFK sample clip. The route returned a good short transcription in under six seconds once warm. A 60 second clip correctly failed fast with a 413 because the audio path is capped at 30 seconds. An audio plus image prompt produced a coherent response grounded in both modalities.

That sidecar is not the end state. It is a bridge. When the standard serving path supports the audio input cleanly, the route can stay and the backend can change.

Again, that is the platform lesson: callers should not care which inference backend made the modality work.

What I Am Actually Testing

The easy version of this post would end with a benchmark table and a confident take.

I do not have that yet, and pretending otherwise would be silly.

What I have is the beginning of a platform trial:

  • Gemma 4 is now in the main chat, priority-chat, multimodal, and prompt-enhance roles.
  • Qwen is still available as a rollback path.
  • Devstral2 remains useful but not a default for this platform.
  • Forge now handles thinking-mode policy for both Qwen3 and Gemma 4.
  • The benchmark harness is better than it was before the experiment.
  • The audio path exists, but it is a sidecar until the normal serving stack catches up.
  • The real evaluation is now happening under actual workloads.

LLM Benchmark

The question I care about over the next week is not:

Is Gemma 4 better than Qwen?

The question is:

Which parts of the platform are better with Gemma 4 in the route, and which parts should move back?

That means watching boring things:

  • error rates
  • 429s and backend saturation
  • latency under background jobs
  • whether agent outputs stay clean
  • whether structured tasks remain reliable
  • whether multimodal routes are useful often enough to stay hot
  • whether the memory footprint is worth it
  • whether fallback behavior is predictable when the boxes are busy

This is less glamorous than a benchmark screenshot. It is also where the real answer lives.

The Takeaway So Far

The first day did not teach me that Gemma 4 is universally better. It taught me something more useful:

A model family becomes valuable when the platform can route it intentionally.

The gateway mattered more than any single model call.

Without Forge, I would have been debugging each app and agent separately. With Forge, the migration became a small number of role changes, a serving-stack rebuild, one generalized thinking-policy fix, and a better set of benchmarks.

That is the part I want to keep building toward: a local AI platform where model families can change without rewriting every caller, and where the system learns from real workloads instead of one-off demos.

This is the start of the Gemma 4 trial in my homelab.

In a week or two, I should have the more honest post: what survived real use, what got demoted, and what I would do differently if I were starting the swap again.