Last Updated on June 8, 2026 by
Author(s): Chew Loong Nian – AI ENGINEER
Originally published on Towards AI.
A 26-billion-parameter model has no business fitting in 15GB of memory and spitting out 193 tokens a second on a single consumer GPU. That is laptop-and-gaming-rig territory, not a datacenter. Yet that is exactly what Google’s new Gemma 4 QAT checkpoints do, and after digging into how they pulled it off, the part that stuck with me is not the speed. It is that the 4-bit version barely loses anything compared to the full-precision original. By every law of quantization I thought I understood, it should be noticeably dumber. It isn’t.

After the lead, the article breaks down why Gemma 4 QAT + Unsloth’s GGUF conversion is unusually effective: it quantizes during training so the model learns to be robust to 4-bit rounding, explains the typical PTQ quality loss, and describes how Unsloth fixes a subtle scale-mismatch bug that otherwise wipes out most of the benefit when converting to llama.cpp formats. It then provides concrete performance and memory numbers for different Gemma 4 variants (especially the 26B-A4B mixture-of-experts model), compares naive vs dynamic conversion accuracy, and summarizes the practical steps to run the model with llama.cpp, plus other deployment options (API server, Ollama/LM Studio, Unsloth Studio, vLLM/SGLang, MLX, and browser ONNX). Finally, it offers guidance on which model to choose based on available hardware, notes the remaining caveat that 4-bit is still 4-bit, and concludes that the usual quality-vs-speed tradeoff is collapsing—making the 26B-A4B feel like a near big-model experience on consumer GPUs.
Read the full blog for free on Medium.
Published via Towards AI
Towards AI Academy
We Build Enterprise-Grade AI. We'll Teach You to Master It Too.
15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.
Start free — no commitment:
→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day
→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages
Our courses:
→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.
→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.
→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.
Note: Article content contains the views of the contributing authors and not Towards AI.











