惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Engineering at Meta
Engineering at Meta
T
Threat Research - Cisco Blogs
V
Vulnerabilities – Threatpost
T
Tor Project blog
T
Troy Hunt's Blog
C
CERT Recently Published Vulnerability Notes
C
Cisco Blogs
W
WeLiveSecurity
Cloudbric
Cloudbric
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
爱范儿
爱范儿
Google Online Security Blog
Google Online Security Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Simon Willison's Weblog
Simon Willison's Weblog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Martin Fowler
Martin Fowler
Cisco Talos Blog
Cisco Talos Blog
F
Full Disclosure
MongoDB | Blog
MongoDB | Blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
I
Intezer
www.infosecurity-magazine.com
www.infosecurity-magazine.com
G
GRAHAM CLULEY
B
Blog RSS Feed
云风的 BLOG
云风的 BLOG
人人都是产品经理
人人都是产品经理
M
MIT News - Artificial intelligence
腾讯CDC
L
LangChain Blog
L
LINUX DO - 热门话题
H
Help Net Security
S
Schneier on Security
N
Netflix TechBlog - Medium
博客园 - Franky
酷 壳 – CoolShell
酷 壳 – CoolShell
Spread Privacy
Spread Privacy
S
Secure Thoughts
T
The Exploit Database - CXSecurity.com
P
Privacy International News Feed
P
Privacy & Cybersecurity Law Blog
Cyberwarzone
Cyberwarzone
A
About on SuperTechFans
NISL@THU
NISL@THU
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
D
DataBreaches.Net
The GitHub Blog
The GitHub Blog
Recorded Future
Recorded Future
雷峰网
雷峰网
AWS News Blog
AWS News Blog
V2EX - 技术
V2EX - 技术

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Run Gemma 4 on Your Laptop — A Hands-On Guide to Google's Latest Open Multimodal LLM
Jubin Soni · 2026-05-15 · via DEV Community

If you've been watching the open-source LLM space, you've probably noticed it's been a great couple of years. Llama, Mistral, Phi, Qwen — a whole zoo of models you can download and run on your own machine. Google's entry into that zoo is Gemma, and the fourth generation, Gemma 4 (released April 2, 2026), is the biggest leap yet: built from Gemini 3 research, multimodal (text + image + video + audio), 256K context, native function calling, configurable "thinking mode," and — finally — a clean Apache 2.0 license.

In this post we're going to:

  1. Understand what Gemma 4 actually is, with an architecture diagram
  2. Get it running on your laptop with Ollama in about 5 minutes
  3. Chat with it from the terminal
  4. Send it an image and ask questions about it
  5. Turn on thinking mode for harder problems
  6. Call it from a Python script like a real API
  7. Build a small project that glues it all together

No GPU rental, no API keys, no telemetry. Let's go.

Heads up: This guide assumes zero ML background. If you can install software and run a terminal command, you can do this.


What is Gemma 4?

Gemma is Google DeepMind's family of open-weight language models. "Open-weight" means the actual neural network weights — the giant matrices of numbers that make the model work — are freely downloadable. You can run them, modify them, fine-tune them, ship them in your product.

Gemma 4 brings several big changes over Gemma 3:

  • Apache 2.0 license. Earlier Gemma releases used a custom license with a Prohibited Use Policy that made some enterprise legal teams nervous. Gemma 4 is plain Apache 2.0 — unlimited commercial use, no MAU caps, no special permissions. This alone is a big deal for production deployments.
  • Mixture-of-Experts. A new 26B MoE variant activates only ~4B parameters per token, giving you 13B-class quality at 4B-class cost.
  • Thinking mode. A configurable reasoning mode where the model thinks step-by-step before answering. Toggle it on for hard problems, off for fast chat.
  • Native function calling. Built-in support for structured tool use — write an agent without needing prompt engineering hacks.
  • More modalities. Image, video frames, and (on the smaller E2B/E4B models) native audio input. Native system prompt support too.
  • Bigger context. 128K on the small models, 256K on the larger ones.

Model sizes at a glance

Model Disk (Ollama) Active params Total params Multimodal Context Best for
E2B ~7.2 GB ~2B ~2.3B text + image + audio 128K Phones, edge devices, browser
E4B ~9.6 GB ~4B ~4.5B text + image + audio 128K Most laptops — the sweet spot
26B A4B (MoE) ~18 GB ~4B 26B text + image 256K Consumer GPUs, agentic workloads
31B Dense ~20 GB 31B 31B text + image 256K Workstations, highest-quality answers

Two naming notes worth understanding:

  • E2B / E4B. The "E" stands for Effective parameters. These are dense edge-first models that use a trick called Per-Layer Embeddings (PLE — more on this below) to do more with fewer active parameters.
  • 26B A4B. This is the Mixture-of-Experts model. 26B parameters total, but only ~4B "activate" per forward pass. Latency and cost behave like a 4B model; quality is closer to a 13B dense model. Caveat: you still need to load all 26B into memory.

For most readers on a laptop, E4B is the right starting point. It runs comfortably on a 16 GB Mac or any modern dev machine.

Gemma 4 vs the rest of the open-model zoo (May 2026)

Model Sizes Multimodal Context License
Gemma 4 E2B / E4B / 26B MoE / 31B text + image + video + audio (small) 128K / 256K Apache 2.0
Llama 4 various text + image 128K+ Llama community license
Qwen 3.5 various text + image 128K+ Apache 2.0
DeepSeek V4 Flash MoE text 128K MIT

Gemma 4's pitch: the only family that spans phones to servers under Apache 2.0, with multimodal and audio in the same release.


The architecture (in plain English)

You don't need this section to use Gemma 4 — feel free to skip to the install steps. But if you've ever wondered what's actually happening when a multimodal model "sees" and "hears," here it is.
gemma4 description

A few pieces worth understanding:

  • Three input paths. Text goes through a SentencePiece tokenizer (shared with Gemini). Images go through a vision encoder that handles variable aspect ratios and resolutions natively (no more square-only inputs like Gemma 3). On the E2B and E4B models, audio goes through a USM-style conformer encoder borrowed from Gemma 3n. All three paths produce tokens that get interleaved in a single stream — so you can freely mix text, images, and audio in any order in one prompt.
  • Alternating local/global attention. Most layers only look at a sliding window of recent tokens (cheap). A subset of layers attend to the full context (expensive but rare). This is the standard trick for keeping the KV cache from blowing up at 256K context.
  • Per-Layer Embeddings (PLE)the small-model secret. In a normal transformer, each token gets one embedding vector at input and that's all the residual stream has to work with. PLE adds a parallel pathway: for each token, every layer gets its own small conditioning vector from a lookup table. The embedding tables are large (lots of memory) but the "active" parameters per token stay small — that's why a 4-billion-active-parameter E4B can punch above its weight.
  • Mixture-of-Experts (26B A4B). The MoE layer has multiple "expert" feed-forward networks. A small router picks 2 of 8 (or similar) for each token. Total params = 26B (all loaded), active params per token = ~4B (only those fire). Pareto-optimal for quality-per-FLOP.
  • Thinking mode. When you include the special <|think|> token at the start of the system prompt, the model emits internal reasoning between <|channel>thought\n...<channel|> markers before the final answer. Disable it for fast chat; enable it for math, code, multi-step reasoning.

That's most of what's worth knowing. Now let's actually run it.


Step 1: Install Ollama

There are a few ways to run Gemma 4 locally, but the easiest by a mile is Ollama. Think of it as "Docker for LLMs" — it handles downloading the model, managing memory, GPU acceleration, and exposing a local API. You don't have to think about CUDA versions or PyTorch.

Install it:

  curl -fsSL https://ollama.com/install.sh | sh

Enter fullscreen mode Exit fullscreen mode

Verify:

ollama --version

Enter fullscreen mode Exit fullscreen mode

You should see a version number. Gemma 4 requires Ollama v0.20.0 or later — if you're on an older version, update first.


Step 2: Pull a Gemma 4 model

Download the default (E4B, ~9.6 GB):

ollama pull gemma4

Enter fullscreen mode Exit fullscreen mode

This downloads about 9.6 GB. Grab a coffee. ☕

Other sizes if you want them:

ollama pull gemma4:e2b   # ~7.2 GB — smallest, for low-RAM machines
ollama pull gemma4:e4b   # ~9.6 GB — the default; same as `gemma4`
ollama pull gemma4:26b   # ~18 GB  — the MoE; 256K context
ollama pull gemma4:31b   # ~20 GB  — biggest dense model

Enter fullscreen mode Exit fullscreen mode

Hardware reality check: On Apple Silicon, 16 GB unified memory handles E4B comfortably. NVIDIA users need the model to fit entirely in VRAM for GPU-accelerated inference. The 26B model fits on 24 GB but leaves very little headroom — treat it as the ceiling, not the target.

List what you've got:

ollama list

Enter fullscreen mode Exit fullscreen mode


Step 3: Chat with it in the terminal

Easiest possible test:

ollama run gemma4

Enter fullscreen mode Exit fullscreen mode

You'll get an interactive prompt:

>>> Explain what a hash map is, like I'm a junior dev.

Enter fullscreen mode Exit fullscreen mode

Hit enter and watch it stream a response. To exit, type /bye.

That's it. You're running a state-of-the-art LLM locally with zero cloud dependency. Try:

  • "Write a Python function that finds duplicates in a list, with three different approaches and their tradeoffs."
  • "What's the difference between TCP and UDP? Use an analogy."
  • "Translate 'Where is the nearest train station?' into Japanese, Spanish, and Hindi."

Step 4: Send it an image

Gemma 4 can see. Drop any image file in your current directory, then:

ollama run gemma4
>>> Describe what's in this image: ./screenshot.png

Enter fullscreen mode Exit fullscreen mode

Ollama loads the image, sends it through the vision encoder, and the model answers. Unlike Gemma 3 (which resized everything to 896×896), Gemma 4 handles variable aspect ratios and resolutions natively — so tall screenshots, wide diagrams, and high-res photos all work without manual cropping.

Try:

  • "What error is shown in this screenshot?" (paste a stack trace)
  • "What's the bounding box for the 'submit' button in this UI?" (Gemma 4 will answer in JSON — natively!)
  • "Read the handwriting in this note and transcribe it."

Step 5: Turn on thinking mode

For harder problems — multi-step math, complex code, logic puzzles — turn on thinking mode. Include the <|think|> token at the very start of your system prompt:

ollama run gemma4
>>> /set system "<|think|>You are a careful, methodical assistant."
>>> Three friends split a $73.42 dinner bill. Alice had a $12 appetizer, Bob had a $9 drink. The rest is shared. What does everyone pay?

Enter fullscreen mode Exit fullscreen mode

The model will emit its reasoning in a <|channel>thought\n...<channel|> block before the final answer. For fast chat, leave the token out and the model answers directly.

🧠 When to use it: Code generation, math, multi-hop reasoning, agentic planning — yes. Single-turn factual questions, summarization, translation — no, it just adds latency.


Step 6: Call Gemma 4 from Python

A chat prompt is nice, but you're a developer — you want to call this thing from code. When Ollama is running, it exposes a local REST API on http://localhost:11434. There's also an official Python client.

Install it:

pip install ollama

Enter fullscreen mode Exit fullscreen mode

Basic chat

import ollama

response = ollama.chat(
    model="gemma4",
    messages=[
        {"role": "system", "content": "You are a senior code reviewer. Be concise and direct."},
        {"role": "user",   "content": "Review this code:\n\ndef add(a, b):\n    return a+b"},
    ],
)

print(response["message"]["content"])

Enter fullscreen mode Exit fullscreen mode

Streaming responses (ChatGPT-style)

import ollama

stream = ollama.chat(
    model="gemma4",
    messages=[{"role": "user", "content": "Write a haiku about debugging."}],
    stream=True,
)

for chunk in stream:
    print(chunk["message"]["content"], end="", flush=True)

Enter fullscreen mode Exit fullscreen mode

Sending an image

import ollama

response = ollama.chat(
    model="gemma4",
    messages=[{
        "role": "user",
        "content": "What's in this image?",
        "images": ["./my_photo.jpg"],
    }],
)

print(response["message"]["content"])

Enter fullscreen mode Exit fullscreen mode

Thinking mode + function calling (the agentic combo)

This is where Gemma 4 actually starts feeling like a "real" agent. You declare your tools as JSON schemas, the model decides when to call them, and you execute the call and pass results back. No prompt engineering hacks needed.

import ollama

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {
                "city": {"type": "string", "description": "City name, e.g. 'Tokyo'"},
            },
            "required": ["city"],
        },
    },
}]

def get_weather(city: str) -> str:
    # Pretend this hits a real API.
    return f"{city}: 22°C, partly cloudy"

response = ollama.chat(
    model="gemma4",
    messages=[
        {"role": "system", "content": "<|think|>You are a helpful weather assistant."},
        {"role": "user",   "content": "Should I bring an umbrella in Tokyo today?"},
    ],
    tools=tools,
)

# If the model wants to call a tool, execute it and feed the result back:
for tool_call in response["message"].get("tool_calls", []):
    name = tool_call["function"]["name"]
    args = tool_call["function"]["arguments"]
    if name == "get_weather":
        result = get_weather(**args)
        # Send result back for the model to finalize its answer
        followup = ollama.chat(
            model="gemma4",
            messages=[
                {"role": "user", "content": "Should I bring an umbrella in Tokyo today?"},
                response["message"],
                {"role": "tool", "content": result, "name": name},
            ],
        )
        print(followup["message"]["content"])

Enter fullscreen mode Exit fullscreen mode

Raw HTTP (no Python client needed)

For any other language:

curl http://localhost:11434/api/chat -d '{
  "model": "gemma4",
  "messages": [{"role": "user", "content": "Hello!"}],
  "stream": false
}'

Enter fullscreen mode Exit fullscreen mode

Same JSON shape works from Node, Go, Rust, your shell — anything that can make an HTTP request.


A small project: folder-watching image describer

Here's a useful ~30-line script. It watches a folder, and any new image dropped in gets automatically described by Gemma 4. Great for accessibility tools, content moderation prototypes, or just learning.

import os, time
import ollama

WATCH_DIR = "./inbox"
os.makedirs(WATCH_DIR, exist_ok=True)
SEEN = set(os.listdir(WATCH_DIR))

print(f"📁 Watching {WATCH_DIR}/ — drop an image in to describe it.")
print("   (Ctrl+C to stop)\n")

IMAGE_EXTS = (".png", ".jpg", ".jpeg", ".webp", ".gif")

try:
    while True:
        current = set(os.listdir(WATCH_DIR))
        new_files = sorted(current - SEEN)

        for filename in new_files:
            if not filename.lower().endswith(IMAGE_EXTS):
                continue

            path = os.path.join(WATCH_DIR, filename)
            print(f"📸 New image: {filename}")

            response = ollama.chat(
                model="gemma4",
                messages=[{
                    "role": "user",
                    "content": (
                        "Describe this image in 2-3 sentences. "
                        "Mention any visible text. Be specific."
                    ),
                    "images": [path],
                }],
            )

            print(f"{response['message']['content']}\n")

        SEEN = current
        time.sleep(2)
except KeyboardInterrupt:
    print("\n Stopped.")

Enter fullscreen mode Exit fullscreen mode

Run it, drag images into the inbox/ folder, and watch descriptions appear. That's a real, useful, completely local AI tool — written in 30 lines.


Things to know before shipping anything serious

A few honest caveats:

Caveat Why it matters
Hallucination Local models still confidently make things up. Don't trust factual claims without verification. Thinking mode reduces this for reasoning tasks but doesn't eliminate it.
CPU latency Expect 1–3 tokens/sec on a CPU-only laptop with E4B. A GPU gives 3–10× speedup.
Context costs RAM 256K context is real, but actually filling it eats memory. Most use cases need <16K tokens.
MoE memory The 26B MoE runs fast (only 4B active per token), but you still need to load all 26B into RAM. Don't confuse active params with memory footprint.
Audio is small-model only E2B/E4B have native audio input. The 26B and 31B models do not.
Apache 2.0 ≠ no responsibilities The license is permissive, but you're still on the hook for safety, bias, and compliance in whatever you ship.

📚 References & further reading