惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

AWS News Blog
AWS News Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - Franky
大猫的无限游戏
大猫的无限游戏
Engineering at Meta
Engineering at Meta
T
Tailwind CSS Blog
T
The Blog of Author Tim Ferriss
L
LangChain Blog
Vercel News
Vercel News
N
Netflix TechBlog - Medium
Hacker News - Newest:
Hacker News - Newest: "LLM"
Spread Privacy
Spread Privacy
小众软件
小众软件
H
Help Net Security
The Last Watchdog
The Last Watchdog
Forbes - Security
Forbes - Security
WordPress大学
WordPress大学
Know Your Adversary
Know Your Adversary
Recent Commits to openclaw:main
Recent Commits to openclaw:main
H
Heimdal Security Blog
GbyAI
GbyAI
P
Privacy International News Feed
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
The Cloudflare Blog
爱范儿
爱范儿
V
V2EX
The Register - Security
The Register - Security
B
Blog RSS Feed
Apple Machine Learning Research
Apple Machine Learning Research
O
OpenAI News
Cisco Talos Blog
Cisco Talos Blog
Cloudbric
Cloudbric
Security Archives - TechRepublic
Security Archives - TechRepublic
S
Secure Thoughts
L
LINUX DO - 最新话题
Recorded Future
Recorded Future
P
Proofpoint News Feed
PCI Perspectives
PCI Perspectives
Hugging Face - Blog
Hugging Face - Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
T
Tor Project blog
Latest news
Latest news
Project Zero
Project Zero
月光博客
月光博客
F
Fortinet All Blogs
A
Arctic Wolf
博客园 - 【当耐特】
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Y
Y Combinator Blog
N
News and Events Feed by Topic

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
How to Build a Private Offline Voice Assistant with Gemma 4 12B: A Complete Local Setup Guide
Christopher Hoeben · 2026-06-17 · via DEV Community

How to Build a Private Offline Voice Assistant with Gemma 4 12B: A Complete Local Setup Guide

A developer’s guide to running Google’s 11.95B-parameter multimodal model with local STT/TTS on a 16 GB laptop under Apache 2.0.

TL;DR: Download Gemma 4 12B (~6.7 GB at 4-bit) into a local runtime such as Google AI Edge Gallery, pair it with a local STT/TTS stack, and expose a local endpoint. The 11.95B-parameter model fits on a 16 GB laptop, runs offline under Apache 2.0, and keeps all voice data on-device.

Check Hardware Constraints and the 30-Second Audio Limit

Before downloading the model, verify your machine has at least 16 GB of RAM and plan your voice pipeline around the model’s strict 30-second audio ceiling. At 4-bit quantization, Gemma 4 12B’s 11.95 billion parameters compress to roughly 6.7 GB. After loading the weights, the remaining ~9 GB must cover the operating system, the inference framework overhead, and any local audio capture or STT services. If you are running other local models or Home Assistant addons concurrently, budget even more conservatively.

Check available memory before launching the stack:

free -h

Aim to have at least 14 GB free at idle; anything less risks swapping during inference.

The model enforces a hard 30-second audio limit. Exceeding it will cause inference to fail or truncate, so your client must enforce a maximum recording duration. A common approach is to chunk incoming streams or fall back to text input for complex multi-part commands. Split existing recordings at the boundary with ffmpeg:

ffmpeg -i input.wav -f segment -segment_time 30 -c copy chunk_%03d.wav

This produces 30-second WAV files that stay within the limit. Feed each chunk separately, or switch to a text fallback when a user’s utterance exceeds one segment.

Install a Local Inference Runtime

To run Gemma 4 12B without external API calls, install a local inference runtime first. The Google AI Edge Gallery is one supported deployment option for both phones and laptops, and releases are delivered as standard OS-specific packages: a Windows .exe installer, a macOS zip bundle, and a Linux package.

Because this runtime serves as the execution backend for your voice pipeline, completing the installation before downloading model weights avoids path and permission errors during setup.

On macOS, download the zip archive, extract it, and drag the resulting application into your system Applications folder. Standard user permissions are sufficient for most local inference workloads when the app resides in the Applications directory. If you prefer the command line, a common approach is to locate the downloaded bundle and move it in one step:

cd ~/Downloads && unzip *.zip && mv *.app /Applications/

On Windows, launch the downloaded .exe installer and proceed through the setup prompts until the wizard finishes. A per-user install is usually adequate, with administrator elevation required only if you explicitly choose a system-wide program directory. You can also trigger the installer non-interactively once it is saved to your Downloads folder:

$exe = Get-ChildItem "$env:USERPROFILE\Downloads" -Filter *.exe | Select-Object -First 1
Start-Process -FilePath $exe.FullName -Wait

Linux users should install the provided package using the distribution’s native package manager; because formats vary by release, refer to the supplied readme for the exact dpkg, rpm, or AppImage command. After installation completes on any platform, open the runtime and verify that the local inference engine is active before pulling the Gemma 4 12B weights. Keeping this layer fully offline ensures voice data never leaves the device.

Load Gemma 4 12B at 4-Bit Quantization

Loading Gemma 4 12B at 4-bit quantization reduces its memory footprint to roughly 6.7 GB, letting the entire model stay resident in RAM on a 16 GB laptop. Select the 4-bit option in your local inference UI or configuration file immediately after importing the model weights.

At 11.95 billion parameters, the full-precision weights would exceed typical consumer memory limits, but 4-bit compression brings private, on-device deployment within reach. In tools like Google AI Edge Gallery, select the 4-bit quantization profile during the model-import step. Because Gemma 4 is encoder-free and processes audio in a single pass, keeping the entire model in RAM is especially critical—any disk access during inference multiplies latency for multimodal inputs. After initialization, verify the process is locked in physical memory and not swapping before you attach speech-to-text or text-to-speech services; even occasional paging destroys the low latency required for conversational voice interaction. On Linux, confirm swap usage is zero with:

grep VmSwap /proc/$(pgrep -f gemma)/status

On macOS, monitor memory pressure while the model loads:

memory_pressure && vm_stat 1

If you see swap growth or pressure warnings, reduce the context window or close other applications until the process stabilizes entirely in RAM. A fully resident model avoids the round-trip disk delay that would otherwise make real-time assistant responses unusable. Treat this verification as a mandatory gate: only after confirming stable, swap-free residency should you layer on the speech pipeline.

Wire Up Local STT and TTS Components

Because current local assistant frameworks still require separate speech and model components, you must bridge a local STT engine and a local TTS service to Gemma 4 12B; the STT text feeds into Gemma’s text context, and the generated reply routes to the TTS service, even though the model natively ingests audio in a single pass. Until front-end conversation agents expose that native audio path, a text pipeline is the only workable architecture, and it conveniently sidesteps Gemma’s hard 30-second audio limit. Splitting the pipeline this way also lets you upgrade either speech component independently of the quantized model.

For the STT layer, a local ONNX/Parakeet model can deliver subsecond transcription latency. Load the ONNX graph and run inference on the captured waveform:

import numpy as np, onnxruntime as ort
session = ort.InferenceSession("parakeet.onnx")
inputs = {session.get_inputs()[0].name: waveform}
text = session.run(None, inputs)[0]

Pass the resulting transcript to your local Gemma endpoint. A common pattern is to POST the text to a local inference server and stream back the response:

import requests, json
r = requests.post("http://localhost:11434/api/generate",
    json={"model": "gemma4:12b", "prompt": text, "stream": False})
reply = r.json()["response"]

Finally, send the reply to a local TTS service. A typical setup pushes the synthesized string to an on-device Piper or similar HTTP endpoint and writes the returned audio to your speaker queue:

audio = requests.post("http://localhost:5000/tts",
    json={"text": reply}, stream=True)
# playback(audio.content)

This keeps the full loop offline: the STT model runs locally, Gemma runs locally, and the TTS service runs locally.

Build the Voice Command Loop

Build the voice command loop by capturing microphone audio, sending it to a local STT service, forwarding the resulting transcript to your local Gemma 4 inference endpoint, and passing the generated reply to a local TTS engine for immediate playback. You must enforce the model’s hard 30-second audio cap by halting microphone capture before that limit; anything longer will exceed the model’s single-pass audio window and trigger truncation or rejection.

A common approach is to record raw PCM audio with sounddevice, flush it to a 16 kHz mono WAV file, and POST it to a local Whisper-compatible STT server listening on port 9000. Once the STT returns the transcript, construct a concise prompt formatted as a direct action command in the style of the Voice Edit pattern—phrasing like “Restructure these notes into an executive summary” or “Translate this into Hindi”—and POST that payload to your local Gemma 4 inference endpoint running under Ollama or llama.cpp on localhost. Avoid conversational preamble; the model executes faster when the instruction is explicit and scoped to a single action.

import sounddevice as sd, requests, wave
frames = sd.rec(int(30 * 16000), samplerate=16000,
                channels=1, dtype='int16')
sd.wait()
with wave.open("cmd.wav", "wb") as f:
    f.setnchannels(1); f.setsampwidth(2); f.setframerate(16000)
    f.writeframes(frames.tobytes())

Submit the recorded file to the STT layer and retrieve the text:

curl -X POST http://localhost:9000/v1/audio/transcriptions \
  -F file="@cmd.wav" -F model="whisper-base"

Forward the transcript to the local Gemma 4 API as an explicit agentic instruction:

r = requests.post("http://localhost:11434/api/generate",
    json={"model": "gemma4:12b",
          "prompt": f"Restructure into an executive summary: {transcript}"})
reply = r.json()["response"]

Finally, push the model’s text output to your local TTS service—such as Piper or Coqui—and play the synthesized audio stream through your default sound device. Keep the loop strictly sequential: record, transcribe, infer, speak, then return to listening so only one audio stream is active at any moment and the pipeline stays synchronized.

Keep the Stack Offline and Private

Keep the stack offline by binding the inference client to a local-only endpoint and dropping all outbound traffic for the process at the system firewall. Because Gemma 4 12B ships under Apache 2.0 with no commercial restrictions and inference runs entirely on-device, no audio or text ever leaves your machine, and all sensitive multimodal data remains on the hardware that owns it.

At the client layer, disable remote base URLs and any automatic fallback to hosted APIs. A common approach is to initialize the SDK against a local inference server so every request stays on the loopback interface and never attempts an external resolver:

from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:11434/v1",
    api_key="not-needed-local"
)

For a hard offline guarantee, add a firewall rule that denies the voice assistant process any outbound route:

sudo iptables -A OUTPUT -m owner --uid-owner assistant -j DROP

The 11.95 billion parameter weights compress to roughly 6.7 GB at 4-bit quantization, so the full audio-to-text pipeline executes in local RAM without cloud encoders or API dependencies. The hard 30-second audio limit also bounds each inference batch to what fits on-device. After starting the assistant, verify isolation by capturing packets during a voice query: if traffic leaves the loopback adapter, the stack is not truly offline.

FAQ

Can I send raw audio directly into Gemma 4 12B and skip STT?

The encoder-free architecture reads audio in a single pass, but most local assistant platforms and conversation agents still require separate speech components today. Until those frameworks natively stream raw audio to the model, you should keep a local STT layer in the pipeline.

Will this run on a laptop with only 16 GB of shared RAM?

Yes. Google DeepMind specifies that Gemma 4 12B runs on a 16 GB laptop. At 4-bit quantization the weights occupy roughly 6.7 GB, leaving headroom for the OS and your STT/TTS services if you manage memory carefully.

What happens if my voice command is longer than 30 seconds?

Gemma 4 12B has a hard 30-second audio limit. A common approach is to chunk long utterances or switch to a text-based prompt once you exceed that boundary.

Do I need an internet connection after the initial setup?

No. The model is Apache 2.0 and open-weight, so after you download the quantized weights and install the local runtime, the entire voice assistant operates offline. No API keys or cloud endpoints are required.

Is there any legal risk using this for a commercial product?

No. The weights ship under Apache 2.0 with no commercial restrictions, removing legal friction for on-device deployments.

References for further reading

Sources consulted while researching this guide, included so you can verify the details and go deeper. Listing them is not a claim that every line was independently fact-checked.


I packaged the setup above into a ready-to-use kit — **Gemma 4 12B Local Multimodal Build Kit (13 Items)* — for anyone who'd rather copy-paste than wire it from scratch: https://unfairhq.gumroad.com/l/nkylsz.*