惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
美团技术团队
Last Week in AI
Last Week in AI
WordPress大学
WordPress大学
博客园 - 三生石上(FineUI控件)
博客园 - 聂微东
雷峰网
雷峰网
阮一峰的网络日志
阮一峰的网络日志
博客园 - 叶小钗
IT之家
IT之家
Google DeepMind News
Google DeepMind News
D
Docker
J
Java Code Geeks
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - 【当耐特】
V
V2EX
Hugging Face - Blog
Hugging Face - Blog
博客园 - Franky
月光博客
月光博客
宝玉的分享
宝玉的分享
酷 壳 – CoolShell
酷 壳 – CoolShell
aimingoo的专栏
aimingoo的专栏
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
Build a Private Voice Assistant with Whisper, Ollama, and...
EveryLocalAI · 2026-06-15 · via DEV Community

EveryLocalAI

Have you ever wanted your own Jarvis? A voice assistant that listens, thinks, and speaks back - all running privately on your own hardware? Here's how to build one with Whisper.cpp, Ollama, and Kokoro TTS.

No cloud, no wake-word fees, no data leaving your machine.

Prerequisites

  • Hardware: Any modern computer with a microphone
  • Software: Python 3.10+, Ollama installed
  • Time: ~30 minutes setup

Installation

1. Install Ollama and Pull a Model

ollama pull qwen3:14b

2. Install Whisper.cpp

git clone https://github.com/ggerganov/whisper.cpp.git
cd whisper.cpp
cmake -B build && cmake --build build --config Release
bash models/download-ggml-model.sh medium

3. Install Kokoro TTS

pip install kokoro pyaudio requests

Wiring It All Together

Save this as voice_assistant.py:

import subprocess
import tempfile
import wave
import pyaudio
import requests
from kokoro import KPipeline

OLLAMA_URL = "http://localhost:11434/api/generate"
MODEL = "qwen3:14b"
WHISPER_BIN = "./whisper.cpp/build/bin/whisper-cli"
WHISPER_MODEL = "./whisper.cpp/models/ggml-medium.bin"
tts_pipeline = KPipeline(lang_code='a')

def record_audio(duration=5, sample_rate=16000):
    p = pyaudio.PyAudio()
    stream = p.open(format=pyaudio.paInt16, channels=1,
                    rate=sample_rate, input=True,
                    frames_per_buffer=1024)
    frames = [stream.read(1024) for _ in range(int(sample_rate / 1024 * duration))]
    stream.close(); p.terminate()
    with tempfile.NamedTemporaryFile(suffix='.wav', delete=False) as f:
        wf = wave.open(f, 'wb')
        wf.setnchannels(1); wf.setsampwidth(2)
        wf.setframerate(sample_rate)
        wf.writeframes(b''.join(frames))
        return f.name

def transcribe(audio_file):
    result = subprocess.run([WHISPER_BIN, '-m', WHISPER_MODEL, '-f', audio_file],
                          capture_output=True, text=True)
    return result.stdout.strip()

def ask_llm(prompt):
    r = requests.post(OLLAMA_URL, json={"model": MODEL, "prompt": prompt, "stream": False})
    return r.json()["response"]

def speak(text):
    for result in tts_pipeline(text):
        with tempfile.NamedTemporaryFile(suffix='.wav', delete=False) as f:
            f.write(result.audio)
        subprocess.run(['ffplay', '-nodisp', '-autoexit', f.name])

# Run it
print("Listening...")
audio_file = record_audio(5)
text = transcribe(audio_file)
print(f"You: {text}")
response = ask_llm(text)
print(f"AI: {response}")
speak(response)

Run it:

python voice_assistant.py

Speak into your mic. Wait 5 seconds. Hear the AI respond.

Performance

  • Whisper medium on CPU: transcribes in 2-4 seconds
  • Qwen3 14B on RTX 3060: responds in 3-5 seconds
  • Kokoro TTS on CPU: speaks in real-time (< 1 second latency)
  • Total round-trip: ~10 seconds on modest hardware

For faster responses, use Whisper tiny or a smaller LLM like Llama 3.1 8B.


Originally published on everylocalai.com