惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
L
LINUX DO - 最新话题
Stack Overflow Blog
Stack Overflow Blog
月光博客
月光博客
雷峰网
雷峰网
Apple Machine Learning Research
Apple Machine Learning Research
V
Visual Studio Blog
Attack and Defense Labs
Attack and Defense Labs
O
OpenAI News
The GitHub Blog
The GitHub Blog
A
About on SuperTechFans
B
Blog RSS Feed
H
Help Net Security
量子位
小众软件
小众软件
SecWiki News
SecWiki News
N
Netflix TechBlog - Medium
TaoSecurity Blog
TaoSecurity Blog
美团技术团队
博客园 - 司徒正美
Hacker News - Newest:
Hacker News - Newest: "LLM"
Recent Commits to openclaw:main
Recent Commits to openclaw:main
The Cloudflare Blog
N
News and Events Feed by Topic
C
Cybersecurity and Infrastructure Security Agency CISA
The Last Watchdog
The Last Watchdog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Scott Helme
Scott Helme
T
The Exploit Database - CXSecurity.com
K
Kaspersky official blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
Threat Research - Cisco Blogs
C
CERT Recently Published Vulnerability Notes
Application and Cybersecurity Blog
Application and Cybersecurity Blog
U
Unit 42
Google DeepMind News
Google DeepMind News
J
Java Code Geeks
Schneier on Security
Schneier on Security
G
Google Developers Blog
Forbes - Security
Forbes - Security
C
CXSECURITY Database RSS Feed - CXSecurity.com
Y
Y Combinator Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
P
Palo Alto Networks Blog
A
Arctic Wolf
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
The Hacker News
The Hacker News
B
Blog
D
DataBreaches.Net
Simon Willison's Weblog
Simon Willison's Weblog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Voice and Camera Input in React: Speech Recognition, Media Devices, and Permissions
reactuse.com · 2026-05-07 · via DEV Community

Voice and Camera Input in React: Speech Recognition, Media Devices, and Permissions

Voice and camera are the two senses that turn a static web app into something that feels alive. A search bar that you can talk to. A note-taking app that transcribes you in real time. A meeting tool that lets you pick which webcam to use. A walkie-talkie that talks when you hold a key. None of these are exotic anymore -- the browser has had the APIs for years -- but every one of them lives behind a gauntlet of permission prompts, vendor prefixes, and lifecycle quirks that make them painful to integrate into a React component.

This post walks through four browser capabilities for voice and camera input: live speech recognition with interim results, enumerating the user's cameras and microphones, querying permissions in a way that survives revocation, and using the Shift key as a push-to-talk modifier. As always, we will start with the manual implementation so you understand the plumbing, then swap it out for a focused hook from ReactUse. At the end, we will combine all four into a complete voice search component with a device picker, a permission gate, and a push-to-talk hold-to-record interaction.

1. Live Speech Recognition

The Manual Way

The Web Speech API is one of the older browser APIs that never really got standardized -- Chrome ships it as webkitSpeechRecognition, and the unprefixed SpeechRecognition is still missing in most engines. The minimal viable React wrapper looks like this:

function ManualSpeechRecognition() {
  const [transcript, setTranscript] = useState("");
  const [listening, setListening] = useState(false);
  const recognitionRef = useRef<any>(null);

  useEffect(() => {
    const SR =
      (window as any).SpeechRecognition ||
      (window as any).webkitSpeechRecognition;
    if (!SR) return;
    const recognition = new SR();
    recognition.continuous = true;
    recognition.interimResults = true;
    recognition.lang = "en-US";
    recognition.onresult = (event: any) => {
      const result = event.results[event.resultIndex];
      setTranscript(result[0].transcript);
    };
    recognition.onend = () => setListening(false);
    recognitionRef.current = recognition;
    return () => recognition.abort();
  }, []);

  const start = () => {
    recognitionRef.current?.start();
    setListening(true);
  };
  const stop = () => {
    recognitionRef.current?.stop();
    setListening(false);
  };

  return (
    <div>
      <button onClick={listening ? stop : start}>
        {listening ? "Stop" : "Start"} listening
      </button>
      <p>{transcript}</p>
    </div>
  );
}

Enter fullscreen mode Exit fullscreen mode

This works, but it ignores the rough edges. There is no isFinal distinction, so the UI cannot tell when the user pauses (the difference between "interim" and "final" results is what makes voice UIs feel responsive). There is no error handling -- if the user denies microphone permission or there is no network, the transcript just silently never updates. There is no language negotiation. And SR typing is awful, because TypeScript does not ship types for webkitSpeechRecognition.

The ReactUse Way: useSpeechRecognition

useSpeechRecognition returns a clean object with the right primitives:

import { useSpeechRecognition } from "@reactuses/core";

function VoiceNote() {
  const { isSupported, isListening, isFinal, result, error, start, stop } =
    useSpeechRecognition({
      lang: "en-US",
      interimResults: true,
      continuous: true,
    });

  if (!isSupported) {
    return <p>Speech recognition is not supported in this browser.</p>;
  }

  return (
    <div>
      <button onClick={isListening ? stop : start}>
        {isListening ? "Stop" : "Start"} dictation
      </button>
      <p
        style={{
          fontStyle: isFinal ? "normal" : "italic",
          color: isFinal ? "#0f172a" : "#64748b",
        }}
      >
        {result || "Say something..."}
      </p>
      {error && <p style={{ color: "#ef4444" }}>Error: {error.error}</p>}
    </div>
  );
}

Enter fullscreen mode Exit fullscreen mode

Things you get for free:

  1. isFinal -- the hook tracks whether the current result is the speech engine's tentative guess (italicized in the example) or the locked-in transcript. This is the single biggest UX improvement over the naive version.
  2. error object -- when permission is denied, the network drops, or the engine fails, you get a typed error you can show the user instead of silently freezing.
  3. Hot config. start({ lang: "fr-FR" }) switches languages mid-session without requiring you to tear down the recognizer.
  4. Cleanup on unmount. The hook calls abort() automatically, so navigating away never leaves the microphone hot.

The most powerful pattern is binding the result to a search input as the user speaks. Because the hook re-renders with each interim result, you can drive a live search query directly from spoken input and let the user see results as they talk.

2. Enumerating Cameras and Microphones

The Manual Way

Listing the user's audio and video devices requires navigator.mediaDevices.enumerateDevices(). There is a catch: until the user has granted permission to some device, the labels come back empty -- you get a list of deviceIds but no label like "FaceTime HD Camera". To get labels, you have to first call getUserMedia to trigger the permission prompt, then re-enumerate.

function ManualDeviceList() {
  const [devices, setDevices] = useState<MediaDeviceInfo[]>([]);

  useEffect(() => {
    let mounted = true;
    const refresh = async () => {
      try {
        // Trigger permission so labels are populated
        const stream = await navigator.mediaDevices.getUserMedia({
          audio: true,
          video: true,
        });
        stream.getTracks().forEach((t) => t.stop());
        const list = await navigator.mediaDevices.enumerateDevices();
        if (mounted) setDevices(list);
      } catch (e) {
        console.error(e);
      }
    };
    refresh();
    navigator.mediaDevices.addEventListener("devicechange", refresh);
    return () => {
      mounted = false;
      navigator.mediaDevices.removeEventListener("devicechange", refresh);
    };
  }, []);

  return (
    <ul>
      {devices.map((d) => (
        <li key={d.deviceId}>
          {d.kind}: {d.label || "(label hidden)"}
        </li>
      ))}
    </ul>
  );
}

Enter fullscreen mode Exit fullscreen mode

That is the right shape, but you have to write the permission-trigger dance, the cleanup of the temporary stream, and the device-change listener every single time.

The ReactUse Way: useMediaDevices

useMediaDevices packages the whole thing:

import { useMediaDevices } from "@reactuses/core";

function CameraPicker({
  selected,
  onSelect,
}: {
  selected: string;
  onSelect: (id: string) => void;
}) {
  const [{ devices }, ensurePermissions] = useMediaDevices({
    requestPermissions: true,
    constraints: { video: true, audio: false },
  });

  const cameras = devices.filter((d) => d.kind === "videoinput");

  return (
    <div>
      <button onClick={() => ensurePermissions()}>Refresh devices</button>
      <select
        value={selected}
        onChange={(e) => onSelect(e.target.value)}
        style={{ marginLeft: 8 }}
      >
        {cameras.map((cam) => (
          <option key={cam.deviceId} value={cam.deviceId}>
            {cam.label || `Camera ${cam.deviceId.slice(0, 6)}`}
          </option>
        ))}
      </select>
    </div>
  );
}

Enter fullscreen mode Exit fullscreen mode

The hook handles three things you would otherwise write yourself:

  • Permission negotiation. Pass requestPermissions: true and the hook will trigger getUserMedia on mount with the constraints you specify, then immediately stop the temporary tracks so the camera light goes off.
  • Live device list. The hook listens for devicechange and re-enumerates automatically -- if a user plugs in a new microphone or unplugs their headphones, the list updates with no extra code.
  • Manual refresh. The returned ensurePermissions lets you re-prompt at any time, useful if the user denied permission once and you want a "Try again" button.

The constraints argument is forwarded directly to getUserMedia, so you can ask only for video (and skip the awkward "do you want microphone access" prompt) when you only need to pick a camera.

3. Querying Permissions Properly

The Manual Way

Checking whether the user has already granted (or denied) microphone or camera permission, without triggering a prompt, requires the Permissions API. It is well-supported but verbose:

function ManualMicPermission() {
  const [state, setState] = useState<PermissionState | "unknown">("unknown");

  useEffect(() => {
    let mounted = true;
    let status: PermissionStatus | null = null;
    (async () => {
      try {
        status = await navigator.permissions.query({
          name: "microphone" as PermissionName,
        });
        if (mounted) setState(status.state);
        status.onchange = () => mounted && status && setState(status.state);
      } catch {
        // Permissions API not available for this name
      }
    })();
    return () => {
      mounted = false;
      if (status) status.onchange = null;
    };
  }, []);

  return <p>Microphone permission: {state}</p>;
}

Enter fullscreen mode Exit fullscreen mode

Three things to notice. First, the API is callback-based via onchange, not React-friendly. Second, you have to feature-detect both the Permissions API itself and the specific name (some browsers do not support "microphone"). Third, the change listener has to be cleaned up explicitly, not via effect return.

The ReactUse Way: usePermission

usePermission reduces the entire dance to a single call:

import { usePermission } from "@reactuses/core";

function MicStatusBadge() {
  const state = usePermission("microphone");

  const color =
    state === "granted"
      ? "#10b981"
      : state === "denied"
      ? "#ef4444"
      : "#f59e0b";

  return (
    <span style={{ color, fontWeight: 600 }}>
      Microphone: {state || "unknown"}
    </span>
  );
}

Enter fullscreen mode Exit fullscreen mode

The state is a React-native string that updates whenever the underlying permission status changes -- including out-of-band, when the user goes into browser settings and revokes the permission, your component's state will flip to "denied" without any action from you.

You can pass a string like "microphone" or "camera", or a full PermissionDescriptor object for permissions like "push" that need extra fields. The shape is identical to navigator.permissions.query, just turned into a hook.

4. Push-To-Talk With useKeyModifier

The Manual Way

A push-to-talk button is harder than it looks. You want to detect when the user is holding a key (say, the Space bar or the Shift key), start recording while it is held, and stop the moment they release it. You also have to handle the case where the user holds the key, switches focus to another window, releases the key while your tab is hidden, and then returns -- otherwise the recorder is stuck on.

function ManualPushToTalk() {
  const [pressed, setPressed] = useState(false);

  useEffect(() => {
    const onDown = (e: KeyboardEvent) => {
      if (e.code === "Space") setPressed(true);
    };
    const onUp = (e: KeyboardEvent) => {
      if (e.code === "Space") setPressed(false);
    };
    const onBlur = () => setPressed(false);
    window.addEventListener("keydown", onDown);
    window.addEventListener("keyup", onUp);
    window.addEventListener("blur", onBlur);
    return () => {
      window.removeEventListener("keydown", onDown);
      window.removeEventListener("keyup", onUp);
      window.removeEventListener("blur", onBlur);
    };
  }, []);

  return <p>{pressed ? "Recording..." : "Hold space to talk"}</p>;
}

Enter fullscreen mode Exit fullscreen mode

This almost works. The bug: if the Space key auto-repeats while held (which it does on most operating systems), you will get a keydown followed by another keydown followed by an eventual keyup. You handled that. But if the user holds Shift instead and uses it as a modifier with another key combination, your manual tracking does not know about it.

The ReactUse Way: useKeyModifier

useKeyModifier exposes the OS-level modifier state (the same values you would get from event.getModifierState) as React state:

import { useKeyModifier } from "@reactuses/core";

function ShiftToRecord({ onTalkStart, onTalkEnd }: {
  onTalkStart: () => void;
  onTalkEnd: () => void;
}) {
  const shift = useKeyModifier("Shift");

  useEffect(() => {
    if (shift) onTalkStart();
    else onTalkEnd();
  }, [shift, onTalkStart, onTalkEnd]);

  return (
    <div
      style={{
        padding: 16,
        background: shift ? "#fef3c7" : "#f1f5f9",
        borderRadius: 8,
        textAlign: "center",
      }}
    >
      {shift ? "Recording (release Shift to stop)" : "Hold Shift to talk"}
    </div>
  );
}

Enter fullscreen mode Exit fullscreen mode

The benefits over the keydown/keyup version:

  • OS-aware. The hook reads getModifierState, which queries the actual modifier state from the OS. It survives auto-repeat, focus loss, and weird key combinations correctly.
  • Works for any modifier. Pass "Control", "Alt", "Meta", "CapsLock", "NumLock" -- anything the browser tracks as a modifier.
  • Initial value. Configure initial: true if you want the React state to start as true (uncommon but useful for debugging).

Putting It All Together: A Voice Search With Device Picker

We will combine all four hooks into a single voice-driven search component. The user can pick which microphone to use, sees a permission badge, holds Shift to start dictating, and watches the transcript update live as they speak. When they release Shift, the final transcript becomes the search query.

import { useEffect, useState } from "react";
import {
  useSpeechRecognition,
  useMediaDevices,
  usePermission,
  useKeyModifier,
} from "@reactuses/core";

function VoiceSearch() {
  const [selectedMic, setSelectedMic] = useState<string>("");
  const [query, setQuery] = useState("");

  const micPermission = usePermission("microphone");
  const [{ devices }, requestDevices] = useMediaDevices({
    requestPermissions: false,
    constraints: { audio: true, video: false },
  });

  const microphones = devices.filter((d) => d.kind === "audioinput");

  const {
    isSupported,
    isListening,
    isFinal,
    result,
    error,
    start,
    stop,
  } = useSpeechRecognition({
    lang: "en-US",
    interimResults: true,
    continuous: false,
  });

  const shiftDown = useKeyModifier("Shift");

  // Push-to-talk: start when Shift is pressed, stop when released
  useEffect(() => {
    if (!isSupported || micPermission !== "granted") return;
    if (shiftDown) {
      start();
    } else if (isListening) {
      stop();
    }
  }, [shiftDown, isSupported, micPermission, start, stop, isListening]);

  // When recognition finalizes, commit the result to the query
  useEffect(() => {
    if (isFinal && result) {
      setQuery(result);
    }
  }, [isFinal, result]);

  const permissionColor =
    micPermission === "granted"
      ? "#10b981"
      : micPermission === "denied"
      ? "#ef4444"
      : "#f59e0b";

  return (
    <div
      style={{
        maxWidth: 640,
        padding: 24,
        background: "#ffffff",
        borderRadius: 16,
        boxShadow: "0 4px 24px rgba(15, 23, 42, 0.06)",
        fontFamily: "system-ui, sans-serif",
      }}
    >
      <header
        style={{
          display: "flex",
          justifyContent: "space-between",
          alignItems: "center",
          marginBottom: 16,
        }}
      >
        <h2 style={{ margin: 0, fontSize: 18 }}>Voice Search</h2>
        <span style={{ color: permissionColor, fontSize: 13, fontWeight: 600 }}>
          ● Mic: {micPermission || "unknown"}
        </span>
      </header>

      {!isSupported && (
        <p style={{ color: "#64748b" }}>
          This browser does not support speech recognition. Try Chrome.
        </p>
      )}

      {isSupported && micPermission !== "granted" && (
        <button
          onClick={requestDevices}
          style={{
            width: "100%",
            padding: 12,
            background: "#3b82f6",
            color: "white",
            border: "none",
            borderRadius: 8,
            cursor: "pointer",
          }}
        >
          Grant microphone access
        </button>
      )}

      {isSupported && micPermission === "granted" && (
        <>
          <div style={{ display: "flex", gap: 12, marginBottom: 12 }}>
            <select
              value={selectedMic}
              onChange={(e) => setSelectedMic(e.target.value)}
              style={{
                flex: 1,
                padding: 8,
                borderRadius: 6,
                border: "1px solid #cbd5e1",
              }}
            >
              <option value="">Default microphone</option>
              {microphones.map((mic) => (
                <option key={mic.deviceId} value={mic.deviceId}>
                  {mic.label || `Mic ${mic.deviceId.slice(0, 6)}`}
                </option>
              ))}
            </select>
          </div>

          <div
            style={{
              padding: 16,
              background: shiftDown ? "#dcfce7" : "#f8fafc",
              borderRadius: 8,
              border: shiftDown
                ? "2px solid #10b981"
                : "2px dashed #cbd5e1",
              textAlign: "center",
              transition: "all 120ms ease",
            }}
          >
            <p style={{ margin: 0, fontWeight: 600, fontSize: 13 }}>
              {shiftDown ? "Listening..." : "Hold Shift to dictate"}
            </p>
            {result && (
              <p
                style={{
                  margin: "8px 0 0",
                  fontStyle: isFinal ? "normal" : "italic",
                  color: isFinal ? "#0f172a" : "#64748b",
                }}
              >
                {result}
              </p>
            )}
          </div>

          {error && (
            <p style={{ color: "#ef4444", fontSize: 13, marginTop: 8 }}>
              Recognition error: {error.error}
            </p>
          )}

          <input
            value={query}
            onChange={(e) => setQuery(e.target.value)}
            placeholder="Search query..."
            style={{
              width: "100%",
              marginTop: 12,
              padding: 10,
              borderRadius: 6,
              border: "1px solid #cbd5e1",
              fontSize: 16,
            }}
          />
        </>
      )}
    </div>
  );
}

Enter fullscreen mode Exit fullscreen mode

Four hooks, four orthogonal concerns:

  • usePermission drives the badge in the header and gates the rest of the UI behind the user's actual decision. Because it is reactive, if the user revokes mic access in browser settings, the badge updates and the input vanishes automatically.
  • useMediaDevices populates the microphone picker without forcing a permission prompt unless the user clicks "Grant".
  • useSpeechRecognition does the actual transcription, distinguishes interim from final results, and surfaces engine errors in a typed way.
  • useKeyModifier turns the Shift key into a push-to-talk trigger that survives focus loss, OS auto-repeat, and weird key combinations.

The component is roughly 130 lines, the vast majority of which is markup. The browser-API plumbing -- the parts that historically take the longest to get right -- is one import line per concern.

A Note on Testing

Voice and camera features are notoriously hard to test because the browser APIs they depend on require human gestures and physical hardware. The hooks all expose isSupported flags so your test environments (jsdom, Vitest, Storybook with mocked navigators) can branch cleanly on the absence of the underlying APIs and render fallback states. If you are building a serious voice UI, dedicate a small layer of integration tests that run in headless Chrome with a fake media stream -- that is the only way to catch the real bugs.

Installation

npm i @reactuses/core

Enter fullscreen mode Exit fullscreen mode

Related Hooks

  • useSpeechRecognition -- Live speech-to-text with interim and final result tracking
  • useMediaDevices -- Enumerate cameras and microphones with permission handling
  • usePermission -- Reactively query the Permissions API for any permission name
  • useKeyModifier -- Track OS-level modifier key state (Shift, Control, etc.)
  • useSupported -- Reactively check whether a browser API is available
  • useEventListener -- Attach event listeners declaratively for custom voice flows
  • useObjectUrl -- Create temporary URLs for recorded audio blobs to preview them

ReactUse provides 100+ hooks for React. Explore them all →