惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Privacy & Cybersecurity Law Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
D
Docker
V
V2EX
GbyAI
GbyAI
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - Franky
Jina AI
Jina AI
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
I
InfoQ
博客园 - 司徒正美
雷峰网
雷峰网
F
Full Disclosure
S
SegmentFault 最新的问题
大猫的无限游戏
大猫的无限游戏
博客园 - 叶小钗
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
D
DataBreaches.Net
M
MIT News - Artificial intelligence
V
Visual Studio Blog
H
Help Net Security
月光博客
月光博客
博客园 - 三生石上(FineUI控件)
博客园_首页
O
OpenAI News
人人都是产品经理
人人都是产品经理
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Attack and Defense Labs
Attack and Defense Labs
Blog — PlanetScale
Blog — PlanetScale
爱范儿
爱范儿
罗磊的独立博客
P
Palo Alto Networks Blog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
博客园 - 聂微东
Last Week in AI
Last Week in AI
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
Lohrmann on Cybersecurity
N
News and Events Feed by Topic
有赞技术团队
有赞技术团队
The Register - Security
The Register - Security
S
Security @ Cisco Blogs
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
酷 壳 – CoolShell
酷 壳 – CoolShell
AWS News Blog
AWS News Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
L
LINUX DO - 最新话题
Hacker News - Newest:
Hacker News - Newest: "LLM"
T
Threat Research - Cisco Blogs

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The Hardest Part of Building an AI-Powered WebRTC Platform Wasn’t WebRTC
Anupam Kumar · 2026-05-13 · via DEV Community

I spent the last few months building SkyMeetAI — a video conferencing platform that watches participants' faces and listens to their voices in real time to figure out how they're feeling. Not via surveys. Live, during the call, every 150–500ms.

Redis returned null into a socket event handler at 4 AM once, and my participant list had quietly vanished. No error. No trace. Just... gone.

That happened more than once. This post is about what that actually took to build — the distributed systems problems, the browser media weirdness, the ML decisions, and the bugs I'm documenting anyway because that's how this stuff works.

Live demo: skymeetai.onrender.com (live preview may not be available due to resource constraints) | Repo: github.com/AnupamKumar-1/skymeetAI


Stack

  • Frontend: React + WebRTC (RTCPeerConnection, AudioWorklet, MediaRecorder)
  • Backend: Node.js + Express + Socket.IO + @socket.io/redis-adapter + MongoDB
  • Emotion service: Python FastAPI + Socket.IO + MediaPipe + Wav2Vec2 + PyTorch + XGBoost
  • Transcription: Python FastAPI + Whisper (small) + DistilRoBERTa

Four services. Two language ecosystems. None share memory, a filesystem, or a database.


What it actually does

Standard WebRTC video call. But while the meeting's running, the host's browser is also:

  • Capturing JPEG frames (720×540) from each remote participant's <video> element
  • Extracting Float32 PCM audio chunks at 16kHz via AudioWorklet
  • Streaming both to a Python inference service over Socket.IO
  • Receiving an emotion label back per participant, every 150–500ms

When the call ends, each participant's recorded audio goes to a separate transcription service — Whisper transcribes it, a fine-tuned DistilRoBERTa model classifies the emotion of each spoken segment, and you get a structured summary: dominant emotion, key topics, per-speaker word counts, speaking pace in WPM.

Architecture: Browser connects to Backend (Socket.IO + REST), Emotion Service (Socket.IO), and Transcription Service (HTTP). Transcription delivers results back to Backend via an HTTP POST callback. Backend uses MongoDB for persistence and Redis for ephemeral state, locks, and pub/sub across its three pm2 processes.

(See architecture diagram above.)


The backend: where the actual pain lives

Three processes, one room

The backend runs as three pm2 processes on ports 8000, 8001, and 8002. Socket.IO's room abstraction is in-process by default — a client on port 8000 can't receive events emitted on port 8001.

Solution: @socket.io/redis-adapter. Room state gets externalised to two Redis keys per room:

meeting:state:<code>         → JSON array of socket IDs (join order)
meeting:participants:<code>  → Map<userId, {socketId, userId, meta}>

Enter fullscreen mode Exit fullscreen mode

No process holds authoritative room state in memory. The tradeoff: every Redis hiccup propagates directly into socket event handlers. I have gaps in null-guard coverage on Redis reads — partial failure produces subtly wrong behaviour rather than a clean error. That's documented in docs/backend.md and still on the fix list.

One deployment constraint the application itself doesn't enforce: the load balancer must provide sticky sessions so each client's Socket.IO handshake and subsequent requests land on the same process every time.

The concurrent join race condition

When two people join simultaneously, every backend process that receives join-call tries to read-modify-write the participant list back to Redis. Classic race. The fix is a distributed lock in redisLock.js:

// Acquire
SET lock:room:<code> token NX PX 10000

// Release — Lua compare-and-swap
// Only deletes if stored value matches caller's token

// Max wait: 8,000ms | Retry: 50ms + up to 50ms jitter

Enter fullscreen mode Exit fullscreen mode

The join handler holds the lock through the full participant state mutation. A crashed process blocks new joins for up to 10 seconds before the TTL expires. Acceptable tradeoff. Room capacity is enforced at 50 participants max.

Auth

JWT via passport-jwt. Per-IP and per-username rate limiting via a Redis Lua script. Account lockout after 10 consecutive failed logins with a 900-second TTL. bcrypt at 10 rounds. Room creation generates an 8-hex-character meetingCode and stores only the sha256 hash of the host secret — the raw secret is returned once at creation and never persisted.

One security gap I know about

The signal event in socket.controller.js forwards SDP/ICE candidates to any target socket ID without checking that both sockets are in the same room. Fine in practice. Real correctness problem on paper. On the fix list.


The browser: one stream, three consumers

Every remote participant's media stream has to simultaneously: play their video locally, stream frames to the emotion service, and record audio for post-meeting transcription.

Sharing the same stream reference across all three caused MediaRecorder and the Web Audio AnalyserNode to interfere in ways that were genuinely hard to reproduce. The fix was three separate tap points from the same underlying track:

  1. captureStream() on the <video> element → JPEG frames for inference (720×540, quality 0.82)
  2. Cloned MediaStreamAudioWorklet node → 1600-sample Float32 PCM chunks at 16kHz, running on a dedicated audio rendering thread, isolated from React's render loop
  3. Standard <video> element → local playback

The AudioWorklet isolation is what mattered. Running audio processing on the main thread caused subtle timing interference with the recorder. Moving it off fixed a whole class of problems at once.

WebRTC peer lifecycle

Perfect negotiation is implemented — polite/impolite roles assigned by the backend via assigned-role. Active speaker detection combines SSRC-based detection where available with an RMS-based AnalyserNode fallback. Both feed a shared score accumulator with decay and cooldown logic to avoid rapid speaker switching.

When video is turned off, a tiny black canvas stream is sent as a placeholder track — this avoids renegotiation on some browsers.

Back-pressure

When the emotion service's face executor queue depth hits 3, it emits a backpressure event to the browser carrying queueDepth, suggestedFps, and a timestamp. The client throttles its capture rate in response. Without this, a slow inference cycle would let the client-side queue grow until the tab ran out of memory. This was not a theoretical concern.

Chat

Messages get sanitised, written to MongoDB (capped at 500 messages), broadcast to the room, and ACK'd back to the sender within 5 seconds via chat-ack. If the ACK doesn't arrive, the message is marked failed and the user can retry. Server-side length enforcement runs independently of the client — you can't bypass the 2000-char limit from the browser.


The emotion pipeline

The core idea

Two models, two modalities, running in parallel and combined.

Video — MediaPipe Face Landmarker extracts 136 key landmarks (nose-centred, eye-rotation-corrected, inter-ocular scale normalised), 51 blendshapes, and 3D head pose angles — 326 dimensions per frame.

Audio — a fine-tuned Wav2Vec2 emotion model (audeering/wav2vec2-large-robust-12-ft-emotion-msp-dim) converts raw PCM into 1024-dimensional embeddings over a centred 0.6-second window.

Both streams are z-score normalised using pre-computed norm_stats.npz before hitting the ensemble.

The ensemble

EmotionTransformer (PyTorch) — dual-stream Transformer encoder with cross-modal attention. Face queries attend to audio keys/values and vice versa, gated by a learned scalar cross_gate. Auxiliary per-modality heads contribute 15% each to training loss. Trained in three curriculum phases. Test accuracy: 74.25% on an actor-disjoint split.

XGBoost — an 8,149-dimensional handcrafted feature vector (sequence stats, temporal deltas, blendshape features, head pose, audio rhythm, face motion energy), reduced to 512 dimensions via PCA. Test accuracy: 66.03% on the same split.

Final blend: calibrated weights, with temperature scaling applied to Transformer logits. Ensemble test accuracy: 74.34%.

The training detail that actually matters

Training used actor-disjoint splits — no actor appears in more than one of train, validation, or test. Without that constraint, the model quietly overfits to speaker identity instead of emotion, and evaluation metrics look fine while the model is doing something useless.

That's the kind of thing that only shows up when you test on genuinely unseen speakers. Most tutorials skip it.

How inference actually runs

The pipeline is audio-driven. When a chunk arrives, _ensure_pump(pid) starts a per-participant async pump coroutine if one isn't running. The pump enforces a minimum 300ms interval between inference cycles.

Single-slot buffers (_LATEST_AUDIO, _LATEST_FRAME) — rapid bursts coalesce to the newest payload instead of queuing. Face and audio embeddings run concurrently in separate 2-worker thread pools. Ensemble inference runs in a single serialised _inference_executor thread to reduce GPU memory contention.

Video frames are opportunistic: if one is available when inference fires, it's used. If not, the system falls back to audio-only. EMA smoothing at α=0.65 with a 2-second history TTL so a participant going quiet doesn't bleed stale emotional context into the next segment.

Modality staleness is checked at inference time — embeddings older than 0.4 seconds are masked out even if their rolling buffer still contains data.

The hard scaling limit: all per-participant state — embedding history, EMA state, pump coroutine handles — lives in Python process memory. One instance, no horizontal scaling. You'd need to externalise that state to Redis before you could run more than one.

Anomaly detection

Four IsolationForest models (one each for both, audio_only, video_only, and global_fallback) score each inference cycle. Anomaly detection doesn't block inference — it sets anomaly: true in the emotion.result payload and lets the client decide what to do with it. Real meetings have bad lighting, partial occlusions, background noise — knowing when not to trust the model matters.


The transcription pipeline (and the bug I'm embarrassed about)

After the call ends, useMeetingLifecycle.js uploads each participant's recorded audio to the transcription service as a multipart POST with an x-host-secret header. The service returns HTTP 202 immediately and does the actual work in a BackgroundTasks task:

  1. Validate extension against {webm, wav, mp3, m4a, ogg, aac, mp4}, sanitise filename via werkzeug
  2. Convert to mono 16kHz WAV via ffmpeg (fallback to original if conversion fails)
  3. Whisper small, English only, for timestamped transcription
  4. Merge consecutive same-speaker segments within 2 seconds and under 60 words
  5. Emotion classification per segment via DistilRoBERTa (segments under 4 words → "neutral" without inference)
  6. Sort all segments chronologically across speakers
  7. Generate structured analysis: summary, up to 5 key points, emotion distribution, top 8 topics by word frequency, per-speaker stats, speaking pace in WPM
  8. POST the whole thing to NODE_API via a single requests.post(…, timeout=None)

Step 8 is the problem.

If that callback fails, the transcript is gone. The frontend's upload call has a silent .catch(() => {}) in runBackgroundTranscript. The only signal that something went wrong is that nothing appears after 10 minutes of polling (30 × 20s, no backoff).

A message queue with at-least-once delivery — Redis streams, RabbitMQ, anything — would fix this entirely. It's the first thing I'd change if starting over, and the gap I'm most annoyed about.

One more thing: multiple audio files in a single request are processed sequentially in a Python for loop — not in parallel. And the Whisper and DistilRoBERTa models are module-level singletons with no explicit locking. Thread safety of concurrent requests depends entirely on the upstream library implementations.


Things I'd fix

Transcript delivery — the silent data loss is indefensible. Queue with retry, minimum.

Externalise emotion service state — in-process participant state means no horizontal scaling. Redis-backed buffers make it stateless.

Server-authoritative host enforcement — the backend validates x-host-secret on the transcript proxy route, but privileged socket events like end-meeting don't have equivalent server-side verification. A client with host:<code> in localStorage can trigger them without additional checks.

Scope the signal relay — room membership check on SDP/ICE forwarding in socket.controller.js.

Dynamic TURN credentialsmeetConfig.js contains hardcoded plaintext openrelayproject credentials. Per-session generation is the right call.

Coordinate the cleanup timer — all three pm2 processes run cleanupOldMeetings via setInterval at module import time. They all fire independently every hour. Redis leader election or distributed cron would fix duplicate execution.

Multilingual transcription — Whisper supports it natively. I hardcoded language="en". Easy fix, real impact.

Reconnect state reconciliation — when a participant reconnects, the backend reconstructs their record from Redis via restoreParticipant. But the emotion service still has their old embedding buffers and EMA state in process memory. No reconciliation step. The two stores drift.


Operational notes if you run this

  • Redis and MongoDB are both required before backend startup. Either missing → process.exit(1).
  • Load balancer must provide sticky sessions — the Socket.IO handshake must land on the same pm2 process every time. Each process has a max_memory_restart of 512 MiB with exponential-backoff restart.
  • The emotion service has no /health or /ready endpoint. Use a TCP probe. /stats and /stats/json expose live P50/P90/P95 inference latency — observability, not health.
  • TURN credentials are hardcoded plaintext in meetConfig.js. Don't deploy without fixing this.
  • CORS origins are hardcoded in app.js (localhost:3000 and skymeetai.onrender.com). Adding an origin requires a code change and redeploy.
  • Latency log is truncated on every backend process restart. No rotation.

Where it stands

It works. You can run a meeting, watch emotion labels update in real time as the conversation moves, end the call, and come back to a transcribed summary with per-segment emotion annotations, dominant emotion, key topics, per-speaker stats, and speaking pace. The distributed backend coordination holds together under load. Inference hits 150–500ms depending on load at 10 concurrent participants, with spikes to ~700ms under heavier load.

It also has real gaps — every one of them documented in the repo rather than hidden.


Want to contribute?

I'm opening good-first-issues soon. If any of the following sounds interesting, keep an eye on the repo:

  • Add /health endpoint to the emotion service — one route, no logic, makes load balancer config actually work
  • Add retry logic on the transcript HTTP callback — the embarrassing one; a queue wrapper fixes the silent data loss
  • Room membership check on the signal relay — small guard, real security improvement
  • Replace hardcoded TURN credentials with a per-session endpoint
  • Distributed cleanup timer — Redis leader election so three pm2 processes stop duplicating it
  • Whisper multilingual support — remove one hardcoded string, add language detection
  • Null-guard the Redis reads in socket.service.js — document identifies the gaps; needs someone to close them

Star the repo and watch for issues. Detailed implementation docs are in docs/frontend.md, docs/backend.md, docs/realTimeEmotionService.md, and docs/transcription-service.md — enough context to dig in before the issues even go up.

Happy to answer questions in the comments — especially about the decisions that turned out to be wrong.