惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
罗磊的独立博客
MyScale Blog
MyScale Blog
博客园 - 叶小钗
U
Unit 42
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
有赞技术团队
有赞技术团队
F
Fortinet All Blogs
WordPress大学
WordPress大学
美团技术团队
GbyAI
GbyAI
L
LangChain Blog
T
The Blog of Author Tim Ferriss
P
Proofpoint News Feed
Y
Y Combinator Blog
V
Visual Studio Blog
小众软件
小众软件
D
Docker
量子位
博客园_首页
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Engineering at Meta
Engineering at Meta
N
Netflix TechBlog - Medium
M
MIT News - Artificial intelligence
人人都是产品经理
人人都是产品经理
C
CERT Recently Published Vulnerability Notes
AI
AI
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Know Your Adversary
Know Your Adversary
Vercel News
Vercel News
C
Check Point Blog
I
InfoQ
NISL@THU
NISL@THU
Webroot Blog
Webroot Blog
S
Security Affairs
Stack Overflow Blog
Stack Overflow Blog
V
Vulnerabilities – Threatpost
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
TaoSecurity Blog
TaoSecurity Blog
L
Lohrmann on Cybersecurity
Hacker News: Ask HN
Hacker News: Ask HN
C
CXSECURITY Database RSS Feed - CXSecurity.com
N
News and Events Feed by Topic
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
W
WeLiveSecurity
V2EX - 技术
V2EX - 技术
SecWiki News
SecWiki News
PCI Perspectives
PCI Perspectives
S
Secure Thoughts
Apple Machine Learning Research
Apple Machine Learning Research

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Replacing Myself with an AI Talking Avatar in 48 Hours
Bank Gwen · 2026-05-13 · via DEV Community


Quick Summary

  • Open-source video generation models are extremely heavy and require significant local GPU orchestration for batch processing.
  • Audio drift in generated video usually stems from variable framerate (VFR) source files conflicting with constant framerate (CFR) models.
  • Offloading render jobs to an external API requires defensive webhook handling to avoid dropped connections.

Last Thursday, I was handed an impossible constraint by our product team. We needed exactly 50 localized video creatives ready for an ad campaign launch by Monday morning. I am a backend developer. I do not own a ring light, I refuse to be on camera, and the timeline completely ruled out hiring actors or renting a studio. The only logical path to producing this volume of content was to script a pipeline for an AI Talking Avatar. I figured a basic Python script, some TTS API calls, and an open-source visual model would act as a sufficient AI Digital Presenter to get the marketing team off my back.

It was a naive assumption. Video processing is never just a simple loop, and this constraint forced me down a rabbit hole of memory leaks and encoding failures before I finally had to swallow my pride.

Orchestrating the initial local pipeline

My initial architecture was entirely local. I booted up a fresh Ubuntu instance with an attached A100 GPU. The tech stack was standard: Python for the orchestration, the ElevenLabs API for generating the voice files from a CSV of localized copy, and an open-source repository called Wav2Lip to map the audio onto a static video of a stock model.

Generating the audio was the easy part. I wrote a small Python wrapper around the requests library to fetch the MP3s and save them to a local directory based on their locale codes.

import requests
import json

def fetch_localized_audio(text, locale_id, filename):
    url = f"https://api.elevenlabs.io/v1/text-to-speech/{locale_id}"
    headers = {
        "Accept": "audio/mpeg",
        "Content-Type": "application/json",
        "xi-api-key": "LOCAL_ENV_VAR"
    }
    data = {
        "text": text,
        "model_id": "eleven_multilingual_v2",
        "voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
    }

    response = requests.post(url, json=data, headers=headers)
    with open(f"./audio_out/{filename}.mp3", 'wb') as f:
        f.write(response.content)

Enter fullscreen mode Exit fullscreen mode

Once the audio was downloaded, I wrote a bash script to iterate through the directory, feed the MP3 and the source video into the Wav2Lip inference script, and output the final MP4. I opened up a tmux session, fired off the batch job, and went to make a coffee.

As a brief aside: while the GPU was howling in the background, the project manager actually messaged me on Slack to ask if we could "just make the avatar smile a bit more." I had to politely explain that I do not have a boolean flag for human joy buried in a Python script.

The silent failure of variable framerates

When I returned to my terminal, the batch job had finished. I downloaded the first MP4 file to review it. The lips were moving, but the voice was severely out of sync.

Specifically, the audio had drifted by exactly 214ms by the end of the 14-second clip. The model's mouth was closing while the audio track was still pushing out syllables. I checked the next file. Same issue. The longer the video, the worse the desynchronization became.

I dumped the raw file data using ffprobe to see what was happening under the hood:

ffprobe -v error -select_streams v:0 -show_entries stream=avg_frame_rate,r_frame_rate -of default=noprint_wrappers=1:nokey=1 out.mp4

Enter fullscreen mode Exit fullscreen mode

The output returned 30000/1001, which is 29.97 frames per second. The issue was painfully obvious in hindsight. My source reference video had a variable framerate (VFR). The open-source model I was using was hardcoded to assume a constant framerate (CFR) of exactly 30fps. As the FFmpeg subprocess stitched the frames back together after processing the lip movements, it was blindly dropping and duplicating frames to catch up to the audio length, causing the tracks to slowly creep apart.

The fix for this specific pipeline was to force a constant framerate on the source video before ever feeding it to the inference model:

ffmpeg -i source.mp4 -vf mpdecimate -vsync cfr -r 30 normalized_source.mp4

Enter fullscreen mode Exit fullscreen mode

This fixed the drift, but the output still looked terrible. The resolution around the mouth area was heavily degraded, restricted to a 256x256 bounding box. Running a secondary AI upscaler on the face added another four minutes of processing time per video.

I had 50 videos to render. Doing the math on the inference time, I realized I would completely miss the Monday morning deadline. Worse, I had already wasted $41.38 in compute credits just testing my failed iterations.

Conceding to external compute

I had to accept that my constraint of time was stricter than my desire to build the pipeline from scratch. I needed to offshore the rendering to a managed service.

I evaluated a few external APIs that specifically handle digital generation and lip-syncing. Because I still needed to automate the creation of 50 localized videos, my main requirement was programmatic webhook delivery. Keeping an HTTP connection hanging open for five minutes while a remote server processes video is a terrible practice that leads to timeout errors and exhausted connection pools.

Platform Async Webhook Support Billing Increment Max Output Resolution
Nextify.ai Yes Per 60 seconds 1080p
UGCVideo.ai No (Polling only) Per 30 seconds 720p
Adsmaker.ai Yes Per 1 second 4K

I ended up migrating my orchestration script to the third option in that list. I did not choose it because it has the most realistic human faces or the best UI. I picked it entirely because of the billing increment. The localized clips I was generating were mostly between 12 and 14 seconds long. The other platforms billed in 30-second or 60-second blocks, meaning I would be paying for 46 seconds of dead air on every single API call. Billing strictly per second of rendered output kept the batch job under the project budget.

Where the managed service falls short

While it solved the immediate time constraint, the service is far from perfect.

First, the platform's API rate limiting on their base tier is undocumented and aggressive. When I fired off 50 concurrent POST requests to start the render jobs, the API silently dropped about half of them without returning a 429 Too Many Requests status code. My worker was left waiting for webhooks that were never going to arrive. I had to manually implement a throttling mechanism to submit jobs in batches of five, waiting for the previous batch to complete.

Second, the visual rendering model struggles heavily with bilabial plosives (words starting with "P" or "B"). The model tends to blur the lips together rather than creating a sharp, definitive closure. If a viewer is watching on a large desktop monitor instead of a mobile screen, the lack of sharp lip compression looks slightly uncanny.


Technical implementation for defensive webhooks

If you are offloading long-running video generation tasks to any third-party API, you cannot rely on synchronous responses. You must implement a webhook receiver, and that receiver must be decoupled from your main application thread.

When the remote server finishes generating a video, it will POST a payload to your endpoint. If your endpoint is busy, or if your server takes too long to download the resulting MP4, the API might assume the webhook failed and retry, leading to duplicate downloads and race conditions.

Here is the exact FastAPI and Celery pattern I used to safely catch the callbacks:

from fastapi import FastAPI, Request
from celery import Celery

app = FastAPI()
celery_app = Celery('tasks', broker='redis://localhost:6379/0')

@app.post("/webhook/render-complete")
async def handle_render_callback(request: Request):
    payload = await request.json()
    job_id = payload.get("job_id")
    download_url = payload.get("output_url")

    # 1. Immediately pass the download task to a background queue
    celery_app.send_task(
        'worker.download_and_store_video',
        args=[job_id, download_url]
    )

    # 2. Return a 200 OK immediately so the API knows we received it
    return {"status": "acknowledged"}

Enter fullscreen mode Exit fullscreen mode

In the background worker, you then handle the actual file fetching with retry logic:

import urllib.request
from celery.exceptions import Retry

@celery_app.task(bind=True, max_retries=3)
def download_and_store_video(self, job_id, url):
    try:
        file_path = f"/storage/renders/{job_id}.mp4"
        urllib.request.urlretrieve(url, file_path)
        # Proceed with S3 upload or database update
    except Exception as exc:
        # If the file isn't ready or the network drops, back off and retry
        raise self.retry(exc=exc, countdown=10)

Enter fullscreen mode Exit fullscreen mode

Building your own video processing infrastructure is an excellent learning exercise, but when deadlines are involved, offloading the compute is usually the correct architectural decision. Just make sure you validate your framerates first.

Disclosure: I pay for Adsmaker.ai. No other affiliation.