惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

The Last Watchdog
The Last Watchdog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
GbyAI
GbyAI
Y
Y Combinator Blog
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
The GitHub Blog
The GitHub Blog
博客园_首页
小众软件
小众软件
I
InfoQ
J
Java Code Geeks
月光博客
月光博客
S
Secure Thoughts
Microsoft Security Blog
Microsoft Security Blog
V
Visual Studio Blog
Hacker News - Newest:
Hacker News - Newest: "LLM"
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Stack Overflow Blog
Stack Overflow Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
N
News and Events Feed by Topic
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
The Cloudflare Blog
T
Threat Research - Cisco Blogs
A
About on SuperTechFans
H
Help Net Security
MongoDB | Blog
MongoDB | Blog
博客园 - 聂微东
人人都是产品经理
人人都是产品经理
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Latest news
Latest news
G
GRAHAM CLULEY
IT之家
IT之家
C
Cisco Blogs
Last Week in AI
Last Week in AI
Engineering at Meta
Engineering at Meta
L
LangChain Blog
The Register - Security
The Register - Security
SecWiki News
SecWiki News
M
MIT News - Artificial intelligence
NISL@THU
NISL@THU
T
Tenable Blog
博客园 - Franky
美团技术团队
I
Intezer
U
Unit 42
雷峰网
雷峰网
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
S
SegmentFault 最新的问题
C
Cyber Attacks, Cyber Crime and Cyber Security

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
On-Device Pose Estimation on iOS: What Actually Works in Production (Not Just Research Papers)
Benjamin Pir · 2026-05-11 · via DEV Community

Research papers on pose estimation show impressive accuracy numbers. Production apps on consumer devices tell a different story. Here's what I learned shipping real-time pose estimation to thousands of users across 22 sports.

I built SportsReflector, an AI coaching app that analyzes athletic form using Apple's Vision framework on-device. The app runs pose estimation at 30fps during live sessions and frame-by-frame during video analysis. This article covers the gap between what the documentation promises and what actually works when real users point their iPhones at themselves in gyms, courts, and living rooms.

The Model Options on iOS
Apple gives you three paths for pose estimation on iOS:

VNDetectHumanBodyPoseRequest (Vision framework)
Extracts 19 body points. Runs on the Neural Engine. No custom model needed. This is what most developers should use.

CreateML trained custom model
Train your own pose model with labeled data. More control over which points you detect. Requires training data which is expensive to create for sports.

Third-party CoreML models (MoveNet, BlazePose, PoseNet)
Converted from TensorFlow/PyTorch. More keypoints (33 with BlazePose vs 19 with Vision). Harder to optimize for Neural Engine. Often slower than Apple's native model.

What I chose and why:
Apple's VNDetectHumanBodyPoseRequest. The 19 keypoints are sufficient for scoring form in every sport I've tested. The performance is dramatically better than converted third-party models because Apple optimized it specifically for their Neural Engine silicon. On an iPhone 13 or newer, single-frame inference is 8-12ms — fast enough for 60fps processing with headroom.

The third-party models give you more keypoints but at 2-3x the inference cost. For a research project where accuracy matters more than speed, use BlazePose. For a consumer app where smooth real-time performance matters more than marginal accuracy, use Apple's native model.

What the Documentation Doesn't Tell You
Problem 1: Confidence scores vary wildly by body position
Apple's pose estimation returns a confidence score (0.0-1.0) for each keypoint. The documentation suggests filtering points below 0.5 confidence. In practice, this threshold is too aggressive for many athletic movements.

During a squat at the bottom position, hip keypoints regularly drop to 0.3-0.4 confidence because the thighs occlude the hip joint from the camera's perspective. During a boxing combination, the rear hand drops below 0.3 when it's behind the torso. During a tennis serve at trophy position, the tossing arm's wrist confidence drops when it crosses behind the head.

The fix was sport-specific confidence thresholds. For squats, I accept hip keypoints down to 0.2 and interpolate position from adjacent frames when confidence is low. For boxing, I track the rear hand's last known position and predict its current position when occluded. For tennis, I use temporal smoothing across 5-frame windows to maintain continuity through low-confidence phases.

What I'd tell other developers: don't use a single confidence threshold globally. Calibrate per body part and per activity type. The wrist needs different thresholds than the shoulder, and both need different thresholds during a squat vs a sprint.

Problem 2: Camera angle dramatically affects accuracy

The documentation shows pose estimation working with a straight-on camera view. Users don't set up cameras at optimal angles. They prop their phone against a water bottle on the gym floor. They lean it against a wall at a 30-degree angle. Their training partner holds it at head height. Their tripod is behind them.
Accuracy degrades significantly at angles beyond 30 degrees from perpendicular. Side views work well for sagittal-plane movements (squats, deadlifts, running gait). Front views work for frontal-plane movements (lateral raises, jumping jacks). But most users don't know which angle to use for which exercise.
The fix was adding a camera setup guide that shows the user exactly where to place their phone for each exercise type, plus an automatic angle quality check that warns users when their camera angle is suboptimal before they start recording.
What I'd tell other developers: never assume optimal camera placement. Your users will find every possible bad angle. Build angle detection into your pipeline and guide users toward good placement before processing.
Problem 3: Lighting conditions in gyms are terrible
Research papers evaluate pose estimation under controlled lighting. Gyms have mixed lighting — overhead fluorescents, natural light from windows, mirror reflections creating secondary light sources, and dark corners near cable machines.
Low-light conditions cause two problems. Frame noise reduces keypoint confidence. And auto-exposure adjustments cause momentary brightness shifts that confuse frame-to-frame tracking.
The fix was pre-processing frames with adaptive histogram equalization before feeding them to the pose estimator, plus locking camera exposure after the initial setup phase to prevent mid-recording brightness shifts.
What I'd tell other developers: test your pose estimation in the worst lighting you can find. If it works under a single dim bulb with mirror reflections, it'll work everywhere.
Problem 4: Clothing and equipment cause occlusion
Loose clothing (baggy gym shorts, hoodies, wide-leg joggers) hides joint positions. The pose estimator can't see the knee through baggy shorts, so it guesses — often incorrectly.
Equipment compounds this. A barbell across the shoulders during a squat occludes the neck and upper back keypoints. A tennis racket in the hand confuses wrist detection. Boxing gloves change hand proportions that the model expects.
There's no clean fix for this. The mitigations are temporal smoothing (use the trajectory of the joint over multiple frames to predict position during occlusion), anatomical constraints (the knee can't be above the hip during a squat, so cap estimates to physiologically possible ranges), and user guidance (suggest form-fitting clothing in the onboarding flow).
What I'd tell other developers: your pose estimation will fail on some percentage of users due to clothing and equipment. Build graceful degradation — lower confidence scores should trigger warnings to the user rather than wildly incorrect analysis.
Problem 5: Multiple people in frame
Gym environments frequently have other people in the background. The pose estimator detects all of them. If you're not careful, your analysis might score the person walking behind your user instead of your user.
The fix was implementing a "primary subject" tracking system. On the first frame, detect all bodies and select the largest (closest to camera) as the primary subject. Track that subject's position across frames using centroid tracking. Ignore all other detected bodies. If the primary subject disappears (walks out of frame), pause analysis and prompt the user to reposition.

What I'd tell other developers: always implement subject isolation if your app runs in environments with multiple people. Single-person pose estimation in multi-person environments is a pipeline problem, not a model problem.

Performance Optimization That Actually Matters
Frame skipping for battery life
Running pose estimation at 30fps continuously drains battery fast. For live AR feedback, you need every frame. For video analysis (deferred path), you don't.

For deferred analysis, I process every 3rd frame (10fps effective) for initial pass, then re-process key frames (phase transitions, lowest position in squat, peak extension in serve) at full resolution. This reduces processing time by 60% with negligible accuracy loss for scoring.

For live AR during workouts, I dynamically adjust processing frequency based on battery level. Above 50% battery, process every frame. Below 50%, process every 2nd frame. Below 20%, process every 3rd frame and show a battery warning.
Memory management during long sessions

A 60-minute workout session at 30fps generates 108,000 frames. Storing pose data for every frame would consume hundreds of megabytes. The fix is streaming analysis — process each frame, extract metrics, store only the aggregate metrics (per-rep scores, phase timestamps, anomaly flags), and discard the raw pose data.
For video analysis where users want to replay with skeleton overlay, store keypoints for key frames only (phase transitions, anomalies) and interpolate between them during playback.
Neural Engine vs GPU vs CPU
Apple's Neural Engine is fastest for pose estimation but isn't always available — it's shared with other system processes. The Vision framework automatically falls back to GPU or CPU when the Neural Engine is busy.
The problem: inference time varies 3-5x between Neural Engine (8ms) and CPU fallback (30-40ms). If your real-time pipeline assumes consistent 8ms inference, CPU fallback frames cause visible stuttering in the AR overlay.
The fix was building a frame budget system that measures actual inference time per frame and dynamically adjusts overlay rendering complexity. Fast inference: full skeleton with joint angles and color-coded feedback. Slow inference: simplified skeleton with key joints only. The user sees smooth animation regardless of which processor handles inference.
The Results
After shipping these optimizations, the production performance profile looks like this:

Real-time AR overlay: 30fps sustained on iPhone 12+, 60fps on iPhone 13 Pro+
Video analysis: 10-15 seconds for a 30-second clip
Battery consumption during 60-minute AR workout: approximately 15-20% battery drain
Crash rate related to pose estimation: 0.0% over 30 days (error boundaries catch all edge cases)
User-reported accuracy satisfaction: tracked through review sentiment

The gap between research demos and production apps is significant. Research optimizes for accuracy on clean datasets. Production optimizes for resilience across terrible camera angles, bad lighting, occluded joints, multiple subjects, and variable processing budgets. Understanding this gap early would have saved me months.

SportsReflector is available on the iOS App Store. Built with Apple's Vision framework, CoreML, and ARKit. If you're shipping pose estimation in production and dealing with similar issues, I'd love to compare notes.