惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

大猫的无限游戏
大猫的无限游戏
U
Unit 42
T
Tailwind CSS Blog
罗磊的独立博客
WordPress大学
WordPress大学
小众软件
小众软件
Recent Announcements
Recent Announcements
博客园 - 聂微东
Jina AI
Jina AI
云风的 BLOG
云风的 BLOG
博客园 - 【当耐特】
爱范儿
爱范儿
Microsoft Azure Blog
Microsoft Azure Blog
GbyAI
GbyAI
V
V2EX
博客园 - 三生石上(FineUI控件)
I
InfoQ
雷峰网
雷峰网
G
Google Developers Blog
阮一峰的网络日志
阮一峰的网络日志
B
Blog
腾讯CDC
A
About on SuperTechFans
博客园 - 叶小钗

Hacker News: Ask HN

The New Window Delete ChatGPT Atlas Spyware Tell HN: Qwen Free Tier Is Discontinued Ask HN: SeedLegals Partnerships in London, worth it? Ask HN: How to highlight talent from untraditional backgrounds? Ask HN: We dont need a programming language now? Durable Object alarm loop: $34k in 8 days, zero users, no platform warning What if Time at the subatomic level has multiple arrows? How to add MidnightBSD Key to UEFI Secure Boot DBX? (Revoked and Forbidden Keys) Ask HN: What's your experience working at xAI as an AI tutor? Any engineers here with experience of clinical data standards? Ask HN: Who is using OpenClaw? Agent Skills for Software Test Automation Ask HN: Who needs contributors? Claude Code is thinking too much Ask HN: What Is the Big-O Order of a Jigsaw Puzzle? Ask HN: Stepping into a new role as a Senior, mentoring dos and dont's? Founder from Zurich heading to SF and Austin for the first time Hacker News No Manual Screenshots: I Built a Scalable Screenshot API Using Cloud Playwright Ask HN: Thought experiment: AGI giving us answers we don't like? Ask HN: I quit my job over weaponized robots to start my own venture 1% Vacancy, 81% Preleased: Where Midmarket Compute Deploys in 2026 Ask HN: Preferred pricing model for sound effects libraries? Copy of the email I sent to my undergraduate professors on Nov 30, 2025 Model API Performance | Hacker News Ask HN: Are open-weight LLMs the new offline encyclopedias? Valgrind 3.27 RC1 is out Claude Code OAuth down for >12 hours Ask HN: What's Better?–Tauri or Electron?
Stera: Open-Source Infra That Turns iPhones into Spatial ...
satpal · 2026-05-16 · via Hacker News: Ask HN

We are releasing Project Stera - an open source, end-to-end pipeline that turns a commodity iPhone into a research-grade capture system for embodied AI training data.

Today, we're open-sourcing the whole stack, along with Stera-10M, a 200+ hour dataset, and 10M+ frames captured entirely through it.

FPV Labs began with one bet - the scaling law for embodied AI will need high-fidelity, multimodal real-world data, and the underlying infrastructure that produces this at scale without compromising downstream quality will determine how fast we build a general-purpose model.

Over the last 12 months, we've seen how high-fidelity data is locked behind gated hardware like Aria, which is out of reach for researchers, builders, and startups that want to work with high-quality multi-modal data, and how every single data lab ends up rebuilding the same harness for their own fleet.

This has led to immense data fragmentation over the last year and a race to collect data that is either low-fidelity or built on a heterogeneous stack, with multiple trade-offs in data quality.

Stera removes the need for gated hardware and turns a commodity iPhone into a high-fidelity spatial data capture engine, and provides open-source tooling via the Stera SDK to read, process, and export the results for downstream eval and training.

It fuses RGB, IMU, Depth, Lidar-guided depth, and 6DoF poses from ARKit out of the box and processes them via the Stera SDK to generate high-fidelity 4D data with spatial, semantic, action, and temporal understanding of the world

Each session from Stera includes a. what the wearer sees b. how the camera moves through space (6-DoF pose) how the hands move (21-joint, anchored in a global frame) c. what the depth geometry looks like (per-frame depth and a session-level room mesh) d. What the IMU measures e. What task, sub-goal, episode, atomic action, and objects are involved (hierarchical instruction tree)

We are also releasing Stera-10M publicly today, which is collected entirely through the Stera stack, so everyone can play with our datasets and reproduce them themselves, without having to rebuild any of the harnesses from scratch.

Think of Stera as a unified interface for the long tail of embodied AI research, allowing anyone to become a data lab today.

Downstream applications include, but are not limited to pre- and mid-training for VLA, World model, World Action Models, action recognition and temporal segmentation models, hand-object interaction modeling, human-to-robot motion retargeting, real-to-sim reconstruction, and so much more.

The data foundation for embodied AI should be open and accessible to every researcher/builder.

Link: https://www.fpvlabs.ai/essays/launching-stera