惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Google DeepMind News
Google DeepMind News
博客园 - 司徒正美
WordPress大学
WordPress大学
爱范儿
爱范儿
小众软件
小众软件
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
罗磊的独立博客
博客园_首页
V
V2EX
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
T
Tailwind CSS Blog
大猫的无限游戏
大猫的无限游戏
The Cloudflare Blog
MyScale Blog
MyScale Blog
IT之家
IT之家
H
Help Net Security
Blog — PlanetScale
Blog — PlanetScale
Microsoft Security Blog
Microsoft Security Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Recent Announcements
Recent Announcements
F
Fortinet All Blogs
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
人人都是产品经理
人人都是产品经理

Hacker News: Ask HN

The New Window Delete ChatGPT Atlas Spyware Tell HN: Qwen Free Tier Is Discontinued Ask HN: SeedLegals Partnerships in London, worth it? Ask HN: How to highlight talent from untraditional backgrounds? Ask HN: We dont need a programming language now? Durable Object alarm loop: $34k in 8 days, zero users, no platform warning What if Time at the subatomic level has multiple arrows? How to add MidnightBSD Key to UEFI Secure Boot DBX? (Revoked and Forbidden Keys) Ask HN: What's your experience working at xAI as an AI tutor? Any engineers here with experience of clinical data standards? Ask HN: Who is using OpenClaw? Agent Skills for Software Test Automation Ask HN: Who needs contributors? Claude Code is thinking too much Ask HN: What Is the Big-O Order of a Jigsaw Puzzle? Ask HN: Stepping into a new role as a Senior, mentoring dos and dont's? Founder from Zurich heading to SF and Austin for the first time Hacker News No Manual Screenshots: I Built a Scalable Screenshot API Using Cloud Playwright Ask HN: Thought experiment: AGI giving us answers we don't like? Ask HN: I quit my job over weaponized robots to start my own venture 1% Vacancy, 81% Preleased: Where Midmarket Compute Deploys in 2026 Ask HN: Preferred pricing model for sound effects libraries? Copy of the email I sent to my undergraduate professors on Nov 30, 2025 Model API Performance | Hacker News Ask HN: Are open-weight LLMs the new offline encyclopedias? Valgrind 3.27 RC1 is out Claude Code OAuth down for >12 hours Ask HN: What's Better?–Tauri or Electron?
Stera: Open-Source Infra That Turns iPhones into Spatial ...
satpal · 2026-05-16 · via Hacker News: Ask HN

We are releasing Project Stera - an open source, end-to-end pipeline that turns a commodity iPhone into a research-grade capture system for embodied AI training data.

Today, we're open-sourcing the whole stack, along with Stera-10M, a 200+ hour dataset, and 10M+ frames captured entirely through it.

FPV Labs began with one bet - the scaling law for embodied AI will need high-fidelity, multimodal real-world data, and the underlying infrastructure that produces this at scale without compromising downstream quality will determine how fast we build a general-purpose model.

Over the last 12 months, we've seen how high-fidelity data is locked behind gated hardware like Aria, which is out of reach for researchers, builders, and startups that want to work with high-quality multi-modal data, and how every single data lab ends up rebuilding the same harness for their own fleet.

This has led to immense data fragmentation over the last year and a race to collect data that is either low-fidelity or built on a heterogeneous stack, with multiple trade-offs in data quality.

Stera removes the need for gated hardware and turns a commodity iPhone into a high-fidelity spatial data capture engine, and provides open-source tooling via the Stera SDK to read, process, and export the results for downstream eval and training.

It fuses RGB, IMU, Depth, Lidar-guided depth, and 6DoF poses from ARKit out of the box and processes them via the Stera SDK to generate high-fidelity 4D data with spatial, semantic, action, and temporal understanding of the world

Each session from Stera includes a. what the wearer sees b. how the camera moves through space (6-DoF pose) how the hands move (21-joint, anchored in a global frame) c. what the depth geometry looks like (per-frame depth and a session-level room mesh) d. What the IMU measures e. What task, sub-goal, episode, atomic action, and objects are involved (hierarchical instruction tree)

We are also releasing Stera-10M publicly today, which is collected entirely through the Stera stack, so everyone can play with our datasets and reproduce them themselves, without having to rebuild any of the harnesses from scratch.

Think of Stera as a unified interface for the long tail of embodied AI research, allowing anyone to become a data lab today.

Downstream applications include, but are not limited to pre- and mid-training for VLA, World model, World Action Models, action recognition and temporal segmentation models, hand-object interaction modeling, human-to-robot motion retargeting, real-to-sim reconstruction, and so much more.

The data foundation for embodied AI should be open and accessible to every researcher/builder.

Link: https://www.fpvlabs.ai/essays/launching-stera