MOVAS-EGO 100K  ·  100,000 HOURS IN INVENTORY
The Data Engine for Physical AI

Data moves robots.
MOVAS moves data.

HUMAN WORKER · HEAD-MOUNTED EIS OFF CONTACT MOVAS-OS EGO VIDEO HAND POSE · MANO CONTACT STATE OBJECT 6-DoF RETARGET · IK SIM VALIDATED ✓ HUMANOID · RETARGETED POLICY REAL WORK → ROBOT-EXECUTABLE BEHAVIOUR
100,000+Hours in inventory
100%Employer-signed
5Layers aligned per clip
3 daysTo an evaluation subset
( 02 )The problem

Three routes to physical data. Three structural defects.

None of them is a small gap that more budget closes. Each fails for a reason built into how the data is produced.

ROUTE 01

Teleoperated robot data

real-robot teleop

One robot, one hour, one trajectory. Scale is capped by hardware count rather than demand — and the scenes are laboratory scenes.

ROUTE 02

Web & crowdsourced video

consumer capture

Electronic image stabilisation rewrites pixel motion and is not invertible. No hand pose, no contact — and no authority over workplace footage.

ROUTE 03

Academic egocentric corpora

perception benchmarks

Built to measure recognition and memory, not to train control. No retargeted trajectories, and licensed research-only.

WHAT SHIPS IN EVERY CLIP

What ships inside a single clip.

COMPONENTS  0 / 6
01

Egocentric video

1080p30 · H.264 CRF 18

Head-mounted, stabilisation disabled at firmware level so ego-motion stays recoverable.

02

Inertial stream

IMU 200 Hz · 9-axis

Hardware-timestamped against the video, drift verified per clip.

03

Hand pose

MANO [T, 51] · L/R

Both hands, every frame, with temporal smoothing.

04

Contact state

Per-joint · per-frame

Derived from MANO surface distance, adjudicated on a sampled subset.

05

Object pose

6-DoF [T, 4, 4] + mask

Open-vocabulary detection, segmentation and pose through occlusion.

06

Retargeted trajectory

5 end-effectors

Physics-replayed before delivery; infeasible sequences rejected, not shipped.

Every clip is self-contained — no cross-referencing required to start training. A catalogue of unrelated datasets cannot deliver these six aligned on the same take.

FLAGSHIP CORPUS · MOVAS-EGO 100K
( 03 )The corpus

MOVAS-Ego 100K

Real people. Real tasks. Real environments.

The largest commercially-licensed egocentric corpus assembled for Physical AI. Captured head-mounted during professional work, with electronic image stabilisation disabled at firmware level so ego-motion stays recoverable downstream. Every hour sits under an employer-signed agreement.

Scene Distribution

Nine scene families, organised by work setting. Contact-rich, bimanual, tool-using work — food preparation, assembly and line work, fine assembly and specialised services — accounts for 31,428 hours, the slice where manipulation policies fail first.

Commercial Operations
25,636 h  25.6%
Food Preparation
17,770 h  17.8%
Domestic & Care Services
17,200 h  17.2%
Industrial & Production
11,000 h  11%
General Retail
9,823 h  9.8%
F&B Storefront
7,466 h  7.5%
Specialised Services
5,397 h  5.4%
Handcraft & Fine Assembly
3,461 h  3.5%
Market, Agri & Mobility
2,247 h  2.2%
PREMISES 80.6%
RESIDENTIAL 17.2%

80.6% commercial and industrial premises · 17.2% residences where the operator is employer-dispatched service staff · 2.2% open-air and in-transit. Every hour is professional work under an employer-signed agreement.

All nine scene families
Commercial Operations — 25.6% Food Preparation — 17.8% Domestic & Care Services — 17.2% Industrial & Production — 11% General Retail — 9.8% F&B Storefront — 7.5% Specialised Services — 5.4% Handcraft & Fine Assembly — 3.5% Market, Agri & Mobility — 2.2% 100K HOURS
  • Commercial Operations 25.6%
  • Food Preparation 17.8%
  • Domestic & Care Services 17.2%
  • Industrial & Production 11%
  • General Retail 9.8%
  • F&B Storefront 7.5%
  • Specialised Services 5.4%
  • Handcraft & Fine Assembly 3.5%
  • Market, Agri & Mobility 2.2%
LAYER ALIGNMENT · FIVE LAYERS ON ONE CLIP
( 04 )The vocabulary

One corpus. Every layer of the same moment.

These are not five datasets. Every layer is derived from — and time-locked to — the same footage, so hand pose, occlusion-free geometry, kinematics and robot-ready trajectories all describe the same moment.

LAYER ALIGNMENT ON A SINGLE CLIP
clip_00841.mp4 · 00:00 — 00:42 FAM-02 FOOD PREPARATIONSYNC DRIFT 0.4 ms / 60 s
SOURCE · RGB
L1–L5 CONTROL
4D GROUND TRUTH
MOCAP
3D ASSETS
RETARGET · 5 EE

Dashed verticals are hardware sync ticks. One take, six aligned streams, no post-hoc registration.

Hardware-synchronised by default. Delivered in UMI, LeRobot, RDT, Open-X-Embodiment or MOVAS native — to your cloud, your VPC, or air-gapped media.

BUILT FOR PHYSICAL INTELLIGENCE
( 05 )Built for physical intelligence

From pixels to actions.

One corpus, four places it lands in the Physical AI stack.

VLA models

vision → language → action

World models

geometry → dynamics → prediction

Robot learning

demonstration → retargeting → policy

Spatial intelligence

depth → 3D → 4D
How MOVAS supports embodied intelligence
AVAILABILITY · INVENTORY VS CAPACITY
( 06 )Availability

Start with existing inventory instead of waiting months for collection.

Both build-to-order and in-inventory are legitimate models. They are not the same purchase, and they do not carry the same schedule risk. We state which one applies on every line.

DELIVERY LEAD TIMES

Start with existing inventory.

Instead of waiting months for collection.

Evaluation subset (≤ 50 h)
on request
3 business days
Single scene family, L1–L3
up to 25,636 h
10 business days
Family with aligned layers, L1–L5
by scope
15 business days
Multi-family programme
by scope
20 business days
Targeted re-capture at a contracted site
by scope
scoped per engagement

All hours quoted are captured, refined and in storage — not collection capacity. Where a figure refers to work not yet captured, it is labelled as such.

Embodied Intelligence

From perceiving the world to executing complex tasks reliably. MOVAS supplies the pre-training corpus, the control-grade supervision on top of it, and the evaluation that proves the lift — then goes back to capture what your model still fails at.

  • Pre-training corpus

    100,000 hours of employer-cleared egocentric video with recoverable ego-motion, weighted toward contact-rich work.

  • Control-grade supervision

    Hand pose, per-joint contact, object 6-DoF and retargeted trajectories — aligned on the same clip.

  • Physics validation

    Every retargeted trajectory replayed in simulation before delivery; infeasible sequences rejected, not shipped.

  • Failure-driven re-capture

    Because the hours come from sites under contract, we can return to the same line or kitchen and capture the exact failure mode your policy is stuck on.

Embodied intelligence
REFINEMENT PIPELINE · MOVAS-OS
( 07 )The engine

We don’t just collect data. We engineer it for Physical AI.

From raw reality to robot-ready data.

Movas-OS turns raw physical footage into trainable robot experience. Six stages, each containerised and version-pinned — every output traces back to the exact model version that produced it. Available for your own footage as well: video containers, ROS bags and MCAP.

01

Candidate filter

view · VLM · hand visibility

02

4D reconstruction

depth · camera path · geometry

03

Hand pose

MANO · bimanual · smoothing

04

Object 6-DoF

detect · segment · pose

05

Retargeting

wrist map · IK · contact-preserving

06

Sim validation

replay · feasibility · score

raw footage → [01] → [06] → UMI / LeRobot / RDT / Open-X-Embodiment / MOVAS native

AI-assisted workflow · sovereign provenance · format export — up to 70% less manual annotation time than a standard pipeline.

Explore Movas-OS
ingest → egocentric/clip_00841.mp4 sync → imu 200Hz drift 0.4ms ✓ hand → HaMeR conf 0.93 L/R ✓ object → 6-DoF mug_blue_001 ✓ contact → frames 218-604 ✓ retarget→ shadow_hand ik ok ✓ sim → isaac replay PASS ✓ score → 87 / 100 → PRO TIER ───────────────────────────── ingest → egocentric/clip_00842.mp4 sync → imu 200Hz drift 0.6ms ✓ hand → HaMeR conf 0.71 L only fallback→ WiLoR conf 0.88 ✓
OPEN EVALUATION · MEVB
( 08 )Evaluation

MEVB

A shared yardstick for embodied policies, not for video understanding. Open tasks, a reproducible harness, sealed splits and a public leaderboard — independently reproduced before anything is published.

Open core · reproducible · academically governed
RIGHTS & PROVENANCE
( 09 )Trust

Who can grant rights to footage shot inside a business?

An individual can consent to their own likeness. They cannot, on their own, grant rights to their employer’s premises, processes, or to colleagues and customers appearing in frame. That is why MOVAS contracts with the employer first and the individual second — and why every clip carries a consent id traceable to both.

Contracted sites, not a contributor pool.

FAQ

Egocentric Data & Physical AI — Common Questions

Egocentric video is captured from a head-mounted camera worn by a person performing a task, so the viewpoint approximates what a humanoid robot would see doing the same work. Because human and humanoid kinematic chains are similar, policies pre-trained on egocentric video transfer to robots with far less adaptation than third-person video requires — and it is roughly a third to a fifth the cost of teleoperated real-robot collection.

100,000 hours across nine scene families and eighteen sub-classes. Commercial operations, home and domestic work, and food preparation are the three largest, together making up roughly 63 percent. The full breakdown is published on the MOVAS-Ego 100K dataset page.

UMI, LeRobot, RDT and Open-X-Embodiment, plus a native MOVAS format that preserves the full annotation stack. Data can be delivered to your cloud, to an air-gapped environment, or accessed through the Movas-OS platform.

Yes. Contributors sign employer-mediated agreements that cover commercial training use, and the corpus carries full provenance and audit records. This is the practical difference between MOVAS and academic egocentric corpora, which usually carry research-only terms.

Your models need the physical world.

Start with training-ready data today. Scale to targeted re-capture when you’re ready.

Request a Subset

Data moves robots. MOVAS moves data.