None of them is a small gap that more budget closes. Each fails for a reason built into how the data is produced.
One robot, one hour, one trajectory. Scale is capped by hardware count rather than demand — and the scenes are laboratory scenes.
Electronic image stabilisation rewrites pixel motion and is not invertible. No hand pose, no contact — and no authority over workplace footage.
Built to measure recognition and memory, not to train control. No retargeted trajectories, and licensed research-only.
Head-mounted, stabilisation disabled at firmware level so ego-motion stays recoverable.
Hardware-timestamped against the video, drift verified per clip.
Both hands, every frame, with temporal smoothing.
Derived from MANO surface distance, adjudicated on a sampled subset.
Open-vocabulary detection, segmentation and pose through occlusion.
Physics-replayed before delivery; infeasible sequences rejected, not shipped.
Every clip is self-contained — no cross-referencing required to start training. A catalogue of unrelated datasets cannot deliver these six aligned on the same take.
Real people. Real tasks. Real environments.
The largest commercially-licensed egocentric corpus assembled for Physical AI. Captured head-mounted during professional work, with electronic image stabilisation disabled at firmware level so ego-motion stays recoverable downstream. Every hour sits under an employer-signed agreement.
Nine scene families, organised by work setting. Contact-rich, bimanual, tool-using work — food preparation, assembly and line work, fine assembly and specialised services — accounts for 31,428 hours, the slice where manipulation policies fail first.
80.6% commercial and industrial premises · 17.2% residences where the operator is employer-dispatched service staff · 2.2% open-air and in-transit. Every hour is professional work under an employer-signed agreement.
All nine scene familiesThese are not five datasets. Every layer is derived from — and time-locked to — the same footage, so hand pose, occlusion-free geometry, kinematics and robot-ready trajectories all describe the same moment.
Dashed verticals are hardware sync ticks. One take, six aligned streams, no post-hoc registration.
The supervision a policy actually trains on — not the labels a video search engine needs.
Time-synchronised multi-camera capture on the same clip as the egocentric view.
Optical and inertial full-body and hand kinematics, aligned frame-for-frame to the egocentric view.
Metric-scaled meshes and environment scans from the same sites the footage was captured in.
Human demonstrations converted into robot-executable trajectories, physics-replayed before delivery.
Hardware-synchronised by default. Delivered in UMI, LeRobot, RDT, Open-X-Embodiment or MOVAS native — to your cloud, your VPC, or air-gapped media.
One corpus, four places it lands in the Physical AI stack.
Both build-to-order and in-inventory are legitimate models. They are not the same purchase, and they do not carry the same schedule risk. We state which one applies on every line.
Instead of waiting months for collection.
All hours quoted are captured, refined and in storage — not collection capacity. Where a figure refers to work not yet captured, it is labelled as such.
From perceiving the world to executing complex tasks reliably. MOVAS supplies the pre-training corpus, the control-grade supervision on top of it, and the evaluation that proves the lift — then goes back to capture what your model still fails at.
100,000 hours of employer-cleared egocentric video with recoverable ego-motion, weighted toward contact-rich work.
Hand pose, per-joint contact, object 6-DoF and retargeted trajectories — aligned on the same clip.
Every retargeted trajectory replayed in simulation before delivery; infeasible sequences rejected, not shipped.
Because the hours come from sites under contract, we can return to the same line or kitchen and capture the exact failure mode your policy is stuck on.
From raw reality to robot-ready data.
Movas-OS turns raw physical footage into trainable robot experience. Six stages, each containerised and version-pinned — every output traces back to the exact model version that produced it. Available for your own footage as well: video containers, ROS bags and MCAP.
view · VLM · hand visibility
depth · camera path · geometry
MANO · bimanual · smoothing
detect · segment · pose
wrist map · IK · contact-preserving
replay · feasibility · score
raw footage → [01] → [06] → UMI / LeRobot / RDT / Open-X-Embodiment / MOVAS native
AI-assisted workflow · sovereign provenance · format export — up to 70% less manual annotation time than a standard pipeline.
Explore Movas-OSA shared yardstick for embodied policies, not for video understanding. Open tasks, a reproducible harness, sealed splits and a public leaderboard — independently reproduced before anything is published.
Open core · reproducible · academically governedAn individual can consent to their own likeness. They cannot, on their own, grant rights to their employer’s premises, processes, or to colleagues and customers appearing in frame. That is why MOVAS contracts with the employer first and the individual second — and why every clip carries a consent id traceable to both.
Contracted sites, not a contributor pool.
Who has the authority to grant commercial training rights to footage shot inside a business? We publish our answer so your legal team can check it before the first call.
Gates are enforced at ingest.
ISO 27001-aligned across capture, processing and delivery — with an air-gapped option where data cannot leave your boundary.
Licensed by scene family × annotation depth.
Egocentric video is captured from a head-mounted camera worn by a person performing a task, so the viewpoint approximates what a humanoid robot would see doing the same work. Because human and humanoid kinematic chains are similar, policies pre-trained on egocentric video transfer to robots with far less adaptation than third-person video requires — and it is roughly a third to a fifth the cost of teleoperated real-robot collection.
100,000 hours across nine scene families and eighteen sub-classes. Commercial operations, home and domestic work, and food preparation are the three largest, together making up roughly 63 percent. The full breakdown is published on the MOVAS-Ego 100K dataset page.
UMI, LeRobot, RDT and Open-X-Embodiment, plus a native MOVAS format that preserves the full annotation stack. Data can be delivered to your cloud, to an air-gapped environment, or accessed through the Movas-OS platform.
Yes. Contributors sign employer-mediated agreements that cover commercial training use, and the corpus carries full provenance and audit records. This is the practical difference between MOVAS and academic egocentric corpora, which usually carry research-only terms.
Start with training-ready data today. Scale to targeted re-capture when you’re ready.
Request a SubsetData moves robots. MOVAS moves data.