Employer-Cleared Egocentric Data for Physical AI

MOVAS-Ego 100K — Egocentric Video Dataset for Physical AI

100,000 hours of head-mounted first-person video captured during professional work — every hour under an employer-signed agreement, with hand pose, contact and retargeted trajectories aligned on the same clip.

Hours in inventory 100,000+
Rights 100% employer-signed
Capture rig Head-mounted, EIS-off
Aligned layers 5 per clip

Overview

MOVAS-Ego 100K is the largest commercially-licensed egocentric corpus assembled specifically for Physical AI. Every hour is captured from a head-mounted rig worn by a worker performing real, unscripted tasks in a live environment — not staged demonstrations and not scraped web video.

Electronic image stabilisation and lens-distortion correction are disabled at capture time. This is deliberate: EIS silently rewrites the relationship between camera motion and scene motion, destroying the ego-motion signal that downstream SLAM, hand-pose and retargeting stages depend on. It cannot be corrected after the fact.

The corpus is weighted toward contact-rich, bimanual, long-horizon work. Food preparation, assembly and line work, fine assembly and specialised services together account for 31,428 hours — the slice where manipulation policies fail first, and the slice that generic crowdsourced or web-scraped video under-represents.

100,000+
Hours in inventory
100% employer-signed
Rights
Head-mounted, EIS-off
Capture rig
5 per clip
Aligned layers

Scene Distribution

The corpus is deliberately unbalanced. Commercial and domestic environments lead because they generate the highest density of contact-rich, bimanual, long-horizon manipulation per recorded hour, and because they are the two settings where humanoid deployment is nearest term. Industrial operations were added as a dedicated family to serve line-side and warehouse policies, which have materially different object statistics from retail.

Commercial Operations — 25.6% Food Preparation — 17.8% Domestic & Care Services — 17.2% Industrial & Production — 11% General Retail — 9.8% F&B Storefront — 7.5% Specialised Services — 5.4% Handcraft & Fine Assembly — 3.5% Market, Agri & Mobility — 2.2% 100K HOURS
  • Commercial Operations 25.6%
  • Food Preparation 17.8%
  • Domestic & Care Services 17.2%
  • Industrial & Production 11%
  • General Retail 9.8%
  • F&B Storefront 7.5%
  • Specialised Services 5.4%
  • Handcraft & Fine Assembly 3.5%
  • Market, Agri & Mobility 2.2%

Full 18 sub-class breakdown

MOVAS-Ego 100K scene class distribution by duration
Sub-ClassScene FamilyDurationShare
Storefront Service & Counter OperationsCommercial Operations25,636 h25.64%
Food Preparation & CookingFood Preparation17,770 h17.77%
Shelf Replenishment & InventoryGeneral Retail9,823 h9.82%
Housekeeping & TidyingDomestic & Care Services8,200 h8.2%
Order, Plating & Table ServiceF&B Storefront7,466 h7.47%
Kitchen & DishwashingDomestic & Care Services5,100 h5.1%
Assembly & Line WorkIndustrial & Production4,800 h4.8%
Laundry & Textile HandlingDomestic & Care Services3,900 h3.9%
Sorting & PackingIndustrial & Production3,600 h3.6%
Handcraft & Fine AssemblyHandcraft & Fine Assembly3,461 h3.46%
Home Repair & Appliance ServiceSpecialised Services2,610 h2.61%
Warehouse & Material HandlingIndustrial & Production2,600 h2.6%
Wet Market & BazaarMarket, Agri & Mobility1,580 h1.58%
Professional RepairSpecialised Services1,523 h1.52%
Automotive ServiceSpecialised Services1,036 h1.04%
Agriculture & FarmingMarket, Agri & Mobility356 h0.36%
Transportation & MobilityMarket, Agri & Mobility311 h0.31%
Beauty & Personal CareSpecialised Services228 h0.23%
Total100,000 h100%

Proportions derived from the MOVAS internal data catalogue and normalised to the 100,000-hour production corpus. Sub-classes are grouped into nine top-level scene families; licensing is available at either level.

Data Structure

Directory layout as delivered. Every clip is self-contained: no cross-referencing required to train on a single sample.

movas-ego-100k/
├── clips/
│   └── {scene_family}/{clip_uuid}/
│       ├── video.mp4              # 1080p30, H.264 CRF 18, EIS off
│       ├── imu.npz                # 200 Hz 9-axis, hw-synced
│       └── meta.json              # scene, duration, rig, consent id
├── hand/
│   └── {clip_uuid}/
│       ├── mano_left.npz          # [T, 51] MANO params
│       ├── mano_right.npz
│       └── contact.json           # per-frame, per-joint contact
├── objects/
│   └── {clip_uuid}/
│       ├── {obj_id}_pose.npz      # [T, 4, 4] 6-DoF
│       └── {obj_id}_mask.npz
├── retarget/
│   └── {clip_uuid}/
│       ├── shadow_hand.npz
│       ├── inspire_hand.npz
│       ├── allegro.npz
│       └── franka_gripper.npz
├── annotations/
│   └── {clip_uuid}.json           # L1-L5 label stack
└── manifest.parquet               # index + per-clip quality score

Annotation Schema

L1

Task

Task identity and success/failure outcome.

L2

Phase

Segmented action phases: approach, pre-grasp, grasp, transport, manipulate, release.

L3

Object

Open-vocabulary object identity, material and state transitions.

L4

Keyframe

1 fps dense keyframes plus every phase-boundary frame.

L5

Contact

Per-hand, per-joint contact points and duration derived from MANO surface distance.

Quality Gates

Published thresholds, enforced at ingest. Records that fail are rejected or flagged in the manifest — never silently included.

GateThresholdMethod
Hand visibility≥ 50% of framesClips below threshold are rejected at ingest.
IMU–video drift< 1 ms / 60 sHardware timestamp sync, verified per clip.
Reprojection error< 3 px @ 1080pGeometric consistency gate after 4D reconstruction.
Composite score≥ 70 / 100Weighted across reconstruction, hand pose, tracking, retarget feasibility, sim replay.

Formats & Access

Export formats

UMILeRobotRDTOpen-X-EmbodimentMOVAS-Ego native

Modalities included

RGB video 1080p30IMU 200 Hz hardware-syncedHand pose (MANO)Per-frame contact stateObject 6-DoFRetargeted trajectories

Cloud delivery

Direct to your S3, GCS or OSS bucket. Manifest-driven incremental sync.

Air-gapped

Physical media transfer for regulated or offline environments.

Via Movas-OS

Stream and re-annotate in place through the Movas-OS platform.

Use Cases

VLA pre-training

Large-scale visuomotor pre-training where scale and action diversity dominate.

World-model pre-training

Next-frame and next-state prediction grounded in real physical dynamics, with recoverable ego-motion.

Human-to-robot transfer

Retargeted trajectories give direct supervision for dexterous and parallel-jaw end effectors.

Affordance & contact learning

L5 contact annotations supervise where and how objects are grasped.

Related Datasets

FAQ

MOVAS-Ego 100K — Frequently Asked Questions

Ego4D and Ego-Exo4D are research corpora built for perception benchmarks — activity recognition, episodic memory and captioning. MOVAS-Ego 100K is built for control. Every clip carries hand pose, contact state and trajectories already retargeted onto common end effectors, capture-side EIS is disabled so ego-motion remains recoverable, and the corpus is cleared for commercial training use rather than research-only.

Inventory. The hours are captured, refined and in storage today. An evaluation subset ships in three business days and a full scene family in ten. Where we quote a figure for work not yet captured — targeted re-capture, for example — we label it as such.

The operating entity of each site, through an employer agreement, with individual consent signed separately by each person filmed. An individual cannot on their own grant rights to their employer’s premises, processes or to colleagues appearing in frame. The full three-layer structure is published on our data rights page.

Yes, at firmware level on every rig. EIS applies a per-frame homography that decouples pixel motion from true camera motion. Any downstream stage that estimates ego-motion — SLAM, hand-pose lifting, object 6-DoF tracking — inherits that error, and it cannot be inverted after the fact.

Five-finger anthropomorphic hands (Shadow, Inspire), four-finger hands (Allegro), parallel-jaw grippers (Franka, UR5) and three-finger underactuated hands (Robotiq 3F). For other embodiments, send a URDF and we generate a custom retarget from the source MANO parameters.

Yes. The corpus is licensed by scene family and by annotation depth. A typical first engagement is a single scene family at L1–L3, expanding to L5 and additional families once training value is demonstrated.

Request Dataset Access

Send us the task you are training for and we will scope the smallest subset that moves your metric.

Request a Subset

Data moves robots. MOVAS moves data.