One corpus. Every layer of the same moment.
Real people. Real tasks. Real environments.
MOVAS-Ego 100K is a single corpus, not a catalogue of unrelated datasets. Every layer below is derived from — and hardware-aligned to — the same footage, so hand pose, contact, 3D structure and robot-ready trajectories all refer to the same moment in time.
Licensed by the environment your policy has to work in.
80.6% commercial and industrial premises · 17.2% residences where the operator is employer-dispatched service staff · 2.2% open-air and in-transit. Every hour is professional work under an employer-signed agreement.
Commercial Operations
Storefront service, customer interaction, counter operations and stock handling in live retail and service sites.
Food Preparation
Cutting, mixing, frying, dough work and plating on deformable materials, in live commercial kitchens.
Domestic & Care Services
Professional housekeeping, laundry, organising and personal care delivered in residences by employer-dispatched staff.
Industrial & Production
Assembly, fastening, routing, dispensing, tool use and fault handling inside live production facilities, plus warehouse and material handling.
General Retail
Shelf replenishment, SKU picking, inventory count, sorting and packing across store and back-of-house areas.
F&B Storefront
Order taking, plating, serving, table turnover and POS operation during real service hours.
Specialised Services
Repair, appliance service, automotive maintenance and personal-care procedures performed by trained practitioners.
Handcraft & Fine Assembly
Fine bimanual assembly, tool use and precision insertion by skilled operators.
Market, Agri & Mobility
Open-air markets, farming operations and in-transit egocentric capture across varied lighting and terrain.
Five layers, on the same clip.
Licensing depth, not separate products. Coverage varies by scene family and is stated on every family page and data card.
Dashed verticals are hardware sync ticks. One take, six aligned streams, no post-hoc registration.
L1–L5 Control Annotation
The supervision a policy actually trains on — not the labels a video search engine needs.
4D Multi-View Ground Truth
Time-synchronised multi-camera capture on the same clip as the egocentric view.
MoCap Kinematics
Optical and inertial full-body and hand kinematics, aligned frame-for-frame to the egocentric view.
3D Scene & Object Assets
Metric-scaled meshes and environment scans from the same sites the footage was captured in.
Human2Robot Retargeting
Human demonstrations converted into robot-executable trajectories, physics-replayed before delivery.
Start with existing inventory.
Instead of waiting months for collection.
All hours quoted are captured, refined and in storage — not collection capacity. Where a figure refers to work not yet captured, it is labelled as such.