One corpus. Every layer of the same moment.
Five layers. One clip. Hardware-aligned.
These are not five datasets. Every layer is derived from — and time-locked to — the same footage, so hand pose, occlusion-free geometry, kinematics and robot-ready trajectories all describe the same moment. A catalogue of unrelated datasets cannot do that.
Dashed verticals are hardware sync ticks. One take, six aligned streams, no post-hoc registration.
L1–L5 Control Annotation
The supervision a policy actually trains on — not the labels a video search engine needs.
4D Multi-View Ground Truth
Time-synchronised multi-camera capture on the same clip as the egocentric view.
MoCap Kinematics
Optical and inertial full-body and hand kinematics, aligned frame-for-frame to the egocentric view.
3D Scene & Object Assets
Metric-scaled meshes and environment scans from the same sites the footage was captured in.
Human2Robot Retargeting
Human demonstrations converted into robot-executable trajectories, physics-replayed before delivery.
Layers are licensing depth, not separate products. Coverage varies by scene family and is stated on each scene family page and on the data card.