Why this data, why now
The strongest open robot-learning models now pretrain on exactly the kind of footage we record, and their builders published the hour counts alongside a simple finding: more human hours, stronger results. The EgoMimic work put it sharpest. In their co-training experiments, an hour of added human-hand video was worth more than an hour of added robot data. None of that data is for sale; the labs that collected it keep it.
So a team that wants more of it has to record it. That is a consent problem, a logistics problem and a quality-control problem before it is a modelling one. Those are the three problems this operation exists to solve.
What the buyers say they need
Their words, not ours, and almost none of it is about size:
- Action labels. NVIDIA, on GR00T N1: human video “do[es] not include explicit action labels”, so their pipeline infers them. Ours are written at export, on every frame, each with a validity flag.
- Hand pose. Apple, explaining why they built their own corpus: “existing large-scale datasets such as Ego4D do not have native hand pose annotations and do not focus on object manipulation”. Here it is a measurement on 73.26% of frames for the left hand and 63.68% for the right.
- A second angle. The HOI4D authors: egocentric video “suffer[s] from more severe occlusion of both human hand and object compared with third-person views”. Every take here is filmed from three synchronised views at once, and the witness views are what the reviewers grade against.
- Curation ahead of volume. Figure, on the data behind their own model: “data quality and consistency matter much more than data quantity”, reporting that a curated set roughly a third the size scored higher, and that a few well-curated hours of demonstration data can carry a dexterous policy. Every episode here is human-reviewed against the witness views before it ships, and the tables below say exactly what is in it.
The rig
Not a lab bench, not a staged set. The archive is filmed where the work actually happens, and it looks like it.
Every clip in this archive is our own capture, filmed with consent.
The process behind every clip is the same three verbs. Film: real people record everyday tasks on a rig of one worn camera and two side views, with consent from everyone in frame. Measure: every frame becomes numbers, both hands in 3D, each value flagged wherever the cameras cannot see. Train: open robot models fine-tune on the result, and every score they earn is published in these essays.
The labels are the part a training pipeline actually consumes. Both hands in 3D on every frame, each value carrying a flag that says whether it was measured or carried forward. On top of that, a person cuts every take and writes its instruction: 408 subtask segments carry a verb, a noun and a sentence, human-adjudicated.

The two corpora
Both read straight out of the shipped artifacts’ own metadata. Their exams differ, so their numbers never share a table.
The ten-hour corpus — current
| Episodes | 456 | every one human-reviewed; 166 distinct tasks, repeated |
| Frames | 318,363 | every frame labelled, built around repetition |
| Hand pose measured | 73.26% L · 63.68% R | share of frames where the pose is a measurement, not a carried value |
| Held out | 55 episodes | four whole blocks, one an entire unseen session, frozen before any training |
| Format | LeRobot v2.1 | loads into the model families’ own training stacks unmodified |
| Verification | 31 checks | all pass; the stamp ships inside the dataset |
The pilot corpus — the one that proved the line
| Episodes | 39 | one complete task each, all human-reviewed |
| Frames | 26,040 | at 30 fps, exact constant rate |
| Numbers per frame | 18 | absolute hand and head pose, with per-frame validity flags |
| Subtask segments | 408 | verb, noun and instruction sentence, human-adjudicated |
Provenance and consent
Consent is captured while the camera is running, not reconstructed afterwards: it covers the person wearing the rig and anyone else who appears in frame, and the paperwork travels with the footage through review and export. Each episode keeps its chain: the masters are never re-encoded, every shipped variant is a derived copy, and the export stamp inside the artifact records which takes were excluded and why. Terms are a live constraint on the public archives rather than a footnote, which is why we describe the mechanism here instead of an adjective.
The exam
There is no industry benchmark for predicting human hand motion from head-mounted video. So we wrote one down before any model existed: whole recording blocks held out, three frozen baselines, and a pass mark of 0.9× the strongest baseline on both hands. It has not moved since, through every run published here. We set it deliberately hard, stricter than any published evaluation we could find for this task, because an easy bar would sell footage rather than measure it.
The strongest baseline is “assume no motion”, and it is brutal on kitchen footage: a hand holding a bowl still is zero motion, frame after frame. The newer exam includes ladling and shaping tasks on purpose, which drops its floor to 0.0687 on the left hand against the pilot exam’s 0.1118. The same model error reads as a worse ratio there because the denominator is stronger. That is also why the two exams’ ratios never share a table.
Our standard
Footage is abundant. Footage a training pipeline can trust is not.
| What we publish | Here | Among ego-video data vendors we surveyed (Sep 2026) |
|---|---|---|
| Pre-registered evaluation, frozen before training | yes | not found |
| Per-frame validity flags on every action value | yes | not found |
| Models trained on the data, scores published either way | yes | one vendor reports internal spot-checks; none publish protocols or results |
| Below-bar results reported as prominently as wins | yes | not found |
- The pass mark was written before the runs and has not moved since.
- Held out as whole recording blocks, never scattered frames, so no test frame shares a scene with a trained one.
- Error is masked to what was genuinely measured. An unmasked metric would look better the worse the data got.
- No comparison crosses a corpus. The two exams’ floors differ by design.
- Results carry the same weight either way. Runs that finish below their bar are published at the same length and the same prominence as the wins.
We would rather be the vendor that grades itself in public than the one with the biggest unverifiable number.
Vendors are unnamed deliberately, and “not found” means not present in their public materials as of September 2026.
The training queue
The model families this corpus is built to reach: three trained and scored, the rest queued behind them. For each, the training stack was read against its official repository before the capture spec was touched.
| Model | Status | Data format | Min GPU | Notes |
|---|---|---|---|---|
| π0.5 open release (openpi) | trained · scored | LeRobot v2.1 (native) | 48 GB+ | The reference model this dataset was shaped for. Four scored runs across both corpora, adapter fine-tunes throughout. |
| SmolVLA Hugging Face | trained · scored | LeRobot v2.1 / v3 | Consumer GPU | Three scored runs, including the closest approach to the pass mark yet recorded here. Consumes the language labels we already write. |
| GR00T N1.7 NVIDIA | trained · scored | LeRobot-style + metadata | 40 GB+ | Pretrained on the largest first-person human-video corpus disclosed to date, and the closest approach yet to the current exam’s floor. One scored run; the H.264 variant and metadata sidecar it needed are now standing exports. |
| RDT-1B Tsinghua | queued | Custom loader | 24 GB | The best structural fit for two-handed data: masked action slots and absolute SI units, which is what our action column already is. Cleanest licence of the set. |
| RDT2 Tsinghua | next capture | WebDataset | 16 GB | Pretrained entirely on human-collected data, strong proof of the market. It accepts only two per-hand wrist views, which the next revision of our rig can add. |
Every entry rides the same masters, and none require re-shooting anything: the LeRobot v3 conversion is one official command, GR00T takes a metadata sidecar plus an H.264 variant, and the relative-motion action column is derived at export. Masters are never re-encoded; every export is a derived copy.
Next steps
Every run essay on the research page trains against this data and this exam. We read training results with buyers directly.
Sources
The model families named above are open projects: openpi (π0.5) · LeRobot + SmolVLA (Hugging Face). Claims about them were verified against these repositories on the dates stated, and licence status is checked per family before any training run.
Quotations in “What the buyers say they need” are their authors’ own published words, read on 11 September 2026: NVIDIA GR00T N1 · Apple EgoDex · HOI4D · Figure, Helix in logistics.