InitAI
InitAI · Human video for robot learning

The data, the rig, and the exam.

The reference piece behind the run essays: what we film, how a take becomes training data, and the bar every model here is measured against.

2026-09-11 · 9 min read
doc NOTE rev A issued

Why this data, why now

NVIDIA · humanoid foundation model 20,854 h of egocentric human video in pretraining. Not released.
Tsinghua · robotics diffusion transformer 10,000 h of human-collected manipulation data. Also not released.

The strongest open robot-learning models now pretrain on exactly the kind of footage we record, and their builders published the hour counts alongside a simple finding: more human hours, stronger results. The EgoMimic work put it sharpest. In their co-training experiments, an hour of added human-hand video was worth more than an hour of added robot data. None of that data is for sale; the labs that collected it keep it.

So a team that wants more of it has to record it. That is a consent problem, a logistics problem and a quality-control problem before it is a modelling one. Those are the three problems this operation exists to solve.

What the buyers say they need

Their words, not ours, and almost none of it is about size:

Every take is one person, one worn camera and two synchronised witness views filming the workspace from either side. Hands are measured in 3D wherever the cameras can see them, and flagged wherever they cannot.

The rig

Not a lab bench, not a staged set. The archive is filmed where the work actually happens, and it looks like it.

kitchen · the fridge run
kitchen · washing up
repair · marking a switch

Every clip in this archive is our own capture, filmed with consent.

The process behind every clip is the same three verbs. Film: real people record everyday tasks on a rig of one worn camera and two side views, with consent from everyone in frame. Measure: every frame becomes numbers, both hands in 3D, each value flagged wherever the cameras cannot see. Train: open robot models fine-tune on the result, and every score they earn is published in these essays.

one take · three synchronised views
Three cameras per take, timecoded to the frame. The worn view trains the model; the side views keep the labels honest.

The labels are the part a training pipeline actually consumes. Both hands in 3D on every frame, each value carrying a flag that says whether it was measured or carried forward. On top of that, a person cuts every take and writes its instruction: 408 subtask segments carry a verb, a noun and a sentence, human-adjudicated.

Hand and head trajectories across one episode: the eighteen numbers the model is asked to predict, with the one-second prediction window marked.
Hand and head trajectories across one episode: the eighteen numbers the model is asked to predict, with the one-second prediction window marked.

The two corpora

Both read straight out of the shipped artifacts’ own metadata. Their exams differ, so their numbers never share a table.

The ten-hour corpus — current

Episodes456every one human-reviewed; 166 distinct tasks, repeated
Frames318,363every frame labelled, built around repetition
Hand pose measured73.26% L · 63.68% Rshare of frames where the pose is a measurement, not a carried value
Held out55 episodesfour whole blocks, one an entire unseen session, frozen before any training
FormatLeRobot v2.1loads into the model families’ own training stacks unmodified
Verification31 checksall pass; the stamp ships inside the dataset

The pilot corpus — the one that proved the line

Episodes39one complete task each, all human-reviewed
Frames26,040at 30 fps, exact constant rate
Numbers per frame18absolute hand and head pose, with per-frame validity flags
Subtask segments408verb, noun and instruction sentence, human-adjudicated

Consent is captured while the camera is running, not reconstructed afterwards: it covers the person wearing the rig and anyone else who appears in frame, and the paperwork travels with the footage through review and export. Each episode keeps its chain: the masters are never re-encoded, every shipped variant is a derived copy, and the export stamp inside the artifact records which takes were excluded and why. Terms are a live constraint on the public archives rather than a footnote, which is why we describe the mechanism here instead of an adjective.

The exam

There is no industry benchmark for predicting human hand motion from head-mounted video. So we wrote one down before any model existed: whole recording blocks held out, three frozen baselines, and a pass mark of 0.9× the strongest baseline on both hands. It has not moved since, through every run published here. We set it deliberately hard, stricter than any published evaluation we could find for this task, because an easy bar would sell footage rather than measure it.

The strongest baseline is “assume no motion”, and it is brutal on kitchen footage: a hand holding a bowl still is zero motion, frame after frame. The newer exam includes ladling and shaping tasks on purpose, which drops its floor to 0.0687 on the left hand against the pilot exam’s 0.1118. The same model error reads as a worse ratio there because the denominator is stronger. That is also why the two exams’ ratios never share a table.

Our standard

Footage is abundant. Footage a training pipeline can trust is not.

What we publishHereAmong ego-video data vendors we surveyed (Sep 2026)
Pre-registered evaluation, frozen before trainingyesnot found
Per-frame validity flags on every action valueyesnot found
Models trained on the data, scores published either wayyesone vendor reports internal spot-checks; none publish protocols or results
Below-bar results reported as prominently as winsyesnot found

We would rather be the vendor that grades itself in public than the one with the biggest unverifiable number.

Vendors are unnamed deliberately, and “not found” means not present in their public materials as of September 2026.

The training queue

The model families this corpus is built to reach: three trained and scored, the rest queued behind them. For each, the training stack was read against its official repository before the capture spec was touched.

ModelStatusData formatMin GPUNotes
π0.5
open release (openpi)
trained · scoredLeRobot v2.1 (native)48 GB+The reference model this dataset was shaped for. Four scored runs across both corpora, adapter fine-tunes throughout.
SmolVLA
Hugging Face
trained · scoredLeRobot v2.1 / v3Consumer GPUThree scored runs, including the closest approach to the pass mark yet recorded here. Consumes the language labels we already write.
GR00T N1.7
NVIDIA
trained · scoredLeRobot-style + metadata40 GB+Pretrained on the largest first-person human-video corpus disclosed to date, and the closest approach yet to the current exam’s floor. One scored run; the H.264 variant and metadata sidecar it needed are now standing exports.
RDT-1B
Tsinghua
queuedCustom loader24 GBThe best structural fit for two-handed data: masked action slots and absolute SI units, which is what our action column already is. Cleanest licence of the set.
RDT2
Tsinghua
next captureWebDataset16 GBPretrained entirely on human-collected data, strong proof of the market. It accepts only two per-hand wrist views, which the next revision of our rig can add.

Every entry rides the same masters, and none require re-shooting anything: the LeRobot v3 conversion is one official command, GR00T takes a metadata sidecar plus an H.264 variant, and the relative-motion action column is derived at export. Masters are never re-encoded; every export is a derived copy.

Next steps

Every run essay on the research page trains against this data and this exam. We read training results with buyers directly.

Sources

The model families named above are open projects: openpi (π0.5) · LeRobot + SmolVLA (Hugging Face). Claims about them were verified against these repositories on the dates stated, and licence status is checked per family before any training run.

Quotations in “What the buyers say they need” are their authors’ own published words, read on 11 September 2026: NVIDIA GR00T N1 · Apple EgoDex · HOI4D · Figure, Helix in logistics.