The two corpora
The same capture-and-label line has shipped two corpora so far: a pilot of one take per task, and the ten-hour corpus built around repetition. Between them, every measure of the data moved up.
Label coverage
The number that matters most is not hours. It is the share of frames where the hand pose is a genuine measurement rather than a value carried forward, because that share is what a training loss can actually trust.
The bar is the current generation; the tick marks where the first generation stood. The detector line was rebuilt between them and verified on a blind, hand-marked answer key.
That rebuild was not free. It was paid for by the detection experiment, the programme’s most instructive experiment, which showed the model pays for position quality, not detection count.
What each run taught
- The first fine-tune proved our own footage trains at all, and that the model reads the frames rather than memorising an average.
- The detection experiment taught that target quality outranks target quantity, and bought the detector rebuild above.
- The rebuilt-targets run fixed the targets, did not move the score, and pointed the question away from the corpus.
- The matched-budget run swapped the architecture with everything else held fixed, and the score moved. The bottleneck was the reader, not the footage.
- The near pass trained that architecture ten times longer and came within touching distance of the registered pass mark.
- The volume test asked what twelve times the footage buys π0.5 and settled the volume question for good. A settled answer at scale is an asset.
- The ten-hour SmolVLA run put the smaller model on the bigger corpus and produced the strongest cells the programme has recorded, including one under the pass margin on a fully unseen session.
- The GR00T run added a third family, the one pretrained on first-person human video like ours, and its right hand came closer to the exam’s floor than any run before it.
What the scores did
Our strongest training numbers to date came from the newest corpus, at a fifth of the training steps of the run before it. That is the trajectory a data operation wants to show: not one lucky number, but each generation of footage carrying the models further under an exam that never moves. The comparison across the two corpora stays in prose deliberately, because their exams’ floors differ by design; the like-for-like tables live inside the run essays.
The next round of footage is already recorded and queued for the same line. We read training results with buyers directly.