Trained on the ten-hour corpus: 456 episodes / 318,363 frames, built around repetition, with its own exam pre-registered before any training on it. About the data and the exam ↗
Highlights
- The first honest architecture row. Same corpus, same ten thousand steps, same frozen exam as the volume test. On identical rows, this run’s masked error is 25.9% lower on the left hand and 21.9% lower on the right.
- The registered bar still stands, with this run at 1.199 left and 1.174 right of the do-nothing floor on the full holdout. We publish the row either way; a scoreboard with missing rows is not a scoreboard.
- On the fully unseen session, the left hand passes the margin. An entire recording session was withheld from training; on it, this run reads 0.871 left. A block-level observation, not a registered pass, and the strongest single cell the programme has produced.
- Where hands move, it already beats doing nothing. On windows with real motion the left hand reads 0.963. The entire headline gap lives in stillness.
- The exam is charging for a skill buyers need. An exam a model cannot yet pass, measuring exactly the behaviour a training pipeline would collide with, is the dataset doing its job.
Overview
The volume test and this run were designed as a pair. Identical footage, identical split, identical step budget, identical frozen exam; the architecture is the only variable. Every earlier cross-model observation in the programme had to be hedged because something else differed too. This pair is the first row that goes into a matrix rather than a narrative.
The run also carries a quieter cross-corpus observation, stated in prose because the two exams’ floors differ by design: this run’s raw left-hand error, 0.0824, is lower than the same architecture’s best pilot-corpus run, 0.0931, on a harder exam at a fifth of the steps. The corpus itself improved the better model.
Results
The designed pair, on the only chart that compares them fairly
Left hand — ratio to this exam’s assume-no-motion floor, lower is better
Right hand
The remaining error, decomposed to its address
Split by how far the ground-truth hand actually travels in the predicted second, the picture is specific. Where there is real motion, the left hand reads 0.963 and the right 1.091: at the do-nothing floor or just above it. On still windows the left hand sits 8.216× above the floor, because a still hand makes doing nothing nearly perfect and the model twitches. About 32.2% of left-hand windows in this exam are still or slow, mass the pilot exam barely carried, which is why this weakness was invisible until this corpus measured it.
The unseen session is its best block
Both runs of the pair score best on the one held-out block whose entire session was withheld from training. For this run that reads 0.871 left and 1.011 right, the left hand under the 0.9× margin. That is a block-level observation, not a registered pass: the bar spans the whole holdout and both hands, and softening it is off the table. Errors that concentrate in a named skill rather than in unfamiliar footage are the pattern a data buyer wants to see.
Both sides of the pair ran a single seed, so the architecture gap is one run ahead of another run once, until the registered seed-variance check runs. Steps were held to the volume test’s budget, just over one pass over the footage against the roughly thirty-three of the August near pass, with loss still falling at cutoff. That is exactly why longer training is the registered next step.
The footage
Limitations
Every caveat stated beside its number above, plus the standing frame: these are offline prediction errors on held-out human video, the stage before hardware, and the stage where dataset quality is actually decided. No cross-corpus number in this essay appears in any chart.
Where this goes next
The registered next step is the same configuration trained longer, since loss had not flattened at cutoff. The stillness gap now has three priced levers against it rather than a narrative. The corpus, the exam and every run in one place: the research page.
Sources
Scored by the programme’s frozen evaluation pipeline against baselines registered before any training; the per-block decomposition reproduces every headline cell before splitting anything. Full scoring artifacts are available to buyers on request. The model families are open projects: SmolVLA / LeRobot (Hugging Face) · openpi (π0.5).