Trained on the ten-hour corpus: 456 episodes / 318,363 frames, built around repetition, with its own exam pre-registered before any training on it. About the data and the exam ↗
Highlights
- The volume question is settled for this configuration. Twelve times the labelled footage and the signature does not change: above the do-nothing floor on both hands, at 1.619 left and 1.503 right.
- The exam got harder on purpose. The new floor is about 38.5% stronger than the pilot exam’s on the left hand, because this corpus deliberately includes tasks that hold a hand still.
- Training was healthy. Loss fell steadily with no instability; the corpus and the pipeline both did their job, and the open gap is a modelling skill.
- Its best block is the unseen one. Of the four held-out blocks, this run scores best on the session it never saw a frame of, at 1.071 left. The errors are a skill gap, not unfamiliarity.
Overview
The ten-hour corpus is 456 episodes and 318,363 frames, built around repetition, the one structural gap the pilot had identified. Its exam was pre-registered before either model trained on it: four whole held-out blocks, one per manipulation family, one of them an entire recording session withheld outright. This run fine-tuned π0.5 with adapters for ten thousand steps, matched step for step with the ten-hour SmolVLA run so that architecture would be the only difference between them.
The floors moved with the corpus, and that is the first thing to read before any ratio. The pilot exam was continuous-motion tasks, where assume-no-motion is weak (0.1118 on the left hand). This exam includes ladling, where one hand holds the vessel still, and small repetitive shaping, so its floor drops to 0.0687. The same model error reads as a worse ratio here because the denominator is stronger. That is also why no number in this essay shares a chart with a pilot number.
Results
Same corpus, same steps, same exam: the only chart this run can fairly sit in
Left hand — ratio to this exam’s assume-no-motion floor, lower is better
Right hand
Where the error concentrates
The scored predictions were kept, so the verdict can be located instead of narrated. On still windows, where the ground-truth hand moves barely at all and doing nothing is nearly perfect, this run sits 14.438× above the floor on the left hand. The single worst cell on the board is the ladling block’s holding hand, at 2.818. Stillness in its purest form is exactly where the model twitches.
Limitations
Single seed, adapters-only fine-tune. Whether a full fine-tune of the same architecture could learn stillness is an untested, priced hypothesis, not a claim. Block-level observations are not passes; the registered bar is defined over the whole holdout. Offline prediction metrics on held-out human video throughout.
Where this goes next
The ten-hour SmolVLA run is the other half of the designed pair: the same corpus and steps under the smaller architecture, and the first like-for-like architecture row the protocol clears. The corpus record is in the data write-up.
Sources
Scored by the programme’s frozen evaluation pipeline against baselines registered before any training; the corpus ships with 31 embedded verification checks, all passing. Full scoring artifacts and the per-block decomposition are available to buyers on request.