Trained on the pilot corpus: 39 episodes / 26,040 frames, one take per task — the corpus that proved the line end to end. About the data and the exam ↗
Highlights
- Every baseline beaten, both hands. 0.833 left and 0.914 right against the assume-no-motion floor. No earlier run had cleared that floor on either hand.
- One hand passes the registered mark. The bar is 0.9× the best baseline on both hands. The left hand is under it; the right sits just above. Both-hands-or-nothing means this is a near pass, reported as exactly that.
- Saturation was watched, not guessed. Held-out error stopped improving at step 20,000 while training loss kept falling. That gap is why the exam, not the loss curve, decides anything here.
- The mark does not move. It was written before any run existed, set deliberately above the baselines, and is reported unmoved rather than adjusted to fit.
Overview
The matched-budget run had shown the smaller architecture reading the footage better at a small budget. The follow-up question was whether it keeps improving with more passes over the same footage, or whether it memorises. This run trained ten times longer on the identical pilot split and was scored on the identical frozen exam.
Results
The pilot board, settled: five runs, two architectures, one exam
Left hand — ratio to the assume-no-motion floor, lower is better
Right hand
Improvement, then a flat line, watched in the open
Checkpoints were scored every few thousand steps. Held-out error fell steadily, went flat around step 20,000, and stayed flat while training loss kept dropping through 49 passes over the footage. Watching for exactly that divergence is how memorisation is kept out of the results: a loss curve alone would have called the run still improving.
The footage
Limitations
The checkpoint-selection bias above is the headline caveat. Single seed throughout, so the gap to the π0.5 cluster is one run beating another until the variance check runs. And the pilot exam is continuous-motion work, where the do-nothing floor is comparatively weak; the ten-hour corpus’s exam later charged for stillness and reset every expectation.
Where this goes next
This run closed the pilot chapter: the footage demonstrably carries more signal than its strongest baseline when the reader is right. The next question was volume, and the volume test bought twelve times the data to ask it. Protocol and corpus live in the data write-up.
Sources
Scored by the programme’s frozen evaluation pipeline and re-scored on a second machine; the full checkpoint curve was kept alongside the headline score. Full scoring artifacts are available to buyers on request. Model: SmolVLA / LeRobot (Hugging Face).