Trained on the pilot corpus: 39 episodes / 26,040 frames, one take per task — the corpus that proved the line end to end. About the data and the exam ↗
Highlights
- First time under the line. The left hand scores 0.907 against the assume-no-motion floor. Every earlier run, on either hand, sat above it.
- Architecture was the only difference. Deliberately matched to the rebuilt-targets run step for step: same footage, same split, same budget, same frozen exam.
- The bottleneck moved. the rebuilt-targets run’s flat line had been read as the corpus’s ceiling. The same corpus under a different model scored lower, so the ceiling belonged to the model at this scale.
- A fraction of the size. SmolVLA is a compact open vision-language-action model, around 450M parameters against π0.5’s billions.
Overview
Three π0.5 runs had clustered tightly, and a cluster invites a conclusion: perhaps this is what the footage supports. Before accepting that, the protocol demanded one control. Train a different architecture on exactly the same inputs and see whether the cluster is a property of the data or of the reader. SmolVLA consumes the same LeRobot-format footage and the same language labels the corpus already carries, which made it the clean candidate.
Results
The cluster belonged to the model, not the data
Left hand — ratio to the assume-no-motion floor, lower is better
Right hand
At 0.907 left and 1.027 right, this budget-matched run rewrote the file’s own working theory. The interesting question stopped being “is the corpus too small” and became “how much is left in this footage that a better reader could extract”. The near pass answers some of that. What the run establishes is narrower than a head-to-head verdict, and still the useful part: one seed each side, the run-to-run spread from seed alone not yet measured here, and the pilot corpus demonstrably not the ceiling.
Limitations
Single seed each side. One exam base behind the pilot figures. Offline prediction error on held-out human video, the stage before hardware.
Where this goes next
The same model trained ten times longer is the near pass, the programme’s closest approach to its own pass mark. Corpus and protocol live in the data write-up.
Sources
Scored by the programme’s frozen evaluation pipeline and re-scored on a second machine, reproducing the verdicts to five decimal places. Full scoring artifacts are available to buyers on request. Model: SmolVLA / LeRobot (Hugging Face).