Trained on the pilot corpus: 39 episodes / 26,040 frames, one take per task — the corpus that proved the line end to end. About the data and the exam ↗
Highlights
- The line holds end to end. Format, loader, optimiser: training loss falls 86% over the run, monotone, with no instability.
- Both hands beat the image-blind baseline. A model that outperforms “predict the training-set average” without ever seeing the frame is impossible; beating that baseline proves the pictures are being read.
- Assume-no-motion stays ahead, at 1.066 of its floor on the left hand and 1.18 on the right. On kitchen footage, doing nothing is a strong opponent.
- The exam predates the run. Held-out blocks, three frozen baselines and the pass mark were all written down before any model existed.
Overview
The question this run existed to answer was the first one a buyer would ask: does this footage train at all? The pilot corpus is 39 episodes and 26,040 frames of everyday kitchen work, one complete task per episode, every value flagged measured or carried. Twenty-nine episodes formed the training split; the rest were held out as whole recording blocks, so no test frame shares a scene with a trained one.
π0.5 is a large vision-language-action model and the reference point of this programme. This run fine-tuned it with adapters for 3,000 steps. Nothing about the exam was adjusted afterwards, and nothing has been since.
Results
Two of the three baselines fall on both hands
Left hand — masked prediction error, lower is better
Right hand
The baseline that survives is “assume the hands do not move for the next second”, and it is stronger than it sounds. Roughly half of all frames carry a pose held from the frame before, and a held pose is zero motion by definition. A model only gets under that line by genuinely extracting movement from the pictures. Nothing did until the matched-budget run, an architecture change.
The footage
Limitations
Training loss falling proves the pipeline, not the data; it is never quoted here as evidence of quality. Every score is a single seed, so gaps between runs are “this run beat that run once” until the variance check runs. And these are offline prediction errors on held-out human video, the stage before any hardware.
Where this goes next
The first fine-tune left an obvious lever: nearly half the pilot’s frames carried a hand pose forward instead of measuring one. The detection experiment pulled that lever under controlled conditions and priced it. The corpus itself is described in the data write-up, and we read training results with buyers directly.
Sources
Scored by the programme’s frozen evaluation pipeline, with the pass mark and baselines registered in writing before any run existed. Full scoring artifacts are available to buyers on request. Model and training stack: openpi (Physical Intelligence).