Why this model
π0.5 is a large vision-language-action model and the reference point this dataset was shaped for. It has four scored runs here: three on the pilot corpus and one on the ten-hour corpus, adapter fine-tunes throughout.
Every score below is a ratio to its exam’s do-nothing floor, measured on held-out recording blocks against baselines frozen before any training. Lower is better, and the floor is one. The two corpora have different exams, and those exams have different floors by design, so their numbers never share a table.
Main findings
- The footage trains. The first fine-tune beat the image-blind baseline on both hands, which is the evidence that the model reads the pictures rather than memorising an average.
- Label quality outranks label quantity. Recovering many more hands per frame made the model worse, not better, because the new positions were single-frame estimates with no tracking behind them.
- Rebuilding the targets did not move the score. At the time we read that as the corpus being the limit. It was not.
- Twelve times the footage did not move it either. Volume is not this configuration’s bottleneck, and repeating the same signature at twelve times the scale is what settles it.
The footage trains
The first fine-tune answered the question that had to come first: footage we filmed ourselves trains a serious model cleanly, and the result beats the image-blind baseline on both hands. Beating that baseline is the difference between a dataset and a pile of video.
Label quality, measured
The detection experiment recovered many more detected hands per frame and tested the obvious hypothesis that more is better. The exam said otherwise. The new detections were located by single-frame estimates with no tracking behind them, and training loss kept falling while held-out error rose, which is what memorising jittery targets looks like. The experiment bought the detector rebuild and the clearest lesson in the programme.
Rebuilt targets, same score
The rebuilt-targets run trained on tracked, side-verified targets that passed blind gates. It erased the regression and landed level with the first fine-tune.
Left hand, ratio to the pilot exam’s do-nothing floor. Lower is better
Right hand
At the time, the flat line read as the corpus being the limit. Two runs later a different architecture on the identical footage scored lower, so the limit was elsewhere. That is why the programme now tests a second model rather than assuming one.
Twelve times the data
The volume test gave the same configuration twelve times the labelled footage and a fresh, deliberately harder exam. The verdict repeated: it beats the image-blind baseline and stays above the do-nothing floor, at 1.619 on the left hand and 1.503 on the right. Training itself was healthy, and its best block was the one whose entire session it never saw. Where it loses is specific: windows where a hand should hold still.
Every π0.5 run here is a single seed and an adapters-only fine-tune. Whether a full fine-tune of the same architecture could learn stillness is a registered, priced hypothesis rather than a claim, and it is one of the queued experiments.
The runs behind this
Each one has its own write-up, published whatever it scored:
- The footage trains · pilot corpus
- Target quality beats target quantity · pilot corpus
- Rebuilt targets, same score · pilot corpus
- Twelve times the data, and what it settled · ten-hour corpus
Next steps
The stillness weakness now has an address, and the registered next steps sit with the ten-hour corpus: longer training on the smaller architecture first, because it is the cheaper test of the same question. The corpus is described in the data write-up.