InitAI
InitAI · Human video for robot learning

What we learned training π0.5.

Four scored runs across both corpora. The footage trains the model cleanly, label quality matters more than label quantity, and twelve times the data did not move the score.

2026-09-11 · 3 min read
doc NOTE rev A issued

Why this model

π0.5 is a large vision-language-action model and the reference point this dataset was shaped for. It has four scored runs here: three on the pilot corpus and one on the ten-hour corpus, adapter fine-tunes throughout.

Every score below is a ratio to its exam’s do-nothing floor, measured on held-out recording blocks against baselines frozen before any training. Lower is better, and the floor is one. The two corpora have different exams, and those exams have different floors by design, so their numbers never share a table.

Main findings

The footage trains

The first fine-tune answered the question that had to come first: footage we filmed ourselves trains a serious model cleanly, and the result beats the image-blind baseline on both hands. Beating that baseline is the difference between a dataset and a pile of video.

Label quality, measured

The detection experiment recovered many more detected hands per frame and tested the obvious hypothesis that more is better. The exam said otherwise. The new detections were located by single-frame estimates with no tracking behind them, and training loss kept falling while held-out error rose, which is what memorising jittery targets looks like. The experiment bought the detector rebuild and the clearest lesson in the programme.

Rebuilt targets, same score

The rebuilt-targets run trained on tracked, side-verified targets that passed blind gates. It erased the regression and landed level with the first fine-tune.

Left hand, ratio to the pilot exam’s do-nothing floor. Lower is better

First fine-tune 1.1
Detection experiment 1.1
Rebuilt targets 1.1
Rebuilt targets recover the ground the detection experiment lost, and land level with the first fine-tune.

Right hand

First fine-tune 1.2
Detection experiment 1.4
Rebuilt targets 1.2
The same pattern on the right hand.

At the time, the flat line read as the corpus being the limit. Two runs later a different architecture on the identical footage scored lower, so the limit was elsewhere. That is why the programme now tests a second model rather than assuming one.

Twelve times the data

The volume test gave the same configuration twelve times the labelled footage and a fresh, deliberately harder exam. The verdict repeated: it beats the image-blind baseline and stays above the do-nothing floor, at 1.619 on the left hand and 1.503 on the right. Training itself was healthy, and its best block was the one whose entire session it never saw. Where it loses is specific: windows where a hand should hold still.

Every π0.5 run here is a single seed and an adapters-only fine-tune. Whether a full fine-tune of the same architecture could learn stillness is a registered, priced hypothesis rather than a claim, and it is one of the queued experiments.

The runs behind this

Each one has its own write-up, published whatever it scored:

Next steps

The stillness weakness now has an address, and the registered next steps sit with the ten-hour corpus: longer training on the smaller architecture first, because it is the cheaper test of the same question. The corpus is described in the data write-up.