InitAI
InitAI · Human video for robot learning

The footage trains.

The first fine-tune of π0.5 on our own pilot footage, scored against three baselines frozen before it existed.

2026-08-17 · 3 min read
doc RUN-01 rev A issued

First fine-tune π0.5 pilot corpus 3,000 steps · 4.9 passes

Trained on the pilot corpus: 39 episodes / 26,040 frames, one take per task — the corpus that proved the line end to end. About the data and the exam ↗

Highlights

Overview

The question this run existed to answer was the first one a buyer would ask: does this footage train at all? The pilot corpus is 39 episodes and 26,040 frames of everyday kitchen work, one complete task per episode, every value flagged measured or carried. Twenty-nine episodes formed the training split; the rest were held out as whole recording blocks, so no test frame shares a scene with a trained one.

π0.5 is a large vision-language-action model and the reference point of this programme. This run fine-tuned it with adapters for 3,000 steps. Nothing about the exam was adjusted afterwards, and nothing has been since.

Results

Two of the three baselines fall on both hands

Left hand — masked prediction error, lower is better

Ignore the image 0.1441
First fine-tune 0.1192
Assume no motion 0.1118
the first fine-tune beats the image-blind baseline and still trails assume-no-motion on the left hand.

Right hand

Ignore the image 0.1114
First fine-tune 0.1026
Assume no motion 0.087
The same ordering on the right hand, which stays the harder of the two throughout the programme.

The baseline that survives is “assume the hands do not move for the next second”, and it is stronger than it sounds. Roughly half of all frames carry a pose held from the frame before, and a held pose is zero motion by definition. A model only gets under that line by genuinely extracting movement from the pictures. Nothing did until the matched-budget run, an architecture change.

The footage

kitchen · the fridge run
The kind of take the corpus is made of: real homes, real mess, our own consented capture.

Limitations

Training loss falling proves the pipeline, not the data; it is never quoted here as evidence of quality. Every score is a single seed, so gaps between runs are “this run beat that run once” until the variance check runs. And these are offline prediction errors on held-out human video, the stage before any hardware.

Where this goes next

The first fine-tune left an obvious lever: nearly half the pilot’s frames carried a hand pose forward instead of measuring one. The detection experiment pulled that lever under controlled conditions and priced it. The corpus itself is described in the data write-up, and we read training results with buyers directly.

Sources

Scored by the programme’s frozen evaluation pipeline, with the pass mark and baselines registered in writing before any run existed. Full scoring artifacts are available to buyers on request. Model and training stack: openpi (Physical Intelligence).