InitAI
InitAI · Human video for robot learning

Twelve times the data, and what it settled.

π0.5 on the ten-hour corpus, against a deliberately harder exam, settles that volume was not this configuration's bottleneck.

2026-09-04 · 3 min read
doc RUN-06 rev A issued

Volume test π0.5 ten-hour corpus 10,000 steps · 1.1 passes

Trained on the ten-hour corpus: 456 episodes / 318,363 frames, built around repetition, with its own exam pre-registered before any training on it. About the data and the exam ↗

Highlights

Overview

The ten-hour corpus is 456 episodes and 318,363 frames, built around repetition, the one structural gap the pilot had identified. Its exam was pre-registered before either model trained on it: four whole held-out blocks, one per manipulation family, one of them an entire recording session withheld outright. This run fine-tuned π0.5 with adapters for ten thousand steps, matched step for step with the ten-hour SmolVLA run so that architecture would be the only difference between them.

The floors moved with the corpus, and that is the first thing to read before any ratio. The pilot exam was continuous-motion tasks, where assume-no-motion is weak (0.1118 on the left hand). This exam includes ladling, where one hand holds the vessel still, and small repetitive shaping, so its floor drops to 0.0687. The same model error reads as a worse ratio here because the denominator is stronger. That is also why no number in this essay shares a chart with a pilot number.

Results

Same corpus, same steps, same exam: the only chart this run can fairly sit in

Left hand — ratio to this exam’s assume-no-motion floor, lower is better

π0.5 (adapters) 1.6
SmolVLA 1.2
Both runs finish above the floor on the full holdout; neither passes the registered bar.

Right hand

π0.5 (adapters) 1.5
SmolVLA 1.2
Same ordering on the right hand.

Where the error concentrates

The scored predictions were kept, so the verdict can be located instead of narrated. On still windows, where the ground-truth hand moves barely at all and doing nothing is nearly perfect, this run sits 14.438× above the floor on the left hand. The single worst cell on the board is the ladling block’s holding hand, at 2.818. Stillness in its purest form is exactly where the model twitches.

Limitations

Single seed, adapters-only fine-tune. Whether a full fine-tune of the same architecture could learn stillness is an untested, priced hypothesis, not a claim. Block-level observations are not passes; the registered bar is defined over the whole holdout. Offline prediction metrics on held-out human video throughout.

Where this goes next

The ten-hour SmolVLA run is the other half of the designed pair: the same corpus and steps under the smaller architecture, and the first like-for-like architecture row the protocol clears. The corpus record is in the data write-up.

Sources

Scored by the programme’s frozen evaluation pipeline against baselines registered before any training; the corpus ships with 31 embedded verification checks, all passing. Full scoring artifacts and the per-block decomposition are available to buyers on request.