InitAI
InitAI · Human video for robot learning

The smaller model reads the footage better.

SmolVLA against π0.5 with corpus, steps and exam all held fixed. It wins every like-for-like comparison the protocol allows, and the registered bar still stands.

2026-09-05 · 4 min read
doc RUN-07 rev A issued

Architecture pair SmolVLA ten-hour corpus 10,000 steps · 1.1 passes

Trained on the ten-hour corpus: 456 episodes / 318,363 frames, built around repetition, with its own exam pre-registered before any training on it. About the data and the exam ↗

Highlights

Overview

The volume test and this run were designed as a pair. Identical footage, identical split, identical step budget, identical frozen exam; the architecture is the only variable. Every earlier cross-model observation in the programme had to be hedged because something else differed too. This pair is the first row that goes into a matrix rather than a narrative.

The run also carries a quieter cross-corpus observation, stated in prose because the two exams’ floors differ by design: this run’s raw left-hand error, 0.0824, is lower than the same architecture’s best pilot-corpus run, 0.0931, on a harder exam at a fifth of the steps. The corpus itself improved the better model.

Results

The designed pair, on the only chart that compares them fairly

Left hand — ratio to this exam’s assume-no-motion floor, lower is better

π0.5 (adapters) 1.6
SmolVLA 1.2
The smaller model reads the same footage with lower error, and both runs stay above the floor on the full holdout.

Right hand

π0.5 (adapters) 1.5
SmolVLA 1.2
Same ordering on the right hand.

The remaining error, decomposed to its address

Split by how far the ground-truth hand actually travels in the predicted second, the picture is specific. Where there is real motion, the left hand reads 0.963 and the right 1.091: at the do-nothing floor or just above it. On still windows the left hand sits 8.216× above the floor, because a still hand makes doing nothing nearly perfect and the model twitches. About 32.2% of left-hand windows in this exam are still or slow, mass the pilot exam barely carried, which is why this weakness was invisible until this corpus measured it.

The unseen session is its best block

Both runs of the pair score best on the one held-out block whose entire session was withheld from training. For this run that reads 0.871 left and 1.011 right, the left hand under the 0.9× margin. That is a block-level observation, not a registered pass: the bar spans the whole holdout and both hands, and softening it is off the table. Errors that concentrate in a named skill rather than in unfamiliar footage are the pattern a data buyer wants to see.

Both sides of the pair ran a single seed, so the architecture gap is one run ahead of another run once, until the registered seed-variance check runs. Steps were held to the volume test’s budget, just over one pass over the footage against the roughly thirty-three of the August near pass, with loss still falling at cutoff. That is exactly why longer training is the registered next step.

The footage

one take · three synchronised views
The rig’s actual output. The worn view trains the model; the witness views keep the labels honest.

Limitations

Every caveat stated beside its number above, plus the standing frame: these are offline prediction errors on held-out human video, the stage before hardware, and the stage where dataset quality is actually decided. No cross-corpus number in this essay appears in any chart.

Where this goes next

The registered next step is the same configuration trained longer, since loss had not flattened at cutoff. The stillness gap now has three priced levers against it rather than a narrative. The corpus, the exam and every run in one place: the research page.

Sources

Scored by the programme’s frozen evaluation pipeline against baselines registered before any training; the per-block decomposition reproduces every headline cell before splitting anything. Full scoring artifacts are available to buyers on request. The model families are open projects: SmolVLA / LeRobot (Hugging Face) · openpi (π0.5).