InitAI
InitAI · Human video for robot learning

A second architecture reads the same footage.

SmolVLA, matched step for step with the rebuilt-targets run, becomes the first run under the do-nothing line on either hand.

2026-08-25 · 2 min read
doc RUN-04 rev A issued

Matched budget SmolVLA pilot corpus 3,000 steps · 4.9 passes

Trained on the pilot corpus: 39 episodes / 26,040 frames, one take per task — the corpus that proved the line end to end. About the data and the exam ↗

Highlights

Overview

Three π0.5 runs had clustered tightly, and a cluster invites a conclusion: perhaps this is what the footage supports. Before accepting that, the protocol demanded one control. Train a different architecture on exactly the same inputs and see whether the cluster is a property of the data or of the reader. SmolVLA consumes the same LeRobot-format footage and the same language labels the corpus already carries, which made it the clean candidate.

Results

The cluster belonged to the model, not the data

Left hand — ratio to the assume-no-motion floor, lower is better

π0.5 1.1
SmolVLA 0.907
Identical inputs and budget: the smaller model crosses under the do-nothing floor on the left hand.

Right hand

π0.5 1.2
SmolVLA 1
The right hand lands just above the floor at this budget; the near pass takes it under.

At 0.907 left and 1.027 right, this budget-matched run rewrote the file’s own working theory. The interesting question stopped being “is the corpus too small” and became “how much is left in this footage that a better reader could extract”. The near pass answers some of that. What the run establishes is narrower than a head-to-head verdict, and still the useful part: one seed each side, the run-to-run spread from seed alone not yet measured here, and the pilot corpus demonstrably not the ceiling.

Limitations

Single seed each side. One exam base behind the pilot figures. Offline prediction error on held-out human video, the stage before hardware.

Where this goes next

The same model trained ten times longer is the near pass, the programme’s closest approach to its own pass mark. Corpus and protocol live in the data write-up.

Sources

Scored by the programme’s frozen evaluation pipeline and re-scored on a second machine, reproducing the verdicts to five decimal places. Full scoring artifacts are available to buyers on request. Model: SmolVLA / LeRobot (Hugging Face).