InitAI
InitAI · Human video for robot learning

What we learned training SmolVLA.

Three scored runs. The compact model reads our footage better than anything else we have trained, it rewards longer training, and it holds the strongest scores in the programme.

2026-09-11 · 4 min read
doc NOTE rev A issued

Why this model

SmolVLA is a compact open vision-language-action model of around 450M parameters, a fraction the size of π0.5. We ran it on the identical footage, the identical held-out blocks and the identical frozen baselines, which is the point of it: holding everything else still measures how much of a score belongs to the data and how much to the model reading it. It turned out to matter a great deal.

Every score below is a ratio to its exam’s do-nothing floor. Lower is better, the floor is one, and the two corpora’s exams never share a table.

Main findings

At a matched budget

The matched-budget run held everything from the rebuilt-targets run fixed except the architecture: same footage, same steps, same exam. The left hand came in under the do-nothing line, the first run in the programme to land there on either hand. That carried real information. If swapping the reader moves the score this much, the pilot corpus was never the binding limit.

Ten times the training

The near pass gave the same architecture ten times the training. Held-out error fell steadily and then levelled off while training loss kept falling, and watching for exactly that gap is how we avoid mistaking memorisation for learning.

Left hand, ratio to the pilot exam’s do-nothing floor. Lower is better

Matched budget 0.907
Ten times longer 0.833
The same architecture, ten times the training. The left hand clears the registered pass mark.

Right hand

Matched budget 1
Ten times longer 0.914
The right hand beats every baseline and stops just above the mark.

That run became the first to beat all three frozen baselines outright on both hands. The right hand stopped just above the registered mark, which is why it is recorded as a near pass, and the mark stays where it was written.

On the ten-hour corpus

The ten-hour SmolVLA run put the model on the larger corpus with the corpus, the steps and the exam all held to the same values as the volume test. On identical rows its masked error is 25.9% lower on the left hand and 21.9% lower on the right, which is the first architecture comparison the protocol accepts as like for like.

On the full holdout it finishes above the do-nothing floor, at 1.199 on the left and 1.174 on the right, because this exam charges for stillness and every model so far predicts motion well and stillness badly.

Two cells are worth separating out.

Unseen session, left hand 0.871 A whole recording session held out of training. Under the pass margin, as a block-level observation rather than a registered pass
Moving windows, left hand 0.963 Windows where the hand genuinely travels. At the floor rather than above it

Both runs on this corpus score best on the one held-out block whose entire session was withheld from training, which is the property a buyer asks about first.

Each of these runs is a single seed, so every ordering above is one run ahead of another run once, until the registered seed-variance check runs. The near pass also selected its checkpoint on part of the held-out data, which is optimistic by construction and stated as such in its own essay. The ten-hour run trained for just over one pass over the footage, with loss still falling at cutoff.

The runs behind this

Each one has its own write-up, published whatever it scored:

Next steps

Loss still falling at cutoff is why longer training on this architecture is the registered next step, ahead of any bigger model. The corpus is described in the data write-up.