Why this model
SmolVLA is a compact open vision-language-action model of around 450M parameters, a fraction the size of π0.5. We ran it on the identical footage, the identical held-out blocks and the identical frozen baselines, which is the point of it: holding everything else still measures how much of a score belongs to the data and how much to the model reading it. It turned out to matter a great deal.
Every score below is a ratio to its exam’s do-nothing floor. Lower is better, the floor is one, and the two corpora’s exams never share a table.
Main findings
- At a matched budget it already scored lower than π0.5. The left hand came in under the do-nothing floor, which no run had managed before on either hand.
- Ten times the training nearly cleared the bar. Both hands beat all three frozen baselines outright, and the left hand cleared the registered mark.
- On the ten-hour corpus it won the first clean architecture comparison. Same corpus, same steps, same exam as the volume test, and lower error on both hands.
- Its best cells are the strongest numbers we have. On a session it had never seen, and on windows where hands actually move, it reaches the floor.
At a matched budget
The matched-budget run held everything from the rebuilt-targets run fixed except the architecture: same footage, same steps, same exam. The left hand came in under the do-nothing line, the first run in the programme to land there on either hand. That carried real information. If swapping the reader moves the score this much, the pilot corpus was never the binding limit.
Ten times the training
The near pass gave the same architecture ten times the training. Held-out error fell steadily and then levelled off while training loss kept falling, and watching for exactly that gap is how we avoid mistaking memorisation for learning.
Left hand, ratio to the pilot exam’s do-nothing floor. Lower is better
Right hand
That run became the first to beat all three frozen baselines outright on both hands. The right hand stopped just above the registered mark, which is why it is recorded as a near pass, and the mark stays where it was written.
On the ten-hour corpus
The ten-hour SmolVLA run put the model on the larger corpus with the corpus, the steps and the exam all held to the same values as the volume test. On identical rows its masked error is 25.9% lower on the left hand and 21.9% lower on the right, which is the first architecture comparison the protocol accepts as like for like.
On the full holdout it finishes above the do-nothing floor, at 1.199 on the left and 1.174 on the right, because this exam charges for stillness and every model so far predicts motion well and stillness badly.
Two cells are worth separating out.
Both runs on this corpus score best on the one held-out block whose entire session was withheld from training, which is the property a buyer asks about first.
Each of these runs is a single seed, so every ordering above is one run ahead of another run once, until the registered seed-variance check runs. The near pass also selected its checkpoint on part of the held-out data, which is optimistic by construction and stated as such in its own essay. The ten-hour run trained for just over one pass over the footage, with loss still falling at cutoff.
The runs behind this
Each one has its own write-up, published whatever it scored:
- A second architecture reads the same footage · pilot corpus
- The near pass · pilot corpus
- The smaller model reads the footage better · ten-hour corpus
Next steps
Loss still falling at cutoff is why longer training on this architecture is the registered next step, ahead of any bigger model. The corpus is described in the data write-up.