InitAI
InitAI · Human video for robot learning

The near pass.

Trained ten times longer, SmolVLA beats every baseline on both hands and stops just above the registered mark.

2026-08-25 · 3 min read
doc RUN-05 rev A issued

Near pass SmolVLA pilot corpus 30,000 steps · 49.1 passes

Trained on the pilot corpus: 39 episodes / 26,040 frames, one take per task — the corpus that proved the line end to end. About the data and the exam ↗

Highlights

Overview

The matched-budget run had shown the smaller architecture reading the footage better at a small budget. The follow-up question was whether it keeps improving with more passes over the same footage, or whether it memorises. This run trained ten times longer on the identical pilot split and was scored on the identical frozen exam.

Results

The pilot board, settled: five runs, two architectures, one exam

Left hand — ratio to the assume-no-motion floor, lower is better

Detection experiment 1.1
Rebuilt targets 1.1
First fine-tune 1.1
Matched budget 0.907
Near pass 0.833
The three π0.5 runs cluster in a band; the two pilot SmolVLA runs are the same footage read by a different model.

Right hand

Detection experiment 1.4
Rebuilt targets 1.2
First fine-tune 1.2
Matched budget 1
Near pass 0.914
The harder hand comes under every baseline for the first time in the programme.

Improvement, then a flat line, watched in the open

Checkpoints were scored every few thousand steps. Held-out error fell steadily, went flat around step 20,000, and stayed flat while training loss kept dropping through 49 passes over the footage. Watching for exactly that divergence is how memorisation is kept out of the results: a loss curve alone would have called the run still improving.

The footage

kitchen · washing up
Continuous two-handed work is what the pilot exam is made of, and where this run earned its score.

Limitations

The checkpoint-selection bias above is the headline caveat. Single seed throughout, so the gap to the π0.5 cluster is one run beating another until the variance check runs. And the pilot exam is continuous-motion work, where the do-nothing floor is comparatively weak; the ten-hour corpus’s exam later charged for stillness and reset every expectation.

Where this goes next

This run closed the pilot chapter: the footage demonstrably carries more signal than its strongest baseline when the reader is right. The next question was volume, and the volume test bought twelve times the data to ask it. Protocol and corpus live in the data write-up.

Sources

Scored by the programme’s frozen evaluation pipeline and re-scored on a second machine; the full checkpoint curve was kept alongside the headline score. Full scoring artifacts are available to buyers on request. Model: SmolVLA / LeRobot (Hugging Face).