Why this model
GR00T N1.7 is NVIDIA’s open vision-language-action model. We included it because of its pretraining data: 20,854 h of action-labelled first-person human video, the largest such corpus disclosed to date, and the only pretraining set among the models we run that matches the kind of footage we produce.
We report one scored run. It was trained on the ten-hour corpus for ten thousand steps and evaluated on the same 55 held-out takes, row for row, as the π0.5 and SmolVLA runs.
Main findings
- A third model family on the same frozen exam. GR00T is the first model we have trained whose pretraining data is first-person human video at this scale.
- The registered bar still stands. Against the do-nothing floor, this run scores 1.225 on the left hand and 1.067 on the right. Both are above the floor.
- Closest approach to the floor so far, on the right hand. At the far end of the one-second prediction window the right hand is within about two percent of doing nothing. The volume test was roughly fifty percent above it on the same task.
- Against the designed pair, on the same scored rows. Right-hand masked error is 9.2% below the ten-hour SmolVLA run and 29% below the volume test. The left hand is about two percent behind the SmolVLA run.
- The prediction we registered was half wrong. We wrote down the expected outcome before the run. The right hand matched it, the left did not, and the miss stays in the record.
Setup
This run follows the design of the earlier pair: hold the corpus, the training budget and the exam fixed, and change only the model. Scoring uses the same (episode, frame) pairs as the volume test’s frozen record, at the same ten thousand steps. Predictions are graded over the exam’s full one-second window, because a shorter window is an easier task and the scorer refuses to accept one.
One caveat keeps this run out of a shared three-way chart. GR00T cannot read the corpus in the form the designed pair trained on: its export uses compressed video rather than lossless frames, carries the hand channels only, and the run sat on a different GPU generation. Those differences are small and were declared before the run, but they travel with every number below, which is why the comparisons here are written in prose rather than drawn on one chart.
The reason to expect something from this model is straightforward. Its pretraining data is our product category, and the ten-hour SmolVLA run had already shown that changing the model moves the score on this footage. A model whose prior comes from first-person human video was the natural next test.
Results against the do-nothing floor
Lower is better. The floor is one, which means predicting that the hands do not move for the next second. The registered bar is 0.9× the floor on both hands.
Both hands are above the floor. The right hand is the closest approach any run of ours has made to it on the full holdout, and that is an approach rather than a claim on the bar.
What is actually attributable to GR00T
A caution before the headline. On this exam the do-nothing floor is weaker on the right hand than the left, for every model we have scored, so a right-hand ratio is intrinsically easier to win.
What is specific to GR00T is the margin on the same hand against the same floor: 9.2% less masked error than the ten-hour SmolVLA run, on identical scored rows, from a single seed on each side. That is one run ahead of another run once, until the registered seed-variance check runs. On the left hand it gives back about two percent, so neither model dominates the other.
The prediction we registered
Before the run, the protocol recorded an expectation: at or modestly better than the SmolVLA run on both hands, and still above the floor. The outcome matched on the right, came in behind on the left, and stayed above the floor as predicted. We record it as half wrong rather than rounding it in our favour, because a page that only confirms its own predictions is not measuring anything.
Longer training is not the cheap lever here
Training loss fell sharply in the first third of the run and had flattened well before cutoff. That is the opposite of the ten-hour SmolVLA run, whose loss was still falling when training stopped, so the cheap next step differs by family: more steps for SmolVLA, something else for GR00T.
Where every model on this exam still loses is unchanged. It is the windows where a hand holds still.
Next steps
Three model families are now scored on one exam, and the ranking questions ahead of us are registered rather than improvised: the seed-variance check, longer SmolVLA training, and training aimed directly at stillness. The corpus is described in the data write-up.
Sources
NVIDIA Isaac GR00T, pinned at the commit recorded in the run log · openpi (π0.5) · LeRobot and SmolVLA. Pretraining-corpus figures are the vendors’ own published statements, read on the dates given in the number ledger.