InitAI
InitAI · Human video for robot learning

What we learned training GR00T N1.7.

One scored run on the ten-hour corpus. Of the models we have trained, GR00T comes closest to the do-nothing baseline on this exam, which is what you would expect from a model pretrained on first-person human video.

2026-09-11 · 4 min read
doc NOTE rev A issued

Why this model

GR00T N1.7 is NVIDIA’s open vision-language-action model. We included it because of its pretraining data: 20,854 h of action-labelled first-person human video, the largest such corpus disclosed to date, and the only pretraining set among the models we run that matches the kind of footage we produce.

We report one scored run. It was trained on the ten-hour corpus for ten thousand steps and evaluated on the same 55 held-out takes, row for row, as the π0.5 and SmolVLA runs.

Main findings

Setup

This run follows the design of the earlier pair: hold the corpus, the training budget and the exam fixed, and change only the model. Scoring uses the same (episode, frame) pairs as the volume test’s frozen record, at the same ten thousand steps. Predictions are graded over the exam’s full one-second window, because a shorter window is an easier task and the scorer refuses to accept one.

One caveat keeps this run out of a shared three-way chart. GR00T cannot read the corpus in the form the designed pair trained on: its export uses compressed video rather than lossless frames, carries the hand channels only, and the run sat on a different GPU generation. Those differences are small and were declared before the run, but they travel with every number below, which is why the comparisons here are written in prose rather than drawn on one chart.

The reason to expect something from this model is straightforward. Its pretraining data is our product category, and the ten-hour SmolVLA run had already shown that changing the model moves the score on this footage. A model whose prior comes from first-person human video was the natural next test.

Results against the do-nothing floor

Lower is better. The floor is one, which means predicting that the hands do not move for the next second. The registered bar is 0.9× the floor on both hands.

Left hand 1.225 Above the floor, and above the registered bar
Right hand 1.067 Closest approach to the floor on the full holdout

Both hands are above the floor. The right hand is the closest approach any run of ours has made to it on the full holdout, and that is an approach rather than a claim on the bar.

What is actually attributable to GR00T

A caution before the headline. On this exam the do-nothing floor is weaker on the right hand than the left, for every model we have scored, so a right-hand ratio is intrinsically easier to win.

What is specific to GR00T is the margin on the same hand against the same floor: 9.2% less masked error than the ten-hour SmolVLA run, on identical scored rows, from a single seed on each side. That is one run ahead of another run once, until the registered seed-variance check runs. On the left hand it gives back about two percent, so neither model dominates the other.

The prediction we registered

Before the run, the protocol recorded an expectation: at or modestly better than the SmolVLA run on both hands, and still above the floor. The outcome matched on the right, came in behind on the left, and stayed above the floor as predicted. We record it as half wrong rather than rounding it in our favour, because a page that only confirms its own predictions is not measuring anything.

Longer training is not the cheap lever here

Training loss fell sharply in the first third of the run and had flattened well before cutoff. That is the opposite of the ten-hour SmolVLA run, whose loss was still falling when training stopped, so the cheap next step differs by family: more steps for SmolVLA, something else for GR00T.

Where every model on this exam still loses is unchanged. It is the windows where a hand holds still.

Next steps

Three model families are now scored on one exam, and the ranking questions ahead of us are registered rather than improvised: the seed-variance check, longer SmolVLA training, and training aimed directly at stillness. The corpus is described in the data write-up.

Sources

NVIDIA Isaac GR00T, pinned at the commit recorded in the run log · openpi (π0.5) · LeRobot and SmolVLA. Pretraining-corpus figures are the vendors’ own published statements, read on the dates given in the number ledger.