InitAI
InitAI · Human video for robot learning

Target quality beats target quantity.

The controlled experiment that priced label quality, and the reason every corpus since ships with a blind-verified detector line.

2026-08-18 · 2 min read
doc RUN-02 rev A issued

Detection experiment π0.5 pilot corpus 3,000 steps · 4.9 passes

Trained on the pilot corpus: 39 episodes / 26,040 frames, one take per task — the corpus that proved the line end to end. About the data and the exam ↗

Highlights

Overview

The first fine-tune left nearly half the pilot’s frames carrying a hand pose forward instead of measuring one, and coverage looked like the obvious lever. A candidate detector layer recovered far more hands per frame. This run trained on those denser targets, on the same footage, with the same step and batch budget as the first fine-tune, so the targets were the only thing that changed.

Results

Denser targets, higher error on both hands

Left hand — ratio to the assume-no-motion floor, lower is better

First fine-tune 1.1
Detection experiment 1.1
Same footage, same budget, denser targets: the left hand moves the wrong way.

Right hand

First fine-tune 1.2
Detection experiment 1.4
The right hand pays the larger penalty for untracked position estimates.

The mechanism is specific. The added detections were genuine hands, so coverage rose. But each new position was a single-frame guess, and a training target that jitters frame to frame teaches jitter. The model fit those targets better every step, which is why the loss curve looked like progress, and predicted held-out motion less well, which is why it was not.

Limitations

One seed, one run per configuration, so the size of the effect is measured once. Its direction was confirmed by the rebuild that followed. These are offline prediction errors on held-out human video.

Where this goes next

The detector line was rebuilt around tracking and side-verification, gated on a blind, hand-marked answer key. The rebuilt-targets run measures what that bought. Questions about how targets are made are ones we like.

Sources

The rebuilt detector’s gates: precision 99%, side accuracy 96.3% on a blind, hand-marked answer key; the corpus record is in the data write-up. Full scoring artifacts are available to buyers on request.