Trained on the pilot corpus: 39 episodes / 26,040 frames, one take per task — the corpus that proved the line end to end. About the data and the exam ↗
Highlights
- The obvious lever, priced under controlled conditions. Many more detected hands, identical model and budget as the first fine-tune, so the targets were the only variable.
- The frozen exam did its job. Training loss kept improving while held-out error moved the other way, which is the exact gap a held-out exam exists to catch, and ours caught it.
- The result was diagnostic. The added hands were real, but their positions came from single-frame estimates with no tracking behind them. Presence quality is not position quality.
- It upgraded the pipeline. This run is the direct cause of the rebuilt, blind-verified detector line whose result the rebuilt-targets run measures.
Overview
The first fine-tune left nearly half the pilot’s frames carrying a hand pose forward instead of measuring one, and coverage looked like the obvious lever. A candidate detector layer recovered far more hands per frame. This run trained on those denser targets, on the same footage, with the same step and batch budget as the first fine-tune, so the targets were the only thing that changed.
Results
Denser targets, higher error on both hands
Left hand — ratio to the assume-no-motion floor, lower is better
Right hand
The mechanism is specific. The added detections were genuine hands, so coverage rose. But each new position was a single-frame guess, and a training target that jitters frame to frame teaches jitter. The model fit those targets better every step, which is why the loss curve looked like progress, and predicted held-out motion less well, which is why it was not.
Limitations
One seed, one run per configuration, so the size of the effect is measured once. Its direction was confirmed by the rebuild that followed. These are offline prediction errors on held-out human video.
Where this goes next
The detector line was rebuilt around tracking and side-verification, gated on a blind, hand-marked answer key. The rebuilt-targets run measures what that bought. Questions about how targets are made are ones we like.
Sources
The rebuilt detector’s gates: precision 99%, side accuracy 96.3% on a blind, hand-marked answer key; the corpus record is in the data write-up. Full scoring artifacts are available to buyers on request.