Trained on the pilot corpus: 39 episodes / 26,040 frames, one take per task — the corpus that proved the line end to end. About the data and the exam ↗
Highlights
- the detection experiment’s ground is fully recovered. Tracked, side-verified targets bring both hands back to the first fine-tune’s level, confirming the detection experiment’s diagnosis.
- The score does not move. Level with the first fine-tune on both hands, despite measurably better targets.
- The flat line was read wrong, usefully. At the time it looked like the corpus was the ceiling. The matched-budget run showed the ceiling belonged to the model, and that correction is why a second architecture is now tested before any conclusion about the data.
- The ruler was checked. The same model was scored on the archived exam base too, because a data rebuild can move the baselines by itself.
Overview
After the detection experiment, the detector line was rebuilt around two rules: no detection becomes a target unless it is tracked across frames, and no left/right assignment survives unless the side view confirms it. The rebuilt line passed a blind gate first: precision 99% and side accuracy 96.3% on frames no earlier round had seen, with genuinely ambiguous sides abstained rather than guessed. This run trained on those targets with the same footage and budget as the two runs before it.
Results
Three runs, one model, only the targets changed
Left hand — ratio to the assume-no-motion floor, lower is better
Right hand
The ruler itself was audited
A dataset rebuild changes the exam as well as the training data, because the baselines are computed on the exam’s own clips. So the rebuilt-targets run was additionally scored on the archived exam base, where it reads 1.096 left and 1.32 right. Any improvement claim has to hold on both bases or it is an artefact of the ruler moving. This one held: level is level on either ruler.
Limitations
Single seed per run, so “level with the first fine-tune” is bounded by unmeasured run-to-run variance. Offline prediction metrics on held-out human video throughout.
Where this goes next
The matched-budget run holds everything fixed except the architecture. The detector gates and the corpus record live in the data write-up.
Sources
Scored by the programme’s frozen evaluation pipeline; the detector gates were graded on a blind 200-frame answer key marked by two independent squads. Full scoring artifacts are available to buyers on request.