What the models actually generated

TRAIN examples are taken from the 6,000 training records; TEST examples are held-out benchmark mazes. The original ID suffix _cot is not a model setting or supervision toggle: both suffixed and unsuffixed source rows contain reasoning. Model choices stay fixed when browsing mazes or checkpoints.

Checkpoint sweep · Earlier evaluations · Exact training targets / supervision

Input maze

Reference is computed from the gold board, never a model prediction. Overlays below are viewer drawings of generated tokens—not images the model generated or re-observed.

Scores for this evaluation