1. GIVEN TO THE MODEL
Exact saved input
The image and question receive no answer-token loss. The saved training row contains no system message; special sequence delimiters are inserted by the loader.
This page: real training examples and saved supervision. Other page: held-out test predictions and accuracy →
All five runs: 3,000 optimizer updates · 4 GPUs · full understanding-branch fine-tuning, not LoRA. Language understanding, vision encoder and connector updated; image-generation expert unchanged. No pixel-generation loss.
| Training variant | Loss applied to | Total job time |
|---|---|---|
| Text reasoning | Complete reasoning + final answer | 3h 15m 51s |
| Primitive reasoning | Text + point/line tags and coordinates + final answer | 3h 14m 22s |
| Points only | Point tags and coordinates only | 3h 18m 38s |
| Prose + points | Points only; supplied gold prose has zero loss | 3h 14m 01s |
| Mixed | Points only; half the examples include gold prose | 3h 13m 45s |
Times include setup, two-step checks, evaluation smoke tests and checkpoint export—not just optimizer compute. Main loops took approximately 2h 58m–3h 01m (see logs for exact boundaries). All jobs completed; low training loss did not translate into good test generation.
24 examples, six per training grid size (3×3–6×6), selected in release order—not by success. These are not the 600 test mazes. No free-running predictions on these training examples have been collected for this view.
The image and question receive no answer-token loss. The saved training row contains no system message; special sequence delimiters are inserted by the loader.
Green = supervised span · unhighlighted = supplied context, zero direct token loss.
These primitive targets were constructed from our recovered grid/path annotations, not released as native point annotations by ThinkMorph. Text targets use their released text-only reasoning. This overlay renders saved target coordinates, not a model prediction.
Highlighting shows character spans for readability. Selective loss is implemented on actual tokenizer tokens; a boundary token may also include adjacent whitespace. Selective runs exclude EOS and final-answer loss. Prose in the teacher-forced target can reveal the complete solution; it is NOT supplied during autonomous test inference.