What did we actually train the model to output?

This page: real training examples and saved supervision. Other page: held-out test predictions and accuracy →

All five runs: 3,000 optimizer updates · 4 GPUs · full understanding-branch fine-tuning, not LoRA. Language understanding, vision encoder and connector updated; image-generation expert unchanged. No pixel-generation loss.

Training variantLoss applied toTotal job time
Text reasoningComplete reasoning + final answer3h 15m 51s
Primitive reasoningText + point/line tags and coordinates + final answer3h 14m 22s
Points onlyPoint tags and coordinates only3h 18m 38s
Prose + pointsPoints only; supplied gold prose has zero loss3h 14m 01s
MixedPoints only; half the examples include gold prose3h 13m 45s

Times include setup, two-step checks, evaluation smoke tests and checkpoint export—not just optimizer compute. Main loops took approximately 2h 58m–3h 01m (see logs for exact boundaries). All jobs completed; low training loss did not translate into good test generation.

24 examples, six per training grid size (3×3–6×6), selected in release order—not by success. These are not the 600 test mazes. No free-running predictions on these training examples have been collected for this view.

1. GIVEN TO THE MODEL

Exact saved input

The image and question receive no answer-token loss. The saved training row contains no system message; special sequence delimiters are inserted by the loader.

2. SAVED TRAINING TARGET

Exact target, with supervision highlighted

Green = supervised span · unhighlighted = supplied context, zero direct token loss.

Exact points/lines and annotation provenance

These primitive targets were constructed from our recovered grid/path annotations, not released as native point annotations by ThinkMorph. Text targets use their released text-only reasoning. This overlay renders saved target coordinates, not a model prediction.

Source file

Highlighting shows character spans for readability. Selective loss is implemented on actual tokenizer tokens; a boundary token may also include adjacent whitespace. Selective runs exclude EOS and final-answer loss. Prose in the teacher-forced target can reveal the complete solution; it is NOT supplied during autonomous test inference.