ThinkMorph visual-SFT archive

The actual BAGEL visual-thought experiment family: base BAGEL-7B-MoT versus released ThinkMorph-7B, which was fine-tuned with interleaved text and visual-thought images.

Visual training samples

Released Chart Refocus, Jigsaw, Spatial Navigation and Visual Search examples, including intermediate visual-thought images and exact text.

START HERE

Base BAGEL vs ThinkMorph

Matched native inference on eight released training examples. Shows the raw generations and where base BAGEL failed.

Balanced greedy probe

Base BAGEL and visual-SFT ThinkMorph on VSP and VisPuzzle, with generated visual thoughts and per-example outputs.

Official sampled decoding

Released stochastic decoding pathway. Useful failure viewer, including nonsensical visual thoughts.

Correct interleaved-input VSP

The crucial input-order correction: 61.1% on the small balanced 18-maze probe.

Visual-thought causal swap

Replace the generated thought image with wrong, blank, noisy, or omitted thoughts and inspect how the answer changes.

Key archived numbers

ProbeBAGEL baseThinkMorph visual-SFT
Balanced greedy VSP, n=180%5.6%
Balanced greedy VisPuzzle, n=1225%58.3%
Correct interleaved-input VSP, n=18not run61.1%
Released training examples, n=85/88/8

The tiny probes are qualitative diagnostics, not substitutes for the paper's full benchmark table. Source archive recovered from $WORK/thinkmorph-reproduction/reports on Galvani.