The actual BAGEL visual-thought experiment family: base BAGEL-7B-MoT versus released ThinkMorph-7B, which was fine-tuned with interleaved text and visual-thought images.
Released Chart Refocus, Jigsaw, Spatial Navigation and Visual Search examples, including intermediate visual-thought images and exact text.
START HEREMatched native inference on eight released training examples. Shows the raw generations and where base BAGEL failed.
Base BAGEL and visual-SFT ThinkMorph on VSP and VisPuzzle, with generated visual thoughts and per-example outputs.
Released stochastic decoding pathway. Useful failure viewer, including nonsensical visual thoughts.
The crucial input-order correction: 61.1% on the small balanced 18-maze probe.
Replace the generated thought image with wrong, blank, noisy, or omitted thoughts and inspect how the answer changes.
| Probe | BAGEL base | ThinkMorph visual-SFT |
|---|---|---|
| Balanced greedy VSP, n=18 | 0% | 5.6% |
| Balanced greedy VisPuzzle, n=12 | 25% | 58.3% |
| Correct interleaved-input VSP, n=18 | not run | 61.1% |
| Released training examples, n=8 | 5/8 | 8/8 |
The tiny probes are qualitative diagnostics, not substitutes for the paper's full benchmark table. Source archive recovered from $WORK/thinkmorph-reproduction/reports on Galvani.