Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning
The paper says latent visual reasoning improves when the model learns its own task-aligned helper representations, rather than relying on a fixed vision encoder.
Scaffolding Minds adds a dedicated encoder to produce better latent targets during supervised fine-tuning. It also changes the reinforcement learning stage by learning both the mean and variance of the latent sampler, giving the system more room to explore latent reasoning paths. The authors report gains over the strongest latent reasoning baseline: +9.5 points on FrozenLake spatial planning, +19 points on 32x32 grids, and +5.6 points on average across nine visual-centric reasoning benchmarks. HF Daily Papers' note
Scaffolding Minds adds a dedicated encoder to produce better latent targets during supervised fine-tuning. It also changes the reinforcement learning stage by learning both the mean and variance of the latent sampler, giving the system more room to explore latent reasoning paths. The authors report gains over the strongest latent reasoning baseline: +9.5 points on FrozenLake spatial planning, +19 points on 32x32 grids, and +5.6 points on average across nine visual-centric reasoning benchmarks. HF Daily Papers' note
score 4