Megadose Built for builders and researchers.

Selective Transfer of RL Updates for Visual Reasoning

· ArXiv · AI/CL/LG ·
The paper argues that only the strongest RL-induced parameter directions transfer well into vision-language models.

The authors isolate the update created during RL post-training, rather than merging whole model endpoints. Their Selective-RL method keeps dominant matrix-wise directions and applies them to a VLM’s language modules. Across three model families and five visual-reasoning benchmarks, it beat full-update interpolation in 12 of 15 comparisons, with an 8.55-point MathVision gain on Qwen. Controls suggest the gains do not come from update size or low rank alone. ArXiv · AI/CL/LG's note

score 5

Categories: Research