Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction
GAS uses generation only during training, then drops that branch so inference cost stays unchanged.
The paper frames visual generation as auxiliary supervision for multimodal representation learning. Its decoupled Mixture-of-Transformers setup shares a lower trunk while keeping upper understanding layers away from direct generation gradients. The authors report gains across scales and training stages, especially on perception and spatial comprehension. HF Daily Papers' note
The paper frames visual generation as auxiliary supervision for multimodal representation learning. Its decoupled Mixture-of-Transformers setup shares a lower trunk while keeping upper understanding layers away from direct generation gradients. The authors report gains across scales and training stages, especially on perception and spatial comprehension. HF Daily Papers' note
score 4