Megadose AI progress, ranked and analyzed.

Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

· HF Daily Papers ·
The paper argues visual tokens can be kept intact while cheaper MLP adapters reconstruct their layer-wise states.

The authors say pruning saves compute but throws away visual evidence later layers may need. Their tests suggest the visual influence needed by text is concentrated in a low-dimensional subspace, and that layer-specific visual states are predictable. They propose `δ-Vision`, which builds lightweight visual memories instead of repeatedly evolving all visual tokens through the Transformer. Across image and video benchmarks, they report better accuracy than pruning baselines at similar or lower compute. HF Daily Papers' note

score 4

Categories: Research