Megadose AI progress, ranked daily.

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

· ArXiv · AI/CL/LG ·
The paper frames token-level visual dependence as both a warning signal for forgetting and a guide for learning from unlabeled multimodal streams.

The authors define MU-CPT as continual post-training for deployed multimodal LLMs on streaming unlabeled data. They argue existing methods treat target tokens too uniformly, missing differences in how much each token depends on visual input. Their VDA framework uses optimal transport to limit distortion of old visual-dependence patterns while preventing a slide into language-only bias. A second component, VMA, weights adaptation toward visually grounded new-task learning, and experiments under their MU-CPT setup are reported as validating the approach. ArXiv · AI/CL/LG's note

score 4

Categories: Research