AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning
A post-hoc interpolation step is pitched as a way to keep multimodal retrieval models from forgetting earlier cross-modal alignment.
AlphaWiSE combines two frozen checkpoints by fitting one scalar interpolation coefficient per aligned parameter tensor. Those coefficients are learned from a smaller exemplar memory, then used to produce a single deployed checkpoint with no added inference cost. The authors report consistent gains over continual-learning baselines on audio-image-text retrieval across multiple retrieval directions and metrics. ArXiv · AI/CL/LG's note
AlphaWiSE combines two frozen checkpoints by fitting one scalar interpolation coefficient per aligned parameter tensor. Those coefficients are learned from a smaller exemplar memory, then used to produce a single deployed checkpoint with no added inference cost. The authors report consistent gains over continual-learning baselines on audio-image-text retrieval across multiple retrieval directions and metrics. ArXiv · AI/CL/LG's note
score 4