Learning Multimodal Embeddings with Evidence-Aligned Readout
EviAlign ties retrieval embeddings to the boundaries of generated evidence, not just to the evidence text itself.
The paper introduces a shared multimodal model that generates five semantic evidence units, reads the contextualized state at each unit boundary, and pools those states into one normalized embedding. A controlled study finds the benefit is strongest when the evidence organization and readout positions are aligned, with a reported 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, the method reaches 76.9 average Recall@1 using 500K training pairs while keeping single-vector indexing and scoring. HF Daily Papers' note
The paper introduces a shared multimodal model that generates five semantic evidence units, reads the contextualized state at each unit boundary, and pools those states into one normalized embedding. A controlled study finds the benefit is strongest when the evidence organization and readout positions are aligned, with a reported 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, the method reaches 76.9 average Recall@1 using 500K training pairs while keeping single-vector indexing and scoring. HF Daily Papers' note
score 4