Architecture-Dependent Fusion Pathways in MLLMs
The paper argues that MLLM fusion works differently depending on the model architecture.
Zhu and Wu compare concatenation-based MLLMs with native multimodal models using alignment, attention, entropy, intrinsic dimensionality, and causal interventions. They find concatenation models tend to route fusion through a text-first, vision-later pathway. Native models show earlier visual-textual co-adaptation and reorganization of the feature space. Source: ArXiv · AI/CL/LG's note
Zhu and Wu compare concatenation-based MLLMs with native multimodal models using alignment, attention, entropy, intrinsic dimensionality, and causal interventions. They find concatenation models tend to route fusion through a text-first, vision-later pathway. Native models show earlier visual-textual co-adaptation and reorganization of the feature space. Source: ArXiv · AI/CL/LG's note
score 4