Test-Time Training for Modality Order Consistency in Vision-Language Models
Putting the image before the question made the same vision-language models perform better.
The paper reports a repeatable order effect across three models and three benchmarks: image-first prompts beat question-first prompts even though the content is semantically unchanged. The authors use that gap to train at test time for consistency between the two prompt orders. Their method narrows the order gap and also improves the stronger image-first path over baseline. Activation patching points to a mid-network region where the two prompt orders diverge, and the adaptation reduces that misalignment. ArXiv · AI/CL/LG's note
The paper reports a repeatable order effect across three models and three benchmarks: image-first prompts beat question-first prompts even though the content is semantically unchanged. The authors use that gap to train at test time for consistency between the two prompt orders. Their method narrows the order gap and also improves the stronger image-first path over baseline. Activation patching points to a mid-network region where the two prompt orders diverge, and the adaptation reduces that misalignment. ArXiv · AI/CL/LG's note
score 4