On Locality and Length Generalization in Visual Reasoning
Local, sequential vision models generalized better on the paper’s visual state-tracking tasks.
The authors test whether models that process images through local glimpses avoid failures seen in global, single-shot vision systems. Their experiments find that vision models can learn global shortcuts and then break when task length or complexity increases. Strictly local recurrent policies reduced those failures in the studied tasks. The paper argues that local attention may be an overlooked requirement for robust compositional generalization. HF Daily Papers' note
The authors test whether models that process images through local glimpses avoid failures seen in global, single-shot vision systems. Their experiments find that vision models can learn global shortcuts and then break when task length or complexity increases. Strictly local recurrent policies reduced those failures in the studied tasks. The paper argues that local attention may be an overlooked requirement for robust compositional generalization. HF Daily Papers' note
score 4