AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
AnswerMap turns a VLM’s own answer probabilities into a spatial rationale, without model internals or training.
The method slices an image into row and column bands, asks the frozen model yes/no relevance questions, and combines the resulting posteriors into a query-conditioned map. The paper says this lets fixed read-outs produce continuous outputs such as location, instead of relying on discrete text tokens. In validation across four models and three query distributions, AnswerMap aligned with the model’s generated point at AUC 0.85, versus 0.38 for attention. Deleting the mapped region flipped 53% of correct answers, compared with 19% for attention-based regions. HF Daily Papers' note
The method slices an image into row and column bands, asks the frozen model yes/no relevance questions, and combines the resulting posteriors into a query-conditioned map. The paper says this lets fixed read-outs produce continuous outputs such as location, instead of relying on discrete text tokens. In validation across four models and three query distributions, AnswerMap aligned with the model’s generated point at AUC 0.85, versus 0.38 for attention. Deleting the mapped region flipped 53% of correct answers, compared with 19% for attention-based regions. HF Daily Papers' note
score 4