Using OCR Heads to Verbalize Image Semantics
The paper says OCR attention heads in VLMs also expose general image semantics, not just text.
Across four models, the authors identify attention heads that are causally necessary for OCR and find they produce interpretable labels for non-text image regions too. In Qwen3-VL-8B, the same mechanism can verbalize a word token as “bike” or a visual region like a bird wing as “feathers.” They compress those attention patterns into a “verbalization lens” that surfaces language-aligned image features from early layers. They also report causal edits using the inverse transform, such as replacing a tractor concept with a revolver in an image representation. ArXiv · AI/CL/LG's note
Across four models, the authors identify attention heads that are causally necessary for OCR and find they produce interpretable labels for non-text image regions too. In Qwen3-VL-8B, the same mechanism can verbalize a word token as “bike” or a visual region like a bird wing as “feathers.” They compress those attention patterns into a “verbalization lens” that surfaces language-aligned image features from early layers. They also report causal edits using the inverse transform, such as replacing a tractor concept with a revolver in an image representation. ArXiv · AI/CL/LG's note
score 6