On the Design Fundamentals of Pixel Text Representation Learning
Pixel Linguist II is trained to read visual text at native resolution and still holds up after 80% token compression.
The paper argues that pixel-text encoders need variable resolutions, natural image-text grounding, layout-aware rendering, and a multilingual curriculum to avoid brittle shortcuts. Its authors train Pixel Linguist II on 280M examples with on-the-fly rendering and unified contrastive grounding. They report state-of-the-art results across English, cross-lingual, and multilingual Visual STS and ViDoRe, plus stronger downstream MLLM evaluation. HF Daily Papers' note
The paper argues that pixel-text encoders need variable resolutions, natural image-text grounding, layout-aware rendering, and a multilingual curriculum to avoid brittle shortcuts. Its authors train Pixel Linguist II on 280M examples with on-the-fly rendering and unified contrastive grounding. They report state-of-the-art results across English, cross-lingual, and multilingual Visual STS and ViDoRe, plus stronger downstream MLLM evaluation. HF Daily Papers' note
score 4