FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
FocusVTC claims adaptive image resolution can compress long text without losing the evidence needed for reasoning.
The paper renders text into low-DPI global views, then selectively enhances regions the model decides are relevant. Its training uses 29.4K reasoning-evidence localization examples tied to page indices and bounding boxes. On RULER v1, it reports 87.4 at 2.9x input compression, far above Glyph’s 57.5 at 3.0x. The authors also report gains on LongBench, MRCR, VTCBench, MMMU, and MME, with a 2.79x online end-to-end speedup over text in the MRCR latency test. HF Daily Papers' note
The paper renders text into low-DPI global views, then selectively enhances regions the model decides are relevant. Its training uses 29.4K reasoning-evidence localization examples tied to page indices and bounding boxes. On RULER v1, it reports 87.4 at 2.9x input compression, far above Glyph’s 57.5 at 3.0x. The authors also report gains on LongBench, MRCR, VTCBench, MMMU, and MME, with a 2.79x online end-to-end speedup over text in the MRCR latency test. HF Daily Papers' note
score 5