NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
NaviDC-OCR is pitched as a single parser for both clean digital pages and distorted camera-captured documents.
The paper says the system adds deformation-aware learning so vision-language models can better handle geometric distortions. It also uses adaptive sampling for complex layouts and separates content from structure when modeling formulas and tables. The authors report state-of-the-art results on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, plus a first-place ranking in the ICDAR 2026 Sci-ImageMiner Challenge. HF Daily Papers' note
The paper says the system adds deformation-aware learning so vision-language models can better handle geometric distortions. It also uses adaptive sampling for complex layouts and separates content from structure when modeling formulas and tables. The authors report state-of-the-art results on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, plus a first-place ranking in the ICDAR 2026 Sci-ImageMiner Challenge. HF Daily Papers' note
score 5