Megadose AI progress, ranked and analyzed.

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

· HF Daily Papers ·
WeVisDoc adds a diagnostic second stage that targets the parser’s remaining weak spots, not just broader training coverage.

The framework first expands document variety with heterogeneous data and degradation synthesis that preserves structure. It then tests the Stage I parser on a held-out probe, groups residual errors by visual-structural clusters, and uses that to redirect data construction and token budget. The 4B model reports 95.38 Overall on OmniDocBench v1.6 and a 75.54 mean Overall across three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Stage II improves both 2B and 4B models, with the largest stated gain a 4.03-point lift for the 4B model on PureDocBench’s Real Degraded track. HF Daily Papers' note

score 5

Categories: Research