ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval
ConceptFormer uses query-conditioned latent concept tokens to align visual document evidence with retrieval relevance.
The paper frames those concepts as an intermediate layer between localized page evidence and semantic matching, avoiding text-only descriptions or raw visual annotations as the main supervision signal. A strong vision-language model sets the number of latent concept tokens dynamically during training. On visual document retrieval benchmarks, the authors report 16.7% average NDCG@10 gains over the strongest visual baseline and 22.1% over the strongest OCR-based text baseline. Code and data are listed as available. HF Daily Papers' note
The paper frames those concepts as an intermediate layer between localized page evidence and semantic matching, avoiding text-only descriptions or raw visual annotations as the main supervision signal. A strong vision-language model sets the number of latent concept tokens dynamically during training. On visual document retrieval benchmarks, the authors report 16.7% average NDCG@10 gains over the strongest visual baseline and 22.1% over the strongest OCR-based text baseline. Code and data are listed as available. HF Daily Papers' note
score 4