CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG
CANOPY compresses retrieved multimodal evidence by choosing different context sizes for different regions, then asks for more retrieval when the evidence still looks thin.
The paper frames the problem as post-retrieval compression for text, tables, images, and video. Its hierarchy-based scorer keeps some regions broad and others narrow, without using LLM calls for node-level pruning. A critic can trigger targeted follow-up retrieval, and the added material is compressed before use. On five QA benchmarks over a 33M-item mixed corpus, the authors report higher average answer accuracy than evaluated retrieval baselines, with token reductions of 14.2-27.7% in one Qwen3-VL-8B-Instruct setting.
HF Daily Papers' note
The paper frames the problem as post-retrieval compression for text, tables, images, and video. Its hierarchy-based scorer keeps some regions broad and others narrow, without using LLM calls for node-level pruning. A critic can trigger targeted follow-up retrieval, and the added material is compressed before use. On five QA benchmarks over a 33M-item mixed corpus, the authors report higher average answer accuracy than evaluated retrieval baselines, with token reductions of 14.2-27.7% in one Qwen3-VL-8B-Instruct setting.
HF Daily Papers' note
score 4