Megadose AI progress, ranked and analyzed.

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

· HF Daily Papers ·
The paper pairs a fast calibrated tagger with a generative VLM to mark hallucinated text spans more reliably.

SpanCalib-VLM uses XLM-RoBERTa-Large with a SigLIP vision encoder, then re-scores spans proposed by a fine-tuned Qwen3.5-4B model. On the SHROOM-Visions English evaluation split, the ensemble reports 0.41 Pearson calibration correlation, 0.39 overall IoU, and 70.7% detection accuracy. The authors say they are releasing the model weights and code. HF Daily Papers' note

score 4

Categories: Research