OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination
OmniConfess tests each generated token against text, image, audio, and video evidence to see what actually supports it.
The method is training-free: it holds a candidate response fixed, then re-scores it under controlled channel-wise evidence changes. That produces a token-by-channel “confession” showing which modality a commitment depends on. The authors use that signal to keep grounded content and revise parts driven by irrelevant or contradictory evidence. They also introduce OmniHalluBench, a 3,540-example benchmark across six datasets and multiple modality/task settings. HF Daily Papers' note
The method is training-free: it holds a candidate response fixed, then re-scores it under controlled channel-wise evidence changes. That produces a token-by-channel “confession” showing which modality a commitment depends on. The authors use that signal to keep grounded content and revise parts driven by irrelevant or contradictory evidence. They also introduce OmniHalluBench, a 3,540-example benchmark across six datasets and multiple modality/task settings. HF Daily Papers' note
score 4