Megadose Built for builders and researchers.

Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models

· HF Daily Papers ·
SECRET steers question-state representations so an AVLLM relies on the modality the question actually asks for.

The paper identifies “source-confused grounding hallucination,” where audio or visual cues from an unused modality pull the model toward unsupported answers. Its analysis points to a “question-relay” mechanism: question states carry both interfering cues and required-source evidence. The proposed training-free method, SECRET, intervenes at that relay using contrastive question representations from modality-pathway interventions. Across CMM and AVHBench on three AVLLMs, it beats prior training-free methods, with reported gains up to 18.0 and 7.1 percentage points over base models. HF Daily Papers' note

score 4

Categories: Research