Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D uses multimodal language models to spot and correct geometry and motion inconsistencies in generated 3D and 4D content.
The paper targets spatial hallucinations like duplicated structures and misaligned geometry, plus 4D failures such as jitter, identity flicker, and structural drift. Its framework renders multiple views and frames, has LMMs summarize inconsistencies, then uses those signals to guide image-space consistency optimization. Candidate corrections are selected through multi-model voting, without retraining or changing the underlying generator. The authors report stronger results than baseline methods across varied 3D and 4D generation settings. HF Daily Papers' note
The paper targets spatial hallucinations like duplicated structures and misaligned geometry, plus 4D failures such as jitter, identity flicker, and structural drift. Its framework renders multiple views and frames, has LMMs summarize inconsistencies, then uses those signals to guide image-space consistency optimization. Candidate corrections are selected through multi-model voting, without retraining or changing the underlying generator. The authors report stronger results than baseline methods across varied 3D and 4D generation settings. HF Daily Papers' note
score 4