Do LiDAR Language Models Really Understand Spatio-temporal Relationships?
The benchmark finds that some LiDAR language model scores can be matched by a fixed answer strategy.
The paper introduces LiDAR-Hallu, a 10,000-question diagnostic benchmark over 150 nuScenes scenes. It tests object existence, position, distance ordering, motion, and temporal localization with controls meant to expose whether models distinguish the queried physical relation. In paired scenes requiring opposite answers, the tested models often gave the same answer, and both configurations missed every positive lateral-motion case. Temporal-shuffle contrastive decoding brought little net gain because fixes were offset by new errors. ArXiv · AI/CL/LG's note
The paper introduces LiDAR-Hallu, a 10,000-question diagnostic benchmark over 150 nuScenes scenes. It tests object existence, position, distance ordering, motion, and temporal localization with controls meant to expose whether models distinguish the queried physical relation. In paired scenes requiring opposite answers, the tested models often gave the same answer, and both configurations missed every positive lateral-motion case. Temporal-shuffle contrastive decoding brought little net gain because fixes were offset by new errors. ArXiv · AI/CL/LG's note
score 4