Megadose AI progress, ranked and analyzed.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

· HF Daily Papers ·
The paper says showing teacher models the ground-truth future path can make their driving explanations less faithful.

The authors call this “trajectory anchoring bias”: the model rationalizes the revealed outcome instead of reasoning from the scene. Their proposed AD-MCQ reframes planning as choosing among explicit trajectory candidates, avoiding open-ended trajectory synthesis. DEFT-RLVR then uses future trajectories after the decision as verification targets, not pre-decision hints. Experiments are reported to improve autonomous-driving reasoning while preserving, and in some cases enhancing, general visual capability. HF Daily Papers' note

score 4

Categories: Research