Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
The paper says showing teacher models the ground-truth future path can make their driving explanations less faithful.
The authors call this “trajectory anchoring bias”: the model rationalizes the revealed outcome instead of reasoning from the scene. Their proposed AD-MCQ reframes planning as choosing among explicit trajectory candidates, avoiding open-ended trajectory synthesis. DEFT-RLVR then uses future trajectories after the decision as verification targets, not pre-decision hints. Experiments are reported to improve autonomous-driving reasoning while preserving, and in some cases enhancing, general visual capability. HF Daily Papers' note
The authors call this “trajectory anchoring bias”: the model rationalizes the revealed outcome instead of reasoning from the scene. Their proposed AD-MCQ reframes planning as choosing among explicit trajectory candidates, avoiding open-ended trajectory synthesis. DEFT-RLVR then uses future trajectories after the decision as verification targets, not pre-decision hints. Experiments are reported to improve autonomous-driving reasoning while preserving, and in some cases enhancing, general visual capability. HF Daily Papers' note
score 4