TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval
TraVEL trains video embeddings to care about ego-motion without needing trajectory data at search time.
The paper fine-tunes Qwen3-VL-Embedding for driving-video retrieval, first with paired clips and reasoning traces, then with trajectory similarity as a training reward. The authors say caption supervision improves retrieval but still misses fine motion distinctions like left versus right turns or acceleration versus braking. On their nuReasoning-derived benchmark, TraVEL improves longitudinal and lateral mAP over SFT at both 2B and 8B scales. ArXiv · AI/CL/LG's note
The paper fine-tunes Qwen3-VL-Embedding for driving-video retrieval, first with paired clips and reasoning traces, then with trajectory similarity as a training reward. The authors say caption supervision improves retrieval but still misses fine motion distinctions like left versus right turns or acceleration versus braking. On their nuReasoning-derived benchmark, TraVEL improves longitudinal and lateral mAP over SFT at both 2B and 8B scales. ArXiv · AI/CL/LG's note
score 4