Megadose AI progress, ranked and analyzed.

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

· HF Daily Papers ·
Qwen-Drive-1.0 adds driving perception and planning heads to a pretrained vision-language model while keeping the base VLM architecture.

The model combines 3D object detection, semantic occupancy prediction, BEV map segmentation, visual question answering, and ego-trajectory planning in one framework. Its BEV perception head is described as an inspectable interface into the scene structure learned from shared representations. The authors say staged training mixes driving supervision with general vision-language data to retain broader visual understanding and instruction following. Evaluations report strong 3D perception and competitive motion-planning results across open-loop, pseudo-closed-loop, and closed-loop settings. HF Daily Papers' note

score 5

Categories: Research