Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Built on Qwen3.5-4B, the model adds an external BEV perception head and a flow-matching trajectory planning module without altering the underlying VLM architecture.
- Performs 3D object detection, semantic occupancy prediction, BEV map segmentation, and visual question answering simultaneously
- Generates 5-second future ego trajectories at 10 Hz via a diffusion transformer Planning Expert, with optional text-reasoning conditioning
- Trained on 2.83 million public samples across unified cross-dataset trajectory formats to mitigate catastrophic forgetting
- Achieved an average driving QA score of 69.43 while retaining general multimodal reasoning performance across standard benchmarks
Autonomous driving researchers can adapt a single unmodified general-purpose VLM to handle end-to-end perception, reasoning, and trajectory planning.
Sources
Read this as text
Back to the AI news