Source-linked AI summary
Vision Language Models in Autonomous Driving: A Survey and Outlook
Xingcheng Zhou, Mingyu Liu, Ekim Yurtsever, Bare Luka Zagar, Walter Zimmer, Hu Cao, Alois C. Knoll
TL;DR
Autonomous driving still faces limitations in complex-scene understanding, explainability, generalization, and robustness. This paper surveys VLM tasks, model categories, applications, datasets, metrics, benefits, challenges, and research directions, concluding that VLM capabilities offer a research direction for addressing these limitations while real-world deployment remains constrained by computation, latency, privacy, and reliability concerns.
Problem
Current autonomous-driving systems face challenges in generalization, interpretability, causal confusion, and robustness, motivating investigation of VLMs for scene understanding and reasoning.
Method
The paper provides a comprehensive survey organized by VLM modality and inter-modality connection, covering tasks, metrics, applications, datasets, benefits, challenges, and future trajectories.
Results
The survey reviews VLM applications across perception, navigation, decision-making, end-to-end driving, and data generation, including zero-shot perception, language-guided tasks, and connected decision-making.
Takeaways & Limitations
Large VLMs offer potential benefits for open-vocabulary recognition, explainable interaction, language-guided planning, and high-level autonomous-driving reasoning.
Takeaways & Limitations
Real-world deployment must address computation and latency requirements, along with privacy, network dependency, hallucination, bias, and reliability risks.
Abstract
from arXiv · showhide
The applications of Vision-Language Models (VLMs) in the field of Autonomous Driving (AD) have attracted widespread attention due to their outstanding performance and the ability to leverage Large Language Models (LLMs). By incorporating language data, driving systems can gain a better understanding of real-world environments, thereby enhancing driving safety and efficiency. In this work, we present a comprehensive and systematic survey of the advances in vision language models in this domain, encompassing perception and understanding, navigation and planning, decision-making and control, end-to-end autonomous driving, and data generation. We introduce the mainstream VLM tasks in AD and the commonly utilized metrics. Additionally, we review current studies and applications in various areas and summarize the existing language-enhanced autonomous driving datasets thoroughly. Lastly, we discuss the benefits and challenges of VLMs in AD and provide researchers with the current research gaps and future trends.
I. INTRODUCTION
The survey motivates VLMs for autonomous driving by identifying limitations in current systems and reviews their tasks, applications, datasets, benefits, challenges, and research gaps.
- Current autonomous-driving systems struggle with complex dynamic environments, decision explanation, and human instructions, often missing details and context.
- VLMs combine visual and linguistic data to support traditional AD tasks and emerging applications such as zero-shot perception and language-guided navigation.
- The survey categorizes existing VLM studies by model types and application domains in autonomous driving.
- It consolidates mainstream vision-language tasks and their commonly used metrics for autonomous driving.
- The paper systematically summarizes classic and language-enhanced autonomous-driving datasets.
- It discusses VLM applications and technological advances, alongside benefits, challenges, and research gaps in autonomous driving.
A. Autonomous Driving
Autonomous driving is organized around modular and end-to-end paradigms, while generalization, interpretability, causal confusion, and robustness remain open challenges that VLMs may address.
- SAE levels range from Level 0 no automation to Level 5 full automation, with higher autonomy reducing human intervention and increasing environmental-understanding requirements.
- Modular autonomous driving separates perception, prediction, planning, and control into independently developed components.
- Perception processes sensor data from cameras, LiDAR, radar, or event-based cameras to extract traffic participants’ locations, sizes, velocities, and related details.
- Planning generates paths using global route search and local real-time adjustments based on vehicle state and environmental information.
- End-to-end autonomous driving integrates separate components into one differentiable system mapping raw sensor inputs directly to control signals.
- Both modular and end-to-end systems face challenges in generalization, interpretability, causal confusion, and robustness.
B. Large Language Models
Large language and vision-language models connect natural-language processing with visual understanding, enabling multimodal inputs, outputs, and cross-modal reasoning for autonomous driving.
- Transformers and BERT established scalable, context-aware language representations that support the development of large models.
- VLMs connect modalities through vision-text matching or vision-text fusion, with fused features supporting downstream tasks.
- Large language models typically contain billions of parameters and exhibit few-shot, zero-shot, and multi-step reasoning abilities.
- Vision-language models bridge language and computer vision by learning relationships between visual content and natural language.
- Autonomous-driving VLMs are categorized as Multimodal-to-Text, Multimodal-to-Vision, or Vision-to-Text according to their input-output modalities.
III. VISION-LANGUAGE TASKS IN AUTONOMOUS DRIVING
Vision-language tasks in autonomous driving extend conventional perception and tracking by using language to specify objects and evaluate scene understanding across modalities and time.
- The survey organizes autonomous-driving vision-language tasks into object referring and tracking, open-vocabulary perception, scene understanding, language-guided navigation, and conditional data generation.
- Object Referring and Tracking: Object referring localizes one or more language-specified objects in 2D images or 3D spaces, generalizing conventional object detection.
- Object Referring and Tracking: Referred object tracking uses language expressions to track specified objects across consecutive video frames, emphasizing referring consistency and robustness.
- Object Referring and Tracking: Object referring is evaluated with mAP, while multiple-object referring and tracking adopts HOTA, MOTA, MOTP, and IDS.
B. Open-Vocabulary Traffic Environment Perception
Open-vocabulary perception targets unseen object categories through language-defined vocabularies, while traffic-scene understanding evaluates visual reasoning and generated language with task-specific metrics.
- Open-Vocabulary 3D Object Detection: Open-vocabulary 3D object detection identifies novel categories absent from training and assigns text-defined semantic labels to 3D bounding boxes.
- Open-Vocabulary 3D Semantic Segmentation: Open-vocabulary 3D segmentation classifies point-cloud or mesh regions into semantically meaningful categories unseen during training.
- Open-Vocabulary Traffic Environment Perception: OV-3D detection uses 3D mAP, whereas OV-3D segmentation uses mIOU and mean Accuracy for evaluation.
- Traffic Scene Understanding: Traffic-scene VQA answers questions from images or video across perception, planning, spatial, temporal, and causal reasoning categories.
- Traffic Scene Understanding: Captioning generates textual descriptions of scenes or objects and can cover scene description, importance ranking, and action explanation.
- Traffic Scene Understanding: Multiple-choice VQA uses Top-N Accuracy, while open-ended VQA and captioning commonly use BLEU, METEOR, ROUGE, and CIDEr.
- Traffic Scene Understanding: Because these metrics have different focuses and limitations, studies often combine them and may evaluate reasoning chains or semantic similarity with language models.
D. Language-Guided Navigation
Language-guided navigation requires vehicles to interpret instructions, localize destinations, plan paths, and reach targets in complex outdoor environments.
- Task Description: Language-guided navigation makes reasoned plans to reach specified locations from natural-language instructions in complex, variable, and uncertain outdoor scenes.
- Task Description: The task commonly includes language-guided localization and path planning, with target-location understanding as a prerequisite for successful navigation.
- Task Description: Existing studies often use Vision-Text Matching VLMs to locate instruction-specified destinations among landmarks in maps.
- Evaluation Metrics: Evaluation considers task completion, target-localization accuracy, and deviation between the target location and final arrival.
- Conditional Autonomous Driving Data Generation: Conditional driving-data generation produces targeted photorealistic synthetic data to diversify and purposefully augment autonomous-driving training databases.
- Conditional Autonomous Driving Data Generation: Generation systems can condition outputs on text prompts, BEV masks, action states, bounding boxes, or HD maps for single- or multiple-view images and videos.
- Evaluation Metrics: FID and FVD evaluate conditional image and video generation, while CLIP score measures visual consistency across generated multiple street views.
IV. MAINSTREAM METHODS AND TECHNIQUES
The survey organizes VLM applications in autonomous driving across perception and understanding tasks, including referring, open-vocabulary perception, retrieval, scene understanding, and visual reasoning. Existing studies also use VLMs for traffic-focused analysis, pedestrian detection, and benchmarked task coverage.
- Perception and Understanding: VLM-based perception research spans object referring, open-vocabulary 3D detection and segmentation, language-guided retrieval, driving scene understanding, and visual scene reasoning.Table I summarizes these task categories alongside related LLM and VLM types.
- Open-Vocabulary Perception: OpenScene and CLIP2Scene align image, text, and point-cloud representations to support zero-shot open-vocabulary 3D scene understanding.UP-VL extends this direction with unsupervised multimodal auto-labeling for class-agnostic 3D detection and tracking.
- Language-Guided Retrieval: Language-guided retrieval uses semantic similarity, BEV features, language augmentation, and multiple imperfect descriptions for scene or vehicle retrieval.Applications include data selection, corner-case retrieval, and natural-language vehicle search.
- Driving Scene Understanding: Traffic VLM studies address image captioning, anomaly recognition, visual question answering, temporal reasoning, and traffic-domain knowledge injection through synthetic captions.NuScenes-QA and NuScenes-MQA provide question-answering benchmarks based on nuScenes, while traffic video-caption pairs support fine-tuning.
- Pedestrian Detection: VLM-based pedestrian detection uses CLIP-derived pixel-wise semantic contexts and pseudo-labeling to address human-like object confusion and scarce border-case samples.The cited approach is designed as an extra-annotation-free method.
- Research Status: The surveyed perception and understanding applications remain in early stages despite covering increasingly diverse language-enhanced capabilities.The survey identifies these directions as promising areas for further work.
B. Navigation and Planning
VLMs extend autonomous-driving navigation from predefined location descriptions to arbitrary natural-language instructions. Surveyed systems connect language and scene representations to waypoint generation, trajectory planning, and motion prediction.
- Language-Guided Navigation: CLIP-driven cross-modal alignment broadens language-guided navigation from predefined locations to free and arbitrary instructions.This development also promotes language-augmented maps.
- Language-Guided Navigation: Talk to the Vehicle maps semantic occupancy and natural-language encodings to local waypoints, which a planning module converts into an executable trajectory.The waypoint generator network performs the language-and-scene mapping.
- Prediction and Planning: GPT-driver reformulates motion planning as language modeling, while CoverNet-T jointly encodes text-based scene descriptions and rasterized scene images for planning-related prediction.These approaches leverage language models or multimodal representations for motion planning and trajectory prediction.
C. Decision-Making and Control
LLMs are being explored as decision-making and control components that use commonsense reasoning, driver interaction, memory, and unified scene representations. Connecting these systems with visual-LLM modules is identified as a promising route toward integrated AD.
- Decision Making: Decision-making studies use LLMs to assist or emulate drivers and to handle complex scenarios requiring human commonsense understanding.Closed-loop systems commonly add memory modules for driving scenarios, experiences, and decision information.
- Decision Making: LanguageMPC uses an LLM for complex-scene decisions, while Drive as You Speak enables direct driver–vehicle communication through an orchestrated LLM framework.Drive as You Speak stores past scenario experiences, decision cues, and reasoning processes in a vector database.
- Unified Representations: BEVGPT unifies scenario prediction, decision-making, and motion planning through a single LLM with BEV input, while DwLLMs generates driving actions from object-level scene vectors.DwLLMs uses two-stage pre-training and fine-tuning to understand driving scenarios and produce actions.
- Vision-LLM Integration: Most existing decision-making and control studies rely on LLMs alone, but visual-LLM connectors can link them with perception for mid-to-mid or end-to-end AD.The survey presents connector design tailored to AD as an open research direction.
D. End-to-End Autonomous Driving
End-to-end AD is a natural application for multimodal-to-text VLMs because both map sensor inputs to plans or control outputs. Early systems demonstrate integrated control, narration, reasoning, and conditional data generation, while autoregressive-model limitations remain unresolved.
- End-to-End AD System: End-to-end AD integrates separate driving components into a differentiable system that maps raw sensor data directly to plans or control signals.This structure aligns naturally with multimodal-to-text VLMs.
- End2End AD System: DriveGPT4, ADAPT, DriveMLM, and VLP demonstrate end-to-end systems that combine sensor inputs with questions, rules, commands, or context to produce controls and explanations.ADAPT continuously outputs control signals, narration, and reasoning descriptions, while DriveMLM supports closed-loop driving in Town05Long.
- Open Challenges: Autoregressive large VLMs and LLMs introduce inherent limitations, leaving many end-to-end AD issues requiring further consideration and resolution.The survey therefore treats current demonstrations as feasibility evidence rather than a completed solution.
- Conditional Data Generation: Conditional generation methods produce large-scale, high-quality autonomous-driving data from controllable visual or textual conditions, supporting data-driven AD development.DrivingDiffusion demonstrates conditional multi-view video generation with layout and text control.
- Conditional Data Generation: DriveGAN controls generated vehicle behavior by disentangling action-dependent and action-independent scene features for high-fidelity neural simulation and data generation.BEVControl uses sketch-style BEV layouts and text prompts to generate street-view multi-view images.
- World Models: World-model approaches learn structured driving-environment representations from video, action, and text to generate predictable or realistic future scenes.DriveDreamer, GAIA-1, and ADriver-I respectively model real-world scenarios, decode token sequences with video diffusion, or fuse vision-action pairs.
V. DATASETS
This section surveys conventional and language-enhanced autonomous-driving datasets, emphasizing how language enriches semantic and contextual understanding and supports interaction-oriented tasks.
- Autonomous-driving datasets underpin the development of safe and efficient perception, prediction, and planning systems.
- Mainstream datasets cover tasks including object detection, tracking, and segmentation across multiple data modalities.
- Language-enhanced datasets extend visual data with natural-language questions, instructions, and descriptions for richer scene comprehension and human–vehicle interaction.
- Table III defines dataset task abbreviations including SS, OT, REID, MP, AE, VSR, VLN, VQA, DM, IC, IR, and related categories.
- Language-oriented datasets support tracking, referring, visual question answering, spatial reasoning, navigation, and related autonomous-driving tasks.Examples include CityFlow-NL, Refer-KITTI, NuPrompt, Talk2Car, Talk2BEV, and language-guided navigation datasets.
VI. DISCUSSION
The discussion presents large VLMs as extending autonomous-driving capabilities across perception, navigation, planning, decision-making, control, interaction, and data generation, while identifying adaptation and data-readiness gaps.
- Advantages of large VLMs in AD: Large VLMs extend classic autonomous-driving modules—including perception, navigation, planning, and control—and offer a new perspective on end-to-end driving.
- Advantages of large VLMs in AD: Zero-shot generalization can improve recognition of corner cases and unknown objects, while scene–text fusion supports open-ended driving-related question answering.
- Advantages of large VLMs in AD: Large VLMs support language-guided localization, navigation, and planning by aligning cross-modal features and interpreting detailed instructions.
- Advantages of large VLMs in AD: Their high-level reasoning, common-sense knowledge, and traffic-rule understanding offer potential for high-level decision-making and control.
- Advantages of large VLMs in AD: Large VLMs can generate realistic, controllable, high-quality synthetic autonomous-driving data with implicit environmental modeling.
- Autonomous Driving Foundation Model: Current foundation-model research motivates Autonomous Driving Foundation Models pretrained on diverse datasets for interpretability, reasoning, forecasting, introspection, and multiple driving tasks.
- Multi-Modality Adaptation: Directly converting sensor data into text can lose context and environmental information and makes motion planning, control, and decision-making dependent on perception quality.
- Public Data Availability and Formatting: Existing autonomous-driving datasets are not yet optimal for direct LLM adaptation, and instruction-tuning formats and large-scale traffic image–text pairs remain underinvestigated.
VII. CHALLENGES
This section identifies deployment, temporal-understanding, safety, ethical, privacy, and security challenges that constrain the effective use of large VLMs in real-world autonomous driving.
- Computation Demands and Deployment Latency: Large VLMs impose substantial computation and latency demands because billion-scale models make fine-tuning and inference resource-intensive.
- Computation Demands and Deployment Latency: On-vehicle deployment offers low latency and local data storage, whereas cloud deployment offers centralized computation but introduces latency, privacy, and network-dependency risks.
- Computation Demands and Deployment Latency: A hybrid architecture could place critical real-time decisions on vehicles and data analytics or model updates in the cloud, but software and hardware requirements remain open questions.
- Temporal Scene Understanding: Image-level VLMs are insufficient for temporal scene understanding because driving dynamics and accident causality require video information.
- Temporal Scene Understanding: Potential remedies include training video-language models or adding temporal adapters while converting video into image-language processing formats.
- Ethical and Societal Concern in Driving Safety: Unfiltered pretraining data can produce toxic, biased, or harmful content, while language-model hallucinations challenge driving-system safety and reliability.
- Data Privacy and Security: Sensitive driving data raises collection, usage, storage, and privacy concerns, while vehicle and adversarial attacks can manipulate systems or cause decision misjudgments.
- Conclusion: The survey synthesizes challenges spanning computation and latency, temporal understanding, ethical and societal safety, and data privacy and security.