Source-linked AI summary
Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City
Adrian Cespedes, Marcelo Chincha, Dunant Cusipuma, Victor Flores-Benites, David Ortega, Arturo Deza
TL;DR
VLMs’ cross-geographic generalization in real-world, chaotic driving environments remains underexplored. Robusto-2 compares Lima and New York City human drivers and VLMs using dashcam VQA, finding geography-related response patterns are generally weak while humans and VLMs diverge by question type.
Problem
Real-world evidence on how VLMs generalize across geographies and chaotic driving environments remains limited.
Method
The study evaluates 10 Lima humans, 10 New York City humans, and 10 VLMs on VQA questions about 200 dashcam clips from both cities.
Results
Humans and VLMs diverge depending on question type, while response patterns show little to no geographic difference between Lima and New York City.
Takeaways & Limitations
Dashcam data from different chaotic cities may help evaluate and develop autonomous-vehicle models despite limited cross-geographic response differences.
Takeaways & Limitations
The study uses only 20 five-second clips and 20 questions per system, yielding a relatively small dataset despite balanced observations.
Abstract
from arXiv · showhide
As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how well will these systems generalize in new settings, in particular out-of-distribution (OOD) edge-case scenarios in new geographies? In this paper, we study this open question by providing a full factorial analysis with human drivers of Lima, human drivers from New York City, and VLMs and showing them dashcam footage collected from Lima and New York City -- prompting them with a variety of questions under a Visual Question Answering (VQA) paradigm. In particular, we pick these two cities as they are highly challenging driving locations where no Self-Driving Car company currently operates in, and ask questions that span 4 categories: Factual, Ratings, Counterfactual and Reasoning. We find that Humans and VLMs diverge in their responses -- though this is modulated by the type of questions asked, and that Humans answer similarly independent of where they are from (Lima/NYC). To our surprise, we did not find a strong difference in terms of answers (Humans or VLMs) that was modulated by geography, likely due to their high out-of-distribution nature. Our dataset is available at: https://huggingface.co/datasets/Artificio/robusto-2
1 Introduction
The paper examines cross-geography and cross-system generalization in autonomous-driving question answering using Lima and New York City dashcam footage, comparing local human drivers with VLMs. Its fully factorial design evaluates how responses vary by respondent population and recording geography, with data and analyses released openly.
- Research question: The study asks how VLMs generalize to new driving scenarios as autonomous-driving systems evolve toward Vision-Language-Action models with VLM cognitive backbones.The introduction frames this question amid a shift from hard-coded heuristics and end-to-end models toward modular end-to-end systems.
- Motivation: Lima and New York City are selected as globally challenging driving environments where self-driving cars remain undeployed.The paper links their lack of deployment to complexity and regulation, contrasting them with cities where several startups already operate.
- Study design: The authors collect same-camera dashcam videos from Lima and New York City and use Visual Question Answering with Lima humans, NYC humans, and VLMs.This setup compares three respondent populations across footage from both geographies.
- Study design: The fully factorial 2 × 3 design tests responses across two recording geographies and three respondent populations, while releasing human, VLM, stimulus, and analysis data.The design makes respondent and recording geography jointly variable rather than evaluating either factor alone.
2 Experimental Setup
The experiment compares 10 Lima humans, 10 New York City humans, and 10 VLMs using English VQA on 200 five-second driving clips. Clips are selected for high response variance across models, and questions cover factual, ratings, counterfactual/hypothetical, and reasoning categories.
- Participants: The study used 30 systems: 10 humans from Lima, 10 humans from New York City, and 10 VLMs, with all questions asked in English.
- OOD Selection: The evaluation selected 10 Lima and 10 New York City clips with the highest cross-model answer variance, treating model disagreement as an indicator of OOD scenes.Responses were embedded, and standard deviations across embedding dimensions were ranked to identify high-dispersion videos.
- Video Dataset: The dataset contained 200 five-second clips, split evenly between New York City and Lima, with metadata covering traffic participants, ego-vehicle actions, signals, sky, and weather.
- VQA Protocol: The 20-question VQA protocol comprised factual questions about driving context, fixed ratings of perceptual judgments, counterfactual or hypothetical scenarios, and reasoning questions.The four blocks were Factual (Q1–Q5), Ratings (Q6–Q10), CounterFactual & Hypothetical (Q11–Q15), and Reasoning (Q16–Q20).
- Reasoning Questions: The reasoning block was newly added relative to Robusto-1 to examine whether humans and VLMs reason similarly, including difficult and sometimes ill-posed questions.
3 General Assessment of Responses for Humans and Machines across Lima and New York City
Across Lima and New York City, response patterns differ little by geography and are instead modulated mainly by question type. Embedding and rating analyses show humans resembling one another more than VLMs, with this human–VLM difference accentuated in Lima.
- Comparative analysis: PCA and Representational Similarity Analysis compare Lima humans, NYC humans, and VLMs across question blocks using sentence-embedding responses.The analyses include 2D PCA projections and representational comparisons of answer embeddings.
- Comparative analysis: Results show little to no geography-dependent pattern difference, with responses instead modulated by the type of questions asked across humans and VLMs.The authors attribute subtle geographic differences to questions steering answers into a narrower subspace.
- Comparative analysis: Ratings have similar geometry across Lima and NYC; humans are more similar to each other than to VLMs, with the human–VLM difference accentuated in Lima.This finding comes from averaged per-geography multidimensional-scaling plots using L1 distances across Block 2 ratings (Q6–Q10).
- Comparative analysis: The general pattern holds when the embedding analysis is replicated with Qwen3-Embedding-4B instead of all-mpnet-base-v2.The replication examines both local block-wise and global answer differences.
4 Analysis of Systematic Bias
Humans cluster together independently of whether videos come from Lima or New York City, while VLMs are farther from humans in Lima than New York City. Across rating analyses, response patterns show little geographic variation, with Question 8 producing the most similar VLM ratings across regions.
- Geography-level differences: Humans cluster together independently of video geography, while VLMs are farther from the human cluster in Lima than New York City.New York City humans also differ less from VLMs than Lima humans in both geographies.
- Rating distributions: Question 8 produced the most similar VLM rating patterns across geographies, indicating similar responses for Lima and New York City videos.This comparison used Wasserstein distances averaged across violin plots and VLMs.
- Geographic variation: Rating patterns showed little column-wise variation by geography, despite a mild tendency for larger differences in Lima than New York City.This pattern appeared in the per-question, per-geography comparison of individual ratings and inter-system differences.
5 Semantic Similarity via LLM-as-a-Judge
The study uses a two-stage Qwen3-4B LLM-as-a-judge pipeline to compare semantic alignment among VLMs and human participants across Lima and NYC videos. Results are summarized with agreement heatmaps across factual, counterfactual, and reasoning blocks, while numeric-rating comparisons are excluded due to judge inconsistency.
- Evaluation framework: Qwen3-4B in thinking mode judged pairwise semantic alignment among VLMs, Lima participants, and NYC participants across videos and questions.The evaluation was designed to compare all participating systems.
- Evaluation framework: The pipeline first filtered semantically comparable answer pairs, then scored the degree of alignment between approved pairs using a rubric.Each comparison used the same video-question pair and two candidate answers.
- Evaluation results: 63,121 pairs passed in Block 1, 57,886 in Block 3, and 61,302 in Block 4, corresponding to passage rates of 86.6%, 79.4%, and 84.1%.Pairs that did not pass the first stage received no rubric score.
- Limitations: Block 2 was excluded because the judge produced inconsistent scores for identical pairs of numeric ratings across runs.The inconsistency arose because Block 2 assessed purely numeric values.
- Evaluation results: Heatmaps report mean agreement scores between every system pair by Lima or NYC video region and question block, covering Blocks 1, 3, and 4.The color scale ranges from strong disagreement to strong agreement.
6 Discussion
The study examines cross-cultural generalization between humans and VLMs using Lima and New York City driving data, finding human–VLM gaps within cities but similar response patterns across geographies. Its conclusions are based on a deliberately small, balanced dataset of 12,000 data points.
- Study scope: The study evaluates humans and VLMs as cognitive backbones for VLAs using NYC and Lima data, including human subgroups from both cities.The authors motivate this assessment by noting that VLM-based VLA backbones require model calibration.
- Limitations: The dataset contains 20 five-second clips and 20 questions per system across 30 subjects, totaling 12’000 data points.The authors describe this as a low number of observations, while emphasizing that the study used the same setup across conditions.
- Main conclusion: A cross-cultural gap appears when comparing humans with VLMs in both cities, but not in the systems’ general response patterns across geographies.The authors characterize humans and VLMs as being chaotic in similar ways across Lima and New York City.
Supplementary material … 6.5 Question Generation and System Prompt
The benchmark used student drivers from Lima and New York City alongside ten VLMs to answer structured questions about standardized dashcam clips. Clips were selected for complex OOD driving events, annotated with metadata, and used to generate and constrain four-block VQA responses.
- Humans of Lima:: 10 Lima participants were full-time students aged 18–30 with valid Peruvian driver’s licenses and English fluency.They were recruited through an open digital call across several universities.
- Humans of New York City:: New York City participants were recruited through a similar open call, but recruitment was harder because many students do not drive and compensation was less attractive.The passage attributes this difficulty partly to Metro use and the international wage gap.
- Vision Language Models (VLMs):: 10 VLMs from companies in the United States and China answered the same English questions as the human participants.The evaluated models included Cosmos Reason 8B, Gemini 3 Flash Preview, Gemini 3 Pro Preview, InternVL3 8B, LLaVA-Video-7B-Qwen2, MiniCPM-o 2.6, PerceptionLM 8B - Phi4 Multimodal Instruct, Qwen3-VL 8B Instruct, and VideoLLama3 7B.
- 6.3 Data Acquisition and Clip Extraction: The dataset used 1080p, 30-fps BlackVue DR970X Plus dashcam recordings from diverse urban environments in New York City and Lima.A custom Python/Qt GUI extracted short segments from several hours of footage.
- 6.3 Data Acquisition and Clip Extraction: Clips targeted OOD or high-complexity events, including anomalous driver behavior, vulnerable road users, dense traffic, complex intersections, and variable weather.These criteria emphasized events such as cutting in, sudden braking, illegal maneuvers, jaywalking, and bicycles in unexpected lanes.
- 6.4 Metadata Taxonomy: Each clip received structured metadata capturing the ego-vehicle’s intent and the surrounding environmental context to support VQA generation.The full annotation taxonomy is provided in Table 2.
- 6.2 List of All VQA Questions: The benchmark organized VQA into four blocks, with Blocks 2–4 predefined and Block 1 dynamically generated for each clip.The question taxonomy is presented in Table 1, and some questions depend on the video.
- 6.5 Question Generation and System Prompt: ChatGPT-4 generated five contextually appropriate questions per JSON-described sample, while block-specific prompts enforced visual-only, uncertainty-aware, concise responses.Rating questions required exactly one predefined option; hypothetical and reasoning questions required conclusions supported by visible evidence.
6.6 VLMs Used and Inference Parameters · Gemini 3 Pro Preview & Gemini 3 Flash Preview
The study evaluated 10 VLMs using Vertex AI Batch Prediction for two closed-source models and a Google Compute Engine A100 instance for eight open-source models. Model-specific video processing and generation settings were fixed for the experiments, with Gemini inputs automatically downsampled to 1 FPS.
- 6.6 VLMs Used and Inference Parameters: 10 VLMs were evaluated: two closed-source models through Google Vertex AI Batch Prediction and eight open-source models on a Google Compute Engine a2-highgpu-1g instance with 1× NVIDIA A100 40 GB.Open-source checkpoints were loaded from HuggingFace and run with the transformers Python library.
- 6.6 VLMs Used and Inference Parameters: The reported inference parameters reflect the settings used in the experiments, not each model’s absolute limits, and Table 3 summarizes these configurations.A dash indicates that a parameter was not applicable or controlled by the API or model internals.
- Gemini 3 Pro Preview & Gemini 3 Flash Preview: Both Gemini models used MEDIA_RESOLUTION_HIGH, temperature 1.0, and Vertex AI Batch Prediction with MP4 videos and one JSONL request file stored in Google Cloud Storage.The Gemini API automatically downsampled all video inputs to 1 FPS server-side.
- Gemini 3 Pro Preview & Gemini 3 Flash Preview: Cosmos-Reason2-8B used float16 with SDPA attention, 4 FPS MP4 input, 256–8192 vision tokens, max_new_tokens=1024, and temperature 0.5.The checkpoint was nvidia/Cosmos-Reason2-8B.
- Gemini 3 Pro Preview & Gemini 3 Flash Preview: InternVL3-8B used bfloat16 with FlashAttention-2, 8 uniformly spaced video segments, dynamic tiling with one tile plus thumbnail, max_new_tokens=3072, and temperature 0.5.Frames were resized to 896×896 and normalized with ImageNet mean and standard deviation.
- Gemini 3 Pro Preview & Gemini 3 Flash Preview: LLaVA-Video-7B-Qwen2 and MiniCPM-o-2.6 decoded up to 50 frames at 10 FPS and used max_new_tokens=1024 with temperature 0.5.MiniCPM-o-2.6 used 336×336 frames; higher resolutions degraded and destabilized its outputs.
- Gemini 3 Pro Preview & Gemini 3 Flash Preview: Perception-LM-8B used 10 FPS with a maximum of 15 frames and max_new_tokens=256, while Qwen3-VL-8B-Instruct and VideoLLaMA3-7B used 10 FPS MP4 inputs with model-specific frame or pixel caps.Perception-LM outputs degraded and became incoherent when frame count or max_new_tokens increased; Qwen3-VL used a 1280 × 28 × 28 pixel limit, and VideoLLaMA3-7B capped inputs at 50 frames.
- Gemini 3 Pro Preview & Gemini 3 Flash Preview: Phi-4-Multimodal used bfloat16 with FlashAttention-2, up to 50 frames at 10 FPS, 428×428 padded frames, max_new_tokens=1024, and temperature 0.5.Higher frame resolutions caused degraded outputs, so 428×428 was kept fixed.
6.7 Embeddings used in the Analysis … 6.11 Embedding Visualisation (PCA)
The analysis used two sentence-embedding models, standardized answer encoding and preprocessing, calibrated VLM temporal-spatial reasoning, and evaluated metrics and PCA visualizations across systems, regions, and question blocks. The embedding-based analyses produced unchanged result patterns across the two encoders.
- 6.7 Embeddings used in the Analysis: all-mpnet-base-v2 was loaded in fp32 native precision for the paper’s embedding analyses.
- Qwen3-Embedding-4B: Qwen/Qwen3-Embedding-4B was loaded in bfloat16 as the alternative embedding model.
- Encoding: Both models used SentenceTransformer for uniform encoding, with embeddings keyed by system, video, question number, and repetition.Batch size one ensured reproducibility because padding multiple answers changes individual embedding computations; the keys index projection, cosine-similarity, and RSA analyses.
- 6.8 VLM Pre-testing: VLM pre-testing used a synthetic ball-and-star video with 10 questions to assess temporal and spatial reasoning before evaluating Lima and NYC driving clips.The protocol also targeted potential mirage behavior, in which confident visual reasoning may lack grounding in the visual input.
- 6.9 Preprocessing step: Answers were whitespace-trimmed and Unicode-NFKC-normalized, while Block 2 free-text responses were extracted as integers in.Both raw and processed datasets were retained; subsequent Block 2 analyses used the processed version, while other blocks remained free-text strings.
- 6.10 Metrics evaluation pipeline: Metrics used cleaned answers from VLMs, Lima human annotators, and NYC human annotators across Lima and NYC videos and four question blocks.The common setup used only answers with t = 1 and represented answers with embeddings from the frozen sentence encoder.
- 6.11 Embedding Visualisation (PCA): PCA was fitted independently per block, projecting centered embeddings into two dimensions and faceting plots by region and block.Color encoded system family and form encoded system type; block-wise and global analyses showed unchanged patterns between allmpnet and Qwen3.
6.12 Pairwise Cosine Similarity · 6.13 Representational Similarity Analysis (RSA) · 6.14 Rating Bias (Block 2)
The paper compares systems using direct embedding similarity and RSA, which captures agreement in how systems structure stimulus items despite paraphrase differences. For Block 2 ratings, it compares VLM distributions with human consensus and quantifies Lima–NYC shifts using Wasserstein and Kolmogorov–Smirnov statistics.
- 6.12 Pairwise Cosine Similarity: For each item and system pair, pairwise cosine similarity measures whether their answers align in embedding space.Scores are averaged within each block and region into symmetric system-by-system matrices with diagonal entries equal to 1.
- 6.12 Pairwise Cosine Similarity: Cosine similarity can understate agreement when systems paraphrase the same content using different wording.This surface-alignment limitation motivates the RSA analysis that follows.
- 6.13 Representational Similarity Analysis (RSA): RSA compares the geometry each system imposes on the same stimulus set rather than directly comparing embedding vectors.Each system’s fingerprint is an m×m Gram matrix of pairwise similarities among its ℓ2-normalised item embeddings.
- 6.13 Representational Similarity Analysis (RSA): For two systems on the same block and region, RSA computes a Pearson correlation between the upper-triangular entries of their similarity matrices.A high correlation indicates similar item-structuring even when the systems’ absolute embeddings occupy different regions of space.
- 6.13 Representational Similarity Analysis (RSA): Cosine asks whether two answers look alike, whereas RSA asks whether two systems find the same items alike.The analyses are complementary and both are plotted.
- 6.14 Rating Bias (Block 2): Block 2 rating-bias analysis compares each VLM’s 1–10 rating distribution for each question and region with consensus ratings from Lima and NYC human groups.Human group members are collapsed to a single mean consensus rating per question and video region, with both human references overlaid on VLM distributions.
- 6.14 Rating Bias (Block 2): Cross-region rating shifts are measured per VLM using 1-Wasserstein distance and the Kolmogorov–Smirnov statistic when videos change from Lima to NYC.Wasserstein distance measures the rating movement needed to match regions, while KS measures the largest CDF gap and provides a same-distribution p-value; per-question VLM-average W1 is reported.
6.15 Systematic-Bias MDS (Block 2)
This section maps pairwise Likert-rating disagreements among systems using Block 2 numeric questions, then embeds those dissimilarities in two dimensions to show global disagreement structure across regions and questions.
- Method: The analysis uses Block 2’s 1–10 numeric questions Q6–Q10, repetition t = 1, and parsed ratings, with unparseable answers treated as NaN.It compares systems directly rather than against a human consensus.
- Pairwise dissimilarities: Pairwise dissimilarity is the mean absolute rating gap between two systems over videos within each question and region.nanmean skips videos where either system is NaN; pairs without shared valid videos remain NaN, while the matrix is symmetric with a zero diagonal.
- Pairwise dissimilarities: Question-pooled dissimilarities average rating gaps across all Block 2 questions for each region, and the resulting matrices are displayed as 2×5 heatmaps on a shared color scale.This layout supports comparisons across cities and questions.
- MDS embedding: Classical Torgerson MDS embeds each system as a point in 2D so Euclidean distances reproduce the pairwise dissimilarities.The method uses precomputed dissimilarities and retains the top two eigenvector dimensions after double-centering.
- MDS embedding: Kruskal’s normalized stress-1 evaluates embedding fit, with lower values better and σ < 0.1 indicating a good fit; only relative point positions are meaningful.Axis orientation can differ across panels because configurations are defined up to rotation and reflection.
6.16 LLM-as-a-Judge Agreement
The paper replaces embedding similarity with a prompt-grounded, two-stage LLM judge that first tests answer comparability and then scores signed agreement. Aggregation reports agreement only for comparable pairs while separately tracking comparability rates and excluding failed parses.
- Two-stage judging: The judge uses Qwen3-4B to perform sequential comparability and agreement decisions, excluding Block 2 numeric ratings.The model uses chain-of-thought with T=1.0, top-p 0.95, top-k 20, and presence penalty 1.5.
- Two-stage judging: Stage 1 marks pairs comparable when answers address the same object or event, including disagreements, and non-comparable when they discuss different referents, attributes, or abstentions.Comparable pairs proceed to agreement scoring; pairs that talk past each other are excluded.
- Two-stage judging: Stage 2 assigns scores in {-2,-1,+1,+2}, ranging from direct contradiction to strong agreement, with intermediate scores for degree disagreements or additional unrelated details.More specific wording receives +2 when the core conclusion matches.
- Aggregation: The agreement matrix excludes non-comparable pairs, while each heatmap cell averages J2 scores over comparable items and reports the comparability rate separately.This distinguishes disagreement among shared answers from failure to answer the same question.
- Reliability: Unordered answer pairs are scored once, and rows with JSON-parse failures after up to 3 retries remain pending rather than entering agreement or comparability aggregates.The resulting heatmap is symmetric.