Source-linked AI summary

The 10th AI City Challenge

Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang, Anuj Sharma, Quan Kong, Munkhjargal Gochoo, Jun-Wei Hsieh, Tomasz Kornuta, Zhedong Zheng, Renran Tian, Judah Goldfeder, Fulgencio Navarro, Yuxing Wang, Yizhou Wang, Sameer Satish Pusegaonkar, Anqi Li, Nalin Dadhich, Ridham Kachhadiya, Dhanishtha Patil, Haoquan Liang, Jiajun Li, Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat, Shuyu Yang, Ashutosh Kumar, Rong Wang, Rafael Martin Nieto, Peter Christiansen, Ahmed Abduljawad, Mohanrasu Shanmugam, Nadeem Shaik, Sujit Biswas, Xunlei Wu, Vidya Murali, Rama Chellappa

arXiv:2608.17044v1cs.CVcs.AI

TL;DR

Urban AI benchmarks must test more than isolated recognition, including multimodal reasoning, geometric grounding, temporal evidence, generation, and domain shift. This paper summarizes the 2026 AI City Challenge’s six tracks and two out-of-domain leaderboards, showing that leading systems commonly used modular, evidence-grounded pipelines while transfer, calibration, and reproducibility remained open gaps.

  • Problem

    Modern urban-AI systems can perform strongly in-domain yet remain vulnerable to geographic, camera, and scene-domain shifts, motivating broader benchmark coverage.

  • Method

    The challenge evaluates six primary tracks and two out-of-domain leaderboards spanning geometry, language, temporal evidence, generation, and cross-city detection.

  • Results

    Accepted-paper results show that leading systems commonly use modular pipelines, with language helping multimodal tasks when grounded in visual evidence.

  • Takeaways & Limitations

    The benchmark identifies practical solution patterns while exposing transfer, calibration, and reproducibility gaps that define directions for future editions.

  • Takeaways & Limitations

    Results remain uneven, with cross-city detection, RGB-only 3D perception, video forecasting, and out-of-domain reasoning exposing gaps in transfer, calibration, and reproducibility.

Abstract

from arXiv · show

The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with vehicle detection, classification, and tracking, the challenge has grown into a broad benchmark suite for multi-camera perception, multimodal reasoning, synthetic-to-real learning, generative forecasting, and privacy-preserving evaluation. The 2026 edition continued this growth with 325 registered teams, up from 245 in 2025, and participation from 26 countries and regions, up from 15. Its six primary tracks cover multi-camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city object detection. Track 3 further includes two out-of-domain leaderboards, submitted as Tracks 7 and 8, for fisheye traffic-violation understanding and pedestrian situated-intent VQA. This paper summarizes the challenge setup, datasets, evaluation protocols, leaderboard results, and workshop papers. Across tracks, successful systems combine foundation models with geometric grounding, retrieval or reranking, synthetic-data design, domain adaptation, and controlled inference.

1 Introduction

The 10th AI City Challenge, an ECCV 2026 workshop, marks a decade-long expansion from vehicle-focused tasks into a broad urban-intelligence benchmark suite. Its 2026 edition combined diverse perception, reasoning, generation, and cross-domain evaluation tasks with record participation.

  • Benchmark evolution: Since 2017, the challenge has expanded from vehicle detection, classification, and tracking to multi-camera 3D perception, multimodal reasoning, synthetic-to-real transfer, generative prediction, and privacy-preserving evaluation.The benchmark reflects the evolution from detecting and tracking objects toward reasoning about events and deploying systems in urban, transportation, and physical AI settings.
  • Challenge motivation: The cross-city detection track demonstrated that strong in-domain performance does not guarantee robustness under geographic, camera, and scene-domain shift.This limitation motivates evaluation beyond isolated recognition tasks.
  • Challenge tracks: Six primary tracks covered 3D perception, safety captioning and VQA, anomaly reasoning, person anomaly search, future-frame generation, and cross-city object detection.Track 3 additionally supplied two out-of-domain leaderboards for fisheye traffic-violation understanding and pedestrian situated-intent VQA.
  • Paper organization: The paper presents the challenge setup, datasets, evaluation protocols, leaderboard system, track results, and technical trends from workshop papers.Dataset and benchmark papers are discussed in the dataset section, while results focus on teams with identifiable evaluation-system entries.

2 Challenge Setup

The 2026 AI City Challenge organized six primary tracks across perception, multimodal understanding, retrieval, generation, and cross-city detection, with Tracks 7 and 8 providing optional out-of-domain evaluations for Track 3. Its setup used public and general leaderboards and encouraged foundation-model systems combined with task-specific constraints.

  • Challenge organization: Six primary tracks and eight evaluation-server leaderboards structured the challenge, with Tracks 7 and 8 serving as optional out-of-domain evaluations for Track 3.Public leaderboards targeted award-eligible submissions, while general leaderboards recorded broader participation.
  • Track 1: Multi-Camera 3D Perception: Track 1 required RGB-only test-time multi-camera 3D tracking of people, robots, forklifts, pallet trucks, and humanoids across synchronized warehouse views.Systems predicted per-frame 3D boxes, class labels, scene-level identities, and world-coordinate locations; depth was restricted to training and validation.
  • Transportation understanding: Tracks 2 and 3 focused on transportation safety captioning, VQA, anomaly detection, causal reasoning, explanations, descriptions, and summarization.Track 2 used synthetic Digital Twin WTS data for real-video captioning and questions, while Track 3 extended to fisheye traffic violations and pedestrian intent through Tracks 7 and 8.
  • Track 4: Text-Based Person Anomaly Search: Track 4 posed synthetic-to-real person anomaly retrieval from natural-language descriptions combining pedestrian appearance and behavior.Synthetic image-text pairs supplied training signal, while hidden real-world query-gallery tests measured generalization to abnormal and routine behaviors.
  • Tracks 5–6: Tracks 5 and 6 evaluated generative traffic forecasting and cross-city object detection under temporal, semantic, geographic, and visual domain-shift constraints.Track 5 conditioned future-frame generation on recent frames and textual behavior descriptions; Track 6 used Hafnia-managed training and hidden mixed-city benchmarks.
  • Evaluation design: Across tracks, classical perception metrics such as HOTA and mAP were combined with language, retrieval, generation, and reasoning metrics.The design encouraged foundation models augmented with domain-specific constraints rather than a single model family for every task.

3 Datasets and Tracks

The 2026 challenge comprises six primary tracks spanning multi-camera perception, traffic safety understanding, anomaly reasoning, person retrieval, video forecasting, and cross-city detection, plus two out-of-domain extensions. Its benchmarks emphasize synthetic-to-real transfer, geometric and action grounding, multimodal reasoning, generation quality, and privacy-preserving evaluation.

  • Challenge structure: Six primary tracks and two out-of-domain leaderboards share an evaluation system; Tracks 7 and 8 extend Track 3 to fisheye violations and situated-intent VQA.The following benchmarks specify distinct data settings, hidden-test assumptions, and evaluation focuses.
  • Track 1: Multi-camera 3D perception: Track 1 evaluates multi-camera 3D perception in warehouses under a strict Sim2Real setting, with synthetic training scenes and real hidden-test layouts.Systems must detect, project into a shared 3D frame, and associate identities under occlusion, viewpoint changes, clutter, and an RGB-only test restriction.
  • Track 2: Traffic safety understanding: Track 2 tests synthetic-to-real traffic safety understanding through complementary captioning and VQA outputs grounded in pedestrian and vehicle behavior.Captioning covers event phases, motion, relations, attributes, weather, geometry, and context; VQA tests whether descriptions are grounded in visual evidence.
  • Tracks 3, 7, and 8: Anomaly reasoning: Track 3 provides TAR and TAR-Bench for ten traffic anomaly reasoning task types, while Tracks 7 and 8 test fisheye violations and ambiguous pedestrian intent.TAR contains 44,040 annotations over 3,670 videos, and TAR-Bench contains 960 human-curated annotations for 80 held-out clips.
  • Track 4: Person anomaly search: Track 4 performs text-based person anomaly search from synthetic image-text training pairs to real query-gallery retrieval, requiring action and context grounding beyond appearance.The benchmark targets behaviors such as falling, lying, being hit, and unsafe crossing, with hard negatives challenging retrieval.
  • Tracks 5 and 6: Generation and cross-city detection: Tracks 5 and 6 address future traffic-video generation and privacy-preserving cross-city detection, respectively, evaluating temporal and semantic consistency alongside geographic domain shift.Track 5 combines low-level, perceptual, semantic, and video-distribution metrics; Track 6 uses managed training, validation, and hidden benchmarking on controlled splits.

4 Evaluation Protocols

The evaluation protocols paired track-specific metrics with hidden-test or managed-platform procedures to measure perception, reasoning, retrieval, forecasting, and cross-city robustness. Optional out-of-domain leaderboards extended Track 3 evaluation to fisheye traffic violations and pedestrian situated intent.

  • Leaderboard procedure: Public and general leaderboards reported interim rankings on limited hidden data, then final scores on full hidden test sets after the deadline.The ranking format “1/21 (1/58)” denotes public-team and general-team ranks, respectively.
  • Main-track metrics: Track 1 used 3D HOTA to jointly evaluate detection, association, and localization across classes and scenes in warehouse-coordinate 3D tracks.Submissions included class labels, frame indices, scene identities, and 3D boxes or equivalent localization fields; RGB-only hidden tests prohibited test-time depth.
  • Main-track metrics: Tracks 2–4 evaluated multimodal semantics and retrieval with captioning and VQA metrics, TAR-Bench task scoring, and text-to-image retrieval mAP under synthetic-to-real transfer.Track 2 averaged caption and VQA components; Track 3 combined accuracy-style and text-generation scores; Track 4 rewarded early correct matches while suppressing hard negatives.
  • Main-track metrics: Tracks 5 and 6 measured text-conditioned video forecasting and cross-city detection using visual, perceptual, semantic, distributional, and object-detection metrics.Track 5 combined PSNR, SSIM, LPIPS, CLIP-S, FID, and FVD; Track 6 used hidden Hafnia images across source and target cities through controlled workflows.
  • Out-of-domain evaluation: Tracks 7 and 8 separately reported out-of-domain Track 3 scores for fisheye traffic-violation understanding and pedestrian situated-intent VQA.Both used VQA- and reasoning-oriented scoring related to Track 3, exposing transfer robustness under fisheye geometry and socially ambiguous pedestrian-intent questions.

5 Challenge Results

Challenge results show that competitive systems paired foundation-model capabilities with geometry, modular reasoning, retrieval, synthetic-data design, and domain adaptation. Across Tracks 1–8, accepted entries increasingly relied on task-specific pipelines rather than generic prompting or detector architecture alone.

  • Track 1: Track 1 leaders combined RGB detection with calibration, 2D-to-3D lifting, geometry-aware tracking, and global multi-camera association under hidden-depth evaluation.EVA led accepted public entries, while SKKU-AL-T1 and Playbox showed that online RGB-only pipelines can remain competitive.
  • Track 2: Track 2 narrowed the synthetic-to-real gap by decomposing scene parsing, phase recognition, captioning, VQA answering, and calibration, with explicit spatial and behavioral semantics.Leading systems used VLM backbones, V-JEPA style features, state-bridging, and modular spatial grounding rather than generic captioning alone.
  • Track 3: Track 3 rewarded evidence-driven agentic pipelines that routed heterogeneous questions, matched metrics, maintained event memory, and aligned answers with visual and causal evidence.Stellarview AI led both public and general TAR rankings among accepted papers.
  • Track 4: Track 4 achieved high retrieval scores through text-image embeddings, action-specific reranking, hard-negative handling, and action semantics jointly learned with identity and context.Xiilab.AIpex was the top accepted public/general entry in the final decision sheet.
  • Track 5: Track 5 used history-frame conditioning, language-guided future descriptions, diffusion or world-model priors, and metric-aware selection to generate continuous traffic futures.Accepted systems included Qyn, SSUPER, Latent Painter - UTE, CHTTL_A30, and VGU_ai_lab; close scores indicated rapid progress while semantic consistency remained difficult to summarize with one metric.
  • Tracks 6–8: Tracks 6–8 emphasized robustness beyond a single domain: cross-city detection used pre-training and targeted augmentation, while OOD reasoning transferred evidence-centric or retrieval-augmented agents across violation and intent tasks.SKKU-AL-T1 led accepted Track 6 entries, and UniTraffic ranked first among accepted public entries on both OOD settings.

6 Discussion and Conclusion

The 2026 AI City Challenge spans six primary tracks and two out-of-domain leaderboards, reflecting a shift toward integrated urban-intelligence systems. Its results reveal practical solution patterns while exposing remaining gaps in transfer, calibration, reproducibility, and evidence-grounded multimodal reasoning.

  • Challenge scope: 325 teams from 26 countries and regions evaluated systems across six primary tracks and two out-of-domain leaderboards.The portfolio covers multi-camera 3D perception, synthetic-to-real safety understanding, anomaly reasoning, text-based search, generative forecasting, and cross-city detection.
  • Challenge scope: The track portfolio moves beyond isolated computer-vision predictions toward systems combining geometry, language, temporal evidence, generation, and domain adaptation.This expansion connects the challenge to broader urban intelligence.
  • Findings and limitations: Foundation-model pipelines achieved strong retrieval and reasoning scores, while cross-city detection, RGB-only 3D perception, video forecasting, and out-of-domain reasoning exposed gaps in transfer, calibration, and reproducibility.These remaining failures define research directions for future benchmark progress.
  • Research agenda: Accepted papers commonly used modular pipelines, treated domain shift as a first-class constraint, and grounded language in visual evidence.These patterns form a shared research agenda for deployable urban AI, particularly in Sim2Real, cross-city, and multimodal settings.
Loading 2608.17044v1…