Source-linked AI summary
RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning
Yanbo Jiang, Haotian Zheng, Jiahao Wang, Hanxiao Ren, Yitao Xu, Yining Xing, Zehong Ke, Hao Cheng, Yiqian Tu, Jinhao Li, Zhiyuan Xuan, Fang Zhang, Jianqiang Wang
TL;DR
Roadside sequence understanding lacks reusable supervision for metric 3D tracking and predictive reasoning across camera layouts and sparse events. RISE addresses this with calibration-guided, training-free multi-view tracking and Oracle-grounded bbox-based predictive QA, achieving 66.9 MOTA on 20 reviewed clips while showing benefits from domain adaptation and temporal context. The benchmark also exposes persistent challenges in spatial grounding, future localization, and interaction reasoning.
Problem
Roadside sequences offer spatial and temporal constraints, but reusable supervision remains limited for camera-only 3D tracking and future-aware predictive VQA.
Method
RISE combines calibration-guided multi-view agreement with video identities for persistent 3D tracks and uses a constrained Oracle to create bbox-grounded predictive QA without future cues.
Results
66.9 MOTA is achieved on 20 reviewed clips, while domain adaptation consistently helps and temporal context generally improves structured reasoning across evaluated tasks.
Takeaways & Limitations
RISE provides a unified framework and benchmark for sequence-level roadside understanding while identifying unresolved challenges in spatial grounding, future localization, and interaction reasoning.
Takeaways & Limitations
The evaluation is limited to 20 sampled clips from six intersections, while image-only tracking depends on accurate calibration, reliable segmentation, and sufficient camera overlap.
Abstract
from arXiv · showhide
We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences. For metric tracking, our image-only method combines SAM3 video identities with calibration-guided mask agreement for multi-view identity association, recovering persistent 3D tracks without LiDAR or task-specific 3D training. Its calibration-conditioned geometry allows the procedure to be instantiated at different calibrated multi-camera intersections without layout-specific retraining. On 20 human-reviewed clips from six intersections, the generated tracks achieve 66.9 MOTA within the defined multi-view evaluation scope. For structured vision-language reasoning, a human-reviewed MLLM pipeline mines high-value clips and uses a constrained full-context Oracle to construct bbox-grounded predictive QA without exposing future evidence to evaluated models. The resulting RISE-VQA dataset contains 33,910 QA pairs from 557 clips across 16 intersections and 61 roadside views. Its intersection-held-out RISE-Bench evaluates semantic choices, coordinates, future boxes, and interaction sets with deterministic task-specific metrics. Experiments show consistent benefits from domain adaptation and generally from temporal context, while revealing persistent challenges in spatial grounding, future localization, and interaction reasoning.
Introduction
RISE frames fixed-camera roadside understanding as complementary metric 3D tracking and structured vision-language reasoning tasks, using spatial geometry and temporal event evidence. It introduces training-free calibrated tracking, Oracle-grounded predictive supervision, and intersection-held-out evaluation to address reusable-supervision challenges.
- Motivation: Fixed roadside cameras provide persistent scene-centric observations, while synchronized overlapping views and complete sequences supply geometric and event-evolution evidence.Converting these complementary signals into reusable supervision remains challenging for camera-only 3D detection and predictive VQA.
- Task framework: RISE formulates metric 3D tracking and structured vision-language reasoning as complementary tasks centered on persistent traffic agents.The framework retains task-specific inputs while organizing both capabilities within roadside sequence understanding.
- Tracking branch: RISE combines calibrated multi-view agreement with video identities to construct persistent metric 3D tracks without LiDAR or task-specific 3D training.Its calibration-conditioned procedure can be instantiated at different calibrated multi-camera intersections.
- Evaluation: RISE-Bench evaluates semantic choices, coordinates, future boxes, and interaction sets on unseen intersections using deterministic task-specific metrics.Baselines quantify domain-adaptation and temporal-context effects while exposing limitations in grounding, prediction, and interaction reasoning.
Related Work
Prior roadside perception datasets provide calibrated infrastructure sensing, 3D annotations, and, in some cases, sequential identities and forecasting. RISE differs through calibration-instantiated camera-only tracking and a grounded benchmark spanning temporal reasoning, heterogeneous outputs, and intersection-held-out evaluation.
- Roadside perception datasets: DAIR-V2X, V2X-Seq, RCooper, and Rope3D provide calibrated infrastructure sensors and 3D annotations for vehicle-infrastructure perception.V2X-Seq additionally includes persistent identities, trajectories, vector maps, traffic-light signals, and VIC3D tracking with online and offline forecasting.
- Camera-only roadside 3D methods: Existing camera-only roadside 3D methods use calibration, ground geometry, or BEV fusion but generally require task-specific supervision.RISE instead combines video-consistent segmentation identities with calibration-guided mask agreement and deployment-specific projection geometry.
- Driving and infrastructure VQA: Driving VQA datasets cover vehicle-centric autonomous driving and mixed-traffic video reasoning, while recent work extends VQA toward infrastructure and cooperative settings.Examples include SUTD-TrafficQA, NuScenes-QA, DriveLM, Talk2BEV, RoadSceneVQA, and TUMTraf VideoQA.
- RISE benchmark design: RISE uses newly collected fixed-roadside-camera clips to model short-term traffic evolution with compact visual context and bounding-box references for queried road users.This design complements existing benchmarks in temporal input, target grounding, output structure, and evaluation protocol.
- RISE-Bench evaluation: RISE-Bench requires visual association across frames and evaluates semantic choices, spatial coordinates, future 2D bounding boxes, and set-valued interactions with deterministic metrics.Complete intersections are held out to test generalization to unseen intersection layouts.
Method · Sequence Task Formulation
RISE formulates sequence understanding through two complementary branches: completed-recording 3D tracking from synchronized multi-view temporal evidence, and partial-observation VQA supervised by complete-event annotations. The VQA instances ground agents in observed 2D boxes and evaluate semantic, spatial, future-box, or interaction predictions.
- Sequence Task Formulation: 3D tracking uses four synchronized roadside views from a completed recording to recover persistent agent tracks.The tracking formulation operates on frame t from fixed camera v.
- Sequence Task Formulation: The two tasks use roadside sequences differently: tracking exploits multi-view temporal identity evidence, whereas VQA evaluates an observed prefix with complete-event supervision.Together, these formulations constitute RISE’s sequence-understanding framework.
- Sequence Task Formulation: Each tracked agent is represented by its metric center, dimensions, heading, and sequence-level identity.The identity identifier links agent i across the sequence.
- Sequence Task Formulation: VQA evaluates reasoning under partial observation, processing each view independently to avoid visual-token and cross-view association burdens.The evaluated VLM receives a 20-frame clip from one view rather than multi-view temporal input.
- Sequence Task Formulation: For each 20-frame clip, the evaluated VLM observes only the initial five frames, while the annotation Oracle may inspect all 20 frames.The Oracle uses complete-sequence access solely to determine supervision targets.
- Sequence Task Formulation: Each QA instance grounds a queried agent with an observed 2D box and requests semantic, spatial, future-box, or interaction-set predictions.These targets define the structured reasoning outputs evaluated in the VQA branch.
Identity-Backbone Multi-View 3D Tracking
RISE builds persistent multi-view 3D tracks by converting view-local SAM3 mask IDs into globally consistent identity backbones using calibrated voxel signatures and MWIS, then fitting and refining metric cuboids. The training-free procedure requires accurate calibration and sufficient view overlap, but no LiDAR or layout-specific retraining.
- Identity-backbone tracking: The identity-first design establishes global object identities before metric cuboid fitting, retaining association evidence when some views are missing.This converts temporally consistent within-camera SAM3 IDs into multi-view 3D tracks rather than relying on frame-level 3D detect-then-associate pipelines.
- Identity hypothesis construction: RISE records SAM3 mask IDs over projected 3D voxels to form calibration-conditioned cross-view semantic signatures.Signatures encode either view-local mask IDs or visibility states and propose cross-view object hypotheses when sufficiently supported.
- Identity hypothesis construction: Visibility-aware support and MWIS select non-conflicting identity backbones that assign each view-local ID to at most one global object.The method accommodates mask fragmentation and missing views while downweighting partially occluded evidence rather than discarding it.
- Deployment assumptions: The pipeline is training-free for target 3D tracking and requires neither 3D box supervision nor detector training on collected intersections.It requires accurate calibration and sufficient overlap, while avoiding LiDAR and layout-specific retraining.
- Metric cuboid fitting and refinement: Metric cuboid fitting uses visibility-aware voxel support, volume penalties for unsupported space, and category-specific size and height bounds.Track-level refinement smooths centers, estimates consensus size, corrects headings using motion, and fills short occlusion gaps.
Oracle-Grounded Structured VQA
RISE-VQA uses lightweight MLLM screening and constrained Oracle grounding to construct structured predictive QA from synchronized roadside clips. The protocol preserves an observation–supervision boundary, while human review verifies question quality and prevents future-derived cues from entering model inputs.
- Data and clip construction: 16 intersections and 61 fixed-camera views provide synchronized roadside data, with each VQA clip containing 20 frames sampled at 5 Hz.Thirteen intersections provide four directional views.
- Clip mining: 25,910 candidate clips are screened by a lightweight MLLM for corner-case severity and interaction risk before selected clips are expanded into full 5-Hz sequences.This avoids full-sequence annotation for routine traffic.
- Oracle grounding and boundary: The constrained Oracle derives realized future boxes, maneuvers, and interactions from complete sequences, while training and evaluation expose only I0–I4 and observed-box-anchored questions.Future frames, references, and Oracle reasoning remain hidden from evaluated models.
- Human review: Five annotators made more than 8,000 edits through add, revise, and delete operations, with secondary checking and separate adjudication for uncertain cases.Review focused primarily on dynamic questions involving subtle inter-frame motion and checked wording for future-derived cues.
- Agent grounding: Released QA identifies traffic agents by observed-frame bounding boxes, requiring models to localize each queried agent and maintain correspondence across observed frames.This avoids referring to agents through attributes such as color or type.
Experiments
Experiments evaluate RISE-Bench under intersection-held-out generalization using three task families, three open-source VLMs, and zero-shot versus five-frame LoRA adaptation. Fine-tuning and temporal context improve key dynamic, localization, and interaction outcomes, while clip-random splitting benefits from shared intersection context.
- Evaluation protocol: RISE-Bench holds out three complete intersections spanning eleven roadside views and evaluates semantic, coordinate, and interaction-set questions.The protocol separates complete intersections to avoid shared layout and camera context and measure cross-intersection generalization.
- Models and training: The evaluation compares Qwen2.5-VL-7B, InternVL3-8B, and MiniCPM-V-4.5-8B under zero-shot and five-frame LoRA fine-tuning.Training and evaluation observe only the first five frames; fine-tuning uses 29,219 training instances with the vision encoder and multimodal projector frozen.
- Adaptation results: Domain-specific adaptation consistently improves all three open-source backbones across answer formats, with Qwen2.5-VL-7B leading aggregate MCQ and Line-F1 and InternVL3-8B leading dynamic, detection, future-box, and interaction metrics.Native bbox serialization and coordinates are normalized before scoring for rows marked by ‡, so the gains are not solely output-format alignment.
- Temporal context: Across all three backbones, 5f improves D-MCQ and T-IoU while reducing C-ADE, whereas S-MCQ changes little for Qwen2.5-VL and MiniCPM-V.InternVL3 shows the largest trajectory improvement, indicating backbone-dependent temporal gains.
- Split sensitivity: On the common 1,451-QA subset, clip-random training exceeds intersection-held-out training on S-MCQ (.841 vs. .772), D-MCQ (.796 vs. .775), and Det-F1 (.654 vs. .574).The result indicates that shared intersection context makes clip-level splitting easier and motivates intersection-held-out evaluation.
3D Tracking Quality Analysis
3D tracking quality is assessed against manually corrected ground truth under a fixed, three-view evaluation scope. The protocol uses calibration-reconfigured tracking across six intersections, with geometric matching separating detection, localization, and identity errors.
- Ground-truth construction: Manual review corrects initial 3D tracks by adding missed vehicles, removing false tracks, and fixing box geometry and identities.Synchronized views and semantic-voxel BEV support the corrections.
- Evaluation scope: Eligible ground-truth vehicles must lie within the voxel region and project within image bounds in at least three synchronized cameras.Tiny or distant vehicles are excluded, and scope exclusions are not counted as tracker false negatives.
- Evaluation scope: The evaluation covers 20 sampled clips from six intersections, with 20 frames reviewed per clip at 2 Hz.The fixed ROI and three-view coverage scope constrain evaluation scale, while calibration-dependent projections and site-specific search regions are reconfigured per site.
- Matching and metrics: Class-aware frame-wise Hungarian matching uses oriented BEV IoU, rejecting matches below 0.5 to count unmatched predictions and ground-truth boxes as FP and FN.Geometric association avoids inherited track IDs and enables separate evaluation of detection, localization, and identity consistency.
- Evaluation scope: A vehicle omitted near an overlap boundary is treated as outside the predefined three-view scope rather than as a tracker false negative.This illustrates how the protocol distinguishes scope limitations from tracking errors.
Conclusion
RISE unifies metric 3D tracking and structured vision-language reasoning for roadside sequences, achieving 66.9 MOTA on 20 clips within its defined multi-view scope. Its limitations include dependence on calibration, segmentation, camera overlap, and short single-view clips, motivating longer-horizon and synchronized multi-view reasoning.
- Conclusion: RISE combines metric 3D tracking with structured vision-language reasoning for roadside sequence understanding.The framework exploits complementary spatial and temporal evidence for persistent 3D tracking and bbox-grounded prediction without exposing future evidence to evaluated models.
- Conclusion: 66.9 MOTA is achieved by generated tracks on 20 clips within the defined multi-view evaluation scope.The evaluation is human-corrected.
- Limitations: Image-only 3D tracking depends on accurate calibration, reliable segmentation, and sufficient camera overlap.Geometric support weakens under heavy occlusion and near shared-view boundaries.
- Limitations: Current 3.8-second clips primarily assess near-term reasoning from individual roadside views.The passage identifies longer observation horizons with preserved temporal resolution and coordinated reasoning across synchronized views as future directions.