Paper deep dive
RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning
Yanbo Jiang, Haotian Zheng, Jiahao Wang, Hanxiao Ren, Yitao Xu, Yining Xing, Zehong Ke, Hao Cheng, Yiqian Tu, Jinhao Li, Zhiyuan Xuan, Fang Zhang, Jianqiang Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/23/2026, 2:14:14 AM
Summary
The paper introduces RISE, a framework for roadside infrastructure sequence understanding that combines metric 3D tracking and structured vision-language reasoning. For tracking, it uses a training-free, image-only method leveraging SAM3 video identities and calibration-guided mask agreement to generate persistent 3D tracks across multi-camera intersections without LiDAR. For reasoning, it presents RISE-VQA, a dataset of 33,910 bbox-grounded QA pairs derived from 557 clips using a constrained Oracle to ensure predictive supervision without future leakage. The framework also introduces RISE-Bench, an intersection-held-out evaluation benchmark assessing semantic choices, coordinates, future boxes, and interactions.
Entities (10)
Relation Signals (8)
RISE â includescomponent â RISE-VQA
confidence 95% ¡ The resulting RISE-VQA dataset contains 33,910 QA pairs... Its intersection-held-out RISE-Bench evaluates...
RISE â includescomponent â RISE-Bench
confidence 95% ¡ Its intersection-held-out RISE-Bench evaluates semantic choices...
RISE â usesmodel â SAM3
confidence 93% ¡ our image-only method combines SAM3 video identities with calibration-guided mask agreement
RISE â developedby â Tsinghua University
confidence 92% ¡ Yanbo Jiang... Tsinghua University... We present RISE
Qwen2.5-VL-7B â evaluatedon â RISE-Bench
confidence 90% ¡ We evaluate three open-source VLMsâQwen2.5-VL-7B... on RISE-Bench
InternVL3-8B â evaluatedon â RISE-Bench
confidence 90% ¡ We evaluate three open-source VLMs... InternVL3-8B... on RISE-Bench
MiniCPM-V-4.5-8B â evaluatedon â RISE-Bench
confidence 90% ¡ We evaluate three open-source VLMs... MiniCPM-V-4.5-8B... on RISE-Bench
RISE â usesalgorithm â MWIS
confidence 88% ¡ uses MWIS to select non-conflicting identity backbones
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences. For metric tracking, our image-only method combines SAM3 video identities with calibration-guided mask agreement for multi-view identity association, recovering persistent 3D tracks without LiDAR or task-specific 3D training. Its calibration-conditioned geometry allows the procedure to be instantiated at different calibrated multi-camera intersections without layout-specific retraining. On 20 human-reviewed clips from six intersections, the generated tracks achieve 66.9 MOTA within the defined multi-view evaluation scope. For structured vision-language reasoning, a human-reviewed MLLM pipeline mines high-value clips and uses a constrained full-context Oracle to construct bbox-grounded predictive QA without exposing future evidence to evaluated models. The resulting RISE-VQA dataset contains 33,910 QA pairs from 557 clips across 16 intersections and 61 roadside views. Its intersection-held-out RISE-Bench evaluates semantic choices, coordinates, future boxes, and interaction sets with deterministic task-specific metrics. Experiments show consistent benefits from domain adaptation and generally from temporal context, while revealing persistent challenges in spatial grounding, future localization, and interaction reasoning.
Tags
Links
- Source: https://arxiv.org/abs/2608.16480v1
- Canonical: https://arxiv.org/abs/2608.16480v1
Trouble viewing inline? Open PDF directly â
Full Text
36,901 characters extracted from source content.
Expand or collapse full text
RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning Yanbo Jiang 1 , Haotian Zheng 1 , Jiahao Wang 1 , Hanxiao Ren 1 , Yitao Xu 1 , Yining Xing 1 , Zehong Ke 1 , Hao Cheng 1 , Yiqian Tu 1 , Jinhao Li 1 , Zhiyuan Xuan 2 , Fang Zhang 1â , Jianqiang Wang 1â 1 School of Vehicle and Mobility, Tsinghua University 2 TsingCloud Abstract We present RISE (Roadside Infrastructure Sequence Under- standing and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in roadside sequences. For metric tracking, our image-only method com- bines SAM3 video identities with calibration-guided mask agreement for multi-view identity association, recovering per- sistent 3D tracks without LiDAR or task-specific 3D train- ing. Its calibration-conditioned geometry allows the proce- dure to be instantiated at different calibrated multi-camera in- tersections without layout-specific retraining. On 20 human- reviewed clips from six intersections, the generated tracks achieve 66.9 MOTA within the defined multi-view evaluation scope. For structured vision-language reasoning, a human- reviewed MLLM pipeline mines high-value clips and uses a constrained full-context Oracle to construct bbox-grounded predictive QA without exposing future evidence to evaluated models. The resulting RISE-VQA dataset contains 33,910 QA pairs from 557 clips across 16 intersections and 61 roadside views. Its intersection-held-out RISE-Bench evaluates seman- tic choices, coordinates, future boxes, and interaction sets with deterministic task-specific metrics. Experiments show consis- tent benefits from domain adaptation and generally from tem- poral context, while revealing persistent challenges in spatial grounding, future localization, and interaction reasoning. Introduction Fixed roadside cameras are a key sensing layer for intelligent transportation systems, supporting intersection monitoring, traffic-safety analysis, and infrastructure-assisted perception through persistent, scene-centric observations (Yu et al. 2023). Turning these sequences into actionable traffic un- derstanding requires two complementary capabilities: metric 3D tracking establishes persistent identities, 3D states, and trajectories, while structured vision-language reasoning pro- vides queryable interpretations of agent states, maneuvers, interactions, and future outcomes. Prior work has largely studied these capabilities separately, emphasizing metric tracking and trajectory forecasting in vehicleâinfrastructure sequences (Yu et al. 2023) or language-based traffic-scene interpretation in driving VQA (Xu, Huang, and Liu 2021; Qian et al. 2024; Sima et al. 2024). These capabilities draw on complementary spatial and temporal evidence in fixed roadside sequences. Spatially, â Corresponding authors: zhang_fang,wjqlws@tsinghua.edu.cn synchronized cameras with overlapping fields of view pro- vide multi-view geometric constraints for metric 3D estima- tion; temporally, complete sequences capture event evolu- tion and outcomes needed for predictive reasoning. Yet con- verting this information into reusable supervision remains challenging. Camera-only roadside 3D detection often re- quires task-specific 3D labels or adaptation to new camera layouts (Yang et al. 2023, 2024; Zhang et al. 2025). For predictive VQA, high-value events are sparse, and realized outcomes must provide supervision without exposing future evidence to evaluated models. Scaling roadside collection across intersections further requires access, calibration, and synchronization at each site. To address these challenges, we introduce RISE (Roadside Infrastructure Sequence Understanding and Evaluation), which exploits complementary spatial and temporal evidence in fixed-camera roadside sequences. Its tracking branch com- bines calibrated multi-view agreement with video identities to construct persistent metric 3D tracks, while its vision- language branch converts realized event outcomes into bbox- grounded predictive supervision through a constrained Ora- cle. Together, the two branches organize roadside sequence understanding around persistent traffic agents, covering met- ric state estimation and semantic event interpretation with task-specific inputs. Our contributions are: 1. A roadside sequence task framework. We formulate metric 3D tracking and structured vision-language rea- soning as complementary tasks centered on persistent traffic agents while retaining task-specific inputs. 2. Training-free multi-view 3D tracking. We combine SAM3 view-local identities with calibration-guided mask agreement, MWIS-based association, and track-level re- finement to recover persistent metric 3D tracks with- out LiDAR or task-specific 3D training. The calibration- conditioned procedure can be instantiated at different cal- ibrated multi-camera intersections, and its outputs initial- ize subsequent human refinement and quality assessment. 3. RISE-VQA with Oracle-grounded predictive supervi- sion. We construct 33,910 bbox-grounded QA pairs from 557 clips across 16 intersections and 61 roadside views. A human-reviewed MLLM pipeline uses a constrained full-context Oracle to generate predictive targets without arXiv:2608.16480v1 [cs.CV] 17 Aug 2026 exposing future evidence to evaluated models. 4. Intersection-held-out structured evaluation. We intro- duce RISE-Bench to evaluate semantic choices, coordi- nates, future boxes, and interaction sets on unseen inter- sections using deterministic task-specific metrics. Base- lines quantify domain-adaptation and temporal-context effects while exposing limitations in grounding, predic- tion, and interaction reasoning. Related Work Roadside 3D Perception and Tracking Roadside perception datasets such as DAIR-V2X (Yu et al. 2022), V2X-Seq (Yu et al. 2023), RCooper (Hao et al. 2024), and Rope3D (Ye et al. 2022) provide calibrated infrastructure sensors and 3D annotations for vehicle-infrastructure percep- tion. V2X-Seq further augments sequential frames with per- sistent identities, trajectories, vector maps, and traffic-light signals, and defines VIC3D tracking together with online and offline forecasting (Yu et al. 2023). I-24 3D provides continuous vehicle trajectories from overlapping calibrated highway cameras (Gloudemans et al. 2023). RISE shares this sequence-centric perspective but studies a different setting: fixed multi-camera, image-only 3D tracking together with structured vision-language reasoning over roadside events. Camera-only roadside 3D methods exploit camera cali- bration, ground geometry, or BEV fusion (Yang et al. 2023, 2024; Zhang et al. 2025), but generally learn 3D percep- tion from task-specific supervision. In contrast, our method uses video-consistent segmentation identities as temporal ev- idence and calibration-guided mask agreement as cross-view geometric evidence. Its projection geometry is instantiated directly from each deploymentâs calibration, allowing the same procedure to operate at a new calibrated multi-camera intersection with sufficient overlap, without training a layout- specific 3D model. Structured Vision-Language Reasoning for Roadside Scenes Driving-oriented VQA datasets include SUTD- TrafficQA (Xu, Huang, and Liu 2021), NuScenes- QA (Qian et al. 2024), DriveLM (Sima et al. 2024), and Talk2BEV (Choudhary et al. 2024). NuScenes- QA and DriveLM construct QA from vehicle-centric autonomous-driving data, while SUTD-TrafficQA targets video reasoning over mixed traffic events. Recent work extends this scope to infrastructure and cooperative settings: RoadSceneVQA (Guan et al. 2026) studies image-level roadside reasoning, TUMTraf VideoQA (Zhou et al. 2025) targets spatiotemporal roadside video understanding, V2X-QA (You et al. 2026) evaluates cooperative VQA on V2X-Seq, and LTD/UniVLT (Huang et al. 2026) considers heterogeneous multi-camera reasoning tasks. As summarized in Table 1, RISE complements existing benchmarks in temporal input, target grounding, output struc- ture, and evaluation protocol. It uses newly collected clips from fixed roadside cameras to capture short-term traffic evolution while keeping the visual context compact. Each queried road user is referenced by its observed bounding box rather than a descriptive phrase, requiring the model to visually ground and associate the agent across frames. The answers include semantic choices, spatial coordinates, fu- ture 2D bounding boxes, and set-valued interactions. RISE- Bench scores these heterogeneous outputs with deterministic task-specific metrics and holds out complete intersections to evaluate generalization to unseen intersection layouts. Method Sequence Task Formulation Let I (v) t denote frame t from fixed camera v. Given a com- pleted roadside recording, the 3D tracking branch uses four synchronized views to recover T =Ď i , Ď i =(b 3D i,t , id i ) t ,(1) whereb 3D i,t contains the metric center, dimensions, and head- ing of agenti, and id i identifies the agent across the sequence. The VQA branch instead evaluates reasoning under partial observation. Each view is processed independently to avoid the visual-token and cross-view association burden of multi- view temporal input. For each 20-frame clip, the evaluated VLM observes only I (v) 0 âI (v) 4 , whereas the annotation Or- acle may inspect the complete sequence I (v) 0 âI (v) 19 solely to determine supervision targets. Each QA instance grounds a queried agent with an observed 2D box and requires a seman- tic choice, spatial coordinates, future 2D boxes, or an inter- action set. The two tasks use roadside sequences differently: 3D tracking exploits multi-view temporal identity evidence over a completed recording, whereas VQA evaluates an ob- served prefix using supervision derived from the complete event. Together, they form RISEâs sequence-understanding framework. Identity-Backbone Multi-View 3D Tracking SAM3 provides temporally consistent mask IDs within each camera, but these IDs are not shared across views. RISE con- verts them into calibration-conditioned cross-view signatures and uses MWIS to select non-conflicting identity backbones, establishing global object identities before metric cuboid fit- ting. Unlike frame-level 3D detect-then-associate pipelines, this identity-first design turns SAM3âs temporal continuity into multi-view 3D tracks and retains association evidence when some views are occluded. Video IDs and Calibrated Semantic Voxels Given four synchronized and calibrated views, SAM3 (Carion et al. 2025) produces category-specific instance masks with tem- porally consistent IDs within each view. We additionally mark static occluders to distinguish occlusion from miss- ing segmentation. We project a cached 3D voxel grid into each image and record the mask ID covering each projected voxel. For voxel n at time t, the cross-view tuple Ď n,t = (s n,1,t ,...,s n,V,t )(2) forms a semantic signature, where s n,v,t denotes either a view-local mask ID or a visibility state in camera v. Suffi- ciently supported signatures propose cross-view object hy- potheses. Spatially adjacent hypotheses sharing at least one Table 1: Comparison with traffic-scene vision-language datasets. âInheritedâ denotes 3D labels from a source dataset; âImg. trackerâ denotes the image-derived tracking component proposed in RISE. DatasetVenue ViewUnitSource QA Gen. #QA Scoring 3D Comp. NuScenes-QA (Qian et al. 2024)AAAIâ24 ego Multi-view nuScenes Template 460k AccuracyInherited DriveLM (Sima et al. 2024)ECCVâ24 egoFrame nuScenes Temp.+Human 443k SPICE/GPT Inherited LingoQA (Marcu et al. 2024)ECCVâ24 egoVideoSelf Human+LLM 420k Lingo-Judgeâ SUTD-TrafficQA (Xu, Huang, and Liu 2021) CVPRâ21 mixed Video MixedHuman 62.5kAcc.â TUMTraf VideoQA (Zhou et al. 2025)ICMLâ25 roadside VideoSelfLLM+Ver. 85.0kAcc.â RoadSceneVQA (Guan et al. 2026)AAAIâ26 roadside Image Rope3D LLM+Ver. 34.7k Text/GPT Inherited V2X-QA (You et al. 2026)arXivâ26 V2X View pair V2X-Seq LLM+Ver. 33.2kAcc.Inherited LTD (Huang et al. 2026)arXivâ26 roadside Multi-img SelfLLM+Ver. 11.6k GPT/Acc.â RISE (Ours)âroadside ClipSelfLLM+Ver. 33.9k Task-specific Img. tracker Training-Free Multi-View 3D Detection Real roadside evidence on the left; geometric reasoning on the right 1 Shared-Coordinate Calibration â˘Cross-view landmark reprojection â˘Reusable fixed-camera calibration 2 SAM3 Video Mask Propagation â˘Prompt-conditioned video segmentation â˘Persistent object tubes through occlusion 3 Ground-Constrained Semantic Voxel Voting â˘Ground-constrained voxel hypotheses â˘Multi-view mask support and contradiction â˘Visibility-weighted semantic consensus 4 View-Signature Association via MWIS â˘View-signature track hypotheses â˘Conflict-aware association graph â˘Globally consistent MWIS selection 5 Track-Level Cuboid Refinement â˘Visibility-weighted track- level fitting â˘Temporal consistency in size, position, and heading Figure 1: Identity-first multi-view 3D tracking in RISE. View-local SAM3 video IDs are grouped into global identity backbones through calibrated voxel signatures and MWIS, followed by metric cuboid fitting and track-level refinement. view-local ID are merged into candidate identity backbones, accommodating mask fragmentation and missing view com- ponents. Visibility-aware voting distinguishes visible, out- of-view, and occluded projections, so partially occluded evi- dence is downweighted rather than removed. Because SAM3 IDs are view-local, different candidate backbones may reuse the same ID and therefore cannot all be valid. We therefore construct a conflict graph and solve a maximum-weight inde- pendent set (MWIS), using visibility-aware voxel support as candidate weights. The selected non-conflicting backbones assign each view-local ID to at most one global object and define the global IDs used for tracking. The pipeline is training-free for the target 3D tracking task, requiring neither 3D box supervision nor detector training on the collected intersections. It requires accurate calibra- tion and sufficient overlap, but no LiDAR or layout-specific retraining. At a new intersection, only the cached voxel pro- jections and site-specific search region are rebuilt from its calibration, while the tracking procedure and parameters re- main unchanged. Frame-Level Metric Cuboid Fitting After MWIS selects a consistent set of identity backbones, we optimize a metric cuboidB for each object using J (B) = X nâB w n â Îť V B â 3 ,(3) where w n is the visibility-aware voxel support, V B is cuboid volume, â is the voxel-grid spacing, and the volume term pe- nalizes unsupported space. Category-specific size and height bounds remove implausible cuboids. Track Construction and Temporal Refinement Each MWIS-selected backbone links the SAM3 identities of one physical object across cameras. Because its component iden- tities are video-consistent within their views, the resulting global identity propagates across frames even when one com- ponent temporarily disappears. Track-level refinement treats per-frame cuboids as noisy observations: it smooths centers, estimates a consensus size, corrects unstable headings using motion, and fills short gaps caused by occlusion. Each output track therefore contains a persistent ID and a time-indexed sequence of metric 3D boxes. Oracle-Grounded Structured VQA Figure 2 summarizes the constrained-Oracle workflow used to construct RISE-VQA. Data Collection and Scene Mining We collect synchro- nized roadside videos from 16 urban intersections, yielding 61 fixed-camera views; 13 intersections provide four direc- tional views. Each selected VQA clip contains 20 frames sampled at 5 Hz. The question and queried-agent box are 1 Roadside Video Pool 2 Reusable Human Priors 3 Mining & Reference Agents 4 Oracle VQA Teacher 5 RISE-VQA + RISE-Bench Diverse Roadside Scenarios Initialized once per camera Lane boundaries Stop lines Traffic-light Agent 20-frame Context 29,219 train 4,691 benchmark Human Review Traffic-light candidates 557 clips 61 roadside views Lane centerlines Fixed Roadside Views N S E W Vehicle & VRU Detector Corner Case Miner clip 4 Key Frames Rare Event Score Interaction Risk Score Shared across its clips +210 added 8,169 revised -689 deleted 25.9K Candidate Clips 11020 Vehicle boxes Light states QAReason Oracle Rules Answer from full context Reason cites first 5 frames No future leakage Expert Oracle VQA Prompt 33,910 QAs Figure 2: Structured VQA production in RISE. The pipeline mines high-value clips, constructs frame-level references, generates Oracle-grounded predictive QA pairs, and applies targeted human review. anchored at I 0 , I 1 âI 4 provide observed motion context, and later key frames supply Oracle-only targets. To avoid spending full-sequence annotation on routine traffic, a lightweight MLLM scores low-rate summaries of 25,910 candidate clips by corner-case severity and interac- tion risk. Selected clips are then expanded into full 5-Hz sequences for structured QA construction. Infrastructure Priors and Structured References Di- rectly asking a single MLLM to localize small traffic sig- nals and bind road users across a long clip is unreliable. Fixed roadside views allow us to reuse infrastructure priors across clips, including signal regions, lane geometry, and stop lines. Specialized reference modules recognize cropped signal states and localize road users with coarse headings on selected key frames. These structured references provide localization and association cues only to the annotation Or- acle; they are never exposed to models during training or evaluation. ObservationâSupervision Separation Question tem- plates are restricted to evidence available in the observed prefix. For predictive QA, the Oracle uses the complete se- quence and structured references to derive realized future boxes, maneuvers, and interactions. Training and evaluation receive onlyI 0 âI 4 and questions anchored to observed boxes; future frames, references, and Oracle reasoning remain hid- den. Human Review and Visual Grounding Five annotators review the generated QA through add, revise, and delete op- erations, producing more than 8,000 edits, primarily on dy- namic questions involving subtle inter-frame motion. Each question type undergoes first-pass review and secondary checking, with uncertain cases adjudicated separately. Hu- man review also enforces the observation boundary by check- Table 2: Observationâsupervision boundary in the con- strained Oracle protocol. Realized future observations de- termine predictive targets but remain hidden from train/test inputs. InformationAnnotation Role Train/Test Input Early frames I 0 âI 4 EvidenceYes Future frames I 5 âI 19 Oracle evidenceNo Future boxesSupervision targetNo Infrastructure priorsReferenceNo ing that question wording contains no future-derived cues. Released QA refers to traffic agents by their observed-frame bounding boxes rather than attributes such as color or type. Models must therefore localize the queried agent from the box and maintain its correspondence across the observed frames before answering. Experiments Dataset and RISE-Bench The reviewed corpus contains 34,545 QA pairs, includ- ing 11,846 static-environment and 22,699 dynamic-behavior questions. We hold out three complete intersections, covering eleven roadside views and both cross- and T-junction layouts. By separating complete intersections rather than clips, this protocol avoids shared layout and camera context and mea- sures cross-intersection generalization (Sekaran et al. 2025), as validated in Table 7. From the 5,326 held-out instances, RISE-Bench retains 4,691 with deterministic task-specific metrics and excludes 635 open-ended or weakly structured instances. Together with 29,219 training instances, the re- leased dataset contains 33,910 QA pairs. Figure 3: Examples of bbox-grounded structured VQA. Boxes identify queried agents and future locations, while col- ored geometry provides infrastructure and interaction con- text. Table 3: RISE-VQA and RISE-Bench statistics. MetricValue MetricValue Intersections16 Roadside views61 4-view / 3-view 13 / 3 Event clips557 Reviewed QA34,545 Released QA33,910 Training QA29,219 Held-out Pool5,326 RISE-Bench QA 4,691 Reviewed Static11,846 (34.3%) Reviewed Dynamic 22,699 (65.7%) Table 4: RISE-Bench composition. Task Family#QAMetric Multiple choice3,794Choice score Coordinate output 745IoU / F1 / C-ADE Set prediction152 Precision / Recall / F1 Total4,691â RISE-Bench has three task families: multiple-choice ques- tions cover semantic traffic understanding; coordinate ques- tions cover infrastructure geometry and future boxes; and set-valued questions identify interacting agents or pairs. Models and Metrics We evaluate three open-source VLMsâQwen2.5-VL-7B, InternVL3-8B, and MiniCPM-V-4.5-8Bâunder zero-shot (ZS) and fine-tuned (FT) settings. FT denotes five-frame LoRA adaptation, and both training and evaluation observe only the first five frames of each clip. LoRA experiments use LLaMA-Factory (Zheng et al. 2024) with rank 8, scaling factor 16, learning rate 5Ă 10 â5 , effective batch size 32, and three epochs on 29,219 training instances; the vision en- coder and multimodal projector remain frozen. Primary runs use direct-answer inference with thinking disabled, while thinking-enabled proprietary models are included as strong zero-shot references marked by â . The ZSâFT comparison measures the effect of domain-specific adaptation, and the 1fâ5f ablation measures the contribution of temporal con- text. Choice score assigns full, partial, or zero credit using pre- defined mappings for each question type; partial credit ap- plies only to ordinal or adjacent choices. Det-F1 matches pre- dicted and reference boxes by IoU, while Line-F1 matches lane and stop-line segments by endpoint distance. Future boxes are evaluated by trajectory IoU (T-IoU), which av- erages box overlap across future steps, and center average displacement error (C-ADE), C-ADE = 1 3N N X i=1 3 X Ď=1 âĽ Ë c i,Ď âc i,Ď âĽ 2 .(4) Herec i,Ď is the ground-truth box center in the common 1000Ă 1000 relative coordinate space. Interaction predic- tions use label-agnostic bbox set matching, and Inter-F1 is macro-averaged over QA samples. Structured Vision-Language Results Table 5 shows that domain-specific adaptation consistently improves all three open-source backbones across answer for- mats. For rows marked by ⥠, native bbox serialization and coordinates are normalized before scoring, so the gains can- not be attributed solely to output-format alignment. Among the fine-tuned models, Qwen2.5-VL-7B leads in aggregate MCQ and Line-F1, whereas InternVL3-8B performs best on dynamic MCQ, detection, future-box, and interaction met- rics. This variation supports separately evaluating semantic choices, geometric grounding, future localization, and in- teraction sets. Thinking improves both paired proprietary baselines but does not dominate every structured task. Temporal Context Ablation Across all three backbones, 5f improves D-MCQ and T-IoU while reducing C-ADE, providing consistent evidence that observed motion benefits dynamic reasoning and future lo- calization. Because static questions are anchored in the first frame shared by both settings, S-MCQ provides a useful ref- erence: it changes little for Qwen2.5-VL and MiniCPM-V, while improving modestly for InternVL3. The gain remains backbone-dependent, with InternVL3 showing the largest trajectory improvement. We retrain Qwen2.5-VL-7B under the intersection-held- out and clip-random train/test partitions. Because the full test sets differ in composition, their rows provide only an overall view. On the common 1,451-QA subset, clip-random train- ing is stronger across all metrics, including S-MCQ (.841 vs. .772), D-MCQ (.796 vs. .775), and Det-F1 (.654 vs. .574). Shared intersection context therefore makes clip-level split- ting easier, motivating intersection-held-out evaluation. 3D Tracking Quality Analysis We evaluate the unedited tracker output as the prediction and the manually corrected tracks as GT. As shown in Fig- ure 5, the tracker provides initial 3D tracks, while synchro- nized views and the semantic-voxel BEV help annotators add missed eligible vehicles, remove false tracks, and correct box geometry and identity. (a) Question word cloud 1234567891011121314 Sub-category index 11.2% 10.8% 5.9% 3.4% 2.4% 0.5% 18.9% 13.8% 9.2% 7.1% 6.9% 4.3% 3.6% 1.8% (b) Question sub-category distribution Static perception 1: Traffic Lights (3,856) 2: Road Structure (3,738) 3: Background Environment (2,054) 4: Special Element Inference (1,178) 5: Visibility and Occlusion (837) 6: Traffic Signs (183) Dynamic reasoning 7: Intent & Trajectory Prediction (6,545) 8: Refined Agent States (4,767) 9: Conflict Anticipation (3,172) 10: Anomaly & Counterfactual Reasoning (2,448) 11: Causal Interaction Graph (2,392) 12: Driving Advice (1,488) 13: Macro Traffic Flow Analysis (1,258) 14: VRU Interaction (629) Figure 4: Distribution of the 34,545 human-reviewed QA pairs before benchmark filtering. Left: frequent concepts in the question corpus. Right: fine-grained sub-category counts. Table 5: Main results on RISE-Bench. S-MCQ and D-MCQ denote static and dynamic MCQ scores. ZS-5f and FT-5f denote five-frame zero-shot inference and LoRA adaptation, respectively. ⥠marks bbox evaluation normalized from model-native seri- alization and coordinate systems; â marks thinking-enabled proprietary zero-shot runs. Other rows use direct-answer inference. Bold indicates the best result within the proprietary-reference or fine-tuned group. ModelAccess SettingS-MCQ D-MCQ MCQDet-F1 Line-F1T-IoU C-ADEâInter-F1 Qwen2.5-VL-7B Open ZS-5f ⥠.479.533 .510.078.001.305 120.4.066 InternVL3-8BOpen ZS-5f ⥠.503.424 .458.029.029.050 306.8.015 MiniCPM-V-4.5-8B Open ZS-5f ⥠.507.571 .544.061.055.257 131.3.042 GPT-5.5Closed ZS-5f.558.593 .578.392.798.248 117.8.106 GPT-5.5 â Closed ZS-5f.610.632 .623.649.920.402 60.8.238 Gemini-3.1-Pro â Closed ZS-5f ⥠.502.491 .495.187.719.323 112.9.214 Qwen3.7-PlusClosed ZS-5f.635.649 .643.382.738.311 99.1.251 Qwen3.7-Plus â Closed ZS-5f.640.666 .655.599.791.374 64.2.332 Qwen2.5-VL-7B Open FT-5f.782.781 .781.643.860.381 74.8.271 InternVL3-8BOpen FT-5f.734.791 .767.672.845.433 62.2.350 MiniCPM-V-4.5-8B Open FT-5f.722.764 .746.648.771.344 91.2.305 Table 6: Temporal-context ablation. Each pair uses the same backbone and LoRA recipe; only the number of input frames differs. ModelFr. S-MCQ D-MCQ T-IoU C-ADEâ Qwen2.5-VL 1f.784.730.312 109.1 5f.782.781.38174.8 InternVL3 1f.698.738.324 104.0 5f.734.791.43362.2 MiniCPM-V- 1f.727.730.322 115.1 4.5-8B5f.722.764.34491.2 All evaluated scenes use the same voxelized world- coordinate region, x,y â [â40, 40] m, with an approxi- mately 0.2 m grid step. We regard a vehicle as observable when it is clearly identifiable at sufficient scale in at least one camera; tiny or distant vehicles are excluded. Among observ- Table 7: Split-policy ablation with Qwen2.5-VL-7B. Held and Clip denote intersection-held-out and clip-random train/test splitting, respectively. All denotes their respective test sets (4,691 and 4,359 QA pairs); Com. denotes the same 1,451 QA pairs absent from both training sets. Eval. Split S-MCQ D-MCQ Det-F1 Line-F1 T-IoU C-ADEâ All Held .782.781 .643 .860 .381 74.8 All Clip .857.806 .754 .944 .379 62.0 Com. Held .772.775 .574 .846 .399 71.0 Com. Clip .841.796 .654 .896 .420 57.0 able vehicles, those inside the voxel region whose projected regions lie within the image bounds of at least three synchro- nized cameras enter the frame-level tracking GT, regardless of occlusion. The remainder are recorded as scope exclusions rather than tracker false negatives; tracker false negatives are unmatched eligible GT boxes. For each 10-second clip, an- Figure 5: Human review interface for initialized 3D tracks. Multi-view projections and the semantic-voxel BEV guide manual correction. Table 8: Statistics of the human-reviewed subset. Dimen- sions are adjusted once per track; a frame requires no frame- wise correction when its XY center and heading remain un- changed. MetricValue MetricValue Generated boxes 3,240 GT boxes3,317 GT tracks251 Scope-excl. inst.1,158 Exclusion rate25.9% No frame-wise edit 1,900 (57.3%) Table 9: 3D tracking quality at a BEV IoU matching threshold of 0.5. MOTP is the mean BEV IoU over valid matches. MetricValueMetricValue MOTAâ66.9MOTP BEV â87.0 IDSâ40F1â83.9 Precision / Recallâ84.9 / 82.9XY err. (m)â.071 notators review 20 frames sampled at 2 Hz. To keep manual correction tractable, the current evaluation uses 20 sampled clips from six intersections within the fixed ROI and three- view coverage scope. All intersections use the same tracking parameters; only calibration-dependent voxel projections and site-specific search regions are reconfigured. This clip-based protocol limits the scale of the present evaluation rather than the temporal duration supported by the tracker. PredictionâGT association is determined geometrically rather than by inherited track IDs, allowing detection errors, localization errors, and identity consistency to be evaluated separately. Specifically, we perform class-aware frame-wise Hungarian matching using oriented BEV IoU; matches with BEV IoU below 0.5 are rejected, and the remaining un- matched predictions and eligible GT boxes are counted as FP and FN, respectively. Figure 6 shows representative projected cuboids across roadside sequences. The peripheral vehicle omitted near an overlap boundary falls outside the predefined three-view evaluation scope rather than being counted as a tracker false negative. Figure 6: Projected cuboids from the multi-view 3D tracker across representative roadside sequences. Conclusion We presented RISE, a framework for roadside sequence un- derstanding through two complementary tasks: metric 3D tracking and structured vision-language reasoning. By ex- ploiting complementary spatial and temporal evidence, RISE supports persistent 3D tracking and bbox-grounded predic- tion without exposing future evidence to evaluated models. Human-corrected evaluation on 20 clips yields 66.9 MOTA for the generated tracks within the defined multi-view scope. Experiments on RISE-Bench demonstrate clear benefits from domain adaptation and generally from temporal context, while revealing remaining challenges in spatial grounding, future localization, and interaction reasoning. These find- ings highlight both the value and the current limitations of sequence-level roadside understanding. Several limitations remain. Image-only 3D tracking de- pends on accurate calibration, reliable segmentation, and suf- ficient camera overlap, with weaker geometric support under heavy occlusion and near shared-view boundaries. On the vision-language side, the current 3.8-second clips primarily assess near-term reasoning from individual roadside views. Extending the observation horizon without sacrificing tem- poral resolution, together with coordinated reasoning across synchronized views, would enable the benchmark to capture longer and more complex traffic evolution. A promising di- rection is to use persistent metric 3D trajectories as explicit spatiotemporal grounding for vision-language models, en- abling more reliable reasoning about long-term motion and multi-agent interactions. References Carion, N.; Gustafson, L.; Hu, Y.-T.; Debnath, S.; Hu, R.; Suris, D.; Ryali, C.; Alwala, K. V.; Khedr, H.; Huang, A.; et al. 2025. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719. Choudhary, T.; Dewangan, V.; Chandhok, S.; Priyadarshan, S.; Jain, A.; Singh, A. K.; Srivastava, S.; Jatavallabhula, K. M.; and Krishna, K. M. 2024. Talk2bev: Language- enhanced birdâs-eye view maps for autonomous driving. In 2024 IEEE International Conference on Robotics and Au- tomation (ICRA), 16345â16352. IEEE. Gloudemans, D.; Work, D.; Wang, Y.; Gumm, G. E.; and Bar- bour, W. 2023. The Interstate-24 3D Dataset: A New Bench- mark for 3D Multi-Camera Vehicle Tracking. In British Machine Vision Conference (BMVC). Guan, R.; Hu, R.; Chen, S.; Xiao, N.; Xia, X.; Liu, J.; Chen, B.; Tang, Z.; Ouyang, N.; Liang, S.; et al. 2026. Road- scenevqa: Benchmarking visual question answering in road- side perception systems for intelligent transportation system. In Proceedings of the AAAI Conference on Artificial Intelli- gence, volume 40, 4366â4375. Hao, R.; Fan, S.; Dai, Y.; Zhang, Z.; Li, C.; Wang, Y.; Yu, H.; Yang, W.; Yuan, J.; and Nie, Z. 2024. Rcooper: A real-world large-scale dataset for roadside cooperative perception. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 22347â22357. Huang, W.; Zhang, S.; Chua, C.; Liang, Y.; Mao, Z.; Yang, H.; and Lv, C. 2026. Towards Safe Mobility: A Unified Transportation Foundation Model enabled by Open-Ended Vision-Language Dataset. arXiv preprint arXiv:2604.22260. Marcu, A.-M.; Chen, L.; HĂźnermann, J.; Karnsund, A.; Hanotte, B.; Chidananda, P.; Nair, S.; Badrinarayanan, V.; Kendall, A.; Shotton, J.; et al. 2024. Lingoqa: Visual question answering for autonomous driving. In European Conference on Computer Vision, 252â269. Springer. Qian, T.; Chen, J.; Zhuo, L.; Jiao, Y.; and Jiang, Y.-G. 2024. Nuscenes-qa: A multi-modal visual question answer- ing benchmark for autonomous driving scenario. In Pro- ceedings of the AAAI Conference on Artificial Intelligence, volume 38, 4542â4550. Sekaran, K. C.; Geisler, M.; RĂśĂle, D.; Mohan, A.; Cre- mers, D.; Utschick, W.; Botsch, M.; Huber, W.; and SchĂśn, T. 2025. UrbanIng-V2X: A Large-Scale Multi-Vehicle, Multi- Infrastructure Dataset Across Multiple Intersections for Co- operative Perception. In The Thirty-ninth Annual Confer- ence on Neural Information Processing Systems Datasets and Benchmarks Track. Sima, C.; Renz, K.; Chitta, K.; Chen, L.; Zhang, H.; Xie, C.; BeiĂwenger, J.; Luo, P.; Geiger, A.; and Li, H. 2024. Drivelm: Driving with graph visual question answering. In European conference on computer vision, 256â274. Springer. Xu, L.; Huang, H.; and Liu, J. 2021. Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events. In Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition, 9878â9888. Yang, L.; Yu, K.; Tang, T.; Li, J.; Yuan, K.; Wang, L.; Zhang, X.; and Chen, P. 2023. Bevheight: A robust framework for vision-based roadside 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 21611â21620. Yang, L.; Zhang, X.; Yu, J.; Li, J.; Zhao, T.; Wang, L.; Huang, Y.; Zhang, C.; Wang, H.; and Li, Y. 2024. MonoGAE: Road- side monocular 3D object detection with ground-aware em- beddings. IEEE Transactions on Intelligent Transportation Systems, 25(11): 17587â17601. Ye, X.; Shu, M.; Li, H.; Shi, Y.; Li, Y.; Wang, G.; Tan, X.; and Ding, E. 2022. Rope3d: The roadside perception dataset for autonomous driving and monocular 3d object detection task. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 21341â21350. You, J.; Li, P.; Jiang, Z.; Tang, W.; Huang, Z.; Gan, R.; Liu, J.; Zhao, Y.; Chen, S.; and Ran, B. 2026. V2x-qa: A comprehensive reasoning dataset and benchmark for multi- modal large language models in autonomous driving across ego, infrastructure, and cooperative views. arXiv preprint arXiv:2604.02710. Yu, H.; Luo, Y.; Shu, M.; Huo, Y.; Yang, Z.; Shi, Y.; Guo, Z.; Li, H.; Hu, X.; Yuan, J.; et al. 2022. Dair-v2x: A large- scale dataset for vehicle-infrastructure cooperative 3d object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 21361â21370. Yu, H.; Yang, W.; Ruan, H.; Yang, Z.; Tang, Y.; Gao, X.; Hao, X.; Shi, Y.; Pan, Y.; Sun, N.; et al. 2023. V2x-seq: A large- scale sequential dataset for vehicle-infrastructure cooperative perception and forecasting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5486â5495. Zhang, Y.; Zheng, Z.; Liu, J.; Huang, Z.; Zhou, Z.; Meng, Z.; Cai, T.; and Ma, J. 2025. MIC-BEV: Multi- Infrastructure Camera Birdâs-Eye-View Transformer with Relation-Aware Fusion for 3D Object Detection. arXiv preprint arXiv:2510.24688. Zheng, Y.; Zhang, R.; Zhang, J.; Ye, Y.; and Luo, Z. 2024. Llamafactory: Unified efficient fine-tuning of 100+ language models. In Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 3: system demonstrations), 400â410. Zhou, X.; Larintzakis, K.; Guo, H.; Zimmer, W.; Liu, M.; Cao, H.; Zhang, J.; Lakshminarasimhan, V.; Strand, L.; and Knoll, A. 2025. TUMTraf VideoQA: Dataset and Bench- mark for Unified Spatio-Temporal Video Understanding in Traffic Scenes. In Forty-second International Conference on Machine Learning.