Paper deep dive
The 10th AI City Challenge
Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang, Anuj Sharma, Quan Kong, Munkhjargal Gochoo, Jun-Wei Hsieh, Tomasz Kornuta, Zhedong Zheng, Renran Tian, Judah Goldfeder, Fulgencio Navarro, Yuxing Wang, Yizhou Wang, Sameer Satish Pusegaonkar, Anqi Li, Nalin Dadhich, Ridham Kachhadiya, Dhanishtha Patil, Haoquan Liang, Jiajun Li, Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat, Shuyu Yang, Ashutosh Kumar, Rong Wang, Rafael Martin Nieto, Peter Christiansen, Ahmed Abduljawad, Mohanrasu Shanmugam, Nadeem Shaik, Sujit Biswas, Xunlei Wu, Vidya Murali, Rama Chellappa
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/19/2026, 4:09:09 AM
Summary
The 10th AI City Challenge, held at ECCV 2026, marks a decade of benchmarking for intelligent transportation and physical AI. The 2026 edition saw significant growth with 325 registered teams from 26 countries. The challenge features six primary tracks covering multi-camera 3D perception, transportation safety captioning/VQA, traffic anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city object detection. It also includes two out-of-domain leaderboards for fisheye traffic violations and pedestrian intent. Successful systems combine foundation models with geometric grounding, synthetic data design, and domain adaptation.
Entities (20)
Relation Signals (18)
AI City Challenge โ heldat โ ECCV 2026
confidence 98% ยท The 10th AI City Challenge, held with ECCV 2026
Zheng Tang โ affiliatedwith โ Santa Clara University
confidence 95% ยท Zheng Tang Affiliation: Shuo Wang Affiliation: David C. Anastasiu Affiliation: Santa Clara University
AI City Challenge โ hastrack โ Track 2
confidence 95% ยท Track 2: Transportation Safety Understanding and Captioning
AI City Challenge โ hastrack โ Track 3
confidence 95% ยท Track 3: Anomalous Events in Transportation
AI City Challenge โ hastrack โ Track 4
confidence 95% ยท Track 4: Text-Based Person Anomaly Search
AI City Challenge โ hastrack โ Track 5
confidence 95% ยท Track 5: Generative Traffic Video Forecasting
AI City Challenge โ hastrack โ Track 6
confidence 95% ยท Track 6: Cross-City Object Detection
AI City Challenge โ hastrack โ Track 1
confidence 95% ยท Its six primary tracks cover multi-camera 3D perception... Track 1: Multi-Camera 3D Perception
Stellarview AI โ โ
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with vehicle detection, classification, and tracking, the challenge has grown into a broad benchmark suite for multi-camera perception, multimodal reasoning, synthetic-to-real learning, generative forecasting, and privacy-preserving evaluation. The 2026 edition continued this growth with 325 registered teams, up from 245 in 2025, and participation from 26 countries and regions, up from 15. Its six primary tracks cover multi-camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city object detection. Track 3 further includes two out-of-domain leaderboards, submitted as Tracks 7 and 8, for fisheye traffic-violation understanding and pedestrian situated-intent VQA. This paper summarizes the challenge setup, datasets, evaluation protocols, leaderboard results, and workshop papers. Across tracks, successful systems combine foundation models with geometric grounding, retrieval or reranking, synthetic-data design, domain adaptation, and controlled inference.
Tags
Links
- Source: https://arxiv.org/abs/2608.17044v1
- Canonical: https://arxiv.org/abs/2608.17044v1
Trouble viewing inline? Open PDF directly โ
Full Text
48,901 characters extracted from source content.
Expand or collapse full text
The 10th AI City Challenge Zheng Tang Affiliation: Shuo Wang Affiliation: David C. Anastasiu Affiliation: Santa Clara University Ming-Ching Chang Affiliation: University at Albany, SUNY Anuj Sharma Affiliation: Iowa State University Quan Kong Affiliation: Woven by Toyota Munkhjargal Gochoo Affiliation: United Arab Emirates University Jun-Wei Hsieh Affiliation: National Yang Ming Chiao Tung University Tomasz Kornuta Affiliation: Zhedong Zheng Affiliation: University of Macau Renran Tian Affiliation: North Carolina State University Judah Goldfeder Affiliation: Columbia University Fulgencio Navarro Affiliation: Milestone Systems Yuxing Wang Affiliation: Yizhou Wang Affiliation: Sameer Satish Pusegaonkar Affiliation: Anqi Li Affiliation: Nalin Dadhich Affiliation: Ridham Kachhadiya Affiliation: Santa Clara University Dhanishtha Patil Affiliation: Santa Clara University Haoquan Liang Affiliation: Jiajun Li Affiliation: Han Zhang Affiliation: Yilin Zhao Affiliation: Zaid Pervaiz Bhat Affiliation: Shuyu Yang Affiliation: Xiโan Jiaotong University Ashutosh Kumar Affiliation: Woven by Toyota Rong Wang Affiliation: Woven by Toyota Rafael Martin Nieto Affiliation: Milestone Systems Peter Christiansen Affiliation: Milestone Systems Ahmed Abduljawad Affiliation: United Arab Emirates University Mohanrasu Shanmugam Affiliation: United Arab Emirates University Nadeem Shaik Affiliation: United Arab Emirates University Sujit Biswas Affiliation: Xunlei Wu Affiliation: Vidya Murali Affiliation: Rama Chellappa Affiliation: Johns Hopkins University Abstract The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with vehicle detection, classification, and tracking, the challenge has grown into a broad benchmark suite for multi-camera perception, multimodal reasoning, synthetic-to-real learning, generative forecasting, and privacy-preserving evaluation. The 2026 edition continued this growth with 325 registered teams, up from 245 in 2025, and participation from 26 countries and regions, up from 15. Its six primary tracks cover multi-camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city object detection. Track 3 further includes two out-of-domain leaderboards, submitted as Tracks 7 and 8, for fisheye traffic-violation understanding and pedestrian situated-intent VQA. This paper summarizes the challenge setup, datasets, evaluation protocols, leaderboard results, and workshop papers. Across tracks, successful systems combine foundation models with geometric grounding, retrieval or reranking, synthetic-data design, domain adaptation, and controlled inference. Keywords: AI City Challenge Synthetic-to-real transfer Video understanding Intelligent transportation Physical AI 1 Introduction The AI City Challenge is a recurring benchmark venue for computer vision systems deployed in urban, transportation, and physical AI settings. Its 10th edition, hosted as an ECCV 2026 workshop, marks a decade of growth from the 2017 vehicle detection, classification, and tracking tasks to a wider testbed for multi-camera 3D perception, multimodal reasoning, synthetic-to-real transfer, generative prediction, and privacy-preserving evaluation. This evolution parallels broader progress in AI. Early tasks reflected deep learning for detection, tracking, and re-identification; recent tasks reflect VQA, foundation models, generative video models, and reasoning systems that must explain events rather than only detect them. The 2026 cross-city object detection track also highlights a persistent limitation of modern AI systems: strong in-domain performance does not guarantee robustness under geographic, camera, and scene-domain shift. Participation also reached a new high. The 2026 challenge registered 325 teams, compared with 245 in 2025, a growth of about 33%, and the participant pool represented 26 countries and regions, compared with 15 in the previous year. The breadth of participation is visible across the leaderboards. On the general and public leaderboards, respectively, Track 1 had 29 and 15 teams, Track 2 had 44 and 22, Track 3 had 76 and 27, Track 4 had 47 and 28, Track 5 had 18 and 12, and Track 6 had 39 and 29. The Track 3 out-of-domain leaderboards also drew participation, with 15 general and 8 public teams for FETV fisheye traffic-violation understanding and 15 general and 7 public teams for PSI-VQA pedestrian situated-intent VQA. This edition therefore introduces a broader set of tasks than a single leaderboard can capture: multi-camera 3D perception, safety-oriented captioning and VQA, anomalous-event reasoning with two OOD leaderboards, text-based person anomaly search, text-conditioned future-frame generation, and cross-city object detection through Milestone Project Hafnia. Together, these tracks form an urban intelligence benchmark suite rather than isolated recognition tasks. This summary paper follows the structure of recent AI City Challenge overview papers [40, 64, 54]: it first presents the challenge setup, then describes the datasets, evaluation protocols, and leaderboard system, and finally summarizes the track results and the main technical trends in workshop papers. Dataset and benchmark papers associated with the 2026 challenge are discussed in the dataset section, while the results section focuses on teams with identifiable evaluation-system entries. 2 Challenge Setup The 2026 AI City Challenge was organized around six primary challenge tracks and eight evaluation-server leaderboards, because the main Track 3 task includes two optional out-of-domain evaluations submitted as Tracks 7 and 8 on the evaluation system. Teams registered, requested submission privileges, and submitted predictions to public or general leaderboards [2]. Public leaderboard entries were intended for award-eligible submissions, while the general leaderboards recorded broader participation. Track 1: Multi-Camera 3D Perception. Participants tracked people, robots, forklifts, pallet trucks, and humanoids across synchronized warehouse cameras. The task required per-frame 3D boxes, class labels, scene-level identities, and world-coordinate localization. Depth was available for training and validation only; hidden real-world testing required RGB-only inference. Track 2: Transportation Safety Understanding and Captioning. Participants used synthetic Digital Twin WTS data to caption and answer questions about real WTS videos. Each event was divided into behaviorally meaningful phases, and systems generated pedestrian and vehicle captions while also answering multiple-choice questions about position, direction, attention, attributes, and context. Track 3: Anomalous Events in Transportation. Participants built one unified system that detects, reasons about, and explains anomalous transportation events through binary and multiple-choice questions, open-ended explanation, causal linkage, scene description, temporal description, and summarization. Tracks 7 and 8 extended the same reasoning theme to fisheye traffic violations and pedestrian intent. Track 4: Text-Based Person Anomaly Search. Participants retrieved real pedestrian images from natural-language descriptions that mention both appearance and behavior. The task is a synthetic-to-real retrieval problem: synthetic image-text pairs provide training signal, while the hidden real-world query-gallery test set measures whether models generalize to realistic abnormal and routine behaviors. Track 5: Generative Traffic Video Forecasting. Participants generated future frames conditioned on recent history frames and textual descriptions of target future behavior. The task stresses temporal consistency, visual fidelity, and whether a generated traffic sequence remains semantically aligned with safety-critical pedestrian and vehicle behavior. Track 6: Cross-City Object Detection. Participants trained object detectors through the Hafnia Training-as-a-Service platform, which exposes managed training and benchmarking workflows without allowing direct extraction of the full real-world traffic corpus. The hidden benchmark mixes source- and target-city samples to measure geographic and visual domain shift. Across the tracks, the challenge combined classical perception metrics such as HOTA and mAP with multimodal language, retrieval, generation, and reasoning metrics. This design encouraged systems that combine foundation models with domain-specific constraints, rather than relying on a single model family across all tasks. 3 Datasets and Tracks The 2026 challenge uses six primary tracks and two out-of-domain leaderboards on the shared evaluation system. The primary tracks are described on the challenge web site [3, 4, 5, 6, 7, 8], while Tracks 7 and 8 extend Track 3 to fisheye traffic-violation understanding and pedestrian situated-intent VQA. The following subsections summarize the data settings, hidden-test assumptions, and evaluation focus for each benchmark. 3.1 Track 1: Physical AI Smart Spaces Track 1 extends the PhysicalAI-SmartSpaces benchmark to multi-camera 3D perception in warehouse environments [3, 43]. The training and validation data were generated from a large set of simulated warehouse scenes with synchronized RGB and depth videos, calibrated camera poses, 2D and 3D annotations, top-down maps, and multi-object identities. The hidden test set used real-world video in layouts that differ from the synthetic training scenes, making the task a strict Sim2Real benchmark rather than a closed-world tracking exercise. The central difficulty is the coupling between geometry and association. A system must detect objects in each camera, lift observations into a shared 3D coordinate frame, and maintain identities despite occlusion, similar instances, changing viewpoints, and cluttered warehouse geometry; Fig. shows the synchronized views and top-down layout. The RGB-only test restriction prevents direct test-time depth use, so successful approaches rely on learned monocular cues, calibration-aware projection, scene priors, and global association. Figure 1: Track 1 sample: synchronized multi-camera warehouse views and a top-down layout for RGB-only multi-camera 3D perception. 3.2 Track 2: Digital Twin WTS Track 2 uses the Digital Twin WTS setting for synthetic-to-real traffic safety understanding [4, 39, 48]. The training and validation data provide synthetic multi-view traffic-safety scenes, annotations, segment-level descriptions, and VQA labels, with representative synthetic and real views shown in Fig. . The hidden test set uses real WTS videos, so the benchmark measures whether methods can transfer from controlled synthetic scenes to real camera footage while preserving fine-grained safety semantics. The task combines two complementary outputs. For captioning, systems describe pedestrian and vehicle behavior across event phases, including movement direction, relative position, gaze, visibility, attributes, weather, road geometry, and traffic context. For VQA, systems answer structured questions that test whether the model has grounded those descriptions in the visual evidence. The track therefore rewards representations that separate stable scene layout from variable appearance and behavior, rather than merely matching surface text patterns. Figure 2: Track 2 sample: synthetic Digital Twin WTS scenes paired with real WTS-style views and structured pedestrian/vehicle descriptions. 3.3 Track 3: Anomalous Events in Transportation Track 3 introduces TAR (Traffic Anomaly Reasoning) and TAR-Bench for traffic anomaly reasoning [5, 67, 44]. The TAR training dataset contains 44,040 annotations over 3,670 transportation videos and covers 10 task types, including binary and multiple-choice reasoning, open-ended QA, captioning, causal explanation, scene description, temporal ordering, and summarization. TAR-Bench provides 960 human-curated annotations for 80 held-out clips. Unlike earlier anomaly-detection tasks, this track emphasizes explanation and evidence grounding: a strong model must identify what happened, infer why it happened, and describe the evidence coherently. Two out-of-domain datasets extend Track 3. Track 7 evaluates FETV traffic-violation understanding from fisheye cameras [1], where geometry changes object scale, direction cues, and perspective. Track 8 evaluates PSI-VQA, a pedestrian situated-intent VQA benchmark for ambiguous pedestrian behavior around automated driving [63]. Fig. illustrates the evidence, question, and explanation structure of the main TAR setting, and the evaluation system reports all three reasoning settings separately as Tracks 3, 7, and 8. Figure 3: Track 3 sample: traffic anomaly reasoning combines video evidence, task-specific questions, event descriptions, and causal explanations. 3.4 Track 4: Pedestrian Anomaly Behavior Track 4 uses the Pedestrian Anomaly Behavior (PAB) benchmark for text-based person anomaly search [6]. The training set is synthetic and provides image-text pairs that describe person appearance, scene context, and action. The real test set contains query descriptions and gallery images. Participants rank gallery images for each query, and the final leaderboard uses retrieval quality on hidden real-world data. The benchmark is difficult because abnormal behavior is often defined by relationships between action, body pose, scene context, and a natural-language query. Appearance-only retrieval is insufficient: examples such as falling, lying, being hit, or unsafe crossing require action grounding, as the synthetic/real and hard-negative examples in Fig. illustrate. Successful methods therefore use text-image contrastive training, action-aware alignment, hard-negative mining, reranking, and query decomposition. Figure 4: Track 4 sample: synthetic PAB training examples and real-gallery retrieval, including hard negatives that share appearance or scene context with the query. 3.5 Track 5: Traffic Video Forecasting Track 5 asks teams to generate future traffic video frames from a short history window and a textual description of expected future behavior [7]. The task builds on WTS-style traffic scenes but changes the output from analysis to generation, as shown by the history and target-frame example in Fig. . Models must synthesize plausible motion, preserve identity and background consistency, and reflect the target pedestrian/vehicle behavior. This task exposes a different failure mode from captioning and retrieval: a generated video may be sharp but semantically wrong, or text-aligned while losing scene consistency and temporal smoothness. The leaderboard therefore combines low-level fidelity, perceptual, semantic, and video-distribution metrics. Accepted methods typically use diffusion or world-model priors, frame-history conditioning, text-guided planning, and post-generation selection. Figure 5: Track 5 sample: future-frame generation from history frames and target behavior descriptions in traffic scenes. 3.6 Track 6: Hafnia Cross-City Detection Track 6 uses Hafnia Training-as-a-Service platform [37] for cross-city object detection [8, 36]. The data consist of real traffic-camera images with object annotations across multiple categories. Unlike tracks where participants directly download the full corpus, this track uses a managed platform for privacy-preserving training, validation, and hidden benchmarking on controlled splits. The track targets a common deployment problem: detectors trained in one city or camera network often lose accuracy when moved to another geography, camera height, lens, weather condition, or traffic pattern, as illustrated in Fig. . The hidden benchmark includes source- and target-city samples, rewarding detector design together with class-aware augmentation, resolution management, domain-shift validation, and confidence calibration. Figure 6: Track 6 sample: privacy-preserved traffic-camera images from Milestone Systems (Project Hafnia), with annotated source-city frames on top and hidden target-city frames below. 4 Evaluation Protocols The evaluation system exposed public and general leaderboards for each track [2]. During the challenge, scores were computed on a subset of hidden test data and only limited ranking information was shown. After the deadline, final scores were recomputed on full hidden test sets and team names were made public for the workshop summary. Public/general ranks in the results tables use the form โ1/21 (1/58)โ, meaning rank 1 among 21 public teams and rank 1 among 58 general teams. 4.1 Track 1 Evaluation Track 1 used 3D HOTA [35], which jointly evaluates detection accuracy, association accuracy, and localization quality over 3D tracks. Submissions provided class labels, frame indices, scene-level identities, and 3D boxes or equivalent localization fields in the warehouse coordinate frame. Scores were averaged across classes and scenes. Because the hidden test set was RGB-only, participants could not use test-time depth; online submissions received an additional ranking bonus when they used only current and past frames. 4.2 Track 2 Evaluation Track 2 combined captioning and VQA. Caption quality used BLEU-4 [46], METEOR [10], ROUGE-L [32], and CIDEr [60]; VQA used answer accuracy over structured questions [9]. The final score averaged caption and VQA components, making the hidden real WTS test set a synthetic-to-real transfer benchmark for semantic grounding. 4.3 Track 3 Evaluation Track 3 evaluated multi-task traffic anomaly reasoning. The official in-domain TAR-Bench mean was computed from nine scored task types, excluding temporal localization. Closed-form tasks used accuracy-style scoring, while text-generation tasks used BLEU, METEOR, ROUGE, and CIDEr [46, 10, 32, 60]. The design rewards models that organize evidence over time and generate semantically correct, visually grounded answers. 4.4 Track 4 Evaluation Track 4 used mean Average Precision (mAP), widely used in detection and retrieval benchmarks [16, 33], for text-to-image retrieval. For each text query, participants returned a ranked gallery list; the score rewards placing correct matches early while suppressing hard negatives. Because training images are synthetic and the test gallery is real, the metric measures both retrieval quality and Sim2Real generalization. 4.5 Track 5 Evaluation Track 5 combined visual, perceptual, and semantic video metrics: PSNR and SSIM [65] for pixel fidelity, LPIPS [68] and CLIP-S [18] for perceptual and text-image alignment, and FID [19] and FVD [59] for distributional realism. The final score rewards realistic future frames that also follow the target behavior described in text. 4.6 Track 6 Evaluation Track 6 used object detection mAP [16, 33] on hidden Hafnia benchmark images. The platform evaluated models on source-city and target-city samples, so the score reflected both in-domain detector quality and robustness to geographic shift. The managed platform also reduced the risk of private-data leakage from the real-world corpus, because participants trained and benchmarked through controlled workflows rather than downloading the full hidden dataset. 4.7 Tracks 7 and 8 Evaluation Tracks 7 and 8 were optional out-of-domain Track 3 leaderboards. Track 7 evaluated FETV traffic-violation understanding from fisheye cameras, while Track 8 evaluated PSI-VQA pedestrian situated-intent questions. Both used VQA- and reasoning-oriented scoring related to the Track 3 protocol [9, 67] and reported final scores separately from the main TAR mean, making transfer robustness visible under fisheye geometry and socially ambiguous pedestrian-intent questions. Award-candidate teams were required to provide reproducible code and models, and the paper-review process encouraged accepted workshop papers to report the final team names and leaderboard values exactly as shown on the evaluation system. This requirement was especially important for tracks with hidden tests or managed data access, where reproducibility depends on both the submitted model and a clearly documented training and inference pipeline. 5 Challenge Results Tables โ summarize accepted-paper entries with identifiable evaluation IDs on the 2026 leaderboards. The tables are not intended to replace the full public evaluation system; instead, they connect leaderboard outcomes to accepted workshop papers that describe the corresponding methods. Public ranks and general ranks are both reported because the public leaderboard captures award-eligible entries, while the general leaderboard captures broader participation. 5.1 Summary for the Track 1 Challenge Track 1 attracted methods that combined object detection, camera calibration, 2D-to-3D lifting, and global multi-camera association. The leading accepted systems relied on strong RGB detectors and then used warehouse geometry to reduce cross-view ambiguity. Because the hidden test set did not expose depth, the best submissions treated depth and 3D layout as learned or calibrated priors rather than as direct test-time inputs. Table 1: Track 1 leaderboard entries connected to accepted papers. Rank Team ID Team 3D HOTA Paper 1/15 (2/30) 289 EVA 56.5447 [66]; online=true 2/15 (3/30) 34 SKKU-AL-T1 52.0118 [55]; online=true 3/15 (4/30) 130 Playbox 38.0105 [53]; online=true 4/15 (5/30) 4 QDTers 34.1845 [29]; online=true 5/15 (6/30) 133 TU-YMLab 25.9712 [21]; online=true The accepted Track 1 papers in Table show several common design choices. EVA [66] led the public leaderboard among accepted papers with geometry-aware tracking. SKKU-AL-T1 [55] and Playbox [53] also performed strongly with online RGB-only pipelines, showing that streaming constraints remain compatible with competitive detection and association. Other papers explored collaborative 2D-3D tracking [29] and uncertainty-aware observation construction [21]. 5.2 Summary for the Track 2 Challenge Track 2 measured transfer from synthetic Digital Twin WTS data to real WTS videos. The strongest systems used VLM backbones, but the accepted papers also show that off-the-shelf prompting was rarely enough. Teams improved transfer by decomposing the task into scene parsing, phase recognition, caption generation, VQA answering, and answer calibration. Table 2: Track 2 leaderboard entries connected to accepted papers. Rank Team ID Team S2 Paper 1/22 (1/44) 47 Latent Painter - UTE 60.0853 [11]; VL-JEPA 2/22 (2/44) 24 UIT - Kitchen 57.3307 [38]; VQA 3/22 (4/44) 266 KZ6 56.7949 [23]; Qwen-3-VL-8B 8/22 (14/44) 127 Team KODE 55.4679 [27]; Qwen The Track 2 results in Table suggest that separating scene structure from behavior helps narrow the synthetic-to-real gap. The leading accepted entries used decoupled semantics, V-JEPA style visual features, state-bridging strategies, and modular spatial grounding [11, 38, 23, 27]. Methods that represented relative position, pedestrian visibility, vehicle motion, and environmental context were better aligned with the task than generic captioning pipelines. 5.3 Summary for the Track 3 Challenge Track 3 was the largest reasoning leaderboard in the 2026 challenge. It required systems to answer heterogeneous questions about anomalous traffic events and to produce explanations that match visual evidence. The ranking emphasized not only timestamp prediction, binary or multiple choice correctness, but also semantic and causal reasoning soundness. Table 3: Track 3 leaderboard entries connected to accepted papers. Rank Team ID Team Mean Paper 1/27 (1/76) 25 Stellarview AI 0.6788 [26]; Qwen 3.5 - (2/76) 45 UOB&UW Team 0.6779 [34]; Qwen3VL-8B 2/27 (5/76) 60 FPT AI Vision 0.6703 [20]; Qwen3-VL-8B 3/27 (7/76) 12 Smart Vision 0.6669 [57]; Qwen 10/27 (26/76) 30 UWIPL_ETRI 0.6185 [52]; Qwen 15/27 (36/76) 122 OptimAI 0.5880 [17]; Qwen 16/27 (41/76) 139 MR-CAS 0.5780 [31]; GPT 24/27 (55/76) 277 Korea Drive 0.4256 [25]; Qwen3 As shown in Table , Stellarview AI [26] led both the public and general TAR ranking among accepted papers. Other strong submissions used evidence-driven chained reasoning, unified multi-leaderboard agents, metric matching, shared event memory, and structured question routing [57, 34, 31, 20, 45]. The accepted papers indicate a shift from simple VLM prompting toward agentic pipelines that first extract visual evidence, then match it to a task-specific answer format. 5.4 Summary for the Track 4 Challenge Track 4 produced high retrieval scores, indicating rapid progress on the PAB synthetic-to-real setting. At the same time, the concentration of strong mAP values made the track sensitive to careful data-use policy and reproducibility checks. The top methods combined strong text-image embeddings with action-specific reranking and hard-negative handling. Table 4: Track 4 leaderboard entries connected to accepted papers. Rank Team ID Team mAP Paper 1/28 (3/47) 59 Xiilab.AIpex 99.3020 [47] 3/28 (6/47) 9 hiensumi 98.3535 [15] 7/28 (9/47) 27 VGU AI LAB 95.4078 [49] 8/28 (11/47) 97 SMART Lab 94.7815 [13] 11/28 (14/47) 76 HCMUS_4CentralVN 93.6715 [61] 13/28 (15/47) 93 SelabHuman 93.6577 [41] 15/28 (22/47) 64 GenAI4E 90.9236 [62] 18/28 (31/47) 29 EMBIA 84.2509 [56] Table lists Xiilab.AIpex [47] as the top accepted public/general entry in the final decision sheet. Other accepted papers explored late consensus, heterogeneous VLM ensembling, embedding prediction, cross-encoder reranking, global assignment, and action-aligned retrieval [49, 56, 41, 13, 15, 62, 61]. The track highlights a broader lesson for text-person search: robust action semantics must be learned together with identity, clothing, and scene context. 5.5 Summary for the Track 5 Challenge Track 5 moved the challenge from recognizing or explaining traffic scenes to generating plausible futures. The leaderboard rewarded systems that maintained scene continuity while following textual descriptions of future behavior. This made the track a natural test of video diffusion models, world models, and text-conditioned forecasting pipelines. Table 5: Track 5 leaderboard entries connected to accepted papers. Rank Team ID Team Final Paper 1/12 (2/18) 209 Qyn 76.4866 [14] 2/12 (3/18) 78 SSUPER 76.0385 [30] 3/12 (4/18) 47 Latent Painter - UTE 75.4302 [42] 4/12 (6/18) 83 CHTTL_A30 74.0544 [12] 5/12 (7/18) 39 VGU_ai_lab 73.3037 [28] The accepted systems summarized in Table , including Qyn [14], SSUPER [30], Latent Painter - UTE [42], CHTTL_A30 [12], and VGU_ai_lab [28], combined history-frame conditioning with language-guided future descriptions, diffusion or world-model priors, and metric-aware selection. The close scores suggest rapid progress, while the metric suite shows that safety-relevant semantic consistency remains hard to capture with one number. 5.6 Summary for the Track 6 Challenge Track 6 emphasized a deployment problem that is often hidden by conventional object-detection benchmarks: a detector tuned for one camera network can lose accuracy when moved to another city with different viewpoints, object scales, weather, compression artifacts, and traffic composition. The Hafnia platform made this setting more realistic by using managed training and evaluation workflows, so teams had to improve domain robustness without directly extracting the full real-world corpus. This makes the track closer to practical cross-site deployment, where data governance and distribution shift are handled together. The accepted papers in Table show several complementary approaches to cross-city detection. SKKU-AL-T1 [50] led the accepted-paper entries with a strategy focused on pre-training and augmentation for zero-shot transfer. BIT-ODL [24] used evidence-conditioned multi-source pretraining, while BK2TheFuture [51] emphasized targeted augmentation and class-aware inference. The lyx submission [58] explored two-stage fusion and scale-aware geometric refinement. Although the final mAP values are lower than the retrieval scores in Track 4, this should be interpreted in light of the task: Track 6 evaluates object localization, class recognition, and confidence ranking under cross-city shift on real imagery. Across the accepted methods, the recurring theme is that detector architecture alone is not sufficient. Successful systems also manage resolution, category imbalance, source-domain bias, validation under distribution shift, and confidence calibration for target-city scenes. Table 6: Track 6 leaderboard entries connected to accepted papers. Rank Team ID Team mAP Paper 1/25 (1/35) 34 SKKU-AL-T1 0.4753 [50] 2/25 (2/35) 265 BIT-ODL 0.4281 [24] 4/25 (5/35) 261 BK2TheFuture 0.4169 [51] - (7/35) 315 lyx 0.4060 [58] 5.7 Summary for the Track 7 and Track 8 OOD Challenges Tracks 7 and 8 extended the Track 3 reasoning task into two out-of-domain settings. Track 7 used FETV traffic-violation understanding from fisheye cameras, where wide-angle projection changes object shape, direction cues, lane geometry, and motion interpretation. Track 8 used PSI-VQA for pedestrian situated-intent reasoning, where answers depend on subtle social signals, occlusion, ambiguity, and the viewpoint of an automated-driving agent. These tracks are therefore not merely auxiliary leaderboards; they test whether traffic reasoning systems transfer from in-domain CCTV anomaly clips to different sensing geometries and interaction questions. Table summarizes accepted-paper entries on both OOD leaderboards. UniTraffic [52] ranked first among accepted public entries on both settings, suggesting that evidence-centric agentic reasoning can transfer across anomaly, violation, and intent tasks. UniTraffic-Agent [31] and Korea Drive [25] also performed strongly with unified or task-routed video-language reasoning. TAU-Agent [34] appears in the general rankings and reflects another design pattern: retrieval-augmented reasoning can help reuse event evidence across related but shifted traffic-understanding tasks. The Track 7 scores are close among the top accepted entries, indicating that fisheye violation understanding remains sensitive to spatial evidence extraction. Track 8 has a wider score range, consistent with the added difficulty of socially ambiguous pedestrian-intent questions. Together, these OOD leaderboards complement TAR by measuring robustness, not only in-domain reasoning accuracy. Table 7: Track 7 and Track 8 OOD leaderboard entries connected to accepted papers. Track Rank Team ID Team Final Paper Track 7: FETV traffic-violation understanding 7 1/8 (1/15) 30 UWIPL_ETRI 0.4891 [52] 7 2/8 (3/15) 139 MR-CAS 0.4884 [31] 7 3/8 (5/15) 277 Korea Drive 0.4634 [25] 7 - (12/15) 45 UOB&UW Team 0.3998 [34] Track 8: PSI-VQA pedestrian situated-intent reasoning 8 1/7 (2/15) 30 UWIPL_ETRI 70.6397 [52] 8 - (5/15) 73 University of Washington 67.9275 [34] 8 4/7 (8/15) 139 MR-CAS 64.4161 [31] 8 5/7 (9/15) 277 Korea Drive 57.0400 [25] 6 Discussion and Conclusion Across a decade, the AI City Challenge has grown from vehicle detection, classification, tracking, and counting to broader urban intelligence. In 2026, 325 teams from 26 countries and regions evaluated systems across six primary tracks and two OOD leaderboards: multi-camera 3D perception, synthetic-to-real safety understanding, anomaly reasoning, text-based search, generative forecasting, and cross-city detection. The track portfolio reflects how the field has moved from isolated CV predictions toward systems that combine geometry, language, temporal evidence, generation, and domain adaptation. The 2026 results also show that progress is uneven in productive ways. Some teams achieved strong retrieval and reasoning scores with foundation-model pipelines, while the cross-city detection, RGB-only 3D perception, video forecasting, and OOD reasoning settings exposed remaining gaps in transfer, calibration, and reproducibility. This is a useful outcome for a benchmark: the best submissions identify practical solution patterns, while the remaining failures define research directions for future editions. We expect subsequent challenges to emphasize reproducible end-to-end systems, more explicit OOD testing, privacy-preserving evaluation platforms, and interaction with related programs such as IARPA Video LINCS [22]. Across the accepted papers, three patterns are especially visible: many high-ranking entries used modular pipelines, the strongest Sim2Real and cross-city methods treated domain shift as a first-class constraint, and the multimodal tracks showed that language is useful only when grounded in visual evidence. Together, these patterns connect the 2026 tracks into a shared research agenda for deployable urban AI. Acknowledgments Rama Chellappa was supported by the VideoLINCS program. This research is based upon work supported in part by the Office of the Director of National Intelligence (ODNI), Intelligence Advanced Research Projects Activity (IARPA), via 56000026C0026. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of ODNI, IARPA, or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright annotation therein. References [1] A. Abduljawad, M. S. S, N. S. Shaik, M. Chang, J. Hsieh, and M. Gochoo (2026) FETV: traffic violation report generation from fisheye camera systems. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง3.3. [2] AI City Challenge (2026) 2026 evaluation system. Note: https://w.aicitychallenge.org/2026-evaluation-system/Accessed 2026-08-05 Cited by: ยง2, ยง4. [3] AI City Challenge (2026) 2026 track 1: multi-camera 3d perception (sim2real). Note: https://w.aicitychallenge.org/2026-track1/Accessed 2026-08-05 Cited by: ยง3.1, ยง3. [4] AI City Challenge (2026) 2026 track 2: transportation safety understanding and captioning (sim2real). Note: https://w.aicitychallenge.org/2026-track2/Accessed 2026-08-05 Cited by: ยง3.2, ยง3. [5] AI City Challenge (2026) 2026 track 3: anomalous events in transportation. Note: https://w.aicitychallenge.org/2026-track3/Accessed 2026-08-05 Cited by: ยง3.3, ยง3. [6] AI City Challenge (2026) 2026 track 4: text-based person anomaly search (sim2real). Note: https://w.aicitychallenge.org/2026-track4/Accessed 2026-08-05 Cited by: ยง3.4, ยง3. [7] AI City Challenge (2026) 2026 track 5: generative traffic video forecasting. Note: https://w.aicitychallenge.org/2026-track5/Accessed 2026-08-05 Cited by: ยง3.5, ยง3. [8] AI City Challenge (2026) 2026 track 6: cross-city object detection (milestone project hafnia). Note: https://w.aicitychallenge.org/2026-track6/Accessed 2026-08-05 Cited by: ยง3.6, ยง3. [9] S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh (2015) VQA: visual question answering. In Proceedings of the IEEE International Conference on Computer Vision, p. 2425โ2433. Cited by: ยง4.2, ยง4.7. [10] S. Banerjee and A. Lavie (2005) METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, p. 65โ72. Cited by: ยง4.2, ยง4.3. [11] N. H. T. Bui, T. N. Vo, T. T. G. Nguyen, and H. D. Bui (2026) Sim-to-real traffic scene understanding by decoupling semantics from caption generation with V-JEPA. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.2, Table 2. [12] Y. Chu, B. Lu, S. Du, and W. Chen (2026) LLM-guided kinematic caption fusion for traffic video forecasting with frozen world foundation models. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.5, Table 5. [13] E. K. O. Denteh, B. A. Kyem, A. Danyo, J. K. Asamoah, R. Dzinyela, S. Owusu-Ansah, D. A. GYIMAH), and A. Aboah (2026) Synthetic-to-real text-based person anomaly search via multi-model retrieval fusion and cross-encoder reranking. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.4, Table 4. [14] Q. M. Dinh and T. Doan (2026) CosmosAlign: adapting a world foundation model for generative traffic video forecasting. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.5, Table 5. [15] H. P. Duy, H. N. P. Gia, B. Tran, T. Nguyen, T. Do, T. D. Ngo, V. Duy-Dinh Le, and S. Satoh (2026) Protocol-aware global assignment for text-based person anomaly search. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.4, Table 4. [16] M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman (2010) The PASCAL visual object classes (VOC) challenge. International Journal of Computer Vision 88, p. 303โ338. Cited by: ยง4.4, ยง4.6. [17] G. Gankhuyag, J. Yoo, J. Park, S. Lee, H. Son, and K. Min (2026) Motion trails as visual prompts: a vision-language pipeline for traffic anomaly reasoning. In ECCV Workshops, Malmรถ, Sweden. Cited by: Table 3. [18] J. Hessel, A. Holtzman, M. Forbes, R. Le Bras, and Y. Choi (2021) CLIPScore: a reference-free evaluation metric for image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 7514โ7528. Cited by: ยง4.5. [19] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter (2017) GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, Cited by: ยง4.5. [20] V. Hoang (2026) Metric-matched test-time selection for traffic anomaly video question answering. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.3, Table 3. [21] Y. Hsu, M. Zielewski, and C. Huang (2026) Uncertainty-guided multi-camera observation construction for online tracking. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.1, Table 1. [22] IARPA (2024) Video lincs: video linking and intelligence from non-collaborative sensors. Note: https://w.iarpa.gov/research-programs/video-lincsAccessed 2026-08-07 Cited by: ยง6. [23] J. Jeong, S. Kim, and G. Gankhuyag (2026) Modular spatial grounding for sim2real traffic-safety VQA. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.2, Table 2. [24] Z. Jia, J. Wang, Y. Zou, and lili (2026) EC-DEIM: evidence-conditioned multi-source pretraining for cross-city fine-grained object detection. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.6, Table 6. [25] H. Kim (2026) Task-routed video-language reasoning across traffic domains: korea drive at ai city challenge 2026. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.7, Table 3, Table 7, Table 7. [26] M. Kulkarni (2026) Reason or recite: complementary specialists with an LLM judge for traffic anomaly reasoning. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.3, Table 3. [27] B. A. Kyem, R. Dzinyela, J. K. Asamoah, E. Denteh, A. Danyo, S. Owusu-Ansah, D. G. Asiedu, and A. Aboah (2026) StateBridge: phase- and track-grounded reasoning for synthetic-to-real traffic video captioning and VQA. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.2, Table 2. [28] K. M. Le, H. D. T. Pham, L. T. Danh, N. Le, H. A. Ngo, P. H. V. Tran, S. N. M. Le, N. T. Nghia, T. T. T. Cam, H. M. N. Nguyen, and C. T. Nguyen (2026) GeoRoute: geometry-aware hybrid inference for traffic future-frame prediction. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.5, Table 5. [29] V. Le, D. L. H. (. Group, V. N. U. Hanoi), H. N. (. A. Institute, V. N. U. Hanoi), H. Tran, and L. Q. Tran (2026) CoTrack3D: a collaborative 2d-3d framework for indoor multi-target multi-camera tracking. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.1, Table 1. [30] J. Lee, D. Kim, and S. Kim (2026) Generative traffic video forecasting with paired future descriptions via parameter-efficient world model adaptation. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.5, Table 5. [31] P. Li, Q. Xu, S. Bao, Y. Jiang, and Q. Huang (2026) UniTraffic-Agent: unified traffic video reasoning for ai city challenge 2026 track 3 with two out-of-domain evaluations. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.3, ยง5.7, Table 3, Table 7, Table 7. [32] C. Lin (2004) ROUGE: a package for automatic evaluation of summaries. In Text Summarization Branches Out, p. 74โ81. Cited by: ยง4.2, ยง4.3. [33] T. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, P. Dollรกr, and C. L. Zitnick (2014) Microsoft COCO: common objects in context. In European Conference on Computer Vision, p. 740โ755. Cited by: ยง4.4, ยง4.6. [34] Y. Lin, Y. Shi, S. Lockyer, H. T. Madabushi, A. N. Evans, W. Li, Y. Wang, and N. Zhang (2026) TAU-Agent: an agentic retrieval-augmented framework for traffic anomaly understanding. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.3, ยง5.7, Table 3, Table 7, Table 7. [35] J. Luiten, A. Osep, P. Dendorfer, P. Torr, A. Geiger, L. Leal-Taixe, and B. Leibe (2021) HOTA: a higher order metric for evaluating multi-object tracking. International Journal of Computer Vision. Cited by: ยง4.1. [36] Milestone Systems (2024) Hafnia dataset: eccv cross city object detection dataset. Note: https://hafnia.milestonesys.com/datasets/ad88fd6e-d608-46e9-8d23-fbe94734cf19Part of the Hafnia project, version 0.0.1 Cited by: ยง3.6. [37] Milestone Systems (2026) Hafnia Training-as-a-Service. Note: https://hafnia.milestonesys.com/Accessed: 2026-08-13 Cited by: ยง3.6. [38] N. D. Minh, M. N. H. Tuan, L. H. L. Kim, L. V. N. Minh, H. P. D. Minh, B. Tran, T. Do, Thanh, D. Ngo, D. Le, and S. Satoh (2026) A semantic state bridge for Sim2Real traffic-safety VQA and captioning. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.2, Table 2. [39] MLCG Lab (2026) SynWTS dataset. Note: https://huggingface.co/datasets/mlcglab/synwtsAI City Challenge Track 2 synthetic traffic-safety dataset Cited by: ยง3.2. [40] M. Naphade, S. Wang, D. C. Anastasiu, Z. Tang, M. Chang, Y. Yao, L. Zheng, M. S. Rahman, M. S. Arya, A. Sharma, Q. Feng, V. Ablavsky, S. Sclaroff, P. Chakraborty, S. Prajapati, A. Li, S. Li, K. Kunadharaju, S. Jiang, and R. Chellappa (2023) The 7th AI city challenge. arXiv preprint arXiv:2304.07500. Cited by: ยง1. [41] T. K. Nguyen, T. Vo, and M. Tran (2026) Action-aligned retrieval with pairwise multimodal reranking for text-based person anomaly search. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.4, Table 4. [42] T. T. G. Nguyen, T. N. Vo, N. H. T. Bui, and H. D. Bui (2026) JEPA guided diffusion: predictive vision-language conditioning for generative traffic forecasting. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.5, Table 5. [43] NVIDIA (2026) PhysicalAI-smartspaces dataset. Note: https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpacesAI City Challenge Track 1 dataset Cited by: ยง3.1. [44] NVIDIA (2026) PhysicalAI-traffic-anomaly-reasoning dataset. Note: https://huggingface.co/datasets/nvidia/PhysicalAI-Traffic-Anomaly-Reasoning Cited by: ยง3.3. [45] A. K. Pandey and M. Chang (2026) TRACE-TAR: reusable video evidence and shared event memory for transportation anomaly reasoning. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.3. [46] K. Papineni, S. Roukos, T. Ward, and W. Zhu (2002) BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, p. 311โ318. Cited by: ยง4.2, ยง4.3. [47] J. Park and J. Lee (2026) Preserving priors, resolving collisions: complementary reranking and injective assignment for text-based person re-identification (sim2real). In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.4, Table 4. [48] D. Patil, R. Kachhadiya, A. Vattuone, H. Liang, J. Li, Y. Wang, Z. Tang, A. Kumar, Q. Kong, and D. C. Anastasiu (2026) SynWTS: multi-view digital-twin dataset for synthetic-to-real traffic safety understanding. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง3.2. [49] H. D. T. Pham, P. H. V. Tran, T. M. Duc, S. N. M. Le, L. M. Khang, H. Vo, M. Phung, H. M. N. Nguyen, and C. T. Nguyen (2026) FaLCon: facet-anchored retrieval with late consensus for sim2real text-based person anomaly search. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.4, Table 4. [50] L. H. Pham, Q. P. Ho, H. Nguyen, D. N. Tran, N. D. Huynh, C. Q. Le, H. Nguyen, H. Jeon, C. D. Tran, S. H. Phan, D. K. Vu, T. L. B. Khanh, and J. W. Jeon (2026) Rethinking pre-training and augmentation for zero-shot cross-city object detection. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.6, Table 6. [51] P. H. Phat, C. T. Bang, K. D. Vu, and D. D. Nguyen) (2026) Cross-city traffic object detection: a targeted augmentation and class-aware inference approach. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.6, Table 6. [52] J. Shangguan, S. Kim, H. Huang, W. Sun, J. Lu, P. Kim, K. Kim, and J. Hwang (2026) UniTraffic: evidence-centric agentic video reasoning across traffic anomalies, violations, and intent. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.7, Table 3, Table 7, Table 7. [53] P. Shrestha, H. Nakayama, and A. Scott (2026) Online multi-camera 3d tracking via id prediction over recurrent sparse queries. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.1, Table 1. [54] Z. Tang, S. Wang, D. C. Anastasiu, M. Chang, A. Sharma, Q. Kong, N. Kobori, M. Gochoo, G. Batnasan, M. Otgonbold, F. Alnajjar, J. Hsieh, T. Kornuta, X. Li, Y. Zhao, H. Zhang, S. Radhakrishnan, A. Jain, R. Kumar, V. N. Murali, Y. Wang, S. S. Pusegaonkar, Y. Wang, S. Biswas, X. Wu, Z. Zheng, P. Chakraborty, and R. Chellappa (2025) The 9th AI city challenge. arXiv preprint arXiv:2508.13564. Cited by: ยง1. [55] D. N. Tran, N. D. Huynh, C. Q. Le, H. Nguyen, L. H. Pham, H. Nguyen, Q. P. Ho, T. L. B. Khanh, C. D. Tran, D. K. Vu, S. H. Phan, H. Jeon, and J. W. Jeon (2026) Syn2RealTrack: bridging the gap between synthetic and real world dataset for online multi-view multi-target tracking. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.1, Table 1. [56] A. Traore, A. Couturier, and E. Hervet (2026) SCOUT: sim-to-real text-based person retrieval by embedding-space prediction over frozen video features. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.4, Table 4. [57] D. Truong, K. D. Dinh, M. H. Hoang, cuong tien nguyen, H. T. Cao, and T. V. Luong (2026) EDCR: evidence-driven chained reasoning for unified traffic anomaly understanding. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.3, Table 3. [58] C. Tsai, Y. X. Li, J. Hsieh, and M. Chang (2026) Decoupled two-stage fusion with scale-aware geometric refinement for cross-city object detection. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.6, Table 6. [59] T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly (2018) Towards accurate generative models of video: a new metric and challenges. arXiv preprint arXiv:1812.01717. Cited by: ยง4.5. [60] R. Vedantam, C. L. Zitnick, and D. Parikh (2015) CIDEr: consensus-based image description evaluation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 4566โ4575. Cited by: ยง4.2, ยง4.3. [61] T. Vo-Lan, P. Nguyen-Ngoc, T. Q. Khanh, and V. T. Thanh (2026) STAR: action-aware two-stage retrieval for text-based person anomaly search. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.4, Table 4. [62] H. Vu, T. T. T. Cam, L. N. T. Toan, H. Vo, D. T. Hieu, H. D. T. Pham, K. M. Le, and H. M. N. Nguyen (2026) Heterogeneous vision-language ensemble with disagreement-aware reranking for text-based person anomaly retrieval. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.4, Table 4. [63] S. Wang, Y. Lu, Z. Tang, H. Zhang, D. C. Anastasiu, and R. Tian (2026) PSI-VQA: evaluating multimodal LLM social reasoning under ambiguous pedestrian intent in automated driving. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง3.3. [64] S. Wang, D. C. Anastasiu, Z. Tang, M. Chang, Y. Yao, L. Zheng, M. S. Rahman, M. S. Arya, A. Sharma, P. Chakraborty, S. Prajapati, Q. Kong, N. Kobori, M. Gochoo, M. Otgonbold, F. Alnajjar, G. Batnasan, P. Chen, J. Hsieh, X. Wu, S. S. Pusegaonkar, Y. Wang, S. Biswas, and R. Chellappa (2024) The 8th AI city challenge. arXiv preprint arXiv:2404.09432. Cited by: ยง1. [65] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli (2004) Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing 13 (4), p. 600โ612. Cited by: ยง4.5. [66] X. Wu, H. Fang, X. Xi, Y. Tang, D. Xu, X. Bai, and D. Liang (2026) Aligning learned spatial priors: gauge canonicalization for multi-camera 3d perception under sim-to-real transfer. In ECCV Workshops, Malmรถ, Sweden. Cited by: ยง5.1, Table 1. [67] H. Zhang, Y. Zhao, Z. P. Bhat, Z. Tang, V. Praveen, V. N. Murali, D. C. Anastasiu, and T. Kornuta (2026) From detection to understanding: tar and tar-bench for multi-task traffic anomaly reasoning. External Links: 2608.10317, Link Cited by: ยง3.3, ยง4.7. [68] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang (2018) The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 586โ595. Cited by: ยง4.5.