Paper deep dive
Is Your Trajectory Displacement Safe in Long-tail?
Qiao Sun, Weicheng Zheng, Yixin Huang, Hang Zhao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 6/20/2026, 8:10:49 AM
Summary
The paper introduces FluidTest, an evaluation pipeline designed to address the limitations of current autonomous driving benchmarks in long-tail scenarios. Existing metrics like Rater Feedback Score (RFS) and closed-loop simulations often saturate or fail to align with human safety preferences. FluidTest utilizes a three-agent verification system (Prosecutor, Detector, and Judge) powered by VLMs/LLMs and a 32-category semantic threat taxonomy to detect 'additional threats' introduced by a planner's trajectory relative to an expert reference. The study demonstrates that state-of-the-art planners (like Poutine and RAP) can exhibit significant safety-relevant failures despite high RFS or low displacement error (ADE), and that the proposed No Additional Threat Rate (NATR) is a more stable and human-aligned metric for identifying long-tail hard cases.
Entities (9)
Relation Signals (4)
FluidTest â comprises â Prosecutor
confidence 100% · The Prosecutor proposes candidate threat classes; the Detector verifies them...
Poutine â evaluatedon â WOD-E2E
confidence 100% · Experiments on the WOD-E2E show that FluidTest produces consistent labels among trained annotators... for Poutine trajectories
FluidTest â usesmetric â NATR
confidence 100% · We define NATR using a taxonomy of 32 semantic threats... and define No Additional Threat Rate (NATR) as the main metric.
Prosecutor â proposesthreats â 32 semantic threats
confidence 90% · The Prosecutor uses VLM semantics to propose candidate threats
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing evaluation pipelines are rarely human-aligned, safety-aware, verifiable, and explainable at the same time: closed-loop metrics often saturate among strong planners, while unstructured human ratings can be noisy without a carefully designed protocol. We formulate planning evaluation as additional-threat detection: given a planner trajectory and an expert reference, does the planner's displacement introduce new unsafe driving behavior? We propose FluidTest, an evaluation pipeline with three components: a pairwise WebUI protocol for reliable human annotation; a taxonomy of 32 semantic threats with evidence-grounded decision graphs; and a three-agent verification system with reflection for precision and auditability. Experiments on the WOD-E2E dataset show that FluidTest produces consistent labels among trained annotators and identifies additional threats in 65% of Poutine trajectories and 51% of RAP trajectories. These results show that state-of-the-art planners can still exhibit substantial safety-relevant failures despite high Rater Feedback Scores (RFS) and low Average Displacement Error (ADE). Additional details, guidance, and code are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2606.16313v1
- Canonical: https://arxiv.org/abs/2606.16313v1
Trouble viewing inline? Open PDF directly â
Full Text
60,640 characters extracted from source content.
Expand or collapse full text
Is Your Trajectory Displacement Safe in Long-tail? Qiao Sun 1 Weicheng Zheng 1,3 Yixin Huang 1,3 Hang Zhao 1,2â 1 Shanghai Qi Zhi Institute 2 Tsinghua University 3 Tongji University Abstract: Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing evaluation pipelines are rarely human-aligned, safety-aware, verifiable, and explainable at the same time: closed-loop metrics often saturate among strong planners, while unstructured human ratings can be noisy without a carefully designed protocol. We formulate planning evaluation as additional-threat detection: given a planner trajectory and an expert reference, does the plannerâs displacement introduce new unsafe driving behavior? We propose FluidTest, an evaluation pipeline with three components: a pairwise WebUI protocol for reliable human annotation; a taxonomy of 32 semantic threats with evidence-grounded decision graphs; and a three-agent verification system with reflection for precision and auditability. Experiments on the WOD-E2E show that FluidTest produces consistent labels among trained annotators, and identifies additional threats in 65% of Poutine, and 51% of RAP trajectories. These results show that state-of-the-art planners can still exhibit substantial safety-relevant failures despite high Rater Feedback Scores (RFS) and low average displacement error (ADE). Keywords: autonomous driving, long-tail scenarios, human-aligned benchmarking 1 Introduction Scaling has improved autonomous driving planning across much of the data distribution, including some tail scenarios [1,2]. As performance on average cases improves, residual risk increasingly concentrates in the deepest tail [3,4]. A common evaluation approach is to run closed-loop simulations initialized from real-world scenarios [5,6,7], using rule-based checks and aggregate metrics. This pipeline is efficient but limited. First, closed-loop metrics are not always aligned with human driving preferences [8], creating interaction-level sim-to-real gaps: a planner may accelerate instead of yielding and still avoid a simulated collision, even though the behavior would be unsafe or undesirable in real traffic. Second, closed-loop benchmarks require accurate 3D reconstruction of maps, traffic controls, and agent corridors to measure events such as red-light violations and collisions, which limits their applicability in long-tail scenes with complex topology, occlusions, or poor lighting. Third, existing metrics often reduce safe driving to a narrow set of event checks, making them too coarse for real-world driving semantics [9]. They can also be vulnerable to reward hacking and score saturation, as shown in Fig. 1. These limitations motivate benchmarks such as WOD-E2E, the Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios [10]. WOD-E2E covers diverse challenging scenarios, and its Rater Feedback Score (RFS) captures human preferences. However, its current testing pipeline does not use the expert trajectory as a reference, even though expert behavior provides crucial scene context [11]. Our experiments identify three limitations: scored samples are sparse relative to the output space of modern planners; scalar scores can be unreliable across planners and scenes; and the scores lack explanations for failure diagnosis. â Corresponding author: hangzhao@mail.tsinghua.edu.cn arXiv:2606.16313v1 [cs.RO] 15 Jun 2026 20252026 92 94 96 98 PDMS NAVSIM v1 navtest PDMS DiffusionDriveV2 DriveSuprim DriveSuprim Top-256 ReflectDrive-2 DriveWorld-VLA SafeDrive SparseDriveV2 ExploreVLA CLOVER Curious-VLAâ ReflectDrive-2â Jun 25Oct 25 7.6 7.8 8.0 8.2 RFS Overall WOD-E2E RFS HMVLM UniPlan Swin-Trajectory DiffusionLTF Poutine RAP-DINO VMA reported resulttopN resulthuman reference 20252026 87 88 89 90 EPDMS NAVSIM v2 EPDMS DriveSuprim WAM-Diff SafeDrive ExploreVLA SparseDriveV2 CLOVER Figure 1: Existing benchmarks are saturated, making it difficult to distinguish among planners and potentially obscuring safety-relevant failures in long-tail driving scenarios. The x-axis indicates the first release date of each methodâs report. Driving Quality Efficiency Issues Defensive Driving Meaningless Actions Blocking traffic without necessity (S1) Low efficiency driving (S2) 41% Stops without escape room (S1) Following too closely (S2) 2% Meaningless lane change (S1) Meaningless lateral drift (S1) Aggressive weaving (S2) < 1% Navigation Plannings Navigation Not following navigation (S1) Not following temporary instructions (S1) Late navigation lane preparation (S1) 3% Collisions & Crashes Collision Risk Collision with static objects (S3) Collision with stationary vehicle (S3) Vehicle collision course (S2) 27% Speeding & Longitudinal Speeding Speeding posted limit (S2) Unsafe speed for condition (S2) Speed contest or exhibition (S2) < 1% Signals & Priority Signal Control Red light violation (S1) Yellow phase noncompliance (S2) Stop sign no full stop (S2) Red light turning no full stop (S2) 3% Right of Way Failure to yield pedestrian (S2) Failure to yield vehicle (S2) 3% Lane Behaviors & Lateral Lane Use Violations Lane straddling (S1) Wrong side of road (S2) Wrong way one way road (S3) 10% Lateral Movements Unsafe lane change (S2) Unsafe passing (S2) Noncommittal merge (S1) < 1% Roadway Position Off road driving (S1) Improper shoulder pass (S2) 17% Special Lane Misuse driving in bike lane (S2) driving in bus only lane (S2) 3% Figure 2: Threat taxonomy used in FluidTest. The taxonomy is substantially more comprehensive than NAVSIM v1 (5 metrics) and NAVSIM v2 (9 metrics). The percentage in each box indicates the violation rate for the Poutine planner on the WOD-E2E val151 subset. We argue that manually designed metrics such as PDMS introduce inductive biases that can weaken reliability on hard cases (Sec. 4.3). Since pairwise comparisons are generally easier for annotators than scalar ratings [12], we ask annotators to directly answer: does the plannerâs displacement from the expert trajectory introduce an additional unsafe driving behavior? Each displacement is classified asNo Threat,Has Threat, orNot Sure. We further organize unsafe driving behavior into 32 semantic threat categories, parsed from the California DMVâs Driver Handbook [13], and define No Additional Threat Rate (NATR) as the main metric. We then build FluidTest, a three-agent verification system with a reflection, revision loop to improve reliability, precision, and explainability than other existing benchmarks. For evaluation, we filter the WOD-E2E validation set using three criteria to reduce noise, yielding 151 scenarios 2 : no obvious future uncertainty, no trivial constant-velocity expert plan, and no severe camera-parameter errors for trajectory projection. We benchmark two state-of-the-art planners with diverse implementations: Poutine 3 [11], a VLM-based planner, and RAP [14], a classical end-to-end planner, as well as an internal high-performing planner, denoted VMA 4 . Our WebUI and labeling 2 Following the WOD-E2E setting, each scene corresponds to the critical frame extracted from a long driving log. This subset is already large compared with the 136 non-long-tail logs in navtest [6]. 3 We use a self-reimplemented version because no official code or checkpoint has been released. Our implementation achieves similar RFS (7.89) and 5sADE (2.77), matching Poutine-Base on the test set. 4 VMA achieves SOTA performance on the official WOD-E2E benchmarks for both RFS (8.06) and 5sADE (2.61). VMA is included only as an additional stress test. 2 Space-excluded poutine scenes (n=108); columns sorted by label pattern Labeller C (0y) Labeller E (3y) Labeller A (researcher) Labeller D (researcher) Labeller B (researcher) A. Binary decisions N Y CEADB C E A D B 1.00 44/44 0.59 30/51 0.61 40/66 0.66 42/64 0.64 44/69 0.59 30/51 1.00 37/37 0.50 33/66 0.57 36/63 0.54 37/69 0.61 40/66 0.50 33/66 1.00 62/62 0.75 53/71 0.87 61/70 0.66 42/64 0.57 36/63 0.75 53/71 1.00 62/62 0.85 60/71 0.64 44/69 0.54 37/69 0.87 61/70 0.85 60/71 1.00 69/69 B. Y-label overlap 0.0 0.2 0.4 0.6 0.8 1.0 Y-label Jaccard Figure 3: Heatmaps of human labelersâ results on the val151 set using Poutineâs planning outputs. The researcher-group results appear at the bottom of the left panel and the lower right of the right panel, showing high overlap and strong labeling consistency. Scenarios labeled "Unsure" are excluded. protocol produce consistent labels across planners, with an overlap rate of 79.31% and a FleissâÎș value of 0.7145 among researcher annotators. The final evaluation results show that both average displacement error (ADE) and RFS are misaligned with human preferences as measured by NATR. Our contributions are threefold: 1.We show that saturated PDMS and RFS fail to capture complex planning threats, and we propose No Additional Threat Rate (NATR) for human-aligned safety evaluation. 2.We introduce FluidTest, a safety-aware, explainable, and comprehensive protocol for benchmarking planning trajectory quality in long-tail scenarios. 3.We release a WebUI and leaderboard for benchmarking new planners, along with labeling code, training code, checkpoints for the threat-detection Prosecutor model, and core multi- agent testing components, including definitions and prompts. 2 Evaluations for Long-tail Planning Autonomous driving planners should perform at least as safely and reliably as human drivers. Yet current planners can still underperform humans in long-tail scenarios, and the observed worst-case distribution can depend strongly on the objectives used during optimization and evaluation. We argue that existing evaluation pipelines are both misaligned with human driving preferences and insufficiently discriminative for separating strong planners from weak ones. Simulation benchmarks are saturated.Closed-loop simulation benchmarks aim to align planner evaluation with driving safety, but several limitations prevent them from fully reflecting real-world complexity. Many rely on hand-crafted or limited scenario sets [15,9]; use simplified behavior models [5] or non-reactive log replay [6]; and reduce safe driving to a small set of predefined events. These choices can create unrealistic interaction patterns, including an implicit âothers always yieldâ bias. As a result, planner scores can quickly saturate near or even above reported human-level performance, leaving dangerous long-tail planning failures undetected. Human rater feedback is useful but not sufficient. Improvements on average-case scenarios do not necessarily reduce rare but safety-critical failures [3,4]. WOD-E2E [10] addresses this issue by collecting challenging long-tail driving segments and using human rater feedback to evaluate end-to-end planning outputs. However, the evaluation protocol becomes the next bottleneck: although RFS is annotated by human raters, our experiments show that it does not consistently reflect human safety preferences and provides limited diagnostic information. VLMs as verifiers and simulators.A practical long-tail planning protocol requires both simulation and verification. Accurate simulation in complex interactive scenarios can be as difficult as the 3 Prosecutor High-Recall Filter A Additional Threats? Classify Threats Detector Reliability & Auditability B Ego States Trajectory Navigation 6 Camera Images w Trajectories VFM Masks (SAM 3 or YOLOPX) Threat label: Off Road Driving Judge High Precision Check C Judge Critique Final Decision Need Changes Ready to Rule Max Rounds Evidence Enough! Q: Does the predicted ego trajectory leave the ordinary roadway and enter a sidewalk, shoulder, curb area, dirt, gore, parking apron, private-property drive/frontage area, or another clearly non-roadway surface? A: Y Q: Is the predicted non-roadway travel explained by emergency recovery, breakdown, directed diversion, or an obviously necessary property-access maneuver? A: N Q: Does the predicted off-road or other non-roadway departure persist long enough to confirm the threat? A: Y No Threats Has Threats The Threat Library ... Human Review No Threats Has Threats Abstain Figure 4: Overview of the FluidTest pipeline. Human annotators, or a VLM, label additional threats. Codex powers a multi-agent critique-and-repair loop that resolves ambiguous cases before producing a confirmed threat set with decision graphs and explanations. Human reviewers make final decisions for abstained cases. planning problem itself [16]. Recent progress suggests that vision-language models (VLMs) and large language models (LLMs) can serve as useful world models [17,18]. Additonally, recent LLM-symbolic systems show that LLMs are more reliable when they translate unstructured inputs into structured programs, logical forms, or planning representations, while symbolic modules perform execution, inference, or validation [19,20,21]. This motivates our design: FluidTest uses VLM/LLM agents for semantic scene understanding and threat proposal, but constrains final decisions through evidence-grounded decision graphs. This hybrid design is especially suitable for driving, where neural perception must be combined with interpretable rule and safety reasoning [22]. 3 FluidTest 3.1 Reasoning and Metrics Driving in complex long-tail scenarios requires a planner to avoid risky maneuvers, even when they do not cause an immediate collision within the planning horizon. This makes explicit simulation both expensive and insufficient, so FluidTest directly evaluates whether a planner trajectory introduces additional threats relative to the expert trajectory. We define NATR using a taxonomy of 32 semantic threats, as shown in Fig. 2, Appendix G, and Appendix G. These threats cover deterministic failures, interaction risks, roadway and lane-use errors, route-compliance failures, and low-quality driving behaviors. Let N =|S| be the number of evaluated scenes. For each scenesâS, FluidTest outputs confirmed threat labelsY s âL, whereL is the full threat set. For any threat subset AâL, we define z s (A) = 1[Y s â© AÌž=â ], r(A) = 1 N X sâS z s (A),NATR(A) = 1â r(A). Here,r(A)is the violation rate for threat setA, andNATR(A)is the corresponding no-violation score. The overall NATR is obtained by setting A =L. 3.2 Human Annotation Protocol and WebUI We collect labels through pairwise additional-threat annotation. Instead of assigning an absolute score, annotators compare the expert trajectoryÏ â s and planner trajectoryËÏ s under the same visual context and answer: Does the planner trajectory introduce any additional driving threat compared 4 0.05660.0979 0.1690.2930.5060.874 1.512.614.51 7.8 3s ADE 0 20 40 60 80 100 0.0999 0.1770.3140.5570.988 1.753.115.519.7717.3 5s ADE 0 20 40 60 80 100 0.1420.2490.4380.772 1.362.394.217.41 1323 FDE 0 20 40 60 80 100 4.34.95.56.16.77.37.98.59.19.7 RFS 20 30 40 50 60 70 Threat rate (%) PoutineRAP Figure 5: Threat rate versus open-loop metrics and RFS. All four panels show the percentage of scenes with confirmed additional threats after binning trajectories by 3sADE, 5sADE, FDE, and RFS, respectively. Non-trivial threat rates still persist in high-RFS regions, showing that these aggregate metrics do not fully capture safety-critical planning failures. See more cases in Appendix C. DiffusionDrive WoTE ReCogDrive DriveVLA-W0 DriveSuprim DiffusionDrive PDMS 88.1 WoTE PDMS 88.3 ReCogDrive PDMS 90.8 DriveVLA-W0 PDMS 93.0 DriveSuprim PDMS 93.5 100.0%44.0%38.6%32.2%24.2% 37.0%100.0%28.0%23.6%17.0% 28.1%24.2%100.0%21.6%20.9% 38.0%33.0%35.0%100.0%21.2% 26.1%21.8%30.9%19.4%100.0% (a) NAVSIM PDMS=0 Poutine RAP Poutine RAP 100.0%56.2% 56.2%100.0% (b) Worst 10% 5s ADE Poutine RAP Poutine RAP 100.0%25.0% 25.0%100.0% (c) Worst 10% RFS Poutine RAP Poutine RAP 100.0%73.0% 57.5%100.0% (d) FluidTest Threats 0.0 0.2 0.4 0.6 0.8 1.0 Share of column model Figure 6: Overlap analysis for long-tail planning evaluation. The heatmaps compare scenario sets selected by different metrics and planners, showing that existing metrics identify planner-dependent hard cases rather than stable long-tail failures. with the expert trajectory? We find that the expert trajectory is essential for both stable human labels and reliable VLM-based explanations. Annotators choose among three labels:Y, meaning the planner introduces at least one additional threat; N, meaning no additional threat despite possible geometric deviation; andNot Sure, used when future uncertainty, unexplained expert behavior, or noisy projection prevents a reliable binary decision. We treatNot Sureas abstention. Appendix A gives the full WebUI guidance, and Appendix B reports labeling-time statistics. 3.3 Testing Protocol and FluidTest Safety Arena (Gold) FluidTest combines human labels with agentic verification and precision control. After binary labeling, the Prosecutor proposes candidate threat classes; the Detector verifies them using evidence- grounded decision graphs; and a reflection loop revises ambiguous cases before producing the final confirmed threat set. Abstained cases are sent to human reviewers for final review. We use Codex as the main agentic pipeline because it can invoke tools such as cropping and zooming when additional visual evidence is needed. The overall pipeline is shown in Fig. 4; additional graph and Prosecutor details are provided in Appendix H and Appendix E. The three-agent loop follows a neuro-symbolic verification design. The Prosecutor uses VLM semantics to propose candidate threats, the Detector verifies them through structured decision graphs, and the Judge performs reflection and consistency checks. This mirrors prior LLM-symbolic systems that use LLMs for flexible semantic translation while delegating reasoning, execution, or validation to structured symbolic modules [19,20,21]. Tool use and reflection further improve auditability: 5 Table 1: Final FluidTest scores on val151 after human review. Entries are no-violation rates; higher is better except for 5sADE. NATR denotes the overall no-additional-threat rate. The low numbers highlight a large gap between AI planners and human driving. MethodQualityâNav.âColli.âLong.âPrior.âLateralâADEâNATRâ Poutine [11]0.670.990.830.990.900.972.440.35 VMA0.721.000.841.000.930.972.460.47 RAP [14]0.710.990.910.990.880.982.620.49 Table 2: Threat-detection comparison for the Prosecutor. Without fine-tuning, large VLMs perform near random; supervised fine-tuning (SFT) results are evaluated on a 10% test split. ModelDatasetAccPrecisionRecallF1F2 Qwen3.5 9BPoutine-Union55%67%46%54%49% ChatGPT 5.4-Mini xHighPoutine-Union50%100%8%15%10% Qwen3.5 4B SFTPoutine-Union83%75%100%85%93% Qwen3.5 9B SFTPoutine-1K79%85%62%71%65% Codex can gather additional visual evidence when needed, and the Judge can request more evidence or reject unsupported conclusions [23, 24]. FluidTest can be deployed locally for indicative evaluation. Controlled cross-planner comparison, however, requires a shared testing protocol. We therefore release a testing server where users submit planning results against a dynamic closed test set. For each submission, several human annotators will label additional threats, after which the full FluidTest pipeline will classify, explain, and score the planner. Results will be reported on the FluidTest Safety Arena (Gold) leaderboard. 4 Experiment Results and Insights We evaluate FluidTest through four questions. First, does the labeling protocol produce consistent human labels? Second, what do the final FluidTest scores reveal about current planners? Third, does a metric identify stable long-tail hard cases across planners? Fourth, can common open-loop and human-feedback metrics substitute for direct NATR evaluation? 4.1 Human Label Consistency and NATR Learning Human consistency requires task-specific prior knowledge. We ask five annotators, including three project researchers, to label the val151 subset with our WebUI for Poutine planning results. As shown in Fig. 3, the researchers produce more consistent labels than the non-researcher group. We find that non-researchers have difficulty reconstructing driving motion from projected trajectories, making it harder to anticipate future motion and identify potential threats. Within the researcher group, an overlap rate of 79.31% and a Fleissâ Îș value of 0.7145 indicate high agreement. Importantly, this agreement is achieved under a deliberately challenging labeling protocol. High consistency is easy to obtain with narrow rules and hard boundaries. Instead, FluidTest asks annotators to conservatively identify all plausible additional threats relative to the expert trajectory, including potential future risks. This requires complex reasoning about all possible worst-case outcomes in defensive driving. Learning human safety preferences requires additional alignment. We use the Poutine union set for evaluation, defined as scenarios labeledYby at least one researcher annotator. After excluding uncertain cases, both Qwen3.5-9B (no thinking) and ChatGPT-5.4-Mini xHigh perform close to random on the binary task of predicting additional threats. In contrast, a fine-tuned Qwen3.5 model 6 RFS samples x 3 RAP Poutine Expert RFS: 8 RFS: 10 RFS: 9 Figure 7: RFS is not guaranteed to reflect critical safety hazards. Left: sparse labels do not cover the expert solution. Right: a small displacement in the wrong direction creates severe collision risk. performs substantially better, suggesting that human safety preferences require additional alignment, as shown in Table 2 and more in Appendix E. 4.2 Final FluidTest Benchmark Results Table 1 summarizes the final FluidTest results. The strongest 5sADE result does not imply the safest behavior: Poutine has the lowest 5sADE (2.44) but the lowest NATR (0.35), meaning that 65% of its trajectories introduce at least one additional threat. RAP has the worst 5sADE (2.62) but the best NATR (0.49), while VMA ranks between them overall and has the best Quality and Priority scores. These results directly support our core claim: displacement quality and scalar human feedback do not reliably measure safety-relevant behavior in long-tail scenes. 4.3 Reliable Hard-Case Metrics Should be Stable Across Planners A useful long-tail benchmark should capture scenario-level difficulty rather than planner or metric related artifacts. Motivated by item-level evaluation theory and cross-model response analyses [25, 26,27,28,29,30], we evaluate failure-case universality: whether hard cases selected by a metric are shared across planners. Using random selectors as a baseline, the expected Jaccard overlap can be low. A metric that captures genuine scenario difficulty should identify substantially overlapping failure sets, even if the detailed failure modes differ. Closed-loop PDMS is weakly universal.We first evaluate five open-source state-of-the-art planners with diverse implementations on NAVSIM [31,32,33,17,34]. For each planner, we define failures as scenarios withPDMS = 0. As shown in Fig. 6, the overlap between planner failure sets is only 17% to 44%, suggesting that PDMS failures are strongly shaped by each plannerâs output distribution rather than by a shared set of universally difficult scenes. RFS is less universal than open-loop displacement.Open-loop displacement provides an objective baseline for testing whether different planners share the same long-tail hard cases. A reliable long-tail metric should not produce lower cross-planner overlap than this baseline. On the val151 subset, the worst 10% 5sADE hard tails of Poutine and RAP overlap by 56.2%. Repeating the analysis with Rater Feedback Score (RFS), using the worst 10% of scenarios for each planner, yields only 25% overlap. This lower overlap indicates that RFS is a less stable hard-case selector, likely due to sparse labeling coverage and noise in scalar ratings. NATR is more reliable and informative. Under NATR, theYsets overlap by 73%, substantially higher than the other metrics. This indicates that safety-relevant hard cases are largely shared across state-of-the-art planner families. Appendix D reports threshold-sensitivity analyses showing that 5sADE and RFS need to cover about half of the dataset to reach comparable overlap, which no longer represents a hard tail. 7 Overall, low scores under existing metrics do not necessarily correspond to universal long-tail hard cases. Closed-loop PDMS and scalar RFS expose some failures, but their hard-case subsets remain heavily planner-dependent. FluidTest provides a more stable, interpretable, and planner-independent benchmark for long-tail planning safety. 4.4 Relationship with Existing Metrics We examine whether existing open-loop displacement metrics and RFS can serve as proxies for direct additional-threat labels used to compute NATR. For each pair, we partition the val151 scenarios into ten bins according to the metric value and report the percentage of human-labeled additional-threat cases in each bin. This analysis tests whether metric scores alone can recover the safety-relevant failures identified by FluidTest. The results show that no existing metric reliably predicts NATR outcomes, especially when comparing across multiple planners. Average displacement errors are limited safety indicators. Fig. 5 shows that 3sADE, 5sADE, and FDE are informative but limited. Displacement errors correlate with hazards within a single plannerâs output distribution, but the relationship weakens across planners. As shown in Fig. 7, similar deviations can either create severe collision risk when they occur in the wrong direction or reflect a safe alternative mode. Thus, displacement errors are useful imitation-quality indicators but cannot serve as standalone safety metrics. Longer-horizon displacement is more threat-sensitive. Among displacement metrics, 5sADE and FDE align better with additional threats than 3sADE. Many failures become visible only after the trajectory commits to a future interaction or route choice, such as entering a conflict zone, ending in an inappropriate lane, failing to prepare for a turn, or continuing toward an obstacle. Extreme displacements are therefore more reliably associated with additional threats. Scalar human feedback does not reliably track additional threats. Although RFS uses human judgment, it performs worse than displacement for this purpose because of noise when comparing trajectories within and across planner distributions. Notably, the highest RFS bin still contains 40% additional threats, as shown in Fig. 5. FluidTest instead directly evaluates whether a planner trajectory introduces additional risk relative to the expert trajectory. 5 Conclusion Planning and testing in long-tail scenarios is difficult because such scenarios directly challenge the plannersâ generalization ability. Reliable evaluation should therefore remain grounded in human driving judgment while providing auditable structure. We introduced FluidTest, a human-aligned, safety-aware, verifiable, and explainable evaluation pipeline for long-tail autonomous driving plan- ning. FluidTest reframes planning evaluation as additional-threat detection using expert behavior as a reference. It combines a pairwise WebUI for human annotation, a taxonomy of 32 safety-relevant threats, evidence-grounded decision graphs, and a Judge with reflection and revision. Across Poutine, RAP, and VMA, our results show that additional threats remain frequent and are often hidden by existing aggregate metrics. In contrast, NATR provides a more stable and semantically meaningful view of long-tail planning safety. 6 Limitations FluidTest has three main limitations. First, although we show that annotators can be trained to produce high-quality labels, the scalability of this procedure for labelers with diverse backgrounds remains to be tested. Second, the pipeline still depends on reliable trajectory projection, visual grounding, and sufficient scene context for human labeling. Third, although we can automatically expand the threat taxonomy from natural language guidelines, the current 32-threat taxonomy is not yet exhaustive. 8 References [1]A. Naumann, X. Gu, T. Dimlioglu, M. Bojarski, A. Degirmenci, A. Popov, D. Bisla, M. Pavone, U. MĂŒller, and B. Ivanovic. Data scaling laws for end-to-end autonomous driving. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2025. [2]Q. Sun, H. Wang, J. Zhan, F. Nie, X. Wen, L. Xu, K. Zhan, P. Jia, X. Lang, and H. Zhao. Generalizing motion planners with mixture of experts for autonomous driving. In 2025 IEEE International Conference on Robotics and Automation (ICRA), pages 6033â6039, 2025. doi: 10.1109/ICRA55743.2025.11127274. [3]M. Hallgarten, J. Zapata, M. Stoll, K. Renz, and A. Zell. Can vehicle motion planning generalize to realistic long-tail scenarios? In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5388â5395, 2024. doi:10.1109/IROS58592.2024.10803052. [4]M. OâKelly*, A. Sinha*, H. Namkoong*, R. Tedrake, and J. Duchi. Scalable end-to-end autonomous vehicle testing via rare-event simulation. In Advances in Neural Information Processing Systems, 2018. [5]H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom. nuplan: A closed-loop ml-based planning benchmark for autonomous vehicles. CoRR, abs/2106.11810, 2021. [6]D. Dauner, M. Hallgarten, A. Geiger, and A. Zell. Navsim: Data-driven non-reactive au- tonomous vehicle simulation and benchmarking. CoRR, abs/2406.15349, 2024. [7]W. Cao, M. Hallgarten, T. Li, D. Dauner, X. Gu, C. Wang, Y. Miron, M. Aiello, H. Li, I. Gilitschenski, B. Ivanovic, M. Pavone, A. Geiger, and K. Chitta. Pseudo-simulation for autonomous driving. In Conference on Robot Learning (CoRL), 2025. [8]D. Dauner, M. Hallgarten, A. Geiger, and K. Chitta. Parting with misconceptions about learning- based vehicle motion planning. In Conference on Robot Learning, pages 1268â1281. PMLR, 2023. [9]X. Jia, Z. Yang, Q. Li, Z. Zhang, and J. Yan. Bench2drive: Towards multi-ability benchmarking of closed-loop end-to-end autonomous driving. Advances in Neural Information Processing Systems, 37:819â844, 2024. [10]R. Xu, H. Lin, W. Jeon, H. Feng, Y. Zou, L. Sun, J. Gorman, E. Tolstaya, S. Tang, B. White, et al. Wod-e2e: Waymo open dataset for end-to-end driving in challenging long-tail scenarios. arXiv preprint arXiv:2510.26125, 2025. [11]L. Rowe, R. de Schaetzen, R. Girgis, C. Pal, and L. Paull. Poutine: Vision-language-trajectory pre-training and reinforcement learning post-training enable robust end-to-end autonomous driving. CoRR, abs/2506.11234, 2025. doi:10.48550/arXiv.2506.11234. [12]S. Kiritchenko and S. Mohammad. Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 465â470, 2017. doi:10.18653/v1/P17-2074. [13]California Department of Motor Vehicles. California driverâs handbook, 2026. URLhttps:// w.dmv.ca.gov/portal/handbook/california-driver-handbook/ . Accessed: 2026- 04-16. [14]L. Feng, Y. Gao, E. Zablocki, Q. Li, W. Li, S. Liu, M. Cord, and A. Alahi. Rap: 3d rasterization augmented end-to-end planning. CoRR, abs/2510.04333, 2025. doi:10.48550/arXiv.2510.04333. 9 [15]A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun. Carla: An open urban driving simulator. In Proceedings of the 1st Annual Conference on Robot Learning, volume 78 of Proceedings of Machine Learning Research, pages 1â16, 2017. [16] Q. Sun, X. Huang, B. C. Williams, and H. Zhao. Intersim: Interactive traffic simulation via explicit relation modeling. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 11416â11423. IEEE, 2022. [17] Y. Li, S. Shang, W. Liu, B. Zhan, H. Wang, Y. Wang, Y. Chen, X. Wang, Y. An, C. Tang, L. Hou, L. Fan, and Z. Zhang. Drivevla-w0: World models amplify data scaling law in autonomous driving. CoRR, abs/2510.12796, 2025. doi:10.48550/arXiv.2510.12796. [18] W. Liang, S. Wang, H.-J. Wang, O. Bastani, D. Jayaraman, and Y. J. Ma. Environment curriculum generation via large language models. In Proceedings of The 8th Conference on Robot Learning, volume 270 of Proceedings of Machine Learning Research, pages 433â454, 2025. [19]L. Pan, A. Albalak, X. Wang, and W. Wang. Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 3806â3824. Association for Computational Linguistics, 2023. [20]L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig. PAL: Program- aided language models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 10764â10799. PMLR, 2023. [21] B. Liu, Y. Jiang, X. Zhang, Q. Liu, S. Zhang, J. Biswas, and P. Stone. LLM+P: Empowering large language models with optimal planning proficiency. arXiv preprint arXiv:2304.11477, 2023. [22]J. Sun, H. Sun, T. Han, and B. Zhou. Neuro-symbolic program search for autonomous driving decision module design. In Proceedings of the 2020 Conference on Robot Learning, volume 155 of Proceedings of Machine Learning Research, pages 21â30. PMLR, 2021. [23]S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao. ReAct: Synergizing rea- soning and acting in language models. In International Conference on Learning Representations, 2023. [24]N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao. Reflexion: Language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems, 2023. [25] F. MartĂnez-Plumed, R. B. C. PrudĂȘncio, A. MartĂnez-UsĂł, and J. HernĂĄndez-Orallo. Item response theory in AI: Analysing machine learning classifiers at the instance level. Artifi- cial Intelligence, 271:18â42, 2019. doi:10.1016/j.artint.2018.09.004. URLhttps://w. sciencedirect.com/science/article/pii/S0004370219300220. [26]J. P. Lalor, H. Wu, and H. Yu. Building an evaluation scale using item response theory. In J. Su, K. Duh, and X. Carreras, editors, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 648â657, Austin, Texas, Nov. 2016. Association for Computational Linguistics. doi:10.18653/v1/D16-1062. URLhttps://aclanthology.org/ D16-1062/. [27]P. Madhyastha. Task-aware evaluation and error-overlap analysis for large language models. In A. Sinha, R. VĂĄzquez, T. Mickus, R. Agarwal, I. Buhnila, P. SchmidtovĂĄ, F. Gamba, D. K. Prasad, and J. Tiedemann, editors, Proceedings of the 1st Workshop on Confabulation, Hallucinations and Overgeneration in Multilingual and Practical Settings (CHOMPS 2025), pages 1â10, Mumbai, India, Dec. 2025. Association for Computational Linguistics. ISBN 979-8-89176- 308-1. doi:10.18653/v1/2025.chomps-main.1. URLhttps://aclanthology.org/2025. chomps-main.1/. 10 [28]O. Uzan and Y. Pinter. CharBench: Evaluating the role of tokenization in character-level tasks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 33296â33304, 2026. doi:10.1609/aaai.v40i39.40615. URLhttps://doi.org/10.1609/ aaai.v40i39.40615. [29]S. Vellamcheti, U. K. Kothapalli, D. Bhowmick, and S. N. Aakur. CVT-Bench: Counterfactual viewpoint transformations reveal unstable spatial representations in multimodal llms, 2026. URL https://arxiv.org/abs/2603.21114. [30] J. Kim, C. Lim, S. H. Gil, and S. Lee. EuraGovExam: A multilingual multimodal benchmark from real-world civil service exams, 2026. URL https://arxiv.org/abs/2603.27223. [31] B. Liao, S. Chen, H. Yin, B. Jiang, C. Wang, S. Yan, X. Zhang, X. Li, Y. Zhang, Q. Zhang, et al. Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 12037â12047, 2025. [32]Y. Li, Y. Wang, Y. Liu, J. He, L. Fan, and Z. Zhang. End-to-end driving with online trajectory evaluation via bev world model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 27137â27146, 2025. [33] Y. Li, K. Xiong, X. Guo, F. Li, S. Yan, G. Xu, L. Zhou, L. Chen, H. Sun, B. Wang, et al. Recogdrive: A reinforced cognitive framework for end-to-end autonomous driving. arXiv preprint arXiv:2506.08052, 2025. [34]W. Yao, Z. Li, S. Lan, Z. Wang, X. Sun, J. M. Alvarez, and Z. Wu. Drivesuprim: Towards precise trajectory selection for end-to-end planning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 11910â11918, 2026. 11 A WebUI for Annotators and Labeling Guidance We design a WebUI for annotators to label planning outputs and find that interface design is crucial for stable labels. Each annotator registers their name and driving experience, reads the labeling guidance, and then labels cases using the interface in Fig. 8. The interface includes scene information, trajectory-state data, a playback controller, three stitched front-camera images, and three rear-camera images. Playback covers the previous 1 s for all cameras. Lazy image loading and a magnifier improve labeling efficiency and accuracy. The full guidance text is summarized below. Additional threat (Y) The predicted trajectory introduces an additional threat compared with the expert trajectory. No additional threat (N) The predicted trajectory does not add any threat beyond the expert trajectory. Not sure (Space) The scene cannot be judged from the available evidence. How to label potential hazards. Annotators decide whether the prediction introduces anything meaningfully worse than the expert trajectory. Use Y only when the prediction creates an additional unnecessary maneuver or an additional potential hazard, including possible future risk. Use N when the prediction does not add anything worse than the expert trajectory. UseSpace/Unsure when the expert trajectory has potential hazards but the available evidence is insufficient to understand why or compare fairly. The core reminder is to label added risk relative to the expert trajectory, not whether the scene itself is perfect. B Human Label Cost Labeling cost analysis. Previous experiments show that our NATR metric is more reliable for human labelers than the previous RFS labels on the val151 subset. In this section, we show that our labeling protocol requires minimal labeling effort. To measure labeling cost, we log the approximate labeling time on the val151 subset of WOD-E2E. The average per-sample labeling time is about 5.292 s, meaning that labeling the full val151 subset requires less than 20 minutes. More details are shown in Fig. 11. Training annotators for additional-threat labels. Additional-threat annotation requires task- specific calibration. Among the three researcher annotators, only Labeler B had direct knowledge of the labeling logic, the complete threat taxonomy and its definitions, and the Prosecutor classifier training pipeline. The other two researcher annotators initially produced noisier labels, closer to those of the non-researcher annotators, especially for ambiguous cases requiring strict worst-case safety judgments. We found that this gap was substantially reduced after expanding the labeling guidance with more representative examples. These examples clarified how to compare the planner trajectory against the expert trajectory, when to label plausible future risks asY, and when to abstain because the expert behavior cannot be reliably explained. This suggests that consistent additional-threat labeling is trainable: annotators do not need to know the implementation details of FluidTest, but they do need calibrated examples that define the intended safety preference. Due to time constraints, we did not perform the same training process for the other two non- researcher annotators. Therefore, the lower consistency observed in the non-researcher group should be interpreted as a limitation of our current annotation study rather than evidence that non- researchers cannot produce reliable labels. In the official FluidTest Safety Arena (Gold), we will include a standardized annotator-training stage with example-based guidance, calibration rounds, and consistency checks before collecting final benchmark labels. 12 Figure 8: Screenshot of our WebUI for pairwise additional-threat annotation. The interface helps annotators compare the semantic consequences of two trajectories directly in image space. Annotator disagreement cases.Human drivers vary in driving style and safety preference, ranging from conservative to aggressive. Annotators may also interpret the same scene differently, especially when road topology, lane direction, or obstacle clearance is ambiguous. In addition, labeling errors can occur when critical visual details are missed. Fig. 9 shows four representative disagreement cases. These examples show that human annotation is not perfect even with our WebUI and labeling protocol. Therefore, reliable NATR evaluation requires multiple trained annotators so that individual preferences, scene-interpretation differences, and occasional missed evidence do not dominate the final result. C RFS Failure Cases In this section, we provide additional case studies showing why RFS fails to fully capture long-tail driving threats. WOD-E2E does not use the expert trajectory when computing RFS, aiming to cover multiple plausible future behaviors. However, we argue that, in many cases, trajectories in the val151 subset with an RFS of 10 do not reflect this design goal. We use âRFS10 reference trajectoryâ to denote the reference future trajectory used by the WOD-E2E RFS procedure that receives the maximum score of 10 for that scene. In Fig. 10, we show three cases from three strong planners: Poutine, RAP, and VMA. The RFS values computed for these scenes are all 10, the maximum possible score. In the upper case, however, both the RFS10 reference trajectory and the planner trajectory choose not to yield to the cyclist on the left, whereas the expert trajectory clearly yields; these two 13 Figure 9: Representative annotator disagreement cases. In the top two examples, some reviewers interpreted the scene as a one-way road or a wide road without visible opposing traffic. Under that interpretation, they judged the plannerâs leftward deviation as not introducing an additional threat, even though the expert trajectory stays more clearly to one side of the road. In the bottom two examples, some reviewers judged the deviation as too small to create a threat or missed the deviation without carefully using the magnifier. The bottom-left case shows a leftward deviation toward traffic cones, while the bottom-right case shows a leftward deviation onto the double-yellow lines. In both cases, the planner introduces an additional threat that is not present in the expert trajectory. trajectories therefore introduce additional collision threats. In the lower-left case, the expert trajectory remains stationary throughout the future 5-second planning horizon, while both the RFS10 reference trajectory and the RAP planner slowly move into the intersection, creating potential collision hazards with crossing vehicles. In the lower-right case, the expert trajectory and the RFS10 trajectory turn left on the right side of the road. However, because the WOD-E2E RFS computation depends heavily on displacement and can fail on high-curvature trajectories, the poor VMA trajectory, which heads toward the road edge, still receives a perfect RFS of 10 despite clear additional threats. These results further explain why the highest RFS bin in Fig. 5 still contains a substantial fraction of additional threats. D Worst-Case Overlap Depends on the Tail Threshold Section 4.3 compares failure-case universality across different metrics using hard-case overlap rates. However, for continuous metrics such as 5sADE and RFS, the overlap rate depends strongly on how the long-tail threshold is defined. In the main text, we use the worst 10% of scenarios to represent hard cases. Increasing the threshold naturally increases overlap because larger portions of the dataset are included. 14 291acc4578630f6231a35e718de21523-147291acc4578630f6231a35e718de21523-147 GT YieldsGT Yields GT StopsGT Stops GT TurnsGT Turns RFS 10RFS 10 RFS 10RFS 10 RFS 10RFS 10 PoutinePoutine RAPRAP VMAVMA 06a63756c3dc03b0a1778312 c6468d58-148 06a63756c3dc03b0a1778312 c6468d58-148 1b000869821271980 2e03983c052d7-148 1b000869821271980 2e03983c052d7-148 Figure 10: We show three poor planning results from the val151 subset that receive perfect RFS scores of 10 for different reasons across three planners. The cyan trajectory is the expert trajectory, the green trajectory is the RFS10 reference trajectory, i.e., the maximum-score reference trajectory used by the RFS procedure for that scene, and the red trajectory is the planner output that receives an RFS of 10. To provide a fairer comparison, we report overlap rates under multiple thresholds in Fig. 12 and Table 4. We observe that both 5sADE and RFS require covering roughly half of the dataset before reaching overlap rates comparable to NATR. However, selecting 50% of the dataset no longer meaningfully represents long-tail hard cases and provides limited diagnostic value for understanding why those scenarios are difficult. We further ask whether 5sADE or RFS can identify scenarios where planners fail under NATR. To study this question, we compare the worst 50% of scenarios selected by each metric against human-labeled additional-threat cases. As shown in Fig. 13, both 5sADE and RFS exhibit limited overlap with the human labeling results. This result further indicates that low 5sADE or RFS scores do not reliably correspond to human-identified safety failures under NATR. 15 0.02.55.07.510.012.515.017.520.0 Time per label (seconds) 0 5 10 15 20 25 Label count P50 4.9sP90 12.7s0-90% avg 5.2sMean 6.9s Figure 11: Detailed labeling-cost results measured by labeling time. The mean labeling time is 6.9 s and the median is 4.9 s, indicating that the interface supports efficient labeling for most long-tail scenarios. The labels were collected for VMA planned trajectories. Figure 12: Directional heatmap extending the previous hard-case overlap analysis. We use the Poutine modelâs predictions for this analysis to remain aligned with the previous analysis. The overlap rate of RFS increases after the threshold is raised above 30%, indicating that RFS is unstable specifically on the 3â10% tail and becomes more stable above 50%. E Prosecutor Threat Classifier Table 2 shows that the threat-difference model reaches 79â83% accuracy on the original evaluation sets. This is already strong because the target is not a hard-bounded geometric predicate, such as checking whether two trajectories overlap. Instead, the model must predict complex human safety preferences end-to-end. The remaining errors mainly arise from preference ambiguity. Some cases are inherently ambiguous even for humans. After filtering out samples on which human annotators disagree, the post-filter evaluation becomes much more deterministic: the 2B variant reaches an accuracy of 0.929, while the 4B variant reaches a perfect accuracy of 1.000 on the held-out evaluation set, as shown in Table 3. 16 Figure 13: This per-scene heatmap shows that poor 5sADE and RFS scenarios do not necessarily produce more threats as measured by human labelers under NATR. This result further emphasizes the value of human labeling with the proposed NATR metric. Per-scenario judge threat status by planner 151 shared samples, columns sorted by pattern across RAP / PoutineSFT / VMA. H = has threats, N = no threats, A = abstain. Has threatsNo threatsAbstain RAP H 69 N 77 A 5 PoutineSFT H 76 N 69 A 6 VMA H 66 N 79 A 6 0255075100125150 Scenario columns sorted by status pattern Figure 14: Per-scenario Judge results across all three planners on the val151 subset. The Judge abstentions indicate that the three-agent pipeline can resolve most cases and produce consistent explanations through the predefined decision graph. This suggests that much of the residual error in the unfiltered setting reflects ambiguity in the human preference labels rather than a failure to learn the main threat-difference rule. F Judge Result Details We provide more detailed results from the Judge agent outputs. As shown in Fig. 14, the full pipeline can produce reliable and consistent evidence and decision-graph results, enabling the Judge to rule on most cases. Furthermore, the ratio of abstained cases remains largely consistent across planners, indicating that abstentions do not change the comparison results across different planners. G Threat Definitions and Boundaries We provide the full definitions of all 32 threats in Table 7. These definitions are used consistently by human annotators, result reviewers, and the agentic Codex server. See the accompanying documents and code for further details. H Threat Decision Graphs and Reusable Nodes We refer readers to this online site for detailed decision graphs for all threats, since the full graphs are too large to include in the paper. Each semantic threat is implemented as a structured graph composed of reusable evidence nodes, routing logic, lawful-exception checks, and terminal decision states. We include one representative graph in full and summarizes the terminal conditions for all authored threat graphs used by the evaluator, as shown in Fig. 15. 17 Table 3: Additional results for the Prosecutor additional-threat classifier. After filtering ambiguous cases in which the three human annotators disagree, the fine-tuned model achieves near-perfect prediction performance. ModelAcc.PrecisionRecallF1F2 Qwen3.5 2B SFT0.72000.64290.81820.72000.7759 Qwen3.5 4B SFT0.72000.62500.90910.74070.8333 Qwen3.5 2B SFT (filtered)0.92861.00000.85710.92310.8824 Qwen3.5 4B SFT (filtered)1.00001.00001.00001.00001.0000 Table 4: Scene overlap analysis between Poutine and RAP under different selection criteria. CriterionUnion scenesShared scenesShared / |RAP set|Shared / |Poutine set| 5sADE worst 3%9120.00%20.00% 5sADE worst 10%23956.25%56.25% 5sADE worst 30%652758.70%58.70% 5sADE worst 50%1064660.53%60.53% RFS worst 3%9120.00%20.00% RFS worst 10%28425.00%25.00% RFS worst 30%672554.35%54.35% RFS worst 50%995369.74%69.74% NATR (Ours)974673.02%57.50% H.1 Graph Design Principles Each threat graph follows four principles: 1.Evidence grounding. Every accepted threat must be supported by explicit observable evidence from trajectory geometry, scene context, visual grounding, or temporal behavior. 2.Composable reusable nodes. Shared checks such as lane occupancy, signal state, pedestrian conflict, collision overlap, and roadway-side validation are implemented once and reused across multiple threat graphs. Details are provided in Table 6. 3.Auditable execution. The graph stores intermediate node outputs, queried visual evidence, routing decisions, and lawful-exception reasoning. 4.Multi-terminal reasoning. Graphs do not produce only binary outputs. Possible terminal states include confirmed, unproven, lawful_exception, routed, and not_applicable. H.2 Terminal States As shown in Table 5, we define five terminal states for each threat graph. Theunproven, lawful_exception, andnot_applicablestates all correspond to different forms of âno threatâ for the currently evaluated threat. Theroutedstate triggers an additional graph check for the routed threat. These final decisions are then provided to the Judge for the final determination. 18 Table 5: Terminal states used by FluidTest threat graphs. StateMeaning confirmedEvidence sufficiently supports the threat. unprovenThreat hypothesis is plausible but unsupported by sufficient evidence. lawful_exceptionBehavior appears risky but is justified by traffic context or legal exception. routedEvidence better matches another sibling threat. not_applicablePreconditions for this threat are absent. Table 6: Core reusable nodes shared across multiple threat graphs. Node IDInputFunction NODE_LANE_OCCUPANCYTrajectory + lane masksDetermines whether the predicted trajectory occupies a valid lane region and measures lane overlap duration. NODE_ROAD_SIDE Trajectory + roadway segmenta- tion Determines whether the trajectory travels on the wrong roadway side or exits drivable space. NODE_SIGNAL_STATETraffic-light cropDetects signal state and movement permissions. NODE_STOP_CHECKEgo velocity profile Determines whether the ego reaches a full stop before entering a controlled area. NODE_PEDESTRIAN_CONFLICTPedestrian detections + trajec- tory Detects whether the trajectory conflicts with pedestrian-priority zones. NODE_VEHICLE_CONFLICTVehicle tracks + trajectoryDetermines whether another vehicle possesses right-of-way conflict priority. NODE_STATIC_COLLISIONStatic-object masks + trajectoryDetects overlap or collision course with barriers, cones, curbs, poles, or other fixed objects. NODE_STATIONARY_VEHICLEParked-vehicle masks + trajec- tory Detects overlap or insufficient clearance to stationary vehicles. NODE_FOLLOWING_DISTANCERelative speed + spacing Estimates whether following distance is unreasonable for visible conditions. NODE_LANE_CHANGETemporal trajectory sequenceDetermines whether the ego commits to a lane change maneuver. NODE_NAVIGATION_MATCHNavigation command + lane structure Checks whether the planner prepares correctly for a required turn or route branch. NODE_SPEED_REASONABLENESSEgo speed + scene contextDetermines whether the speed is unsafe for visible environmental conditions. NODE_WORKZONECone/barrier segmentationDetects temporary work-zone narrowing and clearance risk. NODE_TEMP_CONTROLFlagger/officer/cone cuesDetects temporary traffic-control instructions and diversion paths. N1. Completed lane change visible? N2. Better explained as vehicle right-of-way taking? N3. Better explained as aggressive weaving? N4. Inadequate safety margin to adjacent traffic? N5. Forced reaction, near miss, or collision? N6. Legitimate merge, route, obstacle, or emergency reason? N7. Better explained as noncommittal merge? N8. Lane change lacks traffic or route purpose? Confirmed unsafe_lane_change Routed failure_to_yield_vehicle Routed aggressive_weaving Routed noncommittal_merge Routed meaningless_lane_change Not applicable Not applicable Unproven Unproven yes no abstain no yes abstain no yes abstain yes no abstain yes/no/abstain yes no abstain yes no/abstain yes no abstain Figure 15: Full authored decision graph forunsafe_lane_change. The graph first proves a completed lane change, routes cases better explained by vehicle right-of-way taking or aggressive weaving to narrower sibling threats, confirms unsafe lane change when inadequate margin is present, and suppresses cases explained by legitimate route, merge, obstacle, or emergency needs. 19 GroupThreat labelDefinition Signal/stop controlRed light turning no full stopVehicle makes a turn on red without first coming to a full stop where the red-turn movement is allowed only after stopping and yielding. Signal/stop controlRed light violationVehicle enters or proceeds through a steady red signal or makes a prohibited movement on red. Signal/stop controlStop sign no full stopVehicle fails to come to a full stop where a stop is required. Signal/stop controlYellow phase noncomplianceVehicle enters on a steady yellow when a normal safe stop before the controlled entry was reasonably available. Right-of-wayFailure to yield pedestrianVehicle fails to yield to a pedestrian in a marked or unmarked cross- walk or equivalent pedestrian priority zone. Right-of-wayFailure to yield vehicle Vehicle enters, merges, or turns across another vehicleâs or bicyclistâs right-of-way unsafely. CollisionCollision with static objectsPredicted ego travel remains on a path that would strike, scrape, or is clearly driving into a plausible future collision course with, a static object; judge the ego swept footprint/corridor, not only the centerline overlay, and visible overlap within the current 5s path is not required when ego is aimed toward the object and GT is not. CollisionCollision with stationary vehiclePredicted ego travel remains on a path that would strike a parked or otherwise stationary vehicle. CollisionVehicle collision coursePredicted ego travel remains on a path that would strike another vehicle unless one of the vehicles changes course, outside narrower right-of- way, tailgating, or lane-change explanations. Following distanceFollowing too closely Vehicle follows another vehicle more closely than is reasonable and prudent for the visible speed and traffic context. Lane useLane straddling Vehicle drives for a sustained period between two lanes instead of committing to one lane. Lane useUnsafe lane changeVehicle changes lanes without reasonable safety. Lane useUnsafe passingVehicle passes in a prohibited or unsafe manner. Lane useWrong side of roadVehicle travels on the oncoming or left side of the roadway outside a clearly lawful exception. Lane useWrong way one way roadVehicle travels opposite the designated direction on a clearly one-way roadway. Roadway positionOff road driving Vehicle leaves the ordinary drivable lane or roadway and drives on a sidewalk, curb area, shoulder, dirt, gore, median edge, parking apron, private-property drive/frontage area, landscaped area, or another non- roadway area without a clearly lawful or necessary explanation. Special-lane misuseDriving in bike laneVehicle uses a bicycle lane as a normal travel lane beyond narrow lawful exceptions. Special-lane misuseDriving in bus only laneVehicle uses a lane reserved for public transit buses outside allowed conditions. Special-lane misuseImproper shoulder pass or bypassVehicle uses the shoulder, emergency lane, or other off-main-traveled roadway edge as a pass, bypass, or ordinary travel path outside clear lawful exceptions. SpeedSpeeding posted limitVehicle clearly exceeds a posted speed limit when the limit and excess are visually well supported. SpeedUnsafe speed for conditionsVehicle is traveling faster than is reasonable for the visible scene conditions. Reckless patternSpeed contest or exhibition Vehicle behavior suggests racing, competitive acceleration, or exhibi- tion of speed. PatternAggressive weavingVehicle repeatedly changes lanes or swerves laterally in a rapid, conflict-seeking, or intimidation-like manner that materially raises interaction risk. Route complianceLate navigation lane preparationVehicle stays out of the lane position needed for a near-term instructed turn, exit, or branch until the last reasonable preparation chance is closing or already missed, making future route failure or a forced risky late lane change likely. Route complianceNot following navigation instruction Vehicle materially deviates from a clear near-horizon navigation in- struction by committing to the wrong turn, branch, or lane stack without a visible reason. Temporary controlNot following temporary traffic instruc- tions Vehicle materially disregards a clear temporary traffic-control instruc- tion from an officer, flagger, construction guide, cone-directed diver- sion, or similar active traffic-direction setup. Policy qualityBlocking traffic without necessityVehicle stops or crawls in a way that unnecessarily obstructs a lane, crosswalk, or intersection without a legitimate traffic-control reason. Policy qualityLow efficiency drivingVehicle remains static or crawls at an obviously inefficient pace with- out a visible traffic-control, safety, conflict, or route-based reason. Policy qualityMeaningless lane changeVehicle changes lanes without a reasonable traffic-related purpose or weaves left-right-left without improving progress. Policy qualityMeaningless lateral driftCompared with GT, the predicted vehicle adds or retains material lateral deviation without a clear purpose, while not fully becoming a lane change, lane straddling, or roadway departure. Policy qualityNoncommittal mergeVehicle begins a merge or turn commitment and then hesitates or aborts in a way that reflects noncommittal go-no-go behavior and may disrupt surrounding traffic. Policy qualityStops without escape roomVehicle stops too close to an obstacle or blocked area and leaves insufficient space for a normal forward go-around if the obstacle remains stationary. Table 7: Compact definitions of the semantic threat labels used by FluidTest. 20