Paper deep dive
What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning
Hyeonchang Jeon, Kyungbeom Kim, Eugene Vinitsky, Kyung-Joong Kim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 99%
Last extracted: 7/5/2026, 5:40:28 AM
Summary
The paper investigates the internal representations of autonomous driving policies (Behavior Cloning and Reinforcement Learning) using linear probing and causal interventions. It explores whether improvements in closed-loop performance (scaling) actually reflect better internal prediction of surrounding vehicles and adaptive ego-planning, or merely better behavioral heuristics. The study finds that while scaling improves performance and the ability to ignore irrelevant vehicles, models often fail to form timely predictions during near-collision events. Crucially, the authors demonstrate through causal intervention that correcting internal mispredictions of surrounding vehicles can successfully steer the ego vehicle toward safer trajectories.
Entities (8)
Relation Signals (5)
Linear Probing â assesses â Ego Planning
confidence 100% ¡ We use linear probing... to track when these internal signals emerge, plateau, or fail.
Behavior Cloning (BC) â trainedon â Waymo Open Motion Dataset (WOMD)
confidence 100% ¡ For the BC and RL models, we train our model using the Waymo Open Motion Dataset (WOMD)
Reinforcement Learning (RL) â trainedon â Waymo Open Motion Dataset (WOMD)
confidence 100% ¡ For the BC and RL models, we train our model using the Waymo Open Motion Dataset (WOMD)
SMART â trainedon â Waymo Open Motion Dataset (WOMD)
confidence 100% ¡ We train the IL model using the WOMD dataset
Surrounding-Vehicle Prediction â influences â Ego Planning
confidence 90% ¡ causal intervention shows that correcting mistaken predictions improves ego planning toward safer trajectories.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics. Summary scores from closed-loop simulators do not give significant insight into the policy, making it difficult to determine whether they truly predict the motion of surrounding vehicles, how the ego vehicle generates future plans, or whether they merely rely on brittle heuristics that happen to succeed in nominal scenarios. To better understand the limits and weaknesses of driving policies, we focus on probing for forms of prediction, i.e., where surrounding vehicles will move next, and planning, i.e., understanding how to generate safe trajectories. We focus on these two capabilities because they reflect behaviors expected of effective driving policies, and use their presence or absence to assess policy quality across data-driven behavior cloning and simulation-driven reinforcement learning policies. To evaluate the presence of these capabilities, we investigate them as a function of scale, asking whether the closed-loop gains from larger datasets and longer simulation training reflect stronger prediction and planning or merely better behavioral heuristics. We use linear probing and targeted perturbations in both imitation learning and reinforcement learning models to track when these internal signals emerge, plateau, or fail. Despite good closed-loop performance, policies often fail to form timely surrounding-vehicle predictions during near-collision events, revealing a limitation in the predictive signals available for ego planning. Finally, causal intervention shows that correcting mistaken predictions improves ego planning toward safer trajectories.
Tags
Links
- Source: https://arxiv.org/abs/2606.31106v1
- Canonical: https://arxiv.org/abs/2606.31106v1
Trouble viewing inline? Open PDF directly â
Full Text
114,179 characters extracted from source content.
Expand or collapse full text
What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning Hyeonchang Jeon 1 Kyungbeom Kim 1 Eugene Vinitsky 2,⥠Kyung-Joong Kim 1,⥠1 Gwangju Institute of Science and Technology (GIST) 2 New York University kevinjeon119, kyungbeom8@gm.gist.ac.kr vinitsky.eugene@nyu.edu kjkim@gist.ac.kr Abstract Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics. Summary scores from closed-loop simulators do not give significant insight into the policy, making it difficult to determine whether they truly predict the motion of surrounding vehicles, how the ego vehicle generates future plans, or whether they merely rely on brittle heuristics that happen to succeed in nominal scenarios. To better understand the limits and weaknesses of driving policies, we focus on probing for forms of prediction, i.e., where surrounding vehicles will move next, and planning, i.e., understanding how to generate safe trajectories. We focus on these two capabilities because they reflect behaviors expected of effective driving policies, and use their presence or absence to assess policy quality across data-driven behavior cloning and simulation-driven reinforcement learning policies. To evaluate the presence of these capabilities, we investigate them as a function of scale, asking whether the closed-loop gains from larger datasets and longer simulation training reflect stronger prediction and planning or merely better behavioral heuristics. We use linear probing and targeted perturbations in both imitation learning and reinforcement learning models to track when these internal signals emerge, plateau, or fail. Despite good closed-loop performance, policies often fail to form timely surrounding-vehicle predictions during near-collision events, revealing a limitation in the predictive signals available for ego planning. Finally, causal intervention shows that correcting mistaken predictions improves ego planning toward safer trajectories. 1 Introduction The combination of new, large datasets [Ettinger et al., 2021, Caesar et al., 2020, Chang et al., 2019, Wilson et al., 2021] and GPU-accelerated simulation environments [Kazemkhani et al., 2025, Gulino et al., 2023] has led to rapid advances in data-driven autonomous driving [Shi et al., 2022, Nayakanti et al., 2023, Sima et al., 2024, Philion et al., 2024, Feng et al., 2023, Hu et al., 2022]. Imitation learning and reinforcement learning are core components of many of these advances, being used both for motion prediction, enabling the self-driving car to understand where objects in the scene will move, and for generating suggested trajectories or actions for the self-driving car itself. However, the success of driving policies obscures several potential challenges. Although the available training data is relatively abundant, it remains far smaller than in areas such as NLP and is heavily biased toward nominal, uninteresting scenarios in which drivers simply drive straight [Zhai et al., 2023, Li et al., 2024b, Jeon et al., 2024, Makansi et al., 2021]. Furthermore, scenarios that truly test ⥠Equal advising. Preprint. arXiv:2606.31106v1 [cs.RO] 30 Jun 2026 Figure 1: Recovery of planning with correction of prediction. The ego AV continues straight, while the other vehicle drives in parallel on the ego AVâs left side. Gradient arrows simplify the probing results and show the decoded future motion direction over time. (Original): After the 10-step prediction, the decoded ego plan curves away from its straight ground-truth trajectory, while the other vehicleâs future is also mispredicted. (Changed): Correcting the surrounding-vehicle prediction shifts the decoded ego plan toward the correct straight trajectory. the generalization of driving models are rare, leaving open whether these models learn generalizable rules or instead rely on spurious correlations and shortcut heuristics. These limitations and evaluation gaps raise a question: Are driving agents truly learning the core prediction and planning skills required for safe and robust driving [Makansi et al., 2021, Sun et al., 2024]? There are several natural indicators we could use to assess understanding of driving policies; con- versely, their absence would suggest a failure to internalize key skills. First, a skilled policy should respond sensibly to perturbations in simulation: removing a lead vehicle should allow the ego to accelerate, introducing one should induce deceleration, and reduced interaction should generally make driving easier, yielding increased driving progress and fewer collisions. Second, safe driving requires reasoning about surrounding vehicles, so a good policy would likely attend to a relevant subset rather than treating all vehicles equally or ignoring them. For example, it should respond to a slowing lead vehicle while discounting distant vehicles that do not affect the ego plan. Third, robust safety requires planning for contingencies, which likely entails maintaining internal predictions of safety-critical trajectories. For example, a truly skilled policy should be able to respond to abrupt cut-ins and maintain potential plans to handle unexpected scenarios. From these requirementsâreasoning about surrounding vehicles and maintaining an internal planâwe narrow our focus to two concrete, probeable competencies. We define prediction as inferring the future positions of surrounding vehicles, and planning as the adaptiveness of the ego autonomous vehicle (AV)âs planning in response to changes in predictions during decision-time planning [Bush et al., 2025] as shown in Figure 1. However, a driving model can perform well on driving metrics while failing to internalize these capabilities. Indeed, there is evidence that one can perform well on standard benchmarks while entirely neglecting any information but ego state [Li et al., 2024b] or achieve high-scoring performance while having incorrect causal models [Sun et al., 2024]. Even if these failure modes are avoided, simple rules like âidentify a lead vehicle and follow it at a safe distance" can be effective on most scenes while failing critically on the rare scenarios that make the rule unsafe. In this paper, we investigate the internals and behavior of driving models ranging from simple behavior cloning (BC) to state-of-the-art imitation learning (IL) and reinforcement learning (RL) policies, with the goal of understanding what aspects of safe driving they learn to represent. In particular, we study how surrounding-vehicle information is represented and used, and how prediction and adaptive planning emerge as data scale increases. We train driving models at multiple dataset sizes [Zheng et al., 2024, Naumann et al., 2025, Baniodeh et al., 2025] and evaluate each scale using perturbed closed-loop simulation, probing, and targeted interventions. 2 To test dependency on the presence of often irrelevant surrounding vehicles, we run perturbed simulations in which we randomly remove them. To quantify what the model internally knows about nearby agents, we use surrounding-vehicle linear probing [Alain and Bengio, 2016] to decode future positions of nearby agents from internal representations, and benchmark the quality of the internal representation against probes trained directly on raw inputs. We further analyze simulated near-collision events to examine how probing performance relates to actual collision outcomes. Lastly, to study how surrounding-vehicle information influences planning, we use ego-linear probing to decode the ego AVâs planned future positions and perform a decision-time intervention experiment that modifies internal representations through these probes, testing whether correcting surrounding- vehicle mispredictions leads to safer and more appropriate ego plans [Bush et al., 2025]. From our results, can we say whether âbetter internal predictions translates to better plans?â Often, yes: when provided with correct surrounding-vehicle predictions, the driving models reliably reroute or adjust their speed to produce safe and reasonable plans, and interventions that substitute correct predictions for incorrect ones steer trajectories back on course and reduce collisions in the internal representations. Our contributions are as follows: â˘Prediction/planning probes across scales: We develop linear probes to show that, as training data scale increases, models learn stronger planning representations and better ignore irrelevant vehicles, yet still struggle to identify truly safety-critical ones. â˘Perturbed closed-loop evaluation: We introduce a surrounding-vehicle removal protocol and find that driving models perform worse even when driving should become easier in single AV driving, suggesting that they still partially rely on surrounding vehicles in a spurious way. â˘Near-collision analysis and causal intervention: We use the probing to show that models attend to critical surrounding vehicles too late in near-collision events, and use interven- tions to show that ego planning can adapt with changed predictions and improve when mispredictions are corrected. 2 Related Work 2.1 Planning in Autonomous Driving Autonomous driving research has access to abundant datasets of human behavior [Caesar et al., 2020, Chang et al., 2019, Ettinger et al., 2021]. However, these datasets are heavily skewed toward nominal, low-interaction scenarios, leaving open the question of whether standard closed-loop simulationâtypically evaluated on the same distributionâcan truly measure robust behavior in safety-critical interactions. Standard planning models [Nayakanti et al., 2023, Chai et al., 2020, Seff et al., 2023] are widely used for both prediction and planning, yet current evaluations rarely test whether they use surrounding-vehicle information in a timely and safety-relevant manner. Recent studies [Li et al., 2024b, Zhai et al., 2023] have shown that source datasets are dominated by nominal straight-driving scenarios, in which the ego AVâs state alone often suffices to achieve strong closed-loop performance. As a result, closed-loop metrics can be misleading, indicating high performance even when the model largely ignores other vehicles. This motivates the need for alternative evaluations that directly assess whether models encode and use surrounding-vehicle information for prediction and planning. In this paper, we study how prediction and planning abilities in IL models are reflected in their internal representations and analyze perturbed simulation results using linear probing. 2.2 Data Scaling Laws With the rise of foundation models, scaling laws have become an important research topic [Henighan et al., 2020, Tian et al., 2024, Bharadhwaj et al., 2024]. They study how performance varies with dataset size, model capacity, and compute resources, especially for transformer-based models. This perspective has recently been extended to autonomous driving [Zheng et al., 2024, Naumann et al., 2025, Baniodeh et al., 2025]. For example, [Zheng et al., 2024] introduced ONE-Drive and showed that augmenting rare scenarios is more effective than randomly sampling scenarios. [Naumann et al., 2025] studied end-to-end scaling through larger models and additional camera inputs. [Baniodeh 3 et al., 2025] found stronger alignment between open-loop and closed-loop metrics and identified compute-optimal model sizes. Similarly, we study how autonomous driving performance scales with dataset size and evaluate models with both metrics. Unlike prior work, however, we focus on how prediction and planning capabilities themselves scale with data. 2.3 Interpretability in Deep Learning Interpretable deep learning aims to transform neural representations into forms that can be analyzed and linked to human-understandable concepts. Linear probing is a standard method for measuring what a trained network has learned from data using a probing dataset and a linear classifier [Alain and Bengio, 2016, Mikolov et al., 2013]. It has been used to study planning abilities in both model-free RL [Kim et al., 2018] and supervised agents [Guez et al., 2019]. In autonomous driving, prior work applies linear probing to frozen policy encoders trained from expert actions, using simple heads to evaluate affordance prediction and interpretable control [Xiao et al., 2021]. More recent work [Tas and Wagner, 2025] explores probing-based methods with sparse autoencoders Bricken et al. [2023] to disentangle latent factors and interpret or control ego motion. Our work takes a similar probing-based perspective, but focuses on multi-agent settings, where the probes target not only ego planning but also predictions about surrounding vehicles. 3 Approaches In this section, we analyze how driving policies represent and use surrounding-vehicle information for decision making: (i) whether they can predict safety-relevant agents, (i) whether such predictions are timely in near-collision situations, and (i) how prediction representations causally influence planning via targeted interventions. First, we show that both simple behavior cloning and RL models follow a similar scaling law to prior work [Baniodeh et al., 2025, Naumann et al., 2025] (Section 3.2). This serves as a sanity check that the models we probe exhibit the expected closed-loop improvements with scale, allowing us to ask whether these performance gains are reflected in their internal representations for prediction and planning. Our probing results show that internal representations improve over raw inputs, but the gains are concentrated on simple behaviors. Distance-based analysis further shows that scaling helps models ignore distant vehicles, but not fully identify the most safety-relevant ones. (Section 3.3). To understand how the presence of surrounding vehicles affects the driving modelâs performance, we conduct a perturbation simulation in which vehicles are randomly removed from the scene. We further relate closed-loop behavior to internal representations by analyzing how surrounding-vehicle probing evolves during a near-collision event (Section 3.4). Lastly, we ask: what happens if we correct these mispredictions? We show that, once aligned, the driving model can recover its planning and even exhibit collision-avoidance behavior (Section 3.5). Summary of Results. Scaling experiments with open-loop metrics and closed-loop simulation results improve overall driving performance, but they do not, by themselves, show whether models learn safety-relevant prediction and planning. Our probes show that driving models encode better planning and suppress clearly irrelevant surrounding vehicles, yet they still fail to emphasize truly safety-critical vehicles early enough in near-collision situations. Finally, intervention results suggest that more accurate internal predictions can causally improve ego planning, steering decoded plans toward safer and more appropriate trajectories. 3.1 Experimental Setup Models. We compare a transformer-based behavior cloning policy similar to Gulino et al. [2023] (BC) and a PPO-based RL policy [Schulman et al., 2017, Cornelisse et al., 2025b] (RL) across data scales. We additionally apply our linear probing and causal intervention analyses to SMART [Wu et al., 2024] (IL), a multi-agent trajectory prediction model, demonstrating that our probing-based framework extends to state-of-the-art IL models. Detailed model descriptions are provided in Appendix A. Dataset. For the BC and RL models, we train our model using the Waymo Open Motion Dataset (WOMD) [Ettinger et al., 2021]. The dataset consists of over400Kdriving scenes, each containing 9 seconds of trajectory data sampled at 10 Hz, with up to 128 cars per scene. Due to limitations of 4 the underlying simulator used to replay the trajectories, we exclude scenes containing traffic lights and overpasses, filtering the dataset toâź 90Kscenes. Each observation is partially observable and includes information about the ego vehicle, surrounding vehicles (up to 127), and the map in an ego-centric view. As action labels are not available in the dataset, we derive them through inverse kinematics. For the IL model, we set the same setting in SMART [Wu et al., 2024]. We train using the WOMD dataset with 11 historical steps as input, including global coordinate, heading information, and predict the remaining future trajectories for the target. Since the IL model predicts multi-agent trajectories, it can predict up to 32 vehicles per scenario. More details on the dataset can be found in Appendix B. For a detailed analysis of closed-loop simulations, we first preprocess trajectories to remove those with infeasible kinematics (likely due to noise during data collection), and then classify the remaining trajectories into mutually exclusive categories (Straight, Turn, Reverse, and Uncategorized). (See details in Appendix B.3) Linear Probing. To interpret the driving modelâs internal representation, we perform a linear probing experiment. We select the best model across seeds for each dataset scale and train a linear classifier to predict future positions. The ground truth is set as the vehicleâs positions (ego or surrounding vehicle) 1 to 4 seconds ahead (corresponding to 10, 20, 30, and 40 steps). The positions are discretized into 64 labels by dividing the ego vehicleâs current field of view into an8Ă 8grid along the x- and y-axes, normalized with respect to the ego AVâs current position. To understand what information is retained in the internal network layers, we train linear classifiers on both the raw input and representations from the early and late attention layers. We evaluate performance using the F1 score to account for data imbalance, as well as per-trajectory-type accuracy to investigate how the driving models learn differently across cases. For the IL model, which predicts multi-agent trajectories, there is no predefined ego agent. We therefore sample one vehicle as the ego and select its nearest neighboring vehicle as the surrounding vehicle. Since our goal is to analyze how information about surrounding vehicles contributes to ego-trajectory prediction, this setup provides a natural egocentric perspective for evaluating the use of information from other agents. The linear probing setting is described in Appendix D.1 and Appendix E.1. 3.2 Data Scaling Laws 100500 1k5k 10k20k40k80k Number of Scenes 0.42 0.57 0.77 WOSAC Realism Random WOSAC Realism (â) = . . = . . BC = . RL = . 100500 1k5k 10k20k40k80k Number of Scenes 0.40 0.66 1.08 Goal Progress Ratio Goal Progress Ratio (â) = . . = . . BC = . RL = . 100500 1k5k 10k20k40k80k Number of Scenes 0.03 0.11 0.42 Off-Road Ratio Off-Road (â) = . â . = . â . BC = â . RL = â . 100500 1k5k 10k20k40k80k Number of Scenes 0.02 0.07 0.21 Collision Ratio Veh-Coll (â) = . â . = . â . BC = â . RL = â . BCRL Figure 2: Power law relationships of BC and RL models: We evaluate the WOSAC, collision metrics (Off-Road and Veh-Coll), and goal progress rate. r is the correlation coefficient. To investigate the emergence of planning and prediction ability, we train the BC and RL models with three different random seeds while gradually increasing the dataset size, and evaluate them in the GPUDrive simulator [Kazemkhani et al., 2025] on unseen scenarios. For RL, we train on up to 10K scenes, at which point performance has nearly converged, as shown in prior work [Cornelisse et al., 2025b]. To support our modelâs generality, we evaluate our model performance using the Waymo Open Sim Agents Challenge (WOSAC) metrics [Montali et al., 2023]. For closed-loop simulation, we use three metricsâ vehicle collision (denoted Veh-Coll), the off-road rate, and goal progress ratio which is calculated by1â d final /d initial whered final denotes the distance to the goal at the final timestep and d initial denotes the distance to the goal at the beginning of the trajectory. As shown in Figure 2, metrics decrease with scale, following a curved power law that approaches a plateau afterâ 20Kscenes for the BC model, which has shown similar trends with [Baniodeh et al., 2025]. In the case of the RL model, it achieves much lower overall collision rates (Veh-Coll and Off-Road) while surpassing the goal progress after 1K scenes, ultimately reaching 99% of the goal progress ratio. Both BC and RL models achieve high correlation, indicating that scaling laws hold in closed-loop simulation. See Appendix C.1 and Appendix C.3 for more results. 5 3.3 Linear Probing for Surrounding Vehicles Prediction 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes -0.05 0.00 0.05 0.10 F1 Macro (LP - Raw) F1 Macro Diff (Future Step: 10) 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes F1 Macro Diff (Future Step: 20) 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes F1 Macro Diff (Future Step: 30) 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes F1 Macro Diff (Future Step: 40) 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes 0.0 0.1 Accuracy (LP - Raw) Uncategorized Acc. Diff (Future Step=10) 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes Turn Acc. Diff (Future Step=10) 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes Straight Acc. Diff (Future Step=10) 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes Reverse Acc. Diff (Future Step=10) BC: LP - RawRL: LP - RawSMART (IL): LP - Raw Figure 3: Performance metrics with surrounding-vehicle linear probing. Top: F1 Macro score dif- ference between linear probes on trained model representations and the raw-input baseline (LPâRaw) across future stepsfsâ 10, 20, 30, 40, wherefsdenotes the prediction horizon in timesteps. Bottom: trajectory-type accuracy difference atfs = 10forStraight,Turn,Reverse, andUncategorized cases. Higher values are better for both metrics. TheReversecase of RL is omitted because no Reverse cases are available for the RL setting. To understand how the driving model utilizes surrounding-vehicle information internally, we train linear probes on the intermediate representations of our trained model to evaluate its prediction and planning capabilities, employing the linear probing method [Alain and Bengio, 2016]. We train the linear probes with three different random seeds while gradually increasing the dataset size. Since the models differ in architecture and input dimensionality (i.e., the RL model uses a single timestep, whereas the BC model uses five timesteps), directly comparing raw probing performance across models can be misleading. Therefore, for cross-model comparison, we report the relative probing gain, computed as the difference between the trained modelâs probe and the raw-input probe (LP âRaw). This measures how much additional information is encoded in the learned representation beyond what is easily available from the raw input. As shown in Figure 3, in the case of the BC model, the F1 Macro score difference is increasing until 10K scenes and becomes plateaued. This trend is similar to the closed-loop simulation results in Section 3.2, where surrounding-vehicle probing performance nearly peaks at 10K scenes, then slightly decreases before plateauing. Interestingly, the RL model initially maintains a higher F1 score than BC across all timesteps, but this advantage diminishes substantially by 10K scenes, though it still remains above BC. As shown in Appendix E.2, we find that the RL model gradually places greater emphasis on ego planning. Across timesteps, BC tends to show lower F1 scores as the prediction horizon increases, whereas RL largely maintains its F1 difference over time. Lastly, while SMART (IL) is comparable to 80K BC at the 10-step horizon, it shows a clearer advantage at longer horizons. In particular, its LPâRaw difference remains stable or slightly increases from 10 to 40 steps, whereas BC tends to decline. This suggests that SMART (IL) better preserves information about surrounding vehicles over longer temporal ranges. For type accuracy, BC continues to improve with scale, and this improvement is particularly pro- nounced for the relatively easierUncategorizedandStraightcases, while the gap remains much smaller forReverseandTurncases. In contrast, RL shows a largely similar pattern across scales, with little noticeable change. IL shows relatively stable type-accuracy differences; unlike its F1 Macro trend, it does not clearly outperform BC at the type level. This suggests that IL has an advantage in long-horizon planning, while BC and RL, which directly output actions, tend to focus more on shorter-horizon signals. More detailed results are provided in Appendix D. However, this nominal understanding does not directly assess predictive ability, since only a few vehicles are actually important among the many surrounding vehicles in each scenario. To address this, we further analyze whether the surrounding-vehicle linear probe is stronger for vehicles that will become closer to the ego in the future. In Figure 4, we plot the correlation between linear-probe prediction and future egoâother distance on both BC and RL models. The vertical axis shows the 6 Figure 4: Correlation between linear-probe and future egoâother distance. Vertical: probability gap (LP â raw); horizontal: future egoâother distance. Orange: regression line; r: correlation. difference in predicted probability for the true label between the LP and the raw-input probe, while the horizontal axis shows the future distance 10 steps ahead. Points are colored by relative distance change, currentâfuture current : samples that become more than 40% closer are shown in red, whereas those that become more than 40% farther are shown in blue, with opacity increasing with the magnitude of the change. At a small scale (100 scenes), the linear-probe advantage over the raw-input probe is only weakly associated with future egoâother distance, indicating limited sensitivity to future interaction relevance. With more training data (10,000 and 80,000 scenes), the association becomes markedly more negative, indicating that the representation increasingly favors vehicles that will remain close or move closer to the target. However, among nearby vehicles, the model still attends to both approaching and receding vehicles. This suggests that scale helps the model filter out clearly irrelevant vehicles, but it does not yet fully isolate or predict the truly critical vehicles, i.e., those that will move even closer to the target. 3.4 Randomly Removing Surrounding Vehicles 0.00.20.40.60.81.0 Other Vehicles Removed Ratio 0.02 0.27 0.52 BC Value Off-Road (â) 0.00.20.40.60.8 Other Vehicles Removed Ratio 0.00 0.14 0.29 BC Value Veh-Coll(Adjusted) (â) 0.00.20.40.60.81.0 Other Vehicles Removed Ratio 0.28 0.61 0.94 BC Value Goal Progress Ratio (â) 0.00.20.40.60.81.0 Other Vehicles Removed Ratio 0.01 0.15 0.29 RL Value 0.00.20.40.60.8 Other Vehicles Removed Ratio 0.00 0.10 0.20 RL Value 0.00.20.40.60.81.0 Other Vehicles Removed Ratio 0.02 0.53 1.04 RL Value 0.1K0.5K1K5K10K20K40K80K Figure 5: Perturbed simulation results: We test the IL model by randomly removing the surrounding vehicles with a ratio p. We adjust the ratio by multiplying 1 ratio for vehicle collision. If the driving model does not spuriously depend on the presence of surrounding vehicles, removing them should not degrade performance and may even improve goal-reaching performance by reducing interaction constraints. Moreover, stronger prediction ability should lead to more robust planning 7 across varying numbers of surrounding vehicles. To evaluate this, we randomly remove a fraction of active vehicles from each of 10,000 validation scenes, as shown in Figure 5. In both models, performance drops in the single-AV setting, with increased off-road rates and worse goal-progress trends. This is undesirable: removing surrounding vehicles should make the driving easier, not harder. The degradation, therefore, suggests that the policies partially rely on the presence of the other vehicles in a spurious way, rather than using them as an interaction. This likely reflects WOMDâs bias toward multi-vehicle trajectories. The drop decreases as data scale increases, suggesting that larger datasets partially mitigate this issue. See additional results in Appendix F.1. Figure 6: Near-collision Analysis. Each heatmap shows the normalized difference between the predicted probability of surrounding-vehicle probing in the pre-collision windowwand that of each spatial grid over the full episode, computed as wâa grid a , at 10 to 40 steps before collision. The heatmaps use an ego-centric8Ă 8grid covering[â50, 50]meter along both axes, so each grid cell corresponds to a12.5 mĂ 12.5 mspatial region. Red grids mark areas that become increasingly important to the model near collision, whereas blue grids mark areas that become relatively less important compared with the modelâs average attention over the full episode. The green dot marks the ego vehicle position. Dark gray cells indicate grids that were filtered out due to insufficient collision samples. Top: BC, Bottom: RL Beyond the single AV setting, the perturbed simulation offers only a rough, indirect view of how the model handles surrounding vehicles. To more explicitly examine whether the model predicts the behavior of surrounding vehicles in safety-critical situations, we therefore conduct a near-collision analysis using the surrounding-vehicle linear probe. We define a pre-collision windowWof 10 steps and compare the probing signal within this window against the modelâs average probing signal over the full validation episodes. For each spatial grid cellgrid, we compute the normalized difference (wâ a g rid)/a g rid, wherea g ridis the corresponding full-episode average. We filter out grid cells with fewer than 100 collision cases and mark them in dark gray. Figure 6 visualizes this normalized difference at 40, 30, 20, and 10 steps before collision. Red grids denote spatial regions whose surrounding-vehicle probing signal increases relative to the whole episode average, whereas blue grids denote regions whose signal decreases. Therefore, early anticipation of a dangerous interaction would appear as a strong positive signal around the relevant nearby vehicle at earlier horizons, such as 40, 30, or 20 steps before the event. In contrast, for the BC model, the signal remains weak and diffuse at earlier horizons and becomes clearly concentrated near the ego only in the final 10 steps before collision. Although this late increase indicates that the model eventually represents the nearby colliding vehicle, it is already too late for effective avoidance; therefore, these episodes remain planning failures. The RL model shows a similar qualitative pattern, but collision cases are much rarer, resulting in substantially fewer samples and larger dark-gray filtered regions. Despite this sparsity, the strongest positive signal again appears mainly in the final 10-step window. Overall, both BC and RL appear to 8 use surrounding-vehicle information near a collision, but they fail to emphasize the relevant vehicle early enough to avoid the crash. We analyze near-collision events for off-road in Appendix F.2. 3.5 Intervention Test for Adaptiveness of Ego AV Planning What happens to the modelâs planning when it has more accurate predictions, and how does it adapt to changes in the predictions of surrounding vehicles? To test this, we first applied linear probing to the ego AVâs planning, as in the Section 3.3, which we refer to as ego-linear probing. The result of ego-linear probing is in Appendix E. Inspired by Bush et al. [2025], we intervene on internal representations to test whether ego planning causally adapts to perturbed surrounding-vehicle predictions and recovers when incorrect predictions are corrected. We modify the earlier layer representation of a selected surrounding vehicleoand examine how this perturbation propagates to the egoâs later layer probing. Let the earlier layer feature beg = [g o ; g âo ],g o â R 128 , the feature of surrounding vehicleo, andg âo â R 127Ă128 those of the remaining vehicles including ego vehicle. We obtain the surrounding-linear probing weight w l â R 128 corresponding to the surrounding vehiclesâ future position labellfrom a linear probe trained on earlier layer features of surrounding vehicles at future timesteps â Swhere|S|is the number of future timesteps. Our intervention adds this label direction to g o : g Ⲡo = g o + 1 |S| Îą X s w s l , g Ⲡ= [g Ⲡo ; g âo ].(1) If the model has adaptive planning capability, the later layer activationh Ⲡ= f (g Ⲡ)should change coherently from the baselineh = f (g)along a semantically meaningful directionw l , scaled by a strength parameterÎą. This encourages a change in the ego-linear probing output, fromptop Ⲡ, as predicted by the linear probe layer e. h Ⲡ= f (g Ⲡ), p Ⲡ= e(h Ⲡ).(2) Figure 7: Intervention experiment for adaptiveness and recovery (10 to 40 future timesteps): Examples from BC, RL, and IL models, and each pair of columns compares the original probing result with the probing result after intervention. (a) Adaptiveness: we perturb the surrounding-vehicle prediction so that it overlaps with the egoâs predicted path, and test whether the ego plan changes to avoid the induced conflict. (b) Recovery: we replace an incorrect surrounding-vehicle probe with the correct intervention label, and test whether the ego plan is restored toward a safer or more ground-truth-aligned trajectory. we conduct intervention experiments in two settings: adaptiveness and recovery. Adaptiveness measures whether the ego plan changes in response to a potential collision. To test this, we perturb 9 the surrounding-vehicle probing representation to overlap with the ego representation and observe the resulting change in ego planning. Recovery instead examines whether correcting an incorrect prediction of the surrounding vehicle restores a more appropriate ego plan. For each model, we label 100 validation scenes and remove irrelevant cases such as short trajectories or single-AV scenes. This leaves 59 valid cases for BC (43 adaptiveness, 16 recovery), 58 for RL (53 adaptiveness, 5 recovery), and 64 for IL (51 adaptiveness, 13 recovery). We summarize the overall intervention results across all models in Appendix G. Overall, both models frequently adapt their plans when predictions about surrounding vehicles are updated, and recovery interventions often restore more appropriate planning. The results also reveal model-specific tendencies: BC is more reliable for route-change and slowdown interventions but struggles with speed-up cases, whereas RL shows a more balanced response across intervention types, and the IL shows good at adjusting speed both slower and faster. (See Table 11 in Appendix G for all results) Figure 7 shows that the egoâs planning can shift from an initially unsafe trajectory to a safer one that avoids collisions with surrounding vehicles. Moreover, in failure cases, restoring the incorrect predic- tions of surrounding vehicles led the models to reorient their planning toward the goal, indicating that accurate predictions help generate correct plans. This effect appears across all three models: across both intervention types, the ego-planning probe changes in 36 of 59 BC cases, 37 of 58 RL cases, and 40 of 64 IL cases, suggesting that surrounding-vehicle predictions causally affect ego planning. However, the ego plan sometimes remains roughly aligned with the ground-truth trajectory even with imperfect predictions of surrounding vehicles, suggesting that precise predictions are not always necessary for effective planning. (see Appendix G for more example cases). 4 Conclusion In this paper, we investigate how driving models preserve and use surrounding-vehicle information for ego planning. As the dataset size increases, the models become better at ignoring irrelevant agents. Furthermore, our perturbed closed-loop evaluation reveals that the reliance on surrounding vehicles is not always robust: removing all surrounding vehicles degrades the performance, even though the task should become easier in their absence. This suggests that the policies may partially rely on surrounding vehicles in a spurious way. In particular, during near-collision events, it should anticipate othersâ positions earlier, yet it often fails to do so and remains biased toward ânominalâ vehicles, such as those going straight. Finally, we find that when prediction capability is restored, the planner produces stronger trajectories, suggesting that better prediction can lead to better planning. Moreover, the models not only recover but also actively avoid encroaching vehicles, indicating that their planning is already sufficiently robust to support collision avoidance. These findings suggest that better planning may require not only stronger prediction but also more precise approaches for measuring the surrounding-vehicle information that the models capture. Currently, our probing approach utilizes discretized predictions, which can be coarse and potentially blur fine-grained behavior. Designing probes and visualizations in continuous space is a promising direction. The correlation between surrounding-vehicle probing and future distance suggests that both BC and RL models can suppress some irrelevant agents. However, it remains unclear whether it reliably identifies the agents most critical for safe planning. To address this, incorporating joint future prediction [Luo et al., 2023] or auxiliary tasks [Li et al., 2024a] to improve the model architecture, as well as providing additional language input signals from datasets [Malla et al., 2023, Li et al., 2025, Chang et al., 2025], may help the model develop a richer understanding of surrounding vehicles. References Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644, 2016. Mustafa Baniodeh, Kratarth Goel, Scott Ettinger, Carlos Fuertes, Ari Seff, Tim Shen, Cole Gulino, Chenjie Yang, Ghassen Jerfel, Dokook Choe, et al. Scaling laws of motion forecasting and planningâa technical report. arXiv preprint arXiv:2506.08228, 2025. Homanga Bharadhwaj, Jay Vakil, Mohit Sharma, Abhinav Gupta, Shubham Tulsiani, and Vikash Ku- mar. Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations 10 and action chunking. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 4788â4795. IEEE, 2024. Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, et al. Towards monosemanticity: Decompos- ing language models with dictionary learning. Transformer Circuits Thread, 2, 2023. Thomas Bush, Stephen Chung, Usman Anwar, AdriĂ Garriga-Alonso, and David Krueger. Inter- preting emergent planning in model-free reinforcement learning. In The Thirteenth International Conference on Learning Representations, 2025. Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11621â11631, 2020. Yuning Chai, Benjamin Sapp, Mayank Bansal, and Dragomir Anguelov. Multipath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction. In Conference on Robot Learning, pages 86â99. PMLR, 2020. Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, et al. Argoverse: 3d tracking and forecasting with rich maps. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8748â8757, 2019. Wei-Jer Chang, Wei Zhan, Masayoshi Tomizuka, Manmohan Chandraker, and Francesco Pittaluga. Langtraj: Diffusion model and dataset for language-conditioned trajectory simulation. arXiv preprint arXiv:2504.11521, 2025. Daphne Cornelisse, Spencer Cheng, Pragnay Mandavilli, Julian Hunt, Kevin Joseph, WaĂŤl Doulazmi, Valentin Charraut, Aditya Gupta, Joseph Suarez, and Eugene Vinitsky. PufferDrive: A fast and friendly driving simulator for training and evaluating RL agents.https://github.com/ Emerge-Lab/PufferDrive, 2025a. Version 2.0.0. Equal contribution by the first two authors. Daphne Cornelisse, Aarav Pandya, Kevin Joseph, Joseph SuĂĄrez, and Eugene Vinitsky. Building reliable sim driving agents by scaling self-play. arXiv preprint arXiv:2502.14706, 2025b. Scott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu, Hang Zhao, Sabeek Pradhan, Yuning Chai, Ben Sapp, Charles R Qi, Yin Zhou, et al. Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9710â9719, 2021. Lan Feng, Quanyi Li, Zhenghao Peng, Shuhan Tan, and Bolei Zhou. Trafficgen: Learning to generate diverse and realistic traffic scenarios. In 2023 IEEE international conference on robotics and automation (ICRA), pages 3567â3575. IEEE, 2023. Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, SĂŠbastien Racanière, ThĂŠophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, et al. An investigation of model-free planning. In International conference on machine learning, pages 2464â2473. PMLR, 2019. Cole Gulino, Justin Fu, Wenjie Luo, George Tucker, Eli Bronstein, Yiren Lu, Jean Harb, Xinlei Pan, Yan Wang, Xiangyu Chen, et al. Waymax: An accelerated, data-driven simulator for large-scale autonomous driving research. Advances in Neural Information Processing Systems, 36:7730â7742, 2023. Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al. Scaling laws for autoregressive generative modeling. arXiv preprint arXiv:2010.14701, 2020. Anthony Hu, Gianluca Corrado, Nicolas Griffiths, Zachary Murez, Corina Gurau, Hudson Yeo, Alex Kendall, Roberto Cipolla, and Jamie Shotton. Model-based imitation learning for urban driving. Advances in Neural Information Processing Systems, 35:20703â20716, 2022. 11 Hyeongseok Jeon, Sanmin Kim, Abi Rahman Syamil, Junsoo Kim, and Dongsuk Kum. Beyond the data imbalance: Employing the heterogeneous datasets for vehicle maneuver prediction. In European Conference on Computer Vision, pages 38â53. Springer, 2024. Saman Kazemkhani, Aarav Pandya, Daphne Cornelisse, Brennan Shacklett, and Eugene Vinitsky. GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS. In The Thirteenth International Conference on Learning Representations, 2025. URLhttps://openreview.net/ forum?id=ERv8ptegFi. Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning, pages 2668â2677. PMLR, 2018. Jiachen Li, David Isele, Kanghoon Lee, Jinkyoo Park, Kikuo Fujimura, and Mykel J Kochenderfer. Interactive autonomous navigation with internal state inference and interactivity estimation. IEEE Transactions on Robotics, 40:2932â2949, 2024a. Yiheng Li, Cunxin Fan, Seth Z Zhao, Chenran Li, Chenfeng Xu, Huaxiu Yao, Masayoshi Tomizuka, Bolei Zhou, Chen Tang, Mingyu Ding, et al. Womd-reasoning: A large-scale dataset for interaction reasoning in driving. In Forty-second International Conference on Machine Learning, 2025. Zhiqi Li, Zhiding Yu, Shiyi Lan, Jiahan Li, Jan Kautz, Tong Lu, and Jose M Alvarez. Is ego status all you need for open-loop end-to-end autonomous driving? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14864â14873, 2024b. Wenjie Luo, Cheol Park, Andre Cornman, Benjamin Sapp, and Dragomir Anguelov. Jfp: Joint future prediction with interactive multi-agent modeling for autonomous driving. In Conference on Robot Learning, pages 1457â1467. PMLR, 2023. Osama Makansi, ĂzgĂźn Ăiçek, Yassine Marrakchi, and Thomas Brox. On exposing the challenging long tail in future prediction of traffic actors. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13147â13157, 2021. Srikanth Malla, Chiho Choi, Isht Dwivedi, Joon Hee Choi, and Jiachen Li. Drama: Joint risk localization and captioning in driving. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 1043â1052, 2023. TomĂĄĹĄ Mikolov, Wen-tau Yih, and Geoffrey Zweig. Linguistic regularities in continuous space word representations. In Proceedings of the 2013 conference of the north american chapter of the association for computational linguistics: Human language technologies, pages 746â751, 2013. Nico Montali, John Lambert, Paul Mougin, Alex Kuefler, Nicholas Rhinehart, Michelle Li, Cole Gulino, Tristan Emrich, Zoey Yang, Shimon Whiteson, et al. The waymo open sim agents challenge. Advances in Neural Information Processing Systems, 36:59151â59171, 2023. Alexander Naumann, Xunjiang Gu, Tolga Dimlioglu, Mariusz Bojarski, Alperen Degirmenci, Alexan- der Popov, Devansh Bisla, Marco Pavone, Urs Muller, and Boris Ivanovic. Data scaling laws for end-to-end autonomous driving. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 2571â2582, 2025. Nigamaa Nayakanti, Rami Al-Rfou, Aurick Zhou, Kratarth Goel, Khaled S Refaat, and Benjamin Sapp. Wayformer: Motion forecasting via simple & efficient attention networks. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 2980â2987. IEEE, 2023. Jonah Philion, Xue Bin Peng, and Sanja Fidler. Trajeglish: Traffic modeling as next-token prediction. In The Twelfth International Conference on Learning Representations, 2024. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017. Ari Seff, Brian Cera, Dian Chen, Mason Ng, Aurick Zhou, Nigamaa Nayakanti, Khaled S Refaat, Rami Al-Rfou, and Benjamin Sapp. Motionlm: Multi-agent motion forecasting as language modeling. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8579â8590, 2023. 12 Shaoshuai Shi, Li Jiang, Dengxin Dai, and Bernt Schiele. Motion transformer with global intention localization and local movement refinement. Advances in Neural Information Processing Systems, 35:6531â6543, 2022. Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Jens BeiĂwenger, Ping Luo, Andreas Geiger, and Hongyang Li. Drivelm: Driving with graph vi- sual question answering. In European conference on computer vision, pages 256â274. Springer, 2024. Liting Sun, Rebecca Roelofs, Ben Caine, Khaled S Refaat, Ben Sapp, Scott Ettinger, and Wei Chai. Causalagents: A robustness benchmark for motion forecasting. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6820â6827. IEEE, 2024. Omer Sahin Tas and Royden Wagner. Words in motion: Extracting interpretable control vectors for motion transformers. In International Conference on Learning Representations, volume 2025, pages 87191â87214, 2025. Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang. Visual autoregressive modeling: Scalable image generation via next-scale prediction. Advances in neural information processing systems, 37:84839â84865, 2024. Benjamin Wilson, William Qi, Tanmay Agarwal, John Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, et al. Argoverse 2: Next generation datasets for self-driving perception and forecasting. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. Wei Wu, Xiaoxin Feng, Ziyan Gao, and Yuheng Kan. Smart: Scalable multi-agent real-time motion generation via next-token prediction. Advances in Neural Information Processing Systems, 37: 114048â114071, 2024. Yi Xiao, Felipe Codevilla, Christopher Pal, and Antonio Lopez. Action-based representation learning for autonomous driving. In Conference on Robot Learning, pages 232â246. PMLR, 2021. Jiang-Tian Zhai, Ze Feng, Jinhao Du, Yongqiang Mao, Jiang-Jiang Liu, Zichang Tan, Yifu Zhang, Xiaoqing Ye, and Jingdong Wang. Rethinking the open-loop evaluation of end-to-end autonomous driving in nuscenes. arXiv preprint arXiv:2305.10430, 2023. Yupeng Zheng, Zhongpu Xia, Qichao Zhang, Teng Zhang, Ben Lu, Xiaochuang Huo, Chao Han, Yixian Li, Mengjie Yu, Bu Jin, et al. Preliminary investigation into data scaling laws for imitation learning-based end-to-end autonomous driving. arXiv preprint arXiv:2412.02689, 2024. 13 A Model and Training Details A.1 BC Details : Other vehicles : Ego AV : Road Early Fusion Attention Self-Attention(Obj)Self-Attention(Road) Cross-Attention(Obj) Cross-Attention(Road) Ego Embedding Other Embedding Road Embedding GMM ⢠⢠⢠⢠⢠⢠⢠⢠⢠⢠⢠⢠Figure 8: Behavior Cloning model architecture: Overall architecture of the behavior cloning model. Ego vehicle, other vehicles, and road features are first embedded and fused via early-fusion attention. The fused representations are refined via self- and cross-attention modules and finally modeled with a Gaussian Mixture Model (GMM). In our behavior cloning framework, the model is conditioned on three types of inputs: ego vehicle state, other vehiclesâ state, and road context features in Table 2. Each input modality is first embedded into a latent representation through dedicated encoders: h ego = f ego (x ego ),(3) h other = f other (x other ),(4) h road = f road (x road )(5) wheref ego ,f other ,f road denotes a scene encoder applied to the corresponding features, implemented as a multi-layer perceptron with layer normalization and non-linear activations: z 1 = Wx + b,(Linear)(6) z 2 = Dropout(z 1 ),(Dropout)(7) z 3 = LN(z 2 ),(Layer Normalization)(8) h = Ď(z 3 ),(Activation, e.g., Tanh)(9) The embedded representations are first fused through an early fusion attention module, which enables the ego state to jointly attend to both road and object-level features: H ego ,H other ,H road = Attn([h ego ,h other ,h road ]).(10) To refine the fused representation in a modality-specific manner, self-attention layers are applied independently to object-level and road-level embeddings: Ě H ego , Ě H other = SelfAttn([H ego ,H other ]),(11) Ě H road = SelfAttn(H road ).(12) To enhance modality-specific interaction and enable more informed policy learning, cross-attention layers are applied, allowing the ego representation to selectively attend to road-level and object-level 14 features: Ě H ego other = CrossAttn([ Ě H ego , Ě H other ]),(13) Ě H ego road = CrossAttn([ Ě H ego , Ě H road ]).(14) Finally, the aggregated representationâobtained by concatenating the refined ego embedding with the outputs of the road- and object-conditioned cross-attention modules( Ě H ego , Ě H ego other , Ě H ego road )is fed into an output head parameterizing a Gaussian Mixture Model (GMM): p(y | x) = K X k=1 Ď k N (y | Îź k , ÎŁ k ),(15) whereydenotes the predicted future action of the ego vehicle, andĎ k ,Îź k , ÎŁ k are the mixture weights, means, and covariances of thek-th Gaussian component, respectively. This probabilistic formulation enables the model to capture multimodal action distributions and inherent uncertainties in motion prediction. During training, the model parameters are optimized by minimizing the negative log- likelihood (NLL) of expert demonstrations under the predicted GMM distribution. Specifically, given an expert action y â , the likelihood under the mixture model is p(y â | x) = K X k=1 Ď k N (y â | Îź k , ÎŁ k ),(16) and the loss function is defined as L GMM =âE (x,y â )âźD [logp(y â | x)],(17) whereDdenotes the dataset of expert demonstrations. This objective encourages the model to assign high probability density to expert actions, thereby aligning the predicted distribution with the expert policy. A.2 BC Training Settings Across all experiments of dataset size, the BC model has 1.4M parameters and is trained with a batch size of 512, a hidden layer size of 128, and a 6-component GMM head. We train with a learning rate of0.0005and apply a weight decay. The different settings across different dataset sizes are in Table 1. We used 4 A100 GPUs to train BC models. Table 1: IL Training Details. (a) Small scales #Scenes#Vehicles#SamplesGrad. Steps 1000.7K44K20,000 5003K215K50,000 10007K431K100,000 500033K2.1M250,000 (b) Large scales #Scenes#Vehicles#SamplesGrad. Steps 1000068K4.2M400,000 40000270K16.9M700,000 80000542K33.8M850,000 A.3 RL Training Settings Across all experiments, we train a feed-forward IPPO policy [Schulman et al., 2017, Cornelisse et al., 2025b] with 1.2M parameters and a hidden dimension of 128. Training is performed with a batch size of 65,536 and a minibatch size of 4,096 for 2 update epochs, using a learning rate of 0.0003. We set the discount factor to 0.99, the GAE to 0.95, the clipping coefficient to 0.2, the entropy coefficient to 0.0001, and the value loss coefficient to 0.3. We scale the number of training steps with dataset size, using 0.1B, 0.2B, 0.3B, 0.75B, and 1B steps for 100, 500, 1,000, 5,000, and 10,000 scenes, respectively. We used 4 A100 GPUs to train RL models. 15 B Dataset Details B.1 Observation Features The observation space provides a multi-modal representation composed of four key features: ego state features, describing the ego vehicleâs kinematics and goal-related information; surrounding vehicle features, encoding surrounding dynamic surrounding vehicles; road graph features, capturing static road topology and structural elements (road edge, road line, road lane, crosswalk, speed bump, stop sign, and None as padding). For all experiments, we set the maximum number of agents per scenario to 128, and we consider the nearest 200 road points. The detailed specifications of these observation features, including their type, constituent variables, dimensionality, and description, are summarized in Table 2. Table 2: Observation features specifications. TypeFeatureDimensionDescription ego state Speed(1, )Ego vehicle speed Size(2, )Vehicle length and width Goal position(2, )Relative goal coordinates (x,y) Collision state(1, )1 if collided, 0 otherwise surrounding vehicle Speed(max num agentsâ 1, 1)Partner vehicle speed Size(max num agentsâ 1, 2)Length and width Goal position(max num agentsâ 1, 2)Relative goal coordinates (x,y) Collision state(max num agentsâ 1, 1)1 if collided, 0 otherwise road graph Segment position(top k road points, 2)Road point coordinates Segment size(top k road points, 3)Length, width, and height Segment orientation(top k road points, 1)Orientation of the road segment Segment type(top k road points, 8)One-hot encoded road point type B.2 Inverse Kinematics Model B.2.1 Delta Dynamics Model In our simulation environment, we employ a delta dynamics model [Gulino et al., 2023] to update the agent states. Unlike models that apply displacements directly in the global frame, our approach defines actions in the local coordinate frame of the agent, which aligns with its heading direction. At each timestep t, the action is represented as a t = (âx t , ây t , âĎ t ),(18) whereâx t andây t denote the forward and lateral displacements relative to the agentâs orientation, andâĎ t is the change in yaw angle. To apply the action in the global coordinate frame, we rotate the local displacement using the current yaw Ď t of the agent: " âx global t ây global t # = R(Ď t ) âx t ây t ,âĎ global t = âĎ t ,(19) whereR(Ď t )is the 2D rotation matrix. The global displacements are then added to the agentâs current state(x t ,y t ,Ď t )to obtain the updated trajectory(x t+1 ,y t+1 ,Ď t+1 ). For inverse dynamics, the procedure is reversed: we compute the displacement between consecutive states in the global frame and project it back into the local frame usingR(âĎ t ). This formulation allows actions to be expressed relative to the agentâs forward-facing direction, making them more natural for imitation learning and policy optimization tasks. B.2.2 Bicycle Model Following the kinematic bicycle model [Gulino et al., 2023], we define an agentâs current state information ass = (x,y,θ,v x ,v y ), which includes thex,ypositions in the coordinate space, the yaw angleθ, and the velocities in theXandYdirections. The action is defined as(a,Îş), wherea 16 denotes the longitudinal acceleration andÎşthe steering curvature. Letv = q v 2 x + v 2 y denote the speed magnitude. For the inverse kinematics, given the state information of two consecutive statess = (x,y,θ,v x ,v y ) and s Ⲡ= (x Ⲡ,y Ⲡ,θ Ⲡ,v Ⲡx ,v Ⲡy ), we estimate the acceleration a and steering curvature Îş as a = v Ⲡâ v ât Îş = θ Ⲡâ θ d , d = p (x Ⲡâ x) 2 + (y Ⲡâ y) 2 , wherev Ⲡ= q v Ⲡx 2 + v Ⲡy 2 . This formulation provides a compact approximation of car-like motion for extracting action labels from logged trajectories. B.3 Trajectory Types Distributions ReverseTurnStraightUncategorized 0 10 20 30 40 50 60 Percentage (%) 0.7% 17.9% 17.5% 63.9% 0.7% 17.9% 17.1% 64.2% Training Validation Figure 9: Distributions of 4 trajectory types on training and validation set. Straight refers to trajectories in which changes in âyandâyawremain below predefined thresholds. Turn denotes the presence of a contiguous inter- val with notable variations inâyandâyaw. Re- verse captures cases whereâxis negative for at least half of the trajectory. Uncategorized encom- passes all remaining trajectories that do not sat- isfy the criteria above. The thresholds and criteria employed for this categorization are empirically determined using domain knowledge and a pre- liminary inspection of the dataset. As shown in Figure 9, the trajectory distribution is highly imbal- anced, with nominal scenariosâparticularly Uncat- egorized and Straightâaccounting for more than 80% of the data, whereas Reverse cases constitute less than 1%. To label the trajectory type, we extract the action values (âx, ây, âyaw) of whole trajectories. The training dataset has 542K samples, and the validation dataset has 69 K samples. Empirically, we first filtered out trajectories whose maximum lateral offset,max(ây), exceeded 0.5 and whose maximum yaw change,max(âyaw), exceeded 0.2. We then labeled the remaining trajectories as follows: a trajectory was labeled Straight if the peaks of bothâyandâyawwere below 0.01; Turn if the fraction of timesteps withây > 0.035andâyaw > 0.025was at least 15% of the trajectory; Reverse if the proportion of timesteps withâx < â0.01was at least 50%; and Uncategorized otherwise. 17 C Additional Scaling Results C.1 Training and Validation Scenes Results Table 3: Driving performance metrics across dataset scales of BC model. Num Scenes DatasetGoal RateOff-RoadVeh-Coll Goal Progress Ratio 100 Training0.320Âą 0.0620.306Âą 0.0500.125Âą 0.0310.606Âą 0.149 Validation0.302Âą 0.0340.352Âą 0.0300.183Âą 0.0110.561Âą 0.115 500 Training0.568Âą 0.1330.127Âą 0.0520.092Âą 0.0120.785Âą 0.092 Validation0.521Âą 0.1110.156Âą 0.0330.111Âą 0.0180.745Âą 0.083 1000 Training0.553Âą 0.0380.074Âą 0.0190.083Âą 0.0250.780Âą 0.063 Validation0.543Âą 0.0170.114Âą 0.0110.107Âą 0.0240.748Âą 0.072 5000 Training0.572Âą 0.0970.074Âą 0.0140.070Âą 0.0110.797Âą 0.019 Validation0.573Âą 0.0930.086Âą 0.0130.073Âą 0.0070.795Âą 0.015 10000 Training0.615Âą 0.1670.058Âą 0.0080.076Âą 0.0200.808Âą 0.043 Validation0.615Âą 0.1660.067Âą 0.0080.073Âą 0.0190.809Âą 0.041 20000 Training0.710Âą 0.0200.075Âą 0.0310.062Âą 0.0060.854Âą 0.039 Validation0.709Âą 0.0180.077Âą 0.0300.060Âą 0.0020.857Âą 0.035 40000 Training0.567Âą 0.1630.067Âą 0.0060.073Âą 0.0340.836Âą 0.076 Validation0.566Âą 0.1590.069Âą 0.0100.071Âą 0.0360.839Âą 0.073 80000 (Full) Training0.583Âą 0.0330.066Âą 0.0040.076Âą 0.0140.808Âą 0.074 Validation0.587Âą 0.0340.065Âą 0.0030.073Âą 0.0150.813Âą 0.072 Table 4: Driving performance metrics across dataset scales of the RL model. Num Scenes DatasetGoal RateOff-RoadVeh-Coll Goal Progress Ratio 100Validation0.491Âą 0.1230.088Âą 0.0340.091Âą 0.0200.519Âą 0.119 500Validation0.743Âą 0.0580.081Âą 0.0100.069Âą 0.0120.757Âą 0.067 1000Validation0.875Âą 0.0380.077Âą 0.0170.060Âą 0.0080.885Âą 0.036 5000Validation0.861Âą 0.1010.068Âą 0.0110.040Âą 0.0030.882Âą 0.072 10000Validation0.989Âą 0.0050.034Âą 0.0110.023Âą 0.0050.986Âą 0.005 As shown in Table 3, the BC model improves consistently on almost all metrics up to 10,000 scenes, after which performance largely saturates. The generalization gap between the training and validation sets is also mostly resolved beyond 20,000 scenes. By contrast, for RL, Table 4 shows gradual improvement in all metrics except the goal progress ratio. This indicates a weaker scaling trend than in BC, likely because the RL objective is more directly tied to destination-reaching behavior. C.2 WOSAC metrics of Power-law Relationship WOSAC metrics are widely used to evaluate human likeness in autonomous driving. Following this, we run a data-scaling study to test whether planning quality exhibits a power-law relationship with the amount of training data. As shown in Figure 10, all five metrics are strongly correlated with the number of scenes, and are well-approximated by a power-law fit. The four WOSAC sub- metricsâRealism Meta, Kinematic, Interactive, and Map-basedâimprove steadily with data up toâź 10k scenes, after which gains largely saturate. Our Realism Meta score approaches 0.7, and minADE reachesâź1.1, both within a reasonable range and close to the 2023 leaderboard. This saturation behavior mirrors our surrounding-vehicle linear probing results, which also plateau beyondâź10K scenes. Finally, our random baseline matches that reported in PufferDrive [Cornelisse et al., 2025a], providing a reference point for interpreting the absolute scale of these scores. 18 1005001k5k10k20k40k80k Number of Scenes 0.34 0.36 0.38 0.40 0.42 0.44 0.46 Score = . Kinematic = . . 0.04 0.06 Random 1005001k5k10k20k40k80k Number of Scenes 0.55 0.56 0.57 0.58 0.59 0.60 0.61 Score = . Interactive = . . 0.450 0.475 Random 1005001k5k10k20k40k80k Number of Scenes 0.600 0.625 0.650 0.675 0.700 0.725 0.750 0.775 Score = . Map-based = . . 0.42 0.44 Random 1005001k5k10k20k40k80k Number of Scenes 1.0 1.5 2.0 2.5 3.0 3.5 m = â . minADE = . â . 23.2 24.0 Random Figure 10: Power law relationships for WOSAC metrics: We evaluate the realism meta score, kinematic score, interactive score, map-based score, and minADE for 1,000 scenes in the validation set. C.3 Additional results of Power-law Relationship 1005001k5k10k20k40k80k Number of Scenes 3 Ă 10 â1 4 Ă 10 â1 6 Ă 10 â1 Goal Success Ratio = . Goal Rate = . â . 1005001k5k10k20k40k80k Number of Scenes 10 â1 Offroad Ratio = â . Off-Road = . â â . 1005001k5k10k20k40k80k Number of Scenes 10 â1 4 Ă 10 â2 6 Ă 10 â2 2 Ă 10 â1 Collision Ratio = â . Veh-Coll = . â â . 1005001k5k10k20k40k80k Number of Scenes 5 Ă 10 â1 6 Ă 10 â1 7 Ă 10 â1 8 Ă 10 â1 9 Ă 10 â1 Goal Progress = . Goal Progress Ratio = . â . (a) Closed-loop scaling. Goal/collision vs data; dashed: power-law fit;r: correlation. 1005001k5k10k20k40k80k Number of Scenes 10 1 9.5 Ă 10 0 1.05 Ă 10 1 1.1 Ă 10 1 1.15 Ă 10 1 Negative GMM Loss y = 9.056x 0.021 r = 0.936 Negative GMM Loss 1005001k5k10k20k40k80k Number of Scenes 2 Ă 10 2 3 Ă 10 2 4 Ă 10 2 x Loss (L1) y = 0.026x 0.025 r =0.885 x Loss (L1) 1005001k5k10k20k40k80k Number of Scenes 5.2 Ă 10 3 5.4 Ă 10 3 5.6 Ă 10 3 5.8 Ă 10 3 6 Ă 10 3 6.2 Ă 10 3 6.4 Ă 10 3 6.6 Ă 10 3 y Loss (L1) y = 0.006x 0.018 r =0.911 y Loss (L1) 1005001k5k10k20k40k80k Number of Scenes 2.2 Ă 10 3 2.4 Ă 10 3 2.6 Ă 10 3 2.8 Ă 10 3 3 Ă 10 3 3.2 Ă 10 3 3.4 Ă 10 3 3.6 Ă 10 3 yaw Loss (L1) y = 0.004x 0.057 r =0.899 yaw Loss (L1) (b) Open-loop scaling. GMM and action L1 losses vs data. In this section, we show the additional power-law relationship results for closed-loop metrics as shown in Figure 11a and for open-loop metrics as shown in Figure 11b. To show the performance by cases, we also conduct the power-law relationship by types as in Figure 12. In most cases, the coefficientrwas high, indicating a strong relationship between data scale and performance. In the case ofReverse, the metrics have a high standard deviation because only rare cases occur. However, there are no collisions after 5,000 scenes. 19 1005001K5K10K20K40K80K 0.500 Straight Goal Success Ratio r=0.706 Goal Rate 1005001K5K10K20K40K80K 1.000 r=0.729 Goal Progress Ratio 1005001K5K10K20K40K80K 0.001 0.002 0.005 0.010 0.020 0.050 0.100 0.200 r=-0.923 Off-Road 1005001K5K10K20K40K80K 0.020 0.050 0.100 0.200 r=-0.847 Veh-Coll 1005001K5K10K20K40K80K 0.100 0.200 0.500 Turn Goal Success Ratio r=0.838 1005001K5K10K20K40K80K 0.0000 0.0000 0.0001 0.0002 0.002 0.010 0.100 0.500 r=0.881 1005001K5K10K20K40K80K 0.200 0.500 r=-0.924 1005001K5K10K20K40K80K 0.100 0.200 r=-0.957 1005001K5K10K20K40K80K 0.500 Normal Goal Success Ratio r=0.786 1005001K5K10K20K40K80K 0.500 r=0.819 1005001K5K10K20K40K80K 0.050 0.100 0.200 r=-0.935 1005001K5K10K20K40K80K 0.050 0.100 0.200 r=-0.894 1005001K5K10K20K40K80K Number of Scenes 0.200 0.500 Reverse Goal Success Ratio r=0.860 1005001K5K10K20K40K80K Number of Scenes 0.0000 0.0000 0.0001 0.0002 0.002 0.010 0.100 0.500 r=0.781 1005001K5K10K20K40K80K Number of Scenes 0.001 0.002 0.010 0.020 0.100 0.200 r=-0.909 1005001K5K10K20K40K80K Number of Scenes 0.001 0.002 0.005 0.020 0.050 0.100 r=-0.801 Figure 12: Power law relationships for simulation results by cases (BC): Closed-loop simulation results for BC with data scaling. Missing points correspond to zero-valued metrics. 20 D Additional Surrounding-vehicle probing results D.1 Linear Probing Setting Table 5: Linear Probing Setting Details for Surrounding-Vehicle Prediction. #Scenes Gradient (Step) #Samples (fs@10) #Samples (fs@20) #Samples (fs@30) #Samples (fs@40) 1005K235K184K140K104K 50012K983K768K588K437K 100020K1.8M1.4M1.1M824K 500050K8.6M6.7M5.1M3.8M 10,000 75K17.6M13.7M10.5M7.8M 20,000100K35.2M27.4M21M15.7M 40,000125K70M54.6M41.9M31.2M 80,000150K140.4M110M84.1M62.7M Across all experiments with varying dataset sizes, we train with a learning rate of0.0015and a batch size of 256. The settings across dataset sizes are shown in Table 5.a For SMART (IL), we trained the linear probe for 15,000 gradient steps with the learning rate of0.001and a batch size of 16. In the case of the raw-input model, we use agent-centric features: position, heading, velocity, and type with 11 historical steps. D.2 Detailed Results of Linear Probing 0.1k0.5k 1k5k 10k20k40k80k 0.00 0.05 0.10 F1 Macro F1 Macro (Future Step: 10) 0.1k0.5k 1k5k 10k20k40k80k 0.00 0.05 0.10 F1 Macro (Future Step: 20) 0.1k0.5k 1k5k 10k20k40k80k 0.00 0.05 0.10 F1 Macro (Future Step: 30) 0.1k0.5k 1k5k 10k20k40k80k 0.00 0.05 0.10 F1 Macro (Future Step: 40) 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes 0.0 0.1 0.2 Accuracy Uncategorized Accuracy 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes 0.0 0.1 0.2 Turn Accuracy 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes 0.0 0.1 0.2 Straight Accuracy 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes 0.0 0.1 0.2 Reverse Accuracy Raw-Input (F1) Early LP (F1) Late LP (F1) Raw-Input (Acc) Early LP (Acc) Late LP (Acc) Figure 13: Full performance metrics with surrounding-linear probing (BC): Light colors are raw-input probing and darker colors are BC probing(earlier layer and late layer) with error bars of standard deviation across different seeds. For convenience, we refer to linear probing in the earlier layer of the BC model as early LP and in the later layer as late LP. The top row is the F1 score across future step10, 20, 30, 40 and the bottom row is accuracy with labeled cases (Future step = 10). In this section, we present the full linear probing results for BC and RL, shown in Figure 13 and Figure 14. The BC model does learn about surrounding vehicles as the training scale increases, but its absolute predictive accuracy remains limited. Although both early and late LP consistently outperform the raw-input probe, the overall F1 score remains below 10%, indicating weak predictive power for surrounding vehicles. This weakness is also reflected in type accuracy: BC performs better on relatively easy cases such as Straight and Uncategorized, but degrades substantially on more complex cases such as Turn and Reverse, suggesting that it prioritizes easier agents rather than the safety-critical ones. Interestingly, as the dataset grows, BC probing gains concentrate more on the near horizon, whereas the raw-input probe continues to improve at farther future steps. This suggests 21 0.1k0.5k 1k5k 10k 0.0 0.1 0.2 F1 Macro F1 Macro (Future Step: 10) 0.1k0.5k 1k5k 10k 0.0 0.1 0.2 F1 Macro (Future Step: 20) 0.1k0.5k 1k5k 10k 0.0 0.1 0.2 F1 Macro (Future Step: 30) 0.1k0.5k 1k5k 10k 0.0 0.1 0.2 F1 Macro (Future Step: 40) 0.1k0.5k 1k5k 10k Number of Scenes 0.0 0.1 0.2 Accuracy Uncategorized Accuracy 0.1k0.5k 1k5k 10k Number of Scenes 0.0 0.1 0.2 Turn Accuracy 0.1k0.5k 1k5k 10k Number of Scenes 0.0 0.1 0.2 Straight Accuracy 0.1k0.5k 1k5k 10k Number of Scenes 0.0 0.1 0.2 Reverse Accuracy Raw-Input (F1) LP (F1) Raw-Input (Acc) LP (Acc) Figure 14: Full performance metrics with surrounding-linear probing (RL): Light colors are raw-input probing and darker colors are RL probing with error bars of standard deviation across different seeds. For convenience, we refer to linear probing as LP. The top row is the F1 score across future steps10, 20, 30, 40, and the bottom row is accuracy with labeled cases (Future step = 10). Note that since there are no success cases of Reverse, so no probing results in Reverse. that the BC model increasingly focuses on imminent, collision-prone intervals while paying less attention to longer-range futures. The widening gap between early and late LP at larger scales further indicates that deeper representations increasingly encode near-future behavior. By contrast, the RL model achieves consistently higher F1 scores than BC across all future horizons, suggesting that it learns stronger representations of surrounding-vehicle behavior. The fact that the raw-input probe is also stronger than in BC further suggests that self-play trajectories themselves are more predictable, indicating that RL agents may interact in a more mutually predictable manner. However, this advantage is concentrated on the near future. As the dataset grows, RL largely maintains its performance at the 10-step horizon, while showing weaker gains or slight degradation at longer horizons. This implies that RL increasingly prioritizes short-horizon behavior and discards information about the distant future. One possible explanation is that, unlike BC, which uses stacked observations over multiple past steps, RL operates on single-step inputs and therefore has less access to temporal context for longer-horizon prediction. Overall, these results suggest that RL learns stronger, more short-horizon-oriented representations of surrounding vehicles than BC does. Additionally, we introduce the additional results of linear probing as in the Table 6 and Table 7. Note that the RL model does not have a late LP, since it fuses at an earlier layer. From the Table 6, we report the F1 Macro score across all future timesteps. In Table 7, we report the full results of type accuracy. As shown in Table 6, the BC model shows clear saturation after 10,000 scenes in both Info Loss and LP - Raw. Both values also decrease as the future step increases, indicating that the gain from surrounding-vehicle representations is concentrated on the near horizon. The positive Info Loss further suggests that earlier layers retain more surrounding-vehicle information, while later layers gradually discard it, likely in favor of other signals such as ego planning. By contrast, the RL model exhibits a much larger LP - Raw gap even at 100 scenes, indicating that useful surrounding- vehicle representations emerge much earlier than in BC. This may help explain why RL shows relatively stable collision behavior. However, this advantage is concentrated at short horizons, while longer-horizon gains shrink substantially as scale increases. As shown in Table 7, the BC model exhibits increasingly larger type-wise probing gains with scale, especially for Normal and Straight cases, while Turn remains consistently more difficult and Reverse stays weak and unstable. The positive Info Loss across most types further suggests that earlier layers retain more surrounding-vehicle information, which is gradually discarded in later layers. By contrast, RL shows relatively strong LP - Raw gains even at small scales, indicating earlier emergence of useful surrounding-vehicle representations. However, unlike in BC, these gains do not increase 22 Table 6: F1 Macro differences across future steps. (a) BC, (b) RL.Info Lossis (Early LPâLate LP), andLP â Rawis the difference between probing and the raw-input model. The RL table only has LP â Raw. (a) BC FS=10FS=20FS=30FS=40 Info Loss.LP - RawInfo Loss.LP - RawInfo Loss.LP - RawInfo Loss.LP - Raw 1000.00050.01990.00130.01580.00070.01600.00010.0124 5000.00800.04350.00860.03430.00610.02800.00390.0168 10000.00880.04670.00860.03230.00780.02720.00370.0226 50000.01210.06080.00800.03490.00770.02710.00990.0156 100000.01610.07040.01440.04510.01510.03370.01620.0185 200000.00990.05970.01030.03630.00660.02640.00740.0138 400000.00380.05790.00670.03710.00570.02650.00710.0181 800000.00460.06100.00020.03100.00120.02270.00230.0163 (b) RL #ScenesFS=10FS=20FS=30FS=40 1000.04900.05370.04590.0573 5000.06650.07350.07310.0621 10000.07260.07570.07030.0679 50000.06410.06370.05670.0505 100000.05580.04680.0063-0.0077 Table 7: Accuracy differences by action types. (a) BC at future step 10, (b) RL at future step 10. Info Lossis (Early LPâLate LP), andLP â Rawis the difference between probing and the raw-input model. (a) BC (Future step = 10) NormalTurnStraightReverse Info Loss.LP - RawInfo Loss.LP - RawInfo Loss.LP - RawInfo Loss.LP - Raw 1000.00210.03800.00250.01160.00200.04110.00040.0188 5000.01320.08720.01760.04510.01120.08380.01040.0368 10000.01730.09250.01690.04740.01320.0882-0.00390.0306 50000.02420.11980.02310.07220.02170.12330.02070.0373 100000.02470.12710.02460.07980.02180.12790.01210.0352 200000.01940.10700.02180.08450.02030.11200.00220.0424 400000.01000.10900.01280.07930.01040.1111-0.00470.0488 800000.01090.11270.01100.08260.00440.11510.00020.0516 (b) RL (Future step = 10) NormalTurnStraightReverse LP - RawLP - RawLP - RawLP - Raw 1000.05010.06380.0619â 5000.09650.06330.0905â 10000.07830.07200.0988â 50000.08500.05810.0900â 100000.04840.03480.0815â 23 monotonically with scale; instead, they weaken at 10K scenes for several types. Overall, BC improves more steadily with data, whereas RL learns useful type-specific representations earlier but does not sustain the same scaling trend. D.3 Type Accuracy of Future Steps The figure 15 shows the corresponding results for surrounding-vehicle probing. Here, accuracies are overall lower than in the ego case, and the gap between early and late LP is smaller and less stable, especially for Turn and Reverse. However, there is no tendency for type accuracy in RL as in 16. 100500 10005000 10000200004000080000 0.0 0.1 0.2 Uncategorized Future Step 20 100500 10005000 10000200004000080000 Future Step 30 100500 10005000 10000200004000080000 Future Step 40 100500 10005000 10000200004000080000 0.0 0.1 Turn 100500 10005000 10000200004000080000 100500 10005000 10000200004000080000 100500 10005000 10000200004000080000 0.0 0.1 0.2 Straight 100500 10005000 10000200004000080000 100500 10005000 10000200004000080000 100500 10005000 10000200004000080000 Num Scene 0.00 0.05 0.10 Reverse 100500 10005000 10000200004000080000 Num Scene 100500 10005000 10000200004000080000 Num Scene Raw-Input (Acc)Early LP (Acc)Late LP (Acc) Figure 15: Type accuracy of other future steps (20, 30, 40) (Other) (BC): The row is type (Normal, Turn, Straight, and Reverse) and the column is future step (20, 30, 40). 24 100500 10005000 10000 0.0 0.1 0.2 Uncategorized Future Step 20 100500 10005000 10000 Future Step 30 100500 10005000 10000 Future Step 40 100500 10005000 10000 0.0 0.2 Turn 100500 10005000 10000 100500 10005000 10000 100500 10005000 10000 0.0 0.2 0.4 Straight 100500 10005000 10000 100500 10005000 10000 100500 10005000 10000 Num Scene 0.0 0.5 1.0 Reverse 100500 10005000 10000 Num Scene 100500 10005000 10000 Num Scene Raw-Input (Acc)LP (Acc) Figure 16: Type accuracy of other future steps (20, 30, 40) (Other) (RL): The row is type (Normal, Turn, Straight, and Reverse) and the column is future step (20, 30, 40). 25 E Additional Ego probing results E.1 Linear Probing Setting with Data Scaling Table 8: Linear Probing Setting Details for Ego Prediction. #Scenes Gradient (Step) #Samples (fs@10) #Samples (fs@20) #Samples (fs@30) #Samples (fs@40) 1005K37K31K26K20K 50012K181K152K125K99K 100020K355K297K244K193K 500050K1.7M1.4M1.2M937K 10,000 75K3.5M2.9M2.3M1.9M 20,000100K7M5.8M4.8M3.8M 40,000125K13.9M11.6M9.5M756K 80,000150K27.7M23.2M19M15.1M Across all experiments with varying dataset sizes, we train with a learning rate of0.0015and a batch size of 256. The different settings across different dataset sizes are in Table 8. E.2 Detailed Results of Linear Probing In this section, we report the full results of ego AV linear probing as in Figure 17 and 18. The overall F1 Macro score exceeds 50%, indicating that the BC model successfully learns not only about instant actions but also about long-horizon planning. As the dataset size increases, later-layer probes outperform early LP probes, suggesting that the BC model increasingly integrates map and other vehicle context when selecting actions. This trend contrasts with the linear probing of surrounding vehicles, where early LP probes surpass late LP probes (early LP>late LP). Interestingly, the BC model forms coherent representations for reverse cases, even though such behavior is not reliably learned from raw inputs. 0.1k0.5k 1k5k 10k20k40k80k 0.0 0.2 0.4 0.6 F1 Macro F1 Macro (Future Step: 10) 0.1k0.5k 1k5k 10k20k40k80k 0.0 0.2 0.4 0.6 F1 Macro (Future Step: 20) 0.1k0.5k 1k5k 10k20k40k80k 0.0 0.2 0.4 0.6 F1 Macro (Future Step: 30) 0.1k0.5k 1k5k 10k20k40k80k 0.0 0.2 0.4 0.6 F1 Macro (Future Step: 40) 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes 0.00 0.25 0.50 0.75 Accuracy Uncategorized Accuracy 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes 0.00 0.25 0.50 0.75 Turn Accuracy 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes 0.00 0.25 0.50 0.75 Straight Accuracy 0.1k0.5k 1k5k 10k20k40k80k Number of Scenes 0.00 0.25 0.50 0.75 Reverse Accuracy Raw-Input (F1) Early LP (F1) Late LP (F1) Raw-Input (Acc) Early LP (Acc) Late LP (Acc) Figure 17: Performance metrics with ego AV linear probing (BC): Light colors are raw-input probing and darker colors are BC probing (earlier layer and late layer) with error bars of standard deviation across different seeds. The top row is the F1 score across future steps10, 20, 30, 40and the bottom row is accuracy with labeled cases (Future step = 10). E.3 Type Accuracy of Future Steps The figure 19 summarizes the accuracy of ego linear probing across trajectory types (Uncategorized, Turn,Straight,Reverse), future steps (20, 30, 40), and dataset scales. Across all types and future steps, both early and late LP substantially outperform the raw-input baseline. The gains are most 26 Table 9: F1 Macro differences across future steps. (a) BC, (b) RL.Info Lossis (Early LPâLate LP), and LP â Raw is the difference between trained model probing and the raw-input model. (a) BC FS=10FS=20FS=30FS=40 Info Loss.LP - RawInfo Loss.LP - RawInfo Loss.LP - RawInfo Loss.LP - Raw 1000.00600.08470.01860.15220.02410.18490.03120.1784 5000.00860.16520.01500.24250.02260.22550.03300.2307 10000.00470.15020.01220.23390.02620.21670.04240.2068 5000-0.00780.16590.00120.21900.01300.19920.02510.1946 10000-0.01090.1615-0.00300.22180.00820.19360.02100.1888 20000-0.01920.1731-0.01130.21880.00260.19300.01580.1925 40000-0.01900.1674-0.01000.20630.00320.18870.01860.1799 80000-0.01850.1701-0.01210.2151-0.00040.18880.01430.1800 (b) RL #ScenesFS=10FS=20FS=30FS=40 1000.22070.32340.24870.1218 5000.28720.34940.26160.2799 10000.19980.24760.22400.1567 50000.14220.18590.18230.1801 100000.28930.31910.33940.2771 Table 10: Accuracy differences by action types. (a) BC at future step 10, (b) RL at future step 10. Info Lossis (Early LPâLate LP), andLP â Rawis the difference between model probing and the raw-input model. (a) BC (Future step = 10) NormalTurnStraightReverse Info Loss.LP - RawInfo Loss.LP - RawInfo Loss.LP - RawInfo Loss.LP - Raw 1000.01090.1016-0.00560.01960.01130.07170.06100.1219 5000.01410.1666-0.00360.04310.01720.1589-0.20880.5369 10000.00540.1454-0.00310.03140.00220.12960.01130.5561 5000-0.00470.1505-0.01020.04880.00010.1492-0.01150.6478 10000-0.00900.1479-0.01300.0476-0.00770.1411-0.03650.6517 20000-0.01590.1554-0.01560.0534-0.01320.1563-0.00490.6392 40000-0.01360.1517-0.01550.0541-0.01420.1423-0.00970.6406 80000-0.01450.1556-0.01410.0527-0.01150.14770.01660.6605 (b) RL (Future step = 10) NormalTurnStraightReverse IL - RawIL - RawIL - RawIL - Raw 1000.07120.17890.2450â 5000.19960.23140.1587â 10000.12690.09820.1205â 50000.03530.08440.0604â 100000.15620.33510.0744â 27 0.1k0.5k 1k5k 10k 0.00 0.25 0.50 0.75 F1 Macro F1 Macro (Future Step: 10) 0.1k0.5k 1k5k 10k 0.00 0.25 0.50 0.75 F1 Macro (Future Step: 20) 0.1k0.5k 1k5k 10k 0.00 0.25 0.50 0.75 F1 Macro (Future Step: 30) 0.1k0.5k 1k5k 10k 0.00 0.25 0.50 0.75 F1 Macro (Future Step: 40) 0.1k0.5k 1k5k 10k Number of Scenes 0.0 0.5 1.0 Accuracy Uncategorized Accuracy 0.1k0.5k 1k5k 10k Number of Scenes 0.0 0.5 1.0 Turn Accuracy 0.1k0.5k 1k5k 10k Number of Scenes 0.0 0.5 1.0 Straight Accuracy 0.1k0.5k 1k5k 10k Number of Scenes 0.0 0.5 1.0 Reverse Accuracy Raw-Input (F1) LP (F1) Raw-Input (Acc) LP (Acc) Figure 18: Performance metrics with ego AV linear probing (RL): Light colors are raw-input probing and darker colors are RL probing (earlier layer and late layer) with error bars of standard deviation across different seeds. The top row is the F1 score across future steps10, 20, 30, 40and the bottom row is accuracy with labeled cases (Future step = 10). significant for theStraightandReversecases, and performance gradually improves, then saturates as the number of scenes increases. At the same time, accuracy consistently declines as the future horizon lengthens. The figure 15 shows the corresponding results for surrounding-vehicle probing. Here, accuracies are overall lower than in the ego case, and the gap between early and late LP is smaller and less stable, especially for Turn and Reverse. 100500 10005000 10000200004000080000 0.0 0.5 Uncategorized Future Step 20 100500 10005000 10000200004000080000 Future Step 30 100500 10005000 10000200004000080000 Future Step 40 100500 10005000 10000200004000080000 0.0 0.5 Turn 100500 10005000 10000200004000080000 100500 10005000 10000200004000080000 100500 10005000 10000200004000080000 0.0 0.5 Straight 100500 10005000 10000200004000080000 100500 10005000 10000200004000080000 100500 10005000 10000200004000080000 Num Scene 0.00 0.25 0.50 Reverse 100500 10005000 10000200004000080000 Num Scene 100500 10005000 10000200004000080000 Num Scene Raw-Input (Acc)Early LP (Acc)Late LP (Acc) Figure 19: Type accuracy of other future steps (20, 30, 40 (Ego) (BC): The row is type (Normal, Turn, Straight, and Reverse) and the column is future step (20, 30, 40). 28 100500 10005000 10000 0.0 0.5 Uncategorized Future Step 20 100500 10005000 10000 Future Step 30 100500 10005000 10000 Future Step 40 100500 10005000 10000 0.0 0.5 Turn 100500 10005000 10000 100500 10005000 10000 100500 10005000 10000 0.0 0.5 1.0 Straight 100500 10005000 10000 100500 10005000 10000 100500 10005000 10000 Num Scene 0.0 0.5 1.0 Reverse 100500 10005000 10000 Num Scene 100500 10005000 10000 Num Scene Raw-Input (Acc)LP (Acc) Figure 20: Type accuracy of other future steps (20, 30, 40 (Ego) (RL): The row is type (Normal, Turn, Straight, and Reverse) and the column is future step (20, 30, 40). 29 F Further Analysis F.1 Ego-linear Probing Visualization Comparison. Figure 21: Comparison of ego-linear probing in single AV driving and all vehicles driving with rendered examples: The visualization of ego-linear probing when single AV drives (Top) and all vehicles include AV drive (Bottom). The expert trajectory is shown as a dark green line. The ego AV is blue; it turns yellow when off-road and red when a vehicle collides. To better understand the degradation of single AV driving performance, we visualize ego-linear probing in simulation. We compare ego-linear probing results across the 100 scenes with the largest performance gap. As shown in Figure 21, the top row shows the single AV driving, and the bottom row shows all vehicles driving. In the single AV case, the ego collides aroundt = 40and subsequently loses the ability to plan toward the goal. By contrast, in the all-vehicles-driving case, the model plans slightly below the ground truth path att = 40, leading to a collision aroundt = 60; however, even after the crash, the AV continues to plan accurately toward the goal and recovers the exact trajectory att = 80. From the results, we suspect the cause of this situation is the collapse of ego-linear probing, which also provides strong evidence that the presence of surrounding vehicles affects ego AV planning, given the datasetâs lack of single-AV driving. 30 F.2 Additional Near-Collision Analysis Figure 22: Near-collision Analysis. (Off-road) Each heatmap shows the normalized difference between the predicted probability in the pre-collision windowwand that of each spatial grid over the full episode, computed as wâa grid a , at 10 to 40 steps before collision. The red dot marks the ego vehicle position. Dark gray cells indicate grids that were filtered out due to insufficient collision samples. Top: RL, Bottom: BC 31 G Intervention Results Table 11: Statistics of intervention cases (a) BC: Success-rate cases CaseSuccessFailTotal route change20323 increase speed41115 slow speed415 (b) BC: Adaptiveness and recovery CaseSuccessFailTotal Adaptiveness281543 Recovery8816 (c) RL: Success-rate cases CaseSuccessFailTotal route change12921 increase speed11516 slow speed9716 (d) RL: Adaptiveness and recovery CaseSuccessFailTotal Adaptiveness322153 Recovery505 (e) SMART (IL): Success-rate cases CaseSuccessFailTotal route change131023 increase speed13518 slow speed819 (f) SMART (IL): Adaptiveness and recovery CaseSuccessFailTotal Adaptiveness331851 Recovery7613 We report the detailed results of the intervention experiment in Table 11. For BC, most route change cases are successfully modified to avoid collisions, but interventions requiring increased speed remain difficult, with only 4/15 successful cases. In contrast, RL and SMART (IL) show stronger performance in increase speed cases, achieving 11/16 and 13/18 successes, respectively. IL also performs well in slow speed cases, with 8/9 successful interventions. Overall, the ego-planning probe changes in 36/59 BC cases, 37/58 RL cases, and 40/64 IL cases. These results suggest that the proposed intervention framework applies not only to BC but also to RL and IL models, and that surrounding-vehicle predictions can causally influence ego planning across different driving models. G.1 Failure Case We analyze the failure cases, as illustrated in Figure 23, where there is no change compared to the initial prediction of ego planning. Interestingly, most cases are increasing in speed. As in cases (a) and (d), we make the front vehicles go faster so that the ego vehicle can increase its speed. However, the BC model remains on the same grid as before. Especially in case (d), even if we move multiple vehicles that could affect the egoâs future position, the BC model remains in the same grid. In the case of (b) and (c), we increase the following vehiclesâ speed so that the ego prediction should speed up. However, there is no change in the BC model. In contrast, the RL policies have difficulty changing lanes, as seen in (e), (f), and (g). In (h), we expect the RL policy to increase speed, but it fails to do so. Lastly, in the IL model, as shown in (i), the ego vehicle remains in the grid before intervention, and the other vehiclesâ predictions are completely wrong. We restore the other vehicle to move in the correct direction, but the ego planner changes direction incorrectly. The same situation occurs in (j); we correct the vehicle prediction, but suddenly, ego planning goes the wrong way. G.2 Additional Intervention Cases As shown in Figure 24, panels (a)â(d) present additional BC adaptiveness cases, while panels (e)â(h) present RL adaptiveness cases. For the BC cases, panels (a) and (c) illustrate change lane interventions. 32 In (a), the intervention induces a lane change through a single interaction with the yellow vehicle. In (c), the red vehicle follows from behind while the yellow vehicle changes lanes, requiring the ego to maintain speed and change lanes as well; the BC model successfully performs this behavior. Panels (b) and (d) show slowdown interventions. In (b), when the yellow lead vehicle is kept in the egoâs future grid cells, the ego remains behind it and slows its progress. In (d), we place multiple surrounding-vehicle trajectories in front of the ego, and the BC model responds by slowing down accordingly. For the RL cases, panels (e) and (g) show change lane interventions, whereas panels (f) and (h) show slowdown interventions. In (e), the yellow and red vehicles originally occupy implausible positions, but when they are placed on the egoâs predicted path, the ego changes its trajectory to avoid them. In (f) and (h), we modify previously implausible surrounding-vehicle trajectories to make them more feasible, and the RL model correspondingly reduces speed. Finally, in (g), after aligning an initially unrealistic trajectory, the egoâs predicted motion is adjusted so that it no longer overlaps in time with the intervened vehicle. For the IL cases, panels (i) and (k) show that the ego vehicle successfully adjusts its planning when we intervene in the surrounding-vehicle prediction to align with the original planning. In the cases of (j) and (l), we intervene in the egoâs planning to stay on a further trajectory, so we expect the ego to change plans to slow down. The ego plans successfully to slow down to avoid a collision. G.3 Additional Recovery Case As shown in Figure 25, we describe the additional results of recovery cases. To evaluate the correction of the recovery label, we visualize the correctly recovered cases, which we label based on the surrounding vehicle trajectories and the random labels of other vehicles. In panel (a), the original prediction of the yellow vehicle is totally wrong, and the red vehicle is slightly wrong at 30 and 40 steps after. Therefore, we recover the red and yellow vehicle predictions to straighten and verify that the ego vehicle increases its speed slightly faster than before. Interestingly, when we assign a random label, the ego-vehicle prediction is the same as the initial prediction. Panel (b) examples show that the other vehicle prediction overlaps, so the ego prediction needs to avoid other vehicles. Therefore, we change the other vehicleâs prediction along a straight line and show that the ego prediction changes in the same way, whereas random labeling tends to stay in the same grid for longer. In panel (c), which depicts the RL modelâs intervention, we recover the other vehicleâs trajectory to follow the ego vehicle. Then, the ego vehicle adjusts to stay at the destination and avoid overwhelming it, since it knows the yellow vehicle will not overtake the ego vehicle while the random label has the same planning as before the intervention. In panel (d), the RL completely mispredicts the vehicle in front. When we recover that it is feasible, the model did not predict the off-road area in 20 timesteps after. However, the random label indicates the wrong destination after 40 timesteps. Lastly, in (e), we recover the egoâs trajectory as we approach a nearby vehicle to go straight. However, when we intervene randomly, the ego plan becomes corrupted and zigzags. 33 Figure 23: Failure cases of intervention: The visualization examples for failure cases of changing labels. (a)-(d): BC, (e)-(h): RL, (i)-(j): IL 34 Figure 24: Additional Examples for Adaptiveness cases: The visualization examples for adaptive- ness cases. Change Lane: ego linear probing changes the lane to avoid collision. Slowing Speed: When we disturb the trajectory, the model slows down the speed. (a)-(d): BC, (e)-(h): RL, (i)-(l): IL 35 Figure 25: Additional Examples for Recovery cases: The visualization examples for recovery cases. Correct Recovery: When we change the label to align with the correct expert trajectory of the surrounding vehicle. Incorrect Recovery: When we set the label to a random label. (a)-(b): BC, (c)-(d): RL, (e): IL 36 NeurIPS Paper Checklist 1. Claims Question: Do the main claims made in the abstract and introduction accurately reflect the paperâs contributions and scope? Answer: [Yes] Justification: The abstract and introduction accurately reflect the paperâs scope and contribu- tions. The paper analyzes how autonomous driving models represent and use information about surrounding vehicles, using linear probing, perturbed closed-loop simulations, near- collision analysis, and intervention experiments. These analyses support the central claims that prediction and planning capabilities emerge with scale, that nominal closed-loop suc- cess can mask weaknesses in surrounding-vehicle reasoning, and that improving internal predictions can causally improve ego planning. The experimental sections are well aligned with these stated goals. Guidelines: ⢠The answer [N/A] means that the abstract and introduction do not include the claims made in the paper. â˘The abstract and/or introduction should clearly state the claims made, including the contributions made in the paper and important assumptions and limitations. A [No] or [N/A] answer to this question will not be perceived well by the reviewers. ⢠The claims made should match theoretical and experimental results, and reflect how much the results can be expected to generalize to other settings. â˘It is fine to include aspirational goals as motivation as long as it is clear that these goals are not attained by the paper. 2. Limitations Question: Does the paper discuss the limitations of the work performed by the authors? Answer: [Yes] Justification: We discussed the limitations of the work and future works in Section 4. Guidelines: ⢠The answer [N/A] means that the paper has no limitation while the answer [No] means that the paper has limitations, but those are not discussed in the paper. ⢠The authors are encouraged to create a separate âLimitationsâ section in their paper. â˘The paper should point out any strong assumptions and how robust the results are to violations of these assumptions (e.g., independence assumptions, noiseless settings, model well-specification, asymptotic approximations only holding locally). The authors should reflect on how these assumptions might be violated in practice and what the implications would be. â˘The authors should reflect on the scope of the claims made, e.g., if the approach was only tested on a few datasets or with a few runs. In general, empirical results often depend on implicit assumptions, which should be articulated. â˘The authors should reflect on the factors that influence the performance of the approach. For example, a facial recognition algorithm may perform poorly when image resolution is low or images are taken in low lighting. Or a speech-to-text system might not be used reliably to provide closed captions for online lectures because it fails to handle technical jargon. â˘The authors should discuss the computational efficiency of the proposed algorithms and how they scale with dataset size. â˘If applicable, the authors should discuss possible limitations of their approach to address problems of privacy and fairness. ⢠While the authors might fear that complete honesty about limitations might be used by reviewers as grounds for rejection, a worse outcome might be that reviewers discover limitations that arenât acknowledged in the paper. The authors should use their best judgment and recognize that individual actions in favor of transparency play an impor- tant role in developing norms that preserve the integrity of the community. Reviewers will be specifically instructed to not penalize honesty concerning limitations. 37 3. Theory assumptions and proofs Question: For each theoretical result, does the paper provide the full set of assumptions and a complete (and correct) proof? Answer: [N/A] Justification: In this paper, we do not propose new theorems. In section 3.2, we conduct the scaling-law experiment and show that the driving models follow this scale law. Guidelines: ⢠The answer [N/A] means that the paper does not include theoretical results. ⢠All the theorems, formulas, and proofs in the paper should be numbered and cross- referenced. ⢠All assumptions should be clearly stated or referenced in the statement of any theorems. â˘The proofs can either appear in the main paper or the supplemental material, but if they appear in the supplemental material, the authors are encouraged to provide a short proof sketch to provide intuition. ⢠Inversely, any informal proof provided in the core of the paper should be complemented by formal proofs provided in appendix or supplemental material. ⢠Theorems and Lemmas that the proof relies upon should be properly referenced. 4. Experimental result reproducibility Question: Does the paper fully disclose all the information needed to reproduce the main ex- perimental results of the paper to the extent that it affects the main claims and/or conclusions of the paper (regardless of whether the code and data are provided or not)? Answer: [Yes] Justification: This paper provides model settings in Appendix B and experiment settings in Appendix C.1, D.1, and E.1. Guidelines: ⢠The answer [N/A] means that the paper does not include experiments. ⢠If the paper includes experiments, a [No] answer to this question will not be perceived well by the reviewers: Making the paper reproducible is important, regardless of whether the code and data are provided or not. ⢠If the contribution is a dataset and/or model, the authors should describe the steps taken to make their results reproducible or verifiable. â˘Depending on the contribution, reproducibility can be accomplished in various ways. For example, if the contribution is a novel architecture, describing the architecture fully might suffice, or if the contribution is a specific model and empirical evaluation, it may be necessary to either make it possible for others to replicate the model with the same dataset, or provide access to the model. In general. releasing code and data is often one good way to accomplish this, but reproducibility can also be provided via detailed instructions for how to replicate the results, access to a hosted model (e.g., in the case of a large language model), releasing of a model checkpoint, or other means that are appropriate to the research performed. â˘While NeurIPS does not require releasing code, the conference does require all submis- sions to provide some reasonable avenue for reproducibility, which may depend on the nature of the contribution. For example (a) If the contribution is primarily a new algorithm, the paper should make it clear how to reproduce that algorithm. (b)If the contribution is primarily a new model architecture, the paper should describe the architecture clearly and fully. (c)If the contribution is a new model (e.g., a large language model), then there should either be a way to access this model for reproducing the results or a way to reproduce the model (e.g., with an open-source dataset or instructions for how to construct the dataset). (d) We recognize that reproducibility may be tricky in some cases, in which case authors are welcome to describe the particular way they provide for reproducibility. 38 In the case of closed-source models, it may be that access to the model is limited in some way (e.g., to registered users), but it should be possible for other researchers to have some path to reproducing or verifying the results. 5. Open access to data and code Question: Does the paper provide open access to the data and code, with sufficient instruc- tions to faithfully reproduce the main experimental results, as described in supplemental material? Answer: [No] Justification: Although the data and code will not accompany in the submission version, they will be released upon publication. Guidelines: ⢠The answer [N/A] means that paper does not include experiments requiring code. â˘Please see the NeurIPS code and data submission guidelines (https://neurips.c/ public/guides/CodeSubmissionPolicy) for more details. ⢠While we encourage the release of code and data, we understand that this might not be possible, so [No] is an acceptable answer. Papers cannot be rejected simply for not including code, unless this is central to the contribution (e.g., for a new open-source benchmark). â˘The instructions should contain the exact command and environment needed to run to reproduce the results. See the NeurIPS code and data submission guidelines (https: //neurips.c/public/guides/CodeSubmissionPolicy) for more details. â˘The authors should provide instructions on data access and preparation, including how to access the raw data, preprocessed data, intermediate data, and generated data, etc. ⢠The authors should provide scripts to reproduce all experimental results for the new proposed method and baselines. If only a subset of experiments are reproducible, they should state which ones are omitted from the script and why. â˘At submission time, to preserve anonymity, the authors should release anonymized versions (if applicable). â˘Providing as much information as possible in supplemental material (appended to the paper) is recommended, but including URLs to data and code is permitted. 6. Experimental setting/details Question: Does the paper specify all the training and test details (e.g., data splits, hyperpa- rameters, how they were chosen, type of optimizer) necessary to understand the results? Answer: [Yes] Justification: We briefly introduce the experiment settings in Section 3.1. Detailed settings are in the Appendices. Guidelines: ⢠The answer [N/A] means that the paper does not include experiments. ⢠The experimental setting should be presented in the core of the paper to a level of detail that is necessary to appreciate the results and make sense of them. ⢠The full details can be provided either with the code, in appendix, or as supplemental material. 7. Experiment statistical significance Question: Does the paper report error bars suitably and correctly defined or other appropriate information about the statistical significance of the experiments? Answer: [Yes] Justification: We include the error bar and standard deviation in all experiments. Guidelines: ⢠The answer [N/A] means that the paper does not include experiments. ⢠The authors should answer [Yes] if the results are accompanied by error bars, confidence intervals, or statistical significance tests, at least for the experiments that support the main claims of the paper. 39 â˘The factors of variability that the error bars are capturing should be clearly stated (for example, train/test split, initialization, random drawing of some parameter, or overall run with given experimental conditions). â˘The method for calculating the error bars should be explained (closed form formula, call to a library function, bootstrap, etc.) ⢠The assumptions made should be given (e.g., Normally distributed errors). â˘It should be clear whether the error bar is the standard deviation or the standard error of the mean. â˘It is OK to report 1-sigma error bars, but one should state it. The authors should preferably report a 2-sigma error bar than state that they have a 96% CI, if the hypothesis of Normality of errors is not verified. ⢠For asymmetric distributions, the authors should be careful not to show in tables or figures symmetric error bars that would yield results that are out of range (e.g., negative error rates). ⢠If error bars are reported in tables or plots, the authors should explain in the text how they were calculated and reference the corresponding figures or tables in the text. 8. Experiments compute resources Question: For each experiment, does the paper provide sufficient information on the com- puter resources (type of compute workers, memory, time of execution) needed to reproduce the experiments? Answer: [Yes] Justification: In Appendix A.2 and A.3, we recorded the computing resources of our main experiment. Guidelines: ⢠The answer [N/A] means that the paper does not include experiments. ⢠The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage. â˘The paper should provide the amount of compute required for each of the individual experimental runs as well as estimate the total compute. â˘The paper should disclose whether the full research project required more compute than the experiments reported in the paper (e.g., preliminary or failed experiments that didnât make it into the paper). 9. Code of ethics Question: Does the research conducted in the paper conform, in every respect, with the NeurIPS Code of Ethics https://neurips.c/public/EthicsGuidelines? Answer: [Yes] Justification: Guidelines: â˘The answer [N/A] means that the authors have not reviewed the NeurIPS Code of Ethics. â˘If the authors answer [No], they should explain the special circumstances that require a deviation from the Code of Ethics. â˘The authors should make sure to preserve anonymity (e.g., if there is a special consid- eration due to laws or regulations in their jurisdiction). 10. Broader impacts Question: Does the paper discuss both potential positive societal impacts and negative societal impacts of the work performed? Answer: [Yes] Justification: We summarize the paperâs contribution to the autonomous driving field. We analyzed the black-box deep learning autonomous driving models foucsed on utilization of surrounding vehicle information. 40 Guidelines: ⢠The answer [N/A] means that there is no societal impact of the work performed. ⢠If the authors answer [N/A] or [No], they should explain why their work has no societal impact or why the paper does not address societal impact. â˘Examples of negative societal impacts include potential malicious or unintended uses (e.g., disinformation, generating fake profiles, surveillance), fairness considerations (e.g., deployment of technologies that could make decisions that unfairly impact specific groups), privacy considerations, and security considerations. ⢠The conference expects that many papers will be foundational research and not tied to particular applications, let alone deployments. However, if there is a direct path to any negative applications, the authors should point it out. For example, it is legitimate to point out that an improvement in the quality of generative models could be used to generate Deepfakes for disinformation. On the other hand, it is not needed to point out that a generic algorithm for optimizing neural networks could enable people to train models that generate Deepfakes faster. â˘The authors should consider possible harms that could arise when the technology is being used as intended and functioning correctly, harms that could arise when the technology is being used as intended but gives incorrect results, and harms following from (intentional or unintentional) misuse of the technology. â˘If there are negative societal impacts, the authors could also discuss possible mitigation strategies (e.g., gated release of models, providing defenses in addition to attacks, mechanisms for monitoring misuse, mechanisms to monitor how a system learns from feedback over time, improving the efficiency and accessibility of ML). 11. Safeguards Question: Does the paper describe safeguards that have been put in place for responsible release of data or models that have a high risk for misuse (e.g., pre-trained language models, image generators, or scraped datasets)? Answer: [N/A] Justification: Our experiments were conducted on a simulation, not in the real world. Guidelines: ⢠The answer [N/A] means that the paper poses no such risks. ⢠Released models that have a high risk for misuse or dual-use should be released with necessary safeguards to allow for controlled use of the model, for example by requiring that users adhere to usage guidelines or restrictions to access the model or implementing safety filters. ⢠Datasets that have been scraped from the Internet could pose safety risks. The authors should describe how they avoided releasing unsafe images. ⢠We recognize that providing effective safeguards is challenging, and many papers do not require this, but we encourage authors to take this into account and make a best faith effort. 12. Licenses for existing assets Question: Are the creators or original owners of assets (e.g., code, data, models), used in the paper, properly credited and are the license and terms of use explicitly mentioned and properly respected? Answer: [Yes] Justification: The datasets used in this paper are cited within this paper. Guidelines: ⢠The answer [N/A] means that the paper does not use existing assets. ⢠The authors should cite the original paper that produced the code package or dataset. ⢠The authors should state which version of the asset is used and, if possible, include a URL. ⢠The name of the license (e.g., C-BY 4.0) should be included for each asset. 41 â˘For scraped data from a particular source (e.g., website), the copyright and terms of service of that source should be provided. â˘If assets are released, the license, copyright information, and terms of use in the package should be provided. For popular datasets,paperswithcode.com/datasets has curated licenses for some datasets. Their licensing guide can help determine the license of a dataset. â˘For existing datasets that are re-packaged, both the original license and the license of the derived asset (if it has changed) should be provided. â˘If this information is not available online, the authors are encouraged to reach out to the assetâs creators. 13. New assets Question: Are new assets introduced in the paper well documented and is the documentation provided alongside the assets? Answer: [Yes] Justification: Guidelines: ⢠The answer [N/A] means that the paper does not release new assets. â˘Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates. This includes details about training, license, limitations, etc. ⢠The paper should discuss whether and how consent was obtained from people whose asset is used. â˘At submission time, remember to anonymize your assets (if applicable). You can either create an anonymized URL or include an anonymized zip file. 14. Crowdsourcing and research with human subjects Question: For crowdsourcing experiments and research with human subjects, does the paper include the full text of instructions given to participants and screenshots, if applicable, as well as details about compensation (if any)? Answer: [N/A] Justification: Guidelines: â˘The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects. â˘Including this information in the supplemental material is fine, but if the main contribu- tion of the paper involves human subjects, then as much detail as possible should be included in the main paper. â˘According to the NeurIPS Code of Ethics, workers involved in data collection, curation, or other labor should be paid at least the minimum wage in the country of the data collector. 15.Institutional review board (IRB) approvals or equivalent for research with human subjects Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals (or an equivalent approval/review based on the requirements of your country or institution) were obtained? Answer: [N/A] Justification: Our experiment is conducted in a simulation without human subjects. So, our research does not need for IRB approval. Guidelines: â˘The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects. 42 â˘Depending on the country in which research is conducted, IRB approval (or equivalent) may be required for any human subjects research. If you obtained IRB approval, you should clearly state this in the paper. ⢠We recognize that the procedures for this may vary significantly between institutions and locations, and we expect authors to adhere to the NeurIPS Code of Ethics and the guidelines for their institution. ⢠For initial submissions, do not include any information that would break anonymity (if applicable), such as the institution conducting the review. 16. Declaration of LLM usage Question: Does the paper describe the usage of LLMs if it is an important, original, or non-standard component of the core methods in this research? Note that if the LLM is used only for writing, editing, or formatting purposes and does not impact the core methodology, scientific rigor, or originality of the research, declaration is not required. Answer: [N/A] Justification: Guidelines: â˘The answer [N/A] means that the core method development in this research does not involve LLMs as any important, original, or non-standard components. â˘Please refer to our LLM policy in the NeurIPS handbook for what should or should not be described. 43