Paper deep dive
NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving
Ashkan Yousefi Zadeh, Zishuo Zhu, Xiaomeng Li, Andry Rakotonirainy, Sebastien Glaser, Ronald Schroeter, Patricia Delhomme, Zahra Mehraban
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/18/2026, 4:36:50 AM
Summary
The paper introduces NARRATE, a multimodal real-world Australian driving dataset designed to support human-centred explanations in automated driving. It comprises 2,050 annotated events from 35 experienced drivers and instructors, featuring synchronized visual, LiDAR, and motion data paired with in-vehicle and post-drive free-text explanations. The dataset includes driver-action labels, scenario-context labels, and span-level Situational Awareness (SA) annotations (Perception, Comprehension, Projection). Benchmark tasks demonstrate that while the structure is learnable, fine-grained context recognition and explanation generation remain challenging.
Entities (10)
Relation Signals (8)
NARRATE → collectedfrom → 35 experienced drivers
confidence 100% · comprising 2,050 annotated events from 35 experienced drivers and driving instructors on public roads.
NARRATE → contains → 2,050 annotated events
confidence 100% · NARRATE, a multimodal real-world Australian driving dataset comprising 2,050 annotated events from 35 experienced drivers
NARRATE → locatedin → Brisbane
confidence 95% · public roads in Brisbane, Queensland
NARRATE → supports → Human-Centred Explanations
confidence 95% · NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving
NARRATE → uses → Situational Awareness annotations
confidence 95% · NARRATE provides ... span-level Situational Awareness (SA) annotations over driver explanations for Perception, Comprehension and Projection.
NARRATE → comparedwith → nuScenes
confidence 90% · nuScenes [6] established a widely used multi-sensor benchmark... Table 1: Comparison with related datasets.
NARRATE → comparedwith → BDD100K
confidence 90% · BDD100K [47] demonstrated the value of large-scale front-view video... Table 1: Comparison with related datasets.
Situational Awareness → definedby →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather than elicited from the driver performing the action. We introduce NARRATE, a multimodal real-world Australian driving dataset comprising 2,050 annotated events from 35 experienced drivers and driving instructors on public roads. Each event is grounded in synchronised visual, localisation, motion, and LiDAR streams and paired with in-vehicle and/or post-drive free-text explanations. NARRATE provides action labels, scenario-context labels spanning six high-level and 32 fine-grained categories, and span-level Situational Awareness (SA) annotations over driver explanations for Perception, Comprehension and Projection. Four benchmark tasks (SA, scenario-context, driver-action classification, and explanation generation) show that this structure is learnable from driver language, while fine-grained context recognition and explanation generation remain challenging. NARRATE paves a path towards more human-centred and domain-aware explanation models for automated driving.
Tags
Links
- Source: https://arxiv.org/abs/2608.14767v1
- Canonical: https://arxiv.org/abs/2608.14767v1
Trouble viewing inline? Open PDF directly →
Full Text
51,458 characters extracted from source content.
Expand or collapse full text
NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving Ashkan Yousefi Zadeh Affiliation: ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3), Australia Zishuo Zhu Affiliation: ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3), Australia Xiaomeng Li Affiliation: ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3), Australia Andry Rakotonirainy Affiliation: ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3), Australia Sebastien Glaser Affiliation: ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3), Australia Ronald Schroeter Affiliation: ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3), Australia Patricia Delhomme Affiliation: Université Gustave Eiffel, Laboratory of Applied Psychology and Ergonomics, France Zahra Mehraban Affiliation: ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3), Australia [4pt] Queensland University of Technology (QUT) Faculty of HealthSchool of Psychology and Counselling, Australia Abstract Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather than elicited from the driver performing the action. We introduce NARRATE, a multimodal real-world Australian driving dataset comprising 2,050 annotated events from 35 experienced drivers and driving instructors on public roads. Each event is grounded in synchronised visual, localisation, motion, and LiDAR streams and paired with in-vehicle and/or post-drive free-text explanations. NARRATE provides action labels, scenario-context labels spanning six high-level and 32 fine-grained categories, and span-level Situational Awareness (SA) annotations over driver explanations for Perception, Comprehension and Projection. Four benchmark tasks (SA, scenario-context, driver-action classification, and explanation generation) show that this structure is learnable from driver language, while fine-grained context recognition and explanation generation remain challenging. NARRATE paves a path towards more human-centred and domain-aware explanation models for automated driving. Keywords Multimodal Driving Dataset ⋅· Human-Centred Explanations ⋅· Situational Awareness 1 Introduction As automated vehicles (AVs) move from controlled test tracks to public roads, explaining driving decisions becomes central to safe and trustworthy human–vehicle interaction [18, 50]. Passengers need to understand not only what the vehicle is doing, but why a manoeuvre is appropriate in the current road situation [18, 7]. Explainable AI (XAI) can make model decisions more interpretable and accountable [2], but explanation quality should not be defined only by researchers’ assumptions about what is informative. It should also be grounded in how humans naturally explain events [30], particularly in automated driving, where explanation content, timing, modality, and context affect comprehension, workload, and trust [18, 50, 46]. Despite rapid progress in driving datasets paired with natural-language annotations, most do not capture explanations from the person who performed the driving action. Instead, the accompanying language is often written post-hoc by observers, generated or reconstructed from sensor evidence, structured as question–answer pairs, or collected in simulation [17, 45, 27, 40, 31, 3, 42, 25]. These resources are valuable, but primarily reconstruct plausible reasons from an external viewpoint. They leave open a basic empirical question: how do experienced drivers explain their own decisions in real traffic, and what situational information do they include? In this paper, human-centred explanations originate from the driver performing the task, rather than from external observers or post-hoc annotators. Figure 1: NARRATE data collection protocol: (A) instrumented vehicle and in-cabin setup; (B) real-time explanation elicitation and event tagging during the drive; (C) post-drive video-cued interview using the synchronised multi-view replay dashboard. Driver explanations can go beyond describing an action: drivers may refer to relevant cues, interpret their meaning for the current manoeuvre, and anticipate how the situation may unfold. This corresponds closely to Endsley’s Situational Awareness (SA) framework, which distinguishes Perception, Comprehension, and Projection in human performance in dynamic systems [12, 11]. We do not assume that every effective explanation must contain all three SA levels; rather, SA provides a cognitively grounded lens for analysing which aspects of situational reasoning are expressed in driver-produced explanations, complementing work showing that explanations can support understanding, trust calibration, and situation awareness in automated driving [18, 7, 50]. Explanation timing also matters: in-vehicle explanations capture what is salient under real-time workload, whereas post-drive video-cued reflections allow drivers to elaborate after the event. A dataset for human-centred driving explanations should capture both, rather than treating driver reasoning as a single static text label. Figure 2: Representative events from NARRATE. Each panel shows an annotated driving event with its driver-action label, high-level and fine-grained scenario-context tags, three front-centre camera frames sampled around the tagged event time, and the corresponding driver explanations. The middle frame marks the tagged event, while the neighbouring frames provide pre- and post-event visual context. In-vehicle explanations were recorded during driving, whereas post-drive explanations were recalled during the video-cued debrief. A further gap is geographic and operational coverage. Most large driving datasets are collected in North American, European, Chinese, or simulated settings, whereas Australian driving involves left-hand traffic, right-hand-drive vehicles, domain-specific road rules, signage, lane-use conventions, and road-user expectations. Recent left-hand-driving domain adaptation work shows that transferring autonomous-steering models from U.S. right-hand-driving data to real-world Australian data requires explicit treatment of domain shift [28]. Explanation datasets should therefore be domain-aware: an AV should explain its behaviour using road conventions and situational cues from its operating environment. To address these gaps, we introduce NARRATE: Naturalistic Action Reasoning and Real-World Awareness Through Explanations. NARRATE is a multimodal real-world Australian driving dataset designed to help train and evaluate automated driving explanation models that produce human-centred explanations by learning from experienced drivers’ own decision explanations. It comprises 2,050 annotated events from 35 experienced drivers and driving instructors on public roads in Brisbane, Queensland, paired with in-vehicle and/or post-drive free-text explanations collected through the dual-timing protocol in Fig. 1. NARRATE provides driver-action labels, scenario-context labels, span-level SA annotations, and benchmark tasks; representative annotated events are shown in Fig. 2. Our contributions are threefold: (i) a human-centred dataset of real-world driving explanations produced by experienced drivers during and after driving. (i) Multimodal event grounding with driver-action labels, scenario-context labels across six high-level and 32 fine-grained categories, and span-level SA labels for L1 Perception, L2 Comprehension, and L3 Projection. (i) Participant-disjoint benchmarks for SA classification, context classification, driver-action classification, and explanation generation. NARRATE is not a fleet-scale perception dataset, but a high-fidelity explanation dataset for studying driver-produced explanations of real-world driving decisions. 2 Related Work Large-scale driving datasets have driven substantial progress in AV perception, localisation, prediction, and planning. nuScenes [6] established a widely used multi-sensor benchmark combining multi-camera video, LiDAR, radar, GPS, and IMU, while BDD100K [47] demonstrated the value of large-scale front-view video for diverse visual driving tasks. Behaviour-oriented datasets such as HDD [38] go beyond perception by annotating driver actions, goals, causes, and attention-related labels in naturalistic driving. However, these datasets provide perceptual or behavioural supervision rather than natural-language explanations from the driver who performed the manoeuvre. Table 1: Comparison with related datasets. Columns: Cam. = number of cameras; Lang. = language format; SA = Situational Awareness; Act. = action labels; Ctx. = context labels; Size = headline annotated count (unit as reported); Hrs = hours of driving/video; Drv = distinct human drivers. Sensors: V = video; L = LiDAR; I = IMU; G = GPS; C = CAN; Ra = radar; D = dynamics; Sim. = simulator; var. = variable; mul. = multiple. Language: FT = free text; QA = question–answer; Cmd. = commands; Cat. = categories; Lbl = labels; CF = counterfactual; CoC = chain of causation. Scale: n/r = not reported; n/a = not applicable (built on another dataset or simulated); crowd = crowdsourced (no fixed driver count); der. = derived by us; seg. = segments; scn. = scenes; sess. = sessions. ‡ Reasoning structure, not Endsley L1/L2/L3. ∗* Causal/scenario labels only. † Instruction or advice, not driver explanation. a Alpamayo hours n/r for the 700K CoC set; the paper’s 80,000 h is the fleet pre-training pool, not annotated reasoning hours. b CoVLA also stated as 6M frames / 431K sampled trajectory points; 83.3 h == 30 s × 10,000. c HAD: 45,629 advice annotations (25,549 action / 20,080 attention) over the clips. d NARRATE Hrs given as annotated / recorded: 8.5 h annotated event footage (2,050 × 15 s) and ∼ 23 h total driving recorded (35 × 40 min). Dataset Cam. Sensors Source Lang. SA Act. Ctx. Size Hrs Drv Perception & behaviour nuScenes [6] 6 V,L,Ra,G,I None – ✗ ✗ ✗ 1,000 scn. 5.5 n/r BDD100K [47] 1 V,G,I None – ✗ ✗ ✗ 100K videos ∼ 1,100 der. crowd HDD [38] 3 V,L,I,G,C Labels only Lbl ✗ ✓ ✓∗ 137 sess. 104 n/r Language & explanation Talk2Car [9] 6 V,L,Ra,G Human commands Cmd.† ✗ ✓ ✗ 11,959 cmds. n/a n/a HAD [16] 3 V,C Human advice FT† ✗ ✓ ✓∗ 5,675 clipsc ∼ 32 n/a BDD-X [17] 1 V,G,D Crowd, post-hoc FT ✗ ✓ ✗ 26,228 annot. ∼ 77 n/a BDD-OIA [45] 1 V Crowd, post-hoc Cat. ✗ ✓ ✓∗ 22,924 clips ∼ 32 der. n/a LingoQA [27] 1 V Human + GPT QA ✗ ✓ ✗ 419.9K QA n/r n/r DriveLM [40] 6 V,L,Ra Semi-auto + human Graph QA ✗‡ ✓ ✗ ∼ 1.6M QA n/r n/a Reason2Drive [31] var. V,L Auto + GPT-4 Chain QA ✗‡ ✓ ✗ 632,955 QA n/r n/a CoVLA [3] 1 V,I,G,C,D Auto-generated FT ✗ ✓ ✓∗ 10,000 scn.b 83.3 n/r OmniDrive [42] 6 V,L,Ra Rule+GPT+human QA+CF ✗‡ ✓ ✗ nuScenes keyfr. n/r n/a LaMPilot [25] – Sim. Semi-human Cmd.† ✗ ✓ ✓∗ 4,900 scn. n/r n/a Alpamayo-R1 [43] mul. V,D Auto + human CoC‡ ✗‡ ✓ ✓∗ 700K seg.a n/ra n/r NARRATE (ours) 4 V×4,L,I,G,D Experienced drivers, dual-timing FT (dual) ✓ ✓ ✓ 2,050 events 8.5 / 23d 35 A second line of work introduces language into driving, but with different objectives and language sources. Talk2Car [9] grounds passenger commands to objects in driving scenes, HAD [16] provides human advice for driving videos, and LaMPilot [25] studies language-program instruction following in simulation. Explanation datasets such as BDD-X [17] and BDD-OIA [45] link actions to explanatory text or categories, but use post-hoc observer explanations rather than driver-produced ones. Recent reasoning datasets, including LingoQA [27], DriveLM [40], Reason2Drive [31], OmniDrive [42], CoVLA [3], and Alpamayo-R1 [43], advance language-based driving reasoning through QA, counterfactual, causal, or generated supervision. Yet their language is typically generated, reconstructed, or annotated after the fact from sensor evidence, not elicited from the person who made the driving decision. Table 1 summarises these differences in sensor modalities, language source, annotation format, and action, context, and SA supervision. NARRATE combines real-world multi-sensor grounding, experienced-driver free-text explanations, action and scenario-context labels, and explicit span-level SA annotation. Building on the SA motivation above, NARRATE also differs in how it represents situational reasoning. Originally developed for human performance in dynamic environments [12], SA is relevant to automated driving because reduced driver engagement can weaken situation awareness and takeover readiness [13, 29]. It is central to human–vehicle interfaces that communicate what the vehicle perceives, how it interprets the scene, and what risks may follow [7, 44], and is linked to trust calibration [34]. Existing driving reasoning datasets may include perception-, prediction-, planning-, or causal-reasoning structures, but not span-level Endsley L1/L2/L3 annotation grounded in driver-produced language. NARRATE complements them with span-level SA labels over experienced-driver explanations, enabling models to learn both what drivers do and how their explanations express perception, comprehension, and projection. 3 Data Collection Instrumented vehicle. Data were collected using a Kia Niro SUV with a roof-mounted sensor rig and onboard edge-compute unit running ROS 2 (Humble) [26] (Fig. 1). All streams were recorded as ROS 2 bag files and time-synchronised via GPS pulse-per-second (PPS) through the inertial navigation system. For each event, the platform captured four lossless camera views at 10 fps: front-centre, forward-left, forward-right, and rear-centre, yielding 8,200 camera sequences across 2,050 events. The vehicle also recorded a Leishen CH128X forward-facing LiDAR (120∘×25∘120 ×25 field of view, 160m range, 128 channels), GPS-aided INS/IMU streams with PPS synchronisation, GNSS/GPS position, vehicle speed, and longitudinal acceleration. Participants. Thirty-seven participants were recruited. Two sessions were excluded before dataset construction. One lacked recoverable event timestamps needed to extract time-aligned sensor clips and explanations, and the other involved non-protocol-compliant driving. The final analysed sample comprised 35 participants: 30 experienced non-instructor drivers and 5 licensed driving instructors. All included participants were at least 30 years old, held a Queensland open driver’s licence for at least three years, were familiar with Brisbane CBD, and reported driving at least two hours per week. Driving instructors additionally held a valid instructor licence for at least one year, and individuals involved in an at-fault crash within the previous two years were excluded. The final sample comprised 24 male and 11 female participants, with a mean age of 44.5 years (SD=12.6; range 30–81) and a mean of 23.3 years of driving experience (SD=14.4; range 3–64). All participants provided written informed consent. Driving protocol. Each session comprised approximately 40 minutes of naturalistic driving on public roads in Brisbane. Participants followed a predefined route on an in-vehicle navigation display; before driving, they were briefed on the route, instructed to drive normally while obeying road rules, and reminded to prioritise safe vehicle control. This standardised exposure to the same road infrastructure while preserving natural variation in traffic, road-user behaviour, and signal timing. Route composition and road-feature statistics are reported in Sec. 5. Data were collected between approximately 7 am and 5 pm, under mostly sunny conditions with some cloudy and rainy sessions. Driving was manual, without Advanced Driver Assistance Systems (ADAS) intervention, in a left-hand-traffic, right-hand-drive Australian road context. Three people occupied the vehicle: the participant driver, a front-seat researcher, and a rear-seat data engineer. The engineer monitored sensor recording and tagged candidate events in real time; the researcher elicited in-vehicle explanations only after completed actions and only when safe. Explanation elicitation. NARRATE was designed to capture both immediate and reflective driver explanations. During the drive, events arose from participant think-aloud comments (P), researcher-triggered prompts after completed actions (R), or engineer-flagged deferrals for events unsafe or impractical to discuss while driving (E). For P and R events, the engineer marked whether the in-vehicle explanation was sufficiently informative: short or insufficient explanations were marked S, and informative ones L. Events coded P+S or R+S, together with all E-tagged events, were shown to the participant during the post-drive interview to elicit or clarify the explanation; P+L and R+L events were not shown again. After the drive, participants completed a video-cued semi-structured interview using a purpose-built multi-view replay dashboard. Each event was represented by a 15-second synchronised four-view clip, with the tagged manoeuvre at approximately 10 seconds, providing roughly 10 seconds of pre-event context and 5 seconds of post-event outcome. For each post-drive event, participants were asked up to four pre-specified questions: what action they took, why they took it, what they perceived, and what could have happened otherwise. When a driver did not perform an expected manoeuvre, questions used a why-not framing. These omitted-manoeuvre cases are preserved as distinct non-action classes: not lane change, not stop, not speed up, and not slow down, not conflated with counterfactual responses. Explanation audio was transcribed using automatic speech recognition, manually corrected, and checked during annotation. 4 Annotation Design NARRATE provides two complementary annotation layers: span-level Situational Awareness (SA) labels over driver explanations and event-level scenario-context labels, describing both the reasoning expressed in language and the driving situation. Situational Awareness annotation. SA labels follow Endsley’s three-level framework [12]: L1 Perception denotes information the driver noticed, L2 Comprehension its driving relevance, and L3 Projection anticipation of a future state or risk. Annotators labelled all SA levels expressed in each explanation and highlighted supporting spans. In-vehicle and post-drive explanations were annotated independently, and a single explanation could contain multiple SA levels. For example, in “The light ahead was turning amber [L1], so I knew I needed to slow down [L2], because there was a car close behind me that might not stop in time [L3]”, all three levels appear in one explanation. This span-level design identifies whether an SA level is present and where perceptual, interpretive, and anticipatory reasoning appears in driver-produced language. Scenario-context annotation. Context labels were assigned at the event level using a driving-scenario taxonomy from prior work [48]. Each event could receive multiple fine-grained labels from 32 scenario categories grouped into six high-level contexts: Traffic Compliance, Social Interaction and Traffic Flow, Navigation and Routing, Hazard and Obstacle Management, Special Zones and Stops, and Environmental and Adaptation. Context was annotated at the event level rather than the span level because it may depend on the explanation, visual scene, manoeuvre, and surrounding traffic situation. Annotators and label resolution. Three experienced annotators (A, B, C), with complementary expertise spanning automated driving research, natural language processing, computer vision, and human factors, contributed to the annotation process. Annotator A labelled the full dataset for both SA and context. To estimate reliability, a 303-instance subset (covering approximately 15% of annotated events and all 35 participants) was selected by stratified sampling to match the full-dataset distribution across high-level context categories and train/validation/test splits [19] (Fig. 3(g)), ensuring that agreement estimates were representative of the full corpus rather than an artefact of sample composition. Annotators B and C independently labelled SA on this subset, and Annotator C additionally labelled context. All annotators followed the same guideline document, worked independently in Label Studio [41], and were blinded to each other’s labels. They could view the synchronised event video alongside the explanation text, so labels were grounded in both language and driving context. For the reliability subset, L1 and L2 SA labels were resolved by majority vote across Annotators A, B, and C. L3 SA labels and context labels were resolved by consensus between Annotators A and C, as Annotator B’s markedly lower L3 positive rate (13.2%) relative to A (71.9%) and C (59.7%) indicated a systematically different interpretation of projection that would bias majority-vote outcomes at that level. For the remaining 85% of events, the released labels follow Annotator A’s guideline-based annotations; the substantial A–C agreement at L3 (κ=0.690κ=0.690) reported in Sec. 5 supports the reliability of those labels. Agreement results are reported in Sec. 5. 5 Dataset Statistics Scale, splits, and sensor coverage. From 2,138 raw event tags, 85 events were excluded because the event tag was erroneous (e.g., an accidental or incorrectly timed tag), the response was irrelevant, the participant could not recall the event, or the event could not be reliably located in the transcript. A further 3 events had no valid explanation, yielding 2,050 annotated events from 35 participants. Events were partitioned into participant-disjoint train, validation, and test splits using stratified group splitting to balance action, context, and SA label distributions. Table 2 reports the resulting split sizes. Table 2: NARRATE dataset splits. IV = in-vehicle; Post = post-drive. Component Train Val Test Total Participants 24 5 6 35 Events 1,402 272 376 2,050 IV explanations 875 171 227 1,273 Post explanations 750 155 205 1,110 Dual (both) 223 54 56 333 IV words (mean± ) 17.0±11.417.0±11.4 11.8±9.511.8±9.5 12.9±7.912.9±7.9 15.6±10.815.6±10.8 Post words (mean± ) 26.5±24.326.5±24.3 21.2±15.221.2±15.2 18.7±13.418.7±13.4 24.3±21.824.3±21.8 Route, labels, and language statistics. Fig. 3 summarises route characteristics, label distributions, split design, reliability-sample representativeness, and explanation-language statistics. The predefined 14.5 km Brisbane route covers major arterial roads, motorway/freeway segments, collector roads, and local streets, with speed zones dominated by 60 km/h and 50 km/h sections. The route also includes common traffic and safety features such as pedestrian crossings, signalised and unsignalised junctions, standalone signs, speed bumps, and stop/give-way signs. Raw transcriptions underwent deterministic rule-based cleaning, including filler removal, capitalisation, punctuation correction, and spell-checking with domain-term protection. Overall, 95.5% of explanation fields required no correction, and no semantic content was altered. Post-drive explanations are significantly longer than in-vehicle explanations (median 19 vs. 13 words. Mann–Whitney U=521,171U=521,171, p<0.001p<0.001), consistent with post-drive reflection allowing more elaboration than real-time narration. Content-word differences in Fig. 3(i) further show that in-vehicle explanations emphasise immediate spatial and control terms, whereas post-drive explanations more often include referential or scene-reconstruction terms. (a) Traffic/safety features (b) Top 10 fine-grained context (c) Speed zones (d) Driver action (e) Road type (f) Explanation length (g) Reliability sample (h) Dataset splits (i) Top 20 content words Figure 3: Dataset statistics and distributions. Panels summarise the 14.5 km route, action and context labels, train/validation/test splits, reliability-sample representativeness, explanation length, and content-word differences between in-vehicle and post-drive explanations. Action, context, and SA distributions. NARRATE reflects the long-tailed structure of naturalistic driving. Driver actions are imbalanced (Fig. 3(d)): slow down is the most frequent action (976 events; 47.6%), followed by lane change (451; 22.0%) and speed up (187; 9.1%). The four non-action classes collectively represent 9.0% of events. Context labels show a similar long-tailed pattern (Fig. 3(b)), with the most frequent fine-grained categories being vehicle following (430; 21.0%), route-preparation lane change (372; 18.1%), speed-limit adherence (301; 14.7%), and traffic-signal compliance (232; 11.3%). Together, these top four categories account for 65.1% of events. SA labels are common but not uniform across levels. L2 Comprehension has the highest coverage in both conditions (IV 93.6%, Post 91.3%), followed by L1 Perception (IV 89.0%, Post 83.3%). L3 Projection is less frequent but appears at similar rates in both explanation timings (IV 69.1%, Post 70.2%), indicating that anticipatory reasoning is present even in concise in-vehicle explanations. All three levels co-occur in 59.1% of in-vehicle and 58.6% of post-drive explanations. Explanation trigger source. The final event sources were participant think-aloud (804 events), researcher-triggered prompts (385 events), and engineer-flagged deferrals (861 events). Participant think-aloud and researcher-triggered prompts yielded near-complete in-vehicle narration rates (97.4% and 97.7%, respectively). In contrast, engineer-flagged deferral events produced in-vehicle explanations for only 13.2% of rows but post-drive explanations for 95.9%, reflecting their intended role in capturing events that were unsafe or impractical to discuss during active driving. Table 3: Inter-annotator agreement on the 303-instance reliability subset. CI = bootstrap 95% confidence interval (1,000 resamples); Obs. = observed agreement; Fine/High = fine-/high-level context labels; Exact = exact-match agreement; Jaccard and αMASI _MASI measure multi-label overlap. SA uses annotators A, B, C; context uses A and C. †Prevalence paradox: all annotators >82%>82\% positive, so observed agreement is informative. SA annotation (A, B, C) Three-way Pairwise Level Fleiss κ 95% CI A–C κ A–B κ Obs. (A–C) L1 Perception 0.580 [.49, .67] 0.820 [.70,.91] 0.545 [.38,.70] 0.964 L2 Comp.† 0.050 [-.07,.17] 0.304 [.00,.55] 0.006 [-.08,.10] 0.957 L3 Projection 0.203 [.13, .27] 0.690 [.61,.77] 0.112 [.08,.15] 0.858 Combination — — A–C αMASI=0.660[.58,.73] _MASI=0.660\;[.58,.73] — Context annotation (A, C) Fine High Exact % 82.2 85.8 Jaccard 0.892 0.911 αMASI _MASI 0.852 0.861 95% CI [.81,.89] [.82,.90] Inter-annotator agreement. Table 3 reports agreement on the 303-instance reliability subset (maximum deviation 3.4 p from the full-dataset distribution; Fig. 3(g)). Context labels are robust across both granularities (Krippendorff’s α with MASI distance [33, 4]). For SA, we report Fleiss’ κ [15] as the primary three-way metric; following Landis and Koch [20], L1 Perception shows moderate-to-substantial three-way agreement. L3 Projection is lower, driven by Annotator B’s conservative L3 usage (13.2% positive vs. 71.9% and 59.7% for A and C), while A–C pairwise agreement remains substantial [4]. L2 is dominated by the prevalence paradox [14], so observed three-way agreement (79.2%) is the relevant metric. About 85% of released labels are single-annotator (A), with reliability estimated on this subset and supported by substantial A–C agreement at L3. 6 Benchmarks We define four participant-disjoint baseline tasks to characterise what NARRATE enables and where it remains challenging: Situational Awareness (SA) classification, scenario-context classification, driver-action classification, and explanation generation. These benchmarks support dataset use and comparison rather than proposing a new model, and use the splits in Table 2. Unless otherwise stated, text models use a maximum sequence length of 128, batch size 16, AdamW [24] with learning rate 2×10−52×10^-5 and weight decay 0.01, 5 epochs, linear warmup, and validation macro-F1 for model selection. All trainable baselines were run with five random seeds (1, 2, 3, 4, 42), and results are reported as mean ± standard deviation. Visual features are extracted using frozen CLIP ViT-B/32 [35]: eight frames are sampled uniformly from each 15-second event clip, embedded into 512-dimensional features, and mean-pooled to one event representation. Kinematic features are sampled at the same eight temporal indices from vehicle speed and longitudinal acceleration, yielding a 16-dimensional event vector. 6.1 T1: Situational Awareness Classification T1 evaluates whether the SA structure annotated in NARRATE is recoverable from driver-produced language. The task is framed as three independent binary classifications for L1 Perception, L2 Comprehension, and L3 Projection, evaluated separately for in-vehicle and post-drive explanations. Models use BCEWithLogitsLoss with threshold 0.5. The primary metric is macro-F1 across the three SA levels. Baselines include majority class, TF-IDF+LR, DistilBERT [39], BERT-base [10], and RoBERTa-base [23]. Table 4 shows that the majority baseline is already strong because L1 and L2 have high positive-label prevalence. This indicates that perception and comprehension are frequently expressed in driver explanations, but it also makes these labels less discriminative. L3 Projection provides a more informative test case: it is less frequent and shows clearer model gains over the majority baseline. The most informative result is on L3, where RoBERTa-base beats the majority baseline by the largest margin (0.863 vs. 0.773). The high scores on L1 and L2 mostly reflect their high label prevalence rather than learned signal (RoBERTa 0.914±.007 in-vehicle). For post-drive explanations, DistilBERT gives the highest reported macro-F1 (0.914±.007). Across models, seed variance is low (std ≤.008≤.008), supporting the view that SA structure is a stable and learnable signal in NARRATE. 6.2 T2: Driving-Context Classification T2 evaluates whether scenario context can be inferred from explanation text. We consider two granularities: six high-level multi-label categories for in-vehicle and post-drive explanations, and 32 fine-grained multi-label categories using primary_text. Here, primary_text denotes the in-vehicle explanation when available, and otherwise the post-drive explanation. For the fine-grained setting, macro-F1 is computed over labels with at least five test instances. The 32-class task is evaluated with BERT-base and RoBERTa-base only. As shown in Table 5, high-level context is partially recoverable from explanation text. RoBERTa-base performs best at the six-class level for both in-vehicle and post-drive explanations, reaching 0.407±.023 and 0.418±.004 macro-F1, respectively. The gap between macro-F1 and weighted-F1 reflects the long-tailed label distribution: common contexts are more learnable, while rare categories remain difficult. The 32-class fine-grained task is substantially harder. RoBERTa-base reaches only 0.035±.011 macro-F1 and 0.100±.048 weighted-F1, while BERT-base collapses to near-zero performance on one seed. The near-floor macro-F1 on fine-grained context is itself an informative result: it establishes that 32-class context recognition from explanation text alone is not a solved problem with current text encoders, and that richer grounding in visual scene, map topology, and interaction cues will be necessary, making NARRATE a useful dataset for future multimodal context modelling. Table 4: T1: Situational Awareness classification on the participant-disjoint test set. MF1 = Macro-F1; ↑ higher is better. Trainable models report mean ± std over 5 seeds. In-vehicle Post-drive Model L1 L2 L3 MF1 L1 L2 L3 MF1 Majority 0.906 0.964 0.773 0.881 0.904 0.946 0.801 0.884 TF-IDF+LR 0.908 0.964 0.811 0.894 0.916 0.949 0.857 0.907 DistilBERT 0.906 0.964 0.846 0.905±.004 0.915 0.946 0.882 0.914±.007 BERT-base 0.917 0.963 0.854 0.911±.006 0.915 0.949 0.878 0.914±.005 RoBERTa-base 0.915 0.964 0.863 0.914±.007 0.924 0.949 0.861 0.911±.008 Table 5: T2: Context classification on the participant-disjoint test set. MF1 = Macro-F1; WF1 = Weighted-F1; ↑ higher is better. The 32-class MF1 is computed over labels with at least five test instances. Trainable models report mean ± std over 5 seeds. 6-class IV 6-class Post 32-class Primary Model MF1 WF1 MF1 WF1 MF1 WF1 Majority 0.000 −- 0.000 −- −- −- TF-IDF+LR 0.342 −- 0.377 −- −- −- DistilBERT 0.315±.025 0.562±.040 0.309±.048 0.516±.076 −- −- BERT-base 0.289±.045 0.520±.072 0.317±.049 0.534±.079 0.038±.023 0.086±.061 RoBERTa-base 0.407±.023 0.697±.023 0.418±.004 0.703±.008 0.035±.011 0.100±.048 Table 6: T3: Driver-action classification on the participant-disjoint test set (n=376n=376). Acc = Accuracy; MF1 = Macro-F1; ↑ higher is better. Trainable models report mean ± std over 5 seeds. Model Modality Acc MF1 Majority −- 0.476 0.065 TF-IDF+LR Text 0.681 0.228 CLIP+LR Video 0.468 0.219 Kinematics MLP Motion 0.586±.007 0.241±.007 CLIP+Kin MLP Video+Motion 0.629±.015 0.306±.016 DistilBERT Text 0.708±.018 0.246±.024 BERT-base Text 0.690±.034 0.228±.028 RoBERTa-base Text 0.728±.009 0.289±.018 BERT+Kin Text+Motion 0.725±.011 0.280±.007 Table 7: T4: Explanation generation on the participant-disjoint test set (n=376n=376). B4 = BLEU-4; RL = ROUGE-L; MET = METEOR; BS = BERTScore-F1; ↑ higher is better. Fine-tuned models report mean ± std over 5 seeds. Model Type B4 RL MET BS Random Retrieval 0.010±.001 0.115±.005 0.091±.004 0.872±.001 Action Retrieval 0.012±.004 0.129±.028 0.115±.047 0.872±.003 Action+Context Retrieval 0.019±.007 0.158±.016 0.133±.021 0.879±.005 CLIP-N Retrieval 0.012 0.132 0.108 0.875 CLIP+Kin-N Retrieval 0.013 0.142 0.110 0.876 FLAN-T5-base Fine-tuned 0.034±.001 0.215±.010 0.144±.019 0.892±.001 T5-base Fine-tuned 0.036±.002 0.238±.010 0.176±.010 0.895±.003 BART-base Fine-tuned 0.024±.009 0.169±.033 0.112±.038 0.894±.003 GPT-2 Fine-tuned 0.030±.002 0.236±.013 0.185±.010 0.892±.004 Phi-3-mini 0-shot 0.004 0.063 0.083 0.815 6.3 T3: Driver Action Classification T3 evaluates whether driver actions can be inferred from text, video, motion, or their combinations. The task is a 10-class single-label classification problem, evaluated primarily with macro-F1 and secondarily with accuracy. We compare text baselines, frozen CLIP visual features, kinematic features, and simple fusion baselines: Kinematics MLP uses speed and acceleration, while CLIP+Kin MLP and BERT+Kin concatenate kinematic features with frozen CLIP video features and BERT text embeddings, respectively. Table 6 shows that text is the strongest single modality: RoBERTa-base achieves the highest text-only accuracy (0.728±.009) and macro-F1 (0.289±.018). Kinematic features also provide useful signal, outperforming frozen CLIP video features in macro-F1. The best macro-F1 overall is obtained by the CLIP+Kin MLP (0.306±.016), suggesting that visual and motion cues are complementary for action recognition even when visual features are frozen and simple. In contrast, BERT+Kin achieves strong accuracy and low variance but does not improve macro-F1 over RoBERTa text alone, indicating that simple fusion is not sufficient to resolve rare classes. Overall performance remains limited by class imbalance: dominant actions such as slow down and lane change are easier to classify, while rare and non-action classes drive the macro-F1 shortfall. 6.4 T4: Explanation Generation T4 evaluates the difficulty of generating natural-language driver explanations from structured event information. This task is intended as a structured-input generation baseline, not as an end-to-end multimodal generation model. Each fine-tuned generator conditions on the action label, context category, and kinematic trace. Retrieval baselines provide additional comparisons using random selection, label matching, frozen visual nearest neighbours, and visual+kinematic nearest neighbours. Fine-tuned generators include FLAN-T5-base [8], T5-base [37], BART-base [21], and GPT-2 [36]; Phi-3-mini-4k-instruct [1] is evaluated zero-shot. We evaluate BLEU-4 [32], ROUGE-L [22], METEOR [5], and BERTScore-F1 [49]. Table 7 shows that fine-tuned generators outperform retrieval baselines on most metrics. Among retrieval methods, action+context matching performs best, indicating that symbolic event labels capture more explanation similarity than frozen visual nearest neighbours alone. T5-base achieves the best scores on three metrics: BLEU-4 (0.036±.002), ROUGE-L (0.238±.010), and BERTScore-F1 (0.895±.003). GPT-2 achieves the best METEOR score, with 0.185±.010. BART-base shows higher variance than the other fine-tuned models, indicating sensitivity to random initialisation. Zero-shot Phi-3-mini underperforms all fine-tuned models, suggesting a domain gap between generic instruction following and naturalistic driver explanations. The modest absolute scores are important for interpreting the benchmark. NARRATE explanations are concise, diverse, and context-specific. Multiple valid explanations may describe the same event using different wording or levels of detail. As a result, exact n-gram overlap remains low even when generated explanations are semantically plausible. T4 establishes a text-conditioned lower bound and confirms that naturalistic driver explanation generation remains unsolved under current paradigms, which is an open problem the dataset is designed to support, not one it is expected to solve. 7 Limitations NARRATE is modest in scale, limited to a single daytime route in Brisbane, and subject to long-tailed action and context distributions that leave rare classes thin. The baselines use simplified event-level representations as a deliberate lower bound. Full exploitation of the temporal, LiDAR, and multi-view streams is left to future work. 8 Conclusion We introduced NARRATE, a multimodal real-world Australian driving dataset for human-centred explanations in automated driving. NARRATE contains 2,050 annotated events from 35 experienced drivers and driving instructors on public roads, pairing synchronised visual, LiDAR, localisation, inertial, and kinematic streams with in-vehicle and post-drive explanations produced by the drivers themselves. It provides driver-action labels, scenario-context labels, and span-level Situational Awareness annotations for Perception, Comprehension, and Projection. Baselines show that SA structure is learnable from driver language, while fine-grained context recognition and naturalistic explanation generation remain challenging. NARRATE offers a domain-aware dataset and evaluation testbed for models that reflect how human drivers perceive, interpret, and anticipate driving situations. Data Availability NARRATE is available through mediated access via the QUT Research Data Finder at https://doi.org/10.25912/RDF_1786669427992. Dataset documentation and updates are available at https://github.com/ashkan-zadeh/NARRATE. Acknowledgements This research was supported by the Australian Research Council Discovery Project (DP220102598). References [1] M. Abdin, J. Aneja, H. Awadalla, et al. (2024) Phi-3 technical report: a highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219. Cited by: §6.4. [2] A. Adadi and M. Berrada (2018) Peeking inside the black-box: a survey on explainable artificial intelligence (xai). IEEE access 6, p. 52138–52160. Cited by: §1. [3] H. Arai, K. Miwa, K. Sasaki, K. Watanabe, Y. Yamaguchi, S. Aoki, and I. Yamamoto (2025) Covla: comprehensive vision-language-action dataset for autonomous driving. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), p. 1933–1943. Cited by: §1, Table 1, §2. [4] R. Artstein and M. Poesio (2008) Inter-coder agreement for computational linguistics. Computational Linguistics 34 (4), p. 555–596. Cited by: §5. [5] S. Banerjee and A. Lavie (2005) METEOR: an automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, p. 65–72. Cited by: §6.4. [6] H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom (2020) Nuscenes: a multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 11621–11631. Cited by: Table 1, §2. [7] M. Capallera, L. Angelini, Q. Meteier, O. Abou Khaled, and E. Mugellini (2022) Human-vehicle interaction to support driver’s situation awareness in automated vehicles: a systematic review. IEEE Transactions on Intelligent Vehicles 8 (3), p. 2551–2567. External Links: Document Cited by: §1, §1, §2. [8] H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, et al. (2024) Scaling instruction-finetuned language models. Journal of Machine Learning Research 25 (70), p. 1–53. Cited by: §6.4. [9] T. Deruyttere, S. Vandenhende, D. Grujicic, L. Van Gool, and M. Moens (2019) Talk2Car: taking control of your self-driving car. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, p. 2088–2098. External Links: Document Cited by: Table 1, §2. [10] J. Devlin, M. Chang, K. Lee, and K. Toutanova (2019) Bert: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), p. 4171–4186. Cited by: §6.1. [11] M. R. Endsley and D. J. Garland (Eds.) (2000) Situation awareness analysis and measurement. Lawrence Erlbaum Associates, Mahwah, NJ. Cited by: §1. [12] M. R. Endsley (1995) Toward a theory of situation awareness in dynamic systems. Human Factors 37 (1), p. 32–64. External Links: Document Cited by: §1, §2, §4. [13] M. R. Endsley (2017) From here to autonomy: lessons learned from human–automation research. Human factors 59 (1), p. 5–27. Cited by: §2. [14] A. R. Feinstein and D. V. Cicchetti (1990) High agreement but low kappa: I. the problems of two paradoxes. Journal of Clinical Epidemiology 43 (6), p. 543–549. Cited by: §5. [15] J. L. Fleiss (1971) Measuring nominal scale agreement among many raters.. Psychological bulletin 76 (5), p. 378. Cited by: §5. [16] J. Kim, T. Misu, Y. Chen, A. Tawari, and J. Canny (2019) Grounding human-to-vehicle advice for self-driving vehicles. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 10591–10599. Cited by: Table 1, §2. [17] J. Kim, A. Rohrbach, T. Darrell, J. Canny, and Z. Akata (2018) Textual explanations for self-driving vehicles. In Proceedings of the European Conference on Computer Vision (ECCV), p. 563–578. Cited by: §1, Table 1, §2. [18] J. Koo, J. Kwac, W. Ju, M. Steinert, L. Leifer, and C. Nass (2015) Why did my car just do that? explaining semi-autonomous driving actions to improve driver understanding, trust, and performance. International Journal on Interactive Design and Manufacturing 9 (4), p. 269–275. External Links: Document Cited by: §1, §1. [19] K. Krippendorff (2004) Content analysis: an introduction to its methodology. 2nd edition, Sage Publications, Thousand Oaks, CA. Cited by: §4. [20] J. R. Landis and G. G. Koch (1977) The measurement of observer agreement for categorical data. Biometrics 33 (1), p. 159–174. Cited by: §5. [21] M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer (2020) BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th annual meeting of the association for computational linguistics, p. 7871–7880. Cited by: §6.4. [22] C. Lin (2004) Rouge: a package for automatic evaluation of summaries. In Text summarization branches out, p. 74–81. Cited by: §6.4. [23] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov (2019) Roberta: a robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692. Cited by: §6.1. [24] I. Loshchilov and F. Hutter (2017) Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101. Cited by: §6. [25] Y. Ma, C. Cui, X. Cao, W. Ye, P. Liu, J. Lu, A. Abdelraouf, R. Gupta, K. Han, A. Bera, J. M. Rehg, and Z. Wang (2024) LaMPilot: an open benchmark dataset for autonomous driving with language model programs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 15141–15151. External Links: Document Cited by: §1, Table 1, §2. [26] S. Macenski, T. Foote, B. Gerkey, C. Lalancette, and W. Woodall (2022) Robot operating system 2: design, architecture, and uses in the wild. Science Robotics 7 (66), p. eabm6074. External Links: Document Cited by: §3. [27] A. Marcu, L. Chen, J. Hünermann, A. Karnsund, B. Hanotte, P. Chidananda, S. Nair, V. Badrinarayanan, A. Kendall, J. Shotton, et al. (2024) Lingoqa: visual question answering for autonomous driving. In European Conference on Computer Vision, p. 252–269. Cited by: §1, Table 1, §2. [28] Z. Mehraban, S. Glaser, M. Milford, and R. Schroeter (2025) Saliency-guided domain adaptation for left-hand driving in autonomous steering. In 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 14156–14162. Cited by: §1. [29] N. Merat and A. H. Jamson (2009) Is drivers’ situation awareness influenced by a fully automated driving scenario?. In Human factors, security and safety, Cited by: §2. [30] T. Miller (2019) Explanation in artificial intelligence: insights from the social sciences. Artificial Intelligence 267, p. 1–38. External Links: Document Cited by: §1. [31] M. Nie, R. Peng, C. Wang, X. Cai, J. Han, H. Xu, and L. Zhang (2024) Reason2Drive: towards interpretable and chain-based reasoning for autonomous driving. In Proceedings of the European Conference on Computer Vision (ECCV), p. 292–308. External Links: Document Cited by: §1, Table 1, §2. [32] K. Papineni, S. Roukos, T. Ward, and W. Zhu (2002) Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, p. 311–318. Cited by: §6.4. [33] R. J. Passonneau (2006) Measuring agreement on set-valued items (MASI) for semantic and pragmatic annotation. In Proc. Int. Conf. on Language Resources and Evaluation (LREC), Cited by: §5. [34] L. Petersen, L. Robert, X. J. Yang, and D. M. Tilbury (2019) Situational awareness, drivers trust in automated driving systems and secondary task performance. arXiv preprint arXiv:1903.05251. Cited by: §2. [35] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748–8763. Cited by: §6. [36] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al. (2019) Language models are unsupervised multitask learners. OpenAI blog 1 (8), p. 9. Cited by: §6.4. [37] C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu (2020) Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21 (140), p. 1–67. Cited by: §6.4. [38] V. Ramanishka, Y. Chen, T. Misu, and K. Saenko (2018) Toward driving scene understanding: a dataset for learning driver behavior and causal reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 7699–7707. Cited by: Table 1, §2. [39] V. Sanh, L. Debut, J. Chaumond, and T. Wolf (2019) DistilBERT, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108. Cited by: §6.1. [40] C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beisswenger, P. Luo, A. Geiger, and H. Li (2024) DriveLM: driving with graph visual question answering. In Proceedings of the European Conference on Computer Vision (ECCV), Cited by: §1, Table 1, §2. [41] M. Tkachenko, M. Malyuk, A. Holmanyuk, and N. Liubimov (2020) Label Studio: data labeling software. Note: Open-source software External Links: Link Cited by: §4. [42] S. Wang, Z. Yu, X. Jiang, S. Lan, M. Shi, N. Chang, J. Kautz, Y. Li, and J. M. Alvarez (2025) OmniDrive: a holistic vision-language dataset for autonomous driving with counterfactual reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Cited by: §1, Table 1, §2. [43] Y. Wang, W. Luo, J. Bai, Y. Cao, T. Che, K. Chen, Y. Chen, J. Diamond, Y. Ding, W. Ding, et al. (2025) Alpamayo-r1: bridging reasoning and action prediction for generalizable autonomous driving in the long tail. arXiv preprint arXiv:2511.00088. Cited by: Table 1, §2. [44] H. White, D. R. Large, D. Salanitri, G. Burnett, A. Lawson, and E. Box (2019) Rebuilding drivers’ situation awareness during take-over requests in level 3 automated cars. Contemporary ergonomics and human factors 9. Cited by: §2. [45] Y. Xu, X. Yang, L. Gong, H. Lin, T. Wu, Y. Li, and N. Vasconcelos (2020) Explainable object-induced action decision for autonomous vehicles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 9523–9532. Cited by: §1, Table 1, §2. [46] A. Yousefi Zadeh, X. Li, A. Rakotonirainy, R. Schroeter, and S. Glaser (2025) PsyLingXAV: a psycholinguistics design framework for xai in automated vehicles. In Joint Proceedings of the xAI 2025 Late-breaking Work, Demos and Doctoral Consortium: co-located with the 3rd World Conference on eXplainable Artificial Intelligence (xAI 2025): Istanbul, Turkey, July 09–11, 2025, p. 105–112. Cited by: §1. [47] F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, V. Madhavan, and T. Darrell (2020) Bdd100k: a diverse driving dataset for heterogeneous multitask learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 2636–2645. Cited by: Table 1, §2. [48] A. Y. Zadeh, X. Li, A. Rakotonirainy, R. Schroeter, S. Glaser, and Z. Zhu (2026) X-blocks: linguistic building blocks of natural language explanations for automated vehicles. arXiv preprint arXiv:2602.13248. Cited by: §4. [49] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi (2019) Bertscore: evaluating text generation with bert. arXiv preprint arXiv:1904.09675. Cited by: §6.4. [50] Z. Zhu, X. Li, P. Delhomme, R. Schroeter, S. Glaser, and A. Rakotonirainy (2025) Human-centric explanations for users in automated vehicles: a systematic review. Accident Analysis & Prevention 220, p. 108152. Cited by: §1, §1.