Paper deep dive
From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving
Chaitanya Shinde, Hadi Hajieghrary, Miguel Hurtado
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/12/2026, 2:58:56 AM
Summary
This paper introduces a systematic behavioral taxonomy for Automated Driving Systems (ADS) to bridge the gap between Operational Design Domain (ODD) specifications and behavioral validation. It defines 21 behavioral competencies across Highway (HWY), Urban (URB), and Hub (HUB) domains, derived from the PEGASUS six-layer ODD model. Each behavior is characterized by a four-property framework (Safety, Compliance, Comfort, Efficiency) and decomposed along longitudinal and lateral control axes. The taxonomy aims to support systematic scenario generation and SOTIF coverage evidence, with a specific focus on the underspecified Hub domain.
Entities (17)
Relation Signals (16)
Lane Change → belongstodomain → Highway
confidence 95% · Lane Change HWY Both Safety, Compliance, Comfort
Staging Area Entry and Docking → belongstodomain → Hub
confidence 95% · Staging Area Entry and Docking HUB Both Safety, Efficiency
Unprotected Left Turn → belongstodomain → Urban
confidence 95% · Unprotected Left Turn URB Both Safety, Compliance
Behavioral Competency → characterizedby → Comfort
confidence 95% · Comfort (rider dynamics and trust)
Behavioral Competency → characterizedby → Safety
confidence 95% · characterized against a four-property framework: Safety (gap maintenance, conflict avoidance, kinematic stability)
Behavioral Competency → characterizedby → Compliance
confidence 95% · Compliance (legal rules and behavioral norms)
Behavioral Competency → characterizedby → Efficiency
confidence 95% · Efficiency (mission completion and product-level metrics)
Taxonomy → coversdomain → Urban
confidence 95% · organized across three operational domains-Highway (HWY), Urban (URB), and Hub (HUB)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe what the ADS must demonstrably do once deployed within that domain. This gap between operating condition specification and behavioral validation represents a critical unresolved challenge in ADS safety assurance. This paper presents a structured, standards-grounded taxonomy of 21 behavioral competencies organized across three operational domains-Highway (HWY), Urban (URB), and Hub (HUB)-derived systematically from the PEGASUS six-layer model-based ODD. Each behavior is decomposed along longitudinal and lateral control axes and characterized against a four-property framework: Safety (gap maintenance, conflict avoidance, kinematic stability), Compliance (legal rules and behavioral norms), Comfort (rider dynamics and trust), and Efficiency (mission completion and product-level metrics). We further demonstrate that the crossing of ODD layer parameterizations with behavioral competency specifications yields concrete scenario families suitable for systematic behavioral testing and SOTIF coverage evidence. The taxonomy is grounded in AVSC00008202111, SAE J3237, and SAE J3016, and is validated as an operational specification layer through its deployment in a rule-enforced trajectory optimization system. The Hub domain is identified as a structurally distinct, underspecified domain warranting dedicated research attention.
Tags
Links
- Source: https://arxiv.org/abs/2608.08941v1
- Canonical: https://arxiv.org/abs/2608.08941v1
Trouble viewing inline? Open PDF directly →
Full Text
46,554 characters extracted from source content.
Expand or collapse full text
From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving Chaitanya Shinde1, Hadi Hajieghrary2, Miguel Hurtado3 1Chaitanya Shinde chaitanya.shinde@torc.ai, 2Hadi Hajieghrary hadi.hajieghrary@torc.ai, 3Miguel Hurtado miguel.hurtado@torc.ai 1 Abstract Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe what the ADS must demonstrably do once deployed within that domain. This gap between operating condition specification and behavioral validation represents a critical unresolved challenge in ADS safety assurance. This paper presents a structured, standards-grounded taxonomy of 21 behavioral competencies organized across three operational domains-Highway (HWY), Urban (URB), and Hub (HUB)-derived systematically from the PEGASUS six-layer model-based ODD. Each behavior is decomposed along longitudinal and lateral control axes and characterized against a four-property framework: Safety (gap maintenance, conflict avoidance, kinematic stability), Compliance (legal rules and behavioral norms), Comfort (rider dynamics and trust), and Efficiency (mission completion and product-level metrics). We further demonstrate that the crossing of ODD layer parameterizations with behavioral competency specifications yields concrete scenario families suitable for systematic behavioral testing and SOTIF coverage evidence. The taxonomy is grounded in AVSC00008202111, SAE J3237, and SAE J3016, and is validated as an operational specification layer through its deployment in a rule-enforced trajectory optimization system. The Hub domain is identified as a structurally distinct, underspecified domain warranting dedicated research attention. This paper presents an illustrative behavioral taxonomy developed for research and engineering discussion purposes. It reflects the views of the individual authors and was written based on the state-of-the-art automated driving methodology understood at the time of writing. It is not a complete or comprehensive work, and its contents should be revisited as advances are made in technology, standards, and applicable law. Nothing in this paper is intended to, and nothing herein shall be construed to, establish, create, or represent a minimum or maximum standard of care, a state-of-the-art benchmark, or a complete specification of any product, system, or safety validation process actually deployed by Torc Robotics, Inc. Use of terms such as ‘safety,’ ‘compliance,’ or related terminology refers to defined taxonomic categories used for engineering classification purposes only and does not represent a claim, warranty, or admission regarding the actual safety performance, compliance status, or risk profile of any Torc product or system. 2 Introduction The deployment of automated driving systems (ADS) at commercial scale requires a clear, testable answer to a deceptively simple question: what must the ADS be able to do inside its operational design domain? Operational Design Domain (ODD) specifications-structured via frameworks such as the PEGASUS six-layer model [21]- define the environmental and infrastructural conditions under which an ADS is designed to function. They describe road geometry, traffic infrastructure, moving object classes, environmental conditions, and digital information layers. What they do not, in general, define is the behavioral competency expected of the ADS. This gap has material consequences. Without a systematic enumeration of behavioral competencies, safety validation efforts lack a complete coverage target; behavioral compliance becomes a loosely defined concept varying across teams and jurisdictions; and scenario-based testing lacks a principled generation mechanism. The absence of such a layer is increasingly recognized as a barrier to ADS certification and public trust This paper addresses the gap by presenting a behavioral taxonomy for ADS, grounded in existing standards [4] and structured for practical use by safety, planning, and validation engineers. The taxonomy organizes 21 behaviors across three operationally distinct domains, provides a uniform anatomy for each behavior, and demonstrates a formal mechanism for generating scenario families from the crossing of ODD parameters and behavioral specifications. Note: This behavior taxonomy is just for illustrative purposes and is one of the many ways one can define and group behaviors. 2.1 Paper Objectives This paper pursues three objectives: 1. Derivation: Demonstrate a systematic methodology for deriving operational domains and behaviors from ODD layer composition, using the PEGASUS model as the ODD substrate. 2. Taxonomy: Present a structured, standards-grounded enumeration of 21 behaviors across Highway, Urban, and Hub domains/ use-cases, decomposed along longitudinal/lateral control axes and characterized via a four-property framework. 3. Operationalization: Show that ODD parameter instantiation crossed with behavioral competency specifications yields concrete, testable scenario families aligned with ISO/PAS 21448 (SOTIF) evidence generation. 2.2 Paper Organization Section 2 reviews related work and standards grounding. Section 3 defines the behavioral framework including the four-property characterization. Section 4 presents the full 21-behavior taxonomy across three domains. Section 5 provides deep-dive competency specifications for one behavior per domain. Section 6 formalizes the ODD × Behavior → Scenario mapping. Section 7 discusses the DDT/DDT-F dimension and the Hub domain gap. Section 8 concludes with a summary of contributions and future directions. 3 Background and Related Work 3.1 Standards Foundations The behavioral framing in this paper builds on four standards-track documents. AVSC00008202111 [4] defines a behavioral competency as a demonstrated capability in an ODD context with observable, measurable outcomes. It provides a framework for classifying ADS behaviors and identifying evaluation criteria, and serves as the primary source for the competency structure used in this taxonomy. Applicable Metrics: Dynamic Driving Task Assessment (DA) metrics from SAE J3237 [20] applicable to this behavior, organized by the J3237 taxonomy: Black Box metrics (computable without ADS data), Grey Box metrics (requiring ADS status data), and White Box metrics (requiring full ADS internal data). SAE J3016 [19] provides foundational vocabulary: Dynamic Driving Task (DDT), DDT Fallback (DDT-F), Object and Event Detection and Response (OEDR), and the ODD definition used throughout. ISO/PAS 21448 (SOTIF) [11] defines the safety of the intended functionality framework. The scenario family generation mechanism in Section 6 is explicitly designed to produce SOTIF-aligned coverage evidence. ISO 26262 [10] addresses functional safety of electrical and electronic systems and provides the systematic-fault and random-hardware-fault assurance backbone against which SOTIF and behavioral competency evidence are typically composed in a complete safety case. UL 4600 [25] complements these with a goal-based safety case standard specifically for autonomous products, and ISO/PAS 8800 [12] extends this landscape to safety-related properties of AI/ML elements within road vehicles, addressing insufficiencies not covered by ISO 26262 or SOTIF alone. The taxonomy in this paper is intended to sit alongside, not replace, these standards: it supplies the behavioral competency layer that such safety cases must reference as evidence. 3.2 ODD Modeling The PEGASUS project [21] produced the most widely adopted structured ODD model for ADS scenario-based testing: a six-layer decomposition covering road geometry, road furniture and rules, temporary modifications, moving objects, environmental conditions, and digital information. This model is used directly in Section 3 to derive the three operational domains and in Section 6 as the parameterization substrate for scenario generation. Subsequent work, including the ASAM OpenSCENARIO [3] and ASAM OpenDRIVE formats, operationalize portions of the PEGASUS model in simulation toolchains. Our taxonomy is agnostic to simulation format but maps onto these representations. 3.3 Behavioral Taxonomy and Scenario Generation Existing behavioral taxonomies for ADS tend to be either (a) use-case catalogs organized by traffic situation [26], (b) maneuver-level libraries for motion planning [9], or (c) regulatory compliance checklists [17]. To our knowledge, existing work has not systematically derived behaviors from ODD layer composition, decomposes them along longitudinal/lateral control axes, or characterizes them against a unified property framework. Scenario-based testing literature [16, 18] has addressed the challenge of scenario space coverage, typically operating at the concrete or logical scenario level. Ontology-based approaches to scene generation [5] and keyword-based methods for translating functional scenarios into executable logical scenarios [15] provide complementary machinery for instantiating the scenario families formalized in Section 6, but do not themselves derive the behavioral competencies being tested. Our contribution operates one level above: at the behavioral competency level, providing the intermediate layer between ODD specification and scenario instantiation. Coverage adequacy for scenario-based validation remains an active research area. Amersbach and Winner [2] show a persistent gap between statistically required and computationally feasible scenario counts for highly automated vehicles, and situation- and combinatorial-coverage criteria [1, 23] have been proposed as tractable proxies. The Behavioral Competency Coverage Score identified as future work in Section 8 is intended to provide a coverage criterion at the behavioral level, complementary to these parameter- and situation-level approaches. 3.4 Rule-Based Planning and Behavioral Compliance Recent work on rule-aware ADS planning [8] has demonstrated that explicit behavioral rule catalogs, organized by priority hierarchy, can substantially reduce behavioral violations in closed-loop simulation. The taxonomy presented in this paper provides the behavioral specification layer from which such rule catalogs are derived, forming a complementary pair: the taxonomy defines what must be demonstrated; the rule-enforcement layer defines how it is achieved. 3.5 Formal Safety Models and Safety Argumentation Formal models such as Responsibility-Sensitive Safety (RSS) [22] define mathematically verifiable rules for longitudinal and lateral safe distances and have influenced the kinematic stability sub-component of the Safety property (Section 4.2). RSS operates at the maneuver/control level; the taxonomy in this paper operates at the behavioral competency level and can incorporate RSS-derived thresholds as one implementation choice for the Safety Envelope metrics defined in J3237 [20]. Structuring the resulting evidence into a coherent safety case is a separate concern from generating the evidence itself. Goal Structuring Notation (GSN) [13] provides a graphical safety argument notation widely used to link claims, evidence, and context in ADS safety cases, including in conjunction with standards such as UL 4600 [25]. The behavioral competency coverage evidence generated via Equation 1 is intended to serve as GSN- compatible evidence nodes within such an argument structure, though formalizing that mapping is left as future work (Section 9.1). Testing methodology challenges specific to ADS - non-deterministic algorithms, inductive learning components, and driver-out-of-the-loop operation - are cataloged by Koopman and Wagner [14], and the NHTSA testable-cases framework [24] provides an early structured template for ODD and OEDR-organized test case development that AVSC00008202111 and this taxonomy both build upon. 4 Behavioral Framework 4.1 Defining a Behavioral Competency Following AVSC00008202111 [4], we define a behavioral competency as a demonstrated capability of an ADS to execute a goal-oriented action within a specified ODD context, with observable and measurable outcomes against defined acceptance criteria. This definition distinguishes three levels of behavioral description: • Maneuver: a physical control action (e.g., apply brake, steer left) - atomic, no goal semantics. • Behavior: a goal-oriented action within the DDT (e.g., change lanes to avoid obstacle) - compositional, context-free. • Competency: behavior + ODD context + evaluation criteria (e.g., execute lane change in rain at 105 km h−1105\,km\,h^-1 with a 2 s target gap, measured against J3237 DA metrics including Safety Envelope Violation and Aggressive Acceleration Violation (AAV) thresholds [20]) - the unit of specification and validation. Competency is therefore the appropriate unit for taxonomy construction: it captures what the ADS must do, where it must do it, and how well it must perform. 4.2 The Four-Property Framework Every behavioral competency is simultaneously governed by four co-active properties. These are not a priority ordering; they are concurrent constraints whose mutual tension defines where specification is hardest. [Safety] Encompasses three distinct sub-components operating at different timescales: (1) gap maintenance - preventive, continuous longitudinal and lateral headway management (TTC, THW, minimum clearance); (2) conflict and collision avoidance - reactive OEDR response to dynamic hazards including cut-in vehicles, VRUs, and emergency vehicles; (3) kinematic stability - operation within the safe dynamic envelope defined by lateral/longitudinal acceleration limits, jerk bounds, and tire slip margins. A behavioral specification addressing only one sub-component is incomplete with respect to the full safety property. [Compliance] Encompasses two layers: (1) legal compliance - adherence to traffic laws, signal states, speed limits, right-of-way rules, and regulatory obligations; these are codified and largely testable against ground truth; (2) behavioral norm compliance - courtesy, predictability, and defensive posture toward other road users; these are socially expected but not always legally mandated, and harder to specify formally. An ADS that satisfies legal compliance but violates behavioral norms creates interaction hazards and erodes public trust [7]. [Comfort] Captures the rider’s continuous experience of control predictability across three axes: longitudinal dynamics (acceleration/deceleration smoothness, ramp profiles), lateral dynamics (lateral jerk, steering abruptness, cornering feel), and onset frequency (abruptness of maneuver initiation). Comfort is not a luxury metric: harsh dynamics that are technically safe erode rider trust, reduce adoption, and frequently indicate edge-case behaviors warranting safety review. Comfort therefore serves as a proxy signal for safety margin monitoring. [Efficiency] Captures four sub-dimensions: (1) mission completion - reaching the goal without unnecessary aborts or fallbacks; (2) trip time - route efficiency, unnecessary stops; (3) energy/fuel - smooth dynamics, predictive braking; (4) product-level soft metrics - domain-specific commercial requirements such as bay throughput (Hub), pickup punctuality (robotaxi), and dwell time. Efficiency is the most underspecified property in academic ADS literature but is a first-class requirement in commercial deployment, particularly in the Hub domain. 4.3 Longitudinal and Lateral Decomposition Every behavioral competency is further tagged by the primary control axis it stresses: Longitudinal (Lon), Lateral (Lat), or Both. This decomposition is not merely taxonomic; it reflects the architecture of ADS planning and control stacks, which typically separate longitudinal and lateral controllers. Behaviors stressing both axes simultaneously are, as a class, the hardest to specify and validate: their longitudinal and lateral sub-specifications interact, and failure modes in one axis can propagate to the other. Among our 21 behaviors, all three inter-domain deep-dive examples (Lane Change, Unprotected Left Turn, Staging Entry and Docking) belong to the Both category. SAFETY Gap maintenance ⋅· Conflict & collision avoidance ⋅· Kinematic stability within safe dynamic limits COMPLIANCE Traffic laws & road rules ⋅· Courtesy & predictable defensive driving for other road users COMFORT Longitudinal & lateral dynamics ⋅· Smoothness of onset ⋅· Rider trust-building EFFICIENCY Mission completion ⋅· Trip time ⋅· Energy/fuel ⋅· Product-level soft metrics Co-active on every behavior Figure 1: The four-property framework. Properties are co-active constraints on every behavioral competency; the tension between them defines where specification is hardest. 5 Operational Domain Taxonomy 5.1 Deriving Domains from ODD Layer Composition The three operational domains in this taxonomy are not arbitrary categorizations. They emerge from the characteristic composition of PEGASUS ODD layers that define each deployment context: • Highway (HWY): dominant layers L1 (structured lane geometry), L2 (traffic control infrastructure and applicable traffic laws), L4 (vehicle-only agents), L5 (high-speed conditions). Characterized by continuous DDT execution, structured geometry, and absence of complex agent interactions. • Urban (URB): dominant layers L1-L4 (intersections, mixed agents, VRUs, traffic control infrastructure, and applicable traffic laws), L5. Characterized by heterogeneous agent interactions, negotiated right-of-way, and high behavioral norm demands. • Hub (HUB): dominant layers L1 (facility geometry), L2 (facility-specific rules, not traffic law), L3 (dock events, bay state), L6 (digital coordination, bay assignment). Characterized by geofenced operation, low-speed precision, internal ruleset compliance, and high DDT-F concentration. The Hub domain is structurally distinct in two important ways. First, its compliance layer references facility-internal operational rules in addition to applicable traffic law, meaning that HWY/URB compliance specifications are not directly transferable and must be supplemented with facility-specific requirements (this taxonomic distinction does not affect or limit the application of any applicable federal, state, or local legal or regulatory requirements to hub operations). Second, it has the highest concentration of DDT-F behaviors of any domain, yet almost no published benchmark coverage exists for hub-specific behavioral evaluation (see Section 8). 5.2 Behavior Anatomy Template Each of the 21 behaviors is specified using a uniform five-field anatomy template: 1. Trigger Condition: the ODD state, navigation request, or OEDR detection event that activates the behavioral mode. 2. Entry/Exit Criteria: state transitions defining when the ADS enters and exits the behavioral mode. 3. OEDR Requirements: the object classes, events, and road features that must be detected and responded to. 4. Applicable Metrics: Dynamic Driving Task Assessment (DA) metrics from SAE J3237 [20] applicable to this behavior, selected by observability level (direct measurement, inference from observable proxies, or internal system state access). 5. ODD Sensitivity: how PEGASUS layer parameterizations modulate the specification (e.g., speed range, weather, agent density, visibility). Each behavior is additionally annotated with: (a) its Lon/Lat/Both axis classification, and (b) a four-property profile indicating the relative load on each of the Safety, Compliance, Comfort, and Efficiency properties. 5.3 The 21-Behavior Taxonomy Table 1 enumerates all 21 behavioral competencies organized by domain and Lon/Lat classification. Table 1: 21-Behavior Taxonomy: Domain, Axis Classification, and Primary Property Load. Note: this is for illustration purposes and is not meant to be comprehensive Behavior Dom. Axis Primary Properties Highway Domain (HWY) Speed Limit Compliance HWY Lon Compliance, Safety Following Distance Maintenance HWY Lon Safety, Efficiency Emergency Braking Response HWY Lon Safety Lane Keeping HWY Lat Safety, Comfort Cut-In Response HWY Lat Safety Lane Change HWY Both Safety, Compliance, Comfort Highway Merge HWY Both Safety, Compliance Emergency Vehicle Response HWY Both Compliance, Safety Urban Domain (URB) Traffic Signal Compliance URB Lon Compliance, Safety Stop Sign Compliance URB Lon Compliance Yield / Right-of-Way URB Lon Compliance, Safety Speed Zone Compliance URB Lon Compliance Crosswalk / VRU Avoidance URB Lat Safety Protected Turn Execution URB Lat Compliance, Safety Unprotected Left Turn URB Both Safety, Compliance Intersection Negotiation URB Both Safety, Compliance Hub Domain (HUB) Staging Queue Management HUB Lon Efficiency, Safety Docking Approach Speed Control HUB Lon Safety, Comfort Bay Alignment HUB Lat Safety, Efficiency Geofence Boundary Compliance HUB Lat Compliance Staging Area Entry & Docking HUB Both Safety, Efficiency Figure 2: Taxonomy overview: 21 behaviors across HWY, URB, and HUB domains, with Lon/Lat/Both axis distribution per domain. Several structural observations merit emphasis: • The HWY domain is longitudinal-heavy (3 Lon, 2 Lat, 3 Both), reflecting its structured geometry and speed-dominant challenge space. • The URB domain is longitudinal-heavy in regulatory behaviors (4 Lon) but its highest-complexity behaviors (Unprotected Left, Intersection Negotiation) require both axes simultaneously. • The HUB domain has proportionally the highest Both-axis concentration (1 of 5, but the single Both behavior is its most operationally critical), and its compliance layer references facility rules rather than traffic law. • Behaviors in the Both category across all domains consistently carry the highest Safety property load, confirming the intuition that multi-axis behavioral stress correlates with safety criticality. 6 Behavior Deep Dives To illustrate the behavior anatomy template in practice, this section provides full competency specifications for one representative behavior per domain. The three behaviors - Lane Change (HWY), Unprotected Left Turn (URB), and Staging Area Entry and Docking (HUB) - are selected because each is the most complex Both-axis behavior in its domain and collectively demonstrate the full range of property loading and ODD sensitivity across the taxonomy. 6.1 HWY: Lane Change Trigger: Navigation goal or obstacle avoidance request requiring lateral displacement to an adjacent lane. Entry: Target gap in adjacent lane exceeds minimum TTC threshold; lane change permitted by geometry and signage. Exit: Ego vehicle centered in target lane, lateral acceleration below comfort threshold; original lane clear of following obligations. Lon sub-spec: Gap acceptance against lead vehicle in target lane (TTC≥τminTTC≥ _ , Time Gap ≥Δmin≥ _ ); speed matching to target lane flow prior to merge initiation. Lat sub-spec: Lane departure timing; lateral acceleration profile during merge (|ay|≤ay,max|a_y|≤ a_y, ); lateral jerk onset; time-to-lane-center (TTLC) post-merge. OEDR: Adjacent lane occupancy, TTC with approaching vehicles, turn signal intent of neighboring agents. Metrics (J3237): Lane Departure Violation post-merge (black-box); Aggressive Acceleration Violation (AAV) for lateral excursions during merge (black-box); Safety Envelope Violation for target-lane gap acceptance (black-box). ODD Sensitivity: Minimum gap thresholds (τmin _ , Δmin _ ) increase under L5 degraded conditions (rain, night, fog); lane width modulates acceptable lateral jerk profile; speed range constrains available merge window. 6.2 URB: Unprotected Left Turn Among all 21 behaviors, the Unprotected Left Turn carries the highest aggregate Safety load. It simultaneously activates all three Safety sub-components (gap maintenance for oncoming traffic, conflict/collision avoidance for VRUs in crosswalks, kinematic stability through the turn arc) and engages both Compliance layers (signal phase, right-of-way obligation, and courtesy yielding behavior). Trigger: Navigation goal requires left turn at unprotected intersection (no dedicated left-turn phase). Entry: Traffic signal permits movement; oncoming gap accepted; no VRU in conflict crosswalk. Exit: Ego vehicle centered in receiving lane; intersection cleared; crosswalk no longer in conflict zone. Lon sub-spec: Oncoming vehicle gap acceptance (TTConcoming≥τULTTTC_oncoming≥ _ULT); deceleration profile to yield point; clearance timing through intersection. Lat sub-spec: Turn arc geometry (minimum radius, lane boundary constraints); crosswalk incursion avoidance; lateral positioning in receiving lane. OEDR: Oncoming vehicle TTC, VRU presence and velocity in near and far crosswalks, traffic signal phase and countdown state, right-of-way negotiation with opposing left-turning vehicles. Metrics (J3237): Safety Envelope Violation for oncoming-gap conflict (black-box); Traffic Law Violation for right-of-way/signal compliance (black-box); Crash Instance for VRU/crosswalk conflict outcome (black-box); Event Response Time Violation for reaction to oncoming or VRU hazards (black-box). ODD Sensitivity: Oncoming gap threshold (τULT _ULT) increases significantly under L5 low-visibility conditions; pedestrian density (L4) modulates crosswalk dwell time; unsigned intersections (no L2 signal) require full negotiation via behavioral norms only. 6.3 HUB: Staging Area Entry and Docking The Staging Area Entry and Docking behavior is the most operationally distinctive behavior in the taxonomy for three reasons. First, it is the only behavior in which the Efficiency property load rivals Safety. Second, its compliance layer references facility-internal operational rules in addition to applicable traffic law, meaning that HWY/URB compliance specifications are not directly transferable and must be supplemented with facility-specific requirements (this taxonomic distinction does not affect or limit the application of any applicable federal, state, or local legal or regulatory requirements to hub operations). Third, it contains the highest DDT-to-DDT-F transition risk of any behavior: the handoff zone at bay approach is an area warranting continued engineering attention given the complexity of the DDT-to-DDT-F transition. Trigger: Navigation route endpoint is a hub facility; ego vehicle approaches facility geofence boundary. Entry: Geofence crossed; speed below threshold; assigned bay confirmed via L6 digital channel. Exit: Ego vehicle at rest in assigned bay within docking accuracy tolerance; DDT-F handoff complete. Lon sub-spec: Speed ramp-down profile from facility entry to bay approach; queue gap maintenance to vehicle ahead in staging lane; final docking deceleration profile. Lat sub-spec: Bay alignment (lateral offset from dock marker ≤δmax≤ _ ); geofence boundary compliance; dock marker tracking during final approach. OEDR: Bay occupancy state, ground crew presence and clearance signal, dock alignment marker detection, queue vehicle following distance. Metrics (J3237): [Outside the ODD] ADS DDT Execution Violation for DDT-F handoff correctness at bay approach (grey-box); Intervention Request/Prompt to Take Over Violation for handoff signaling (grey-box); ODD Recognition Violation for geofence boundary detection (white-box); DDT-Relevant Object Distance Calculation Error Rate for dock-marker/bay-alignment accuracy (white-box). ODD Sensitivity: Indoor vs. outdoor bay geometry modulates sensor availability; structured vs. unstructured bay layout modulates marker detection reliability; dock crew protocols vary by facility operator (L2 layer). 7 ODD × Behavior → Scenario Families 7.1 Formal Parameterization A key operational property of the taxonomy is that its intersection with PEGASUS ODD layer parameterizations directly generates concrete scenario families for behavioral testing and evaluation. We formalize this as follows. Let ℬ=b1,…,b21B=\b_1,…,b_21\ denote the set of 21 behavioral competencies, and let ℒ=L1,…,L6L=\L_1,…,L_6\ denote the PEGASUS ODD layer set. Each layer LiL_i is associated with a parameter space Ωi _i (e.g., Ω5 _5 includes weather state, lighting condition, and road surface friction). A scenario family (bi,ω)S(b_i,ω) is defined as: (bi,ω)=bi⊗ω,ω∈Ω1×⋯×Ω6S(b_i,ω)\;=\;b_i\; \;ω, ω∈ _1×·s× _6 (1) where ⊗ denotes the contextualization of behavior bib_i under ODD parameter instantiation ω. Each S is associated with: • A set of entry conditions derived from the behavior’s trigger and entry criteria (Section 4), modulated by ω. • A set of pass/fail acceptance criteria derived from the behavior’s applicable J3237 metrics, with thresholds modulated by ω (e.g., TTC thresholds increase under low visibility). • An ODD sensitivity profile indicating which layers have the strongest modulating effect on the scenario’s difficulty and metric thresholds. 7.2 Worked Example Behavior: Lane Change (HWY, Section 6.1) ODD instantiation ω: • L1L_1: three-lane divided highway, 3.6 m3.6\,m lanes • L4L_4: cut-in agent in target lane, initial TTC = 1.5 s • L5L_5: rain, nighttime, road surface friction coefficient 0.6 Scenario family S: • Entry: ego navigating in left lane; obstacle in path triggers lane change request to center lane; cut-in agent approaches at closing speed. • Pass criteria: TTCadj≥τmin(ω)TTC_adj≥ _ (ω) where τmin _ is increased under L5 rain/night conditions per ODD sensitivity profile (an illustrative 20% shown here for concreteness; actual factors are ODD and implementation-specific); |ay|≤ay,max|a_y|≤ a_y, at road friction coefficient; lateral deviation from lane center ≤δTTLC≤ _TTLC within 3 s3\,s of completion. • Fail modes: merge initiated below minimum gap; lateral jerk exceeds comfort envelope; lane center not achieved within TTLC threshold. This single behavior × ODD instantiation yields a concrete, executable scenario with defined pass/fail criteria traceable to J3237 metrics (Safety Envelope Violation, AAV) and supplementary comfort-oriented criteria (TTLC), modulated by PEGASUS layer parameters. 7.3 SOTIF Alignment The scenario family generation mechanism is designed to support the ISO 21448 evidence structure [11]. SOTIF describes identification of triggering conditions for known unsafe scenarios; the ODD sensitivity profiles in each behavior specification identify which layer parameters most strongly modulate behavioral stress, intended to inform triggering condition analysis. SOTIF also describes verification and validation based on scenarios; scenario families generated via Equation 1 are suitable for use within a SOTIF verification and validation, with the behavioral competency coverage metric potentially serving as one input to a coverage argument. PEGASUS ODDL1: Road geometryL2: InfrastructureL3: Temp. eventsL4: Moving objectsL5: EnvironmentL6: Digital info ×Behavior21 behaviors × 3 domainsLon ⋅· Lat ⋅· Both axis4-property specJ3237 metricsODD sensitivity Scenario Families for Behavioral Testing & Evaluation Entry conditions ⋅· J3237 pass/fail thresholds (modulated by ω) ⋅· SOTIF Clause 8/9 coverage evidence Example: Lane Change (HWY) × L5: rain + night × L4: cut-in at 1.5 s TTC ⇒\; \; Concrete scenario with increased τmin _ and Safety Envelope Violation/AAV pass criteria Figure 3: ODD × Behavior → Scenario Families. Each PEGASUS layer parameterization crossed with a behavior specifications yields a concrete scenario family with J3237-traceable pass/fail criteria. 7.4 Fleet Log Coverage Verification An important complementary use of the taxonomy is as a coverage query schema for AV fleet log mining. Given a fleet log corpus and the 21-behavior taxonomy, each log segment can be labeled against the behavior it exercises and the ODD parameters under which it was collected, enabling computation of behavioral competency coverage scores across the corpus. This application is a direct extension of the scenario family generation mechanism and is identified as future work in Section 9.1. 8 Discussion 8.1 The DDT / DDT-F Dimension The taxonomy reveals a structural dimension that is largely absent from existing behavioral frameworks: the DDT/DDT-F distribution across domains. Behaviors in which the ADS maintains full continuous Dynamic Driving Task execution (DDT behaviors) are well-represented in existing benchmarks and simulation environments. Behaviors involving DDT Fallback (DDT-F) - transitions, handoffs, minimal risk condition activation - are poorly covered. The three domains differ substantially in their DDT/DDT-F concentration. The Highway domain is DDT-dominant: the ADS is in continuous control throughout all eight behaviors. The Urban domain is mixed: most behaviors are DDT, but intersection negotiation and emergency scenarios involve partial DDT-F elements. The Hub domain is DDT-F heavy: the Staging Area Entry and Docking behavior involves a mandatory DDT-to-DDT-F transition at bay approach, and this transition represents the highest-complexity operational handoff in the domain. To our knowledge, no published benchmark specifically targets hub DDT-F transitions. 8.2 The Hub Domain Gap The Hub domain represents the most structurally distinct and least researched operational context in the taxonomy. Its distinguishing characteristics - geofenced operation, facility-rule-based compliance, low-speed precision maneuvering, high DDT-F concentration, and Efficiency as a first-class property - mean that HWY and URB behavioral specifications are not directly transferable. Yet the published literature treating hub-specific behavioral requirements remains limited. Public datasets and benchmarks for hub behavioral evaluation are not currently available at the scale of those for highway and urban domains (e.g., Argoverse [27], Waymo Open Motion [6]). This gap is consequential: commercial AV deployments increasingly rely on hub operations as the operational anchor for service (depot-to-service-area dispatch, automated last-mile logistics). The five Hub behaviors in this taxonomy, and particularly Staging Area Entry and Docking, represent a high-priority research target. 8.3 Open Problems in the Published Literature The taxonomy surfaces five gaps in the published academic and industry literature on ADS behavioral specification and validation. Note that these open problems describe gaps in the published research literature and are not intended to characterize the completeness or adequacy of any specific organization’s internal validation or safety assurance processes, which may address these areas through methods not described in the public literature. 1. Inter-behavior transition specifications: Formal specification of behavioral state transitions at domain boundaries (e.g., HWY to URB entry) and between behavioral modes within a domain does not appear in the published literature. The behavior anatomy template specifies entry and exit criteria per behavior but does not formalize the transition graph between behaviors. 2. Quantitative competency coverage metrics: To our knowledge, no publicly documented metric currently exists for measuring the degree to which a scenario suite covers a behavioral specification space. A Behavioral Competency Coverage Score (BCCS) over a taxonomy such as the one presented here would provide a tractable, standards-aligned coverage argument suitable for public disclosure. 3. Hub domain benchmarks: The development of publicly available behavioral benchmarks and annotated datasets for hub operational behaviors represents a gap in the published literature with direct deployment relevance. 4. Closed-loop behavioral compliance at scale: Open-loop metric evaluation does not capture interaction dynamics between ADS and other agents. To our knowledge, no publicly documented infrastructure exists for behavior-level compliance evaluation in reactive multi-agent scenarios at deployment scale. 5. Four-property weighting by ODD context: The relative weighting of Safety, Compliance, Comfort, and Efficiency properties is likely ODD-dependent. To our knowledge, no publicly documented framework for context-dependent property weighting in multi-objective behavioral specification has been proposed in the literature. 9 Summary and Conclusions This paper presented a structured, standards-grounded taxonomy of 21 behavioral competencies for automated driving systems, organized across three operationally distinct domains derived from PEGASUS ODD layer composition. The primary contributions are: 1. A derivation methodology: Operational domains and behavioral competencies are derived systematically from ODD layer composition, providing a principled bridge from operating condition specification to behavioral enumeration. 2. A 21-behavior taxonomy: Behaviors are specified using a uniform five-field anatomy template, decomposed along longitudinal and lateral control axes, and characterized against a four-property framework (Safety, Compliance, Comfort, Efficiency) that is simultaneously active on every behavior. 3. A scenario generation mechanism: The formal crossing of PEGASUS ODD layer parameterizations with behavioral competency specifications yields concrete scenario families with J3237-traceable pass/fail criteria, directly supporting SOTIF Clause 8/9 coverage evidence generation. The taxonomy identifies the Hub domain as the most structurally distinct and least benchmarked operational context in commercial ADS deployment, with the DDT-F transition in Staging Area Entry and Docking representing a high-priority gap for both specification and validation research. 9.1 Future Research Directions This taxonomy serves as the behavioral specification layer for a broader research program. The following directions are identified as extensions warranted by gaps in the published literature, and are not intended to suggest that these capabilities are absent from any organization’s internal engineering or validation processes. Immediate directions include: (1) formal specification of the inter-behavior transition graph at domain boundaries; (2) development and public documentation of a Behavioral Competency Coverage Score (BCCS) metric for scenario suite adequacy assessment, of which to our knowledge no publicly documented version currently exists; (3) publication of fleet log mining methodology for behavioral coverage verification using the taxonomy as a query schema; (4) closed-loop behavioral compliance evaluation in reactive simulation environments; and (5) hub-specific benchmark development for public release. The taxonomy has been deployed as the specification layer for a rule-enforced trajectory optimization system [8], establishing its operational validity as an engineering artifact. For the avoidance of doubt, this publication is not intended and shall not be construed to establish, create, or be deemed to have created a required or expected standard, methodology, or benchmark regarding the design, testing, or validation of automated driving systems, nor shall it prevent, hinder, or restrict Torc Robotics, Inc. or any affiliate from adopting, rejecting, or deviating from any concept, framework, or open problem discussed herein in its actual engineering practices. References to open research problems, future work, or gaps in published literature describe the state of the broader academic and industry research landscape and are not representations regarding the completeness, adequacy, or sufficiency of Torc’s internal validation, testing, or safety assurance processes for any specific deployed product. References [1] R. Alexander, H. Hawkins, and D. Rae (2015) Situation Coverage – A Coverage Criterion for Testing Autonomous Robots. Technical report Department of Computer Science, University of York. Cited by: §3.3. [2] C. Amersbach and H. Winner (2019) Defining Required and Feasible Test Coverage for Scenario-Based Validation of Highly Automated Vehicles. In Proc. IEEE Intelligent Transportation Systems Conference (ITSC), p. 425–430. External Links: Document Cited by: §3.3. [3] ASAM e.V. (2023) ASAM OpenSCENARIO 2.0: User Guide. Note: https://w.asam.net/standards/detail/openscenario/ Cited by: §3.2. [4] Automated Vehicle Safety Consortium (2021) Best Practices for ADS Behavioral Competency Evaluation (AVSC00008202111). Best Practice Document Technical Report AVSC00008202111, SAE International. Cited by: §2, §3.1, §4.1. [5] G. Bagschik, T. Menzel, and M. Maurer (2018) Ontology based Scene Creation for the Development of Automated Vehicles. In Proc. IEEE Intelligent Vehicles Symposium (IV), p. 1813–1820. External Links: Document Cited by: §3.3. [6] S. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y. Chai, B. Sapp, C. R. Qi, Y. Zhou, et al. (2021) Large scale interactive motion forecasting for autonomous driving: the waymo open motion dataset. In Proceedings of the IEEE/CVF international conference on computer vision, p. 9710–9719. Cited by: §8.2. [7] L. Fridman, B. Reimer, B. Mehler, and W. T. Freeman (2018) Cognitive load estimation in the wild. In Proceedings of the 2018 chi conference on human factors in computing systems, p. 1–9. Cited by: item [Compliance]. [8] H. Hajieghrary, B. Walter, C. Shinde, P. Schmitt, and M. Hurtado (2026) RECTOR: priority-aware rule-based reranking for compliance-aware autonomous driving trajectory selection. arXiv preprint arXiv:2605.25095. Cited by: §3.4, §9.1. [9] C. Hubmann, J. Schulz, M. Becker, D. Althoff, and C. Stiller (2018) Automated driving in uncertain environments: planning with interaction and uncertain maneuver prediction. IEEE transactions on intelligent vehicles 3 (1), p. 5–17. Cited by: §3.3. [10] International Organization for Standardization (2018) ISO 26262:2018 — Road Vehicles — Functional Safety. Technical report Technical Report ISO 26262, ISO. Cited by: §3.1. [11] International Organization for Standardization (2022) ISO 21448:2022 — Road Vehicles — Safety of the Intended Functionality. Technical report Technical Report ISO 21448, ISO. Cited by: §3.1, §7.3. [12] International Organization for Standardization (2024) ISO/PAS 8800:2024 – Road Vehicles – Safety and Artificial Intelligence. Technical report Technical Report ISO/PAS 8800, ISO. Cited by: §3.1. [13] T. Kelly and R. Weaver (2004) The Goal Structuring Notation – A Safety Argument Notation. In Proc. Dependable Systems and Networks (DSN) Workshop on Assurance Cases, Cited by: §3.5. [14] P. Koopman and M. Wagner (2016) Challenges in autonomous vehicle testing and validation. SAE International journal of transportation safety 4 (2016-01-0128), p. 15–24. Cited by: §3.5. [15] T. Menzel, G. Bagschik, L. Isensee, A. Schomburg, and M. Maurer (2019) From Functional to Logical Scenarios: Detailing a Keyword-Based Scenario Description for Execution in a Simulation Environment. In Proc. IEEE Intelligent Vehicles Symposium (IV), p. 2383–2390. External Links: Document Cited by: §3.3. [16] T. Menzel, G. Bagschik, and M. Maurer (2018) Scenarios for development, test and validation of automated vehicles. In 2018 IEEE intelligent vehicles symposium (IV), p. 1821–1827. Cited by: §3.3. [17] National Highway Traffic Safety Administration (2017) Automated Driving Systems 2.0: A Vision for Safety. Technical report Technical Report DOT HS 812 442, US Department of Transportation, NHTSA. Cited by: §3.3. [18] S. Riedmaier, T. Ponn, D. Ludwig, B. Schick, and F. Diermeyer (2020) Survey on scenario-based safety assessment of automated vehicles. IEEE access 8, p. 87456–87477. Cited by: §3.3. [19] SAE International (2021) Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles. Surface Vehicle Recommended Practice Technical Report J3016_202104, SAE International. Note: https://doi.org/10.4271/J3016_202104 External Links: Document Cited by: §3.1. [20] SAE International (2025-08) Dynamic Driving Task Assessment (DA) Metrics for Automated Driving Systems. SAE Recommended Practice Technical Report J3237_202508, SAE International. Note: https://doi.org/10.4271/J3237_202508 External Links: Document Cited by: §3.1, §3.5, 3rd item, item 4. [21] M. Scholtes, L. Westhofen, L. R. Turner, K. Lotto, M. Schuldes, H. Weber, N. Wagener, C. Neurohr, M. H. Bollmann, F. Körtke, J. Hiller, M. Hoss, J. Bock, and L. Eckstein (2021) 6-layer model for a structured description and categorization of urban traffic and environment. IEEE Access 9 (), p. 59131–59147. External Links: Document Cited by: §2, §3.2. [22] S. Shalev-Shwartz, S. Shammah, and A. Shashua (2017) On a Formal Model of Safe and Scalable Self-driving Cars. Note: arXiv:1708.06374 External Links: 1708.06374 Cited by: §3.5. [23] J. Tao, Y. Li, F. Wotawa, H. Felbinger, and M. Nica (2019) On the industrial application of combinatorial testing for autonomous driving functions. In 2019 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW), p. 234–240. Cited by: §3.3. [24] E. Thorn, S. C. Kimmel, and M. Chaka (2018) A Framework for Automated Driving System Testable Cases and Scenarios. Technical report Technical Report DOT HS 812 623, National Highway Traffic Safety Administration. Cited by: §3.5. [25] UL Standards & Engagement (2023) UL 4600: Standard for Safety for the Evaluation of Autonomous Products. Technical report Technical Report UL 4600, Underwriters Laboratories. Cited by: §3.1, §3.5. [26] S. Ulbrich, T. Menzel, A. Reschka, F. Schuldt, and M. Maurer (2015) Defining and substantiating the terms scene, situation, and scenario for automated driving. In 2015 IEEE 18th international conference on intelligent transportation systems, p. 982–988. Cited by: §3.3. [27] B. Wilson, W. Qi, T. Agarwal, J. Lambert, J. Singh, S. Khandelwal, B. Pan, R. Kumar, A. Hartnett, J. K. Pontes, et al. (2023) Argoverse 2: next generation datasets for self-driving perception and forecasting. arXiv preprint arXiv:2301.00493. Cited by: §8.2.