Paper deep dive
MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure
Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/19/2026, 5:29:50 AM
Summary
The paper introduces MotoSafety, an edge-AI architecture for assessing collision risk in powered two-wheeler (PTW) riders under time pressure (TP). Utilizing a large-scale dataset of 129,000 multivariate time-series sequences from 51 participants in a simulator, the model employs Learned Temporal Importance (LTI) to achieve high accuracy (94.97%) and low latency (0.135 ms), making it suitable for edge deployment. The study highlights the impact of TP on rider behavior and demonstrates the model's transferability to other domains.
Entities (10)
Relation Signals (6)
Time Pressure → influences → Collision Risk
confidence 95% · limited studies exist on how cognitive stressors such as Time Pressure influence collision risk.
MotoSafety → usesprinciple → Learned Temporal Importance
confidence 95% · MotoSafety, a novel edge-AI architecture grounded in the Learned Temporal Importance principle.
MotoSafety → isdeployableon → Edge devices
confidence 90% · it is suitable for edge deployment on low-cost CPU hardware.
MotoSafety → outperforms → TimesNet
confidence 90% · MotoSafety achieves 94.97% accuracy... outperforming ten baselines, including TimesNet
MotoSafety → outperforms → LLM4TS
confidence 90% · MotoSafety achieves 94.97% accuracy... outperforming ten baselines, including TimesNet and LLM4TS
MotoSafety → supports → Safe System Approach
confidence 90% · This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, and High TP, capturing 64 features across vehicle dynamics, control inputs, proximity, and behavioral violations. Building on this dataset, we propose MotoSafety, a novel edge-AI architecture grounded in the Learned Temporal Importance principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4x lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems.
Tags
Links
- Source: https://arxiv.org/abs/2608.17823v1
- Canonical: https://arxiv.org/abs/2608.17823v1
Trouble viewing inline? Open PDF directly →
Full Text
73,276 characters extracted from source content.
Expand or collapse full text
MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure Sumit S. Shevtekar Affiliation: Department of Computer Science and Engineering, Indian Institute of Technology Indore, Khandwa Road, Simrol, Indore, 453552, Madhya Pradesh, India Corresponding author: Corresponding author. E-mail: sumit.shevtekar@gmail.com Chandresh K. Maurya Affiliation: Department of Computer Science and Engineering, Indian Institute of Technology Indore, Khandwa Road, Simrol, Indore, 453552, Madhya Pradesh, India Gourab Sil Affiliation: Department of Civil Engineering, Indian Institute of Technology Indore, Khandwa Road, Simrol, Indore, 453552, Madhya Pradesh, India Subasish Das Affiliation: Civil Engineering Program, Ingram School of Engineering, Texas State University, RFM 5202, San Marcos, 78666, TX, USA Abstract Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, and High TP, capturing 64 features across vehicle dynamics, control inputs, proximity, and behavioral violations. Building on this dataset, we propose MotoSafety, a novel edge-AI architecture grounded in the Learned Temporal Importance principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4× lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems. Keywords: Powered two-wheeler safety , Time pressure , Collision risk assessment , Deep learning , Intelligent transportation systems †graphicalabstract: 1 Introduction Road traffic accidents are the leading cause of death worldwide, particularly in Low- and Middle-Income Countries (LMICs) [1]. Powered two-wheeler (PTW) riders represent one of the most vulnerable groups, facing disproportionately high fatality rates due to limited protection, complex traffic, and socio-economic pressures to meet mobility and livelihood deadlines. Globally, PTWs account for approximately 21% of road traffic fatalities, a figure that rises to nearly 38% in LMICs, where PTWs serve as an affordable and widely adopted mode of transport [1]. In India alone, PTWs constituted 74.4% of registered vehicles as of 2022 shown in Fig. 2 [2]. As of 2023, two-wheelers constitute the largest segment of the India’s vehicle population with more than 263 million registered units and were involved in 44.5% of total road traffic fatalities, as shown in Fig. 2 [3, 4]. According to official MoRTH reports on road accidents in India, road fatalities in India are predominantly male, ranging from 85.2% to 87.3% between 2019 and 2023 [5, 6, 7, 8, 9], with most fatalities occurring in the 18–45 age group. These statistics show the high risk associated with male PTW riding behavior. Despite advances in infrastructure design and traffic enforcement, human error remains the dominant contributor to PTW crashes, with risky behaviors such as overspeeding, abrupt maneuvering, and inadequate hazard response accounting for a large share of fatalities [10]. Naturalistic and simulator-based studies consistently show that riding speed, control variability, and situational context strongly influence crash risk [11]. PTW riders frequently travel at speeds up to 2.3 times higher than four-wheeler drivers, and overspeeding alone contributes to approximately 34% of PTW fatalities [12]. Figure 1: Registered Vehicles Share in India. Figure 2: Road Crash Fatalities in India. Time Pressure (TP) refers to the subjective perception of insufficient time to complete a riding task, commonly arising from risk-prone behaviors, situational urgency, or externally imposed constraints [12, 13]. This is especially common for app-based delivery riders in LMICs such as India. App-generated delivery deadlines, performance incentives, and intense customer demands for fast, flawless service combine to create persistent and often extreme TP [13, 14, 15]. Elevated TP affects cognitive processing and motor control, narrowing attentional bandwidth and accelerating decision-making. This in turn promotes risk-prone behaviors such as overspeeding, inconsistent braking, and delayed hazard response [12, 16, 10]. Simulator studies on car drivers report increased accident risk of up to 181% under high time pressure (HTP) compared to no time pressure (NTP) [12]. Related work [17] demonstrates that cognitive and emotional stress amplifies speed variability and control instability. However, a key cause of these behaviors, TP-induced cognitive stress, is still not well understood within the PTW-driving situation. Recent AI-driven traffic crash prediction leverages large-scale datasets, classical ML/DL architectures to improve safety [18, 19, 20, 21, 22]. Although effective in predicting four-wheeler crashes, these approaches inadequately model the unique dynamics and highly unstable control characteristics of PTWs. For example, [19] achieved ∼ 90% accuracy using deep neural networks on highway data, [20] improved driver behavior classification with feature selection on large datasets, [21] used Random Forest to predict injury severity (92% accuracy). However, these studies focus primarily on four-wheelers or general traffic datasets and do not incorporate fine-grained temporal modeling of PTW riders or cognitive stress factors such as TP. To the best of our knowledge, this is the first work to address the prediction of collision risk for two-wheeler riders under TP. To bridge this gap, we propose a novel MotoSafety model specifically designed for multivariate time series data and to quantify and predict collision risk by modeling the interplay between a riders cognitive state under TP and their operational maneuvers. Our work addresses a critical public safety exigency: the protection of vulnerable PTW riders operating under chronic TP—a phenomenon increasingly prevalent in the rapidly urbanizing and dense traffic environments of emerging economies like India. The proposed system can be deployed as an edge-based alert system, where a handlebar-mounted device or smart helmet provides haptic/audio alerts when collision risk exceeds a threshold, enabling proactive safety interventions for riders under TP (e.g., emergency commuting, delivery deadlines, or urgent travel). This methodology was co-developed with road safety authorities and transportation experts to address PTW fatalities in LMICs. By integrating stakeholder input on crash-prone scenarios and demographic vulnerability, we ensure that our edge-deployable model satisfies real-world operational requirements. Consequently, our approach transitions from theoretical inference to a practitioner-validated solution for tangible public safety impact. The key contributions of this work are summarized as follows: 1.1 Key Contributions We make the following key contributions. 1. Large-Scale PTW Simulator Data Under TP: To study the effect of TP on PTW riders, we collect 129,000 labeled multivariate time series data from 51 participants under No, Low, and High TP. Each input sequence comprises 64 features categorized into vehicle dynamics, control inputs, headway/tailway distances, and behavioral violations. 2. Novel Architecture for Collision Risk Forecasting and Classification: We propose MotoSafety, a novel architecture grounded in the Learned Temporal Importance (LTI) principle. The architecture integrates: (i) parallel multi-scale dilated convolutions for local-to-medium pattern extraction; (i) Bidirectional Long Short-Term Memory (Bi-LSTM) for long-range dependencies; (i) Temporal Importance Pooling (TIP) for content-aware temporal collapse, applied independently to both CNN and BiLSTM branches, which reduces downstream fusion, SE-recalibration, and attention stages to (1)O(1) with respect to sequence length, while the front-end encoder retains (L)O(L) complexity, giving an overall (L)O(L) model (Table 4); and (iv) self-gated fusion with squeeze-and-excitation (SE) and multi-head attention (MHA) for dynamic representation refinement. 3. SOTA Performance and Deployability: MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, showing improved performance over ten baselines including TimesNet, Time-LLM and LLM4TS. For long-term forecasting, it achieves 0.039 MSE and 0.094 MAE (4.4× lower error than Time-LLM and iTransformer). Notably, MotoSafety requires only ∼ 1.15M parameters and achieves an inference latency of 0.135 ms, making it 21.9× and 71.9× faster than TimesNet and LLM4TS, respectively, and makes it suitable for edge devices. 4. TP Inductive Bias, Real-World Feasibility, and Architecture Transferability: Explicit TP prediction improves collision accuracy from 94.09% to 94.97% (ground truth) and 94.82% (predicted TP). Using only 21 real-world features (IMU+GPS), MotoSafety achieves 93.91% accuracy, indicating practical deployment potential; hardware validation remains future work. Beyond PTW safety, the architecture demonstrates strong transferability to human activity (97.66%) and clinical exercise (99.65%) domains when retrained from scratch. 2 Related Work 2.1 Time Pressure as a Latent Risk Factor in Driving Safety Time Pressure (TP) is widely recognized as a cognitive stressor that degrades driver and rider decision-making and elevates crash risk. Under TP, individuals experience a perceived urgency to complete the driving task within insufficient time, which narrows attentional focus, accelerates cognitive processing, and increases reliance on risk-oriented heuristics [12, 16]. Empirical studies consistently show that TP leads to higher speeds, shorter accepted gaps, reduced safety margins, and increased crash likelihood across diverse traffic contexts [23, 10, 13]. Experimental evidence from simulator-based studies indicates that TP significantly alters longitudinal and lateral control behavior, including increased acceleration variability, unstable steering, and abrupt braking [24]. These behavioral deviations are strongly associated with near-crash and collision events, suggesting that TP acts as an upstream cognitive trigger for unsafe driving outcomes rather than merely a contextual modifier. PTWs are particularly susceptible to TP-induced risk due to unstable dynamics and high dependence on rider motor control compared to four-wheelers [23, 10, 25]. At unsignalized intersections, [12] show that drivers with low time pressure (LTP) and high time pressure (HTP) accept significantly smaller gaps, resulting in elevated crash likelihoods of 127% in LTP and 181% in HTP relative to no time pressure (NTP). TP increases cognitive load and unintentional attentional lapses under stress [26, 27], reducing situational awareness and leading to errors that precede collisions. These findings highlight TP as a critical contributor to risk in dense, heterogeneous traffic flows common in LMICs. 2.2 Simulator-Based Collision and Risk Analysis Driving simulators provide a safe and controlled environment for analyzing collision risk under cognitively demanding conditions such as TP, eliminating real-world crash risks while enabling high-resolution data collection for both collision risk classification and kinematic state forecasting [24, 28, 29]. Prior work demonstrates strong relative validity between simulated and real-world behavior, particularly for risk-related metrics including speed selection, braking intensity, and control variability [24, 28, 29]. Although absolute behavioral magnitudes may differ between simulated and naturalistic settings, the fundamental patterns of TP-induced changes remain consistent: increased speeds, more frequent lane changes, amplified braking intensity, and greater vehicular variability occur in both environments [24]. This supports the ecological validity of using simulators to investigate TP-induced behavioral changes in PTW rider behavior, enabling the current study to train and evaluate MotoSafety’s classification and forecasting tasks under controlled yet realistic conditions. 2.3 Machine Learning for Collision and Risk Prediction The intelligent transportation systems (ITS) literature increasingly employs machine learning and deep learning techniques to predict risky events, collisions, and hazardous driving events. Classical approaches using Random Forests, Support Vector Machines, and feed-forward neural networks report moderate success in identifying high-risk maneuvers [30, 31]. More recently, sequence models such as Temporal Transformers, Informer, and TimesNet have shown strong performance in modeling long-range temporal dependencies inherent in driving dynamics [32, 33]. However, most existing models focus on detecting externally observable outcomes (e.g., collisions or near-crashes) after risk has already manifested. They largely overlook latent cognitive precursors such as TP that shape behavior well before unsafe actions occur. Furthermore, few studies address PTW-specific dynamics, which are characterized by inherent instability and high dependence on rider motor control. 2.4 Research Gap and Motivation Although TP is a critical contributor to unsafe driving behavior, it is rarely integrated into collision risk assessment frameworks as a latent risk factor. This gap is particularly pronounced for PTWs, where safety margins are narrow and early intervention is essential. Motivated by these limitations, this work advances collision risk assessment under TP by leveraging high-resolution multivariate time-series data to infer risk directly from behavioral dynamics. By focusing on collision risk in cognitively stressed riding conditions, this study aligns with the Safe System Approach (SSA) [34] and the goal of Vision Zero [35], enabling the detection of high-risk states and supporting the development of intelligent rider-assistance systems and safer PTW mobility. 3 Methodology In this section, we provide details of our data collection protocol, followed by MotoSafety architecture. 3.1 Experimental Design and Data Collection 3.1.1 Simulator Setup All experimental trials use a static PTW simulator housed at the Institute (Fig. 3). The system integrates an instrumented Honda motorcycle frame with fully operational throttle, brake, clutch, and gear controls, replicating the ergonomics of real-world riding. A three-screen immersive display surrounds the rider to deliver a wide field of view with realistic roadway visualization. High-precision industrial-grade sensors continuously record rider inputs, vehicle dynamics, and environmental parameters through a real-time data acquisition interface, complying with IEC 393 accuracy standards [36]. While the frame is static, a 4-actuator motion system provides vibration and acceleration. Complete technical specifications, including sensor accuracy ratings, appear in Table 1. This setup provides a high-fidelity, controlled environment that supports systematic investigation of rider responses across diverse traffic and roadway scenarios. Table 1: Specifications of the two-wheeler simulator Component Specification Platform Static PTW frame, ISO-compliant Controls Throttle, clutch, hydraulic brakes, 5-speed gearbox, steering, handbrake Sensors Servo Potentiometers (Wire Wound – 50W): Resistance: 5 kΩ (±10%); Independent Linearity: ±0.5% (IEC 393), Electrical Angle: 355° ± 3°; Mechanical Rotation: 360° Power Rating: 3 W (@ 70°C); Rotational Life: 2,000,000 revolutions Operating Temperature: –40°C to +105°C Load Cells: Brake force measurement Optical Encoders: Gear position and control tracking, Limit Switches (SPDT): Gear shift detection Visual System 3 × 50" inch LED displays, 180° field of view Motion System 3-DOF Electric Actuators (×4): Max acceleration ±1 G, 0–100 Hz bandwidth Pitch: ±1.75°, Roll: ±3.5°, Heave: 38.1 m; Stroke: 1.5 inch Load Capacity: 114 kg per actuator Computer Systems Rider Station: Intel i7, 32GB RAM, GTX 1650 Instructor Station: Intel i5, 16GB RAM, GTX 1650 Software TechnoSim (AI traffic, environment modelling, scenario scripting) Data Logging Real-time data acquisition, sampling frequency fs≥100f_s≥ 100 Hz This ISO-compliant simulator provides a controlled, scalable, and safe environment to investigate rider behavior, cognitive load, and collision risk under realistic operational conditions. 3.1.2 Simulated Scenarios Figure 3: Participants performing the simulator test. (a) Primary route with intersections and conflict zones (b) AI vehicle triggers simulating dynamic interactions (c) Route direction changes and detour prompts (d) Speed change triggers near sensitive areas Figure 4: Sample snapshots of the simulated riding environment. We design a 4.8 km urban scenario on a high-fidelity two-wheeler simulator to investigate rider behavior under cognitive stress (TP). The route incorporates both undivided two-lane and four-lane sections with a 50 km/h speed limit, reflecting typical Indian urban conditions. The scenario embeds diverse events—including pedestrian crossings, obstacle overtaking, intersections, bike and car following segments to elicit real-time decision-making, adaptive control, and risk-taking under varying TP. The key elements of the scenario are illustrated in Fig. 4: Fig. 4(a): primary route with intersections and conflict zones, Fig. 4(b): AI vehicle triggers simulating dynamic interactions, Fig. 4(c): route direction changes and detour prompts, and Fig. 4(d): speed change triggers near sensitive areas. 3.1.3 Participants Table 2: Rider Characteristics (N = 51) Variable Level / Description Mean (SD) or % Age (yrs) 18–42 26.4 (5.7) Experience (yrs) ≥2≥ 2 5.6 (3.1) License Valid two-wheeler 100% Education Graduate / Final-year B.Tech 58.8% / 41.2% Simulator Exp No / Yes 92.2% / 7.8% Health Medically fit 100% The vast majority of two-wheeler riders in India are male, as reflected in MoRTH road fatality statistics report [5, 6, 7, 8, 9]. According to official MoRTH reports, males accounted for 86.0%, 87.3%, 86.4%, 86.2%, and 85.2% of PTW fatalities between 2019 and 2023, respectively, with most deaths in the 18–45 age group. This overwhelming male predominance provides a strong empirical basis for focusing on male participants in this study. Accordingly, fifty-one male riders (ages 18–42) with at least two years of riding experience participated. Detailed rider characteristics are presented in Table 2, and Fig. 3 shows participants performing simulator tasks. The male-only cohort is justified by three converging factors: First, national crash data shows male riders as the dominant high-risk group. According to MoRTH reports, male PTW fatalities accounted for 85.2–87.3% of all PTW deaths from 2019 to 2023 [5, 6, 7, 8, 9], with fatalities concentrated in the 18–45 age group. Statistical tests confirm this pattern: a binomial test shows male fatalities (86%) significantly exceed equal representation (p<0.0001p<0.0001), and a chi-square test confirms consistency across years (χ2=262.3χ^2=262.3, df = 5, p<0.0001p<0.0001). Second, female exposure to manually geared PTWs is extremely limited. Women hold only about 6% of motor vehicle licenses in India [37, 38, 39], significantly below equal representation (p<0.0001p<0.0001). The proportion of women actually riding manually geared motorcycles is even lower due to socio-cultural norms and practical barriers [38]. Consequently, the pool of female riders with manual transmission experience suitable for simulator-based research is extremely small. Third, the simulator’s manual transmission requirement creates a practical recruitment constraint. The high-fidelity PTW simulator replicates a manually geared motorcycle with functional clutch and gear controls, requiring prior manual transmission experience. Given the extremely low prevalence of female riders with such experience in India, recruiting a statistically sufficient female sample for robust analysis is infeasible. Thus, male participants were selected to ensure a sufficiently large and statistically robust sample for meaningful spatial risk analysis, consistent with the dominant high-risk demographic. Future studies will include female riders as their PTW participation increases. 3.1.4 Time Pressure (TP) Scenarios To ethically evoke graduated cognitive load mimicking emergency commuting, participants were tasked with reaching an examination hall under three conditions: (i) NTP: Baseline with ample time; (i) LTP: Restricted to 90% of baseline duration (Try not to be late); and (i) HTP: Restricted to 80% of baseline with active urgency prompts (e.g., Exam gate closing). This standardized exam-deadline scenario simulates high-arousal stress states common in urban transit, providing a controlled environment to study behavioral degradation. 3.1.5 Experimental Design and Protocol Figure 5: Flow of dataset collection and experimental design. A structured four-phase experimental protocol (Fig. 5) was implemented to ensure data consistency: (i) Briefing: Standardized orientation and informed consent; (i) Practice: A 5–10 minute familiarization ride (excluded from analysis); (i) Main Task: Three riding sessions under NTP, LTP, and HTP conditions, with session order counterbalanced to eliminate sequence bias; and (iv) Rest: 5–8 minute inter-session breaks to stabilize cognitive load and mitigate fatigue. Ethics and Informed Consent: All participants provided written informed consent before participating in the study. The study was approved by the Institutional Ethics Committee of the Indian Institute of Technology Indore (Approval No. BSBE/IITI/IHEC-11/2025/11). Consent to Publish: Written informed consent was obtained from all participants to publish anonymized data and findings from this study. All data have been anonymized to protect the privacy of the participants. 3.1.6 Behavioral Transition : Why LTP Matters The inclusion of a LTP condition captures the intermediate behavioral transition between safe (NTP) and risky (HTP) riding, representing the early stages of behavioral degradation. This continuum follows the progression: NTP → LTP → HTP. While NTP reflects calm and stable riding behavior, and HTP indicates high-stress operation characterized by significant overspeeding and traffic violations, LTP acts as a critical threshold that reveals subtle, early signs of risk—such as mild overspeeding or minor lane deviations. By incorporating LTP, the model improves its sensitivity to rising stress levels, supporting earlier detection and enabling more timely safety interventions. Moreover, LTP data inform policymakers and trainers by defining practical stress thresholds, facilitating preventive strategies before riders reach the high-risk HTP state. 3.2 Feature Engineering and Preprocessing From raw time-series sensor data, we derived 64 domain-informed features that capture vehicle dynamics, rider behavior, and contextual factors as summarized in Table 3. These features include statistics such as mean and standard deviation of speed and acceleration, counts of braking events, lane-change frequency, and TP levels. All features were normalized to the [0,1][0,1] range using Min–Max scaling. We segment the time series data into fixed-length overlapping windows to preserve temporal dependencies, enabling the model to capture both gradual risk accumulation and abrupt behavioral deviations. Detailed training configurations are provided in Section 4.1. 3.3 Dataset and Feature Set The experimental protocol yielded 153 multivariate time-series sessions (51 participants × 3 conditions). Data were sampled at 100 Hz. The 100 Hz streams were segmented into 960 ms windows (96 time steps) with 50% overlap, yielding 129,209 windows slid continuously. A window is labeled as collision (1) if a collision occurs within it, otherwise 0. The dataset exhibits a class distribution of 21% collision and 79% non-collision events. A single collision spans multiple windows (impact+sliding+aftermath, due to overlap), reflecting the inherent imbalance of safety-critical scenarios. This is window-level, not event-level labeling. Each input sequence comprises 64 features (summarized in Table 3) categorized into: (i) Vehicle Dynamics: Speed, acceleration, and 3D rotation; (i) Control Inputs: Throttle, hydraulic brake force, and gear transitions; (i) Proximity: Headway/tailway distances and lane offsets; (iv) Time context and scenario: Time Stamp, TP (0=HTP, 1=LTP, 2=NTP); and (v) Behavioral Violations: Overspeeding, improper gap maintenance and traffic rule infractions. All collisions were aggregated into a single binary target variable, with leakage prevented by excluding impact-related features from the input. Table 3: Summary of simulator features used in this study. Category Feature Names Count Vehicle Controls Ignition, Engine, Accelerator, Brake, Clutch, Handbrake, Steering, Gear, Headlight, Horn Violation 10 Vehicle Performance Speed, RPM, Fuel Economy, Distance Travelled 4 Lighting and Indicators Indicator, Indicated before moving off, Indicated while turning at junction, Indicated while changing lanes, Failed to use headlights 5 Behavioral Violations Over-speeding, Incorrect speed at intersections/junctions, Incorrect speed on speed breakers, Improper gap maintenance, Dangerous overtaking, Turned without indication, Incorrect lane driving, Wrong-side driving, Driving with handbrake applied, Clutch riding, Incorrect gear change sequence, Improper clutch release, Gear shift without clutch, Correct gear before moving off, Smooth releasing of clutch 15 Traffic Rule Violations Crossed white line, Crossed yellow line, Crossed stop line, Signal jumping, No-entry violation, U-turn violation, No-parking violation 7 Time Context and Scenario Time Stamp, TP (0=HTP, 1=LTP, 2=NTP) 2 Spatial Position Position (X, Y, Z), Rotation (X, Y, Z), Lane No., Left Lane Offset, Right Lane Offset 9 Motion and Proximity Lateral Velocity, Longitudinal Velocity, Headway Distance, Headway Time, Tailway Distance, Tailway Time, Leftway Distance, Rightway Distance, Steering Angle 9 Brake Force Brake test done, Front Tire Brake Force, Rear Tire Brake Force 3 Note: All types of collisions features (with vehicles, objects, obstacles, etc.) were combined into a single binary target variable ‘Target‘, with 0 indicating no collision and 1 indicating collision. These collision events were excluded from the input features to prevent label leakage. 3.4 Scope of the Classification and Forecasting Tasks We define two complementary but distinct tasks, and are explicit about what each does and does not establish. Classification Task (Same-Window Risk Detection): Given a 960 ms observation window, the model predicts whether a collision event occurs within that window. Because a subset of positive windows overlaps the collision event itself (impact, sliding, or aftermath phases, as noted in Section 3.3), this task is best characterized as same-window collision risk detection rather than a fixed-lead-time forecast: it evaluates whether pre-collision and in-progress behavioral signatures are jointly separable from normal riding, not how far in advance a collision can be anticipated. Forecasting Task (Prospective Kinematic Prediction): Given a lookback window of L=96L=96 timesteps, a separately trained instance of the architecture predicts future kinematic states (e.g., speed, lean angle, longitudinal force) up to H∈96,192,336,720H∈\96,192,336,720\ timesteps (up to 7.2s) ahead. This task, evaluated in Section 5.2, is a genuine prospective-forecasting result with an explicit lead time, but its target is future sensor state, not future collision probability; it demonstrates that the architecture can anticipate the kinematic precursors of hazardous states, which is complementary to, but distinct from, a direct lead-time collision-probability forecast. We view a labeled, fixed-lead-time collision-probability forecast (i.e., excluding all windows that overlap a collision event and labeling remaining windows by whether a collision begins within a defined horizon τ afterward) as an important direction for follow-up work, and report the current classification results with this scope explicitly stated rather than implied. 3.5 Proposed Architecture: MotoSafety 3.5.1 Architecture Overview Figure 6: The proposed MotoSafety architecture. The proposed MotoSafety architecture, shown in Fig. 6, is a DL architecture based on the Learned Temporal Importance (LTI) principle, developed for the prediction of PTW collisions and forecasting. It employs a multi-stage pipeline: (i) a parallel multi-scale dilated convolutional encoder with three branches (dilation factors 1, 2, 4) extracts local to medium-range temporal patterns; (i) a 2-layer Bi-LSTM (128 hidden units per direction) captures long-range sequential dependencies; (i) a novel TIP module (which uses a lightweight Multi-Layer Perceptron (MLP) as a learned scoring function to evaluate the importance of each timestep) is applied independently to the outputs of both the CNN and LSTM paths, learning to adaptively weight the most critical pre-collision timesteps for each feature type and collapsing them into compact vectors; (iv) these vectors are concatenated and processed by a self-gated fusion stage, refined by channel-wise recalibration via a SE block and MHA (4 heads). The resulting representation is normalized, regularized with dropout, and linearly projected to produce binary collision probabilities. 3.5.2 Comment on the Novelty We distinguish clearly between architectural components adapted from prior work and the components that constitute this paper’s novel contribution. Adapted components: The parallel dilated-convolution encoder, Bi-LSTM, SE recalibration, and MHA are established building blocks, each individually well-studied in time-series and sequence modeling literature. Novel contribution: TIP. The core contribution of this work is the TIP module and the LTI principle it instantiates. Rather than collapsing the temporal dimension with a fixed operator (mean-pool, max-pool, or final-hidden-state extraction, as is standard practice in CNN/RNN-based time-series classifiers), TIP learns a content-aware scoring function that assigns a data-dependent importance weight to every timestep before collapsing the sequence. This is applied independently to both the CNN and BiLSTM representations, allowing the model to learn which pre-collision moments matter most for each feature pathway, rather than treating all timesteps as equally informative. Unlike full self-attention pooling (which computes pairwise timestep interactions at (L2)O(L^2) cost), TIP scores each timestep independently against a shared learned criterion, at (L)O(L) cost, and produces a fixed-size vector consumed by the downstream fusion stage. Applied independently to both CNN and BiLSTM branches, TIP reduces the downstream fusion, SE-recalibration, and attention stages to (1)O(1) with respect to sequence length, while the front-end encoder retains (L)O(L) complexity, giving an overall (L)O(L) model (Table 4). Empirical validation of TIP. The effectiveness of TIP is validated in Section 5.8 (Table 14), where it outperforms standard fixed-pooling strategies (mean-pooling, max-pooling, last-timestep) under an identical architecture, indicating that the performance gain stems from learned, content-aware weighting rather than fixed pooling. This strategic “content-aware collapse” means the fusion, SE-recalibration, and MHA stages operate on a fixed-size 320-dimensional vector regardless of sequence length, i.e., these downstream stages cost (1)O(1) per sample once pooling is complete. The upstream CNN and Bi-LSTM encoder remains (L)O(L), so the model’s end-to-end complexity is (L)O(L) overall, not (1)O(1); the practical benefit of TIP is that it removes the (L2)O(L^2) cost a full-sequence self-attention mechanism would otherwise add on top of the encoder. Table 4 details the per-module accounting. This design ensures the low-latency performance required for robust real-time deployment on edge devices. Table 4: Per-module computational complexity. L: sequence length, C: feature channels, H: hidden dimension, k: kernel size. Module Complexity Dilated CNN encoder (×3 branches) (L⋅k⋅C)O(L· k· C) Bi-LSTM (2 layers) (L⋅H2)O(L· H^2) TIP scoring + weighted sum (L⋅H)O(L· H) Gated fusion + SE block (1)O(1) (fixed-size input) Multi-head attention (post-pooling) (1)O(1) (fixed-size input) MotoSafety total (L)O(L) 3.5.3 MotoSafety: Mathematical Formulation We address the problems of binary collision prediction and multi-horizon time-series forecasting from multivariate sensor data under TP. Our input consists of sensor readings X∈ℝB×L×CX ^B× L× C from a two-wheeler simulator, with B batch size, lookback window L and C feature channels. Forward Process: The input X is processed by two parallel branches: CNN and BiLSTM. The CNN branch captures local patterns using three dilated convolutional branches with kernel size k=3k=3. To accommodate the channel-first requirement of 1D convolutions, we apply a permutation function π(⋅)π(·) to X such that X^=π(X)∈ℝB×C×L X=π(X) ^B× C× L: Hd H_d =ReLU(Conv1Dd(X^)) =ReLU(Conv1D_d( X)) (1) where d=1,2,4d=\1,2,4\ is the dilation factor. These views are merged into a unified 64-channel representation via a 1×11× 1 convolution: Hms=Conv1D1×1([H1;H2;H4])∈ℝB×64×LH_ms=Conv1D_1× 1([H_1;H_2;H_4]) ^B× 64× L. In parallel, a bidirectional LSTM captures long-range dynamics: Hlstm=BiLSTM(X)∈ℝB×L×256.H_lstm=BiLSTM(X) ^B× L× 256. (2) To identify safety-critical moments, we utilize TIP to learn importance weights across the temporal dimension. Two distinct networks, MLPθ1MLP_ _1 and MLPθ2MLP_ _2, assign scores to the CNN and LSTM features respectively: wc w_c =Softmax(MLPθ1(HmsT)),fcnn=∑t=1Lwc(t)Hms(t) =Softmax(MLP_ _1(H_ms^T)), 9.24994ptf_cnn= _t=1^Lw_c^(t)H_ms^(t) (3) wl w_l =Softmax(MLPθ2(Hlstm)),flstm=∑t=1Lwl(t)Hlstm(t) =Softmax(MLP_ _2(H_lstm)), 9.24994ptf_lstm= _t=1^Lw_l^(t)H_lstm^(t) (4) The resulting vectors, fcnn∈ℝ64f_cnn ^64 and flstm∈ℝ256f_lstm ^256, are concatenated to form the joint representation f=[fcnn,flstm]∈ℝ320f=[f_cnn,f_lstm] ^320. The fused representation undergoes refinement through a gating mechanism and a channel-wise recalibration block (SE-Block). Finally, a MHA block performs high-order feature refinement: fgate f_gate =f⊙σ(Wgf+bg), =f σ(W_gf+b_g), (5) fse f_se =SEBlock(fgate), =SEBlock(f_gate), (6) fattn f_attn =MHA(fse). =MHA(f_se). (7) The final prediction ∈ℝ2y ^2 is obtained via a linear projection of the normalized and regularized context: =Softmax(Linear(LayerNorm(Dropout(fattn)))CLOSE.y=Softmax (Linear (LayerNorm(Dropout(f_attn) ) ). (8) To address the extreme class imbalance (rare collision events), we utilize Focal Loss with a focusing parameter γ=2γ=2: ℒcls=−1B∑i=1B(1−Pi,yi)γlog(Pi,yi),L_cls=- 1B _i=1^B(1-P_i,y_i)^γ (P_i,y_i), (9) where Pi,yiP_i,y_i is the predicted probability for the true binary collision label yi∈0,1y_i∈\0,1\. Multi-Horizon Forecasting Module Our forecasting module predicts future target values Y^∈ℝB×τ Y ^B×τ over horizon τ. It uses a separate but architecturally similar network with identical CNN and BiLSTM branches, TIP, gating, SE-Block, and MHA. Both the classification and forecasting modules use the same L=96L=96 timestep lookback window; they differ only in their prediction target (same-window collision label vs. future kinematic states) and are trained separately, as described above. This produces a refined context vector attn∈ℝB×320f_attn ^B× 320. The final forecast is obtained via: Y^=Linearforecast(LayerNorm(attn)). Y=Linear_forecast(LayerNorm(f_attn)). (10) The module is trained to minimize the Mean Squared Error (MSE) loss: ℒforecast=1B⋅τ∑i=1B∑k=1τ(yi,L+k−Y^i,k)2,L_forecast= 1B·τ _i=1^B _k=1^τ (y_i,L+k- Y_i,k )^2, (11) where yi,L+ky_i,L+k is the true target value at future step k. Both models share identical architectural components but are trained separately for their respective tasks. 4 Experimental Setup and Baselines 4.0.1 Ground Truth and Human Validation To ensure the reliability of the simulator-generated labels, a systematic validation protocol was conducted. Two road experts in road safety and human factors (male, post-graduate) reviewed a representative subset of 2,000 randomly selected segments from the total of 129,000-sequence dataset. Experts independently annotated the onset and severity of high-risk episodes (e.g. abrupt swerves, loss of balance, critical near-misses) using a standardized coding scheme based on kinematic thresholds and behavioral cues. The inter-annotator agreement, measured via Cohen’s Kappa coefficient (κ), was 0.87, indicating almost perfect agreement according to conventional benchmarks [40]. This high agreement validates the dataset’s labels as a robust ground truth for model training. 4.0.2 Baselines We benchmark MotoSafety against ML (RF, CNN, RNN) and modern DL architectures: Informer [32] and iTransformer [41] for Transformer-based long-range modeling. TimesNet [42] for multi-scale frequency analysis; PatchTST [43] for patching-based local semantics. Time-LLM [44] and LLM4TS (GPT-2) [45] as a representative fine-tuned LLM baseline. 4.1 Training and Evaluation 4.1.1 Training Protocol The MotoSafety model is implemented in PyTorch and optimized using AdamW (lr=3×10−4lr=3× 10^-4, weight decay 5×10−55× 10^-5) with Focal Loss (γ=2γ=2) for classification and MSE loss for forecasting. To enhance generalization, we employed Mixup augmentation, Exponential Moving Average (EMA) of weights, and Monte Carlo (MC) Dropout (5 samples) for uncertainty calibration. Training was conducted on an NVIDIA T400 GPU (4GB VRAM) with a batch size of 64 for 50 epochs. Key training parameters are summarized in Table 5. All experiments used a participant-wise split (80/10/10), where all sessions from the same rider are kept together in the same split to prevent data leakage. Table 5: Training configuration for MotoSafety classification and forecasting models. Component Setting Classification Model Input sequence length 96 timesteps (960 ms) Feature dimension 64 CNN hidden channels 32 per branch, 64 fused BiLSTM hidden dimensions 128 (bidirectional) Total model parameters ∼ 1.15 million Dropout rate 0.3 Loss function Focal Loss (γ=2γ=2) Forecasting Model Input lookback length 96 timesteps Prediction horizons H∈96,192,336,720H∈\96,192,336,720\ Architecture Identical to classification model Loss function MSE Shared Settings Optimizer AdamW (lr=3×10−43× 10^-4, weight decay=5×10−55× 10^-5) Batch size 64 Training epochs 50 Regularization EMA, Mixup, MC Dropout (5 samples) Hardware NVIDIA T400 GPU (4GB VRAM) 4.1.2 Evaluation Protocol Performance was measured using Accuracy, F1-Score, and ROC-AUC for classification, and MSE and mean absolute error (MAE) for forecasting. Statistical significance was assessed via paired Wilcoxon signed-rank tests over 5 runs, with the Holm–Bonferroni correction (α=0.05α=0.05). 5 Results 5.1 Impact of Time Pressure on Riding Behavior Table 6: Behavioral differences across time pressure conditions: Normal (NTP), Low (LTP), and High (HTP) (Mean ± Std). Feature NTP LTP HTP Speed (km/h) 33.12 ± 19.42 39.69 ± 21.67 49.00 ± 26.48 Sudden braking 1.32 ± 2.54 1.38 ± 2.36 1.80 ± 2.01 Dangerous overtaking 1.07 ± 1.18 1.21 ± 1.27 1.36 ± 1.52 Headway distance (m) 8.99 ± 14.38 8.20 ± 14.34 7.75 ± 13.88 Analysis of our PTW dataset reveals that HTP is the primary driver of behavioral volatility, with a clear transition to high-risk strategies as TP increases, as shown in Table 6. Riders under HTP exhibited 48% higher mean speed, 36% increase in sudden braking, and 14% shorter headway distance compared to NTP, along with a 27% increase in dangerous overtaking. These statistically supported trends define a hurry-up strategy of excessive speed, abrupt control, and reduced spacing, highlighting TP as a key behavioral stressor that increases the risk of PTW collisions. Table 7: Multivariate long-term forecasting results (mean ± std over 5 runs). Input lookback window L=96L=96. Prediction horizons H∈96,192,336,720H∈\96,192,336,720\. Avg is averaged over all four prediction horizons. Models MotoSafety Time-LLM iTransformer PatchTST Metric MSE MAE MSE MAE MSE MAE MSE MAE 96 0.033 ± 0.005 0.082 ± 0.007 0.136 0.148 0.161 0.237 0.188 0.208 192 0.037 ± 0.006 0.086 ± 0.008 0.145 0.153 0.170 0.246 0.205 0.223 336 0.041 ± 0.006 0.097 ± 0.010 0.192 0.215 0.177 0.258 0.227 0.272 720 0.045 ± 0.007 0.111 ± 0.012 0.212 0.293 0.183 0.262 0.267 0.338 Avg 0.039 0.094 0.171 0.210 0.173 0.251 0.222 0.260 5.2 Multivariate Long-Term Forecasting Performance To evaluate the proactive safety capabilities of MotoSafety, we conducted multivariate long-term forecasting across four horizons (H∈96,192,336,720H∈\96,192,336,720\) on our PTW data. This experiment assesses the models ability to project future kinematic states such as lean angles and longitudinal forces, essential for anticipating hazardous transitions before they culminate in a collision. As shown in Table 7, MotoSafety achieves an average MSE of 0.039 and MAE of 0.094, representing a 4.4× reduction in error over the Time-LLM (0.171) and iTransformer (0.173) baselines. This high forecasting fidelity under TP is critical for proactive safety; by accurately projecting states up to H=720H=720, the model provides a vital safety buffer for early risk mitigation before a maneuver reaches a non-recoverable limit. 5.3 Comparative Classification Performance Table 8 summarizes the performance of the models on the PTW simulation data under TP. The proposed MotoSafety achieves 94.97% accuracy, 93.7% F1-score and 99.33% ROC AUC, outperforming ten baselines including TimesNet, PatchTST, iTransformer, Time-LLM and LLM4TS. To assess the robustness of the performance gains, statistical significance is evaluated using paired Wilcoxon signed-rank tests over 5 independent runs (Table 8). To account for the nine pairwise comparisons against baseline models, we employ the Holm–Bonferroni correction to control the family-wise error rate (FWER) at α=0.05α=0.05. The high ROC AUC indicates that MotoSafety can maintain a low false-alarm rate, which is a prerequisite for real-world Advanced Rider Assistance Systems (ARAS). Table 8: Performance comparison of all models (mean ± std over 5 runs). Asterisks indicate significant differences (paired Wilcoxon signed-rank with Holm–Bonferroni correction, α=0.05α=0.05): ∗p<0.001^***p<0.001, p∗∗<0.01^**p<0.01, ∗p<0.05^*p<0.05. All tests compare each baseline model against MotoSafety. Model Acc. (%) Prec. Rec. F1-Score (%) AUC (%) Mean ± Std Result Result Mean ± Std Mean ± Std RF 84.93±0.16∗84.93± 0.16 85.24 84.87 85.02±0.23∗85.02± 0.23 93.68±0.31∗93.68± 0.31 CNN 90.19±0.11∗90.19± 0.11 93.82 80.68 86.74±0.21∗86.74± 0.21 97.62±0.22∗97.62± 0.22 RNN 90.94±0.12∗90.94± 0.12 83.41 96.38 89.43±0.32∗89.43± 0.32 97.27±0.34∗97.27± 0.34 LLM4TS 90.06±0.14∗90.06± 0.14 90.58 83.23 86.81±0.31∗86.81± 0.31 97.31±0.42∗97.31± 0.42 Informer 93.61±0.07∗93.61± 0.07 89.72 94.88 92.21±0.22∗∗92.21± 0.22 99.12±0.21∗∗99.12± 0.21 TST 93.86±0.09∗93.86± 0.09 93.67 90.74 92.18±0.29∗∗92.18± 0.29 99.23±0.28∗99.23± 0.28 TimesNet 94.06±0.05∗94.06± 0.05 89.23 96.69 92.77±0.38∗∗92.77± 0.38 99.18±0.24∗99.18± 0.24 iTransformer 94.39±0.06∗∗94.39± 0.06 92.41 93.57 93.12±0.42∗93.12± 0.42 99.22±0.32∗99.22± 0.32 PatchTST 94.41±0.05∗∗94.41± 0.05 90.73 96.18 93.28±0.24∗93.28± 0.24 99.31±0.2299.31± 0.22 Time-LLM 94.46±0.07∗94.46± 0.07 93.12 93.79 93.41±0.33∗93.41± 0.33 99.24±0.31∗99.24± 0.31 MotoSafety 94.97±0.0894.97± 0.08 93.8093.80 93.9093.90 93.70±0.2093.70± 0.20 99.33±0.3099.33± 0.30 5.4 Forecasting Horizon Analysis To characterize how forecasting accuracy varies with prediction horizon, Table 9 reports MSE and MAE for the kinematic-state forecasting task (Section 3.5.3) across horizons up to 7.2s. Error increases gradually with horizon (MSE from 0.033 at H=96H=96 to 0.045 at H=720H=720), indicating the model retains meaningful predictive signal even at longer horizons rather than degrading sharply. We note this result characterizes future kinematic-state prediction and is reported separately from the same-window collision classification task described in Section 3.3; the two tasks use different targets and should not be read as a single fused collision-lead-time result. Table 9: Forecasting performance by prediction horizon (mean over 5 runs) Horizon H Time (s) MSE MAE 96 0.96 0.033 ± 0.005 0.082 ± 0.007 192 1.92 0.037 ± 0.006 0.086 ± 0.008 336 3.36 0.041 ± 0.006 0.097 ± 0.010 720 7.20 0.045 ± 0.007 0.111 ± 0.012 5.5 Downstream Task: Impact of TP on Model Performance Table 10: Performance comparison of MotoSafety showing improvement from adding TP features. Significance markers indicate improvement over the Without TP baseline: ∗p<0.001^***p<0.001 Model Accuracy (%) AUC (%) Without TP Feature 94.09±0.0494.09± 0.04 99.10±0.0199.10± 0.01 With Predicted TP Feature 94.82±0.02∗94.82± 0.02^*** 99.24±0.01∗99.24± 0.01^*** With Ground Truth TP Feature 94.97±0.01∗94.97± 0.01^*** 99.33±0.01∗99.33± 0.01^*** In practical deployment, direct and accurate measurement of ground truth TP labels, denoted TPgtTP_gt, may be infeasible due to sensor limitations or real-time constraints. In the proposed data set, the TP characteristic is a categorical variable representing three distinct TP conditions: HTP = 0, LTP = 1, NTP = 2. To address missing or unavailable GT TP labels, we employed a Time Series Transformer (TST) [46] model to predict these categorical TP labels from the multivariate time series input data. The TST achieved 89.26% accuracy for 3-class TP prediction (0=HTP, 1=LTP, 2=NTP). Table 11 shows the confusion matrix (raw counts) per-class performance. Table 11: Confusion matrix (raw counts) for TST-based TP prediction. True / Predicted 0: HTP 1: LTP 2: NTP 0: HTP 8878 229 163 1: LTP 76 7326 568 2: NTP 82 1661 6859 Let =(1,2,…,T)∈ℝT×DX=(x_1,x_2,…,x_T) ^T× D denote the multivariate time series input from the PTW simulator excluding any explicit TP feature, where T represents the sequence length and D=63D=63 is the number of telemetry features. Each vector t∈ℝDx_t ^D represents the synchronized sensor states at time step t. The TST model is defined as: TPpred=fTST(),TP_pred=f_TST(X), (12) where TPpred∈0,1,2TP_pred∈\0,1,2\ is the predicted TP label. The predicted TP feature TPpredTP_pred, along with all other relevant sensor and behavioral features of the data set, forms the complete input to our collision risk prediction model, MotoSafety. Formally, the model can be expressed as: y^=gMotoSafety(,TP), y=g_MotoSafety(X,TP), (13) where TPTP is the ground truth TPgtTP_gt or predicted TPpredTP_pred, and y y is the predicted safety risk label for the rider. We evaluated the robustness of model by comparing the predictions when using: y^gt y_gt =gMotoSafety(,TPgt), =g_MotoSafety(X,TP_gt), (14) y^pred y_pred =gMotoSafety(,TPpred). =g_MotoSafety(X,TP_pred). (15) The experimental results in Table 10 show the importance of TP in collision risk prediction. While the baseline model fMotosafety(X)f_Motosafety(X) achieves an accuracy 94.09% using only PTW simulation data without TP feature. The integration of predicted TP features fMotosafety(X⊕TPpred)f_Motosafety(X TP_pred) improved this to 94.82%, nearly matching the 94.97% Oracle performance fMotosafety(X⊕TPgt)f_Motosafety(X TP_gt). This performance gain shows that TPpredTP_pred captures latent psychological context, such as rider stress, which is not fully represented by telemetry alone. The marginal 0.15% gap between the predicted (94.82%) and Oracle (94.97%) results validates MotoSafety for real-world deployment where ground truth states are not available. 5.6 Model Performance and Reliability 5.6.1 Confusion Matrix Analysis Figure 7: Confusion matrix for the MotoSafety model on the primary PTW simulator dataset under time pressure. The confusion matrix provides a detailed view of the predictive performance of MotoSafety across collision and non-collision events. Fig. 7 visualizes the distribution of correct and incorrect predictions. 5.6.2 Calibration Curve and Risk Scores The MST calibration curve on the primary PTW simulator dataset under TP (Fig. 9) shows that predicted collision probabilities closely match observed event frequencies, indicating statistically reliable and well-calibrated outputs. In safety-critical contexts, such calibration enables risk-aware adaptive interventions rather than reliance on fixed thresholds. The distribution of predicted risk scores by true outcome (Fig. 9) further illustrates clear separation: collision events consistently receive higher predicted probabilities, while non-collision samples remain concentrated at lower values. This separation reinforces both discriminative ability and calibration quality, supporting the potential for real-world deployment. Figure 8: Calibration curve. Figure 9: Distribution of predicted collision risk. 5.7 Architecture Transferability Across Safety Domains Table 12: Cross-Safety-Domain Evaluation of MotoSafety Accuracy (%) and Prior Work Accuracy (%). MST: MotoSafety. Dataset Details Safety Domain Application Country MST Prior Work Accuracy Proposed Dataset PTW Simulator with GT TP PTW Safety Collision Prediction India 94.97 – [47] PTW High-Fidelity Simulator PTW Safety Collision Prediction Germany 99.49 91.0 [RF, GB [31]] [48] PTW Real Time Fall Event Data PTW Safety Fall Detection France 96.10 91.59 [DT [49]] [50] Human Activity Recognition Human Safety Activity Recognition Italy 97.66 83.35 [K-means + NB [51]] [52] Wearable-Sensor Movement Human Clinical Safety Exercise Quality Assessment Turkey 99.65 88.65[MTMM-DTW [53]] Note that each dataset below is used to train and evaluate a separate MotoSafety instance from scratch (i.e., no weights are transferred from the PTW model); this experiment therefore evaluates the transferability of the architecture to other safety-critical sequence classification tasks, rather than knowledge transfer from the PTW domain itself. To evaluate the versatility of the MotoSafety framework beyond motorcycle simulators, we performed cross-dataset evaluations across distinct safety-critical domains as shown in Table 12. In the PTW Safety domain, MotoSafety achieved 99.49% accuracy on the high-fidelity German simulator dataset [47] for collision prediction, outperforming traditional ensemble methods (RF/GB) by 8.49%. Similarly, on real-world fall event data [48], achieved improved 96.10% accuracy, a significant improvement over the 91.59% reported for Decision Tree (DT) benchmarks [49]. Beyond vehicle dynamics, the model exhibited exceptional robustness in Human and Clinical Safety. On the UCI HAR human activity recognition dataset [50], MotoSafety surpassed Naive Bayes (83.35%) by 14.31%, achieving 97.66% accuracy. Most notably, in clinical exercise quality assessment [52], the framework achieved near-perfect classification (99.65%), marking an 11% gain over DTW-based clinical algorithms (88.65%) [53]. These results demonstrate that the MotoSafety architecture is not overfit to PTW-simulator-specific artifacts and performs competitively when retrained on diverse safety-critical sequence classification tasks, supporting the architecture’s applicability beyond its original design domain. 5.8 Ablation Study We evaluated individual contributions of MotoSafety components through a comprehensive ablation study shown in Table 13. The study was carried out on PTW simulation data under TP. The results show that the full model achieves an accuracy of 94.97%. The removal of BiLSTM and MHA resulted in significant performance degradation of 4.55% and 4.19%, respectively, highlighting their role in capturing long-range dependencies. Similarly, the exclusion of the CNN branch and Gated Fusion mechanism led to accuracy drops of 2.31% and 2.07%, underscoring the necessity of multi-scale feature extraction. Furthermore, removing the SE block decreases accuracy by 3.73%, validating its effectiveness in channel-wise feature recalibration. Collectively, these findings justify the necessity of each integrated module to maintain high predictive precision. Table 13: Ablation Study of MotoSafety Components. Variant Accuracy (%) Δ Acc. Full Model 94.97 – No BiLSTM 90.42 −4.55-4.55 No SE Block 91.24 −3.73-3.73 No Attention 90.78 −4.19-4.19 No CNN Branch 92.66 −2.31-2.31 No Gated Fusion 92.90 −2.07-2.07 Table 14: Comparison of MotoSafety with TIP against standard fixed temporal-pooling strategies (mean ± std over 5 runs). Pooling Strategy Accuracy (%) AUC (%) Mean-pooling 92.23 ± 0.15 98.80 ± 0.22 Max-pooling 93.41 ± 0.12 99.11 ± 0.18 Last-timestep 92.87 ± 0.14 98.97 ± 0.20 TIP 94.97 ± 0.08 99.33 ± 0.30 To further isolate the effect of TIP, we compare it against three standard temporal-pooling strategies—mean-pooling, max-pooling, and last-timestep extraction—under an otherwise identical architecture. As shown in Table 14, MotoSafety with TIP achieves 94.97% accuracy, outperforming mean-pooling by 2.74%, max-pooling by 1.56%, and last-timestep by 2.10%. This indicates that the performance gain stems from the learned, content-aware weighting rather than from the surrounding architecture alone. 5.9 Real-World Feature Feasibility To assess real-world deployment feasibility, we identified that 21 of the 64 features used in this study are available on standard motorcycles via low-cost IMU+GPS (e.g., speed, acceleration, yaw rate, position, brake force, throttle, steering angle). The remaining features (e.g., precise gap metrics, rule violations) are simulator-only. As shown in Table 15, accuracy increases with more features, achieving 93.91% with 21 features (compared to 94.97% with all 64 features), indicating that real-world deployment is feasible with onboard sensors. Validation on physical hardware in real traffic remains necessary and is identified as future work in Section 6.1. Table 15: MotoSafety accuracy with real-bike features (IMU+GPS). Features Accuracy (%) 3 80.22 5 83.11 7 88.78 9 90.16 21 93.91 All 64 (simulator) 94.97 5.10 Model Complexity and Edge-AI Deployability The MotoSafety model is designed for real-time deployment in resource-constrained environments. Latency was measured on an Intel Core i7 CPU (2.8 GHz) with batch size 32, PyTorch 2.0, CPU-only inference, averaged over 1000 batches. With only 1.15M parameters and a 4.38 MB memory footprint (Table 16A), it maintains a low computational profile without sacrificing accuracy. Benchmarking on our PTW simulation dataset reveals that MotoSafety achieves a per-sample inference latency of 0.135 ms, outperforming all baseline and SOTA architectures (Table 16B). Specifically, MotoSafety is 5.3×, 9.5×, 21.9×, and 71.9× faster than PatchTST, Time-LLM, TimesNet, and LLM4TS (GPT-2), respectively. Even when accounting for TST preprocessing (36.88ms overhead) to predict TP labels when ground truth is unavailable, MotoSafety (37.01ms) remains faster than iTransformer (37.47ms), PatchTST (37.60ms), Time-LLM (38.16ms), TimesNet (39.84ms), and LLM4TS (46.59ms). This ultra-low latency ensures near-instantaneous risk assessment, allowing maximum time for emergency interventions. These efficiency metrics satisfy the stringent requirements for real-time deployment on low-cost on-board units (OBUs). Unlike GPU-dependent Transformer baselines, complexity of MotoSafety is (L)O(L). This balance of efficiency and performance makes it a viable candidate for large-scale ITS deployment in developing regions, providing an accessible AI-driven safety net for vulnerable riders. Table 16: MotoSafety Efficiency Metrics and Comparative Latency. A. MotoSafety Efficiency & Deployability Metric Value Significance Total Parameters 1,149,048 Low Complexity Model Size 4.38 MB Edge-ready Storage Throughput >7,400>7,400 s/sec High-speed Inference Arch. Complexity (L)O(L) Linear Scalability Edge-AI Deployability Yes Suitable for wearable Deployment Regions Global Operate in low-resource B. Latency Comparison (PTW Simulator Dataset Under TP) Model (Year) Latency (ms) Speed Gap Parameters MotoSafety 0.135 1.0× 1,149,048 iTransformer (2024) 0.587 4.3× ↓ 275,714 PatchTST (2023) 0.720 5.3× ↓ 187,330 Time-LLM (2024) 1.280 9.5× ↓ 3,183,618 TimesNet (2023) 2.960 21.9× ↓ 4,708,226 LLM4TS (2025) 9.710 71.9× ↓ 124,506,626 6 Conclusion This work studied the effect of TP on PTW riders and its influence on collision risk. To address this gap, we introduced a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides involving 51 participants under no, low, and high TP conditions. Each sequence captures 64 features spanning vehicle dynamics, control inputs, proximity, temporal context, and behavioral violations. Building on this dataset, we proposed MotoSafety, a novel deep learning architecture grounded in the LTI principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC for collision risk assessment, outperforming ten baselines including TimesNet, PatchTST, iTransformer, Time-LLM, and LLM4TS. For long-term forecasting, it achieves 0.039 MSE and 0.094 MAE (4.4× lower error than Time-LLM and iTransformer). With only 1.15 million parameters and 0.135 ms inference latency, the model is 21.9× and 71.9× faster than TimesNet and LLM4TS. Our findings show that explicit TP prediction provides a critical inductive bias, improving collision accuracy from 94.09% to 94.97% (ground truth) and 94.82% with predicted TP. Furthermore, using only 21 real-world features (available via low-cost IMU+GPS), MotoSafety achieves 93.91% accuracy, indicating practical deployment potential; real-world hardware validation remains future work. Beyond PTW safety, the MotoSafety architecture shows improved transferability when retrained on human activity recognition (97.66%) and clinical exercise monitoring (99.65%) datasets. The model’s small size (1.15M parameters, 4.38 MB) and low latency (0.135 ms) make it well-suited for future edge deployment on handlebar-mounted devices or smart helmets as a potential edge-deployable black box for two-wheelers, enabling real-time collision risk alerts without relying on cloud connectivity. 6.1 Limitations and Future Work This study has several limitations that should be acknowledged. First, the participant cohort is entirely male, consistent with Indian statistics where males constitute 85.2–87.3% of PTW fatalities. However, this restricts generalizability to female riders, as TP perception and riding behavior may differ across genders. Future studies should include female riders to systematically examine whether the observed relationships hold across genders. Second, data were collected in a controlled simulator environment, which may not fully represent real-world conditions such as fatigue, weather variations, or social pressures. Real-world TP collision data cannot be collected safely, so simulators provide the only ethical controlled setting. To validate in real conditions, we plan phased testing: closed-course trials followed by naturalistic data collection through commercial riding platforms. Third, we did not measure physiological stress indicators (e.g., heart rate variability, electrodermal activity). Behavioral markers such as speed variability and control inputs were used as indicators of cognitive load; however, the absence of physiological validation limits direct assessment of the underlying stress response. Future work should incorporate physiological measures to complement behavioral markers. Fourth, the proposed safety system requires prospective validation in real-world operational settings before practical deployment. Hardware validation on handlebar-mounted devices or smart helmets remains future work. Future work will validate MotoSafety across diverse populations and road environments using both simulator and naturalistic data. Multimodal signals (physiological, behavioral, contextual) will be integrated to enhance robustness. Adaptive ITS interventions (haptics, throttle modulation, context-aware alerts) will be developed to manage crash risk and kinetic energy transfer. Transfer learning and domain adaptation will be applied for scalability across vehicle types, regions, and cultures. Data and Code Availability All data and code used will be made available from the corresponding author on a reasonable request upon publication. CRediT Authorship Contribution Statement Sumit S. Shevtekar: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing. Dr. Chandresh K. Maurya: Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Supervision, Validation, Writing – review & editing. Dr. Gourab Sil: Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Resources, Supervision, Validation, Writing – review & editing. Dr. Subasish Das: Methodology, Supervision, Validation, Writing – review & editing. Funding This work was supported by the IIT Indore Young Faculty Research Catalyzing Grant (YFRCG) Scheme [Project ID: IITI/YFRCG/2023-24/01]. Acknowledgments The authors sincerely thank all volunteers who participated in this study. Special thanks to Manvendra T. for his support in data collection. Declaration of Competing Interest The authors declare no conflict of interest. References [1] World Health Organization. Global status report on road safety 2023. Geneva: WHO; 2023. [2] Ministry of Road Transport and Highways. Road accidents in India 2023-24. New Delhi: MoRTH; 2024. [3] Ministry of Road Transport and Highways. Road accidents in India 2024-25. New Delhi: MoRTH; 2025. [4] Ministry of Road Transport and Highways. Road accidents in India 2022-23. New Delhi: MoRTH; 2023. [5] Ministry of Road Transport and Highways. Road accidents in India 2019. New Delhi: Transport Research Wing; 2019. [6] Ministry of Road Transport and Highways. Road accidents in India 2020. New Delhi: Transport Research Wing; 2020. [7] Ministry of Road Transport and Highways. Road accidents in India 2021. New Delhi: Transport Research Wing; 2021. [8] Ministry of Road Transport and Highways. Road accidents in India 2022. New Delhi: Transport Research Wing; 2022. [9] Ministry of Road Transport and Highways. Road accidents in India 2023. New Delhi: Transport Research Wing; 2023. [10] Sharma I, Gupta M, Mishra S, Velaga NR. Exploring the impact of time pressure on motorized two-wheeler riders’ over-speeding behavior. Transp Lett. 2025;17(4):595–611. [11] Kong X, Das S, Jha K, Zhang Y. Understanding speeding behavior from naturalistic driving data: Applying classification based association rule mining. Accid Anal Prev. 2020;144:105620. [12] Pawar NM, Velaga NR. Analyzing the impact of time pressure on drivers’ safety by assessing gap-acceptance behavior at un-signalized intersections. Saf Sci. 2022;147:105582. [13] Fortune India Staff. Racing against time: Quick commerce is pushing delivery riders to the edge, claims study. Fortune India. 2025 Feb. [14] The Hindu Staff Reporter. Rash driving, fall in income, more pressure: Behind the scenes of 10-minute delivery. The Hindu. 2022 Mar. [15] Reuters. Delivery race brings road safety risks in India. Jakarta Post. 2022 Jan. [16] Gupta M, Velaga NR. Motorized two-wheeler riders’ rear brake application in sudden hazardous event of animal crossing. Transp Lett. 2024;16(10):1268–1275. [17] Pavlidis I, Dcosta M, Taamneh S, Manser M, Ferris T, Wunderlich R, et al. Dissecting driver behaviors under cognitive, emotional, sensorimotor, and mixed stressors. Sci Rep. 2016;6:25651. [18] Mostafa AM, Aldughayfiq B, Tarek M, Alaerjan AS, Allahem H, Elbashir MK, et al. AI-based prediction of traffic crash severity for improving road safety and transportation efficiency. Sci Rep. 2025;15:27468. [19] Pan G, Wang G, Wei H, Chen Q, Zhang A. Development of an automated global crash prediction model with adaptive feature selection of deep neural networks. IEEE Trans Ind Inform. 2024;20:12010–12020. [20] Bouhsissin S, Sael N, Benabbou F, Soultana A. Enhancing machine learning algorithm performance through feature selection for driver behavior classification. Indones J Electr Eng Comput Sci. 2024;35(1):354–365. [21] Acı Çİ, Mutlu G, Ozen M, Acı M. Enhanced multi-class driver injury severity prediction using a hybrid deep learning and random forest approach. Appl Sci. 2025;15. [22] Jiang Y, Qu X, Zhang W, Guo W, Xu J, Yu W, et al. Analyzing crash severity: Human injury severity prediction method based on transformer model. Vehicles. 2025;7(1):5. [23] Pawar NM, Velaga NR, Mishra S. Impact of time pressure on acceleration behavior and crossing decision at the onset of yellow signal. Transp Res Part F Traffic Psychol Behav. 2022;87:1–18. [24] Pawar NM, Velaga NR, Sharmila RB. Exploring behavioral validity of driving simulator under time pressure driving conditions of professional drivers. Transp Res Part F Traffic Psychol Behav. 2022;89:29–52. [25] Gashaw S, Goatin P, Härri J. Modeling and analysis of mixed flow of cars and powered two wheelers. Transp Res Part C Emerg Technol. 2018;89:148–167. [26] Gupta M, Pawar NM, Velaga NR, Mishra S. Modeling distraction tendency of motorized two-wheeler drivers in time pressure situations. Saf Sci. 2022;154:105820. [27] Leung S, Croft RJ, Jackson ML, Howard ME, McKenzie RJ. A comparison of the effect of mobile phone use and alcohol consumption on driving simulation performance. Traffic Inj Prev. 2012;13(6):566–574. [28] Bham GH, Leu MC. A driving simulator study to analyze the effects of portable changeable message signs on mean speeds of drivers. J Transp Saf Secur. 2018;10(1-2):45–71. [29] Li X, Oviedo-Trespalacios O, Rakotonirainy A, Yan X. Collision risk management of cognitively distracted drivers in a car-following situation. Transp Res Part F Traffic Psychol Behav. 2019;60:288–298. [30] Bouhsissin S, Sael N, Benabbou F, Soultana A. Enhancing machine learning algorithm performance through feature selection for driver behavior classification. Int J Electr Comput Eng Syst. 2023;35(1):354–365. [31] Rodegast P, Maier S, Kneifl J, Fehr J. On using machine learning algorithms for motorcycle collision detection. Discov Appl Sci. 2024;6(6):326. [32] Zhou H, Zhang S, Peng J, Zhang S, Li J, Xiong H, et al. Informer: Beyond efficient transformer for long sequence time-series forecasting. In: Proceedings of the AAAI Conference on Artificial Intelligence. 2021;35:11106–11115. [33] Zeng A, Chen M, Zhang L, Xu Q. Are transformers effective for time series forecasting? In: Proceedings of the Thirty-Seventh AAAI Conference on AI and Thirty-Fifth Conference on Innovative Applications of AI and Thirteenth Symposium (AAAI’23/IAAI’23/EAAI’23). AAAI Press; 2023. [34] World Health Organization. Road traffic injuries. Geneva: WHO; 2023. [35] Krug E. It’s time to end deaths on our roads. Geneva: WHO; 2022. [36] International Electrotechnical Commission. Potentiometers for use in electronic equipment. Standard IEC 60393. Geneva: IEC; 2023. [37] Gupta M, Velaga NR. Dynamic dilemma zone at signalized intersection: attention allocation patterns using cure survival analysis for male riders. Accid Anal Prev. 2026;228:108408. [38] Transportation Research and Injury Prevention Centre. Road safety in India. New Delhi: IIT Delhi; 2023. [39] Atif M, Sil G. Modeling the effects of driver and road geometric characteristics on consecutive horizontal curve perception. Transp Res Rec. 2026;2680(3):286–310. [40] Landis JR, Koch G. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159–174. [41] Liu Y, Hu T, Zhang H, Wu H, Wang S, Ma L, et al. iTransformer: Inverted transformers are effective for time series forecasting. arXiv. 2024. [42] Wu H, Hu T, Liu Y, Zhou H, Wang J, Long M. TimesNet: Temporal 2D-variation modeling for general time series analysis. arXiv. 2023. [43] Nie Y, Nguyen NH, Sinthong P, Kalagnanam J. A time series is worth 64 words: Long-term forecasting with transformers. arXiv. 2023. [44] Jin M, Wang S, Ma L, Chu Z, Zhang JY, Shi X, et al. Time-LLM: Time series forecasting by reprogramming large language models. arXiv. 2024. [45] Chang C, Wang WY, Peng WC, Chen TF. LLM4TS: Aligning pre-trained LLMs as data-efficient time-series forecasters. ACM Trans Intell Syst Technol. 2025;16(3). [46] Zerveas G, Jayaraman S, Patel D, Bhamidipaty A, Eickhoff C. A transformer-based framework for multivariate time series representation learning. In: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2021. p. 2114–2124. [47] Rodegast M, et al. Motorcycle collision dataset. 2024. [48] Boubezoul A, Dufour F, Bouaziz S, Larnaudie B, Espié S. Dataset on powered two wheelers fall and critical events detection. Data Brief. 2019;23:103828. [49] Elwy F, Aburukba R, Al-Ali AR, Al Nabulsi A, Tarek A, Ayub A, et al. Data-driven safe deliveries: The synergy of IoT and machine learning in shared mobility. Future Internet. 2023;15(10). [50] Reyes-Ortiz J, Anguita D, Ghio A, Oneto L, Parra X. Human Activity Recognition Using Smartphones [dataset]. UCI Machine Learning Repository. 2013. [51] Ismi DP, Panchoo S, Murinto M. K-means clustering based filter feature selection on high dimensional data. Int J Adv Intell Inform. 2016;2:38–45. [52] Yurtman A, Barshan B. Physical Therapy Exercises [dataset]. 2014. [53] Yurtman A, Barshan B. Automated evaluation of physical therapy exercises using multi-template dynamic time warping on wearable sensor signals. Comput Methods Programs Biomed. 2014;117(2):189–207.