Paper deep dive
MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure
Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/20/2026, 4:05:17 AM
Summary
The paper introduces MotoSafety, an edge-AI framework for assessing collision risk in powered two-wheeler (PTW) riders under time pressure (TP). Utilizing a novel dataset of 129,000 labeled multivariate time-series samples from 51 participants across No, Low, and High TP scenarios, the study proposes a lightweight architecture based on Learned Temporal Importance (LTI). MotoSafety achieves high accuracy (94.97%) and low latency (0.135 ms), making it suitable for edge deployment on low-cost hardware. The framework outperforms existing baselines and demonstrates transferability to other domains like human activity and clinical monitoring.
Entities (10)
Relation Signals (8)
MotoSafety → hasaccuracy → 94.97%
confidence 98% · MotoSafety achieves 94.97% accuracy
MotoSafety → assessesriskfor → Powered Two-Wheeler
confidence 96% · MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment
MotoSafety → haslatency → 0.135 ms
confidence 95% · 0.135 ms latency, it is suitable for edge deployment
MotoSafety → usesconcept → Learned Temporal Importance
confidence 95% · MotoSafety, a new edge-AI framework built on the Learned Temporal Importance (LTI) concept.
Time Pressure → influences → Collision Risk
confidence 94% · limited studies exist on how cognitive stressors such as Time Pressure influence collision risk.
MotoSafety → outperforms → LLM4TS
confidence 90% · outperforming ten baselines, including TimesNet and LLM4TS
MotoSafety → outperforms → TimesNet
confidence 90% · outperforming ten baselines, including TimesNet and LLM4TS
MotoSafety → supports → Safe System Approach
confidence 85% · supporting the Safe System Approach for Intelligent Transportation Systems
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. We address this gap by introducing a comprehensive dataset consisting of over 129,000 labeled multivariate time-series samples, gathered across 153 simulator rides from 51 participants under No, Low, and High TP scenarios. Across each sequence, we capture 64 distinct attributes covering vehicle motion, rider control actions, spatial proximity, and rule compliance indicators. Using this dataset, we introduce MotoSafety, a new edge-AI framework built on the Learned Temporal Importance (LTI) concept. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4x lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems.
Tags
Links
- Source: https://arxiv.org/abs/2608.17823v2
- Canonical: https://arxiv.org/abs/2608.17823v2
Trouble viewing inline? Open PDF directly →
Full Text
75,903 characters extracted from source content.
Expand or collapse full text
MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure Sumit S. Shevtekar Affiliation: Department of Computer Science and Engineering, Indian Institute of Technology Indore, Khandwa Road, Simrol, Indore, 453552, Madhya Pradesh, India Corresponding author: Corresponding author. E-mail: sumit.shevtekar@gmail.com Chandresh K. Maurya Affiliation: Department of Computer Science and Engineering, Indian Institute of Technology Indore, Khandwa Road, Simrol, Indore, 453552, Madhya Pradesh, India Gourab Sil Affiliation: Department of Civil Engineering, Indian Institute of Technology Indore, Khandwa Road, Simrol, Indore, 453552, Madhya Pradesh, India Subasish Das Affiliation: Civil Engineering Program, Ingram School of Engineering, Texas State University, RFM 5202, San Marcos, 78666, TX, USA Abstract Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. We address this gap by introducing a comprehensive dataset consisting of over 129,000 labeled multivariate time-series samples, gathered across 153 simulator rides from 51 participants under No, Low, and High TP scenarios. Across each sequence, we capture 64 distinct attributes covering vehicle motion, rider control actions, spatial proximity, and rule compliance indicators. Using this dataset, we introduce MotoSafety, a new edge-AI framework built on the Learned Temporal Importance (LTI) concept. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4× lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems. Keywords: Powered two-wheeler safety , Time pressure , Collision risk assessment , Deep learning , Intelligent transportation systems †graphicalabstract: 1 Introduction Road traffic crashes are a major global health concern, particularly affecting Low- and Middle-Income Countries (LMICs), where the mortality burden is high [46]. Powered two-wheeler (PTW) riders represent one of the most vulnerable groups, facing disproportionately high fatality rates due to limited protection, complex traffic, and socio-economic pressures to meet mobility and livelihood deadlines. Globally, PTWs account for approximately 21% of road traffic fatalities, a figure that rises to nearly 38% in LMICs, where PTWs serve as an affordable and widely adopted mode of transport [46]. In India alone, PTWs constituted 74.4% of registered vehicles as of 2022 shown in Fig. 2 [29]. As of 2023, two-wheelers constitute the largest segment of the India’s vehicle population with more than 263 million registered units and were involved in 44.5% of total road traffic fatalities, as shown in Fig. 2 [30, 27]. Official MoRTH data from 2019 to 2023 consistently indicates that male riders account for the vast majority of road fatalities in India, with proportions varying between 85.2% and 87.3% across these years [23, 24, 25, 26, 28]. The highest concentration of these fatalities is observed within the 18–45 age demographic. These patterns underscore the elevated vulnerability of male PTW riders in India’s traffic environment. Despite advances in infrastructure design and traffic enforcement, human error remains the dominant contributor to PTW crashes, with risky behaviors such as overspeeding, abrupt maneuvering, and inadequate hazard response accounting for a large share of fatalities [43]. Naturalistic and simulator-based studies consistently show that riding speed, control variability, and situational context strongly influence crash risk [17]. The operating speeds of PTW riders are considerably higher than those of four-wheeler drivers, often reaching up to 2.3 times the speeds observed in cars [37]. Figure 1: Registered Vehicles Share. Figure 2: Road Crash Fatalities in India. Time Pressure (TP) refers to the subjective perception of insufficient time to complete a riding task, commonly arising from risk-prone behaviors, situational urgency, or externally imposed constraints [37, 44]. This is especially common for app-based delivery riders in LMICs such as India. App-generated delivery deadlines, performance incentives, and intense customer demands for fast, flawless service combine to create persistent and often extreme TP [44, 38, 39]. Elevated TP affects cognitive processing and motor control, narrowing attentional bandwidth and accelerating decision-making. This in turn promotes risk-prone behaviors such as overspeeding, inconsistent braking, and delayed hazard response [37, 11, 43]. Simulator studies on car drivers report increased accident risk of up to 181% under high time pressure (HTP) compared to no time pressure (NTP) [37]. Related work [34] demonstrates that cognitive and emotional stress amplifies speed variability and control instability. However, a key cause of these behaviors, TP-induced cognitive stress, is still not well understood within the PTW-driving situation. Recent AI-driven traffic crash prediction leverages large-scale datasets, classical ML/DL architectures to improve safety [31, 33, 5, 1, 15]. Although effective in predicting four-wheeler crashes, these approaches inadequately model the unique dynamics and highly unstable control characteristics of PTWs. For example, [33] achieved ∼ 90% accuracy using deep neural networks on highway data, [5] improved driver behavior classification with feature selection on large datasets, [1] used Random Forest to predict injury severity (92% accuracy). However, these studies focus primarily on four-wheelers or general traffic datasets and do not incorporate fine-grained temporal modeling of PTW riders or cognitive stress factors such as TP. To our knowledge, this is the first study to focus on predicting collision risk for two-wheeler riders under TP conditions. To address this gap, we introduce the MotoSafety specifically designed for multivariate time series data and to quantify and predict collision risk by modeling the interplay between a riders cognitive state under TP and their operational maneuvers. Our work addresses a critical public safety exigency: the protection of vulnerable PTW riders operating under chronic TP—a phenomenon increasingly prevalent in the rapidly urbanizing and dense traffic environments of emerging economies like India. The proposed system can be deployed as an edge-based alert system, where a handlebar-mounted device or smart helmet provides haptic/audio alerts when collision risk exceeds a threshold, enabling proactive safety interventions for riders under TP (e.g., emergency commuting, delivery deadlines, or urgent travel). This methodology was co-developed with road safety authorities and transportation experts to address PTW fatalities in LMICs. By integrating stakeholder input on crash-prone scenarios and demographic vulnerability, we ensure that our edge-deployable model satisfies real-world operational requirements. Consequently, our approach transitions from theoretical inference to a practitioner-validated solution for tangible public safety impact. The major contributions of this research are summarized as: 1.1 Key Contributions We make the following key contributions. 1. Large-Scale PTW Simulator Data Under TP: To study the effect of TP on PTW riders, we collected 129,000 labeled samples from 51 participants under No/Low/High TP. Each input sequence comprises 64 features categorized into vehicle dynamics, control inputs, headway/tailway distances, and behavioral violations. 2. Novel Architecture for Collision Risk Forecasting and Classification: We propose MotoSafety, a novel architecture grounded in the Learned Temporal Importance (LTI) principle. The architecture integrates: (i) parallel multi-scale dilated convolutions for local-to-medium pattern extraction; (i) Bidirectional Long Short-Term Memory (Bi-LSTM) for long-range dependencies; (i) Temporal Importance Pooling (TIP) for content-aware temporal collapse, applied independently to both CNN and BiLSTM branches, which reduces downstream fusion, SE-recalibration, and attention stages to (1)O(1) relative to the sequence length, while the front-end encoder operates at (L)O(L) complexity, resulting in an overall (L)O(L) framework (Table 4); and (iv) a self-gated fusion mechanism that leverages squeeze-and-excitation (SE) blocks and multi-head attention (MHA) for adaptive representation refinement. 3. SOTA Performance and Deployability: MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, showing improved performance over ten baselines including TimesNet, Time-LLM and LLM4TS. For long-term forecasting, it achieves 0.039 MSE and 0.094 MAE (4.4× lower error than Time-LLM and iTransformer). Notably, MotoSafety requires only ∼ 1.15M parameters and achieves an inference latency of 0.135 ms, making it 21.9× and 71.9× faster than TimesNet and LLM4TS, respectively, and makes it suitable for edge devices. 4. TP Inductive Bias, Real-World Feasibility, and Architecture Transferability: Explicit TP prediction improves collision accuracy from 94.09% to 94.97% (ground truth) and 94.82% (predicted TP). Using only 21 real-world features (IMU+GPS), MotoSafety achieves 93.91% accuracy, indicating practical deployment potential; hardware validation remains future work. Beyond PTW safety, the architecture demonstrates strong transferability to human activity (97.66%) and clinical exercise (99.65%) domains when retrained from scratch. 2 Related Work 2.1 Time Pressure as a Latent Risk Factor in Driving Safety Time Pressure (TP) is widely recognized as a cognitive stressor that degrades driver and rider decision-making and elevates crash risk. Under TP, individuals experience a perceived urgency to complete the driving task within insufficient time, which narrows attentional focus, accelerates cognitive processing, and increases reliance on risk-oriented heuristics [37, 11]. Empirical studies consistently show that TP leads to higher speeds, shorter accepted gaps, reduced safety margins, and increased crash likelihood across diverse traffic contexts [35, 43, 44]. Experimental evidence from simulator-based studies indicates that TP significantly alters longitudinal and lateral control behavior, including increased acceleration variability, unstable steering, and abrupt braking [36]. These behavioral deviations are strongly associated with near-crash and collision events, suggesting that TP acts as an upstream cognitive trigger for unsafe driving outcomes rather than merely a contextual modifier. PTWs are particularly susceptible to TP-induced risk due to unstable dynamics and high dependence on rider motor control compared to four-wheelers [35, 43, 9]. At unsignalized intersections, drivers experiencing low time pressure (LTP) and high time pressure (HTP) tend to accept substantially shorter gaps, leading to crash probability increases of 127% under LTP and 181% under HTP compared to no time pressure (NTP) conditions [37]. TP increases cognitive load and unintentional attentional lapses under stress [10, 20], reducing situational awareness and leading to errors that precede collisions. These findings highlight TP as a critical contributor to risk in dense, heterogeneous traffic flows common in LMICs. 2.2 Simulator-Based Collision and Risk Analysis Driving simulators offer a secure and regulated setting for examining collision risk under cognitively demanding scenarios, including TP. They eliminate the hazards associated with real-world driving while facilitating high-resolution data acquisition for both collision risk classification and kinematic state prediction [36, 3, 21]. Prior work demonstrates strong relative validity between simulated and real-world behavior, particularly for risk-related metrics including speed selection, braking intensity, and control variability [36, 3, 21]. Although the magnitude of observed behaviors may differ between simulator and naturalistic conditions, the core behavioral responses to TP are preserved: both environments exhibit higher speeds, more frequent lane changes, stronger braking, and greater variability in vehicle control [36]. This supports the ecological validity of using simulators to investigate TP-induced behavioral changes in PTW rider behavior, enabling the current study to train and evaluate MotoSafety’s classification and forecasting tasks under controlled yet realistic conditions. 2.3 Machine Learning for Collision and Risk Prediction The application of ML and DL methods in intelligent transportation systems (ITS) has grown considerably, particularly for predicting risky events, collisions, and hazardous driving behaviors. Traditional techniques such as Decision Tree, Support Vector Machines, and feed-forward neural networks have shown moderate effectiveness in detecting high-risk maneuvers [6, 42]. In recent years, sequence-based architectures like Temporal Transformers, Informer, and TimesNet have demonstrated considerable effectiveness in capturing long-range temporal patterns in driving behavior data [53, 51]. However, most existing models focus on detecting externally observable outcomes (e.g., collisions or near-crashes) after risk has already manifested. They largely overlook latent cognitive precursors such as TP that shape behavior well before unsafe actions occur. Furthermore, few studies address PTW-specific dynamics, which are characterized by inherent instability and high dependence on rider motor control. 2.4 Research Gap and Motivation Although TP is a critical contributor to unsafe driving behavior, it is rarely integrated into collision risk assessment frameworks as a latent risk factor. This gap is particularly pronounced for PTWs, where safety margins are narrow and early intervention is essential. Motivated by these limitations, this work advances collision risk assessment under TP by leveraging high-resolution multivariate time-series data to infer risk directly from behavioral dynamics. By focusing on collision risk in cognitively stressed riding conditions, this study aligns with the Safe System Approach (SSA) [47] and the goal of Vision Zero [18], enabling the detection of high-risk states and supporting the development of intelligent rider-assistance systems and enhanced safety for PTW users. 3 Methodology In this section, we provide details of our data collection protocol, followed by MotoSafety architecture. 3.1 Experimental Design and Data Collection 3.1.1 Simulator Setup We carried out all experimental trials on a fixed-base PTW simulator at our institute (Fig. 3). The simulator features a Honda motorcycle frame fitted with working throttle, brake, clutch, and gear controls, designed to replicate the riding experience of a real motorcycle. Riders are surrounded by a three-screen setup that delivers a wide viewing angle and lifelike road imagery. High-precision industrial-grade sensors continuously record rider inputs, vehicle dynamics, and environmental parameters through a real-time data acquisition interface, complying with IEC 393 accuracy standards [13]. While the frame is static, a 4-actuator motion system provides vibration and acceleration. Complete technical specifications, including sensor accuracy ratings, appear in Table 1. This configuration offers a realistic yet controlled setting, allowing us to systematically study how riders behave across various traffic and road conditions. Table 1: Two-wheeler simulator: key technical parameters Component Details Platform Fixed-base PTW frame with ISO certification Controls Acceleration, clutch, hydraulic brake unit, 5-speed shifting mechanism, handlebar steering, parking brake Sensors Wire-wound servo potentiometers (type 50W): Resistance: 5 kΩ (±10%); Linearity error: ±0.5% (IEC 393 compliant), Angular range: 355° ± 3°; Full mechanical rotation: 360° Maximum power: 3 W (at 70°C); Expected operational life: 2 million cycles Safe operating limits: –40°C to +105°C Force transducers: Brake force sensing Rotary encoders: Gear position detection; SPDT microswitches: Gear shift sensing Visual System Three 50-inch LED panels with 180° horizontal coverage Motion System Four 3-DOF electric actuators capable of ±1 G max acceleration, 0–100 Hz response Motion ranges: Pitch ±1.75°, Roll ±3.5°, Heave 38.1 m; Actuator stroke: 1.5 inches Per-actuator load rating: 114 kg Computer Machine Rider station: Intel i7, 32 GB RAM, GTX 1650 graphics card Operator station: Intel i5 processor, 16GB RAM, GTX 1650 GPU Software TechnoSim (AI-powered traffic, environment, and scenario simulation) Data Logging Continuous data recording with sampling frequency fs≥100f_s≥ 100 Hz This ISO-compliant simulator provides a controlled, scalable, and safe environment to investigate rider behavior, cognitive load, and collision risk under realistic operational conditions. 3.1.2 Simulated Scenarios Figure 3: Participants performing the simulator test. We design a 4.8 km urban scenario on a high-fidelity two-wheeler simulator to investigate rider behavior under cognitive stress (TP). The route comprises a combination of two-lane undivided and four-lane road segments, with a maximum speed of 50 km/h, representing common road conditions found in Indian cities. The scenario embeds diverse events—including pedestrian crossings, obstacle overtaking, intersections, bike and car following segments to elicit real-time decision-making, adaptive control, and risk-taking under varying TP. The key elements of the scenario are illustrated in Fig. 4: Fig. 4(a): primary route with intersections and conflict zones, Fig. 4(b): AI vehicle triggers simulating dynamic interactions, Fig. 4(c): route direction changes and detour prompts, and Fig. 4(d): speed change triggers near sensitive areas. (a) Primary route with intersections and conflict zones (b) AI vehicle triggers simulating dynamic interactions (c) Route direction changes and detour prompts (d) Speed change triggers near sensitive areas Figure 4: Sample snapshots of the simulated riding environment. 3.1.3 Participants Table 2: Demographic and riding profile of participants (N = 51) Attribute Category / Range Mean (SD) or % Age (years) 18–42 26.4 (5.7) Riding experience (years) ≥2≥ 2 5.6 (3.1) License status Valid PTW license 100% Education level Graduate / B.Tech 58.8% / 41.2% Simulator exposure No / Yes 92.2% / 7.8% Health status Medically fit 100% MoRTH fatality data from 2019 to 2023 consistently shows that male riders constitute the overwhelming share of two-wheeler deaths in India, with yearly proportions of 86.0%, 87.3%, 86.4%, 86.2%, and 85.2%, respectively [23, 24, 25, 26, 28]. The 18–45 age group accounts for the largest share of these deaths. This strong demographic pattern supports our choice to include only male riders in this study. We therefore recruited 51 male riders aged 18 to 42 years, all with a minimum of two years of riding experience. Table 2 summarizes their characteristics, while Fig. 3 shows participants during simulator trials. The male-only cohort is justified by three converging factors: First, national crash data shows male riders as the dominant high-risk group. According to MoRTH reports, male PTW fatalities accounted for 85.2–87.3% of all PTW deaths from 2019 to 2023 [23, 24, 25, 26, 28], with fatalities concentrated in the 18–45 age group. Statistical tests confirm this pattern: a binomial test shows male fatalities (86%) significantly exceed equal representation (p<0.0001p<0.0001), and a chi-square test confirms consistency across years (χ2=262.3χ^2=262.3, df = 5, p<0.0001p<0.0001). Second, female exposure to manually geared PTWs is extremely limited. Women hold only about 6% of motor vehicle licenses in India [12, 45, 2], significantly below equal representation (p<0.0001p<0.0001). The number of women riding manually geared motorcycles is even smaller, largely due to socio-cultural factors and traditional norms and practical barriers [45]. Consequently, the pool of female riders with manual transmission experience suitable for simulator-based research is extremely small. Third, the simulator’s manual transmission requirement creates a practical recruitment constraint. The high-fidelity PTW simulator replicates a manually geared motorcycle with functional clutch and gear controls, requiring prior manual transmission experience. Given the extremely low prevalence of female riders with such experience in India, recruiting a statistically sufficient female sample for robust analysis is infeasible. Thus, male participants were selected to ensure a sufficiently large and statistically robust sample for meaningful spatial risk analysis, consistent with the dominant high-risk demographic. Future studies will include female riders as their PTW participation increases. 3.1.4 Time Pressure (TP) Scenarios To ethically evoke graduated cognitive load mimicking emergency commuting, participants were tasked with reaching an examination hall under three conditions: (i) NTP: Baseline with ample time; (i) LTP: Limited to 90% of the baseline time (with a reminder to avoid being late); and (i) HTP: Restricted to 80% of the baseline time, accompanied by active urgency cues (e.g., exam hall closing shortly). This standardized exam-deadline scenario simulates high-arousal stress states common in urban transit, providing a controlled environment to study behavioral degradation. 3.1.5 Experimental Design and Protocol Figure 5: Flow of dataset collection and experimental design. A structured four-phase experimental protocol (Fig. 5) was implemented to ensure data consistency: (i) Briefing: Standardized orientation and informed consent; (i) Practice: A 5–10 minute familiarization ride (excluded from analysis); (i) Main Task: Participants completed three riding sessions under NTP, LTP, and HTP conditions, with the session sequence randomized across participants to avoid order effects; and (iv) Rest: 5–8 minute inter-session breaks to stabilize cognitive load and mitigate fatigue. Ethics and Informed Consent: All participants signed a written consent form before taking part in the research. The research protocol received ethical clearance from the IE Committee at the IIT Indore (Approval No. BSBE/IITI/IHEC-11/2025/11). Consent to Publish: All participants gave their written consent for the publication of anonymized data and findings derived from this study. To safeguard participant privacy, all data have been fully anonymized. 3.1.6 Behavioral Transition : Why LTP Matters Adding the LTP condition allows us to capture the transitional phase from safe riding (NTP) to risky riding (HTP), representing the early stages of behavioral decline. The progression NTP → LTP → HTP illustrates the transition from stable, safe riding to high-stress operation characterized by speeding and infractions. LTP serves as a critical indicator where subtle warning signs—such as minor speeding or slight lane drifts—first manifest. By incorporating LTP, the model improves its sensitivity to rising stress levels, supporting earlier detection and enabling more timely safety interventions. Moreover, LTP data inform policymakers and trainers by defining practical stress thresholds, facilitating preventive strategies before riders reach the high-risk HTP state. 3.2 Feature Engineering and Preprocessing From raw time-series sensor data, we derived 64 domain-informed features that capture vehicle dynamics, rider behavior, and contextual factors as summarized in Table 3. The features include things like average and variation in speed and acceleration, how often the rider brakes, how frequently they change lanes, and the TP level. Min-Max normalization was used to bring all feature values within the [0,1][0,1] range. We segment the time series data into fixed-length overlapping windows to preserve temporal dependencies, enabling the model to capture both gradual risk accumulation and abrupt behavioral deviations. Detailed training configurations are provided in Section 4.1. 3.3 Dataset and Feature Set The experimental protocol yielded 153 multivariate time-series sessions (51 participants × 3 conditions). Data were sampled at 100 Hz. The 100 Hz streams were segmented into 960 ms windows (96 time steps) with 50% overlap, yielding 129,209 windows slid continuously. A window is labeled as collision (1) if a collision occurs within it, otherwise 0. The dataset exhibits a class distribution of 21% collision and 79% non-collision events. A single collision spans multiple windows (impact+sliding+aftermath, due to overlap), reflecting the inherent imbalance of safety-critical scenarios. This is window-level, not event-level labeling. Each input sequence comprises 64 features (summarized in Table 3) categorized into: (i) Vehicle Dynamics: Speed, acceleration, and 3D rotation; (i) Control Inputs: Throttle, hydraulic brake force, and gear transitions; (i) Proximity: Headway/tailway distances and lane offsets; (iv) Time context and scenario: Time Stamp, TP (0=HTP, 1=LTP, 2=NTP); and (v) Behavioral Violations: Overspeeding, improper gap maintenance and traffic rule infractions. All collisions were aggregated into a single binary target variable, with leakage prevented by excluding impact-related features from the input. Table 3: Overview of simulator-derived features and their categories Feature Category Included Parameters Count Vehicle Controls Ignition status, Engine state, Throttle position, Brake application, Clutch engagement, Handbrake, Steering angle, Gear position, Headlight usage, Horn activation 10 Vehicle Performance Speed, Engine RPM, Fuel consumption, Distance covered 4 Lighting & Indicators Indicator signal, Pre-move indication, Junction turning signal, Lane-change signal, Headlight non-compliance 5 Behavioral Violations Speeding, Incorrect junction/intersection speed, Speed violations on speed breakers, Gap maintenance errors, Hazardous overtaking, Unsigned turns, Lane misuse, Driving on the wrong side, Handbrake use while moving, Riding the clutch, Faulty gear shift order, Improper clutch release, Shifting without clutch engagement, Proper gear selection before start, Gradual clutch engagement 15 Traffic Infractions White line violations, Yellow line crossings, Stop line breaches, Signal jumping, Entering no-entry zones, Illegal U-turns, Parking infringements 7 Temporal Context Time stamp, Time Pressure condition (0=HTP, 1=LTP, 2=NTP) 2 Spatial Attributes 3D position coordinates, 3D rotation angles, Lane number, Left/Right lane offsets 9 Motion and Proximity Lateral/Longitudinal velocity, Headway (distance & time), Tailway (distance & time), Left/Rightway distance, Steering angle 9 Brake System Brake test completion, Front and Rear braking force 3 Note: All features associated with collisions (vehicles, objects, obstacles, etc.) were merged into a single binary target variable ‘Target‘, where 0 = no collision and 1 = collision. These collision events were excluded from the input feature set to prevent label leakage. 3.4 Scope of the Classification and Forecasting Tasks We define two complementary but distinct tasks, and are explicit about what each does and does not establish. Classification Task (Same-Window Risk Detection): Given a 960 ms observation window, the model predicts whether a collision event occurs within that window. Because a subset of positive windows overlaps the collision event itself (impact, sliding, or aftermath phases, as noted in Section 3.3), this task is best characterized as same-window collision risk detection rather than a fixed-lead-time forecast: it evaluates whether pre-collision and in-progress behavioral signatures are jointly separable from normal riding, not how far in advance a collision can be anticipated. Forecasting Task (Prospective Kinematic Prediction): Given a lookback window of L=96L=96 timesteps, a separately trained instance of the architecture predicts future kinematic states (e.g., speed, lean angle, longitudinal force) up to H∈96,192,336,720H∈\96,192,336,720\ timesteps (up to 7.2s) ahead. This task, evaluated in Section 5.2, is a genuine prospective-forecasting result with an explicit lead time, but its target is future sensor state, not future collision probability; it demonstrates that the architecture can anticipate the kinematic precursors of hazardous states, which is complementary to, but distinct from, a direct lead-time collision-probability forecast. We view a labeled, fixed-lead-time collision-probability forecast (i.e., excluding all windows that overlap a collision event and labeling remaining windows by whether a collision begins within a defined horizon τ afterward) as an important direction for follow-up work, and report the current classification results with this scope explicitly stated rather than implied. 3.5 MotoSafety: Proposed Architecture 3.5.1 Architecture Overview Figure 6: The proposed MotoSafety architecture. The proposed MotoSafety architecture, shown in Fig. 6, is a DL architecture based on the Learned Temporal Importance (LTI) principle, developed for the prediction of PTW collisions and forecasting. It employs a multi-stage pipeline: (i) a parallel multi-scale dilated convolutional encoder with three branches (dilation factors 1, 2, 4) extracts local to medium-range temporal patterns; (i) a 2-layer Bi-LSTM (128 hidden units per direction) captures long-range sequential dependencies; (i) a novel TIP module (which uses a lightweight Multi-Layer Perceptron (MLP) as a learned scoring function to evaluate the importance of each timestep) is applied independently to the outputs of both the CNN and LSTM paths, learning to adaptively weight the most critical pre-collision timesteps for each feature type and collapsing them into compact vectors; (iv) these vectors are concatenated and processed by a self-gated fusion stage, refined by channel-wise recalibration via a SE block and MHA (4 heads). This representation undergoes normalization, dropout-based regularization, and linear projection to output binary collision risk scores. 3.5.2 Comment on the Novelty We distinguish clearly between architectural components adapted from prior work and the components that constitute this paper’s novel contribution. Adapted components: The parallel dilated-convolution encoder, Bi-LSTM, SE recalibration, and MHA are established building blocks, each individually well-studied in time-series and sequence modeling literature. Novel contribution: TIP. The core contribution of this work is the TIP module and the LTI principle it instantiates. Rather than collapsing the temporal dimension with a fixed operator (mean-pool, max-pool, or final-hidden-state extraction, as is standard practice in CNN/RNN-based time-series classifiers), TIP learns a content-aware scoring function that assigns a data-dependent importance weight to every timestep before collapsing the sequence. This is applied independently to both the CNN and BiLSTM representations, allowing the model to learn which pre-collision moments matter most for each feature pathway, rather than treating all timesteps as equally informative. Unlike full self-attention pooling (which computes pairwise timestep interactions at (L2)O(L^2) cost), TIP scores each timestep independently against a shared learned criterion, at (L)O(L) cost, and produces a fixed-size vector consumed by the downstream fusion stage. Applied independently to both CNN and BiLSTM branches, TIP reduces the downstream fusion, SE-recalibration, and attention stages to (1)O(1) relative to the sequence length, while the front-end encoder operates at (L)O(L) complexity, resulting in an overall (L)O(L) framework (Table 4). Empirical validation of TIP. The effectiveness of TIP is validated in Section 5.8 (Table 14), where it outperforms standard fixed-pooling strategies (mean-pooling, max-pooling, last-timestep) under an identical architecture, indicating that the performance gain stems from learned, content-aware weighting rather than fixed pooling. This strategic “content-aware collapse” means the fusion, SE-recalibration, and MHA stages operate on a fixed-size 320-dimensional vector regardless of sequence length, i.e., these downstream stages cost (1)O(1) per sample once pooling is complete. The upstream CNN and Bi-LSTM encoder remains (L)O(L), so the model’s end-to-end complexity is (L)O(L) overall, not (1)O(1); the practical benefit of TIP is that it removes the (L2)O(L^2) cost a full-sequence self-attention mechanism would otherwise add on top of the encoder. Table 4 details the per-module accounting. This design ensures the low-latency performance required for robust real-time deployment on edge devices. Table 4: Per-module computational complexity. L: sequence length, C: feature channels, H: hidden dimension, k: kernel size. Module Complexity Dilated CNN encoder (×3 branches) (L⋅k⋅C)O(L· k· C) Bi-LSTM (2 layers) (L⋅H2)O(L· H^2) TIP scoring + weighted sum (L⋅H)O(L· H) Gated fusion + SE block (1)O(1) (fixed-size input) Multi-head attention (post-pooling) (1)O(1) (fixed-size input) MotoSafety total (L)O(L) 3.5.3 MotoSafety: Mathematical Formulation We address the problems of binary collision prediction and multi-horizon time-series forecasting from multivariate sensor data under TP. Our input consists of sensor readings X∈ℝB×L×CX ^B× L× C from a two-wheeler simulator, with B batch size, lookback window L and C feature channels. Forward Process: The input X is processed by two parallel branches: CNN and BiLSTM. The CNN branch captures local patterns using three dilated convolutional branches with kernel size k=3k=3. To accommodate the channel-first requirement of 1D convolutions, we apply a permutation function π(⋅)π(·) to X such that X^=π(X)∈ℝB×C×L X=π(X) ^B× C× L: Hd H_d =ReLU(Conv1Dd(X^)) =ReLU(Conv1D_d( X)) (1) where d=1,2,4d=\1,2,4\ is the dilation factor. These views are merged into a unified 64-channel representation via a 1×11× 1 convolution: Hms=Conv1D1×1([H1;H2;H4])∈ℝB×64×LH_ms=Conv1D_1× 1([H_1;H_2;H_4]) ^B× 64× L. In parallel, a bidirectional LSTM captures long-range dynamics: Hlstm=BiLSTM(X)∈ℝB×L×256.H_lstm=BiLSTM(X) ^B× L× 256. (2) To identify safety-critical moments, we utilize TIP to learn importance weights across the temporal dimension. Two distinct networks, MLPθ1MLP_ _1 and MLPθ2MLP_ _2, assign scores to the CNN and LSTM features respectively: wc w_c =Softmax(MLPθ1(HmsT)),fcnn=∑t=1Lwc(t)Hms(t) =Softmax(MLP_ _1(H_ms^T)), 9.24994ptf_cnn= _t=1^Lw_c^(t)H_ms^(t) (3) wl w_l =Softmax(MLPθ2(Hlstm)),flstm=∑t=1Lwl(t)Hlstm(t) =Softmax(MLP_ _2(H_lstm)), 9.24994ptf_lstm= _t=1^Lw_l^(t)H_lstm^(t) (4) The resulting vectors, fcnn∈ℝ64f_cnn ^64 and flstm∈ℝ256f_lstm ^256, are concatenated to form the joint representation f=[fcnn,flstm]∈ℝ320f=[f_cnn,f_lstm] ^320. The fused representation undergoes refinement through a gating mechanism and a channel-wise recalibration block (SE-Block). Finally, a MHA block performs high-order feature refinement: fgate f_gate =f⊙σ(Wgf+bg), =f σ(W_gf+b_g), (5) fse f_se =SEBlock(fgate), =SEBlock(f_gate), (6) fattn f_attn =MHA(fse). =MHA(f_se). (7) The final output ∈ℝ2y ^2 is produced by applying a linear layer to the normalized and regularized context representation: =Softmax(Linear(LayerNorm(Dropout(fattn)))CLOSE.y=Softmax (Linear (LayerNorm(Dropout(f_attn) ) ). (8) To address the extreme class imbalance (rare collision events), we utilize Focal Loss with a focusing parameter γ=2γ=2: ℒcls=−1B∑i=1B(1−Pi,yi)γlog(Pi,yi),L_cls=- 1B _i=1^B(1-P_i,y_i)^γ (P_i,y_i), (9) where Pi,yiP_i,y_i denotes the model’s estimated probability for the actual binary collision outcome yi∈0,1y_i∈\0,1\. Multi-Horizon Forecasting Module Our forecasting module predicts future target values Y^∈ℝB×τ Y ^B×τ over horizon τ. It uses a separate but architecturally similar network with identical CNN and BiLSTM branches, TIP, gating, SE-Block, and MHA. Both the classification and forecasting modules use the same L=96L=96 timestep lookback window; they differ only in their prediction target (same-window collision label vs. future kinematic states) and are trained separately, as described above. This produces a refined context vector attn∈ℝB×320f_attn ^B× 320. The final forecast is obtained via: Y^=Linearforecast(LayerNorm(attn)). Y=Linear_forecast(LayerNorm(f_attn)). (10) The module is trained to minimize the Mean Squared Error (MSE) loss: ℒforecast=1B⋅τ∑i=1B∑k=1τ(yi,L+k−Y^i,k)2,L_forecast= 1B·τ _i=1^B _k=1^τ (y_i,L+k- Y_i,k )^2, (11) where yi,L+ky_i,L+k is the true target value at future step k. Both models share identical architectural components but are trained separately for their respective tasks. 4 Experimental Setup and Baselines 4.0.1 Ground Truth and Human Validation To ensure the reliability of the simulator-generated labels, a systematic validation protocol was conducted. Two road experts in road safety and human factors (male, post-graduate) reviewed a representative subset of 2,000 randomly selected segments from the total of 129,000-sequence dataset. Experts independently annotated the onset and severity of high-risk episodes (e.g. abrupt swerves, loss of balance, critical near-misses) using a standardized coding scheme based on kinematic thresholds and behavioral cues. We used Cohen’s Kappa (κ) to check how well the two annotators agreed with each other, which yielded a value of 0.87. This indicates near-perfect agreement based on established interpretation standards [19]. This high agreement validates the dataset’s labels as a robust ground truth for model training. 4.0.2 Baselines We benchmark MotoSafety against ML (RF, CNN, RNN) and modern DL architectures: Informer [53] and iTransformer [22] for Transformer-based long-range modeling. TimesNet [48] for multi-scale frequency analysis; PatchTST [32] for patching-based local semantics. Time-LLM [16] and LLM4TS (GPT-2) [7] as a representative fine-tuned LLM baseline. 4.1 Training and Evaluation 4.1.1 Training Protocol The MotoSafety model is implemented in PyTorch and optimized using AdamW (initial learning rate: 3×10−43× 10^-4, L2 weight decay: 5×10−55× 10^-5) with Focal Loss (γ=2γ=2) for classification and MSE loss for forecasting. To enhance generalization, we employed Mixup augmentation, Exponential Moving Average (EMA) of weights, and Monte Carlo (MC) Dropout (5 samples) for uncertainty calibration. For model training, we utilized an NVIDIA T400 GPU (4 GB VRAM) and set the batch size to 64, running the process for 50 epochs. The main training settings are listed in Table 5. For all experiments, we used a participant-wise data split of 80% for training, 10% for validation, and 10% for testing, where all sessions from the same rider are kept together in the same split to prevent data leakage. Table 5: Training configuration for MotoSafety classification and forecasting models. Component Setting Classification Model Input sequence length 96 timesteps (960 ms) Feature dimension 64 CNN hidden channels 32 per branch, 64 fused BiLSTM hidden dimensions 128 (bidirectional) Total model parameters ∼ 1.15 million Dropout rate 0.3 Loss function Focal Loss (γ=2γ=2) Forecasting Model Input lookback length 96 timesteps Prediction horizons H∈96,192,336,720H∈\96,192,336,720\ Architecture Identical to classification model Loss function MSE Shared Settings Optimizer AdamW (initial learning rate: 3×10−43× 10^-4, L2 weight decay: 5×10−55× 10^-5) Batch size 64 Training epochs 50 Regularization EMA, Mixup, MC Dropout (5 samples) Hardware NVIDIA T400 GPU (4GB VRAM) 4.1.2 Evaluation Protocol We evaluated classification performance using Accuracy, F1-Score, and ROC-AUC, while for the forecasting task, we used MSE and Mean Absolute Error (MAE) as evaluation metrics. Statistical significance was assessed via paired Wilcoxon signed-rank tests over 5 runs, with the Holm–Bonferroni correction (α=0.05α=0.05). 5 Results 5.1 Impact of Time Pressure on Riding Behavior Table 6: Behavioral differences across TP: NTP/LTP/HTP (Mean ± Std). Feature NTP LTP HTP Speed (km/h) 33.12 ± 19.42 39.69 ± 21.67 49.00 ± 26.48 Sudden braking 1.32 ± 2.54 1.38 ± 2.36 1.80 ± 2.01 Dangerous overtaking 1.07 ± 1.18 1.21 ± 1.27 1.36 ± 1.52 Headway distance (m) 8.99 ± 14.38 8.20 ± 14.34 7.75 ± 13.88 Analysis of our PTW dataset reveals that HTP is the primary driver of behavioral volatility, with a clear transition to high-risk strategies as TP increases, as shown in Table 6. Riders under HTP exhibited 48% higher mean speed, 36% increase in sudden braking, and 14% shorter headway distance compared to NTP, along with a 27% increase in dangerous overtaking. These statistically supported trends define a hurry-up strategy of excessive speed, abrupt control, and reduced spacing, highlighting TP as a key behavioral stressor that increases the risk of PTW collisions. Table 7: Multivariate long-term forecasting results (mean ± std over 5 runs). Input lookback length L=96L=96. Forecast horizons H∈96,192,336,720H∈\96,192,336,720\. Avg is averaged over all four prediction horizons. Models MotoSafety Time-LLM iTransformer PatchTST Metric MSE MAE MSE MAE MSE MAE MSE MAE 96 0.033 ± 0.005 0.082 ± 0.007 0.136 0.148 0.161 0.237 0.188 0.208 192 0.037 ± 0.006 0.086 ± 0.008 0.145 0.153 0.170 0.246 0.205 0.223 336 0.041 ± 0.006 0.097 ± 0.010 0.192 0.215 0.177 0.258 0.227 0.272 720 0.045 ± 0.007 0.111 ± 0.012 0.212 0.293 0.183 0.262 0.267 0.338 Avg 0.039 0.094 0.171 0.210 0.173 0.251 0.222 0.260 5.2 Multivariate Long-Term Forecasting Performance To evaluate the proactive safety capabilities of MotoSafety, we conducted multivariate forecasting across four different time horizons (H∈96,192,336,720H∈\96,192,336,720\) on our PTW data. This experiment assesses the models ability to project future kinematic states such as lean angles and longitudinal forces, essential for anticipating hazardous transitions before they culminate in a collision. As shown in Table 7, MotoSafety attains a mean MSE of 0.039 and an MAE of 0.094, representing a 4.4× reduction in error over the Time-LLM (0.171) and iTransformer (0.173) baselines. This high forecasting fidelity under TP is critical for proactive safety; by accurately projecting states up to H=720H=720, the model provides a vital safety buffer for early risk mitigation before a maneuver reaches a non-recoverable limit. 5.3 Comparative Classification Performance The model performance results on the PTW simulation dataset under TP are presented in Table 8. The proposed MotoSafety achieves 94.97% accuracy, 93.7% F1-score and 99.33% ROC AUC, outperforming ten baselines including TimesNet, PatchTST, iTransformer, Time-LLM and LLM4TS. To assess the robustness of the performance gains, statistical significance is evaluated using paired Wilcoxon signed-rank tests over 5 independent runs (Table 8). For the nine pairwise comparisons against baselines, we applied the Holm–Bonferroni method (α=0.05α=0.05) to regulate the family-wise error rate (FWER). The high ROC AUC indicates that MotoSafety can maintain a low false-alarm rate, which is a prerequisite for real-world Advanced Rider Assistance Systems (ARAS). Table 8: Comparison of model performance (mean ± standard deviation over 5 runs). We used the paired Wilcoxon signed-rank test with Holm–Bonferroni adjustment to determine statistical significance (α=0.05α=0.05). Significance is denoted by: ∗p<0.001^***p<0.001, p∗∗<0.01^**p<0.01, ∗p<0.05^*p<0.05. All baseline models were compared against the proposed MotoSafety framework. Model Acc. (%) Prec. Rec. F1-Score (%) AUC (%) Mean ± Std Result Result Mean ± Std Mean ± Std RF 84.93±0.16∗84.93± 0.16 85.24 84.87 85.02±0.23∗85.02± 0.23 93.68±0.31∗93.68± 0.31 CNN 90.19±0.11∗90.19± 0.11 93.82 80.68 86.74±0.21∗86.74± 0.21 97.62±0.22∗97.62± 0.22 RNN 90.94±0.12∗90.94± 0.12 83.41 96.38 89.43±0.32∗89.43± 0.32 97.27±0.34∗97.27± 0.34 LLM4TS 90.06±0.14∗90.06± 0.14 90.58 83.23 86.81±0.31∗86.81± 0.31 97.31±0.42∗97.31± 0.42 Informer 93.61±0.07∗93.61± 0.07 89.72 94.88 92.21±0.22∗∗92.21± 0.22 99.12±0.21∗∗99.12± 0.21 TST 93.86±0.09∗93.86± 0.09 93.67 90.74 92.18±0.29∗∗92.18± 0.29 99.23±0.28∗99.23± 0.28 TimesNet 94.06±0.05∗94.06± 0.05 89.23 96.69 92.77±0.38∗∗92.77± 0.38 99.18±0.24∗99.18± 0.24 iTransformer 94.39±0.06∗∗94.39± 0.06 92.41 93.57 93.12±0.42∗93.12± 0.42 99.22±0.32∗99.22± 0.32 PatchTST 94.41±0.05∗∗94.41± 0.05 90.73 96.18 93.28±0.24∗93.28± 0.24 99.31±0.2299.31± 0.22 Time-LLM 94.46±0.07∗94.46± 0.07 93.12 93.79 93.41±0.33∗93.41± 0.33 99.24±0.31∗99.24± 0.31 MotoSafety 94.97±0.0894.97± 0.08 93.8093.80 93.9093.90 93.70±0.2093.70± 0.20 99.33±0.3099.33± 0.30 5.4 Forecasting Horizon Analysis To characterize how forecasting accuracy varies with prediction horizon, Table 9 reports MSE and MAE for the kinematic-state forecasting task (Section 3.5.3) across horizons up to 7.2s. Error increases gradually with horizon (MSE from 0.033 at H=96H=96 to 0.045 at H=720H=720), indicating the model retains meaningful predictive signal even at longer horizons rather than degrading sharply. We note this result characterizes future kinematic-state prediction and is reported separately from the same-window collision classification task described in Section 3.3; the two tasks use different targets and should not be read as a single fused collision-lead-time result. Table 9: Forecasting performance by prediction horizon (mean over 5 runs) Horizon H Time (s) MSE MAE 96 0.96 0.033 ± 0.005 0.082 ± 0.007 192 1.92 0.037 ± 0.006 0.086 ± 0.008 336 3.36 0.041 ± 0.006 0.097 ± 0.010 720 7.20 0.045 ± 0.007 0.111 ± 0.012 5.5 Downstream Task: Impact of TP on Model Performance Table 10: Performance comparison of MotoSafety showing improvement from adding TP features. Significance markers indicate improvement over the Without TP baseline: ∗p<0.001^***p<0.001 Model Accuracy (%) AUC (%) Without TP Feature 94.09±0.0494.09± 0.04 99.10±0.0199.10± 0.01 With Predicted TP Feature 94.82±0.02∗94.82± 0.02^*** 99.24±0.01∗99.24± 0.01^*** With Ground Truth TP Feature 94.97±0.01∗94.97± 0.01^*** 99.33±0.01∗99.33± 0.01^*** In practical deployment, direct and accurate measurement of ground truth TP labels, denoted TPgtTP_gt, may be infeasible due to sensor limitations or real-time constraints. In the proposed data set, the TP characteristic is a categorical variable representing three distinct TP conditions: HTP = 0, LTP = 1, NTP = 2. To address missing or unavailable GT TP labels, we employed a Time Series Transformer (TST) [52] model to predict these categorical TP labels from the multivariate time series input data. The TST achieved 89.26% accuracy for 3-class TP prediction (0=HTP, 1=LTP, 2=NTP). Table 11 shows the confusion matrix (raw counts) per-class performance. Table 11: Confusion matrix (raw counts) for TST-based TP prediction. True / Predicted 0: HTP 1: LTP 2: NTP 0: HTP 8878 229 163 1: LTP 76 7326 568 2: NTP 82 1661 6859 Let =(1,2,…,T)∈ℝT×DX=(x_1,x_2,…,x_T) ^T× D represent the multivariate time series input from the PTW simulator, which includes all features except the explicit TP indicator. In this notation, T refers to the sequence length, while D=63D=63 specifies the number of telemetry features. Each t∈ℝDx_t ^D captures the sensor states at the t-th time instant. The TST model is then expressed as: TPpred=fTST(),TP_pred=f_TST(X), (12) where TPpred∈0,1,2TP_pred∈\0,1,2\ is the predicted TP label. The predicted TP feature TPpredTP_pred, along with all other relevant sensor and behavioral features of the data set, forms the complete input to our collision risk prediction model, MotoSafety. Formally, the model can be expressed as: y^=gMotoSafety(,TP), y=g_MotoSafety(X,TP), (13) where TPTP is the ground truth TPgtTP_gt or predicted TPpredTP_pred, and y y is the predicted safety risk label for the rider. We evaluated the robustness of model by comparing the predictions when using: y^gt y_gt =gMotoSafety(,TPgt), =g_MotoSafety(X,TP_gt), (14) y^pred y_pred =gMotoSafety(,TPpred). =g_MotoSafety(X,TP_pred). (15) The results presented in Table 10 highlight the critical role of TP in predicting collision risk. While the baseline model fMotosafety(X)f_Motosafety(X) achieves an accuracy 94.09% using only PTW simulation data without TP feature. The integration of predicted TP features fMotosafety(X⊕TPpred)f_Motosafety(X TP_pred) improved this to 94.82%, nearly matching the 94.97% Oracle performance fMotosafety(X⊕TPgt)f_Motosafety(X TP_gt). This performance gain shows that TPpredTP_pred captures latent psychological context, such as rider stress, which is not fully represented by telemetry alone. The marginal 0.15% gap between the predicted (94.82%) and Oracle (94.97%) results validates MotoSafety for real-world deployment where ground truth states are not available. 5.6 Model Performance and Reliability 5.6.1 Confusion Matrix Analysis Figure 7: Confusion matrix of MotoSafety predictions on the main PTW simulator dataset under TP. The confusion matrix provides a detailed summary of MotoSafety model classification outcomes for collision and non-collision cases. Figure 7 shows the distribution of accurate and inaccurate predictions. 5.6.2 Calibration Curve and Risk Scores The MST calibration curve on the primary PTW simulator dataset under TP (Fig. 9) shows that predicted collision probabilities closely match observed event frequencies, indicating statistically reliable and well-calibrated outputs. In safety-critical contexts, such calibration enables risk-aware adaptive interventions rather than reliance on fixed thresholds. The distribution of predicted risk scores by true outcome (Fig. 9) further illustrates clear separation: collision events consistently receive higher predicted probabilities, while non-collision samples remain concentrated at lower values. This separation reinforces both discriminative ability and calibration quality, supporting the potential for real-world deployment. Figure 8: Calibration curve. Figure 9: Distribution of predicted collision risk. 5.7 Architecture Transferability Across Safety Domains Table 12: Cross-Safety-Domain Evaluation of MotoSafety Accuracy (%) and Prior Work Accuracy (%). MST: MotoSafety. Dataset Details Safety Domain Application Country MST Prior Work Accuracy Proposed Dataset PTW Simulator with GT TP PTW Safety Collision Prediction India 94.97 – [41] PTW High-Fidelity Simulator PTW Safety Collision Prediction Germany 99.49 91.0 [RF, GB [42]] [4] PTW Real Time Fall Event Data PTW Safety Fall Detection France 96.10 91.59 [DT [8]] [40] Human Activity Recognition Human Safety Activity Recognition Italy 97.66 83.35 [K-means + NB [14]] [50] Wearable-Sensor Movement Human Clinical Safety Exercise Quality Assessment Turkey 99.65 88.65[MTMM-DTW [49]] Note that each dataset below is used to train and evaluate a separate MotoSafety instance from scratch (i.e., no weights are transferred from the PTW model); this experiment therefore evaluates the transferability of the architecture to other safety-critical sequence classification tasks, rather than knowledge transfer from the PTW domain itself. To evaluate the versatility of the MotoSafety framework beyond motorcycle simulators, we performed cross-dataset evaluations across distinct safety-critical domains as shown in Table 12. In the PTW Safety domain, MotoSafety achieved 99.49% accuracy on the high-fidelity German simulator dataset [41] for collision prediction, outperforming traditional ensemble methods (RF/GB) by 8.49%. Similarly, on real-world fall event data [4], achieved improved 96.10% accuracy, a significant improvement over the 91.59% reported for Decision Tree (DT) benchmarks [8]. Beyond vehicle dynamics, the model exhibited exceptional robustness in Human and Clinical Safety. On the UCI HAR human activity recognition dataset [40], MotoSafety surpassed Naive Bayes (83.35%) by 14.31%, achieving 97.66% accuracy. Most notably, in clinical exercise quality assessment [50], the framework achieved near-perfect classification (99.65%), marking an 11% gain over DTW-based clinical algorithms (88.65%) [49]. These results demonstrate that the MotoSafety architecture is not overfit to PTW-simulator-specific artifacts and performs competitively when retrained on diverse safety-critical sequence classification tasks, supporting the architecture’s applicability beyond its original design domain. 5.8 Ablation Study We evaluated individual contributions of MotoSafety components through a comprehensive ablation study shown in Table 13. The study was carried out on PTW simulation data under TP. The complete model attains an accuracy of 94.97%. The removal of BiLSTM and MHA resulted in significant performance degradation of 4.55% and 4.19%, respectively, highlighting their role in capturing long-range dependencies. Similarly, the exclusion of the CNN branch and Gated Fusion mechanism led to accuracy drops of 2.31% and 2.07%, underscoring the necessity of multi-scale feature extraction. Furthermore, removing the SE block decreases accuracy by 3.73%, validating its effectiveness in channel-wise feature recalibration. Collectively, these findings justify the necessity of each integrated module to maintain high predictive precision. Table 13: Ablation Study of MotoSafety Components. Variant Accuracy (%) Δ Acc. Full Model 94.97 – No BiLSTM 90.42 −4.55-4.55 No SE Block 91.24 −3.73-3.73 No Attention 90.78 −4.19-4.19 No CNN Branch 92.66 −2.31-2.31 No Gated Fusion 92.90 −2.07-2.07 Table 14: Comparison of MotoSafety with TIP against standard fixed temporal-pooling strategies (mean ± std over 5 runs). Pooling Strategy Accuracy (%) AUC (%) Mean-pooling 92.23 ± 0.15 98.80 ± 0.22 Max-pooling 93.41 ± 0.12 99.11 ± 0.18 Last-timestep 92.87 ± 0.14 98.97 ± 0.20 TIP 94.97 ± 0.08 99.33 ± 0.30 To further isolate the effect of TIP, we compare it against three standard temporal-pooling strategies—mean-pooling, max-pooling, and last-timestep extraction—under an otherwise identical architecture. As shown in Table 14, MotoSafety with TIP achieves 94.97% accuracy, outperforming mean-pooling by 2.74%, max-pooling by 1.56%, and last-timestep by 2.10%. This indicates that the performance gain stems from the learned, content-aware weighting rather than from the surrounding architecture alone. 5.9 Real-World Feature Feasibility To assess real-world deployment feasibility, we identified that 21 of the 64 features used in this study are available on standard motorcycles via low-cost IMU+GPS (e.g., speed, acceleration, yaw rate, position, brake force, throttle, steering angle). The remaining features (e.g., precise gap metrics, rule violations) are simulator-only. As shown in Table 15, accuracy increases with more features, achieving 93.91% with 21 features (compared to 94.97% with all 64 features), indicating that real-world deployment is feasible with onboard sensors. Testing on actual hardware in real traffic conditions remains essential and is therefore deferred to future work, as discussed in Section 6.1. Table 15: MotoSafety accuracy with real-bike features (IMU+GPS). Features Accuracy (%) 3 80.22 5 83.11 7 88.78 9 90.16 21 93.91 All 64 (simulator) 94.97 5.10 Model Complexity and Edge-AI Deployability The MotoSafety model is designed for real-time deployment in resource-constrained environments. Latency was measured on an Intel Core i7 CPU (2.8 GHz) with batch size 32, PyTorch 2.0, CPU-only inference, averaged over 1000 batches. With only 1.15M parameters and a 4.38 MB memory footprint (Table 16A), it maintains a low computational profile without sacrificing accuracy. Benchmarking on our PTW simulation dataset reveals that MotoSafety achieves a per-sample inference latency of 0.135 ms, outperforming all baseline and SOTA architectures (Table 16B). Specifically, MotoSafety is 5.3×, 9.5×, 21.9×, and 71.9× faster than PatchTST, Time-LLM, TimesNet, and LLM4TS (GPT-2), respectively. Even when accounting for TST preprocessing (36.88ms overhead) to predict TP labels when ground truth is unavailable, MotoSafety (37.01ms) remains faster than iTransformer (37.47ms), PatchTST (37.60ms), Time-LLM (38.16ms), TimesNet (39.84ms), and LLM4TS (46.59ms). This ultra-low latency ensures near-instantaneous risk assessment, allowing maximum time for emergency interventions. These efficiency metrics satisfy the stringent requirements for real-time deployment on low-cost on-board units (OBUs). Unlike GPU-dependent Transformer baselines, complexity of MotoSafety is (L)O(L). This balance of efficiency and performance makes it a viable candidate for large-scale ITS deployment in developing regions, providing an accessible AI-driven safety net for vulnerable riders. Table 16: MotoSafety Efficiency Metrics and Comparative Latency. A. MotoSafety Efficiency & Deployability Metric Value Significance Total Parameters 1,149,048 Low Complexity Model Size 4.38 MB Edge-ready Storage Throughput >7,400>7,400 s/sec High-speed Inference Arch. Complexity (L)O(L) Linear Scalability Edge-AI Deployability Yes Suitable for wearable Deployment Regions Global Operate in low-resource B. Latency Comparison (PTW Simulator Dataset Under TP) Model (Year) Latency (ms) Speed Gap Parameters MotoSafety 0.135 1.0× 1,149,048 iTransformer (2024) 0.587 4.3× ↓ 275,714 PatchTST (2023) 0.720 5.3× ↓ 187,330 Time-LLM (2024) 1.280 9.5× ↓ 3,183,618 TimesNet (2023) 2.960 21.9× ↓ 4,708,226 LLM4TS (2025) 9.710 71.9× ↓ 124,506,626 6 Conclusion In this work, we investigated how TP affects PTW riders and contributes to collision risk. To fill this gap, we collected a large-scale dataset comprising over 129,000 labeled multivariate time-series sequences from 153 simulator rides conducted by 51 participants under no, low, and high TP conditions. Each sequence includes 64 features covering vehicle dynamics, control inputs, proximity measures, temporal context, and behavioral violations. Building on this dataset, we proposed MotoSafety, a novel deep learning architecture grounded in the LTI principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC for collision risk assessment, outperforming ten baselines including TimesNet, PatchTST, iTransformer, Time-LLM, and LLM4TS. For long-term forecasting, it achieves 0.039 MSE and 0.094 MAE (4.4× lower error than Time-LLM and iTransformer). With only 1.15 million parameters and 0.135 ms inference latency, the model is 21.9× and 71.9× faster than TimesNet and LLM4TS. Our findings show that explicit TP prediction provides a critical inductive bias, improving collision accuracy from 94.09% to 94.97% (ground truth) and 94.82% with predicted TP. Furthermore, using only 21 real-world features (available via low-cost IMU+GPS), MotoSafety achieves 93.91% accuracy, indicating practical deployment potential; real-world hardware validation remains future work. Beyond PTW safety, the MotoSafety architecture shows improved transferability when retrained on human activity recognition (97.66%) and clinical exercise monitoring (99.65%) datasets. The model’s small size (1.15M parameters, 4.38 MB) and low latency (0.135 ms) make it well-suited for future edge deployment on handlebar-mounted devices or smart helmets as a potential edge-deployable black box for two-wheelers, enabling real-time collision risk alerts without relying on cloud connectivity. 6.1 Limitations and Future Work We acknowledge several limitations in the present study. First, our participant pool consisted solely of male riders, which aligns with Indian fatality statistics showing that males account for 85.2–87.3% of PTW deaths. However, this limits the generalizability of our findings to female riders, since perceptions of TP and riding behavior may vary between genders. Future studies should include female riders to systematically examine whether the observed relationships hold across genders. Second, the data were gathered in a controlled simulator setting, which may not completely capture real-world factors like rider fatigue, changing weather, or social pressures. Real-world TP collision data cannot be collected safely, so simulators provide the only ethical controlled setting. To validate in real conditions, we plan phased testing: closed-course trials followed by naturalistic data collection through commercial riding platforms. Third, we did not record physiological stress markers, such as heart rate variability or electrodermal activity. Behavioral markers such as speed variability and control inputs were used as indicators of cognitive load; however, the absence of physiological validation limits direct assessment of the underlying stress response. Future work should incorporate physiological measures to complement behavioral markers. Fourth, the proposed safety system requires prospective validation in real-world operational settings before practical deployment. Hardware validation on handlebar-mounted devices or smart helmets remains future work. Future research will assess MotoSafety across varied rider populations and road conditions using both simulator-based and naturalistic datasets. We plan to incorporate multimodal signals—including physiological, behavioral, and contextual data—to improve robustness. Adaptive ITS interventions, such as haptic feedback, throttle modulation, and context-sensitive alerts, will be explored to mitigate crash risk and manage kinetic energy transfer. Transfer learning and domain adaptation techniques will also be applied to ensure the framework is scalable across different vehicle types, regions, and cultural contexts. Data and Code Availability The data and code supporting this study will be made available by the corresponding author upon reasonable request after publication. Funding This research received funding from the IIT Indore Young Faculty Research Catalyzing Grant (YFRCG) Scheme under Project ID: IITI/YFRCG/2023-24/01. Acknowledgments The authors gratefully acknowledge all the volunteers who took part in this study. We also extend our sincere thanks to Manvendra T. for his valuable assistance with data collection. Declaration of Competing Interest The authors declare that they have no known competing financial interests or personal relationships that could have influenced the work reported in this study. References [1] Ç.İ. Acı, G. Mutlu, M. Ozen, and M. Acı (2025) Enhanced multi-class driver injury severity prediction using a hybrid deep learning and random forest approach. Applied Sciences 15. Cited by: §1. [2] M. Atif and G. Sil (2026) Modeling the effects of driver and road geometric characteristics on consecutive horizontal curve perception. Transportation Research Record 2680 (3), p. 286–310. Cited by: §3.1.3. [3] G. H. Bham and M. C. Leu (2018) A driving simulator study to analyze the effects of portable changeable message signs on mean speeds of drivers. Journal of Transportation Safety & Security 10 (1-2), p. 45–71. Cited by: §2.2. [4] A. Boubezoul, F. Dufour, S. Bouaziz, B. Larnaudie, and S. Espié (2019) Dataset on powered two wheelers fall and critical events detection. Data in Brief 23, p. 103828. External Links: ISSN 2352-3409 Cited by: §5.7, Table 12. [5] S. Bouhsissin, N. Sael, F. Benabbou, and A. Soultana (2024) Enhancing machine learning algorithm performance through feature selection for driver behavior classification. Indonesian Journal of Electrical Engineering and Computer Science 35 (1), p. 354–365. Cited by: §1. [6] S. Bouhsissin, N. Sael, F. Benabbou, and A. Soultana (2023) Enhancing machine learning algorithm performance through feature selection for driver behavior classification. International Journal of Electrical and Computer Engineering Systems 35 (1), p. 354–365. Cited by: §2.3. [7] C. Chang, W. Wang, W. Peng, and T. Chen (2025) LLM4TS: aligning pre-trained llms as data-efficient time-series forecasters. ACM Trans. Intell. Syst. Technol. 16 (3). External Links: ISSN 2157-6904 Cited by: §4.0.2. [8] F. Elwy, R. Aburukba, A. R. Al-Ali, A. A. Nabulsi, A. Tarek, A. Ayub, and M. Elsayeh (2023) Data-driven safe deliveries: the synergy of iot and machine learning in shared mobility. Future Internet 15 (10). External Links: ISSN 1999-5903 Cited by: §5.7, Table 12. [9] S. Gashaw, P. Goatin, and J. Härri (2018) Modeling and analysis of mixed flow of cars and powered two wheelers. Transportation Research Part C: Emerging Technologies 89, p. 148–167. External Links: ISSN 0968-090X Cited by: §2.1. [10] M. Gupta, N. M. Pawar, N. R. Velaga, and S. Mishra (2022) Modeling distraction tendency of motorized two-wheeler drivers in time pressure situations. Safety Science 154, p. 105820. Cited by: §2.1. [11] M. Gupta and N. R. Velaga (2024) Motorized two-wheeler riders’ rear brake application in sudden hazardous event of animal crossing. Transportation Letters 16 (10), p. 1268–1275. Cited by: §1, §2.1. [12] M. Gupta and N. R. Velaga (2026) Dynamic dilemma zone at signalized intersection: attention allocation patterns using cure survival analysis for male riders. Accident Analysis & Prevention 228, p. 108408. External Links: ISSN 0001-4575 Cited by: §3.1.3. [13] International Electrotechnical Commission (2023) Potentiometers for use in electronic equipment. Standard Technical Report IEC 60393, International Electrotechnical Commission. Cited by: §3.1.1. [14] D. P. Ismi, S. Panchoo, and M. Murinto (2016) K-means clustering based filter feature selection on high dimensional data. International Journal of Advances in Intelligent Informatics 2, p. 38–45. Cited by: Table 12. [15] Y. Jiang, X. Qu, W. Zhang, W. Guo, J. Xu, W. Yu, and Y. Chen (2025) Analyzing crash severity: human injury severity prediction method based on transformer model. Vehicles 7 (1), p. 5. Cited by: §1. [16] M. Jin, S. Wang, L. Ma, Z. Chu, J. Y. Zhang, X. Shi, P. Chen, Y. Liang, Y. Li, S. Pan, and Q. Wen (2024) Time-llm: time series forecasting by reprogramming large language models. Cited by: §4.0.2. [17] X. Kong, S. Das, K. Jha, and Y. Zhang (2020) Understanding speeding behavior from naturalistic driving data: applying classification based association rule mining. Accident Analysis & Prevention 144, p. 105620. External Links: ISSN 0001-4575 Cited by: §1. [18] E. Krug (2022) It’s time to end deaths on our roads. World Health Organization. Cited by: §2.4. [19] J. R. Landis and G. G. Koch (1977) The measurement of observer agreement for categorical data. Biometrics 33 (1), p. 159–174. Cited by: §4.0.1. [20] S. Leung, R. J. Croft, M. L. Jackson, M. E. Howard, and R. J. Mckenzie (2012) A comparison of the effect of mobile phone use and alcohol consumption on driving simulation performance. Traffic Injury Prevention 13 (6), p. 566–574. Cited by: §2.1. [21] X. Li, O. Oviedo-Trespalacios, A. Rakotonirainy, and X. Yan (2019) Collision risk management of cognitively distracted drivers in a car-following situation. Transportation Research Part F: Traffic Psychology and Behaviour 60, p. 288–298. External Links: ISSN 1369-8478 Cited by: §2.2. [22] Y. Liu, T. Hu, H. Zhang, H. Wu, S. Wang, L. Ma, and M. Long (2024) ITransformer: inverted transformers are effective for time series forecasting. Cited by: §4.0.2. [23] MORTH (2019) Road accidents in india – 2019. Technical report Transport Research. Cited by: §1, §3.1.3, §3.1.3. [24] MoRTH (2020) Road accidents in india – 2020. Technical report Transport Research. Cited by: §1, §3.1.3, §3.1.3. [25] MoRTH (2021) Road accidents in india – 2021. Technical report Transport Research. Cited by: §1, §3.1.3, §3.1.3. [26] MORTH (2022) Road accidents in india – 2022. Technical report Transport Research. Cited by: §1, §3.1.3, §3.1.3. [27] MORTH (2023) Ministry of road transport & highway, 2022-23.. Book 2022–2023. External Links: ISSN Cited by: §1. [28] MoRTH (2023) Road accidents in india – 2023. Technical report Transport Research. Cited by: §1, §3.1.3, §3.1.3. [29] MORTH (2024) Ministry of road transport & highway, 2023-24.. Book 2023–2024. External Links: ISSN Cited by: §1. [30] MORTH (2025) Ministry of road transport & highway, 2024-25.. Book 2024–2025. Cited by: §1. [31] A. M. Mostafa, B. Aldughayfiq, M. Tarek, A. S. Alaerjan, H. Allahem, M. K. Elbashir, M. Ezz, and E. Hamouda (2025) AI-based prediction of traffic crash severity for improving road safety and transportation efficiency. Scientific Reports 15, p. 27468. Cited by: §1. [32] Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam (2023) A time series is worth 64 words: long-term forecasting with transformers. Cited by: §4.0.2. [33] G. Pan, G. Wang, H. Wei, Q. Chen, and A. Zhang (2024) Development of an automated global crash prediction model with adaptive feature selection of deep neural networks. IEEE Transactions on Industrial Informatics 20, p. 12010–12020. Cited by: §1. [34] I. Pavlidis, M. Dcosta, S. Taamneh, M. Manser, T. Ferris, R. Wunderlich, E. Akleman, and P. Tsiamyrtzis (2016) Dissecting driver behaviors under cognitive, emotional, sensorimotor, and mixed stressors. Scientific Reports 6, p. 25651. Cited by: §1. [35] N. M. Pawar, N. R. Velaga, and S. Mishra (2022) Impact of time pressure on acceleration behavior and crossing decision at the onset of yellow signal. Transportation Research Part F: Traffic Psychology and Behaviour 87, p. 1–18. External Links: ISSN 1369-8478 Cited by: §2.1, §2.1. [36] N. M. Pawar, N. R. Velaga, and R.B. Sharmila (2022) Exploring behavioral validity of driving simulator under time pressure driving conditions of professional drivers. Transportation Research Part F: Traffic Psychology and Behaviour 89, p. 29–52. External Links: ISSN 1369-8478 Cited by: §2.1, §2.2. [37] N. M. Pawar and N. R. Velaga (2022) Analyzing the impact of time pressure on drivers’ safety by assessing gap-acceptance behavior at un-signalized intersections. Safety Science 147, p. 105582. Cited by: §1, §1, §2.1, §2.1. [38] T. H. S. Reporter (2022) Rash driving, fall in income, more pressure: behind the scenes of 10-minute delivery. The Hindu. Cited by: §1. [39] Reuters (2022) Delivery race brings road safety risks in india. The Jakarta Post. Cited by: §1. [40] J. Reyes-Ortiz, D. Anguita, A. Ghio, L. Oneto, and X. Parra (2013) Human Activity Recognition Using Smartphones. Note: UCI Machine Learning RepositoryDOI: https://doi.org/10.24432/C54S4K Cited by: §5.7, Table 12. [41] M. Rodegast et al. (2024) Motorcycle collision dataset. Universität Stuttgart. Cited by: §5.7, Table 12. [42] P. Rodegast, S. Maier, J. Kneifl, and J. Fehr (2024) On using machine learning algorithms for motorcycle collision detection. Discover Applied Sciences 6 (6), p. 326. External Links: ISSN 3004-9261 Cited by: §2.3, Table 12. [43] I. Sharma, M. Gupta, S. Mishra, and N. R. Velaga (2025) Exploring the impact of time pressure on motorized two-wheeler riders’ over-speeding behavior. Transportation Letters 17 (4), p. 595–611. Cited by: §1, §1, §2.1, §2.1. [44] F. I. Staff (2025) Racing against time: quick commerce is pushing delivery riders to the edge, claims study. Fortune India. Cited by: §1, §2.1. [45] Transportation Research and Injury Prevention Centre (2023) Road safety in india. Technical Report Indian Institute of Technology Delhi, New Delhi, India. Cited by: §3.1.3. [46] WHO (2023) Global status report on road safety, 2023.. World Health Organization.. Cited by: §1. [47] World Health Organization (2023) Road traffic injuries. Cited by: §2.4. [48] H. Wu, T. Hu, Y. Liu, H. Zhou, J. Wang, and M. Long (2023) TimesNet: temporal 2d-variation modeling for general time series analysis. Cited by: §4.0.2. [49] A. Yurtman and B. Barshan (2014) Automated evaluation of physical therapy exercises using multi-template dynamic time warping on wearable sensor signals. Computer Methods and Programs in Biomedicine 117 (2), p. 189–207. External Links: ISSN 0169-2607 Cited by: §5.7, Table 12. [50] A. Yurtman and B. Barshan (2014) Physical Therapy Exercises. Cited by: §5.7, Table 12. [51] A. Zeng, M. Chen, L. Zhang, and Q. Xu (2023) Are transformers effective for time series forecasting?. In Proceedings of the Thirty-Seventh AAAI Conference on AI and Thirty-Fifth Conference on Innovative Applications of AI and Thirteenth Symposium, AAAI’23/IAAI’23/EAAI’23. External Links: ISBN 978-1-57735-880-0 Cited by: §2.3. [52] G. Zerveas, S. Jayaraman, D. Patel, A. Bhamidipaty, and C. Eickhoff (2020) A transformer-based framework for multivariate time series representation learning. Cited by: §5.5. [53] H. Zhou, S. Zhang, J. Peng, S. Zhang, J. Li, H. Xiong, and W. Zhang (2021) Informer: beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, p. 11106–11115. Cited by: §2.3, §4.0.2.