1. Introduction
Conditionally automated driving can reduce driver workload during routine traffic operations, yet safe system use still depends on timely and appropriate human takeover when automation reaches its operational limits or when a safety-critical event emerges [
1]. In such situations, takeover quality is not defined solely by whether drivers respond to a request, but also by whether they can regain control with sufficient stability, situational awareness, and safety margin [
2]. This issue becomes more critical when takeover requests occur under degraded weather conditions or in scenarios involving elevated traffic conflict, where drivers must rapidly recover task-relevant attention and execute control actions within a limited response window [
3].
Among the available human–machine interface (HMI) strategies for takeover support, auditory alerts are widely used because they can capture driver attention without directly competing with the visual demands of roadway monitoring [
4]. However, the effectiveness of auditory alerts depends not only on their presence, but also on how urgent they are perceived to be [
5]. Alerts that are too weak may fail to promote timely control transition, whereas overly urgent alerts may increase tension or induce abrupt responses [
6]. Accordingly, alert urgency should be evaluated not only as a perceptual characteristic, but also as a design factor that may directly influence takeover quality and safety performance under realistic operational conditions [
7].
Weather and scenario complexity further complicate this problem [
8]. Reduced visual clarity in rainy environments may increase perceptual difficulty and delay the recovery of situation awareness [
9]. Likewise, different conflict structures may impose different takeover demands [
10]. A lead-brake event primarily challenges longitudinal response and distance control, whereas a cut-in event introduces stronger spatial conflict and time pressure, thereby increasing the need for rapid hazard interpretation and control adaptation [
11]. From an engineering perspective, takeover support should therefore be assessed within a combined framework in which alert design is examined together with environmental and scenario complexity rather than in isolation [
12].
Recent takeover-support research has increasingly moved beyond the simple presence or absence of a takeover request and has examined how warning modality, signal attributes, driver state, and adaptive support logic jointly influence takeover quality. Studies on warning modality have shown that auditory, visual, haptic, and multimodal requests may differ in their ability to capture attention, support control transition, and avoid competition with the visual monitoring task [
4,
5,
6,
7]. Research on auditory warning design further indicates that perceived urgency is shaped by signal attributes rather than by warning availability alone, suggesting that alert urgency should be treated as a manipulable design factor rather than as a merely perceptual outcome [
4,
6,
7].
A second stream of research emphasizes that takeover performance is also conditioned by driver state and task demand before the request. Driver state- and workload-oriented studies have shown that attention allocation, fatigue, trust, cognitive load, and prior monitoring behavior may affect takeover timing, control initiation, and safety-related response quality [
6,
9,
13]. Studies considering driver characteristics and takeover-time budgets further indicate that takeover performance is not a fixed response to a warning stimulus, but is shaped by individual background, available response time, and the operational context in which the request occurs [
8,
13,
14]. These findings support the need to interpret auditory warning effects in relation to both driver readiness and task demand.
A third line of work has begun to move toward adaptive takeover support, in which warning timing, warning intensity, or warning modality is adjusted according to situational risk, uncertainty, driver state, or conflict severity [
6,
7,
12,
14]. This direction is important because it shifts takeover support from a one-level warning design toward a context-sensitive assistance logic. However, much of the existing evidence still treats warning design, driver state, environmental burden, and scenario conflict as separate analytical topics. Fewer studies have examined how auditory alert urgency interacts with weather condition and scenario demand within a single simulator framework centered on takeover timing, behavioral-control response, and safety-margin outcomes. In addition, although eye-tracking data can help explain whether drivers visually reorient to task-relevant regions after an alert, such evidence is retained in the present study only as limited supportive information rather than as a primary evaluation basis [
15].
To address this issue, the present study investigates the effects of auditory alert urgency on takeover safety under different weather and scenario conditions in conditionally automated driving. The study is based on a controlled driving-simulator experiment involving 40 female licensed drivers. Participants completed a 3 × 2 × 2 within-subject design in OpenDS 4.5 4.5, covering three alert urgency levels, two weather conditions, and two representative takeover scenarios, resulting in 12 takeover tasks per participant.
The contribution of the present study is not the proposal of a new takeover algorithm or a new warning modality. Instead, it lies in providing a controlled context-sensitive evaluation of auditory warning urgency under combined environmental and scenario demands. First, the study examines auditory alert urgency together with weather condition and conflict scenario within the same 3 × 2 × 2 within-subject framework, rather than treating warning urgency as a context-free warning property. Second, it organizes the evidence into a hierarchical structure in which OpenDS 4.5-derived takeover timing, behavioral-control indicators, and safety-margin measures form the primary evaluation layer, whereas eye-tracking and questionnaire measures are retained as supporting evidence for visual reorientation and condition appraisal. Third, it translates the empirical patterns into a proof-of-concept warning-configuration logic, while explicitly treating this logic as simulator-based design guidance rather than as a production-ready HMI specification.
To structure the analysis, the present study adopts a staged framework linking task conditions, takeover process, safety-related outcomes, and supplementary appraisal evidence. Specifically, auditory alert urgency, weather condition, and takeover scenario are treated as task conditions; takeover timing, control-related behavior, and visual-attention recovery are treated as process-level indicators; safety margins are treated as downstream outcome measures; and questionnaire evaluations are used as supplementary appraisal evidence.
The study addresses the following research questions:
RQ1 (Primary takeover outcomes): How does auditory alert urgency affect primary takeover execution and safety-related outcomes in conditionally automated driving, as reflected in takeover timing, control-related behavior, and retained safety-margin indicators?
RQ2 (Context dependence): Do these urgency-related effects vary across weather and scenario conditions, indicating that takeover support should be interpreted in relation to both perceptual demand and conflict structure rather than as a context-free warning property?
RQ3 (Supporting explanatory evidence): Do secondary eye-tracking indicators and supplementary questionnaire appraisals provide supporting evidence for interpreting the observed variation in primary takeover behavior and safety outcomes, without displacing the analytical priority of behavior- and safety-based measures?
2. Materials and Methods
2.1. Participants
A total of 40 female licensed drivers participated in the experiment. All participants were adult drivers with valid driving licenses and recent real-world driving experience. Before the formal experiment, each participant completed a background questionnaire covering basic demographic information, driving experience, recent driving exposure, prior use of advanced driver assistance systems (ADAS), and familiarity with automated driving (AD)-related functions.
Participants were recruited on a voluntary basis and were screened before the formal session. To ensure suitability for simulator-based and eye-tracking-based testing, all participants were required to have normal or corrected-to-normal vision and to report no severe motion sickness history or medical conditions that could interfere with driving-related operations or visual tasks. The present study intentionally adopted a female-only sample in order to establish a more targeted empirical baseline for this driver group while reducing additional between-subject variance associated with driver-characteristic heterogeneity in takeover-related behavior and response patterns [
16,
17]. This sampling strategy was not intended to represent the full driving population, but rather to support a more internally consistent examination of takeover safety and related supplementary appraisal within a clearly defined participant group. This design choice should therefore be interpreted as a controlled sampling strategy rather than as a claim of population-level representativeness. Previous takeover studies and reviews have suggested that driver characteristics may contribute to variability in takeover timing, control initiation, steering and braking responses, and risk-related behavior [
16,
17]. However, the findings should not be directly generalized to male drivers or mixed-gender driver populations without further validation. Participant characteristics are summarized in
Table 1.
2.2. Experimental Design
The experiment adopted a 3 × 2 × 2 within-subject design. The first independent variable was auditory alert urgency, with three levels: low urgency, medium urgency, and high urgency. The second independent variable was weather condition, with two levels: sunny and rainy. The third independent variable was takeover scenario, with two levels: a lead-brake scenario and a cut-in scenario. The full combination of these factors yielded 12 takeover tasks for each participant, and each participant completed all 12 tasks.
The low-, medium-, and high-urgency alerts were designed to induce different levels of perceived urgency while retaining the same functional role as takeover-support warnings [
18].
To improve the reproducibility of the auditory-warning manipulation, the main acoustic characteristics of the three alert sounds are summarized in
Table 2. The three alerts served the same functional role as takeover requests but differed primarily in dominant frequency and perceived pitch. The relative RMS amplitude of the three sounds was comparable, and the playback volume setting was kept constant across participants and experimental conditions.
The sunny and rainy conditions represented two levels of environmental visibility and perceptual demand [
19]. The lead-brake scenario represented a longitudinal conflict in which the participant was required to resume control in response to a decelerating lead vehicle [
20]. The cut-in scenario represented a lateral conflict in which another vehicle intruded into the participant’s path, thereby creating stronger spatial conflict and time pressure [
21].
A within-subject design was adopted to reduce the influence of inter-individual variability on takeover performance, visual attention, and subjective appraisal. All factor combinations were included once for each participant. The task schedule was predefined at the trial-plan level, but the presentation order of the 12 takeover tasks was varied across participants rather than administered in a single fixed sequence. This participant-specific ordering was used to reduce simple sequence-related bias, including learning, habituation, and fatigue effects, while preserving one-to-one matching among OpenDS 4.5 records, supplementary records, and event logs during subsequent analysis.
Figure 1 presents the overall experimental framework of the study, including the participant sample, the 3 × 2 × 2 within-subject design, the synchronized multimodal data streams, and the hierarchical evidence structure adopted for analysis.
As shown in
Figure 1, the present study treated OpenDS 4.5-derived driving and safety indicators as the primary outcome layer, while only limited supporting measures were retained for interpretation and condition appraisal. The kinematic parameters of the two hazard scenarios were fixed at the scenario-template level and were not adjusted across participants within the same experimental condition. Detailed kinematic configuration information for the lead-brake and cut-in events is provided in
Supplementary Table S1.
2.3. Apparatus and Experimental Environment
The driving experiment was conducted using OpenDS 4.5 4.5 as the simulation platform in a controlled laboratory setting. Participants completed the takeover tasks using a conventional keyboard-and-mouse interface. Steering and braking variables were therefore interpreted as normalized control-input measures for within-experiment comparison. Accordingly, these control-input measures were not interpreted as direct physical equivalents of steering-wheel angle, pedal force, or pedal displacement in a real vehicle. Their role in the present analysis was to support controlled within-experiment comparisons across alert-urgency, weather, and scenario conditions under the same input interface. The simulator environment was configured to implement the predefined takeover tasks under the corresponding weather and scenario conditions. Vehicle trajectory and control data were recorded throughout each task, including variables related to speed, steering, braking, and surrounding-vehicle interactions. These OpenDS 4.5-derived records served as the primary source for behavioral and safety-related analysis.
Eye-tracking data were collected simultaneously using the D-Lab system (DL3) based on Dikablis Glasses 3 (Ergoneers GmbH, Manching, Germany), recording at approximately 60 Hz with binocular eye tracking and a forward-facing scene camera. In line with the analytical focus of the present study, the eye-tracking analysis was restricted to a limited set of area of interest (AOI)- and glance-based indicators associated with the main task-relevant display region. These measures were retained only as secondary support for interpreting visual reorientation during takeover.
All experiments were conducted in a controlled laboratory environment. Standard device fitting and calibration procedures were completed before formal data collection. The same simulation environment, task structure, and device setup were used for all participants. This consistency was important for maintaining comparability across participants and for preserving the validity of the multimodal dataset.
Figure 2 illustrates the experimental conditions used in the takeover tasks, including the two weather settings, the two representative scenario types, and the three auditory alert-urgency levels.
These condition combinations formed the 12 matched takeover tasks completed by each participant in the formal experiment.
2.4. Measures
The dependent measures were organized into three layers. First, behavioral- and control-related indicators were treated as the primary evidence layer. To preserve a clearer process-to-safety interpretation of takeover quality, the primary OpenDS 4.5-derived outcome set was organized into three components: takeover-process indicators, strict same-lane safety-margin indicators, and broader behavioral-control outcomes. The process indicators included response occurrence and response time to first effective control input (RTcontrol); the safety-margin indicators included minimum time-to-collision (TTC) and minimum time headway (THW) under the strict same-lane criterion; and the broader behavioral-control outcomes included takeover duration, mean speed, braking response, steering-related behavior, and minimum center distance. These OpenDS 4.5-derived measures were evaluated as the primary basis for condition effects.
Eye-tracking measures derived from the D-Lab system were retained only as limited secondary indicators of visual reorientation after the alert [
22]. Consistent with this restricted role, only a small set of task-relevant AOI- and glance-based measures was included: first glance latency, glance count, total glance duration, mean glance duration, and AOI ratio. These measures were used only to provide supportive evidence on whether participants visually reoriented toward the task-relevant region during takeover. Because several gaze-derived indicators were sensitive to gaze-continuity quality and event-alignment uncertainty, these measures were not used as primary endpoints. They were retained only to provide conservative support for interpreting whether participants visually reoriented toward the task-relevant display region after the alert.
Post-experiment questionnaire responses were treated as supplementary measures and were used to provide an additional interpretive layer for the experimental findings. The questionnaire covered several categories of subjective evaluation, including perceived alert urgency, perceived helpfulness and acceptability of the auditory alerts, perceived effects of weather and scenario difficulty, and overall evaluations of trust, acceptance, and perceived workload. Specifically, the questionnaire covered alert-related evaluations, comparative appraisals of weather- and scenario-related difficulty, and broader retrospective impressions of reliability, trust, acceptance, perceived safety benefit, burden, and task-related tension. Because the questionnaire was administered after completion of all takeover tasks, these ratings were treated as retrospective condition evaluations and overall system impressions rather than as trial-level subjective responses. Accordingly, the questionnaire results were used to complement the interpretation of the behavioral and eye-tracking findings, not to replace the behavior- and safety-based evidence. The operational definitions of the retained measures are summarized in
Table 3. Detailed kinematic configuration information for the fixed hazard scenarios is provided in
Supplementary Table S1. The overall experimental design and retained dependent measures are summarized in
Table 4.
2.5. Takeover-Process and Safety-Margin Indicators
To preserve a more explicit process-to-safety interpretation of takeover quality, the present study additionally retained response occurrence and RTcontrol as takeover-process indicators, and minTTC and minTHW as strict same-lane surrogate safety-margin indicators, following the operational definitions summarized in
Table 3. The retained post-alert analysis window was defined as the 10 s interval after alert onset and was truncated at trial termination where necessary. Under this framework, the takeover-process indicators captured whether and how quickly effective manual control was initiated, whereas the strict same-lane safety-margin indicators characterized the downstream temporal safety margins that remained available under the retained conflict geometry. The
Supplementary Materials provide the corresponding strict same-lane computability rules, field-name mapping, missingness summaries, and sensitivity-oriented summaries for the process and safety-margin indicators.
2.6. Experimental Procedure
Upon arrival, each participant received an explanation of the study and completed the background questionnaire. Informed consent was obtained before the formal experiment began. The experimenter then introduced the simulator task, the takeover procedure, and the general operation of the test system. After the eye-tracking device had been fitted and calibrated, the experimenter checked the quality of the binocular recording and the stability of the scene-camera view.
Before the formal session, participants completed a practice phase to become familiar with the simulator environment, the driving controls, and the general structure of the takeover task. This familiarization stage was intended to reduce unnecessary learning effects during the formal trials and to ensure that participants understood the operational meaning of the takeover request.
During the formal experiment, each participant completed 12 takeover tasks corresponding to the full combination of the three alert urgency levels, two weather conditions, and two scenario types. In each task, the participant experienced an automated-driving period followed by a takeover request presented through an auditory alert. After the alert, the participant was required to resume manual control and respond appropriately to the traffic situation. OpenDS 4.5 recorded the full driving process throughout each task, while supporting eye-tracking data were recorded simultaneously for matched trials.
The procedure was designed to preserve the continuity of the simulator and eye-tracking session. In particular, repeated questionnaire interruptions were not inserted between individual takeover tasks. This design choice reduced task fragmentation and helped limit possible head movement or eye-tracking instability caused by frequent session interruptions. After all 12 tasks had been completed, participants answered the post-experiment questionnaire covering urgency perception, environmental difficulty, scenario appraisal, trust, acceptance, and perceived workload. At the end of the session, participants were debriefed and received the planned experimental compensation.
2.7. Data Processing and Analysis
Before formal analysis, a multimodal consistency check was conducted across the OpenDS 4.5 records, eye-tracking files, event logs, and questionnaire responses. Participant identifiers and trial identifiers were cross-checked across the three main data sources. The dataset used in the present study contained 40 questionnaire records, 480 OpenDS 4.5 takeover tasks, 480 record-start entries, 480 alert-onset entries, and 480 corresponding D-Lab trial files. Only matched participant-level and trial-level records were retained for joint analysis. For cases involving repeated or redundant event anchors, the earliest valid alert-onset entry was retained according to the cleaned event log.
OpenDS data were then processed to derive behavior- and safety-related indicators for each task. Eye-tracking data were screened to identify usable task files before the retained AOI- and glance-based measures were summarized. Data quality control focused on whether usable gaze information was available within the key takeover-analysis window for the limited retained indicators. Trials with insufficient recording quality, unstable binocular tracking, or inadequate usable gaze information in the critical response window were excluded from the corresponding eye-tracking analysis. Questionnaire responses were coded as retrospective condition evaluations and overall subjective assessments. High-risk and failure-related tasks were retained as valid safety-relevant observations rather than removed as outliers. In particular, collision-terminated trials and near-contact or contact-risk trials were treated as meaningful lower-tail safety outcomes of the takeover process. These events were not discarded during data cleaning; instead, they were retained in the archived dataset and considered during event-level safety audit.
The statistical analysis was organized around the central question of how alert urgency, weather condition, and scenario condition influenced takeover safety. For the continuous OpenDS 4.5-derived primary outcomes, linear mixed models were used to test fixed effects of alert urgency, weather condition, scenario condition, and their interactions under the repeated-measures design. Participant was treated as a random intercept to account for within-driver correlation across repeated trials. This modeling strategy was adopted to provide a unified inferential framework for the OpenDS 4.5-derived primary outcomes and, in particular, to accommodate the structurally unbalanced repeated-measures structure of the strict same-lane safety-margin indicators without requiring participant-level listwise deletion. Linear mixed models were fitted to the OpenDS 4.5-derived primary continuous outcomes, and fixed effects were evaluated using Wald chi-square tests.
For the three alert-related questionnaire ratings, namely perceived urgency, perceived helpfulness, and subjective acceptability, Friedman tests were used because these repeated single-item responses were collected on ordinal Likert-type scales. When the omnibus Friedman test was significant, post hoc pairwise comparisons were conducted using Wilcoxon signed-rank tests with Bonferroni correction. The remaining questionnaire items were summarized descriptively in the current presentation.
The eye-tracking measures were analyzed only as limited secondary indicators and were interpreted conservatively when gaze continuity or event alignment was unstable. Because the retained outcomes represented different mechanism-oriented dimensions of takeover performance, including initiation timing, strict same-lane safety margins, broader behavioral-control responses, and supplementary appraisal measures, the present study analyzed them separately rather than collapsing them into a single multivariate omnibus test. At the same time, to reduce overinterpretation under repeated outcome testing, particular emphasis was placed on effects that were theoretically central, statistically robust, and supported by non-trivial effect sizes rather than on isolated marginal p-values.
3. Results
In line with the analytical hierarchy defined in the
Section 2, the results are presented from the primary behavior- and safety-based outcomes to the secondary eye-tracking indicators and then to the supplementary questionnaire appraisals. This organization preserves the central role of takeover-related driving performance while using visual-attention and subjective measures as supporting evidence for interpretation.
3.1. Overall Experimental Results
After participant-level and trial-level matching across the three data sources, the final dataset retained all 40 participants and all 480 formal takeover tasks. Specifically, the analysis dataset comprised 480 OpenDS 4.5 task records, 480 D-Lab eye-tracking trial files, 480 record-start anchors, 480 alert-onset anchors, and 40 post-experiment questionnaires. No participant-level exclusions were made after the consistency check, and no trial was removed solely because of elevated risk or unfavorable outcome, because such trials represented meaningful safety-related observations within the takeover process.
At the overall level, the objective results showed that scenario condition remained the most stable determinant across the retained OpenDS 4.5-derived primary outcomes, whereas urgency-related effects were more selective and weather effects were generally weaker in the objective measures. The limited eye-tracking results showed only a narrow retained pattern, with total glance duration providing the clearest secondary difference. In the questionnaire-based supplementary outcomes, high-urgency alerts were rated as the most urgent and the most helpful, whereas medium-urgency alerts received the highest acceptability ratings.
3.2. Collision and Contact-Risk Audit
A post hoc event-level audit was conducted for the archived OpenDS 4.5 trials to identify collision-terminated and contact-risk events. Within the 480 retained takeover tasks, one trial was classified as a high-confidence collision-terminated event, namely P012 trial_09 under the lead-brake, sunny, and low-urgency condition. In addition, P030 trial_08 under the cut-in, sunny, and high-urgency condition was retained as a high-risk near-contact event because it showed elevated contact risk but was not conservatively classified here as a definitive collision-terminated trial. No additional trials were treated as confirmed collisions in the current presentation. These events were retained as valid safety-relevant observations rather than excluded from the dataset.
3.3. Results of Takeover-Process Indicators
The takeover-process indicators showed that condition effects were more evident in takeover initiation timing than in the mere occurrence of effective control. Across the 480 retained takeover tasks, effective control input was detected within the retained post-alert window in 96.9% of trials. Response occurrence remained high overall, but the lowest rate was observed under low-urgency alerts (95.0%), compared with 98.1% under medium urgency and 97.5% under high urgency. By scenario, response occurrence was slightly lower in cut-in trials (95.4%) than in lead-brake trials (98.3%); by weather, it was also slightly lower under rainy conditions (95.8%) than under sunny conditions (97.9%).
By contrast, RTcontrol showed clear condition-related variation. Across response-valid trials, the overall mean RTcontrol was 4.88 ± 0.78 s. The longest RTcontrol was observed under low urgency (M = 6.15 ± 1.05 s), followed by medium urgency (M = 4.77 ± 0.97 s), whereas high urgency was associated with the shortest RTcontrol (M = 3.75 ± 1.21 s). Analyzed under a sum-coded linear mixed-model framework with overall term tests, RTcontrol showed a strong fixed effect of alert urgency, χ
2(2) = 232.98,
p < 0.001, together with significant fixed effects of weather, χ
2(1) = 14.69,
p < 0.001, and scenario, χ
2(1) = 3.96,
p = 0.047. In the participant-level collapsed means shown in
Table 5, low urgency was associated with a 2.40 s longer RTcontrol than high urgency and a 1.38 s longer RTcontrol than medium urgency, indicating a substantial timing penalty under weaker warning support. The interaction terms were not retained as significant in the model. Taken together, these process-level results indicate that takeover initiation timing was most strongly affected by alert urgency, while weather-related burden and scenario structure also contributed to variation in effective control entry. The descriptive results for the takeover-process indicators are summarized in
Table 5.
3.4. Results of Strict Same-Lane Safety-Margin Indicators
Under the strict same-lane criterion, the safety-margin indicators provided a more explicit downstream interpretation of takeover quality. Across all 480 trials, strict same-lane minTTC and minTHW were computable in 384 trials (80.0%). Computability differed markedly by scenario: lead-brake trials yielded computable strict same-lane conflicts in 239 of 240 trials (99.6%), whereas only 145 of 240 cut-in trials (60.4%) satisfied the strict same-lane criterion. Missingness under the strict definition was therefore driven primarily by conflict geometry rather than by generic data loss, with most missing cases attributable to the absence of a sustained same-lane relation. Consistent with the archived event-level audit, the dataset also contained one confirmed collision-terminated trial and one additional near-contact high-risk trial; these events were retained as valid safety-relevant observations, but the continuous strict same-lane indicators were still summarized according to their computability rules rather than recoded as a separate binary crash endpoint in the present inferential analysis.
Among the computable strict same-lane trials, the overall mean minimum TTC was 1.76 ± 0.54 s and the overall mean minimum THW was 1.49 ± 0.45 s. Alert urgency showed the clearest descriptive pattern. High urgency was associated with the largest safety margins (minTTC: M = 2.77 ± 1.04 s; minTHW: M = 2.51 ± 0.99 s), medium urgency showed intermediate values (minTTC: M = 1.94 ± 0.98 s; minTHW: M = 1.67 ± 0.77 s), and low urgency was associated with the smallest safety margins (minTTC: M = 0.95 ± 0.43 s; minTHW: M = 0.72 ± 0.39 s). Linear mixed-model analysis confirmed a significant fixed effect of alert urgency on both minTTC, χ
2 = 18.21,
p < 0.001, and minTHW, χ
2 = 35.16,
p < 0.001. Descriptively, the participant-level collapsed means indicate that the high-urgency condition exceeded the low-urgency condition by approximately 1.82 s in minTTC and 1.79 s in minTHW, reflecting a substantial increase in downstream temporal safety margin under stronger warning support. For minTHW, scenario condition also showed a significant fixed effect, χ
2 = 4.74,
p = 0.029, whereas the corresponding scenario effect for minTTC was not significant under the mixed-model framework. In interpretive terms, these strict same-lane indicators support a process-to-safety linkage in which lower urgency tended to coincide with tighter downstream temporal safety margins. The descriptive results for the strict same-lane safety-margin indicators are summarized in
Table 6.
3.5. Results of OpenDS 4.5-Derived Behavioral-Control Outcomes
Beyond the takeover-process and strict same-lane safety-margin indicators, the OpenDS 4.5-derived behavioral-control outcomes were retained to characterize how condition effects were behaviorally expressed during the takeover episode. The inferential results for the OpenDS 4.5-derived primary continuous outcomes under the linear mixed-model framework are summarized in
Table 7.
For takeover duration, linear mixed-model analysis revealed a significant fixed effect of scenario condition, χ2 = 21.51, p < 0.001. Descriptively, collapsed across weather condition, the mean takeover duration decreased from the low-urgency condition (M = 37.27 s, SD = 2.87) to the medium-urgency condition (M = 36.40 s, SD = 2.51) and further to the high-urgency condition (M = 35.73 s, SD = 2.55). Across scenario conditions, cut-in trials showed longer durations than lead-brake trials (M = 37.54 s, SD = 2.90 versus M = 35.39 s, SD = 2.33).
Descriptively, mean speed increased from the low-urgency condition (M = 35.47 km/h, SD = 2.45) to the medium-urgency condition (M = 36.55 km/h, SD = 2.31) and the high-urgency condition (M = 37.21 km/h, SD = 2.46). Across scenarios, lead-brake trials showed a higher mean speed than cut-in trials (M = 36.99 km/h, SD = 1.69 versus M = 35.84 km/h, SD = 2.90).
For braking response, linear mixed-model analysis showed significant fixed effects of alert urgency, χ2 = 7.82, p = 0.020, and scenario condition, χ2 = 23.46, p < 0.001. Descriptively, the maximum braking input was highest in the low-urgency condition (M = 0.362, SD = 0.320), followed by the medium-urgency condition (M = 0.156, SD = 0.231) and the high-urgency condition (M = 0.113, SD = 0.187). Across scenarios, cut-in trials produced larger braking responses than lead-brake trials (M = 0.325, SD = 0.313 versus M = 0.096, SD = 0.150).
For mean absolute steering input, linear mixed-model analysis showed a significant fixed effect of scenario, χ2 = 8.92, p = 0.003, together with a significant urgency-by-scenario interaction, χ2 = 7.30, p = 0.026. Descriptively, in the lead-brake scenario, the low-urgency condition produced a larger mean absolute steering input value (M = 0.0103) than the medium- and high-urgency conditions (M = 0.0087 and M = 0.0086). In contrast, in the cut-in scenario, the low-urgency condition produced a smaller mean absolute steering input value (M = 0.0066) than the medium- and high-urgency conditions (M = 0.0104 and M = 0.0103).
Steering-input variability likewise showed a significant fixed effect of scenario, χ2 = 15.18, p < 0.001, together with a significant urgency-by-scenario interaction, χ2 = 9.68, p = 0.008. In the cut-in scenario, the low-urgency condition was associated with lower steering-input variability (M = 0.0392) than the medium- and high-urgency conditions (M = 0.0506 and M = 0.0519), whereas the reverse ordering was observed in the lead-brake scenario.
For the minimum center-distance indicator, linear mixed-model analysis revealed a significant fixed effect of scenario condition, χ
2 = 35.98,
p < 0.001. Descriptively, the mean minimum center distance was 3.18 m (SD = 0.21) in the lead-brake scenario and 4.68 m (SD = 1.57) in the cut-in scenario. Within the cut-in scenario, the low-urgency condition showed the largest mean minimum center distance (M = 5.11 m), followed by the medium-urgency condition (M = 4.55 m) and the high-urgency condition (M = 4.37 m).
Figure 3 visualizes the main OpenDS 4.5-derived outcome patterns across alert urgency and scenario condition for the key takeover-related indicators.
As shown in
Figure 3, scenario condition accounted for the clearest variation in the broader OpenDS 4.5-derived behavioral-control outcomes, whereas urgency-related effects were more selective and weather condition did not show a comparably strong pattern in the selected OpenDS 4.5 measures.
3.6. Limited Eye-Tracking Results as Secondary Support
Among the retained variables, total glance duration showed the clearest condition-related pattern, whereas first glance latency was treated descriptively and the remaining measures were non-significant.
First glance latency did not show a stable pattern suitable for strong inferential interpretation in the final raw-data-based recheck. Descriptively, the mean first glance latency was 192.13 ms (SD = 795.28) under rainy conditions and 85.81 ms (SD = 165.06) under sunny conditions. Under the high-urgency condition, the corresponding means were 62.64 ms (SD = 182.16) in rain and 86.14 ms (SD = 262.14) in sun. Given its substantial dispersion and sensitivity to gaze continuity and event alignment, first glance latency was retained only as a descriptive secondary measure.
Total glance duration showed significant main effects of alert urgency, F(2, 78) = 9.297, p < 0.001, ηp2 = 0.193, and scenario condition, F(1, 39) = 42.556, p < 0.001, ηp2 = 0.522. The longest total glance duration was observed under low urgency and in cut-in trials. By contrast, glance count, mean glance duration, and AOI ratio did not show significant main or interaction effects in the current dataset.
3.7. Results of Questionnaire-Derived Supplementary Measures
The questionnaire results showed a clear pattern for alert-related retrospective evaluations. Because these repeated ratings were collected on ordinal Likert-type scales, Friedman tests were used for the three alert-related questionnaire items. For perceived urgency, the omnibus Friedman test showed a significant effect of alert urgency, χ2(2) = 79.00, p < 0.001, Kendall’s W = 0.988. Mean urgency ratings increased monotonically from low (M = 2.05, SD = 0.22) to medium (M = 3.45, SD = 0.55) and high urgency (M = 5.00, SD = 0.00). The same pattern was observed for perceived helpfulness, with a significant omnibus Friedman effect, χ2(2) = 74.44, p < 0.001, Kendall’s W = 0.931; high urgency received the highest helpfulness rating (M = 4.28, SD = 0.45), followed by medium (M = 3.90, SD = 0.30) and low urgency (M = 2.48, SD = 0.51). For subjective acceptability, the omnibus Friedman test also remained significant, χ2(2) = 69.87, p < 0.001, Kendall’s W = 0.873, but the rank ordering differed: medium urgency received the highest acceptability rating (M = 4.33, SD = 0.47), followed by low (M = 3.45, SD = 0.50) and high urgency (M = 2.98, SD = 0.16). Bonferroni-adjusted Wilcoxon signed-rank tests further showed that all three alert levels differed significantly from one another for perceived urgency, perceived helpfulness, and subjective acceptability (all adjusted p < 0.001).
The questionnaire-based descriptive and inferential results are summarized in
Table 8.
When participants compared rainy weather with sunny weather, they rated rainy conditions as more difficult for takeover completion (M = 4.25, SD = 0.44), more difficult for rapid recognition of key roadway information (M = 3.63, SD = 0.54), and more likely to increase takeover risk (M = 4.23, SD = 0.48). When participants compared the cut-in scenario with the lead-brake scenario, they rated cut-in as more difficult to handle (M = 4.20, SD = 0.52), more dangerous (M = 4.20, SD = 0.56), and more time-pressuring (M = 3.85, SD = 0.62). These retrospective ratings indicated that both rain and cut-in were consistently perceived as more demanding conditions than their corresponding comparison conditions.
The overall system-evaluation items showed generally positive but moderate subjective appraisal. The mean ratings were 3.70 (SD = 0.56) for overall system reliability, 3.88 (SD = 0.33) for the system’s ability to identify risk and remind drivers to take over, 3.68 (SD = 0.53) for overall trust, 3.83 (SD = 0.50) for perceived safety benefit, 3.68 (SD = 0.47) for willingness to use a similar system in one’s own vehicle, and 3.63 (SD = 0.54) for willingness to recommend a similar system to others. For overall burden, the mean rating was 2.98 (SD = 0.77), which was close to the scale midpoint, whereas the mean tension rating for the takeover task was 3.55 (SD = 0.68). These results indicated that the system was generally evaluated positively, while the overall task burden remained moderate and the takeover task was accompanied by noticeable subjective tension.
Figure 4 summarizes the main questionnaire patterns, including alert-related evaluations and the perceived difficulty associated with weather and scenario conditions.
The subjective patterns shown in
Figure 4 broadly support the objective findings, while also indicating that perceived helpfulness, acceptability, and environmental difficulty should not be treated as equivalent to objective takeover performance.
3.8. Summary of Main Findings
First, in the takeover-process indicators, condition effects were more evident in takeover initiation timing than in the mere occurrence of effective control. Response occurrence remained high overall, whereas RTcontrol showed clear condition-related variation. Analysis under an overall fixed-effect mixed-model framework indicated a strong urgency effect, together with significant weather and scenario effects, showing that lower urgency was associated with slower entry into effective manual control.
Second, in the strict same-lane safety-margin indicators, higher urgency was associated with larger minTTC and minTHW values, whereas lower urgency tended to coincide with tighter downstream temporal safety margins. Scenario-related differences were also evident, with lead-brake trials generally retaining larger strict same-lane safety margins than cut-in trials.
Third, in the broader OpenDS 4.5-derived behavioral-control outcomes, scenario condition remained the most stable determinant across takeover duration, steering-related behavior, and minimum center distance, whereas urgency-related effects were more selective and were particularly evident in braking response and in the scenario-dependent steering measures.
Fourth, only limited supplementary evidence was retained beyond the primary OpenDS 4.5 outcomes. Within this layer, total glance duration showed the only clear retained eye-tracking pattern, whereas the remaining eye-tracking measures were descriptive or non-significant. Questionnaire results, by contrast, provided clearer support for condition appraisal and user evaluation. In the questionnaire-based supplementary outcomes, the three alert levels were clearly differentiated by perceived urgency, helpfulness, and acceptability: high-urgency alerts were rated as the most urgent and the most helpful, whereas medium-urgency alerts received the highest acceptability ratings. Rainy weather and cut-in scenarios were also consistently rated as more demanding than their respective comparison conditions. Taken together, the retained evidence supports a staged interpretation in which condition manipulations first affect takeover initiation, then shape downstream safety margins and broader behavioral-control outcomes, and are finally reflected in visual-attention recovery and post-task appraisal.
4. Discussion
4.1. Overview of Key Findings
The present study examined how auditory alert urgency influenced takeover safety under different weather and scenario conditions within a unified 3 × 2 × 2 within-subject framework integrating primary OpenDS 4.5 outcomes and limited supplementary evidence. Across the matched dataset, the most stable effects were associated with scenario structure and selective warning-related differences rather than with weather condition alone. In the process layer, RTcontrol showed clear descriptive variation across alert-urgency levels, and the mixed-model results confirmed alert urgency as the strongest inferential determinant of takeover initiation timing; descriptively, low urgency was associated with a 2.40 s longer RTcontrol than high urgency in the participant-level collapsed means. In the safety-margin layer, higher urgency was associated with larger strict same-lane minTTC and minTHW values, a pattern that is consistent with an interpretation in which slower takeover initiation under lower urgency was accompanied by tighter downstream temporal safety margins. In the broader behavioral-control layer, scenario condition remained the most stable determinant across takeover duration, steering-related behavior, and minimum center distance, whereas urgency-related effects were more selective and were most evident in braking response and in the scenario-dependent steering measures.
Taken together, these findings support the central premise stated in the Introduction that takeover warning design should not be evaluated in isolation [
23]. Instead, its effect should be interpreted together with scenario complexity and environmental condition, while giving priority to behavior- and safety-based evidence [
24]. OpenDS 4.5-derived indicators captured the primary variation in takeover quality, whereas limited supplementary evidence was retained to support interpretation of attention recovery, condition appraisal, and system acceptance [
25]. This process-oriented interpretation is also consistent with integrated accounts of takeover time, which emphasize that takeover performance should be understood as the result of multiple interacting driver-, task-, and traffic-related factors rather than as a single isolated response measure [
26]. The overall interpretation of the retained findings is summarized in
Figure 5.
The present findings should be interpreted within the boundary of a controlled simulator-based experiment. The participant sample consisted exclusively of female drivers, and the eye-tracking measures were used as a supporting layer rather than as the primary evaluation basis. Accordingly, the main conclusions remain anchored in behavior- and safety-based outcomes.
4.2. Interpretation of Primary Results (OpenDS)
The OpenDS 4.5-derived primary results indicate that takeover performance was shaped by the combined influence of alert urgency, scenario condition, and, in some cases, weather-related burden, although the stability of these effects differed across specific outcome types [
27,
28]. In the takeover-process layer, RTcontrol showed clear descriptive variation across alert-urgency levels, and the overall-term mixed model confirmed alert urgency as the strongest inferential determinant of takeover initiation timing. Weather and scenario also contributed significantly, but the dominant pattern in RTcontrol remained the urgency gradient from low to high alert levels.
The braking results add further nuance to this interpretation. Maximum braking input was largest in the low-urgency condition, with substantially lower peak braking under medium- and high-urgency alerts. One plausible interpretation is that weaker warning urgency delayed effective takeover engagement, thereby increasing the need for abrupt longitudinal correction once drivers recognized the conflict [
29,
30]. Under more urgent alerts, drivers may have initiated control transition earlier and distributed the response across a longer portion of the event, reducing the need for a sharp brake peak. In this sense, the braking results are compatible with the temporal findings and suggest that higher alert urgency may support smoother longitudinal correction even when speed-related differences were not retained as significant fixed effects in the mixed-model framework.
The broader OpenDS 4.5-derived behavioral-control outcomes should therefore be read as behavioral manifestations of this process-to-safety chain rather than as stand-alone substitutes for it. Scenario condition exerted a particularly strong influence on the primary outcomes, which is consistent with the conflict structures embedded in the experiment. Compared with lead-brake trials, cut-in trials produced longer takeover duration, lower mean speed, larger braking response, and larger minimum center distance at the descriptive level. These differences indicate that cut-in scenarios imposed a more demanding control problem involving stronger spatial conflict and greater need for integrated longitudinal and lateral adaptation. The mixed-model results indicate that scenario remained the most stable inferential determinant across the broader behavioral-control outcomes, whereas urgency-related effects were more selective and were concentrated primarily in braking response and in the scenario-dependent steering measures.
The steering-related interaction pattern deserves particular attention. Mean absolute steering input and steering-input variability did not show a simple monotonic relationship with alert urgency across both scenarios. Instead, the direction of the urgency effect differed between lead-brake and cut-in conditions. This result suggests that lateral control response was not determined by warning urgency alone, but by the way in which urgency interacted with scenario-specific maneuvering demand. In a lead-brake event, the main requirement is often longitudinal correction with relatively limited lateral avoidance, whereas a cut-in event may require stronger or more variable steering adjustment to restore a safe path. The observed interaction is therefore consistent with a scenario-sensitive interpretation of warning effectiveness rather than a uniform “more urgent is always better” account.
The minimum center-distance result further supports this point. Scenario condition had the strongest influence on this safety indicator, indicating that conflict structure remained the primary determinant of lateral proximity patterns in the present dataset. Importantly, the larger minimum center distance observed in cut-in trials should not be interpreted in isolation as evidence of better safety. In the same scenario, drivers also showed longer takeover duration, stronger braking response, and longer total glance duration, all of which suggest higher takeover demand. The larger center distance in cut-in trials is more plausibly understood as part of a broader evasive-control pattern under spatially constrained conflict rather than as a direct indicator of reduced task difficulty. This is precisely why takeover safety in the present study was evaluated through a set of complementary primary measures rather than through a single metric.
By contrast, weather condition showed limited influence across the selected behavioral-control outcomes. This does not mean that weather was behaviorally irrelevant. Rather, it suggests that within the present simulator configuration, weather-related degradation did not produce effects as robust as those associated with scenario structure and selective warning-related differences in the main driving and safety outcomes. This pattern becomes important when interpreted together with the questionnaire data, because it indicates that weather may have operated more strongly at the level of perceived difficulty than at the level of the selected objective measures.
A plausible interpretation is that alert urgency, weather degradation, and scenario complexity do not operate at the same stage of the takeover process. Alert urgency is more likely to affect attentional capture and the perceived criticality of the takeover request; adverse weather increases the burden of visual-information reconstruction; and scenario structure changes the time pressure and conflict geometry under which control must be resumed. In this sense, the observed condition effects are better understood as process-level differences in reorientation and execution rather than as isolated shifts in a single outcome metric.
4.3. Interpretation of Secondary Results (Eye-Tracking)
The eye-tracking findings were retained only as a limited secondary evidence layer and therefore should not be interpreted with the same weight as the primary driving results [
31]. In the final presentation, only total glance duration provided a stable retained pattern, whereas first glance latency was reported descriptively because of its substantial dispersion and the remaining measures were non-significant.
The longer total glance duration observed under low urgency and in cut-in trials suggests that weaker warning salience and greater scenario demand were associated with greater sustained visual engagement with the task-relevant region during takeover. This pattern is broadly consistent with the primary results showing greater takeover demand in cut-in conditions and less favorable temporal characteristics under lower urgency.
At the same time, the largely non-significant pattern across the remaining eye-tracking variables indicates that the visual-attention evidence in the present dataset was selective rather than comprehensive. The eye-tracking component was therefore not intended to serve as an independent primary basis for evaluating takeover safety. Instead, the gaze-derived indicators were used only to provide limited process-level support for interpreting visual reorientation during takeover. The main conclusions of the study remain anchored in OpenDS 4.5-derived takeover timing, behavioral-control indicators, and safety-margin measures.
4.4. Interpretation of Supplementary Results (Questionnaire)
The questionnaire results provide an important subjective complement to the objective findings [
32]. The three alert levels were clearly differentiated by perceived urgency, helpfulness, and acceptability, showing that the warning manipulation was successful from the participant perspective. High-urgency alerts were rated as the most urgent and the most helpful, which is consistent with the primary results showing more favorable urgency-related patterns in braking response and in the strict same-lane temporal safety margins. At the same time, medium-urgency alerts received the highest acceptability ratings, indicating that the subjectively preferred warning was not the same as the warning perceived as most effective for immediate takeover support [
33]. This distinction is important for warning design because it highlights a practical tradeoff between response facilitation and long-term user acceptance. From an engineering perspective, the present findings do not support a single fixed warning setting as optimal across all conditions. Rather, they suggest that warning urgency should be calibrated adaptively according to situational demand [
34].
The weather- and scenario-comparison items also help clarify the relationship between subjective appraisal and objective performance. Rainy weather was consistently rated as more difficult, more visually challenging, and more likely to increase takeover risk, even though weather effects were not dominant in the main OpenDS 4.5-derived primary indicators [
35]. This mismatch is analytically important rather than problematic. It suggests that environmental degradation in the present experiment was strongly perceived by participants, but its behavioral manifestation in the selected objective measures was more limited or more indirect. This pattern supports the argument that subjective difficulty and objective safety response should be analyzed together rather than treated as interchangeable constructs [
36].
The scenario-comparison items showed stronger alignment with the primary results. Participants consistently rated cut-in scenarios as more difficult, more dangerous, and more time-pressuring than lead-brake scenarios. This subjective pattern corresponds well with the objective findings showing longer takeover duration, larger braking response, stronger steering-related interactions, and longer total glance duration in cut-in trials. In this case, the supplementary results reinforce the primary interpretation that cut-in scenarios imposed a more demanding takeover problem than lead-brake scenarios.
The overall system-evaluation items further indicate that the tested warning-supported takeover framework was viewed positively, but not without tension. Participants reported generally favorable evaluations of system reliability, warning effectiveness, trust, willingness to use a similar system, and willingness to recommend it, while also indicating moderate workload and noticeable task-related tension. This pattern is consistent with the experimental context. Participants did not reject the system, and they did not evaluate the warning-supported takeover process as excessively burdensome, but neither did they view it as effortless. Such a response profile is compatible with a realistic appraisal of takeover as a demanding but manageable human–machine transition [
37]. From an engineering perspective, this distinction is important because it indicates that warnings perceived as most effective for immediate takeover support need not coincide with warnings rated as most acceptable for repeated use, thereby reinforcing the need for context-sensitive rather than single-level warning design [
38,
39].
The interpretive framework developed in this study is not only useful for explaining condition-related differences in takeover behavior, but also for translating empirical findings into warning-policy design logic. Specifically, the takeover-process indicators identify when warning urgency influences entry into effective manual control, whereas the strict same-lane safety-margin indicators clarify whether these process-level differences are propagated into downstream temporal safety buffers. This layered structure enables an evidence-to-policy translation in which scenario demand and environmental burden can be used to adapt warning urgency rather than relying on a fixed one-level alert strategy.
4.5. Theoretical Implications
First, it addresses the gap identified in the Introduction by evaluating auditory warning urgency together with weather and scenario condition in a single experimental framework. Rather than treating warning design, environmental complexity, and scenario structure as separate research topics, the present study shows how these factors jointly shape takeover quality. The results suggest that warning urgency is not a fixed-effect interface feature whose consequences are independent of context. Its influence becomes more or less consequential depending on the conflict structure of the scenario and, to a lesser extent in the present dataset, on environmental condition.
Second, the study supports a hierarchical view of multimodal evidence in takeover research. The primary OpenDS 4.5 indicators captured the main variation in takeover quality and safety. Limited supplementary evidence provided selective support for interpreting condition-related differences beyond the primary OpenDS 4.5 outcomes. The questionnaire responses added a complementary layer of perceived urgency, difficulty, acceptance, trust, and workload. This structure is theoretically useful because it avoids collapsing behavior, attention, and subjective appraisal into a single undifferentiated outcome domain. Instead, it clarifies that these data types answer related but different questions about the takeover process.
Accordingly, the originality of the present work should be understood as an integrative experimental and interpretive contribution. The study does not claim that high-urgency alerts are universally optimal, nor does it propose a new takeover-control algorithm. Rather, it shows how urgency-related effects can be interpreted within the combined structure of warning salience, environmental burden, and scenario conflict, and how these effects can be organized into a bounded context-sensitive warning-evaluation framework.
4.6. Practical and Engineering Implications
The present findings provide preliminary implications for warning-configuration design in conditionally automated driving [
40]. The first implication is that warning urgency should be calibrated with explicit attention to scenario demand. Because scenario condition remained the most stable determinant across the primary OpenDS 4.5-derived outcomes, while urgency-related effects remained especially relevant for braking response and strict same-lane temporal safety margins, the results suggest that warning strategies should be evaluated under conflict-specific conditions rather than through context-free perceptual testing alone. In particular, cut-in events appear to be more sensitive to warning design because they impose stronger combined spatial and temporal demand. This is precisely why the present study is relevant to ITS-oriented design: the contribution lies not only in identifying condition effects, but in showing how those effects can inform adaptive warning allocation under different operational demands.
A related implication is that warning design should not be optimized on urgency alone. High-urgency alerts were rated as most urgent and most helpful, but medium-urgency alerts were rated as most acceptable [
41]. From an engineering perspective, this pattern suggests that no single static urgency level is likely to be optimal across all operating conditions. A more practical strategy would be to adopt a context-sensitive, adaptive, and multi-tier urgency logic rather than relying on a single static warning level. The present results therefore support adaptive alert tuning rather than a one-level-fits-all warning design.
To make this design implication more operational,
Figure 6 reformulates the present findings as a proof-of-concept warning-configuration logic linking condition recognition, configuration selection, and HMI presentation. Rather than serving only as an interpretive extension of the empirical findings, this engineering-oriented view shows how empirical results can be translated into a structured adaptation pathway for context-sensitive takeover support.
Table 9 further translates the current evidence into a proof-of-concept adaptive auditory-warning policy across the four scenario-and-weather combinations examined in the present study. The table is not intended to prescribe fixed production parameters; rather, it shows how the current 40-participant evidence base can be converted into condition-specific urgency recommendations, supporting rationale, and implementation-oriented notes.
The current evidence supports a bounded escalation logic in which medium urgency serves as the default option under lower-demand conditions, whereas higher urgency may be preferred when scenario conflict and environmental burden jointly increase takeover demand. This translation should be understood as a proof-of-concept policy layer derived from simulator evidence, not as a calibrated production-ready HMI specification.
From a scenario-risk assessment and future system-design perspective, the value of this matrix is not limited to post hoc interpretation [
42]. It provides a proof-of-concept policy layer that can be embedded into a tiered warning strategy or context-aware warning management logic. In such a design, scenario classification and environmental burden assessment would serve as upstream inputs, urgency selection would function as the intermediate policy decision, and warning presentation would constitute the downstream HMI output. It shows how empirical takeover evidence can be translated into a structured adaptation rule for future warning-support systems, while also indicating the need for subsequent closed-loop or deployment-oriented validation before such a policy can be treated as an operational control rule.
4.7. Limitations
First, the study was conducted in a controlled OpenDS 4.5 simulation environment rather than in an on-road setting. Although this setup allowed consistent manipulation of alert urgency, weather condition, and scenario structure, the findings should be interpreted within the bounds of simulator-based takeover behavior. In addition, participants used a keyboard-and-mouse interface rather than a wheel-and-pedal cockpit. This limits ecological validity because the interface cannot reproduce the continuous haptic feedback, steering-wheel dynamics, pedal travel, or physical workload associated with real-vehicle control. Therefore, steering- and braking-related variables in this study should be interpreted as normalized behavioral-control inputs within a controlled simulator environment, not as direct proxies for real-vehicle control actions. The OpenDS 4.5 scenarios were also implemented using fixed scenario templates and predefined hazard kinematics. This design improved experimental control and comparability across participants, but it did not reproduce the full variability of natural traffic flow, surrounding-driver interaction, or real-road weather degradation.
Second, the participant sample consisted exclusively of female drivers. This design choice was made to obtain a more homogeneous dataset and to strengthen the internal consistency of the present analysis for a driver group that remains comparatively underrepresented in the takeover literature. At the same time, this participant-selection strategy limits direct generalization of the current findings to male drivers, mixed-gender driver populations, and broader driver groups with different behavioral-response patterns. In addition, the participant group was relatively homogeneous in age, with most participants falling within the 25–35-year range. This further limits the extent to which the findings can be generalized to older drivers, novice drivers, or broader driver populations with different perceptual, cognitive, auditory, and motor-response characteristics. Accordingly, the policy implications derived from
Figure 6 and
Table 9 should be interpreted as a proof-of-concept evidence-to-policy framework rather than as calibrated production parameters for direct deployment across the full driving population.
Third, although the 12 takeover tasks were not administered in a single fixed sequence and participant-specific ordering was used to reduce simple sequence-related bias, residual learning, habituation, or fatigue effects associated with repeated takeover exposure cannot be entirely ruled out in a within-subject design [
43].
Fourth, the questionnaire was administered after the completion of all 12 tasks rather than after each individual trial. As a result, the supplementary subjective measures capture retrospective condition appraisal and overall system impressions rather than trial-level subjective variation.
Fifth, the eye-tracking results were selective rather than comprehensive. Only total glance duration showed a clear retained condition-related pattern, whereas several other gaze-derived measures were non-significant or descriptively unstable. This limits the explanatory scope of the eye-tracking layer. Therefore, the D-Lab-derived indicators were retained only as secondary process-level support and were not used to support the main safety conclusions independently.
Finally, weather effects were more evident in subjective appraisal than in the main objective indicators. This pattern is informative, but it also suggests that the behavioral consequences of environmental degradation may depend on the specific metric set and simulator implementation used in the present experiment.
In addition, the strict same-lane safety-margin indicators should be interpreted as surrogate conflict measures within the present simulator-based geometry and rule set rather than as direct substitutes for real-world crash risk estimates.
4.8. Future Work
Future work can extend the present study in several directions. A first priority is to examine whether the current findings remain stable in broader participant groups and under more diverse driver characteristics. This would help determine which aspects of the present warning- and scenario-related effects are sample-specific and which are more general. A second direction is to test similar warning manipulations under expanded scenario sets and more varied environmental conditions in order to determine whether the dominance of scenario effects over weather effects persists across a wider operational design space.
A third direction is to refine multimodal interpretation by incorporating additional trial-level supporting measures while preserving the current hierarchy in which driving and safety indicators remain primary. This extension is also consistent with recent Applied Sciences studies showing that takeover-request responses can vary with signal modality, signal attributes, and individual background factors [
44], and that takeover-support design may further depend on driver in situ state, feedforward timing, and information modality [
45]. Trial-level subjective measurements, physiological measures, or extended visual metrics may help explain residual variance in takeover behavior, provided that they do not displace the central role of safety-related driving outcomes. A fourth direction is to translate the present findings into adaptive warning design strategies that tune alert urgency according to scenario demand and environmental difficulty. Such work would provide a more direct bridge between simulator-based experimental evidence and practical takeover-support implementation in intelligent transportation systems.
5. Conclusions
This study investigated the effects of auditory alert urgency on takeover safety under different weather and scenario conditions in conditionally automated driving. Using a 3 × 2 × 2 within-subject design, 40 female licensed drivers completed 12 matched takeover tasks in OpenDS 4.5 under three alert-urgency levels, two weather conditions, and two representative takeover scenarios. By integrating takeover-process indicators, strict same-lane safety-margin indicators, broader OpenDS 4.5-derived behavioral-control outcomes, D-Lab-derived eye-tracking measures, and questionnaire-based subjective evaluations within a unified analytical framework, the study examined how warning urgency operated together with environmental difficulty and conflict structure rather than in isolation.
The results showed that the most stable condition effects were associated with scenario structure and selective warning-related differences rather than with weather condition alone. In the takeover-process layer, RTcontrol showed clear descriptive variation across alert-urgency levels, and the mixed-model results confirmed alert urgency as the strongest inferential determinant of takeover initiation timing, with weather and scenario also contributing significantly to variation in effective control entry. In the strict same-lane safety-margin layer, higher urgency was associated with larger minTTC and minTHW values, whereas lower urgency tended to coincide with tighter downstream temporal safety margins. In the broader behavioral-control layer, scenario condition remained the most stable determinant across takeover duration, steering-related behavior, and minimum center distance, whereas urgency-related effects were more selective and were most evident in braking response and in the scenario-dependent steering measures. In the secondary and supplementary layers, limited eye-tracking evidence showed only a narrow retained pattern in total glance duration, while questionnaire responses showed that high-urgency alerts were perceived as the most urgent and the most helpful, whereas medium-urgency alerts received the highest acceptability ratings.
Taken together, these findings support a context-sensitive interpretation of auditory takeover warnings. Warning urgency should not be evaluated as an isolated interface property, but in relation to scenario complexity and environmental demand. Under more demanding conditions, especially when conflict structure is tighter or perceptual burden is higher, higher urgency may better support takeover timeliness and safety-margin preservation. Under less demanding conditions, medium urgency may provide a more balanced compromise between response support and long-term acceptability. These recommendations should be interpreted as evidence-based design guidance derived from a controlled simulator context rather than as directly calibrated production HMI specifications. More broadly, the present study suggests that takeover-warning evaluation should move beyond context-free urgency testing toward evidence-guided warning policies that explicitly account for scenario demand and environmental burden. In this sense, the present findings offer not only a human-factors explanation of takeover behavior, but also a proof-of-concept basis for context-aware warning-control logic in future intelligent transportation and automated-driving support systems.
Several limitations should be noted. The findings were derived from a controlled simulator setting with a keyboard-and-mouse interface and fixed OpenDS 4.5 scenario templates, and the safety-margin indicators should be interpreted as simulator-based surrogate measures rather than direct estimates of real-world crash risk. The participant sample was limited to female drivers and was relatively concentrated in the 25–35-year age range. Therefore, the findings should not be directly generalized to mixed-gender populations, older drivers, or more heterogeneous driving groups without further validation. Future research should examine whether the present warning- and scenario-related patterns remain stable across broader participant groups, more immersive wheel-and-pedal simulator settings, richer traffic interactions, and extended multimodal measurement configurations.