Next Article in Journal
Numerical Investigation of Creasing Instability in Compression Packer Rubber Cylinders: Effects of Geometry, Friction, and Meshing Strategy
Previous Article in Journal
Early Prediction of Epileptic Seizures Based on Multifractal Analysis and Optimized Graph Neural Networks Using Scalp EEG Data
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Information Sources and Incremental Value in Short-Horizon Prediction of a Multimodal Driving Index in Extra-Long Tunnels

1
CCCC Mechanical and Electrical Engineering Bureau Limited Company, Beijing 101300, China
2
School of Transportation, Shijiazhuang Tiedao University, Shijiazhuang 050043, China
*
Authors to whom correspondence should be addressed.
Appl. Sci. 2026, 16(18), 8998; https://doi.org/10.3390/app16188998
Submission received: 8 August 2026 / Revised: 5 September 2026 / Accepted: 8 September 2026 / Published: 10 September 2026
(This article belongs to the Section Transportation and Future Mobility)

Abstract

Predicting driver-state evolution in extra-long tunnel corridors remains challenging because of prolonged spatial confinement and repeated lighting transitions. This study uses a statistical human–vehicle composite, the comprehensive driving index (CDI), as a reproducible quantitative target for predictive auditing. In this study, multimodal information denotes synchronized ocular, physiological, vehicle-motion, and environmental sensor signals; the objective is to quantify their incremental predictive value rather than introduce a new fusion architecture. Fully nested leave-one-driver-out cross-validation with a prespecified 120 s unsupervised initialization estimated all preprocessing, scaling, PCA, model-selection, and calibration parameters from training data only. The five components explained 60.36% of target variance. In the original-range 30 s task (4835 evaluation windows), history-only ridge regression achieved an RMSE of 0.08294 and an R2 of 0.166, while directly tuned AR achieved an RMSE of 0.08261 and an R2 of 0.168. On the common 4259-window sample, expanded ridge and AR achieved RMSEs of 0.08123 (R2 0.187) and 0.08153 (R2 0.182). Adding coarse scene information produced an ΔRMSE = +0.00002 (95% CI −0.00026 to 0.00027), whereas external environmental summaries produced an ΔRMSE = −0.00019 (95% CI −0.00037 to −0.00003). HistGradientBoosting did not improve performance. The primary contribution is a leakage-controlled predictive-audit framework for screening candidate information sources before deployment decisions.

1. Introduction

Extra-long tunnels on mountainous highways impose abrupt brightness transitions, spatial closure, and repeated route changes that challenge short-horizon prediction of driver-state evolution. Prior studies have documented visual-load differences, fatigue trajectories, physiological responses, and attention changes in tunnel environments [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16]. Whether onboard multimodal signals improve prediction for an unseen driver beyond recent driver-specific history remains unresolved. The 36.2 km Tianshan Shengli corridor therefore provides a demanding test case for evaluating the incremental value of candidate information sources.
The relevant evidence base spans tunnel murals and EEG-related load [1,2], entrance/exit attention and speed-adaptation effects [3,4,5], visual-load, fatigue, and light-environment responses [6,7,8], and bridge–tunnel workload and sign-information effects [9,10]. Studies of mental workload, ocular and physiological signals, multimodal distraction, and drowsiness prediction extend this background [11,12,13,14,15,16,17,18,19,20,21,22,23,24].
Driving in these environments induces physiological and vehicle-motion responses that provide candidate signals for multimodal prediction. Existing studies have advanced detection, classification, and behavior modeling [17,18,19,20,21,22,23,24], but the incremental value of these signals for cross-driver forecasting beyond strong historical baselines remains to be quantified.
The credibility of predictive studies depends critically on validation design. Methodological studies have distinguished explanatory from predictive objectives [21] and warned against optimism bias in joint model selection and error estimation [25]. They have also recommended group-based or time-dependent cross-validation [26], clarified the validity conditions for autoregressive time-series evaluation [27], emphasized scenario-consistent validation [28], and highlighted the risk of data leakage during preprocessing or feature construction [29]. We therefore combine strong historical baselines with fully nested leave-one-driver-out cross-validation. Pairwise driver-level comparisons and prediction-interval calibration then separate target continuity from the out-of-fold contributions of route, scene, and environmental information.
This article does not propose a new fusion architecture. Instead, it asks whether synchronized multimodal information adds reproducible value beyond optimized historical baselines. We ask two questions: (1) how much reproducible cross-driver, short-horizon predictability remains for the five-component human–vehicle CDI beyond persistence, historical-mean, and autoregressive baselines, and (2) whether the currently collected coarse-scene encodings or external CO2 and illuminance summaries reduce prediction error for held-out drivers after the history-only CDI model is optimized. The CDI comprises 15 ocular, physiological, and vehicle-motion variables; CO2 and illuminance are used only as external predictors. Five components define the primary target, and seven components are retained for sensitivity analysis. All preprocessing, PCA, scaling, model selection, and interval calibration were estimated within nested training folds.

2. Materials and Methods

2.1. Study Corridor and Analytical Design

This corridor-specific methodological audit addressed target construction, strong-baseline comparison, information isolation, expanded hyperparameter searches, and scenario robustness. The 36.2 km corridor included 12 drivers, 24 directional trips, and synchronized ocular, physiological, vehicle-motion, and environmental observations. Predictive performance was evaluated for held-out drivers. The CDI is a statistical composite built from 15 non-environmental variables, while CO2 and illuminance were retained as external predictors. Scenario analyses tested whether low-resolution route proxies provided reproducible predictive increments.

2.2. Experimental Protocol

2.2.1. Road Segment

The field test covered the 36.2 km section from Haxionggou Tunnel to Tianshan Shengli Tunnel on the Urumqi–Yuli Expressway (Figure 1). The continuous tunnels preceding the main tunnel were coded as the tunnel group (TG), and the Tianshan Shengli main tunnel was coded as TS. These labels combine differences in the number of portals, brightness continuity, road alignment, altitude profile, and travel sequence; they are therefore coarse composite route proxies rather than isolated physical attributes. The TG/TS comparison was used to audit whether route labels added independent predictive information. A null increment for these labels cannot be extrapolated to continuous geometry, portal transitions, traffic states, or higher-resolution scene measurements.

2.2.2. Instrumentation

The test platform consisted of a test vehicle, sensing system, recording equipment, and synchronized output (Figure 2). The vehicle was a Great Wall pickup truck. Sensing devices included Tobii Glasses 3 eye trackers (manufactured by Tobii AB (publ) and headquartered in Danderyd, Sweden), ErgoLAB wearable physiological recorders (produced by Beijing Jinfa Technology Co., Ltd., Beijing, China), onboard diagnostic units (Shenzhen Daotong Technology Co., Ltd., Shenzhen, China), and light meters (Jiangsu Ankerui Electric Appliance, Jiangyin, China). Eye-movement data were sampled at 50 or 100 Hz, and scene video at 25 frames/s. Physiological, vehicle-motion, and environmental signals were synchronized through the ErgoLAB multichannel module (Beijing Jinfa Technology Co., Ltd., Beijing, China). A forward-facing recorder mounted inside the windshield captured road and traffic conditions for subsequent blinded driving-performance annotation.

2.2.3. Experimental Procedure

A 1–2-day pre-test was carried out in Shijiazhuang to familiarize the research team with the test plan, calibrate the equipment, and standardize the recording process. Because the altitude, climate, and road environment differed between the pre-test and field-test sites, field-test data quality was assessed using simultaneous recordings, anomaly screening, and missingness statistics rather than the pre-test thresholds.
The field experiment was conducted over four days in Urumqi County, Xinjiang, during an approximately six-day field campaign. Each driver completed two to four test runs per day. A test included one outbound and one return trip and lasted approximately one hour. After each test, eye-movement, physiological, vehicle-motion, trajectory, and environmental records were exported, backed up, and checked for completeness, continuity, validity, anomalies, and missing observations.

2.2.4. Participants

Twelve licensed drivers participated (9 males and 3 females; age 23–44 years, mean 35.6 years). Visual acuity was recorded using the Chinese Standard Logarithmic Visual Acuity Chart with five-point (Miao) notation. On this scale, 5.0 corresponds to decimal acuity 1.0 (approximately Snellen 20/20), 4.6 corresponds to approximately decimal 0.4 (approximately Snellen 20/50), and higher values indicate better acuity. The available individual records do not distinguish unaided from corrected measurements; the study eligibility materials specified corrected visual acuity of at least 4.6. The participant records showed five drivers with 5.0, six with 4.9, and one with 4.6; no driver had 4.8. Seven reported no previous traffic accidents, four reported one, and one reported two. The pre-study questionnaire excluded severe mental or cardiovascular conditions that could interfere with driving. Written informed consent was obtained from every participant. With 12 independent drivers, the results should not be interpreted as population-level estimates for all drivers or tunnel environments.

2.3. Synchronization, Cleaning, and Scaling

Eye-movement, physiological, and vehicle-motion variables were aligned to a unified 1 s timestamp. Data completeness Q was defined as the proportion of the 15 non-environmental CDI variables observed at each timestamp. CO2 and illuminance were retained as external variables and were not counted in the CDI target. Before imputation, missing cases, consecutive gaps, and differences between directions and TG/TS labels were summarized. Q was used only for data-quality description, quality adjustment, and prespecified sensitivity analyses; it was not used to construct the primary CDI. Before imputation, 11,440 of 51,741 synchronized 1 s observations (22.11%) had a Q < 0.5, and 14,240 (27.52%) had a Q < 0.7 under this 15-variable definition.
Within each driver trip, gaps no longer than 5 s were filled forward using the most recent observation. Remaining missing values were replaced by the median estimated from the current training set. For the 15 CDI variables, the first 120 s in each direction were used to establish a prespecified, unsupervised driver-specific amplitude baseline, and this calibration period was excluded from error assessment. Imputation, population scaling, PCA loadings, CDI weights, scaling bounds, and model parameters were estimated from the corresponding training data only; CO2 and illuminance were standardized as external predictors within the same folds.
The PCA composite index is termed the CDI, the model inputs are termed predictor variables, and each held-out driver’s outer-loop error is termed the out-of-fold error. Forecast samples were evaluated every 10 s, and all predictors were restricted to information available at time t. The CDI is a statistical summary of joint variation in ocular, physiological, and vehicle-motion observations. It is not prespecified as a measure of workload, fatigue, safety, clinical status, or operational risk. Because no independent lane-position, TTC, safety-event, or blinded video labels were analyzed, construct validity and relationships with external outcomes remain open questions for future study. Accordingly, references to the CDI throughout the Results, Discussion, and Conclusions denote this statistical index only and do not constitute direct observations of driver state or safety.

2.4. Construction of the Comprehensive Driving Index

The number of candidate components was examined using time-preserving parallel analysis based on within-driver cyclic displacements, and loading reproducibility was assessed with 1000 driver-level bootstrap samples and Tucker congruence coefficients. Parallel analysis formally retained PCs 1–7 (with an isolated higher-order anomaly), but eigenvalue separation and resampling stability declined after PC5. We therefore used five components for the primary CDI based on parsimony and resampling stability; this choice does not imply that five is uniquely correct. Six- and seven-component targets were analyzed as dimensionality sensitivities. Each sensitivity repeated imputation, scaling, PCA, CDI weighting and scaling, model tuning, and held-out-driver evaluation within the same information barrier.
The five-, six-, and seven-component CDIs use different loadings, weights, and scaling boundaries. Their absolute values, RMSEs, and regression coefficients are therefore not directly compared. Sensitivity analyses compare model ranking, incremental direction, and qualitative consistency under each target definition.
When driver k was held-out in the outer loop, superscript (−k) indicates that the relevant imputation, standardization, PCA, CDI-scaling, and prediction parameters were estimated only from the remaining drivers. Training-standardized values, component scores, and the CDI were calculated as follows:
Z i j ( k ) = X i j i m p M j ( k ) S j ( k ) ;
P C i m ( k ) = j Z i j ( k ) v j m ( k ) ;
C D I i ( k ) = M i n M a x t r a i n m w m ( k ) P C i m ( k ) .
For every outer- or inner-loop validation driver, the first 120 s in the corresponding direction were used for the prespecified unsupervised driver-specific baseline transformation. Population scaling parameters Mj/Sj, PCA loadings, component weights, and CDI scaling bounds estimated from the corresponding training set were then applied without modification. CO2 and illuminance were never used in the target construction and were standardized only as external predictors within the relevant training fold.
Information sets were decomposed into CDI history (CDI(t), trailing mean, and ordinary-least-squares trend), coarse scene information (direction, TG/TS, direction × TG/TS, segment duration, and Q), and external environmental information (trailing means and OLS trends for CO2 and illuminance). Variables used to construct the CDI were not entered as primary predictors because of the part–whole relationship; they were examined only in a target-source sensitivity analysis. Q was used for quality adjustment and quality-threshold sensitivity, not as a component of the CDI.

2.5. Short-Horizon Prediction, Strong Baselines, and Nested Validation

2.5.1. Baselines and Information Sources

Three simple baselines established the minimum level of temporal predictability. Persistence used CDI(t) to predict CDI(t + h), the historical-mean baseline used the recent window mean, and the training-mean baseline used the outer training-set mean. The main linear models were history-only ridge regression, ridge regression with the coarse-scene set, and ridge regression with external environmental summaries. A directly tuned multi-step AR(p) model served as a classical time-series benchmark. HistGradientBoostingRegressor models were tuned in the same nested driver folds to challenge the linear specification. The primary information audit therefore compares historical continuity, coarse-scene proxies, external environmental predictors, and a nonlinear benchmark under the same target definition. Failure of the evaluated ridge or gradient-boosting models to extract stable incremental value is not interpreted as evidence that multimodal or environmental information is generally non-predictive (Figure 3). All statistical analyses and modeling were performed using Python (version 3.10.9) with scikit-learn (version 1.2.1), pandas (version 1.5.3), and NumPy (version 1.24.1). Figures were generated using Matplotlib (version 3.7.1).

2.5.2. Input Windows, Forecast Horizons, and Feature Engineering

One driver was reserved for each outer loop, and the remaining 11 were used for training. The inner loop used three folds grouped by drivers, with approximately 7–8 drivers per fold for training and 3–4 for validation, to select the input window, ridge penalty, AR order, and nonlinear-model settings. The primary ridge window grid was {30, 60, 90, 120, 180, 240} s, and AR orders were {6, 12, 18, 24, 36}; the original-range comparison was retained for the common 30 s task. Imputation, scaling, PCA, CDI scaling, predictor scaling, and all model parameters were re-estimated within each inner training subset. These grids were used to audit temporal memory under the present data and were not defined as candidate deployment cache lengths or update cycles.
The first 120 s in each direction constituted unsupervised personalization at test time and used no future CDI values, prediction errors, scenario endings, or tuning feedback. The primary results therefore evaluate cross-driver prediction after 120 s of initialization rather than a zero-sample cold start. Initialization sensitivity was additionally examined at 30, 60, 120, and 180 s on a common 4547-window set. Because changing initialization changes the driver-specific target calibration, absolute RMSEs across these settings are not directly comparable; the 180 s result was treated as target-definition-dependent rather than as evidence of a universally optimal initialization length.
Ridge models selected the window and α from the expanded nested grids (windows {30, 60, 90, 120, 180, 240} s; α ∈ {0.1, 1, 3, 10, 30}). Direct AR(p) selected p from {6, 12, 18, 24, 36}, and HistGradientBoostingRegressor settings were selected within the same grouped folds. After selection, preprocessing, the CDI, and model parameters were re-estimated from the complete outer training set, and the held-out driver was evaluated once. The trailing mean, OLS trend, and h-step target were defined as follows:
x ¯ t , w = 1 w i = 0 w 1 x t i ;
s t , w = i = 0 w 1 ( i ı ¯ ) ( x t i x ¯ t , w ) i = 0 w 1 ( i ı ¯ ) 2 ;
y t ( h ) = C D I t + h .
The input window contains observations from t − w + 1 to t, and the target is CDI(t + h); OLS trends and all other predictors use no future observations. The main forecast horizon remained 30 s. The expanded search used windows of 30, 60, 90, 120, 180, and 240 s and AR orders of 6, 12, 18, 24, and 36. For fair comparison of extended settings, performance was summarized on the common set of 4259 evaluation windows available to all extended candidates. Even after expansion, a selected boundary or near-boundary value is reported as a within-grid choice, not as a globally optimal or operational setting.
The incremental value per driver was defined as the RMSE(history-only CDI) − RMSE(additional-information model), so positive values represent error reduction. The primary comparison was the driver-by-driver difference between the coarse scene model and the history-only model; environmental and nonlinear comparisons were secondary audits. For each model contrast, the 12 paired driver-level RMSE differences were sampled with replacement 20,000 times, and the 2.5th and 97.5th percentiles of the bootstrap mean formed the descriptive 95% interval. These were ordinary percentile, not BCa, intervals. Secondary information-source comparisons were exploratory and were not corrected for multiple testing; no isolated comparison, including Q, was interpreted as a confirmatory significant gain.
Q-threshold sensitivity analyses used the full sample, Q ≥ 0.5, and Q ≥ 0.7, constraining both the mean Q in the input window and Q at the target time. The window and α were reselected for every threshold. Among the 4835 main 30 s evaluation windows, Q ≥ 0.5 retained 3729 and excluded 1106 (22.87%); Q ≥ 0.7 retained 3379 and excluded 1456 (30.11%). These analyses were interpreted as data-quality robustness checks rather than as independent evidence for a CDI construct or a safety effect.

2.5.3. Model Estimation and Nested Tuning

All linear information sets used standardized ridge regression with the same search space. Direct AR(p) used only discrete CDI lag terms, and its PCA/CDI construction and order selection were repeated within each inner training subset. HistGradientBoostingRegressor was tuned as a nonlinear challenge model, not as a post hoc replacement selected after seeing the outer test errors. Prediction-interval calibration was retained as a secondary uncertainty benchmark; point-prediction and incremental comparisons were the primary endpoints of this revision.

2.5.4. Incremental and Driver-Level Evaluation

Continuous forecast results are summarized via the R2, mean absolute error (MAE), and RMSE. Driver-by-driver RMSE is the primary endpoint for incremental comparisons. We also report the coverage and average width of the 90% prediction intervals. All uncertainty statements and incremental comparisons are interpreted at the driver level; the dense windows are repeated observations rather than independent inferential units.
The unit of inference was the driver rather than the dense time window: the study contains 12 major independent driver clusters. Time points improve trajectory resolution but cannot replace independent driver samples. Accordingly, bootstrap intervals and driver-by-driver changes are interpreted as uncertainty summaries, not as proof of driver subgroups or individualized warning thresholds. Generalization is therefore restricted to this small driver sample and study corridor; broader population validity requires independent drivers, routes, and external replication.

2.6. Scene Analysis and Serial-Correlation Adjustment

For descriptive interpretation of coarse-grained route labels, the five-component CDI was aggregated into 10 s observations and analyzed with a hierarchical linear mixed-effects model. The fixed effects were TS, elapsed time within the road segment, direction, all second-order interactions, the TS × elapsed-time × direction interaction, Q, and standardized external CO2 and illuminance. This full hierarchical interaction structure was prespecified as a saturated descriptive model; it was not selected by QIC, stepwise testing, or outer prediction error, and all lower-order terms were retained according to the hierarchy principle. The random structure included a driver intercept, a driver-specific direction slope, and a trip-level variance component. The model describes contemporaneous associations in 4931 repeated observations from 12 drivers and 24 directional trips; out-of-fold prediction remains the sole analysis of predictive increments.
The hierarchical mixed-effects model and out-of-fold incremental analysis have different objectives. The mixed model estimates descriptive conditional associations of coarse scene labels and external environmental signals with contemporaneous CDI distributions while accounting for driver and trip clustering. The nested out-of-fold analysis tests whether information available before prediction improves the error for unseen drivers. Neither analysis is causal, and the mixed-model coefficients should not be interpreted as independent validation of the CDI or as evidence of safety effects.

3. Results

3.1. Data Overview and Index Construction

Time-preserving parallel analysis formally retained PCs 1–7, while higher-order eigenvalue separation and loading stability weakened after PC5. The five principal components explained 60.36% of the variance (Table 1; Figure 4). Tucker medians were 0.920, 0.875, 0.810, 0.658, and 0.531 for PC1–PC5, respectively. The five-component target was retained as a parsimonious statistical projection.
These loadings and explained variances provide a statistical target baseline for the study corridor. CO2 and illuminance were excluded from target construction so their predictive contribution could be audited as external information. The CDI is a statistical composite of the measured human–vehicle signals; validation against independent driving-performance outcomes is reserved for subsequent studies.
PC1 was mainly determined by fixation proportion, gaze-loss proportion, and fixation angular velocity; PC2 by RR interval, instantaneous heart rate, and mean heart rate; and PC3 by angular-velocity magnitude, acceleration magnitude, and SDNN. PC4 was mainly determined by RMSSD, physiological relaxation, and pupil diameter, whereas PC5 was mainly determined by respiration, pupil-area change, and pupil diameter. Loading signs indicate relative statistical directions only. Although CO2 and illuminance were excluded from the CDI target, the lower Tucker congruence of PC4 (0.658) and PC5 (0.531) indicates limited cross-driver reproducibility of higher-order physiological/ocular structure. This instability may add target noise and attenuate small incremental gains; therefore, the external-information null results cannot be interpreted as evidence that environmental information is intrinsically uninformative.
The six- and seven-component sensitivity targets used their own nested loadings and scaling rules. They were used to check whether the ranking of history, scene, environmental, and nonlinear models was qualitatively preserved, not to compare absolute RMSEs across differently calibrated targets.
The five-, six-, and seven-component sensitivity analyses compared model ordering, incremental direction, and qualitative consistency, but did not establish loading equivalence or external construct validity. In particular, vehicle-motion variables included in the CDI cannot serve as independent validation outcomes. Future work should match component symbols and orders, report reconstruction error, and evaluate independent blinded video, lane-position, TTC, or safety-event outcomes with target construction that avoids part–whole overlap. No result in this section establishes a safety threshold, operational risk category, or direct driver-state measurement.

3.2. Direction Sensitivity and Serial-Correlation Adjustment

Before examining scene associations, we describe the four direction × scene groups. The scene analysis contained 4931 10 s aggregate observations from 12 drivers and 24 directional trips. The mean CDI values of outbound TG, outbound TS, return TG, and return TS were 0.3713, 0.3627, 0.3292, and 0.3424, respectively. The overall TG and TS means were 0.3522 and 0.3517, indicating strong overlap and motivating a hierarchical rather than a simple pooled comparison (Figure 5).
In the hierarchical mixed-effects model, the TS coefficient was β = 0.02759 (p < 0.001), the elapsed segment time was β = −0.00137 (p < 0.001), and the outbound direction was β = 0.03856 (p = 0.017). The TS × direction interaction was β = −0.02228 (p = 0.004), and time × direction was β = 0.00205 (p < 0.001); the TS × time and three-way interaction were not significant. The external CO2 association was β = −0.00374 (p = 0.016), whereas illuminance was β = 0.00198 (p = 0.062) and Q was not significant (p = 0.125). These are descriptive contemporaneous associations after accounting for driver and trip clustering and do not imply causal or predictive effects (Table 2). The Q association was exploratory and did not support the earlier isolated increment interpretation; no multiplicity-adjusted or confirmatory Q effect is claimed.

3.3. Strong Historical Baselines and Incremental Prediction

Primary 30 s Task

Table 3 reports the simple and tuned historical baselines for the original-range 30 s task across 4835 evaluation windows. History-only ridge and directly tuned AR(p) improved on the historical-mean, persistence, and outer-training-mean baselines. The paired AR-minus-ridge RMSE difference was 0.00034 (95% driver-bootstrap CI −0.00090 to 0.00177), indicating practically similar performance for the two tuned history models under the current target calibration.
The history-only model remained the main comparison reference. The scene ridge model achieved an RMSE = 0.08292 and an R2 = 0.165, while the environment-augmented ridge model achieved an RMSE = 0.08314 and an R2 = 0.162. A nonlinear HistGradientBoostingRegressor achieved an RMSE = 0.08403 with history alone and 0.08500 with all information (R2 = 0.143 and 0.115, respectively). Thus, the nonlinear challenge did not improve performance when coarse-scene or external environmental predictors were added.
Relative to history-only ridge regression, adding the coarse scene set yielded an ΔRMSE = +0.00002 (95% CI −0.00026 to 0.00027), with improvement for 8 of 12 drivers. Adding external CO2/illuminance summaries yielded an ΔRMSE = −0.00019 (95% CI −0.00037 to −0.00003), with improvement for 3 of 12 drivers.
On the common 4259-window sample used for expanded candidates, ridge regression achieved an RMSE = 0.08123 (R2 = 0.187) and extended AR achieved an RMSE = 0.08153 (R2 = 0.182). Outer-loop selections covered 90–240 s windows and AR orders 12–36. Initialization sensitivities were evaluated at 30, 60, 120, and 180 s (Figure 6).
Under the revised 15-variable completeness definition, a Q ≥ 0.5 retained 3729 of the 4835 main evaluation windows and excluded 1106 (22.87%). A Q ≥ 0.7 retained 3379 and excluded 1456 (30.11%). These exclusions quantify the sample impact of the prespecified data-quality sensitivities.
A fully nested six-component sensitivity explained 67.40% of the variance in the complete-data structure (mean training-fold estimate across the 12 outer loops, 68.14%). History-only ridge and direct AR achieved RMSEs of 0.08117 and 0.08114, respectively; the scene and environmental models achieved RMSEs of 0.08125 and 0.08133. The model ordering was qualitatively consistent with the five-component analysis.

3.4. Cross-Driver Generalization and Predictive Uncertainty

This section reports the same contrasts at the independent-driver level: Figure 7 shows the direction and magnitude of each driver’s change, while Figure 8 illustrates the point-prediction and interval-calibration display for one out-of-fold sequence.
The coarse-scene model improved the RMSE for 8 of 12 drivers, but the driver-bootstrap confidence interval for the paired difference included zero, so the average effect was not statistically distinguishable from zero. The external environmental model improved performance for 3 of 12 drivers and increased average error. The nonlinear model improved 2 of 12 drivers and performed worse on average. Driver-level counts describe the present sample.
Ridge windows selected across drivers were 90 s (n = 1), 120 s (n = 2), 180 s (n = 4), and 240 s (n = 5). AR orders were p = 12 (n = 1), p = 18 (n = 4), p = 24 (n = 3), and p = 36 (n = 4). These selections are sensitivity results for the present sample and tested grid.
The common-sample comparison contained 4259 windows and showed only a small difference between extended ridge and extended AR (RMSE 0.08123 versus 0.08153). Because the primary revision changed target construction and calibration, the previous interval-width and coverage values are not carried forward as validated safety metrics. Interval calibration remains a secondary uncertainty benchmark to be recomputed alongside any future externally validated target.
Initialization length materially changes the calibration of the driver-specific target and therefore the absolute scale of the RMSE. The 30, 60, and 120 s sensitivity values were reported as a robustness audit, while the anomalously low 180 s value was flagged as target-definition-dependent. Future deployment studies should prespecify initialization, quantify performance during the initialization period, and compare population, personalized, and switching models on a common externally meaningful outcome.

4. Discussion

4.1. Historical Continuity Within Limited Predictability

History-only ridge regression and tuned AR produced very similar errors, and both improved on the simple baselines. Recent CDI history constituted the strongest and most reproducible reference under the evaluated conditions. The fully nested audit further showed that information sources with descriptive relevance do not necessarily provide incremental forward-prediction value after historical continuity is controlled.
Within the evaluated feature space and model classes, recent history was the most reproducible reference for cross-driver forecasting. Higher-resolution scene, environmental, and geometric predictors remain to be tested.
The CDI is a statistical human–vehicle composite, not a validated workload, fatigue, safety, clinical, or operational scale. Acceleration and angular-velocity variables participate in target construction and therefore cannot be used as independent external outcomes. Future construct-validity work should prespecify blinded video labels, lane-position variability, TTC, sudden-deceleration events, or other outcomes collected independently of the CDI. It should also use leave-one-modality-out or otherwise non-overlapping target construction when testing part–whole relationships.

4.2. Boundaries of Incremental Value and Implications for Study Design

The hierarchical model showed descriptive scene and direction associations, whereas the out-of-fold audit found no stable mean predictive increment from the coarse-scene set; external CO2 and illuminance summaries slightly increased the average error. These analyses answer different questions. The mixed model describes contemporaneous associations in the CDI, whereas the predictive audit tests whether the same variables reduce future error for a held-out driver after recent CDI history is optimized. The significant mixed-model associations may partly reflect temporal confounding or co-trending, because scene labels, elapsed route time, direction, and contemporaneous CDI can vary together. Once recent CDI history and driver-specific initialization are included, this shared temporal structure may be absorbed by historical predictors, leaving limited independent information for scene variables. The lower cross-driver stability of PC4 and PC5 may further increase target noise and reduce the detectability of small incremental effects. Continuous curvature, grade, portal distance, illumination-change rate, traffic flow, and car-following status should be evaluated under the same protocol (Table 4).
Across the full sample and Q-threshold sensitivities, the coarse route and external environmental sets showed no stable positive predictive increment. These results characterize the tested low-resolution representations; higher-resolution predictors and independently annotated driving-performance outcomes remain priorities for future data collection.

4.3. Methodological Contribution of Fully Nested Validation

The methodological contribution is a leakage-controlled predictive-audit framework that integrates target construction, classical time-series baselines, model selection, nonlinear challenge models, and driver-level outer-loop evaluation. This protocol quantifies incremental information value before deployment decisions.
Prediction intervals complement point metrics and provide a secondary uncertainty benchmark. Future analyses should recompute interval coverage and scores after prespecifying an independently validated outcome, including multiple nominal coverage levels, driver-level interval scores, conditional coverage, and interval width versus absolute error.

4.4. Applicability and Future Validation

The Tianshan Shengli Tunnel provides a challenging stress-test scenario for cross-driver prediction. Applying the audit protocol to other settings will require continuous geometry, traffic flow, and independently measured driving-performance outcomes in future studies.
These findings define what can and cannot be inferred from TG/TS, direction, and segment-duration encodings at their current resolution. With 12 drivers as independent clusters, confidence intervals reflect limited estimation precision. Future work should expand the driver sample and evaluate continuous geometry, traffic flow, and independent driving-performance outcomes. Visual acuity was not entered as a predictor, and this sample does not support a meaningful acuity-stratified analysis. We therefore make no claim that the observed acuity differences influenced the results; a larger, prospectively characterized sample could examine whether mild acuity differences modify initialization or prediction stability.
The fully nested leave-one-driver-out design, strong time-series baselines, nonlinear challenge, paired driver-level comparisons, and explicit target-validity caveat form a transferable audit protocol. The next stage should prespecify initialization and model-switching rules, extend the history and AR grids, and add independent blinded driving-performance outcomes.
For Tianshan Shengli Tunnel operations, historical and AR models provide reference baselines before additional sensors or environmental predictors are considered. Continuous geometry, opening distance, illumination changes, traffic conditions, and independently annotated driving performance are priority next-stage data.

4.5. Reproducibility and Open Science

Reproducibility requires documenting the 15-variable target definition, the exclusion of CO2 and illuminance from PCA, preprocessing order, inner and outer fold assignments, expanded hyperparameter grids, initialization sensitivity, random seeds, software versions, and package dependencies. Raw naturalistic-driving data should not be disclosed when restricted by consent or privacy. Analysis code, parameter files, environment specifications, derived data dictionaries, and a clear statement that no independent outcome validation was performed should accompany the revised materials. Reproducible reporting should also include the explicit simple-baseline table, the six-component nested sensitivity, the 20,000-resample driver-percentile bootstrap procedure, and the numbers of observations and evaluation windows excluded by each Q threshold.

5. Conclusions

This study provides a corridor-specific, leakage-controlled benchmark for auditing short-horizon prediction and the incremental value of multimodal information. Rather than introducing a new fusion architecture, the framework tests whether synchronized sensor streams add reproducible value beyond optimized historical baselines. Based on naturalistic driving data from 12 drivers, the five-component CDI explained 60.36% of the variance. History-only ridge and direct AR performed similarly; coarse-scene information produced negligible average change, while external environmental summaries slightly increased the error. The findings apply to the present corridor, sample, target definition, feature representations, and evaluated model classes. Future work should apply this protocol to continuous geometric, traffic, and independently measured driving-performance outcomes.
(1)
The five principal components explain 60.36% of the human–vehicle CDI variance. Time-preserving parallel analysis retained more components formally, but weaker separation and loading stability after PC5 supported the parsimonious five-component target.
(2)
Tuned history-only ridge and direct AR performed similarly, and both improved on the simple baselines. Expanded searches showed that selected history windows and AR orders varied across outer folds, while initialization length materially affected target calibration and absolute error. These settings should therefore be re-estimated in future audits rather than treated as fixed operational parameters.
(3)
The hierarchical mixed-effects model describes contemporaneous associations. The nested prediction audit found no stable increment from the coarse scene or external environmental summaries, and the nonlinear all-information benchmark was worse than history-only ridge. The six-component sensitivity produced the same qualitative ordering.
(4)
Initialization length materially changed target calibration and absolute RMSEs, including an anomalous 180 s sensitivity result. This methodological sensitivity underscores the need to re-estimate initialization and search settings within each audit rather than treat them as fixed operational parameters. Independent blinded video, lane-position, TTC, sudden-deceleration, or other external outcomes are needed before the CDI can be interpreted in behavioral or operational terms.

Author Contributions

Conceptualization, C.S., X.K., L.N. and Y.Z.; methodology, C.S., X.K., L.N., Y.Z. and Y.L.; software, X.K.; validation, Y.Z., L.N. and Y.L.; formal analysis, L.N. and Y.L.; investigation, L.N., Y.Z., C.S., X.K. and Y.L.; resources, C.S.; data curation, L.N. and X.K.; writing—original draft preparation, L.N. and X.K.; writing—review and editing, L.N., Y.Z., C.S., X.K. and Y.L.; visualization, X.K., L.N., Y.Z., C.S. and Y.L.; supervision, Y.Z., L.N., C.S. and Y.L.; project administration, C.S. and X.K. All authors have read and agreed to the published version of the manuscript.

Funding

This study was funded by CCCC Electromechanical Engineering Bureau Co., Ltd. 2025-JDJKJ-008 and BJK2024102 of Hebei Provincial Department of Education.

Institutional Review Board Statement

Ethical review was waived in accordance with Article 32 of the Measures for Ethical Review of Life Science and Medical Research Involving Humans (Order No. 4 of the National Health Commission of the People’s Republic of China, 2023). The study used coded human information data, caused no harm beyond the naturalistic driving protocol, did not involve sensitive personal information, and was not conducted for commercial purposes. Under this provision, no ethics-committee review or separate institutional waiver application was required; therefore, no approval or waiver number or date applies.

Informed Consent Statement

Written informed consent was obtained from all participants.

Data Availability Statement

Due to participant privacy and informed-consent restrictions, raw naturalistic-driving data are not publicly available. Analysis code, parameter specifications, derived variable definitions, and data-processing documentation will be made available from the corresponding author upon reasonable request, subject to applicable privacy restrictions.

Acknowledgments

The authors would like to express their gratitude to CCCC Mechanical and Electrical Engineering Bureau Co., Ltd., for their support.

Conflicts of Interest

Authors Chunhui Shi and Yu Zhang were employed by the company CCCC Mechanical and Electrical Engineering Bureau Co., Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Dong, W.H.; Zhang, Z.Q.; Luan, S.; Zhao, X.; Zhang, F.; Han, W. Exploring the impact of discontinuous murals in extra-long tunnels on drivers’ attention. J. Transp. Eng. Part A Syst. 2026, 152, 04026048. [Google Scholar] [CrossRef] [Scilit]
  2. Lu, F.T.; Zheng, X. EEG characterization of driving load in long highway tunnels and circadian differences. In International Conference on Frontiers of Traffic and Transportation Engineering (FTTE 2025); SPIE: Bellingham, WA, USA, 2026; Volume 35. [Google Scholar] [CrossRef] [Scilit]
  3. Yang, Y.Z.; Du, Z.G.; Alonso, F.; Faus, M. Why does driver attention abnormally decrease? An experimental analysis on the slack effect at highway tunnel entrances and exits. Transp. Res. Part F Traffic Psychol. Behav. 2025, 111, 145–161. [Google Scholar] [CrossRef] [Scilit]
  4. Mei, J.L.; Wang, S.S.; He, S.M.; Du, Z.; Jiao, F. Speed control, visual adaptation, and mental workload in urban short underpass tunnels: A naturalistic driving study. Traffic Inj. Prev. 2025, 27, 707–715. [Google Scholar] [CrossRef] [Scilit]
  5. Yang, Y.Z.; Alonso, F.; Faus, M.; Du, Z.; Mei, J. Exploring the causes of frequent accidents at highway tunnel exits: Coupling analysis of the slack effect and white hole effect in extra-long tunnels. Transp. Res. Part F Traffic Psychol. Behav. 2024, 106, 288–305. [Google Scholar] [CrossRef] [Scilit]
  6. Wang, S.S.; Du, Z.G.; Jiao, F.T.; Zheng, H.; Ni, Y. Drivers’ visual load at different time periods in entrance and exit zones of extra-long tunnel. Traffic Inj. Prev. 2020, 21, 539–544. [Google Scholar] [CrossRef] [Scilit]
  7. Qin, P.C.; Wang, M.N.; Chen, Z.W.; Yan, G.; Yan, T.; Han, C.; Bao, Y.; Wang, X. Characteristics of driver fatigue and fatigue-relieving effect of special light belt in extra-long highway tunnel: A real-road driving study. Tunn. Undergr. Space Technol. 2021, 114, 103990. [Google Scholar] [CrossRef] [Scilit]
  8. Peng, L.; Weng, J.; Yang, Y.; Wen, H. Impact of light environment on driver’s physiology and psychology in interior zone of long tunnel. Front. Public Health 2022, 10, 842750. [Google Scholar] [CrossRef] [Scilit]
  9. Zhang, B.; Bai, J.R.; Yin, Z.W.; Zhou, A.; Li, J. Study on the driver visual workload of bridge-tunnel groups on mountainous expressways. Appl. Sci. 2023, 13, 10186. [Google Scholar] [CrossRef] [Scilit]
  10. Han, L.; Du, Z.G.; Ma, A.J. Evaluation of traffic signs information volume at highway tunnel entrance zone based on the visual sample entropy of novice and experienced drivers. Traffic Inj. Prev. 2024, 25, 499–509. [Google Scholar] [CrossRef] [Scilit]
  11. De Waard, D. The Measurement of Drivers’ Mental Workload; University of Groningen: Groningen, The Netherlands, 1996. [Google Scholar]
  12. Brookhuis, K.A.; de Waard, D. Monitoring drivers’ mental workload in driving simulators using physiological measures. Accid. Anal. Prev. 2010, 42, 898–903. [Google Scholar] [CrossRef] [Scilit]
  13. Recarte, M.A.; Nunes, L.M. Mental workload while driving: Effects on visual search, discrimination, and decision making. J. Exp. Psychol. Appl. 2003, 9, 119–137. [Google Scholar] [CrossRef] [Scilit]
  14. He, D.B.; Wang, Z.Q.; Khalil, E.B.; Donmez, B.; Qiao, G.; Kumar, S. Classification of driver cognitive load: Exploring the benefits of fusing eye-tracking and physiological measures. Transp. Res. Rec. 2022, 2676, 670–681. [Google Scholar] [CrossRef] [Scilit]
  15. Angkan, P.; Behinaein, B.; Mahmud, Z.; Bhatti, A.; Rodenburg, D.; Hungler, P.; Etemad, A. Multimodal brain-computer interface for in-vehicle driver cognitive load measurement: Dataset and baselines. IEEE Trans. Intell. Transp. Syst. 2024, 25, 5949–5964. [Google Scholar] [CrossRef] [Scilit]
  16. Zhou, Y.; Chen, Y.X.; Zhang, Y.X. Driver distraction detection in conditionally automated driving using multimodal physiological and ocular signals. Electronics 2025, 14, 3811. [Google Scholar] [CrossRef] [Scilit]
  17. Fresta, M.; Bellotti, F.; Bochenko, I.; Lazzaroni, L.; Merlhiot, G.; Tango, F.; Berta, R. Deep learning-based real-time driver cognitive distraction detection. IEEE Access 2025, 13, 26589–26607. [Google Scholar] [CrossRef] [Scilit]
  18. Navarro, J.; Lappi, O.; Osiurak, F.; Hernout, E.; Gabaude, C.; Reynaud, E. Dynamic scan paths investigations under manual and highly automated driving. Sci. Rep. 2021, 11, 3776. [Google Scholar] [CrossRef] [Scilit]
  19. Majdi, M.S.; Ram, S.; Gill, J.T.; Rodriguez, J.J. Drive-Net: Convolutional network for driver distraction detection. In Proceedings of the 2018 IEEE Southwest Symposium on Image Analysis and Interpretation, Las Vegas, NV, USA, 8–10 April 2018; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  20. Lorenzo-Seva, U.; ten Berge, J.M.F. Tucker’s congruence coefficient as a meaningful index of factor similarity. Methodology 2006, 2, 57–64. [Google Scholar] [CrossRef] [Scilit]
  21. Shmueli, G. To explain or to predict? Stat. Sci. 2010, 25, 289–310. [Google Scholar] [CrossRef] [Scilit]
  22. Jacobé de Naurois, C.; Bourdin, C.; Stratulat, A.; Diaz, E.; Vercher, J.-L. Detection and prediction of driver drowsiness using artificial neural network models. Accid. Anal. Prev. 2019, 126, 95–104. [Google Scholar] [CrossRef] [Scilit]
  23. Zhou, X.F.; Kundu, S. Prediction of driver drowsiness level using recurrent neural networks and multi-time-scale fusion. In SAE Technical Paper 2021-01-0909; SAE International: Warrendale, PA, USA, 2021. [Google Scholar] [CrossRef] [Scilit]
  24. Gao, J.; Yi, J.G.; Murphey, Y.L. An efficient driving behavior prediction approach using physiological auxiliary and adaptive LSTM. Mach. Vis. Appl. 2024, 35, 113. [Google Scholar] [CrossRef] [Scilit]
  25. Varma, S.; Simon, R. Bias in error estimation when using cross-validation for model selection. BMC Bioinform. 2006, 7, 91. [Google Scholar] [CrossRef] [Scilit]
  26. Roberts, D.R.; Bahn, V.; Ciuti, S.; Boyce, M.S.; Elith, J.; Guillera-Arroita, G.; Hauenstein, S.; Lahoz-Monfort, J.J.; Schröder, B.; Thuiller, W.; et al. Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography 2017, 40, 913–929. [Google Scholar] [CrossRef] [Scilit]
  27. Bergmeir, C.; Hyndman, R.J.; Koo, B. A note on the validity of cross-validation for evaluating autoregressive time series prediction. Comput. Stat. Data Anal. 2018, 120, 70–83. [Google Scholar] [CrossRef] [Scilit]
  28. Saeb, S.; Lonini, L.; Jayaraman, A.; Mohr, D.C.; Kording, K.P. The need to approximate the use-case in clinical machine learning. GigaScience 2017, 6, gix019. [Google Scholar] [CrossRef] [Scilit]
  29. Kapoor, S.; Narayanan, A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 2023, 4, 100804. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Experimental corridor from Haxionggou Tunnel to Tianshan Shengli Tunnel. TG denotes the tunnel-group composite route proxy and TS denotes the Tianshan Shengli main tunnel; inset photographs show representative corridor segments.
Figure 1. Experimental corridor from Haxionggou Tunnel to Tianshan Shengli Tunnel. TG denotes the tunnel-group composite route proxy and TS denotes the Tianshan Shengli main tunnel; inset photographs show representative corridor segments.
Applsci 16 08998 g001
Figure 2. Experimental vehicle and data-acquisition equipment. Panels (ad) show the experimental vehicle, recording interface, Tobii Glasses 3, and onboard diagnostic units, respectively.
Figure 2. Experimental vehicle and data-acquisition equipment. Panels (ad) show the experimental vehicle, recording interface, Tobii Glasses 3, and onboard diagnostic units, respectively.
Applsci 16 08998 g002
Figure 3. Fully nested leave-one-driver-out validation, three-fold grouped inner validation, and parameter isolation. Arrows indicate the information flow; the 120 s initialization is unsupervised, and all fitted parameters are estimated inside the relevant training fold before held-out-driver evaluation.
Figure 3. Fully nested leave-one-driver-out validation, three-fold grouped inner validation, and parameter isolation. Arrows indicate the information flow; the 120 s initialization is unsupervised, and all fitted parameters are estimated inside the relevant training fold before held-out-driver evaluation.
Applsci 16 08998 g003
Figure 4. PCA component selection and driver-bootstrap stability. The dashed selection marker identifies the parsimonious five-component target; the stability panel reports driver-bootstrap congruence with uncertainty bars. In panel (b), the green dashed line indicates the 0.85 Tucker congruence reference threshold.
Figure 4. PCA component selection and driver-bootstrap stability. The dashed selection marker identifies the parsimonious five-component target; the stability panel reports driver-bootstrap congruence with uncertainty bars. In panel (b), the green dashed line indicates the 0.85 Tucker congruence reference threshold.
Applsci 16 08998 g004
Figure 5. Normalized CDI distributions by direction and composite TG/TS scene at the 10 s aggregation scale. Boxes show the interquartile range, center lines show medians, and whiskers extend to 1.5 times the interquartile range.
Figure 5. Normalized CDI distributions by direction and composite TG/TS scene at the 10 s aggregation scale. Boxes show the interquartile range, center lines show medians, and whiskers extend to 1.5 times the interquartile range.
Applsci 16 08998 g005
Figure 6. (a) Expanded inner-validation error curves and outer-loop selections for history windows and AR orders. Panels (i,ii) show inner-validation error curves and panels (iii,iv) show outer-loop selection counts for ridge windows and AR orders. (b) Incremental predictive value of added information in the original-range 30 s task. Positive ΔRMSEs denote error reduction; whiskers are 95% driver-bootstrap ordinary-percentile intervals.
Figure 6. (a) Expanded inner-validation error curves and outer-loop selections for history windows and AR orders. Panels (i,ii) show inner-validation error curves and panels (iii,iv) show outer-loop selection counts for ridge windows and AR orders. (b) Incremental predictive value of added information in the original-range 30 s task. Positive ΔRMSEs denote error reduction; whiskers are 95% driver-bootstrap ordinary-percentile intervals.
Applsci 16 08998 g006
Figure 7. Driver-specific incremental value for the original-range 30 s task. Green bars indicate lower error after adding the information source and magenta bars indicate higher error; the unit is the driver.
Figure 7. Driver-specific incremental value for the original-range 30 s task. Green bars indicate lower error after adding the information source and magenta bars indicate higher error; the unit is the driver.
Applsci 16 08998 g007
Figure 8. Example out-of-fold prediction and prediction-interval calibration benchmark for the history-only CDI model. The black line is the observed CDI, the blue line is the out-of-fold prediction, and the shaded band is the 90% prediction interval.
Figure 8. Example out-of-fold prediction and prediction-interval calibration benchmark for the history-only CDI model. The black line is the observed CDI, the blue line is the out-of-fold prediction, and the shaded band is the 90% prediction interval.
Applsci 16 08998 g008
Table 1. Explained variance, dominant loadings, and bootstrap congruence for the five-component human–vehicle CDI. Percentages are based on the 15-variable human–vehicle target; the Tucker median is the driver-bootstrap congruence summary.
Table 1. Explained variance, dominant loadings, and bootstrap congruence for the five-component human–vehicle CDI. Percentages are based on the 15-variable human–vehicle target; the Tucker median is the driver-bootstrap congruence summary.
PCEigenvalueVarianceCumulativeTucker MedianDominant Loadings (Signed)
PC12.61817.05%17.05%0.920Fixation proportion −0.576; gaze-loss proportion −0.555; fixation angular velocity −0.473
PC22.40415.65%32.71%0.875RR interval −0.599; instantaneous heart rate −0.596; mean heart rate −0.481
PC31.72711.25%43.95%0.810Angular-velocity magnitude −0.590; acceleration magnitude −0.565; SDNN +0.385
PC41.3959.08%53.04%0.658RMSSD −0.580; physiological relaxation −0.473; pupil diameter +0.437
PC51.1247.32%60.36%0.531Respiration −0.550; pupil-area change +0.482; pupil diameter −0.385
Table 2. Hierarchical mixed-effects estimates for descriptive scene and external-environment associations with the five-component CDI. Estimates are descriptive conditional associations from the prespecified saturated mixed-effects model over 10 s observations, not predictive increments.
Table 2. Hierarchical mixed-effects estimates for descriptive scene and external-environment associations with the five-component CDI. Estimates are descriptive conditional associations from the prespecified saturated mixed-effects model over 10 s observations, not predictive increments.
TermβSEp
TS composite scene0.027590.00567<0.001
Within-segment time−0.001370.00039<0.001
Outbound direction0.038560.016090.017
TS × segment time−0.000230.000520.661
TS × direction−0.022280.007830.004
Segment time × direction0.002050.00050<0.001
TS × segment time × direction−0.000950.000770.217
CO2 (standardized external)−0.003740.001550.016
Illuminance (standardized external)0.001980.001060.062
Completeness Q0.012910.008420.125
Table 3. Simple and tuned historical baselines for the five-component CDI in the original-range 30 s task. Metrics are evaluated on the same 4835 windows in the original-range 30 s task; a lower RMSE/MAE and higher R2 indicate better point prediction.
Table 3. Simple and tuned historical baselines for the five-component CDI in the original-range 30 s task. Metrics are evaluated on the same 4835 windows in the original-range 30 s task; a lower RMSE/MAE and higher R2 indicate better point prediction.
ModelRMSEMAER2
30 s historical mean0.086950.065820.082
Persistence0.102200.07808−0.289
Outer-training mean0.102370.07937−0.294
History-only ridge0.082940.062770.166
Direct AR(p)0.082610.062870.168
Table 4. Overview of the Tianshan Shengli Tunnel study scene, target definition, and information resolution. Resolution describes the present data representation; the implications delimit the audit scope and are not deployment or safety claims.
Table 4. Overview of the Tianshan Shengli Tunnel study scene, target definition, and information resolution. Resolution describes the present data representation; the implications delimit the audit scope and are not deployment or safety claims.
Information LevelCurrently Reportable
Content
ResolutionModel RepresentationImplication for Subsequent Audits
Study corridorHaxionggou Tunnel to Tianshan Shengli Tunnel, 36.2 kmRoute extentTG/TS, outbound/returnDefines the naturalistic-driving scene and route transitions.
Corridor-specific evidence from 12 drivers; not a population-valid effect estimate.
Human–vehicle observations12 drivers, 24 directional trips, 4931 10 s observations; 15 ocular, physiological, and vehicle-motion variables1 s synchronization; 10 s evaluationFive-component CDI, trailing means, and OLS trendsEstablishes a cross-driver statistical target and prediction baseline.
Statistical prediction target only; no direct driver-state or safety interpretation.
External environmental predictorsCO2 and illuminance summaries excluded from PCA target1 s synchronization; window summariesExternal means and OLS trendsTheir incremental value must be audited without target contamination.
Results apply to the evaluated environmental summaries and models, not to all environmental information.
Coarse spatial proxiesTG/TS, direction, segment durationCoarseLabels and interactionsNo stable mean predictive increment observed.
The null increment is specific to the present coarse labels and tested models.
Engineering variables and outcomes to addContinuous grade, curvature, portal distance, illuminance-change rate, traffic flow, car-following state, and independent driving-performance outcomesNot collected or not used in the present analysisNext-stage ARX, modality ablation, and external-validity analyses.Connects tunnel attributes and independently measured outcomes with predictive performance.
Future variables must pass the same audit before any operational or safety claim.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Shi, C.; Kang, X.; Nie, L.; Zhang, Y.; Li, Y. Information Sources and Incremental Value in Short-Horizon Prediction of a Multimodal Driving Index in Extra-Long Tunnels. Appl. Sci. 2026, 16, 8998. https://doi.org/10.3390/app16188998

AMA Style

Shi C, Kang X, Nie L, Zhang Y, Li Y. Information Sources and Incremental Value in Short-Horizon Prediction of a Multimodal Driving Index in Extra-Long Tunnels. Applied Sciences. 2026; 16(18):8998. https://doi.org/10.3390/app16188998

Chicago/Turabian Style

Shi, Chunhui, Xuejian Kang, Liangtao Nie, Yu Zhang, and Yuner Li. 2026. "Information Sources and Incremental Value in Short-Horizon Prediction of a Multimodal Driving Index in Extra-Long Tunnels" Applied Sciences 16, no. 18: 8998. https://doi.org/10.3390/app16188998

APA Style

Shi, C., Kang, X., Nie, L., Zhang, Y., & Li, Y. (2026). Information Sources and Incremental Value in Short-Horizon Prediction of a Multimodal Driving Index in Extra-Long Tunnels. Applied Sciences, 16(18), 8998. https://doi.org/10.3390/app16188998

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop