Research on Effectiveness of Vehicle Driving Simulation System Based on Coupling Modeling of Driving Behavior and Psychology
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe manuscript talks about an important topic in driving simulation research. It combines biopsychological indicators into simulator evaluation. The idea of using LSTM modeling and Bayesian optimization is new and could be very useful for the field.. The authors need to make it clearer how their "Mul-Bayes-LSTM" framework is different from other methods that are already out there for validating simulators.
The literature review is thorough. It needs to be more critical. Now it just lists a lot of references without really explaining how they relate to the research gap that this study is trying to fill. The authors should also talk about traditional studies and what they found. For example they could mention studies like the ones at https://doi.org/10.1016/j.ijtst.2022.06.001 and https://doi.org/10.3390/infrastructures11030106 .
The methodology section needs clarification. Specifically the authors need to explain how they created the two- feature matrix and how they combined CNN and LSTM. If they provided mathematical and step-by-step details it would be easier for others to repeat the study.
The number of participants in the experiment is relatively small with 38 people taking part. The authors should discuss how this might affect the models ability to generalize and whether they considered using data or trying different validation strategies.
The results show that the model performs well especially when it comes to predicting ECG and EDA. However the authors do not fully explain why the model does not perform well with EMG. If they did analysis of how sensitive the features are or how much the physiological signals vary it would make their conclusions stronger.
The paper would be more convincing if it compared the driving simulation research to real-world driving or naturalistic data. Even a small-scale comparison with on-road experiments would make the proposed evaluation framework more practical and credible. The authors do acknowledge this limitation in the conclusion. It would be better if they actually did the comparison.
The manuscript needs to make the Mul-Bayes-LSTM framework and the driving simulation research more clear and connected to real-world driving. This would make the study more useful and relevant to the field. The authors should focus on making the methodology. Results more detailed and easy to understand and they should try to compare their findings to real-world driving data. This would make the manuscript more convincing and useful, to readers.
Comments on the Quality of English LanguageModerate changes are needed.
Author Response
Comment 1: The manuscript talks about an important topic in driving simulation research. It combines biopsychological indicators into simulator evaluation. The idea of using LSTM modeling and Bayesian optimization is new and could be very useful for the field.. The authors need to make it clearer how their "Mul-Bayes-LSTM" framework is different from other methods that are already out there for validating simulators.
Authors’ reply:We appreciate the expert's affirmation of the value of the research topic and methodology of this paper. We have added an explanation of the Mul-Bayes-LSTM framework at the end of the introduction and at the beginning of Chapter 3, and further clarified its differences from traditional simulator validation methods. Traditional driving simulator validation methods mainly focus on physical validity and behavioral validity, such as simulator hardware, vehicle dynamics, visual display, vehicle speed, braking, reaction time, lane position, and vehicle trajectory indicators. This paper further introduces physiological and psychological indicators such as EDA, ECG, and EMG to evaluate whether the driving simulator can elicit drivers' internal perception and physiological-psychological responses in risk conflict scenarios.
Authors’ revision:Add at the end of the introduction:This paper constructs a Mul-Bayes-LSTM evaluation model, taking vehicle, road, and conflict scene parameters as inputs, and EDA, ECG, and EMG as outputs, to extract the correlation features between driving behavior and physiological and psychological responses, and to evaluate the effectiveness of the driving simulator.
Supplement at the beginning of Chapter 3:This chapter constructs a Mul-Bayes-LSTM evaluation model based on the coupling of driving behavior and physiological-psychological responses. The model takes vehicle operation parameters, road environment parameters, and conflict scenario parameters as inputs, and EDA, ECG, and EMG as outputs. It extracts attribute correlations and temporal dynamic features through CNN-LSTM, and uses Bayesian optimization to improve the efficiency of model parameter selection.
Revision notes: For the specific modification locations, see the revised draft page 4, lines 173-177; page 8, lines 250-255.
Comment 2: The literature review is thorough. It needs to be more critical. Now it just lists a lot of references without really explaining how they relate to the research gap that this study is trying to fill. The authors should also talk about traditional studies and what they found. For example they could mention studies like the ones at https://doi.org/10.1016/j.ijtst.2022.06.001 and https://doi.org/10.3390/infrastructures11030106 .
Authors’ reply:We appreciate the expert for pointing out the lack of logic and criticality in the literature review section. We agree that the literature review in the original manuscript mostly lists existing studies, without adequately explaining the logical relationships between different studies or the research gaps that this paper attempts to address.
Based on the expert's suggestions, we have organized and optimized the literature review in the introduction. We further summarized the relevant literature along three main lines: first, early research mainly focused on the physical validity of driving simulators, that is, whether the simulator hardware, vehicle dynamics, motion feedback, visual display, and control interfaces can replicate real driving environments; second, with the increased application of driving simulators in traffic behavior research, studies gradually shifted toward behavioral validity, mainly verifying the simulator's effectiveness through observable behavioral indicators such as speed, braking response, reaction time, trajectory deviation, and parking compliance; third, in recent years, research has begun to focus on driver human factor responses in complex traffic scenarios and high-risk environments, but studies on the relationship between physiological and psychological responses and driving simulator validity remain relatively insufficient.
At the same time, we unified the reference formatting and appropriately updated and supplemented the relevant literature according to the research lines, especially studies related to driving simulator behavioral validity, complex road environments, human factor safety assessment, and physiological and psychological responses. Through the above revisions, the revised manuscript can more clearly explain how existing research supports the study background, and further clarifies the research gap of this paper: existing studies largely rely on physical performance or observable behavioral indicators, lacking a systematic evaluation framework that combines driving behavior, physiological and psychological responses, and comprehensive weighted assessment.
Authors’ revision:The introduction is revised as follows: In addition, if the risk conflict scenarios designed in this study were to be conducted in real vehicle experiments on actual roads, they could involve high safety risks and restrictions related to road test approvals. Therefore, this study chooses to construct controlled conflict scenarios in a driving simulator, aiming to evaluate the simulator's ability to replicate driver human factor responses under safe, controllable, and repeatable conditions.
To more clearly illustrate the relationship between existing research and the research gap in this paper, the study on the effectiveness of driving simulators is summarized into three stages: physical validity, behavioral validity, and human factors response validity. The first stage mainly focuses on the physical validity of the simulator, with its core question being whether the simulator's hardware system, vehicle dynamics, motion feedback, visual display, and control interface can closely approximate real vehicles and real road environments.
Add to the summary of the first phase: Overall, studies on physical validity provide an important basis for the verification of basic performance of driving simulators, but their focus is mainly on device performance, motion feedback, and scene realism, with insufficient attention to drivers' behavioral responses and physiological and psychological changes in complex risk scenarios.
The research focus of the second stage gradually shifted from the authenticity at the equipment level to the effectiveness of driving behavior. This stage mainly evaluates whether drivers' explicit behaviors are consistent in the two environments by comparing indicators such as vehicle speed, acceleration, braking response, reaction time, lane position, trajectory deviation, and compliance with stopping in simulated environments and real road environments.
In recent years, driving simulators have also gradually been applied to studies of more complex road and infrastructure scenarios, such as underground interchanges, rural road curves, road markings, and traffic safety facility evaluations. Liu et al. used driving simulators and data mining methods to analyze the impact of traffic safety facilities at underground interchanges on driving safety and comfort; Martin-Castresana et al. studied the role of different road markings in controlling vehicle speed on rural road curves through driving simulators; Bosurgi et al. proposed new behavioral evaluation indicators from the perspective of driver steering behavior on curves.
Summary of the second phase: Overall, behavioral effectiveness research can fairly well evaluate the simulator's ability to replicate driving behaviors, but its evaluation still mainly focuses on observable behavioral aspects such as speed, braking, reaction time, and trajectory, making it difficult to fully reflect the driver's psychological tension, risk perception, and physiological arousal processes in risky scenarios.
Supplement to the third stage: The third stage begins to focus on drivers' human factor responses in risky traffic environments. With the development of human factors engineering and intelligent transportation research, merely analyzing overt driving behavior is no longer sufficient to fully explain the internal state of drivers. Therefore, physiological and psychological indicators have gradually been introduced into driving behavior analysis and the effectiveness evaluation of driving simulators.
Healey and Picard collected physiological signals such as ECG, EMG, electrodermal activity, and respiration during real driving tasks to identify driving stress levels. Nacpil et al. evaluated the validity of driving simulator data by analyzing surface electromyography signals during the steering process, and Nacpil and Nakano further combined surface EMG signals with machine learning classification methods for research on driver assistance interfaces. In addition, Kierzkowski et al. involved eye movement and electrodermal response analysis in a tram simulator study, reflecting the development trend of human factors data collection in simulator research. Paliotto and Meocci also constructed road safety assessment procedures from the perspective of human factors principles, indicating the importance of human factors in road safety evaluation.
These studies indicate that physiological and psychological indicators can provide a useful perspective for understanding drivers' internal states. However, existing research mostly focuses on a single physiological signal or specific driving scenarios, and there is still a lack of a systematic framework that combines driving behavior, physiological and psychological responses, machine learning modeling, and comprehensive empowerment evaluation.
Add to the beginning of the last paragraph: From the above research, it can be seen that studies on the validity of driving simulators have gradually expanded from verifying the physical realism of the equipment to analyzing behavioral consistency and evaluating human factors responses. However, different stages of research focus on different aspects, and it is still difficult to form a comprehensive evaluation of the validity of driving simulators.
Revision Notes: Specific modification locations can be found on page 2 of the revised manuscript, lines 47–52, lines 59–65, lines 81–91; page 3, lines 105–113, lines 126–137, lines 138–page 4, line 154.
Supplementary References:
- Bella, F. Driving Simulator for Speed Research on Two-Lane Rural Roads. Accident Analysis & Prevention2008, 40, 1078–1087. https://doi.org/10.1016/j.aap.2007.10.015.
- Liu, Z.; Yang, Q.; Wang, A.; Gu, X. Vehicle Driving Safety of Underground Interchanges Using a Driving Simulator and Data Mining Analysis. Infrastructures2024, 9, 28. https://doi.org/10.3390/infrastructures9020028.
- Martin-Castresana, S.; Alvarez, D.; Andrade-Cataño, F.; Castro, M. Effect of Road Markings on Speed through Curves on Rural Roads: A Driving Simulator Study in Spain. Infrastructures2025, 10, 94. https://doi.org/10.3390/infrastructures10040094.
- Paliotto, A.; Meocci, M. Development of a Network-Level Road Safety Assessment Procedure Based on Human Factors Principles. Infrastructures2024, 9, 35. https://doi.org/10.3390/infrastructures9020035.
- Kierzkowski, A.; Wolniewicz, Ł.; Danilevičius, A.; Mardeusz, E.; Kin, M.; Bakinowski, Ł.; Barabasz, D.; Wielkopolan, P. The Concept of a Universal Tram Driver Console with Interchangeable Panels for a Polish Tram Simulator. Infrastructures2024, 9, 41. https://doi.org/10.3390/infrastructures9030041.
- Bosurgi, G.; Di Perna, M.; Pellegrino, O.; Sollazzo, G.; Ruggeri, A. Drivers’ Steering Behavior in Curve by Means of New Indicators. Infrastructures2024, 9, 43. https://doi.org/10.3390/infrastructures9030043.
- Healey, J. A.; Picard, R. W. Detecting Stress during Real-World Driving Tasks Using Physiological Sensors.IEEE Transactions on Intelligent Transportation Systems, 2005, 6(2), 156–166.DOI: 10.1109/TITS.2005.848368.
Comment 3: The methodology section needs clarification. Specifically the authors need to explain how they created the two- feature matrix and how they combined CNN and LSTM. If they provided mathematical and step-by-step details it would be easier for others to repeat the study.
Authors’ reply:We appreciate the expert's suggestion regarding the reproducibility of the methods. Although the original manuscript mentioned the two-dimensional feature matrix, CNN, and LSTM, it did not provide sufficient explanation about the input variables, output variables, matrix structure, and the connection between CNN and LSTM. Based on this suggestion, we have made key revisions to Section 3.2.1, supplementing the description of the CNN attribute feature extraction process and the LSTM temporal feature extraction process.
Authors’ revision:At the end of Section 3.2.1, we added:
CNNs are mainly composed of convolutional layers, pooling layers, and fully connected layers, with the descriptions of the three layers shown in the Formula 2 and Formula 3:
In the equation,is the input of the convolutional layer; is the output of the convolutional layer, which is the input of the activation layer;is the output of the activation layer; is the weight between the l-th layer's nth unit and the i-th unit of the previous layer; is the bias term; is a nonlinear activation function; is the pooling function. Simulator data has characteristics such as randomness and time-variability. After constructing the feature matrix, the feature matrix is input into the CNN to successively extract the attribute features of the data, that is, horizontally traversing the two-dimensional feature matrix. After using CNN to extract data attribute features, it is necessary to extract temporal features; the paper uses an LSTM network to capture temporal features, that is, vertically traversing the two-dimensional feature matrix.
Modification Note: The specific modification locations can be found on page 10 of the revised draft, lines 304–319.
Comment 4: The number of participants in the experiment is relatively small with 38 people taking part. The authors should discuss how this might affect the models ability to generalize and whether they considered using data or trying different validation strategies.
Authors’ reply:We appreciate the expert's comments regarding the sample size and the model's generalization capability. We agree that the number of participants can affect the model's ability to generalize to a broader driving population. In our experiments, a total of 38 participants were included. Although the sample size is relatively limited, it still provides a certain reference value for controlled driving simulator experiments and preliminary model validation. Previous driving simulator and human factors studies have also often used similar sample sizes for controlled experimental validation; therefore, the sample size in this study can support the preliminary evaluation of the proposed method.
It should be noted that the primary purpose of this study is not to develop a universal predictive model for all drivers but to verify whether the Mul-Bayes-LSTM framework can effectively characterize the relationship between driving behavior and physiological and psychological responses in controlled risk-conflict scenarios. To mitigate the impact of the limited sample size, we employed a unified experimental scenario, consistent data collection procedures, and consistent sampling intervals in the experimental design. Furthermore, we supplemented the model stability assessment using K-fold cross-validation.
Based on the expert's suggestion, we have provided a more cautious explanation of this issue in the revised manuscript. The revision additionally emphasizes that the current sample size is sufficient for preliminary model validation, but the model's generalization capability across groups with different ages, driving experience, risk preferences, and driving styles still needs further verification. Future research will expand the sample size and combine data from more types of drivers to validate the model's external applicability.
Authors’ revision:Added at the end of Section 2.2:Although the current sample size is sufficient to support preliminary validation of the proposed Mul-Bayes-LSTM framework under controlled simulator conditions, it may still limit the model's generalizability to a broader driver population. Differences in age, driving experience, risk preference, and driving style may result in varying behaviors and psychophysiological responses. To mitigate the impact of the limited sample size, this study employed standardized experimental scenarios, unified data collection procedures, and K-fold cross-validation. Future work will further expand the participant sample and include a more diverse group of drivers to enhance the external applicability of the proposed framework.
Modification Note: The specific modification locations can be found on page 8 of the revised draft, lines 248-256.
Comment 5: The results show that the model performs well especially when it comes to predicting ECG and EDA. However the authors do not fully explain why the model does not perform well with EMG. If they did analysis of how sensitive the features are or how much the physiological signals vary it would make their conclusions stronger.
Authors’ reply:Thanks to the experts for pointing out the issue of insufficient explanation of the results. The original manuscript mainly described the prediction results of EDA, ECG, and EMG, but the reasons for the differences in the fitting effects among the three types of physiological signals were not adequately explained. According to this suggestion, we have added an analysis of the differences in features of EDA, ECG, and EMG signals in Section 4.1, and focused on explaining the reasons for the weaker fitting effect of EMG.
Authors’ revision:Supplementary content for Section 4.1:The model prediction results show that the fitting performance of EDA and ECG is better than that of EMG, which is mainly related to the stability and sensitive characteristics of different physiological signals. EDA is closely related to emotional arousal, psychological tension, and risk perception, and changes more noticeably in traffic conflict scenarios; ECG can reflect cardiovascular responses under psychological load and driving stress, with an overall relatively stable trend. In comparison, EMG is more easily affected by factors such as the grip strength on the steering wheel, arm posture, pedal operation, sensor attachment position, and individual driving habits, leading to greater short-term fluctuations and noise, and therefore the model fitting effect is relatively weak. This indicates that EMG is more suitable as an auxiliary indicator in the evaluation of driving simulator effectiveness.
Modification Note: The specific modification locations can be found on page 15 of the revised draft, lines 434–444.
Comment 6: The paper would be more convincing if it compared the driving simulation research to real-world driving or naturalistic data. Even a small-scale comparison with on-road experiments would make the proposed evaluation framework more practical and credible. The authors do acknowledge this limitation in the conclusion. It would be better if they actually did the comparison.
Authors’ reply:We appreciate the expert's suggestion to compare data from real roads or natural driving. We fully agree that such comparisons can further enhance the external validity and practical credibility of the evaluation framework presented in this paper. However, the experimental scenarios designed in this study involve risk conflict situations, and directly replicating them on real roads could pose safety risks to participants and other road users. Additionally, it would involve limitations related to road test approvals, controlling real traffic interference, and synchronizing physiological signal collection. Therefore, considering safety, controllability, and repeatability, this study chooses to conduct experiments in a driving simulator. The purpose of this study is not to prove that simulated driving is exactly equivalent to real-world driving, but to evaluate whether the driving simulator can elicit stable and explainable human-factor responses in controlled risk scenarios. Based on the expert's suggestion, we have further elaborated on this issue in the limitations and future research section, and proposed that subsequent studies will conduct external verification of indicators such as vehicle speed, braking response, trajectory characteristics, and EDA, ECG, EMG under conditions of closed-site testing, small-scale real vehicle comparison experiments, or natural driving data.
Authors’ revision:Revise the limitations in the conclusion to:This paper has several limitations, and future research can be carried out in the following aspects: First, this study mainly focuses on the biopsychological characteristics of experienced drivers. Future research could increase the sample size to examine the differences in characteristics among drivers of different genders, ages, and personality types. Second, the driver characteristic parameter system can be further expanded. In the future, visual perception, EEG signals, vehicle speed parameters, and other sensory parameter data could be integrated to build a comprehensive evaluation model and improve the overall performance of the model. Third, the results of this paper are mainly based on controlled risk conflict scenarios in driving simulators, reflecting the simulator's ability to reproduce drivers' behavioral and physiological-psychological responses in such scenarios. Therefore, the evaluation results in this paper should be understood as the internal validity of the simulator under controlled experimental conditions and cannot be directly equated with the external validity under real road driving conditions. Future research still needs to combine real road experiments, closed-field experiments, or natural driving data to further compare and verify the simulator experiment results. Finally, this paper has not yet conducted a direct comparison with real road or natural driving data. Future research could use closed-field tests, real road driving tests, or natural driving data to compare simulator data with real driving data in terms of speed, braking behavior, reaction time, vehicle trajectory, EDA, ECG, and EMG, in order to further verify the external validity of the evaluation framework proposed in this paper.
Modification Note: The specific modification locations can be found on page 21-22 of the revised draft, lines 644–664.
Comment 7: The manuscript needs to make the Mul-Bayes-LSTM framework and the driving simulation research more clear and connected to real-world driving. This would make the study more useful and relevant to the field. The authors should focus on making the methodology. Results more detailed and easy to understand and they should try to compare their findings to real-world driving data. This would make the manuscript more convincing and useful, to readers.
Authors’ reply:We thank the experts for their comprehensive suggestions on the entire manuscript. Based on these comments, we have adjusted the overall logic of the paper. First, in the introduction, we added the practical significance of evaluating the effectiveness of driving simulators, explaining that the simulator does not represent fully real driving, and therefore its effectiveness needs to be assessed from both driver behavior and physiological-psychological responses. Second, in the methods section, we supplemented the Mul-Bayes-LSTM framework, including data preprocessing, CNN-LSTM feature extraction, Bayesian hyperparameter optimization, and integrated weighting. Finally, in the conclusions and limitations, we explicitly stated that this study primarily evaluates internal validity under controlled conflict scenarios, and future work should incorporate real-world road data for external validation.
Authors’ revision:The introduction section has been revised:In addition, if the risk conflict scenarios designed in this study were conducted directly in real-road experiments with actual vehicles, they could involve higher safety risks and road testing approval constraints. Therefore, this study chooses to construct controlled conflict scenarios in a driving simulator, with the aim of evaluating the simulator's ability to reproduce driver human factor responses under safe, controllable, and repeatable conditions.
Add after Chapter 3 title:This chapter constructs a Mul-Bayes-LSTM evaluation model based on the coupling of driving behavior and physiological-psychological responses. The model takes vehicle operation parameters, road environment parameters, and conflict scenario parameters as inputs, and EDA, ECG, and EMG as outputs. It extracts attribute correlations and temporal dynamic features through CNN-LSTM, and uses Bayesian optimization to improve the efficiency of model parameter selection.
Limitations of the conclusion revised to:This paper has several limitations, and future research can be carried out in the following aspects: First, this study mainly focuses on the biopsychological characteristics of experienced drivers. Future research could increase the sample size to examine the differences in characteristics among drivers of different genders, ages, and personality types. Second, the driver characteristic parameter system can be further expanded. In the future, visual perception, EEG signals, vehicle speed parameters, and other sensory parameter data could be integrated to build a comprehensive evaluation model and improve the overall performance of the model. Third, the results of this paper are mainly based on controlled risk conflict scenarios in driving simulators, reflecting the simulator's ability to reproduce drivers' behavioral and physiological-psychological responses in such scenarios. Therefore, the evaluation results in this paper should be understood as the internal validity of the simulator under controlled experimental conditions and cannot be directly equated with the external validity under real road driving conditions. Future research still needs to combine real road experiments, closed-field experiments, or natural driving data to further compare and verify the simulator experiment results. Finally, this paper has not yet conducted a direct comparison with real road or natural driving data. Future research could use closed-field tests, real road driving tests, or natural driving data to compare simulator data with real driving data in terms of speed, braking behavior, reaction time, vehicle trajectory, EDA, ECG, and EMG, in order to further verify the external validity of the evaluation framework proposed in this paper.
Revision Notes: Specific modification locations can be found on page 2 of the revised draft, lines 47–52; page 8, lines 259–264; pages 21–22, lines 644–664.
Acknowledgements at the End
Once again, I sincerely thank the reviewers for their valuable suggestions for revisions. All issues have been addressed one by one. The revised paper is logically rigorous, content-wise complete, and well-supported by experiments. I respectfully request the experts to review it and consider it for acceptance!
Sincerely,
Respectfully yours!
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsDear authors,
Hereafter some comments from my side:
- Overall, can you check the text throughout the whole document? In particular, many formulas, from 8) on, are incomplete, with some parts missing.
- Similarly, you named figures as Tables: can you check it?
- In addition, can put some summary text between title level 1 and title level 2 (e.g., section 3 and sub-section 3.1)? I think this can help the reader…
- On page 2, at the beginning of the section, the problem is also when you use the driving simulator with subjects, because they know it is not real: this can affect their behavior.
- On page 4, towards the end, how did you select these values of speed? Can they be dynamic, meaning they can be changed dynamically?
- On page 5, section 2.2., what is the age of users?
- Same page, at the beginning of section 3.1, can you detail a little bit more about the problem of “weight influence”?
- On page 7, at the end of section 3.2.1, usually a dataset is divided in three parts: training, checking / validation and testing sets; have you considered training and checking together?
- Same page, towards the end, how do you consider Bayesian parameters with previous parameters’ information?
- On page 11, formulas (10) and (11), you wrote twice that “The positive indicator is shown in formula …”: is this correct?
I hope my comments can help to improve your work.
Author Response
Response to Reviewers’ comments for ID infrastructures-4318835: Research on Effectiveness of Vehicle Driving Simulation System Based on Coupling Modeling of Driving Behavior and Psychology
Dear Editor ,
We would like to thank you very much for your help in handling our paper. We would also like to thank the reviewers for their valuable time in providing constructive feedback. Please find below our responses to the comments. For your convenience, the comments are in blue. Our reply is in black. Our revision is in red. A marked version of the manuscript has been submitted. We are confident that the quality of the paper has been significantly improved by incorporating the comments from the reviewers. We hope the revision will be satisfactory. Thank you once again, and your support is greatly appreciated.
Sincerely,
Liang Chen, JiaLin Yang,JiMing Xie,Fengbo Liu,Mingli Li.
Comment 1:Overall, can you check the text throughout the whole document? In particular, many formulas, from 8) on, are incomplete, with some parts missing.
Authors’ reply: We appreciate the expert for pointing out the issue of incomplete display of the full text and formulas. We have conducted a comprehensive review of the entire paper, focusing on checking the formulas, variable explanations, and formula numbering in Chapter 3 (Model Construction), Chapter 4 (Model Performance Evaluation), and Chapter 5 (Comprehensive Evaluation). Upon inspection, some formulas in the original manuscript indeed had issues such as incomplete display, missing superscripts and subscripts, unclear matrix expressions, and insufficient explanation before and after the formulas. Based on this feedback, we have made the following main modifications: First, we supplemented the process of CNN attribute feature extraction and the process of LSTM temporal feature extraction; second, we added an explanation in Section 3.2.2 on how Bayesian optimization uses historical parameter information for the next round of parameter selection; third, we corrected the explanation of the normalization formulas for positive and negative indicators in Section 5.1; fourth, we reorganized the calculation logic of the entropy weighting method, grey relational analysis, comprehensive weights, and comprehensive evaluation index in Chapter 5. Through these modifications, the formula expressions in the paper are now more complete, and the model construction and comprehensive evaluation process are clearer.
Authors’ revision: The following changes have been made in the revised draft:
In Section 3.2.1 we added:CNNs are mainly composed of convolutional layers, pooling layers, and fully connected layers, with the descriptions of the three layers shown in the Formula 2 and Formula 3:
In the equation,is the input of the convolutional layer; is the output of the convolutional layer, which is the input of the activation layer;is the output of the activation layer; is the weight between the l-th layer's nth unit and the i-th unit of the previous layer; is the bias term; is a nonlinear activation function; is the pooling function. Simulator data has characteristics such as randomness and time-variability. After constructing the feature matrix, the feature matrix is input into the CNN to successively extract the attribute features of the data, that is, horizontally traversing the two-dimensional feature matrix. After using CNN to extract data attribute features, it is necessary to extract temporal features; the paper uses an LSTM network to capture temporal features, that is, vertically traversing the two-dimensional feature matrix.
In Section 3.2.2, we supplemented the Bayesian optimization process:
To improve the efficiency and stability of LSTM model hyperparameter selection, this paper employs Bayesian optimization to search for the model's key parameters. Unlike grid search and random search, Bayesian optimization can leverage existing parameter combinations and their model error results, continuously update the surrogate model, and guide the next round of parameter selection.
Let the historical sample set that has been evaluated before the t-th iteration be:
Here, represents the i-th combination of hyperparameters, and represents the model loss value or validation error corresponding to that set of parameters. The hyperparameters in this paper mainly include the number of LSTM hidden layer units, initial learning rate, L2 regularization coefficient, number of training iterations, and so on. The model optimization objective can be expressed as:
Here, Ωrepresents the hyperparameter search space, and represents the optimal hyperparameter combination that minimizes the model error.
In each iteration, Bayesian optimization first builds or updates a Gaussian process surrogate model based on the historical sample set :
Here, is the mean function, and is the covariance function. This surrogate model is used to estimate the posterior distribution of the objective function values corresponding to different combinations of hyperparameters. Subsequently, the next set of candidate hyperparameters is selected using the expected improvement function:
Here, represents the optimal objective function value obtained from the current historical samples.
It can be seen that Bayesian optimization can use previous parameter combinations and their model errors as prior information, balancing the use of already good parameter regions and exploring uncertain parameter regions, thereby improving the efficiency of hyperparameter search.
During the model training process, the test set is only used for the final model performance evaluation and does not participate in hyperparameter selection. If a separate validation set is used, the validation set error is taken as Lθ; if the sample size is limited, the internal cross-validation error of the training set is used as Lθ, to reduce the impact of a single data split on optimization results. When the objective function value no longer decreases significantly after several consecutive iterations, the Bayesian optimization process is considered basically converged, and the currently optimal hyperparameter combination is taken as the final model parameters.
In Section 5.1, we corrected the normalization formulas for positive and negative indicators:For positive indicators, the larger the indicator value, the better the evaluation effect, and its standardization formula is:
For negative indicators, the smaller the indicator value, the better the evaluation effect, and its standardization formula is:
Here, represents the original value of the j-th indicator for the i-th evaluation object, and represents the standardized value of the indicator. In this study, R² and Rank Correlation are considered positive indicators, while MSE, RMSE, NRMSE, ErrorMean, and ErrorStd are considered negative indicators. After standardization, all indicators are converted into a unified direction where 'the larger the value, the better the evaluation performance,' providing a basis for subsequent CRITIC weighting and comprehensive evaluation calculations.
In Sections 5.2 and 5.4, the formulas for grey relational analysis, comprehensive weight calculation, and comprehensive evaluation index calculation were uniformly checked and reformatted, with the meanings of the formula variables and calculation logic supplemented:
5.2. Evaluation parameter weighting based on entropy method
The entropy method measures the information content of indicator data through their degree of dispersion and determines the weights of the indicators accordingly. The greater the difference in indicators, the higher the information content and the larger the weight; the smaller the difference in indicators, the weaker the distinguishing ability and the smaller the weight.
- Data standardization. Suppose there are n evaluation subjects and m evaluation indicators. First, standardize the positive and negative indicators.
- The calculation of the entropy value of the i-th evaluation index at level j in the simulator test evaluation is shown in the formula:
In the formula:, represents the probability of the i-th indicator at the j-th level, and n represents the number of states of a single indicator.
- The entropy calculation of the weight of the i-th indicator is shown in the formula:
According to the properties of entropy, it can be obtained that:,,Finally, the set of evaluation indicator weights can be obtained:.
According to the calculation method, substituting the model data yields the entropy method weight calculation results as shown in the table 4:
|
Evaluation item |
Entropy value |
Difference coefficient |
Entropy weight |
|
MMS_ECG |
0.9321 |
0.0679 |
0.2353 |
|
MMS_EDA |
0.8510 |
0.1490 |
0.5166 |
|
MMS_EMG |
0.9285 |
0.0715 |
0.2481 |
Table 4. Entropy weight calculation results.
Weights of MMS_ECG, MMS_EDA, and MMS_EMG were calculated using the entropy method, and their weight values are 0.235, 0.517, and 0.248, respectively. The weights among the items are about 0.333, which is relatively uniform.
5.3:
|
Appraisal items |
Correlation degree |
Rank |
|
MMS_ECG |
0.669 |
3 |
|
MMS_EDA |
0.870 |
1 |
|
MMS_EMG |
0.737 |
2 |
Table 4. Correlation result
Revision notes: Specific modification locations can be found on page 10, lines 304-319; page 12, lines 366-400; page 16, lines 486-493; pages 17-18, lines 516-539.
Comment 2:Similarly, you named figures as Tables: can you check it?
Authors’ reply: Thank you very much to the experts for pointing out this issue. We appreciate the experts' reminder. We have rechecked all figure titles, numbering, and text references throughout the manuscript. Indeed, the original manuscript mistakenly labeled some images as "Table," such as diagrams of experimental equipment, sensor wearing schematics, model fitting graphs, error comparison charts, and mean deviation graphs. These items are essentially images or graphical results, and therefore should be uniformly labeled as "Figure" rather than "Table".
Authors’ revision:The following revisions have been made in the revised draft:
|
|
|
|
(a) |
(b) |
Figure 2. Driving simulator for testing, True car cab (a), self-built cockpit (b).
Figure 3. Diagram of wearing wireless EDA, ECG, and EMG sensors.
|
|
|
|
|
|
(a) |
(b) |
(c) |
|
|
|
|
||
|
(e) |
(f) |
||
Figure 4. Experimental dynamic and static scene diagram: (a) No signal and no occlusion conflict port, (b) No signal and occlusion conflict port, (c) No signal and low flow conflict port, (d) No signal and heavy flow conflict port, (e) Dynamic virtual scene.
|
Driver No. |
Gender |
Driving Experience / Mileage |
Age |
Accident in the Past Three Years (Y/N) |
Traffic Violation in the Past Three Years (Y/N) |
|
1 |
Female |
9 years / 50,000–60,000 km |
35 |
N |
N |
|
2 |
Male |
24 years / 250,000 km |
56 |
N |
Y |
|
3 |
Male |
9 years / 100,000 km |
38 |
N |
Y |
|
4 |
Male |
9 years / 100,000 km |
38 |
N |
N |
|
5 |
Male |
10 years / 80,000 km |
43 |
N |
Y |
|
6 |
Male |
4 years / 30,000 km |
28 |
N |
Y |
|
7 |
Male |
8 years / 140,000 km |
41 |
Y |
Y |
|
8 |
Female |
7 years / 10,000 km |
41 |
N |
N |
|
9 |
Female |
10 years / 50,000 km |
38 |
N |
N |
|
10 |
Male |
24 years / 200,000 km |
46 |
N |
Y |
|
11 |
Female |
8 years / 70,000 km |
39 |
N |
Y |
|
12 |
Female |
6 years / 50,000 km |
48 |
N |
Y |
|
13 |
Female |
12 years / 100,000 km |
39 |
N |
Y |
|
14 |
Female |
5 years / 50,000 km |
32 |
N |
N |
|
15 |
Female |
14 years / 60,000 km |
33 |
N |
Y |
|
16 |
Female |
14 years / 60,000 km |
38 |
N |
Y |
|
17 |
Male |
15 years / 150,000 km |
37 |
N |
Y |
|
18 |
Male |
5 years / 60,000 km |
24 |
N |
N |
|
19 |
Male |
4 years / 100,000 km |
23 |
N |
N |
|
20 |
Male |
3 years / 10,000 km |
23 |
N |
N |
|
21 |
Male |
20 years / 300,000 km |
51 |
N |
N |
|
22 |
Male |
3 years / 50,000 km |
33 |
N |
N |
|
23 |
Male |
4 years / 50,000 km |
23 |
N |
Y |
|
24 |
Male |
3 years / 20,000 km |
24 |
N |
N |
|
25 |
Male |
3 years / 20,000 km |
24 |
N |
N |
|
26 |
Male |
4 years / 80,000 km |
22 |
N |
N |
|
27 |
Male |
3 years / 20,000 km |
21 |
N |
N |
Table 1. Driver detailed information
Figure 5. The schematic flow of evaluation model.
Figure 6. Bayesian hyperparameter optimization algorithm.
|
|
|
MSE |
RMSE |
NRMSE |
Rank Correlation |
Error Mean |
Error Std |
|
EMG |
0.7122 |
212.2497 |
14.5688 |
-0.9109 |
0.7792 |
-1.8282 |
14.4826 |
|
ECG |
0.9739 |
358.0546 |
18.9223 |
-0.3277 |
0.9054 |
-1.8729 |
18.8672 |
|
EDA |
0.9662 |
0.0013 |
0.0360 |
0.0033 |
0.9775 |
-0.0046 |
0.0358 |
Table 2. The fitting value at the intersection of predicted value and original value.
|
|
|
|
|
(a) |
(b) |
(c) |
Figure 7. Comparation of R2: (a) EMG ,(b) ECG, (c) EDA.
|
|
|
|
|
(a) |
(b) |
(c) |
Figure 8. Error comparison diagram: (a) EMG, (b) ECG, (c) EDA.
|
|
|
|
|
(a) |
(b) |
(c) |
Figure 9. Comparation of Rank Correlation: (a) EMG, (b) ECG, (c) EDA.
|
|
|
|
|
(a) |
(b) |
(c) |
Figure 10. Comparison of mean deviation: (a) EMG, (b) ECG, (c) EDA.
|
|
|
|
|
(a) |
(b) |
(c) |
Figure 11. Optimization effect comparison (a) EMG (b) ECG (c) EDA.
|
Items |
Variability of indicators |
Conflict of indicators |
Information content |
Weight |
|
MMS_ECG |
0.108 |
1.976 |
0.213 |
14.51% |
|
MMS_EDA |
0.470 |
1.998 |
0.939 |
64.06% |
|
MMS_EMG |
0.153 |
2.048 |
0.314 |
21.43% |
Table 3. Weight calculation results of CRITIC.
|
Evaluation item |
Entropy value |
Difference coefficient |
Entropy weight |
|
MMS_ECG |
0.9321 |
0.0679 |
0.2353 |
|
MMS_EDA |
0.8510 |
0.1490 |
0.5166 |
|
MMS_EMG |
0.9285 |
0.0715 |
0.2481 |
Table 4. Entropy weight calculation results.
|
Appraisal items |
Correlation degree |
Rank |
|
MMS_ECG |
0.669 |
3 |
|
MMS_EDA |
0.870 |
1 |
|
MMS_EMG |
0.737 |
2 |
Table 5. Correlation result
Figure 12. CRITIC-Entropy method weight diagram.
|
Title 1 |
|
RankCorrelation |
w |
ξ |
|
EMG |
71.22% |
77.92% |
80.08% |
12.71% |
|
ECG |
97.39% |
90.54% |
73.57% |
8.16% |
|
EDA |
96.62% |
97.75% |
64.90% |
79.12% |
Table 6. Model prediction error evaluation index summary Table.
Comment 3:In addition, can put some summary text between title level 1 and title level 2 (e.g., section 3 and sub-section 3.1)? I think this can help the reader…
Authors’ reply: We appreciate the expert's constructive suggestions. In the original manuscript, the primary headings are followed directly by secondary headings, lacking introductory text for the chapters, which can indeed affect the reader's understanding of the chapter logic. Based on this feedback, we have added introductory paragraphs after the headings in Chapters 2, 3, 4, and 5 to explain the research objectives, main content, and logical relationships of each chapter.
Authors’ revision:The revised draft has added the following chapter introduction content.
Add after Chapter 2 title:This chapter mainly introduces the driving simulation experimental platform, the design of traffic conflict scenarios, data collection methods, and the selection criteria for experimental subjects. To ensure the comparability of experimental results, this study constructs a standardized unsignalized intersection conflict scenario and simultaneously collects vehicle operation parameters and drivers' physiological and psychological data, providing a foundation for subsequent model development.
Add after Chapter 3 title:This chapter constructs a Mul-Bayes-LSTM evaluation model based on the coupling of driving behavior and physiological-psychological responses. The model takes vehicle operation parameters, road environment parameters, and conflict scenario parameters as inputs, and EDA, ECG, and EMG as outputs. It extracts attribute correlations and temporal dynamic features through CNN-LSTM, and uses Bayesian optimization to improve the efficiency of model parameter selection.
Add after Chapter 4 title:This chapter analyzes the prediction performance and optimization effect of the Mul-Bayes-LSTM model. This paper selects indicators such as (R^2), Rank Correlation, MSE, RMSE, NRMSE, ErrorMean, and ErrorStd to evaluate the model's performance from the perspectives of fitting degree, correlation, and error distribution, and further explains the differences in EDA, ECG, and EMG results.
Add after Chapter 5 title:Based on the model evaluation results, this chapter conducts a comprehensive assessment of the validity of the driving simulator. Considering the differences in information content, variability, and indicator contribution of EDA, ECG, and EMG, this study uses the CRITIC method, entropy weight method, and grey relational analysis to determine the indicator weights and form a comprehensive evaluation index.
Revision notes: Specific modification locations can be found on page 4, lines 179–184; page 8, lines 259–264; page 13, lines 402–406; page 16, lines 467–471 of the revised draft.
Comment 4:On page 2, at the beginning of the section, the problem is also when you use the driving simulator with subjects, because they know it is not real: this can affect their behavior.
Authors’ reply:We sincerely appreciate the expert for raising this important point. We fully agree that driving simulator experiments are not entirely equivalent to real-world driving. Since participants are aware that the current experimental environment is not real traffic, their risk perception, speed judgment, braking response, steering operations, and level of psychological tension may all differ from those under actual driving conditions.
It should be noted that the traffic conflict scenarios designed in this study carry certain risk characteristics. If the same conditions were tested with real vehicles on actual roads, there could be significant safety risks and restrictions related to road test approvals. Therefore, we chose to construct controlled-risk conflict scenarios in a driving simulator. The purpose of this study is not to assume that a simulator can fully replace real-world driving, but to evaluate whether a driving simulator can elicit stable, interpretable, and behaviorally meaningful driver responses under safe, controllable, and repeatable experimental conditions.
Accordingly, this study did not rely solely on explicit driving behavior indicators such as speed, braking, and trajectory, but also introduced physiological and psychological measures including EDA, ECG, and EMG to analyze drivers' internal perception and physiological-psychological responses in simulated conflict scenarios. By combining driving behavior data with physiological and psychological indicators, this study aims to evaluate the human factors response reproducibility of the driving simulator at both the explicit behavior and internal physiological response levels.
Therefore, the results of this study should be understood as an evaluation of the simulator's internal validity under controlled experimental conditions, rather than proof that simulated driving is completely equivalent to real-world driving. We again thank the expert for the reminder; we have carefully considered this issue in the definition of research objectives and interpretation of results.
Authors’ revision:In response to the problem of differences between driving simulators and real driving, we revised the introduction:If the risk conflict scenarios designed in this article were to be directly tested with real vehicles on actual roads, it could involve high safety risks and road testing approval restrictions. Therefore, this article chooses to construct controlled conflict scenarios in a driving simulator, with the aim of evaluating the simulator's ability to reproduce drivers' human factor responses under safe, controllable, and repeatable conditions.
Modification Note: The specific modification locations can be found on page 2 of the revised draft, lines 47–52.
Comment 5:On page 4, towards the end, how did you select these values of speed? Can they be dynamic, meaning they can be changed dynamically?
Authors’ reply:We appreciate the expert's valuable comments on the speed parameter settings for conflict vehicles. The original manuscript did not provide sufficient explanation for the basis of the speed values, which may affect readers' understanding of the rationality of the experimental scenario design. In this study, the speed settings for conflict vehicles mainly refer to the relevant requirements for operating speeds on urban roads and intersections as specified in the "Urban Road Design Code" (CJJ 37-2012) and the "Code for Design of Urban Road Intersections" (CJJ 152-2010), and are set in conjunction with the urban road unsignalized intersection conflict scenarios constructed in this paper. In this research, the speeds of conflict vehicles are set as fixed values primarily to ensure that different participants face the same external traffic stimuli during the experiment, thereby improving the consistency, repeatability, and comparability of the experimental conditions. It should be noted that these speed values are not fixed constraints in the model or evaluation framework, but are reference parameters within the experimental scenarios of this study. In subsequent research or other experimental scenarios, the speed of conflict vehicles can be dynamically adjusted according to the research objectives, for example, being set in accelerating, decelerating, or randomly perturbed forms, to further analyze changes in driving behavior and physiological and psychological responses under different conflict intensities. According to the expert's suggestion, we have supplemented the explanation of the basis for selecting speed parameters and their potential for dynamic adjustment in Section 2.1 of the revised manuscript.
Authors’ revision:The following content has been added at the end of Section 2.1 in the text:The speed parameters of conflicting vehicles mainly refer to the relevant requirements for urban road and intersection operating speeds in the 'Urban Road Design Code' (CJJ 37-2012) and the 'Urban Road Intersection Design Specifications' (CJJ 152-2010), and are set in conjunction with the unsignalized intersection conflict scenarios constructed in this study. A fixed speed is adopted herein mainly to ensure that different participants are exposed to consistent experimental stimuli, thereby improving experimental repeatability and data comparability; in subsequent studies, the speed of conflicting vehicles can be set to dynamically change, such as accelerating, decelerating, or random perturbations, according to the research objectives.
Modification Note: The specific modification locations can be found on page 7 of the revised draft, lines 222–230.
Comment 6:On page 5, section 2.2., what is the age of users?
Authors’ reply:Thank you very much for the expert's reminder. The original manuscript only indicated the number of participants, their gender composition, and driving experience, but did not provide the age range and average age, which is indeed incomplete. We have supplemented the participants' age information in Section 2.2.
Authors’ revision:Revised draft Section 2.2 adds content:
According to the subject information statistics table, a total of 27 valid subjects were included in this experiment. The age of the subjects ranged from 21 to 56 years, with an average age of 34.74 years and a standard deviation of 9.65 years, covering the young to middle-aged driver population. Specific details on subjects' gender, driving experience, mileage, accident history, and traffic violations are shown in Table 1.
|
Driver No. |
Gender |
Driving Experience / Mileage |
Age |
Accident in the Past Three Years (Y/N) |
Traffic Violation in the Past Three Years (Y/N) |
|
1 |
Female |
9 years / 50,000–60,000 km |
35 |
N |
N |
|
2 |
Male |
24 years / 250,000 km |
56 |
N |
Y |
|
3 |
Male |
9 years / 100,000 km |
38 |
N |
Y |
|
4 |
Male |
9 years / 100,000 km |
38 |
N |
N |
|
5 |
Male |
10 years / 80,000 km |
43 |
N |
Y |
|
6 |
Male |
4 years / 30,000 km |
28 |
N |
Y |
|
7 |
Male |
8 years / 140,000 km |
41 |
Y |
Y |
|
8 |
Female |
7 years / 10,000 km |
41 |
N |
N |
|
9 |
Female |
10 years / 50,000 km |
38 |
N |
N |
|
10 |
Male |
24 years / 200,000 km |
46 |
N |
Y |
|
11 |
Female |
8 years / 70,000 km |
39 |
N |
Y |
|
12 |
Female |
6 years / 50,000 km |
48 |
N |
Y |
|
13 |
Female |
12 years / 100,000 km |
39 |
N |
Y |
|
14 |
Female |
5 years / 50,000 km |
32 |
N |
N |
|
15 |
Female |
14 years / 60,000 km |
33 |
N |
Y |
|
16 |
Female |
14 years / 60,000 km |
38 |
N |
Y |
|
17 |
Male |
15 years / 150,000 km |
37 |
N |
Y |
|
18 |
Male |
5 years / 60,000 km |
24 |
N |
N |
|
19 |
Male |
4 years / 100,000 km |
23 |
N |
N |
|
20 |
Male |
3 years / 10,000 km |
23 |
N |
N |
|
21 |
Male |
20 years / 300,000 km |
51 |
N |
N |
|
22 |
Male |
3 years / 50,000 km |
33 |
N |
N |
|
23 |
Male |
4 years / 50,000 km |
23 |
N |
Y |
|
24 |
Male |
3 years / 20,000 km |
24 |
N |
N |
|
25 |
Male |
3 years / 20,000 km |
24 |
N |
N |
|
26 |
Male |
4 years / 80,000 km |
22 |
N |
N |
|
27 |
Male |
3 years / 20,000 km |
21 |
N |
N |
Table1. Driver detailed information
Modification Note: The specific modification locations can be found on page 7 of the revised draft, lines 231–240.
Comment 7:Same page, at the beginning of section 3.1, can you detail a little bit more about the problem of “weight influence”?
Authors’ reply:We appreciate the expert pointing out that this part of the statement was not clear enough. Upon re-examination, we found that the expression "weight influence" in the original manuscript was not accurate and did not clearly convey the model issues discussed in this paper. The original intention was not to discuss the influence of subjects' weight on driving behavior or model results, but to indicate that traditional RNNs may have issues such as insufficient long-term dependency capture, vanishing gradients, or exploding gradients when processing longer time series data. Following the expert's suggestion, we have revised the beginning of Section 3.1 by removing the inaccurate expression "weight influence" and adding an explanation of the limitations of traditional RNNs in long-sequence modeling, as well as the reason for using the LSTM model in this paper. After revision, this section can more clearly explain why LSTMs are suitable for modeling driving behavior and psychophysiological time series data.
Authors’ revision:At the beginning of Section 3.1, replace the original wording with:Traditional recurrent neural networks can handle time series data, but when dealing with longer time spans and complex sequence data, they are prone to issues such as insufficient capture of long-term dependencies, vanishing gradients, or exploding gradients. In driving simulation experiments, drivers' behavioral responses and physiological-psychological responses exhibit obvious continuity and temporal dependence. The driving response at a certain moment is influenced not only by the current traffic conflict stimulus but also possibly by previous changes in speed, steering operations, braking behavior, and psychological arousal states. Therefore, it is necessary to use models with stronger temporal memory capabilities to model driving behavior and physiological-psychological data.
Modification Note: The specific modification locations can be found on page 8 of the revised draft, lines 266–275.
Comment 8:On page 7, at the end of section 3.2.1, usually a dataset is divided in three parts: training, checking / validation and testing sets; have you considered training and checking together?
Authors’ reply:We appreciate the expert's suggestions regarding dataset splitting and validation strategies. The paper only mentions using 80% of the samples as the training set and 20% as the test set, without fully explaining the validation process and methods for testing model stability. Considering the relatively small sample size in our experiments, forcibly re-dividing the training, validation, and test sets could further reduce the number of training samples. Therefore, in the revised manuscript, while retaining the original 80% training set and 20% test set split, we have added the use of K-fold cross-validation as a method to test model stability.
Authors’ revision:Add at the end of section 3.2.1:The test set is used only for final evaluation and does not participate in model training or hyperparameter selection. Considering the relatively limited sample size, this paper further uses the K-fold cross-validation method to perform supplementary testing of model stability, in order to reduce the impact of a single random split on the model evaluation results.
Add at the end of Section 4.2:To further verify the adaptability of the Mul-Bayes-LSTM model to driving simulator data, this study sets K=2 to K=10 for cross-validation and calculates the R² values under different folds to analyze model stability and overfitting risk. As shown in Figure 13, the driving simulation dataset is divided into 2 to 10 folds, and the R² value for each fold is calculated to prevent overfitting. The results indicate that when K=6, the model achieves a relatively high and stable R² value, suggesting that the model has good stability and generalization performance on the current driving simulator dataset.
Figure 12. to 10-fold cross validation R2
Modification Note: The specific modification locations can be found on page 10 of the revised draft, lines 304–319;Page 15, lines 432-440.
Comment 9:Same page, towards the end, how do you consider Bayesian parameters with previous parameters’ information?
Authors’ reply:We thank the expert for pointing out the insufficient explanation of the Bayesian optimization process. Although the original manuscript provided the formulas for the Gaussian process and the acquisition function, it did not explain how each iteration uses historical parameter information. We have added, in Section 3.2.2, an explanation of the process of using the set of historical samples, the objective function, the surrogate model update, and the expected improvement function to select the next set of hyperparameters, making the Bayesian optimization process more complete.
Authors’ revision:Supplementary content in Section 3.2.2:
To improve the efficiency and stability of LSTM model hyperparameter selection, this paper employs Bayesian optimization to search for the model's key parameters. Unlike grid search and random search, Bayesian optimization can leverage existing parameter combinations and their model error results, continuously update the surrogate model, and guide the next round of parameter selection.
Let the historical sample set that has been evaluated before the t-th iteration be:
Here, represents the i-th combination of hyperparameters, and represents the model loss value or validation error corresponding to that set of parameters. The hyperparameters in this paper mainly include the number of LSTM hidden layer units, initial learning rate, L2 regularization coefficient, number of training iterations, and so on. The model optimization objective can be expressed as:
Here, Ωrepresents the hyperparameter search space, and represents the optimal hyperparameter combination that minimizes the model error.
In each iteration, Bayesian optimization first builds or updates a Gaussian process surrogate model based on the historical sample set :
Here, is the mean function, and is the covariance function. This surrogate model is used to estimate the posterior distribution of the objective function values corresponding to different combinations of hyperparameters. Subsequently, the next set of candidate hyperparameters is selected using the expected improvement function:
Here, represents the optimal objective function value obtained from the current historical samples.
It can be seen that Bayesian optimization can use previous parameter combinations and their model errors as prior information, balancing the use of already good parameter regions and exploring uncertain parameter regions, thereby improving the efficiency of hyperparameter search.
During the model training process, the test set is only used for the final model performance evaluation and does not participate in hyperparameter selection. If a separate validation set is used, the validation set error is taken as Lθ; if the sample size is limited, the internal cross-validation error of the training set is used as Lθ, to reduce the impact of a single data split on optimization results. When the objective function value no longer decreases significantly after several consecutive iterations, the Bayesian optimization process is considered basically converged, and the currently optimal hyperparameter combination is taken as the final model parameters.
Revision Notes: The specific modification locations can be found on pages 12-13 of the revised draft, lines 366-400.
Comment 10:On page 11, formulas (10) and (11), you wrote twice that “The positive indicator is shown in formula …”: is this correct?
Authors’ reply:Thanks to the expert for pointing out the error at this point. Upon checking, formula (10) should be the standardization of a positive indicator, and formula (11) should be the standardization of a negative indicator. In the original manuscript, both were written as 'positive indicator,' which is inaccurate. We have revised the wording before formula (11) to 'negative indicator'.
Authors’ revision:Explanation before and after modifying formulas:
The positive indicator is shown in formula :
The negative indicator is shown in formula :
Revision notes: For the specific modification location, see page 16, line 486 of the revised draft.
Acknowledgements at the End
Once again, I sincerely thank the reviewers for their valuable suggestions for revisions. All issues have been addressed one by one. The revised paper is logically rigorous, content-wise complete, and well-supported by experiments. I respectfully request the experts to review it and consider it for acceptance!
Sincerely,
Respectfully yours!
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsNONE
Comments on the Quality of English LanguageModerate changes are needed.
Reviewer 2 Report
Comments and Suggestions for AuthorsDear authors,
Thanks a lot for addressing all my comments in detail.
My only concern is about the age of the subjects. which ranged from 21 to 56 years, with an average age of 34.74 years. It is true that it covers the young and middle-aged driver’s population, but it is also quite large. For the future, it could be interesting to investigate the influence and the impact of gender and age on your findings.

