Review Reports
- Sungmo Ku 1 and
- Jinho Lee 2,*
Reviewer 1: Anonymous Reviewer 2: Anonymous Reviewer 3: Anonymous
Round 1
Reviewer 1 Report
Comments and Suggestions for Authors(1) The abstract of this manuscript requires sufficient data support, at least the main quantitative findings related should be listed.
(2) The timing module in this manuscript only compares the LSTM series and does not introduce cutting-edge efficient sequence models such as Mamba, xLSTM, and timing Transformer.
(3) In this manuscript, multi task loss weights are fixed without using adaptive dynamic weighting, which can lead to uneven task difficulty.
Author Response
We thank the reviewer for the constructive comments on quantitative reporting, sequence-model comparisons, and adaptive task weighting. These suggestions guided additional experiments and revisions that strengthened the evaluation and presentation of the framework.
Q1. The abstract of this manuscript requires sufficient data support, at least the main quantitative findings related should be listed.
A1. Thank you for this comment. We revised the abstract to include the main quantitative findings, including the highest Full Match Accuracy of 91.61% and trajectory accuracy of 99.97% under fixed weighting. The interpretation of perception performance relative to the majority-class baseline is provided in the Results section.
Q2. The timing module in this manuscript only compares the LSTM series and does not introduce cutting-edge efficient sequence models such as Mamba, xLSTM, and timing Transformer.
A2. Thank you for this suggestion. We added Transformer-AE, xLSTM-AE, and Mamba-AE to the experiments, expanding the evaluation to six trajectory encoders and 24 text–trajectory encoder combinations. Their architectures and results are now included in the revised manuscript.
Q3. In this manuscript, multi task loss weights are fixed without using adaptive dynamic weighting, which can lead to uneven task difficulty.
A3. We added uncertainty-based adaptive task weighting and compared it with fixed weighting across all encoder combinations. The revised manuscript presents the weighting formulation and analyzes its effects on individual-task and joint classification performance.
Reviewer 2 Report
Comments and Suggestions for AuthorsThis study is generally well-organized. The topic is relevant and timely, and the proposed multimodal approach that combines text embeddings with trajectory representations is interesting. However, the manuscript has several significant limitations that need to be addressed before it can be considered for publication. I recommend major revision.
- The decision-disturbance label is based solely on collision occurrence, which represents only the final outcome and ignores the continuous spectrum of pre-collision interaction severity. This binary labeling approach oversimplifies the complex nature of hazardous driving situations and limits the practical utility of the framework for safety assessment.
- The simulation dataset is generated from a highly constrained intersection environment with only three trajectory types and a limited set of environmental conditions. This controlled setup may artificially inflate classification performance and raises serious questions about the generalizability of the proposed framework to diverse real-world driving scenarios.
- The control disturbance category is discussed conceptually but completely excluded from quantitative evaluation, yet the paper claims to jointly consider disturbances affecting perception, decision-making, and control. This discrepancy between the stated contribution and the actual implementation undermines the core claim of the framework..
- The latest advances on autonomous driving should be reviewed for completeness. For instance, Learning safety-critical scenarios from real-world pre-crash data for autonomous driving safety validation; Multi-objective autonomous eco-driving strategy: A pathway to future green mobility; The investigation of reinforcement learning-based end-to-end decision-making algorithms for autonomous driving on the road with consecutive sharp turns.
- The LSTM Autoencoder is trained solely on the SinD dataset while the classification model is evaluated on simulation data, but the paper does not adequately address the domain gap between real-world and simulated vehicle motion distributions. The reported transferability metrics are insufficient to demonstrate robust cross-domain generalization.
- The text encoder selection experiment compares four models but does not justify why these specific encoders were chosen or how their representational capabilities differ for structured OpenSCENARIO data. The computational efficiency analysis suggests MiniLM is fastest, but the corresponding classification performance is not the highest across all tasks.
- The paper claims to address the limitation of existing layer-based and ontology-based classification approaches, but the proposed framework still relies on predefined trajectory types and disturbance labels rather than learning a more flexible scenario representation. The novelty of the approach is therefore incremental rather than transformative.
Author Response
We thank the reviewer for the thoughtful comments. We recognize the concerns regarding the labeling criteria, evaluation scope, generalizability, and contribution of the framework. We have carefully considered these points and revised the manuscript to clarify what the experiments demonstrate and where limitations remain.
Q1. The decision-disturbance label is based solely on collision occurrence, which represents only the final outcome and ignores the continuous spectrum of pre-collision interaction severity. This binary labeling approach oversimplifies the complex nature of hazardous driving situations and limits the practical utility of the framework for safety assessment.
A1. We agree that collision occurrence alone cannot capture continuous pre-collision risk. The supplementary indicators in Appendix A remain under investigation and require positional and motion information from both the Ego and surrounding vehicles. To explicitly represent these interactions when predicting the revised risk-based labels, we would need to extend the current Ego-only motion input to include relative positions and surrounding-vehicle trajectories, modify the model architecture accordingly, and perform retraining and validation. These changes could not be adequately validated within the revision period, so experiments using the new risk-based labels are not included in this revision. We have clarified the limitation of collision-based labeling and identified interaction-aware risk assessment as future work.
Q2. The simulation dataset is generated from a highly constrained intersection environment with only three trajectory types and a limited set of environmental conditions. This controlled setup may artificially inflate classification performance and raises serious questions about the generalizability of the proposed framework to diverse real-world driving scenarios.
A2. We have acknowledged that the constrained scenario configurations may make some classification tasks easier. We also compared validation performance with results on a separately generated test dataset using modified temporal and environmental parameters and reported the perception majority-class baseline. We clarify that this evaluation concerns parameter changes within the same scenario-generation framework and does not establish generalization to diverse road environments or real-world driving.
Q3. The control disturbance category is discussed conceptually but completely excluded from quantitative evaluation, yet the paper claims to jointly consider disturbances affecting perception, decision-making, and control. This discrepancy between the stated contribution and the actual implementation undermines the core claim of the framework.
A3. We revised the Related Work, labeling description, model architecture, and Conclusion to align the stated contribution with the actual evaluation scope. The quantitative targets are limited to trajectory type, perception disturbance, and decision disturbance; control disturbances are discussed only conceptually. Although road friction is included in the environmental information, the available data are insufficient to establish its relationship with control-disturbance occurrence. Establishing such labels would require additional information linking driving commands to vehicle responses. Control disturbances are therefore excluded from classifier training and evaluation.
Q4. The latest advances on autonomous driving should be reviewed for completeness. For instance, Learning safety-critical scenarios from real-world pre-crash data for autonomous driving safety validation; Multi-objective autonomous eco-driving strategy: A pathway to future green mobility; The investigation of reinforcement learning-based end-to-end decision-making algorithms for autonomous driving on the road with consecutive sharp turns.
A4. Thank you for the suggested references. We have included the study on learning safety-critical scenarios from real-world pre-crash data in the Introduction to broaden the discussion of data-driven scenario generation. The other two suggested studies focus on driving-policy optimization, whereas the present work addresses scenario classification. We recognize their relevance to the broader autonomous-driving context, but they are not included in the current revision.
Q5. The LSTM Autoencoder is trained solely on the SinD dataset while the classification model is evaluated on simulation data, but the paper does not adequately address the domain gap between real-world and simulated vehicle motion distributions. The reported transferability metrics are insufficient to demonstrate robust cross-domain generalization.
A5. We agree that domain metrics alone cannot demonstrate robust cross-domain generalization. In the revised manuscript, the SinD–simulation silhouette score and MMD² are interpreted as descriptive measures of distributional differences. We explicitly state that lower values do not directly establish domain alignment or improved generalization. Classification claims are restricted to the evaluated simulation configurations, while domain adaptation and joint representation learning are identified as future research directions.
Q6. The text encoder selection experiment compares four models but does not justify why these specific encoders were chosen or how their representational capabilities differ for structured OpenSCENARIO data. The computational efficiency analysis suggests MiniLM is fastest, but the corresponding classification performance is not the highest across all tasks.
A6. The four encoders were selected to compare scenario-classification performance across models with different sizes and embedding dimensions, as explained in Section 3.3.2 and Table 4. The revised Results section focuses on classification performance across encoder combinations rather than treating computational speed as evidence of classification quality. These experiments evaluate the encoders within the complete classification pipeline; they do not directly isolate the preservation of individual OpenSCENARIO attributes or relationships.
Q7. The paper claims to address the limitation of existing layer-based and ontology-based classification approaches, but the proposed framework still relies on predefined trajectory types and disturbance labels rather than learning a more flexible scenario representation. The novelty of the approach is therefore incremental rather than transformative.
A7. We agree that the framework relies on predefined labels and does not automatically discover new scenario categories. We have therefore clarified the contribution as an integrated framework for jointly representing trajectory and disturbance labels, supported by a systematic evaluation. The revised discussion emphasizes its complementary role alongside layer-based and ontology-based approaches. We also added evaluation on separately generated test scenarios with modified parameters to examine performance beyond the training parameter configuration.
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript presents an interesting multi-task framework for classifying autonomous-driving scenarios by combining vehicle trajectories with structural information extracted from OpenSCENARIO files. The integration of real-world trajectory information from SinD with simulation-based scenarios is particularly valuable, and the manuscript is generally well structured. However, I believe that several aspects should be strengthened before publication.
My main concern relates to the very high classification performance and the construction of the disturbance labels. Perception disturbance is directly defined from environmental conditions such as rain, snow, fog, or illumination, while these same characteristics are represented in the OpenSCENARIO information used as model input. The authors should therefore demonstrate that the model is learning meaningful patterns rather than simply recovering the rules used to construct the labels. A relatively simple ablation comparing trajectory-only, scenario-only, and combined inputs would considerably strengthen the paper.
Similarly, decision disturbance is defined through collision occurrence. While this provides a clear criterion, it represents a rather restrictive interpretation of disturbance. The authors already discuss indicators such as TTC, headway, TCPA, and DCPA, and some consideration of near-collision or continuous risk conditions would provide a richer assessment of safety-critical scenarios.
I would also encourage the authors to reconsider the validation strategy. Since a large number of scenarios are systematically generated from combinations of a limited number of parameters, randomly separating them into training and validation samples may result in highly similar scenarios appearing in both subsets. Testing the framework on genuinely unseen scenario configurations or environmental conditions would provide stronger evidence of generalization.
More generally, the manuscript should better discuss the weight and meaning of results obtained from simulation environments. Simulators are extremely useful for controlled and reproducible experimentation, particularly for safety-critical situations, but transferring results from simulation to real driving conditions remains challenging. The inclusion of SinD partially addresses this issue, although real-world data are used mainly to learn trajectory representations rather than to validate the complete disturbance-classification framework. This distinction should be made clearer and the generalization claims moderated accordingly.
In this context, the authors may consider discussing approaches based on naturalistic driving data, which capture vehicle behavior under actual traffic, infrastructure, environmental, and driver conditions. Some recent studies about Extraction of Naturalistic Driving Patterns with Geographic Information Systems, provides a useful example of how fine-grained real-world driving data can be used to characterize driving patterns across different situations and spatial scales. This perspective could help contextualize both the strengths and limitations of simulation-based validation.
Author Response
We thank the reviewer for the positive assessment of the framework and the constructive comments. We have carefully considered the concerns regarding label construction, input dependence, and generalization, and revised the manuscript to clarify the evidence provided by the experiments and the remaining limitations.
Q1. My main concern relates to the very high classification performance and the construction of the disturbance labels. Perception disturbance is directly defined from environmental conditions such as rain, snow, fog, or illumination, while these same characteristics are represented in the OpenSCENARIO information used as model input. The authors should therefore demonstrate that the model is learning meaningful patterns rather than simply recovering the rules used to construct the labels. A relatively simple ablation comparing trajectory-only, scenario-only, and combined inputs would considerably strengthen the paper.
A1. We clarified that perception-disturbance labels represent environmental conditions rather than directly measured sensor degradation. We added ablations that set either the text or trajectory latent vector to zero in the trained models, as described in Section 4.5, and reported the perception majority-class baseline in Section 4. These experiments assess input dependence rather than compare separately trained unimodal models. We acknowledge that they do not establish learning beyond recovery of the labeling rules.
Q2. Similarly, decision disturbance is defined through collision occurrence. While this provides a clear criterion, it represents a rather restrictive interpretation of disturbance. The authors already discuss indicators such as TTC, headway, TCPA, and DCPA, and some consideration of near-collision or continuous risk conditions would provide a richer assessment of safety-critical scenarios.
A2. We agree that collision occurrence alone cannot capture continuous pre-collision risk. The supplementary indicators in Appendix A remain under investigation and require positional and motion information from both the Ego and surrounding vehicles. To explicitly represent these interactions when predicting the revised risk-based labels, we would need to extend the current Ego-only motion input to include relative positions and surrounding-vehicle trajectories, modify the model architecture accordingly, and perform retraining and validation. These changes could not be adequately validated within the revision period, so experiments using the new risk-based labels are not included in this revision. We have clarified the limitation of collision-based labeling and identified interaction-aware risk assessment as future work.
Q3. I would also encourage the authors to reconsider the validation strategy. Since a large number of scenarios are systematically generated from combinations of a limited number of parameters, randomly separating them into training and validation samples may result in highly similar scenarios appearing in both subsets. Testing the framework on genuinely unseen scenario configurations or environmental conditions would provide stronger evidence of generalization.
A3. We acknowledged that training and validation samples share the same generation settings, limiting the evidence that validation performance provides for generalization. We therefore presented results on a separately generated test dataset with modified temporal and environmental parameters and compared them with validation results. Because the test dataset retains the same intersection and scenario types and shares some parameter values, we restrict the interpretation to performance under modified parameter configurations.
Q4. More generally, the manuscript should better discuss the weight and meaning of results obtained from simulation environments. Simulators are extremely useful for controlled and reproducible experimentation, particularly for safety-critical situations, but transferring results from simulation to real driving conditions remains challenging. The inclusion of SinD partially addresses this issue, although real-world data are used mainly to learn trajectory representations rather than to validate the complete disturbance-classification framework. This distinction should be made clearer and the generalization claims moderated accordingly.
A4. We clarified that SinD is used to train the trajectory Autoencoders, whereas the complete classifier is trained and evaluated using simulation data. Neither the domain-distribution metrics nor the simulation classification results establish real-world disturbance-classification performance. We have accordingly moderated the generalization claims and identified broader evaluation, domain adaptation, and joint representation learning with real-world and simulation data as future research directions.
Q5. In this context, the authors may consider discussing approaches based on naturalistic driving data, which capture vehicle behavior under actual traffic, infrastructure, environmental, and driver conditions. Some recent studies about Extraction of Naturalistic Driving Patterns with Geographic Information Systems, provides a useful example of how fine-grained real-world driving data can be used to characterize driving patterns across different situations and spatial scales. This perspective could help contextualize both the strengths and limitations of simulation-based validation.
A5. Thank you for recommending this reference. We added the study to the scenario-classification subsection of the Related Work, discussing how naturalistic driving patterns can be examined through kinematic measurements and geographic context across different spatial scales. This addition highlights the importance of interpreting vehicle behavior within its road context and helps explain the complementary roles of controlled simulation and naturalistic driving analysis.
Round 2
Reviewer 2 Report
Comments and Suggestions for AuthorsThe revised manuscript is acceptable for publishing.
Reviewer 3 Report
Comments and Suggestions for AuthorsDear Authors,
Thank you for the careful revision. I appreciate the authors’ efforts in addressing my previous comments. The manuscript has improved substantially, and my main concerns have been satisfactorily addressed.
I have no further comments and support publication in its present form.
Best regards,
Reviewer