Next Article in Journal
Uncovering Static–Dynamic Interaction Patterns Between Commercial and Residential Spaces Within Beijing’s Sixth Ring Road Using POI and Human Mobility Trajectory Data
Previous Article in Journal
Two-Hourly Urban Vitality Dynamics and Spatiotemporal Heterogeneity of Built-Environment–Vitality Associations at the Parcel Scale: Evidence from Zhengzhou, China
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Multi-Task Framework for Vehicle Trajectory and Disturbance Classification in Autonomous Driving Scenarios

by
Sungmo Ku
1 and
Jinho Lee
2,*
1
Department of Smart Manufacturing Engineering, Tech University of Korea, Siheung 15073, Gyeonggi-do, Republic of Korea
2
Department of Computer Engineering, Tech University of Korea, Siheung 15073, Gyeonggi-do, Republic of Korea
*
Author to whom correspondence should be addressed.
ISPRS Int. J. Geo-Inf. 2026, 15(10), 443; https://doi.org/10.3390/ijgi15100443
Submission received: 12 August 2026 / Revised: 10 September 2026 / Accepted: 24 September 2026 / Published: 27 September 2026

Abstract

Systematic scenario classification supports the characterization of test conditions in simulation-based autonomous driving validation. This paper proposes a multi-task framework that combines OpenSCENARIO text embeddings with Ego-vehicle motion representations learned by Autoencoders trained on the SinD real-world dataset. The framework jointly predicts trajectory type and binary perception and decision disturbances, defined using environmental conditions and collision occurrence, respectively. We evaluate 24 combinations of four text encoders and six trajectory Autoencoders under fixed and uncertainty-based task weighting. On generated test scenarios with modified temporal and environmental parameters, Qwen3-Embedding-0.6B combined with Mamba-AE achieves the highest Full Match Accuracy of 91.61%, requiring all three labels to be predicted correctly, and trajectory accuracy of 99.97% under fixed weighting. Perception classification limits joint performance, and high validation accuracy does not consistently extend to the modified test conditions. Uncertainty-based weighting provides no consistent improvement, while latent-vector ablations show task-dependent contributions from text and motion representations. These findings support offline classification within the evaluated scenario configurations while identifying limitations in robustness to parameter changes.

1. Introduction

With the rapid advancement of autonomous driving technology, the importance of validating vehicle safety and reliability has grown significantly [1,2,3]. In particular, testing autonomous vehicles in real-world environments requires ensuring safety in hazardous situations, a requirement that is also emphasized in functional safety standards such as ISO 26262 [4]. However, real-world validation entails substantial time and economic costs, and repeated experimentation is inherently difficult. To address these challenges, simulation-based validation methods that replicate real-world environments have been actively studied. Simulators such as CARLA (https://carla.org/ (accessed on 11 August 2026)), esmini (https://github.com/esmini (accessed on 11 August 2026)), and MORAI (https://www.morai.ai/ (accessed on 11 August 2026)) enable repeated reproduction of diverse driving situations with user-defined environment configurations, facilitating efficient generation of driving and accident scenarios across a wide range of spatiotemporal conditions [1,5,6].
However, when simulation-based scenarios are generated without a systematic classification framework, characterizing individual scenarios and analyzing validation results becomes difficult. In particular, identifying vulnerabilities in autonomous driving systems requires structural comparison and systematic categorization of scenarios. Accordingly, there is a clear need for a method that can define and classify scenarios according to consistent criteria. Existing approaches include the layer-based classification framework proposed by the Pegasus project and ontology-based methods. However, since these approaches primarily define scenarios around environmental conditions and objects, they suffer from increasing complexity and limited scalability when new elements are introduced.
Widely used standards in autonomous driving simulation include OpenSCENARIO (https://www.asam.net/standards/detail/openscenario-xml/ (accessed on 11 August 2026)) and OpenDRIVE (https://www.asam.net/standards/detail/opendrive/ (accessed on 11 August 2026)), both proposed by the Pegasus project (https://www.pegasusprojekt.de/en/home (accessed on 11 August 2026)). OpenSCENARIO employs an XML-based structure to define environmental settings, object creation and control, and success/failure conditions within a scenario, and is used to represent dynamic elements in simulation. OpenDRIVE, on the other hand, defines the static environment including road networks, lane structures, and intersection information. The overall structure of these standards is illustrated in Figure 1. Various simulators adopt different map representations based on these standards; for example, CARLA and esmini use OpenDRIVE, while Autoware Foundation (https://github.com/autowarefoundation (accessed on 11 August 2026)) utilizes Lanelet2-based maps [7]. These systems commonly define the space in which vehicles and pedestrians can move, and construct driving routes based on road connectivity.
Previous studies have investigated the use of real-world driving data for scenario extraction, simulation scenario construction, and comparison of simulated and observed vehicle behavior [1,5,8,9,10]. Recent work has also explored the generation of safety-critical scenarios from real-world pre-crash trajectories [11]. These efforts provide foundations for constructing test cases and motivate complementary classification methods for organizing the resulting scenarios. Our prior work classified scenarios using vehicle trajectories [12]. However, trajectory types alone do not describe the environmental conditions or collision-related outcomes associated with a maneuver. This study proposes an integrated framework for classifying intersection scenarios by combining structured OpenSCENARIO information with Ego-vehicle motion representations learned from the SinD dataset. The framework jointly predicts trajectory type, perception-related environmental conditions, and collision-related outcomes used as decision-disturbance labels. These outputs characterize scenario conditions and observed outcomes rather than directly diagnosing failures in individual autonomous driving components.
The contributions of this study are threefold. First, a multi-task classification framework combines trajectory and disturbance labels within a shared representation and prediction pipeline. Second, the framework is evaluated across 24 text–trajectory encoder combinations and two task-weighting strategies to examine their effects on individual-task and joint classification performance. Third, separately generated test scenarios with modified parameters and latent-vector ablations are used to evaluate performance under changed parameter configurations and assess input dependence within the trained classifier. The remainder of the paper is organized as follows. Section 2 introduces related work. Section 3 describes the dataset and proposed method. Section 4 presents experimental results. Section 5 concludes the paper and discusses future research directions.

2. Related Work

This section reviews prior studies related to scenario classification, disturbance characterization, and standardized safety testing for autonomous driving. Text encoding and variable-length time-series modeling are then reviewed as approaches for representing scenario context and vehicle motion.

2.1. Scenario Classification for Autonomous Driving

Various studies have been conducted to systematically define and classify autonomous driving scenarios. Layer-based approaches organize scenario components into categories describing road infrastructure, environmental conditions, and dynamic objects [13,14]. Such structures provide a systematic description of scenario composition, while further characterization is needed to distinguish the behaviors and interactions occurring within these components. Ontology-based approaches have also been proposed to represent relationships between surrounding objects and describe scenario interactions in a structured manner [15]. Naturalistic driving studies have also examined vehicle behavior in relation to its spatial context. Geographic Information Systems have been used to extract driving patterns from kinematic measurements across different spatial scales, relating vehicle behavior to the road context in which it occurs [16].
Our prior work proposed a scenario classification method based on vehicle driving trajectories and behavioral patterns, demonstrating that trajectory-centric features can serve as a practical basis for scenario classification [12]. However, its reliance on rule-based criteria required additional rules when new scenario types or conditions were introduced, and environmental disturbance factors were not incorporated into the classification. Trajectory types describe the vehicle maneuver, but do not fully characterize the conditions under which it is performed. Scenarios involving the same maneuver may differ in visibility, road conditions, and surrounding-vehicle interactions, placing different demands on autonomous driving functions. Joint characterization of vehicle behavior and disturbance factors therefore allows test scenarios to be organized according to both the maneuver and its associated conditions. The present study extends trajectory-based classification by jointly considering vehicle maneuvers, perception-related environmental conditions, and collision-related outcomes used as decision-disturbance labels.

2.2. Disturbance Factors in Autonomous Driving

Autonomous driving systems rely on sensing, decision-making, and control processes, each of which can be affected by different disturbance factors. Commonly used sensors include cameras, LiDAR, Radar, GNSS, and IMU, and their performance can vary with environmental conditions [17]. Camera-based perception is sensitive to illumination and adverse weather [18], while LiDAR and Radar can also exhibit degraded performance under unfavorable environmental conditions [19,20,21]. Previous studies have therefore investigated low illumination, adverse weather, and sensor contamination as sources of perception uncertainty [22,23,24]. These factors motivate identifying scenario conditions that may affect the reliability of perception inputs.
At the decision-making stage, surrounding-vehicle interactions can create situations requiring appropriate maneuver selection or collision avoidance. Such interactions are commonly characterized using surrogate safety measures, including Time-To-Collision (TTC), headway, and collision-related indicators [25,26,27]. These measures describe aspects of interaction risk and provide a basis for distinguishing potentially hazardous scenarios. They characterize the traffic situation, however, rather than directly establishing that an error occurred within a vehicle’s decision-making module.
In the vehicle control stage, discrepancies may arise between intended commands and actual vehicle behavior because of vehicle dynamics and road surface conditions [28]. Low-friction surfaces, for example, can affect the motion produced by a given command. Characterizing control disturbances therefore requires consideration of vehicle response and operating conditions. The same environmental factor may also affect multiple functions: adverse weather can influence both sensor observations and road friction. Consequently, functional categories describe potentially affected processes and need not be mutually exclusive.
These studies motivate organizing disturbance factors according to the autonomous driving functions they may affect. This categorization characterizes scenario conditions rather than directly diagnosing failures in individual system components. In the present study, perception and decision-related conditions are included in the classification tasks because they can be labeled using the available scenario and simulation data. Control disturbances are reviewed as part of the conceptual categorization but are excluded from classifier training and quantitative evaluation, as a validated control-disturbance labeling criterion has not been established for the present dataset. Additional continuous risk indicators for characterizing pre-collision interactions are summarized in Appendix A.

2.3. Standardized Testing and Digital Twin-Based Validation for Autonomous Vehicles

In addition to identifying disturbance factors, autonomous driving evaluation requires systematic and reproducible testing conditions. Standardized testing frameworks specify road configurations, operating conditions, and safety-critical interactions to support consistent assessment. Programs such as Euro NCAP provide defined test conditions for vehicle safety evaluation [29]. Intersection test procedures and accident-based scenario definitions provide further foundations for evaluating interactions involving crossing and turning vehicles [30,31]. These frameworks provide a basis for constructing reproducible test scenarios, while scenario classification supports organizing the resulting cases according to vehicle behavior and disturbance conditions. The two functions are complementary: standardized conditions define how scenarios are instantiated, and classification describes which maneuvers and functional challenges are represented in a test collection.
Digital twin frameworks have also been developed to support autonomous vehicle validation across different vehicle scales and operational design domains [32]. Their integration with systematic test definition, automated simulation, and performance analysis connects digital representations to validation workflows [33]. These applications highlight the potential of digital twins to serve as infrastructure for validation decisions, extending their role beyond digital replication. Scenario classification could contribute to this role by organizing simulated cases according to vehicle maneuvers, environmental conditions, and interaction outcomes. Such descriptions could support the retrieval of comparable cases, examination of scenario-category coverage, and selection of cases for further testing.

2.4. Text Encoding for Scenario Representation

Beyond defining physical and behavioral conditions, data-driven scenario classification requires contextual information to be represented in a machine-processable form. OpenSCENARIO descriptions encode scenario entities, actions, and environmental settings in structured XML, providing contextual information that complements observed vehicle motion. Transformer-based models capture contextual relationships within sequential text [34]. Pre-trained language models, including BERT, RoBERTa, and the GPT family, have demonstrated their applicability to text representation and language understanding [35,36,37]. Dedicated text embedding models have also been developed to produce vector representations for downstream tasks, offering different model sizes, embedding dimensions, and computational requirements [38,39,40,41]. These approaches provide a basis for representing textual scenario descriptions as feature vectors. For structured scenario files, the relevance of an embedding depends on its ability to retain information about scenario entities, parameters, and relationships. Text encoding can therefore provide contextual features to complement temporal representations of vehicle behavior, while input structure and sequence length remain relevant considerations.

2.5. Preprocessing for Variable-Length Time Series Data

Vehicle trajectories vary in duration across scenarios, requiring appropriate handling of variable-length time-series data. Previous studies have investigated padding, interpolation, and frequency-based compression methods to accommodate such sequences as model inputs [42,43,44]. Zero-padding and edge-padding extend sequences with additional values, while interpolation and methods such as Spectral Pooling transform the sequence representation [43,44]. These approaches differ in how they retain temporal information. Interpolation or compression may modify temporal characteristics, and fixed observation windows may exclude observations outside the selected interval. Padding retains the original valid observations but introduces artificial values that may affect learning if they are not appropriately excluded. Sequence-length information and masking can therefore be used with padding to distinguish valid observations from padded regions in relevant model computations and reconstruction losses.

2.6. Time-Series Modeling for Driving Trajectories

Beyond preprocessing variable-length sequences, an appropriate temporal model is required to capture the dynamic patterns embedded in vehicle trajectories. Various deep learning-based approaches have been proposed for modeling time-series data. Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) networks have been widely used for sequential data modeling because they can capture temporal dependencies across observations [45,46]. In particular, LSTMs alleviate the difficulty of learning long-term dependencies in conventional RNNs, making them suitable for representing complex temporal patterns in driving trajectories. Transformer-based models have also been applied to time-series modeling because of their ability to capture long-range dependencies and support parallel computation [47,48]. More recently, xLSTM and Mamba have been introduced as alternative sequence-modeling architectures, extending recurrent memory mechanisms and employing selective state-space modeling, respectively [49,50]. These developments broaden the range of temporal modeling approaches available for sequential data, with different mechanisms for capturing temporal dependencies and processing long sequences.
Previous studies provide important foundations for scenario classification, disturbance characterization, and temporal representation learning. Together, they motivate a scenario representation that combines vehicle maneuvers with the conditions affecting autonomous driving functions. Scenario descriptions provide structural and environmental context, while vehicle trajectories describe the resulting motion over time. Integrating these complementary sources provides a basis for joint scenario characterization. Accordingly, this study focuses on simultaneous classification of trajectory types and perception and decision disturbances, with control disturbances retained only in the conceptual discussion.

3. Methodology

This section describes the dataset construction, model architecture, and training procedure used for scenario classification. Two data sources are used: simulation scenarios generated and labeled under predefined configurations for classifier training and evaluation, and real-world intersection trajectories from the SinD dataset [51] for training the trajectory Autoencoders. The proposed multimodal framework separately encodes structured scenario information and Ego-vehicle motion, then combines the resulting representations for joint classification of trajectory type, perception-related conditions, and collision-related outcomes used as decision-disturbance labels. Control disturbances are excluded from the quantitative classification targets because the available data are insufficient to establish a reliable labeling criterion. Section 3.1 describes the simulation environment, scenario generation, labeling, and simulation data analysis. Section 3.2 introduces the SinD dataset and its preprocessing procedure. Section 3.3 presents the model architecture and training procedure, covering Autoencoder-based variable-length time-series encoding in Section 3.3.1, text encoder selection in Section 3.3.2, and the integrated multi-task classifier, training objectives, and optimization settings in Section 3.3.3.

3.1. Scenario Data Generation

To reproduce safety-critical situations that may occur in real intersection environments, a scenario-based dataset was constructed. The scenarios are represented in OpenSCENARIO (.xosc) format, and simulation execution provides the corresponding vehicle driving data and collision outcomes.

3.1.1. Scenario Environment Configuration

The road environment considered in this study is a four-lane bidirectional intersection, as illustrated in Figure 2. Following the principle of standardized safety testing, the road geometry, initial conditions, and scenario parameters are explicitly defined so that the same scenario configuration can be reproduced across repeated simulation runs. A situation is assumed in which an Ego vehicle and an NPC vehicle simultaneously enter the intersection. No commanded acceleration or deceleration is applied during scenario execution; therefore, each vehicle is intended to maintain its assigned constant speed unless its motion is affected by the simulated interaction or termination condition. Various driving patterns were defined with reference to the types of accidents that may occur at intersections [31]. The simulation map was configured with road-width and intersection-structure information, and each road segment was assigned a unique identifier. Because the OpenDRIVE road network used in this study does not provide all possible intersection connections, particularly U-turn connectivity, the generated scenarios are restricted to ST, LT, and RT maneuvers. Consequently, not all possible intersection accident configurations are represented in the present dataset.

3.1.2. OpenSCENARIO-Based Scenario Structure

In simulation, the map and scenario have complementary roles: the map defines the spatial structure in which vehicles move, while the scenario specifies vehicle behavior within that space. The map consists of predefined elements, and rendering methods and environmental factors (e.g., lighting, vegetation, structures, physics engine) vary across simulators. As these differences affect sensor data such as LiDAR and Radar, all experiments were conducted within a single simulator environment. OpenSCENARIO definitions used in this study follow v1.2 (https://www.asam.net/static_downloads/ASAM_OpenSCENARIO_V1.2.0_Model_Documentation/modelDocumentation/RenderedXsdOutput.html, accessed on 11 August 2026). The scenarios consist of four main components:
  • Entity definition (Ego vehicle, NPC vehicle);
  • Environment configuration and entity initial state;
  • Scenario action definition;
  • Success and failure conditions.
Specifically, the Ego and NPC vehicles are defined, their initial positions and speeds are set, and the direction of travel is specified through the scenario. The scenario is considered successful if the Ego vehicle reaches the goal point within the time limit (30 s), and a failure if a collision with another object occurs or the goal is not reached within the time limit.

3.1.3. Scenario Types and Data Generation

Scenarios for straight (ST), left-turn (LT), and right-turn (RT) situations were constructed based on possible accident types at a four-lane bidirectional intersection [31]. Each logical scenario specifies the starting positions and travel directions of the Ego and NPC vehicles. In the training and validation configuration, both vehicles initially start 25 m from the intersection. Concrete scenarios were generated by combining temporal, initial-speed, and environmental parameters with lane-level variations of the road network. This procedure produced 10,976 ST, 5488 LT, and 784 RT simulation instances, which were subsequently divided into training and validation subsets.
The test dataset was generated separately using a modified parameter configuration shared across all scenario types. As summarized in Table 1, test scenarios use different times of day and modified road friction and visibility settings. The initial-speed range remains unchanged, while the sampling interval increases from approximately 5 to 10 km/h. The test dataset therefore evaluates classification under new temporal and environmental parameter settings within the same scenario generation framework, with some individual parameter values shared with the training and validation configuration.
For both dataset configurations, environmental parameters are specified as coupled sets, preserving the prescribed combination of precipitation type, road friction, and visibility for each condition. These sets are combined with temporal and initial-speed parameters and the lane-level configurations of each logical scenario.

3.1.4. Data Labeling

The generated data were assigned trajectory and disturbance labels. Trajectory labels identify the Ego vehicle maneuver as straight (ST), left-turn (LT), or right-turn (RT). Following the conceptual categorization of perception, decision-making, and control discussed in Section 2, the present study uses perception and decision disturbances as quantitative classification targets. These labels characterize scenario conditions and simulation outcomes rather than directly diagnosing failures in individual autonomous driving components. The same labeling criteria are applied to the training, validation, and separately generated test datasets.
For perception disturbances, environmental conditions associated with potential sensor-performance degradation are used as labeling criteria [17,18,19,20,21,22,23,28]. The binary label is set to 1 when a scenario contains adverse weather (rain, snow, or fog) or a designated low-illumination condition, and to 0 otherwise. The low-illumination settings correspond to 06:00 and 22:00 in the training and validation configuration and 03:00 and 23:00 in the test configuration. Thus, the labeling rule remains consistent across datasets despite differences in their parameter values. Representative environmental conditions are illustrated in Figure 3.
For decision disturbances, an explicit event-based criterion is used: the binary label is set to 1 when a collision between the Ego and NPC vehicles occurs during simulation, and to 0 otherwise. The present study uses the Ego trajectory as the motion input without explicitly incorporating relative-motion features between vehicles into the trajectory encoder. Collision occurrence provides an observable outcome for labeling, while the Ego trajectory represents the vehicle motion used for classification. The label therefore identifies a collision-related outcome rather than directly quantifying pre-collision interaction risk. To characterize interactions with other vehicles, surrogate safety measures such as TTC, headway, TCPA, and DCPA can be considered using relative-state information and appropriate thresholds or risk-mapping rules. These supplementary indicators are summarized in Appendix A as possible extensions for pre-collision risk analysis.
For control disturbances, the scenario configuration includes the road-friction parameter, which can influence vehicle response to driving commands. However, the available data are insufficient to establish a reliable relationship between the specified friction values and the occurrence of a control disturbance. Road friction alone is therefore not used as a labeling criterion in the present experiments. It is retained as an environmental parameter, while control disturbances remain part of the conceptual categorization and are excluded from classifier training and quantitative evaluation. Each scenario is consequently represented by a joint label tuple,
y = y traj , y perc , y dec ,
where y traj ∈ { ST , LT , RT } identifies the Ego maneuver, and y perc , y dec ∈ { 0 ,   1 } describe the perception-related condition and collision-related outcome, respectively. This representation distinguishes scenarios that share the same maneuver but differ in disturbance labels. It provides a common basis for grouping and comparing simulation cases across both behavioral and disturbance-related categories.

3.1.5. Analysis of Simulation Data

Various driving data were collected by executing the generated scenarios through simulation. Key features of the acquired data are summarized in Table 2. Variables such as position coordinates, velocity, and acceleration directly reflect the vehicle’s driving state and play an important role in driving behavior analysis.
The motion representation uses velocity and acceleration rather than absolute position to describe vehicle behavior without directly encoding its location within the simulation map. Position data remain available in the simulation logs but are not included in the Autoencoder input. The simulation data were collected at 50 Hz.

3.2. SinD Dataset

This section describes the SinD dataset used to train the variable-length Autoencoders, including the dataset structure, preprocessing procedures, and statistical properties.

3.2.1. SinD Dataset Overview

Our previous work classified vehicle trajectories into seven classes using simulation-generated data [12]. While this approach enabled classification within the evaluated simulation scenarios, its applicability to real-world trajectory data was not established. In the present study, the SinD dataset is used to train the trajectory Autoencoders using observations collected from real intersection environments [51]. SinD provides positional information for vehicles and other road users extracted from drone recordings at urban intersections in China, as illustrated in Figure 4. These observations introduce real-world motion variations into representation learning, without implying generalization to driving environments beyond those evaluated in this study.

3.2.2. Data Preprocessing

SinD includes a variety of intersection structures with different geometric configurations across locations. Most data was collected in intersection environments, with vehicle behaviors including straight driving, left turns, right turns, and U-turns. The dataset also includes diverse traffic participants such as trucks and motorcycles, as well as pedestrians. The following preprocessing steps were applied to align SinD with the simulation data format:
  • Vehicle objects were isolated and pedestrian and other non-vehicle data were excluded.
  • The temporal resolution was resampled to 50 Hz via interpolation to match the simulation data, ensuring temporal consistency between the two datasets.
  • Time-series data were separated per vehicle and U-turn (UT) data were excluded as the simulation map does not define U-turn routing.
  • Unlike the simulation data, SinD does not contain the environmental scenario attributes or explicit simulation collision outcomes used to define the perception- and decision-disturbance labels in this study. Therefore, SinD is used only for learning vehicle-motion representations and is not assigned disturbance labels.

3.3. Model Architecture

This section describes the model architecture and training procedure. The proposed method is based on a multimodal structure that independently encodes structural environmental information from the scenario file and vehicle motion information from simulation data, then combines them for classification.

3.3.1. Autoencoder-Based Variable-Length Time Series Encoding

Autoencoders were used to learn compressed latent representations of vehicle motion at intersections [52,53]. The input consists of vehicle state time series from the SinD dataset, represented by six features at each time step: velocities ( v e l x , v e l y , v e l z ) and accelerations ( a c c x , a c c y , a c c z ) in the three coordinate directions [12,54]. The features were normalized using per-feature means and standard deviations obtained from the training subset, with the same statistics applied to the validation subset. A moving average with a window size of five was applied to smooth temporal fluctuations. To limit memory requirements, the maximum input sequence length was set to 15,000 time steps, corresponding to approximately five minutes at 50 Hz. Vehicle trajectories vary in duration across scenarios. The training pipeline processes sequences in mini-batches together with their valid lengths. A binary mask derived from these lengths excludes padded time steps from the reconstruction loss, so that only valid observations contribute to the reconstruction objective.
Six Autoencoder configurations were considered: Simple-AE, BiLSTM-AE, Stacked-AE, Transformer-AE, xLSTM-AE, and Mamba-AE. The first three employ conventional LSTM encoders, while the remaining configurations use Transformer, xLSTM, and Mamba temporal modeling components, respectively [47,49,50]. All models follow the Encoder–Bottleneck–Decoder structure illustrated in Figure 5. The encoder maps the input sequence to a fixed-dimensional latent representation, and the decoder reconstructs the sequence from this representation. After training, the encoder provides vehicle-motion features for downstream scenario classification.
The reconstruction objective is the masked mean squared error,
MSE masked = ∑ b , t , d m b , t x b , t , d − x ^ b , t , d 2 D ∑ b , t m b , t ,
where x b , t , d and x ^ b , t , d denote the observed and reconstructed values of feature d at time step t in sequence b, respectively. The mask m b , t equals one for valid observations and zero for padding, and D is the feature dimensionality. All Autoencoder models were optimized for 300 epochs using AdamW with a learning rate of 10 − 3 and weight decay of 10 − 4 .
The learned latent representations were first processed using PCA for dimensionality reduction and then projected into two dimensions using t-SNE, as shown in Figure 6. Colors and marker shapes identify trajectory types, while darker and lighter shades distinguish the binary decision labels. The visualization provides a qualitative view of local grouping and overlap in the latent space. The t-SNE projections show local groups associated with trajectory types and decision labels. In several encoders, decision-0 samples form multiple compact groups, while decision-1 samples occupy regions with overlapping trajectory types. Mamba-AE shows comparatively distinct ST and LT groups among the decision-1 samples, although complete separation is not observed across all classes. These projections provide qualitative evidence of latent structure, but their two-dimensional grouping alone does not establish downstream classification performance. To complement the visual analysis, Table 3 reports the SinD–simulation domain silhouette score and squared maximum mean discrepancy (MMD2). The silhouette score measures separation according to dataset membership, whereas MMD2 measures discrepancy between the two distributions. These metrics characterize domain differences rather than trajectory- or decision-label separability.
All encoders produced positive domain silhouette scores, indicating that domain-related structure remains in the latent spaces. Several encoders yielded lower silhouette scores than the log-summary reference, while Stacked-AE remained nearly unchanged and Transformer-AE showed a slight increase. All six encoders also yielded lower reported MMD2 values than the reference. These comparisons are descriptive because the log-summary vectors and learned representations differ in feature construction and dimensionality. Lower reported values do not independently establish domain alignment, preservation of task-relevant information, or improved cross-domain generalization. The downstream classification results provide the direct assessment of the representations’ utility within the evaluated scenario configurations.

3.3.2. Text Encoder Selection

OpenSCENARIO uses a structured XML-based format in which environmental conditions, entity definitions, initial states, and scenario actions are explicitly represented. Because the amount and structural complexity of this information increase with scenario complexity, the scenario file is converted into a machine-processable textual representation rather than manually reducing it to a small set of numerical attributes. In the proposed pipeline, the OpenSCENARIO file is parsed and serialized as a JSON string before text encoding. This preserves the hierarchical scenario information while providing a consistent input form for pre-trained text encoders.
Large generative language models could also be used for this purpose, but repeated scenario-level embedding would introduce substantial computational overhead. Therefore, pre-trained Transformer-based embedding models were selected to balance representation capability and computational cost. Four encoders with different model sizes and embedding dimensions were evaluated: EmbeddingGemma-300M, all-MiniLM-L6-v2, BGE-M3, and Qwen3-Embedding-0.6B. Their main characteristics are summarized in Table 4. The encoders are used as pre-trained feature extractors, and their output embeddings are subsequently combined with the trajectory latent vectors produced by the Autoencoders.

3.3.3. Full Model Architecture

The overall model integrates a pre-trained AE and a pre-trained text encoder to perform multi-task classification of intersection scenarios. Vehicle motion features are extracted by the AE encoder, yielding a 64-dimensional latent vector from the Ego vehicle’s velocity and acceleration time series. Scenario structural features are obtained by parsing the OpenSCENARIO file, serializing the structured information as a JSON string, and encoding it with one of the four pre-trained text encoders described in Section 3.3.1. The final classifier input is formed by concatenating the trajectory latent vector and the text embedding. All 24 combinations (6 trajectory AEs × 4 text encoders) are independently trained and compared. The overall processing pipeline is illustrated in Figure 7.
The classifier consists of a shared MLP backbone and three task-specific heads. The backbone combines scenario context and observed vehicle motion by processing the concatenated features through fully connected layers with LayerNorm, ReLU, and Dropout. The task-specific heads predict trajectory type (ST/LT/RT), perception disturbance, and decision disturbance, forming the joint scenario label defined in Section 3.1.4. Control disturbance is excluded from the prediction targets, consistent with the labeling scope. Per-task metrics evaluate the individual outputs, while Full Match Accuracy measures whether the complete scenario label is predicted correctly.
Each task loss is defined using cross-entropy:
L task = − ∑ c = 1 C y c log p ^ c ,
where y c denotes the ground-truth label for class c and p ^ c denotes the predicted probability for class c.
For the fixed-weight configuration, the relative task weights are assigned as 1 : 0.5 : 0.5 for trajectory, perception disturbance, and decision disturbance, respectively. The total loss is therefore defined as
L fixed = L traj + 1 2 L perc + L dec .
For the uncertainty-weighted configuration, each task is assigned a learnable log-variance parameter
s i = log σ i 2 , i ∈ { traj , perc , dec } .
For numerical stability, the log-variance is clipped to the interval [ − 2 ,   2 ] :
s ˜ i = clip ( s i , − 2 ,   2 ) .
The uncertainty-weighted objective is defined as
L unc = ∑ i ∈ { traj , perc , dec } exp ( − s ˜ i ) L i + s ˜ i ,
Here, L i denotes the loss of the corresponding task. The factor exp ( − s ˜ i ) determines the task weight from the clipped log-variance, while the additive s ˜ i term regularizes the learned weighting. The classifier and log-variance parameters were optimized jointly using AdamW. The quantitative evaluation considered trajectory type, perception disturbance, and decision disturbance. The simulation dataset was stratified into training and validation subsets according to trajectory type and decision-disturbance labels, and the same split was used for all model configurations.

4. Results

We evaluate the proposed framework across 24 combinations of four text encoders (EmbeddingGemma, MiniLM, BGE-M3, and Qwen-Emb) and six trajectory autoencoder structures (BiLSTM-AE, Simple-AE, Stacked-AE, Transformer-AE, xLSTM-AE, and Mamba-AE). Classification performance is assessed on trajectory type (ST, LT, or RT), perception disturbance, and decision disturbance using per-task accuracy, trajectory macro-averaged F1, and Full Match Accuracy. Full Match Accuracy requires simultaneous correctness for all three labels. The test set contains 3840 samples, including 3520 perception-positive and 320 perception-negative samples. Predicting every sample as perception-positive yields an accuracy of 0.9167 and a positive-class F1 score of 0.9565. These values therefore serve as reference baselines and do not, by themselves, demonstrate discrimination between the two perception classes. Binary F1 values for perception and decision are positive-class F1 scores, whereas trajectory F1 is the macro-averaged F1 over ST, LT, and RT.
The main analysis focuses on the separately generated test dataset. Reported validation results are included to contextualize performance under the parameter configuration used for training, with detailed validation tables provided in Appendix B. The evaluation addresses three questions: how accurately the framework predicts individual targets and complete scenario labels; how encoder selection and task weighting affect these predictions; and how performance changes under modified scenario parameters and removal of either input modality. The following analyses therefore consider classification performance, configuration sensitivity, and input contributions within the integrated framework.

4.1. Fixed Task Weighting

Classification performance under fixed task weighting is presented in Table 5. Mamba-AE achieved the strongest trajectory classification performance across text encoder pairings, with consistently high trajectory accuracy and macro F1. Perception and decision performance varied more substantially across encoder pairings, indicating that joint performance depends on the interaction between the text and trajectory representations. Under fixed weighting, Qwen-Emb with Mamba-AE achieved the highest Full Match Accuracy of 0.9161. Although this configuration obtained near-perfect trajectory and decision scores, its perception accuracy matched the majority-class baseline. The Full Match result therefore indicates high simultaneous label agreement on the evaluated test distribution, but does not establish balanced discrimination of perception-positive and perception-negative scenarios.

4.2. Uncertainty-Based Task Weighting

Uncertainty-based weighting adjusts the contribution of each task through learnable log-variance parameters. Table 6 shows the corresponding test results. The best Full Match Accuracy was 0.9161, achieved by Qwen-Emb with Mamba-AE, while performance varied across other text–trajectory encoder pairings. Compared with fixed weighting, uncertainty-based weighting did not provide a consistent improvement in Full Match Accuracy. Its effect depended on the encoder pairing and task: some configurations improved trajectory F1, whereas disturbance accuracy or Full Match Accuracy decreased. These results indicate that task weighting and encoder selection should be considered jointly.

4.3. Encoder-Wise Comparison

To compare architectures independently of a single text encoder pairing, Table 7 reports the mean performance of each trajectory encoder across the four text encoders and of each text encoder across the six trajectory encoders. Mamba-AE achieved the highest mean trajectory accuracy, trajectory F1, and Full Match Accuracy under fixed weighting. Among text encoders, Qwen-Emb achieved the highest fixed-weight mean Full Match Accuracy. The results also show that the best individual pairing does not necessarily imply uniformly superior performance across all tasks.

4.4. Validation and Test Comparison

Table 8 summarizes the reported validation and test performance, averaged over the six trajectory encoders for each text encoder. The validation subset was drawn from the scenario collection generated under the training parameter configuration, whereas the test dataset was generated separately using modified times of day and environmental settings. Detailed validation results for all encoder combinations are provided in Appendix B.
Validation perception accuracy was generally higher than test accuracy under both weighting strategies. High performance within the training parameter configuration therefore did not consistently extend to the modified test configuration. The magnitude of this difference depended on the encoder pairing, and not every configuration showed a decrease. The corresponding Full Match results also indicate that strong validation performance did not ensure equally high simultaneous classification accuracy on the test set. Perception labels are defined using weather and time-of-day conditions described in the scenario text. The validation–test differences show that performance under the training parameter configuration did not consistently extend to the modified test configuration. However, the comparison does not isolate the effects of individual parameter changes or identify text-encoder representation quality as the cause. Differences in class composition and scenario coverage may also influence aggregate accuracy. The results should therefore be interpreted as performance under a changed test configuration rather than as a controlled estimate of the effect of weather or time-of-day changes alone.

4.5. Latent-Vector Ablation

To examine the contribution of the two modalities, the trained checkpoints were evaluated after setting either the text latent vector or the trajectory latent vector to zero. Table 9 reports the results averaged over the four text encoder pairings. These experiments measure input dependence within the trained multimodal classifier and are not comparisons with separately trained unimodal models. Removing the text latent vector substantially reduced trajectory accuracy for most encoders, whereas Mamba-AE retained nearly the same trajectory accuracy. Removing the trajectory latent vector consistently reduced decision accuracy, indicating that the decision heads relied strongly on trajectory information. Perception accuracy changed less consistently, demonstrating that the contribution of each modality was task- and encoder-dependent.
The results indicate task-dependent use of the two representations. In particular, trajectory information was important for decision disturbance classification, while the effect of text and trajectory removal on perception classification varied across encoder structures. Because these are zero-vector evaluations of already trained multimodal classifiers, they measure input dependence rather than the standalone performance of a newly trained unimodal model.

5. Conclusions

This section summarizes the main findings of the proposed multimodal scenario classification framework and discusses its limitations and future research directions.

5.1. Summary of Experimental Results

This study presented an integrated framework for classifying intersection scenarios through a joint description of vehicle maneuver, perception-related conditions, and collision-related outcomes. By combining OpenSCENARIO text embeddings with SinD-trained Ego-motion representations, the framework produces three classification outputs within a common multi-task pipeline. The evaluation across 24 encoder combinations and two weighting strategies showed that encoder selection affects both individual-task and joint performance. Qwen-Emb combined with Mamba-AE achieved the highest Full Match Accuracy of 0.9161, while uncertainty-based weighting did not consistently improve the results. Perception classification remained a limiting factor, demonstrating the importance of examining individual targets alongside the complete scenario label. The validation–test comparison and latent-vector ablations further characterized the framework’s behavior. High validation performance did not consistently extend to modified test parameters, and the contribution of each modality differed across tasks and encoder configurations. In particular, removing trajectory information substantially reduced decision-label classification performance, whereas Mamba-AE retained high trajectory accuracy after text removal. These findings establish the evaluated performance and input dependence of the proposed classification pipeline within the considered scenario family. The framework provides a basis for organizing simulation cases by both maneuver and disturbance labels, with broader scenario coverage and additional labeling criteria remaining directions for further development.
The proposed classification framework could support decision-making in digital twin-based autonomous driving validation by organizing simulated scenarios according to vehicle maneuvers, perception-related conditions, and collision-related outcomes. These joint labels provide a basis for retrieving comparable cases, examining the coverage of evaluated scenario categories, and selecting cases for further review. In this role, the framework provides scenario-level information for validation planning and analysis. Integration with an operational digital twin and assessment of its benefits to validation workflows remain to be investigated.

5.2. Limitations and Future Work

Several limitations remain in the current framework. First, the simulation dataset is generated from a constrained intersection environment with predefined vehicle behaviors and environmental conditions. Although this setup supports systematic scenario generation and reproducible evaluation, some classification targets may be easier to distinguish than in more diverse driving environments. Expanding the dataset to cover additional road geometries, maneuvers, and traffic participants, including lane changes, U-turns, roundabouts, and pedestrian interactions, would enable a broader assessment of the framework. Second, perception classification showed differences between validation and test performance under modified temporal and environmental settings. High validation accuracy therefore did not consistently indicate reliable classification under new parameter configurations. Further investigation should examine the representation of weather and illumination information, broaden environmental coverage, and assess performance using class-wise metrics alongside overall accuracy. Third, the decision-disturbance label is based on collision occurrence and therefore captures an interaction outcome rather than the severity of pre-collision risk. The current motion representation also uses only the Ego trajectory, without explicitly modeling relative motion between vehicles. Incorporating surrounding-vehicle trajectories and continuous safety indicators, such as TTC, headway, TCPA, and DCPA, could support more detailed interaction analysis and decision-disturbance labeling. Fourth, control disturbances were excluded from classifier training and quantitative evaluation. Although road friction is specified in the scenario configuration, the available data are insufficient to establish a reliable relationship between friction values and the occurrence of a control disturbance. Developing and validating control-disturbance labeling criteria will require additional vehicle-dynamics information and control-performance indicators, including tracking error and lateral stability. Finally, evaluation on larger and more diverse datasets, together with repeated runs, would help assess the robustness and performance variability of the encoder combinations and task-weighting strategies. Domain adaptation and joint representation learning with real-world and simulation data also warrant investigation to address the remaining distributional differences. These extensions will help determine the framework’s applicability beyond the scenario configurations evaluated in this study.

Author Contributions

Conceptualization, Sungmo Ku; methodology, Sungmo Ku; software, Sungmo Ku; validation, Jinho Lee; investigation, Sungmo Ku; writing—original draft preparation, Sungmo Ku; writing—review and editing, Sungmo Ku and Jinho Lee; visualization, Sungmo Ku; supervision, Jinho Lee; project administration, Jinho Lee; funding acquisition, Jinho Lee. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Institute of Information Communications Technology Planning Evaluation (IITP)-Innovative Human Resource Development for Local Intellectualization program grant funded by the Korea government (MSIT) (IITP-2026-RS-2020-II201741).

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The generated OpenSCENARIO data, excluding the map data, are available from the corresponding author upon reasonable request. The map data used in this study cannot be publicly shared because they are proprietary data owned by the simulation company.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT (5.5, 5.6 sol) for grammar correction, translation, and assistance with table and figure formatting. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ADAutonomous Driving
AEAutoencoder
GNSSGlobal Navigation Satellite System
IMUInertial Measurement Unit
MMDMaximum Mean Discrepancy
PCAPrincipal Component Analysis
t-SNEt-distributed Stochastic Neighbor Embedding

Appendix A. Continuous Risk Indicators for Pre-Collision Interaction Analysis

This appendix summarizes supplementary continuous risk indicators that can be used to characterize the severity of vehicle interactions before an actual collision occurs. These indicators are not used as decision-disturbance labels in the quantitative experiments of this study; the decision-disturbance label is defined by collision occurrence as described in Section 3.1.4. Instead, the following formulation provides an extensible basis for future risk-based labeling and post-simulation safety analysis.

Appendix A.1. Vehicle State and Relative Variables

Let the Ego vehicle and an interacting NPC vehicle be denoted by the subscripts E and N, respectively. Their planar positions at time t are defined in a common Cartesian coordinate system as
p E ( t ) = x E ( t ) y E ( t ) , p N ( t ) = x N ( t ) y N ( t ) ,
and their velocity vectors are
v E ( t ) = v x , E ( t ) v y , E ( t ) , v N ( t ) = v x , N ( t ) v y , N ( t ) .
The relative position and relative velocity are then defined as
r ( t ) = p N ( t ) − p E ( t ) ,
v r ( t ) = v N ( t ) − v E ( t ) .
The Euclidean distance between the two vehicles and the magnitude of the relative velocity are
d ( t ) = r ( t ) ,
v rel ( t ) = v r ( t ) .
To distinguish approaching motion from separating motion, the closing speed is defined as the component of the relative velocity directed toward the other vehicle:
v c ( t ) = max 0 , − r ( t ) · v r ( t ) r ( t ) .
Accordingly, v c ( t ) > 0 indicates that the inter-vehicle distance is decreasing, whereas v c ( t ) = 0 indicates that the vehicles are not approaching along the current relative-position direction.

Appendix A.2. Time-to-Collision

A current-state TTC can be approximated using the instantaneous inter-vehicle distance and closing speed. For v c ( t ) > 0 ,
T T C ( t ) = d ( t ) v c ( t ) .
To avoid assigning a finite collision time when the vehicles are not approaching, the complete definition is
T T C ( t ) = d ( t ) v c ( t ) , v c ( t ) > 0 , ∞ , v c ( t ) = 0 .
This TTC formulation provides an instantaneous temporal proximity measure. Because it is derived from the current distance and radial closing speed, it should be interpreted as a local kinematic indicator rather than an exact collision time for turning trajectories.

Appendix A.3. Time and Distance to the Closest Point of Approach

For interactions in which the relative velocity is assumed to remain constant over a short prediction interval, the time to the closest point of approach (TCPA) is
T C P A ( t ) = − r ( t ) · v r ( t ) v r ( t ) 2 , v r ( t ) 2 > 0 .
A positive TCPA indicates that the closest approach is expected in the future, T C P A ( t ) = 0 indicates that the vehicles are currently at their closest approach under the constant-velocity assumption, and a negative TCPA indicates that the closest point has already been passed.
The relative position at the closest point of approach is
r CPA ( t ) = r ( t ) + v r ( t ) T C P A ( t ) ,
and the corresponding distance at the closest point of approach (DCPA) is
D C P A ( t ) = r ( t ) + v r ( t ) T C P A ( t ) .
TCPA and DCPA can be combined into a continuous CPA-based risk term:
R CPA ( t ) = exp − D C P A ( t ) D 0 exp − T C P A ( t ) T 0 , T C P A ( t ) ≥ 0 , 0 , T C P A ( t ) < 0 ,
where D 0 and T 0 are distance and time scale parameters, respectively. This formulation assigns higher risk when the predicted closest approach occurs both soon and at a short distance.

Appendix A.4. Dynamic Safety Distance

A fixed distance threshold does not reflect changes in interaction severity caused by different relative speeds. Therefore, a dynamic safety distance is defined using the minimum standstill distance, closing speed, and relative speed:
d safe ( t ) = d 0 + T h v c ( t ) + k v v rel ( t ) ,
where d 0 is the base minimum distance, T h is a time-headway coefficient, and k v is a relative-speed correction coefficient. An initial parameter setting used for interpretation can be expressed as
d 0 = 3.0 m , T h = 1.0 s , k v = 0.2 s .
The current proximity risk is then normalized with respect to the dynamic safety distance:
R prox ( t ) = exp − d ( t ) d safe ( t ) .
Because d ( t ) ≥ 0 and d safe ( t ) > 0 , R prox ( t ) lies in the interval ( 0 ,   1 ] , with values approaching one as the vehicles become spatially closer.

Appendix A.5. Risk Synthesis and Temporal Memory

The instantaneous risk terms describe different aspects of the interaction. R prox ( t ) represents the current spatial proximity, whereas R CPA ( t ) represents short-term predictive risk under a constant-relative-velocity assumption. To prevent a high-risk indication in one component from being diluted by a lower value in another component, the instantaneous base risk is defined using the maximum operator:
R base ( t ) = max R prox ( t ) , R CPA ( t ) .
A memory term is introduced because the risk should not immediately drop to a safe state immediately after a near-pass event. For two consecutive observations separated by Δ t = t k − t k − 1 , the decayed previous risk is
R memory ( t k ) = R ( t k − 1 ) exp − Δ t τ m ,
where τ m is the memory decay constant. The temporally persistent risk is therefore
R ( t k ) = max R base ( t k ) , R ( t k − 1 ) exp − Δ t τ m .
For example, τ m = 2.0 s causes the previous risk contribution to decay exponentially rather than disappear immediately. When the trajectory data are sampled at 50 Hz, Δ t is nominally 0.02 s , although the actual timestamp difference should be used when the sampling interval is not perfectly uniform.

Appendix A.6. Future Trajectory Coordinates

The constant-velocity assumption used by TCPA and DCPA can be inaccurate for intersection maneuvers involving turning or rapidly changing motion. To incorporate the actual or predicted future trajectory, let the future positions of the Ego and NPC vehicles over a prediction horizon H be
p E ( t + τ ) , p N ( t + τ ) , τ ∈ [ 0 , H ] .
The minimum future inter-vehicle distance is defined as
D future ( t ) = min τ ∈ [ 0 , H ] p E ( t + τ ) − p N ( t + τ ) ,
and the time at which this minimum occurs is
τ min ( t ) = arg min τ ∈ [ 0 , H ] p E ( t + τ ) − p N ( t + τ ) .
A horizon such as H = 3.0 s may be used as an initial setting. If future coordinates are obtained directly from recorded simulation trajectories, D future ( t ) is an offline, ground-truth future indicator. For online use, the same formulation requires predicted future coordinates generated by a motion-prediction or vehicle-motion model.

Appendix A.7. Future-Aware Composite Risk Indicator

The future minimum distance can be normalized by the dynamic safety distance to obtain a future-trajectory risk term:
R future ( t ) = exp − D future ( t ) d safe ( t ) .
The final instantaneous risk can then incorporate current proximity, CPA-based prediction, and future-trajectory information:
R base + ( t ) = max R prox ( t ) , R CPA ( t ) , R future ( t ) .
Including temporal memory gives the future-aware persistent risk indicator
R + ( t k ) = max R base + ( t k ) , R + ( t k − 1 ) exp − Δ t τ m
This formulation is intended to represent three complementary risk conditions: (i) the vehicles are currently close relative to the dynamic safety distance, (ii) their current relative motion predicts a close encounter in the near future, or (iii) the actual or predicted future trajectories indicate a close pass that may not be captured by a constant-velocity model.
For discrete SAFE/RISK state assignment, hysteresis can be added using separate entry and exit thresholds,
R enter > R exit ,
with the state entering RISK when R + ( t ) ≥ R enter and returning to SAFE only when
R + ( t ) ≤ R exit ∧ d ( t ) > d safe ( t ) .
The second condition prevents an immediate SAFE transition while the vehicles remain inside the dynamic safety distance even if the scalar risk has already decayed below the exit threshold.
The parameter values in Table A1 are intended only as initial settings. For quantitative use, the parameters and decision thresholds should be calibrated using the distributions of collision, near-miss, and safe interactions in the target dataset.
Table A1. Example Initial Parameters for the Supplementary Risk Indicators.
Table A1. Example Initial Parameters for the Supplementary Risk Indicators.
ParameterExample ValueDescription
d 0 3.0 m Base minimum safety distance
T h 1.0 s Time-headway coefficient
k v 0.2 s Relative-speed correction coefficient
D 0 3.0 m DCPA risk scale
T 0 3.0 s TCPA risk scale
H 3.0 s Future-trajectory horizon
τ m 2.0 s       Risk-memory decay constant      
R enter 0.60 Risk-state entry threshold
R exit 0.30 Risk-state release threshold

Appendix B. Detailed Validation Results

Table A2 and Table A3 report the validation metrics for all 24 text–trajectory encoder combinations under fixed and uncertainty-based weighting, respectively. The validation subset comes from the parameter configuration used for training and is distinct from the separately generated test dataset. Trajectory F1 is macro-averaged over ST, LT, and RT; perception and decision F1 are positive-class scores. Full Match Accuracy requires all three classification targets to be predicted correctly for the same sample. These tables supplement the test results in Section 4.
Table A2. Validation classification performance with fixed task weighting. Acc. denotes accuracy and FM denotes Full Match Accuracy.
Table A2. Validation classification performance with fixed task weighting. Acc. denotes accuracy and FM denotes Full Match Accuracy.
Text EncoderTrajectory AETrajectoryPerceptionDecisionFM
Acc.F1Acc.F1Acc.F1
EmbeddingGemmaBiLSTM-AE0.94320.86581.00001.00000.99250.99050.9426
EmbeddingGemmaSimple-AE0.93800.85971.00001.00000.98490.98080.9252
EmbeddingGemmaStacked-AE0.98780.98451.00001.00000.99880.99850.9867
EmbeddingGemmaTransformer-AE1.00001.00001.00001.00001.00001.00001.0000
EmbeddingGemmaxLSTM-AE1.00001.00001.00001.00001.00001.00001.0000
EmbeddingGemmaMamba-AE1.00001.00001.00001.00001.00001.00001.0000
MiniLMBiLSTM-AE0.94960.87580.98380.99060.99650.99560.9328
MiniLMSimple-AE0.94430.87820.99070.99460.98780.98460.9270
MiniLMStacked-AE0.98900.99130.99010.99430.99880.99850.9786
MiniLMTransformer-AE0.99940.99950.99830.99901.00001.00000.9977
MiniLMxLSTM-AE1.00001.00001.00001.00001.00001.00001.0000
MiniLMMamba-AE1.00001.00001.00001.00001.00001.00001.0000
BGE-M3BiLSTM-AE0.94200.86041.00001.00000.99360.99200.9403
BGE-M3Simple-AE0.94140.86891.00001.00000.98610.98230.9304
BGE-M3Stacked-AE0.98610.98311.00001.00000.99830.99780.9843
BGE-M3Transformer-AE1.00001.00001.00001.00001.00001.00001.0000
BGE-M3xLSTM-AE1.00001.00001.00001.00001.00001.00001.0000
BGE-M3Mamba-AE1.00001.00001.00001.00001.00001.00001.0000
Qwen-EmbBiLSTM-AE0.95770.89530.90780.94580.99710.99630.8672
Qwen-EmbSimple-AE0.95710.90840.94960.97100.98550.98160.8962
Qwen-EmbStacked-AE0.98780.98860.90430.94360.99830.99780.8922
Qwen-EmbTransformer-AE0.99770.99820.99070.99471.00001.00000.9884
Qwen-EmbxLSTM-AE1.00001.00000.99360.99631.00001.00000.9936
Qwen-EmbMamba-AE1.00001.00000.99830.99901.00001.00000.9983
Table A3. Validation classification performance with uncertainty-based task weighting. Acc. denotes accuracy and FM denotes Full Match Accuracy.
Table A3. Validation classification performance with uncertainty-based task weighting. Acc. denotes accuracy and FM denotes Full Match Accuracy.
Text EncoderTrajectory AETrajectoryPerceptionDecisionFM
Acc.F1Acc.F1Acc.F1
EmbeddingGemmaBiLSTM-AE0.83070.61390.99880.99930.97910.97310.8249
EmbeddingGemmaSimple-AE0.89620.78700.99940.99970.97910.97310.8783
EmbeddingGemmaStacked-AE0.95770.93290.99940.99970.99590.99490.9542
EmbeddingGemmaTransformer-AE0.99420.99551.00001.00000.99880.99850.9930
EmbeddingGemmaxLSTM-AE0.99940.99951.00001.00001.00001.00000.9994
EmbeddingGemmaMamba-AE1.00001.00001.00001.00001.00001.00001.0000
MiniLMBiLSTM-AE0.83540.60790.86610.92820.98320.97840.7183
MiniLMSimple-AE0.92120.81210.96120.97740.98490.98080.8736
MiniLMStacked-AE0.94490.87740.86610.92820.99360.99190.8104
MiniLMTransformer-AE1.00001.00000.98960.99401.00001.00000.9896
MiniLMxLSTM-AE1.00001.00001.00001.00001.00001.00001.0000
MiniLMMamba-AE1.00001.00001.00001.00001.00001.00001.0000
BGE-M3BiLSTM-AE0.85510.59730.94030.96520.98380.97930.8012
BGE-M3Simple-AE0.90550.78780.99540.99730.98090.97530.8864
BGE-M3Stacked-AE0.94090.90810.90380.94440.99250.99040.8429
BGE-M3Transformer-AE1.00001.00001.00001.00001.00001.00001.0000
BGE-M3xLSTM-AE1.00001.00001.00001.00001.00001.00001.0000
BGE-M3Mamba-AE1.00001.00001.00001.00001.00001.00001.0000
Qwen-EmbBiLSTM-AE0.82840.56920.86610.92820.99250.99050.7165
Qwen-EmbSimple-AE0.89620.76520.86610.92820.97220.96390.7588
Qwen-EmbStacked-AE0.94320.92680.86610.92820.99070.98810.8058
Qwen-EmbTransformer-AE0.99010.98260.86610.92821.00001.00000.8562
Qwen-EmbxLSTM-AE1.00001.00000.99540.99731.00001.00000.9954
Qwen-EmbMamba-AE1.00001.00000.99300.99601.00001.00000.9930

References

  1. Chen, W.; Li, A.; Jiang, H. Risk Assessment of Roundabout Scenarios in Virtual Testing Based on an Improved Driving Safety Field. Sensors 2024, 24, 5539. [Google Scholar] [CrossRef] [Scilit]
  2. Thal, S.; Wallis, P.; Henze, R.; Hasegawa, R.; Nakamura, H.; Kitajima, S.; Abe, G. Towards Realistic, Safety-Critical and Complete Test Case Catalogs for Safe Automated Driving in Urban Scenarios. In Proceedings of the 2023 IEEE Intelligent Vehicles Symposium (IV); IEEE: New York, NY, USA, 2023; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
  3. Barabás, I.; Todoruţ, A.; Cordoş, N.; Molea, A. Current Challenges in Autonomous Driving. IOP Conf. Ser. Mater. Sci. Eng. 2017, 252, 012096. [Google Scholar] [CrossRef] [Scilit]
  4. ISO 26262-1:2018; Road Vehicles—Functional Safety. International Organization for Standardization: Geneva, Switzerland, 2018. Available online: https://www.iso.org/standard/68383.html (accessed on 23 September 2026).
  5. De Gelder, E.; Manders, J.; Grappiolo, C.; Paardekooper, J.P.; Den Camp, O.O.; De Schutter, B. Real-world scenario mining for the assessment of automated vehicles. In Proceedings of the 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC); IEEE: New York, NY, USA, 2020; pp. 1–8. [Google Scholar]
  6. Zhang, X.; Tao, J.; Tan, K.; Törngren, M.; Sánchez, J.M.G.; Ramli, M.R.; Tao, X.; Gyllenhammar, M.; Wotawa, F.; Mohan, N.; et al. Finding critical scenarios for automated driving systems: A systematic mapping study. IEEE Trans. Softw. Eng. 2022, 49, 991–1026. [Google Scholar] [CrossRef] [Scilit]
  7. Dosovitskiy, A.; Ros, G.; Codevilla, F.; Lopez, A.; Koltun, V. CARLA: An Open Urban Driving Simulator. In Proceedings of the 1st Annual Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2017; pp. 1–16. [Google Scholar]
  8. Erdogan, A.; Ugranli, B.; Adali, E.; Sentas, A.; Mungan, E.; Kaplan, E.; Leitner, A. Real-world maneuver extraction for autonomous vehicle validation: A comparative study. In Proceedings of the 2019 IEEE Intelligent Vehicles Symposium (IV); IEEE: New York, NY, USA, 2019; pp. 267–272. [Google Scholar]
  9. Krajewski, R.; Bock, J.; Kloeker, L.; Eckstein, L. The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for Validation of Highly Automated Driving Systems. In Proceedings of the 2018 21st International Conference on Intelligent Transportation Systems (ITSC); IEEE: New York, NY, USA, 2018; pp. 2118–2125. [Google Scholar] [CrossRef] [Scilit]
  10. Tenbrock, A.; König, A.; Keutgens, T.; Weber, H. The conscend dataset: Concrete scenarios from the highd dataset according to alks regulation unece r157 in openx. In Proceedings of the 2021 IEEE Intelligent Vehicles Symposium Workshops (IV Workshops); IEEE: New York, NY, USA, 2021; pp. 174–181. [Google Scholar]
  11. Liu, M.; Bian, J.; Liu, X.; Huang, H.; Gui, G.; Zhou, R.; Gui, W. Learning Safety-Critical Scenarios from Real-World Pre-Crash Data for Autonomous Driving Safety Validation. Green Energy Intell. Transp. 2026, 100436. [Google Scholar] [CrossRef] [Scilit]
  12. Ku, S.; Lee, J. Rule-Based Scenario Classification Using Vehicle Trajectories. ISPRS Int. J. Geo-Inf. 2026, 15, 37. [Google Scholar] [CrossRef] [Scilit]
  13. Thorn, E.; Kimmel, S.C.; Chaka, M. A Framework for Automated Driving System Testable Cases and Scenarios; National Highway Traffic Safety Administration: Washington, DC, USA, 2018. [Google Scholar]
  14. Zhong, Z.; Tang, Y.; Zhou, Y.; Neves, V.D.O.; Liu, Y.; Ray, B. A survey on scenario-based testing for automated driving systems in high-fidelity simulation. arXiv 2021, arXiv:2112.00964. [Google Scholar]
  15. Bagschik, G.; Menzel, T.; Maurer, M. Ontology based scene creation for the development of automated vehicles. In Proceedings of the 2018 IEEE Intelligent Vehicles Symposium (IV); IEEE: New York, NY, USA, 2018; pp. 1813–1820. [Google Scholar]
  16. Balsa-Barreiro, J.; Valero-Mora, P.M.; Menéndez, M.; Mehmood, R. Extraction of Naturalistic Driving Patterns with Geographic Information Systems. Mob. Netw. Appl. 2023, 28, 619–635. [Google Scholar] [CrossRef] [Scilit]
  17. Routray, S.K. Visualization and Visual Analytics in Autonomous Driving. IEEE Comput. Graph. Appl. 2024, 44, 43–53. [Google Scholar] [CrossRef] [Scilit]
  18. Li, H.; Wang, J.; Yuan, J.; Li, Y.; Weng, W.; Peng, Y.; Zhang, Y.; Xiong, Z.; Sun, X. Event-assisted Low-Light Video Object Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 3250–3259. [Google Scholar]
  19. Park, J.I.; Jo, S.; Seo, H.T.; Park, J. LiDAR Denoising Methods in Adverse Environments: A Review. IEEE Sens. J. 2025, 25, 7916–7932. [Google Scholar] [CrossRef] [Scilit]
  20. Gourova, R.; Krasnov, O.; Yarovoy, A. Analysis of Rain Clutter Detections in Commercial 77 GHz Automotive Radar. In Proceedings of the 2017 European Radar Conference (EURAD); IEEE: New York, NY, USA, 2017; pp. 25–28. [Google Scholar] [CrossRef] [Scilit]
  21. Sezgin, F.; Vriesman, D.; Steinhauser, D.; Lugner, R.; Brandmeier, T. Safe Autonomous Driving in Adverse Weather: Sensor Evaluation and Performance Monitoring. In Proceedings of the 2023 IEEE Intelligent Vehicles Symposium (IV); IEEE: New York, NY, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  22. Liu, T.; Zhu, H.; Shen, Y.; Xu, L.; Feng, S. Impact of Tunnel Lighting on Driver Perception and Safety in Foggy Conditions. Build. Environ. 2025, 285, 113540. [Google Scholar] [CrossRef] [Scilit]
  23. Chen, S.; Lin, Y.; Zhao, H. Effects of Flicker with Various Brightness Contrasts on Visual Fatigue in Road Lighting Using Fixed Low-Mounting-Height Luminaires. Tunn. Undergr. Space Technol. 2023, 136, 105091. [Google Scholar] [CrossRef] [Scilit]
  24. Uricar, M.; Krizek, P.; Sistu, G.; Yogamani, S. SoilingNet: Soiling Detection on Automotive Surround-View Cameras. In Proceedings of the 2019 IEEE Intelligent Transportation Systems Conference (ITSC); IEEE: New York, NY, USA, 2019; pp. 67–72. [Google Scholar] [CrossRef] [Scilit]
  25. Lee, M.; Jo, K.; Sunwoo, M. Collision Risk Assessment for Possible Collision Vehicle in Occluded Area Based on Precise Map. In Proceedings of the 2017 IEEE 20th International Conference on Intelligent Transportation Systems (ITSC); IEEE: New York, NY, USA, 2017; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  26. Fu, C.; Lu, Z.; Ding, N.; Bai, W. Distance Headway-Based Safety Evaluation of Emerging Mixed Traffic Flow under Snowy Weather. Phys. A Stat. Mech. Its Appl. 2024, 642, 129792. [Google Scholar] [CrossRef] [Scilit]
  27. Shetty, A.; Tavafoghi, H.; Kurzhanskiy, A.; Poolla, K.; Varaiya, P. Risk Assessment of Autonomous Vehicles across Diverse Driving Contexts. In Proceedings of the 2021 IEEE International Intelligent Transportation Systems Conference (ITSC); IEEE: New York, NY, USA, 2021; pp. 712–719. [Google Scholar] [CrossRef] [Scilit]
  28. Hou, Y.; Wang, C.; Wang, J.; Xue, X.; Zhang, X.L.; Zhu, J.; Wang, D.; Chen, S. Visual Evaluation for Autonomous Driving. IEEE Trans. Vis. Comput. Graph. 2021, 28, 1030–1039. [Google Scholar] [CrossRef] [Scilit]
  29. Euro NCAP. AEB Car-to-Car Systems Test Protocol; Technical report; European New Car Assessment Programme: Leuven, Belgium, 2022. [Google Scholar]
  30. Ministry of Land, Infrastructure and Transport. Intersection Design Guidelines; Technical Report 11-1613000-00162-01; Ministry of Land, Infrastructure and Transport: Sejong City, Republic of Korea, 2025. (In Korean) [Google Scholar]
  31. So, J.J.; Park, I.; Wee, J.; Park, S.; Yun, I. Generating Traffic Safety Test Scenarios for Automated Vehicles Using a Big Data Technique. KSCE J. Civ. Eng. 2019, 23, 2702–2712. [Google Scholar] [CrossRef] [Scilit]
  32. Samak, T.V.; Samak, C.V.; Krovi, V.N. Towards Validation of Autonomous Vehicles Across Scales Using an Integrated Digital Twin Framework. In Proceedings of the 2024 IEEE International Conference on Advanced Intelligent Mechatronics (AIM), Boston, MA, USA, 15–19 July 2024; IEEE: New York, NY, USA, 2024; pp. 1068–1075. [Google Scholar] [CrossRef] [Scilit]
  33. Samak, T.; Samak, C.; Brault, J.; Harber, C.; McCane, K.; Smereka, J.; Brudnak, M.; Gorsich, D.; Krovi, V. A Systematic Digital Engineering Approach to Verification & Validation of Autonomous Ground Vehicles in Off-Road Environments. IFAC-PapersOnLine 2025, 59, 797–802. [Google Scholar] [CrossRef] [Scilit]
  34. Casola, S.; Lauriola, I.; Lavelli, A. Pre-Trained Transformers: An Empirical Comparison. Mach. Learn. Appl. 2022, 9, 100334. [Google Scholar] [CrossRef] [Scilit]
  35. Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Minneapolis, MN, USA, 2019; Volume 1 (Long and Short Papers), pp. 4171–4186. [Google Scholar]
  36. Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
  37. Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I. Improving Language Understanding by Generative Pre-Training; OpenAI: San Francisco, CA, USA, 2018. [Google Scholar]
  38. Vera, H.S.; Dua, S.; Zhang, B.; Salz, D.; Mullins, R.; Panyam, S.R.; Smoot, S.; Naim, I.; Zou, J.; Chen, F.; et al. EmbeddingGemma: Powerful and Lightweight Text Representations. arXiv 2025, arXiv:2509.20354. [Google Scholar] [CrossRef] [Scilit]
  39. Wang, W.; Wei, F.; Dong, L.; Bao, H.; Yang, N.; Zhou, M. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. Adv. Neural Inf. Process. Syst. 2020, 33, 5776–5788. [Google Scholar]
  40. Chen, J.; Xiao, S.; Zhang, P.; Luo, K.; Lian, D.; Liu, Z. BGE M3-Embedding: Multi-Lingual, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. arXiv 2024, arXiv:2402.03216. [Google Scholar]
  41. Zhang, Y.; Li, M.; Long, D.; Zhang, X.; Lin, H.; Yang, B.; Xie, P.; Yang, A.; Liu, D.; Lin, J.; et al. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv 2025, arXiv:2506.05176. [Google Scholar]
  42. Irani, H.; Ghahremani, Y.; Kermani, A.; Metsis, V. Time series embedding methods for classification tasks: A review. Expert Syst. 2025, 42, e70148. [Google Scholar] [CrossRef] [Scilit]
  43. Wu, S.; Song, S.; Deng, S.; Xie, W.; Shen, L. Variable-length time series classification: Benchmarking, analysis and effective spectral pooling strategy. Inf. Fusion 2025, 126, 103584. [Google Scholar] [CrossRef] [Scilit]
  44. Lee, H.; Shin, D. Beyond Information Distortion: Imaging Variable-Length Time Series Data for Classification. Sensors 2025, 25, 621. [Google Scholar] [CrossRef] [Scilit]
  45. Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit]
  46. Rossi, L.; Ajmar, A.; Paolanti, M.; Pierdicca, R. Vehicle trajectory prediction and generation using LSTM models and GANs. PLoS ONE 2021, 16, e0253868. [Google Scholar] [CrossRef] [Scilit]
  47. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. arXiv 2023, arXiv:1706.03762. [Google Scholar] [CrossRef] [Scilit]
  48. Zerveas, G.; Jayaraman, S.; Patel, D.; Bhamidipaty, A.; Eickhoff, C. A Transformer-based Framework for Multivariate Time Series Representation Learning. arXiv 2020, arXiv:2010.02803. [Google Scholar] [CrossRef] [Scilit]
  49. Beck, M.; Pöppel, K.; Spanring, M.; Auer, A.; Prudnikova, O.; Kopp, M.; Klambauer, G.; Brandstetter, J.; Hochreiter, S. xLSTM: Extended Long Short-Term Memory. arXiv 2024, arXiv:2405.04517. [Google Scholar] [CrossRef] [Scilit]
  50. Gu, A.; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv 2024, arXiv:2312.00752. [Google Scholar] [CrossRef] [Scilit]
  51. Xu, Y.; Shao, W.; Li, J.; Yang, K.; Wang, W.; Huang, H.; Lv, C.; Wang, H. SIND: A drone dataset at signalized intersection in China. arXiv 2022, arXiv:2209.02297. [Google Scholar]
  52. Hinton, G.E.; Salakhutdinov, R.R. Reducing the dimensionality of data with neural networks. Science 2006, 313, 504–507. [Google Scholar] [CrossRef] [Scilit]
  53. Narmadha, S.; Balaji, N. Improved network anomaly detection system using optimized autoencoder- LSTM. Expert Syst. With Appl. 2025, 273, 126854. [Google Scholar] [CrossRef] [Scilit]
  54. Carter, N.; Beier, S.; Cordero, R. Lateral and Tangential Accelerations of Left Turning Vehicles from Naturalistic Observations; SAE International: Detroit, MI, USA, 2019. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Illustration of the relationship between OpenDRIVE, OpenSCENARIO, and the simulator in a simulation environment. OpenDRIVE provides road geometry, lane structure, and traffic signal information, forming the basis on which driving scenarios are executed. OpenSCENARIO defines the scenario entities, their initial states, behaviors, goals, and interactions within the OpenDRIVE road network. Based on the road and scenario configurations, the simulator executes the scenario using its physics engine and outputs the resulting vehicle trajectories and motion data.
Figure 1. Illustration of the relationship between OpenDRIVE, OpenSCENARIO, and the simulator in a simulation environment. OpenDRIVE provides road geometry, lane structure, and traffic signal information, forming the basis on which driving scenarios are executed. OpenSCENARIO defines the scenario entities, their initial states, behaviors, goals, and interactions within the OpenDRIVE road network. Based on the road and scenario configurations, the simulator executes the scenario using its physics engine and outputs the resulting vehicle trajectories and motion data.
Ijgi 15 00443 g001
Figure 2. Road network used for simulation. White markings indicate lane boundaries, stop lines, and pedestrian crossings. Red dots indicate the endpoints of road segments. Red dashed lines represent straight connections between road segments, while orange dashed lines indicate connections involving turns or direction changes. Because the map does not include U-turn connections, U-turn maneuvers are not considered in the generated simulation scenarios.
Figure 2. Road network used for simulation. White markings indicate lane boundaries, stop lines, and pedestrian crossings. Red dots indicate the endpoints of road segments. Red dashed lines represent straight connections between road segments, while orange dashed lines indicate connections involving turns or direction changes. Because the map does not include U-turn connections, U-turn maneuvers are not considered in the generated simulation scenarios.
Ijgi 15 00443 g002
Figure 3. Environmental variations in the simulation environment. The conditions include dawn, afternoon, evening, night, sunny, rain, snow, and fog. Among these, dawn and night are treated as low-illumination conditions, while rain, snow, and fog are treated as adverse weather conditions associated with perception disturbances.
Figure 3. Environmental variations in the simulation environment. The conditions include dawn, afternoon, evening, night, sunny, rain, snow, and fog. Among these, dawn and night are treated as low-illumination conditions, while rain, snow, and fog are treated as adverse weather conditions associated with perception disturbances.
Ijgi 15 00443 g003
Figure 4. Visualization of vehicle trajectories and background in the SinD dataset (Changchun). Vehicle trajectories collected from the Changchun scenario in the SinD dataset are overlaid on a background map. Each trajectory is categorized by maneuver type: straight (green), left turn (red), right turn (blue), and U-turn (orange). Due to imperfect alignment between the background image and the coordinate system, minor visual discrepancies may be present.
Figure 4. Visualization of vehicle trajectories and background in the SinD dataset (Changchun). Vehicle trajectories collected from the Changchun scenario in the SinD dataset are overlaid on a background map. Each trajectory is categorized by maneuver type: straight (green), left turn (red), right turn (blue), and U-turn (orange). Due to imperfect alignment between the background image and the coordinate system, minor visual discrepancies may be present.
Ijgi 15 00443 g004
Figure 5. Encoder–Bottleneck–Decoder framework shared by the evaluated Autoencoders for variable-length trajectory representation learning. The encoder compresses the input sequence into a latent vector, and the decoder reconstructs the sequence. Padded time steps are excluded from the reconstruction loss through masking.
Figure 5. Encoder–Bottleneck–Decoder framework shared by the evaluated Autoencoders for variable-length trajectory representation learning. The encoder compresses the input sequence into a latent vector, and the decoder reconstructs the sequence. Padded time steps are excluded from the reconstruction loss through masking.
Ijgi 15 00443 g005
Figure 6. Two-dimensional t-SNE projections of latent representations from six trajectory Autoencoders after PCA preprocessing. Orange circles, green squares, and purple triangles denote ST, LT, and RT, respectively. Darker and lighter shades indicate decision labels of 0 and 1.
Figure 6. Two-dimensional t-SNE projections of latent representations from six trajectory Autoencoders after PCA preprocessing. Orange circles, green squares, and purple triangles denote ST, LT, and RT, respectively. Darker and lighter shades indicate decision labels of 0 and 1.
Ijgi 15 00443 g006
Figure 7. Overall model architecture for scenario classification. The framework is composed of 24 comparative configurations, combining six types of autoencoders trained on variable-length time-series data from the SinD dataset and four types of text encoders. Time-series data obtained from simulation are encoded into latent vectors, while OpenSCENARIO inputs are preprocessed and embedded using a text encoder. The resulting representations are concatenated and fed into a classifier for joint prediction of three targets: trajectory type, perception disturbance, and decision disturbance. Blue denotes the vehicle-motion target (trajectory), while red denotes the disturbance-related targets (perception and decision).
Figure 7. Overall model architecture for scenario classification. The framework is composed of 24 comparative configurations, combining six types of autoencoders trained on variable-length time-series data from the SinD dataset and four types of text encoders. Time-series data obtained from simulation are encoded into latent vectors, while OpenSCENARIO inputs are preprocessed and embedded using a text encoder. The resulting representations are concatenated and fed into a classifier for joint prediction of three targets: trajectory type, perception disturbance, and decision disturbance. Blue denotes the vehicle-motion target (trajectory), while red denotes the disturbance-related targets (perception and decision).
Ijgi 15 00443 g007
Table 1. Parameter configurations for generating the training and validation data and the separately generated test data. Environmental values are shown as paired road friction coefficient ( μ ) and visibility range (V).
Table 1. Parameter configurations for generating the training and validation data and the separately generated test data. Environmental values are shown as paired road friction coefficient ( μ ) and visibility range (V).
CategoryParameterTraining and ValidationTest
TimeTime of day06:00, 14:00, 19:00, 22:0003:00, 12:00, 23:00
Initial SpeedEgo speed8.33–16.67 m/s;
step: ≈1.39 m/s (5 km/h)
8.33–16.67 m/s;
step: ≈2.78 m/s (10 km/h)
NPC speed8.33–16.67 m/s;
step: ≈1.39 m/s (5 km/h)
8.33–16.67 m/s;
step: ≈2.78 m/s (10 km/h)
EnvironmentRain μ = 0.7 , V = 200 m μ = 0.8 , V = 250 m
Snow μ = 0.4 , V = 200 m μ = 0.6 , V = 150 m
Sunny μ = 1.0 , V = 5000 m μ = 1.0 , V = 4000 m
Fog μ = 1.0 , V = 50 m μ = 1.0 , V = 80 m
Table 2. Major Features in the Driving Trajectory Dataset.
Table 2. Major Features in the Driving Trajectory Dataset.
FeatureType/Unit
Times (sec)
EntityString
Velocity X (entity coordinate)km/h
Velocity Y (entity coordinate)km/h
Velocity Z (entity coordinate)km/h
Acceleration X (entity coordinate)m/s2
Acceleration Y (entity coordinate)m/s2
Acceleration Z (entity coordinate)m/s2
Table 3. Domain distribution metrics between SinD and simulation data. The reference consists of standardized per-log summaries of velocity and acceleration features. Comparisons across representations are descriptive.
Table 3. Domain distribution metrics between SinD and simulation data. The reference consists of standardized per-log summaries of velocity and acceleration features. Comparisons across representations are descriptive.
RepresentationDomain SilhouetteMMD2
Log-summary reference0.21120.3618
BiLSTM-AE0.17390.2165
Simple-AE0.12790.1821
Stacked-AE0.21110.2368
Transformer-AE0.22240.2376
xLSTM-AE0.15250.1827
Mamba-AE0.18260.2610
Table 4. Text Encoders Used for Comparative Evaluation.
Table 4. Text Encoders Used for Comparative Evaluation.
ModelEmbedding Dim.Parameters
EmbeddingGemma-300M768307.6M
all-MiniLM-L6-v238422.7M
BGE-M31024567.8M
Qwen3-Embedding-0.6B1024595.8M
Table 5. Test classification performance with fixed task weighting. F1 denotes trajectory macro F1; FM denotes Full Match Accuracy. Bold values identify the highest reported score in each column.
Table 5. Test classification performance with fixed task weighting. F1 denotes trajectory macro F1; FM denotes Full Match Accuracy. Bold values identify the highest reported score in each column.
Text EncoderTrajectory AETraj. Acc.F1Perc. Acc.Dec. Acc.FM
EmbeddingGemmaBiLSTM-AE0.83540.85550.91670.99920.7641
EmbeddingGemmaSimple-AE0.95830.93390.91670.99900.8771
EmbeddingGemmaStacked-AE0.74740.75540.91670.99970.6849
EmbeddingGemmaTransformer-AE0.82760.82120.91670.98440.7435
EmbeddingGemmaxLSTM-AE0.61930.34710.91670.73650.4081
EmbeddingGemmaMamba-AE0.99970.99900.91670.99920.9156
MiniLMBiLSTM-AE0.87140.77480.60860.99970.5414
MiniLMSimple-AE0.97920.96600.29530.99970.2917
MiniLMStacked-AE0.92010.93570.80960.99970.7516
MiniLMTransformer-AE0.88050.90710.89790.99790.7893
MiniLMxLSTM-AE0.65570.53440.85160.93520.5169
MiniLMMamba-AE0.99970.99900.57210.99920.5719
BGE-M3BiLSTM-AE0.97240.96920.75441.00000.7310
BGE-M3Simple-AE0.97630.96920.79400.99900.7729
BGE-M3Stacked-AE0.87400.85280.83980.99900.7312
BGE-M3Transformer-AE0.88070.91270.83230.99480.7299
BGE-M3xLSTM-AE0.66250.52890.80080.84450.4792
BGE-M3Mamba-AE0.99950.99800.86740.99900.8659
Qwen-EmbBiLSTM-AE0.96410.95980.91671.00000.8831
Qwen-EmbSimple-AE0.96410.96370.91670.99970.8833
Qwen-EmbStacked-AE0.88230.85960.91670.99870.8078
Qwen-EmbTransformer-AE0.83390.75730.91670.99840.7625
Qwen-EmbxLSTM-AE0.65940.52840.91670.88540.5357
Qwen-EmbMamba-AE0.99970.99900.91670.99970.9161
Table 6. Test classification performance with uncertainty-based task weighting. F1 denotes trajectory macro F1; FM denotes Full Match Accuracy. Bold values identify the highest reported score in each column.
Table 6. Test classification performance with uncertainty-based task weighting. F1 denotes trajectory macro F1; FM denotes Full Match Accuracy. Bold values identify the highest reported score in each column.
Text EncoderTrajectory AETraj. Acc.F1Perc. Acc.Dec. Acc.FM
EmbeddingGemmaBiLSTM-AE0.85100.87450.91670.99970.7784
EmbeddingGemmaSimple-AE0.93720.88020.91671.00000.8594
EmbeddingGemmaStacked-AE0.76720.77620.91670.99900.7008
EmbeddingGemmaTransformer-AE0.81350.81150.91640.97920.7263
EmbeddingGemmaxLSTM-AE0.64400.52580.91590.76020.4315
EmbeddingGemmaMamba-AE0.99950.99800.91670.99920.9154
MiniLMBiLSTM-AE0.86220.89130.26900.99970.2195
MiniLMSimple-AE0.98180.97870.60000.99970.5888
MiniLMStacked-AE0.89190.89570.40760.99970.3773
MiniLMTransformer-AE0.86200.87730.89300.99610.7693
MiniLMxLSTM-AE0.66280.54780.82860.87290.5255
MiniLMMamba-AE0.99950.99800.30520.99900.3036
BGE-M3BiLSTM-AE0.91460.92870.81591.00000.7531
BGE-M3Simple-AE0.97920.97600.84220.99970.8224
BGE-M3Stacked-AE0.89010.90900.86090.99900.7643
BGE-M3Transformer-AE0.87710.90810.70730.99580.6174
BGE-M3xLSTM-AE0.69240.64210.73700.83750.4401
BGE-M3Mamba-AE0.99950.99800.85780.99900.8562
Qwen-EmbBiLSTM-AE0.96820.96640.91671.00000.8867
Qwen-EmbSimple-AE0.96800.96600.91670.99970.8870
Qwen-EmbStacked-AE0.82400.81720.91670.99870.7570
Qwen-EmbTransformer-AE0.86430.89480.91670.98020.7732
Qwen-EmbxLSTM-AE0.66670.55560.91670.71850.4190
Qwen-EmbMamba-AE0.99950.99800.91671.00000.9161
Table 7. Mean classification performance by trajectory encoder and text encoder under fixed and uncertainty-based weighting. Each trajectory encoder mean is computed over four text encoders; each text encoder mean is computed over six trajectory encoders. Trajectory F1 is macro F1 over ST/LT/RT; perception and decision F1 are positive-class F1 scores. Bold values identify the highest reported score in each column.
Table 7. Mean classification performance by trajectory encoder and text encoder under fixed and uncertainty-based weighting. Each trajectory encoder mean is computed over four text encoders; each text encoder mean is computed over six trajectory encoders. Trajectory F1 is macro F1 over ST/LT/RT; perception and decision F1 are positive-class F1 scores. Bold values identify the highest reported score in each column.
EncoderWeightingTrajectoryPerceptionDecisionFM
Acc.F1Acc.F1Acc.F1
BiLSTM-AEFixed0.91080.88980.79910.87370.99970.99960.7299
Uncertainty0.89900.91520.72960.78470.99990.99980.6594
Simple-AEFixed0.96950.95820.73070.79060.99930.99910.7063
Uncertainty0.96650.95020.81890.88680.99980.99970.7894
Stacked-AEFixed0.85590.85090.87070.92690.99930.99900.7439
Uncertainty0.84330.84950.77550.84030.99910.99870.6499
Transformer-AEFixed0.85570.84960.89090.93960.99390.99140.7563
Uncertainty0.85420.87290.85830.91620.98780.98310.7215
xLSTM-AEFixed0.64920.48470.87140.92820.85040.72430.4850
Uncertainty0.66650.56780.84950.91330.79730.67800.4540
Mamba-AEFixed0.99970.99880.81820.88400.99930.99900.8174
Uncertainty0.99950.99800.74910.80480.99930.99900.7479
EmbeddingGemmaFixed0.83130.78530.91670.95650.95300.90050.7322
Uncertainty0.83540.81100.91650.95640.95620.90990.7353
MiniLMFixed0.88440.85280.67250.76090.98860.98250.5771
Uncertainty0.87670.86480.55060.63860.97790.96980.4640
BGE-M3Fixed0.89420.87180.81480.88800.97270.95540.7184
Uncertainty0.89210.89360.80350.87920.97180.94940.7089
Qwen-EmbFixed0.88390.84470.91670.95650.98030.96990.7981
Uncertainty0.88180.86630.91670.95650.94950.94320.7732
Table 8. Reported validation and test performance averaged over six trajectory encoders for each text encoder. All denotes the unweighted mean over all 24 encoder combinations. Perc. denotes perception accuracy and FM denotes Full Match Accuracy.
Table 8. Reported validation and test performance averaged over six trajectory encoders for each text encoder. All denotes the unweighted mean over all 24 encoder combinations. Perc. denotes perception accuracy and FM denotes Full Match Accuracy.
Text EncoderWeightingPerc. Val.Perc. TestFM Val.FM Test
EmbeddingGemmaFixed1.00000.91670.97570.7322
EmbeddingGemmaUncertainty0.99960.91650.94160.7353
MiniLMFixed0.99380.67250.97270.5771
MiniLMUncertainty0.94710.55060.89860.4640
BGE-M3Fixed1.00000.81480.97580.7184
BGE-M3Uncertainty0.97320.80350.92170.7089
Qwen-EmbFixed0.95740.91670.93930.7981
Qwen-EmbUncertainty0.90880.91670.85430.7732
AllFixed0.98780.83020.96590.7064
AllUncertainty0.95720.79680.90410.6704
Table 9. Latent-vector ablation results averaged over four text encoder pairings. M: full multimodal input; ZT: text latent vector set to zero; ZR: trajectory latent vector set to zero. Bold values identify the highest reported score in each column. All reported values are accuracies.
Table 9. Latent-vector ablation results averaged over four text encoder pairings. M: full multimodal input; ZT: text latent vector set to zero; ZR: trajectory latent vector set to zero. Bold values identify the highest reported score in each column. All reported values are accuracies.
WeightingTrajectory AETraj. MTraj. ZTPerc. MPerc. ZRDec. MDec. ZR
Fixed            BiLSTM-AE0.91080.50330.79910.89060.99970.3952
FixedSimple-AE0.96950.72010.73070.84110.99930.3555
FixedStacked-AE0.85590.38310.87070.91670.99930.3555
FixedTransformer-AE0.85570.41980.89090.91930.99390.3555
FixedxLSTM-AE0.64920.60010.87140.91670.85040.3555
FixedMamba-AE0.99970.99970.81820.75520.99930.6445
UncertaintyBiLSTM-AE0.89900.39000.72960.80860.99990.3555
UncertaintySimple-AE0.96650.65650.81890.92320.99980.3555
UncertaintyStacked-AE0.84330.37620.77550.91670.99910.3555
UncertaintyTransformer-AE0.85420.44220.85830.87370.98780.3555
UncertaintyxLSTM-AE0.66650.61330.84950.87370.79730.3555
UncertaintyMamba-AE0.99950.99940.74910.70050.99930.6445
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ku, S.; Lee, J. A Multi-Task Framework for Vehicle Trajectory and Disturbance Classification in Autonomous Driving Scenarios. ISPRS Int. J. Geo-Inf. 2026, 15, 443. https://doi.org/10.3390/ijgi15100443

AMA Style

Ku S, Lee J. A Multi-Task Framework for Vehicle Trajectory and Disturbance Classification in Autonomous Driving Scenarios. ISPRS International Journal of Geo-Information. 2026; 15(10):443. https://doi.org/10.3390/ijgi15100443

Chicago/Turabian Style

Ku, Sungmo, and Jinho Lee. 2026. "A Multi-Task Framework for Vehicle Trajectory and Disturbance Classification in Autonomous Driving Scenarios" ISPRS International Journal of Geo-Information 15, no. 10: 443. https://doi.org/10.3390/ijgi15100443

APA Style

Ku, S., & Lee, J. (2026). A Multi-Task Framework for Vehicle Trajectory and Disturbance Classification in Autonomous Driving Scenarios. ISPRS International Journal of Geo-Information, 15(10), 443. https://doi.org/10.3390/ijgi15100443

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop