Figure 1.
Illustration of the relationship between OpenDRIVE, OpenSCENARIO, and the simulator in a simulation environment. OpenDRIVE provides road geometry, lane structure, and traffic signal information, forming the basis on which driving scenarios are executed. OpenSCENARIO defines the scenario entities, their initial states, behaviors, goals, and interactions within the OpenDRIVE road network. Based on the road and scenario configurations, the simulator executes the scenario using its physics engine and outputs the resulting vehicle trajectories and motion data.
Figure 1.
Illustration of the relationship between OpenDRIVE, OpenSCENARIO, and the simulator in a simulation environment. OpenDRIVE provides road geometry, lane structure, and traffic signal information, forming the basis on which driving scenarios are executed. OpenSCENARIO defines the scenario entities, their initial states, behaviors, goals, and interactions within the OpenDRIVE road network. Based on the road and scenario configurations, the simulator executes the scenario using its physics engine and outputs the resulting vehicle trajectories and motion data.
Figure 2.
Road network used for simulation. White markings indicate lane boundaries, stop lines, and pedestrian crossings. Red dots indicate the endpoints of road segments. Red dashed lines represent straight connections between road segments, while orange dashed lines indicate connections involving turns or direction changes. Because the map does not include U-turn connections, U-turn maneuvers are not considered in the generated simulation scenarios.
Figure 2.
Road network used for simulation. White markings indicate lane boundaries, stop lines, and pedestrian crossings. Red dots indicate the endpoints of road segments. Red dashed lines represent straight connections between road segments, while orange dashed lines indicate connections involving turns or direction changes. Because the map does not include U-turn connections, U-turn maneuvers are not considered in the generated simulation scenarios.
Figure 3.
Environmental variations in the simulation environment. The conditions include dawn, afternoon, evening, night, sunny, rain, snow, and fog. Among these, dawn and night are treated as low-illumination conditions, while rain, snow, and fog are treated as adverse weather conditions associated with perception disturbances.
Figure 3.
Environmental variations in the simulation environment. The conditions include dawn, afternoon, evening, night, sunny, rain, snow, and fog. Among these, dawn and night are treated as low-illumination conditions, while rain, snow, and fog are treated as adverse weather conditions associated with perception disturbances.
Figure 4.
Visualization of vehicle trajectories and background in the SinD dataset (Changchun). Vehicle trajectories collected from the Changchun scenario in the SinD dataset are overlaid on a background map. Each trajectory is categorized by maneuver type: straight (green), left turn (red), right turn (blue), and U-turn (orange). Due to imperfect alignment between the background image and the coordinate system, minor visual discrepancies may be present.
Figure 4.
Visualization of vehicle trajectories and background in the SinD dataset (Changchun). Vehicle trajectories collected from the Changchun scenario in the SinD dataset are overlaid on a background map. Each trajectory is categorized by maneuver type: straight (green), left turn (red), right turn (blue), and U-turn (orange). Due to imperfect alignment between the background image and the coordinate system, minor visual discrepancies may be present.
Figure 5.
Encoder–Bottleneck–Decoder framework shared by the evaluated Autoencoders for variable-length trajectory representation learning. The encoder compresses the input sequence into a latent vector, and the decoder reconstructs the sequence. Padded time steps are excluded from the reconstruction loss through masking.
Figure 5.
Encoder–Bottleneck–Decoder framework shared by the evaluated Autoencoders for variable-length trajectory representation learning. The encoder compresses the input sequence into a latent vector, and the decoder reconstructs the sequence. Padded time steps are excluded from the reconstruction loss through masking.
Figure 6.
Two-dimensional t-SNE projections of latent representations from six trajectory Autoencoders after PCA preprocessing. Orange circles, green squares, and purple triangles denote ST, LT, and RT, respectively. Darker and lighter shades indicate decision labels of 0 and 1.
Figure 6.
Two-dimensional t-SNE projections of latent representations from six trajectory Autoencoders after PCA preprocessing. Orange circles, green squares, and purple triangles denote ST, LT, and RT, respectively. Darker and lighter shades indicate decision labels of 0 and 1.
Figure 7.
Overall model architecture for scenario classification. The framework is composed of 24 comparative configurations, combining six types of autoencoders trained on variable-length time-series data from the SinD dataset and four types of text encoders. Time-series data obtained from simulation are encoded into latent vectors, while OpenSCENARIO inputs are preprocessed and embedded using a text encoder. The resulting representations are concatenated and fed into a classifier for joint prediction of three targets: trajectory type, perception disturbance, and decision disturbance. Blue denotes the vehicle-motion target (trajectory), while red denotes the disturbance-related targets (perception and decision).
Figure 7.
Overall model architecture for scenario classification. The framework is composed of 24 comparative configurations, combining six types of autoencoders trained on variable-length time-series data from the SinD dataset and four types of text encoders. Time-series data obtained from simulation are encoded into latent vectors, while OpenSCENARIO inputs are preprocessed and embedded using a text encoder. The resulting representations are concatenated and fed into a classifier for joint prediction of three targets: trajectory type, perception disturbance, and decision disturbance. Blue denotes the vehicle-motion target (trajectory), while red denotes the disturbance-related targets (perception and decision).
Table 1.
Parameter configurations for generating the training and validation data and the separately generated test data. Environmental values are shown as paired road friction coefficient () and visibility range (V).
Table 1.
Parameter configurations for generating the training and validation data and the separately generated test data. Environmental values are shown as paired road friction coefficient () and visibility range (V).
| Category | Parameter | Training and Validation | Test |
|---|
| Time | Time of day | 06:00, 14:00, 19:00, 22:00 | 03:00, 12:00, 23:00 |
| Initial Speed | Ego speed | 8.33–16.67 m/s; step: ≈1.39 m/s (5 km/h) | 8.33–16.67 m/s; step: ≈2.78 m/s (10 km/h) |
| NPC speed | 8.33–16.67 m/s; step: ≈1.39 m/s (5 km/h) | 8.33–16.67 m/s; step: ≈2.78 m/s (10 km/h) |
| Environment | Rain | , m | , m |
| Snow | , m | , m |
| Sunny | , m | , m |
| Fog | , m | , m |
Table 2.
Major Features in the Driving Trajectory Dataset.
Table 2.
Major Features in the Driving Trajectory Dataset.
| Feature | Type/Unit |
|---|
| Time | s (sec) |
| Entity | String |
| Velocity X (entity coordinate) | km/h |
| Velocity Y (entity coordinate) | km/h |
| Velocity Z (entity coordinate) | km/h |
| Acceleration X (entity coordinate) | m/s2 |
| Acceleration Y (entity coordinate) | m/s2 |
| Acceleration Z (entity coordinate) | m/s2 |
Table 3.
Domain distribution metrics between SinD and simulation data. The reference consists of standardized per-log summaries of velocity and acceleration features. Comparisons across representations are descriptive.
Table 3.
Domain distribution metrics between SinD and simulation data. The reference consists of standardized per-log summaries of velocity and acceleration features. Comparisons across representations are descriptive.
| Representation | Domain Silhouette | MMD2 |
|---|
| Log-summary reference | 0.2112 | 0.3618 |
| BiLSTM-AE | 0.1739 | 0.2165 |
| Simple-AE | 0.1279 | 0.1821 |
| Stacked-AE | 0.2111 | 0.2368 |
| Transformer-AE | 0.2224 | 0.2376 |
| xLSTM-AE | 0.1525 | 0.1827 |
| Mamba-AE | 0.1826 | 0.2610 |
Table 4.
Text Encoders Used for Comparative Evaluation.
Table 4.
Text Encoders Used for Comparative Evaluation.
| Model | Embedding Dim. | Parameters |
|---|
| EmbeddingGemma-300M | 768 | 307.6M |
| all-MiniLM-L6-v2 | 384 | 22.7M |
| BGE-M3 | 1024 | 567.8M |
| Qwen3-Embedding-0.6B | 1024 | 595.8M |
Table 5.
Test classification performance with fixed task weighting. F1 denotes trajectory macro F1; FM denotes Full Match Accuracy. Bold values identify the highest reported score in each column.
Table 5.
Test classification performance with fixed task weighting. F1 denotes trajectory macro F1; FM denotes Full Match Accuracy. Bold values identify the highest reported score in each column.
| Text Encoder | Trajectory AE | Traj. Acc. | F1 | Perc. Acc. | Dec. Acc. | FM |
|---|
| EmbeddingGemma | BiLSTM-AE | 0.8354 | 0.8555 | 0.9167 | 0.9992 | 0.7641 |
| EmbeddingGemma | Simple-AE | 0.9583 | 0.9339 | 0.9167 | 0.9990 | 0.8771 |
| EmbeddingGemma | Stacked-AE | 0.7474 | 0.7554 | 0.9167 | 0.9997 | 0.6849 |
| EmbeddingGemma | Transformer-AE | 0.8276 | 0.8212 | 0.9167 | 0.9844 | 0.7435 |
| EmbeddingGemma | xLSTM-AE | 0.6193 | 0.3471 | 0.9167 | 0.7365 | 0.4081 |
| EmbeddingGemma | Mamba-AE | 0.9997 | 0.9990 | 0.9167 | 0.9992 | 0.9156 |
| MiniLM | BiLSTM-AE | 0.8714 | 0.7748 | 0.6086 | 0.9997 | 0.5414 |
| MiniLM | Simple-AE | 0.9792 | 0.9660 | 0.2953 | 0.9997 | 0.2917 |
| MiniLM | Stacked-AE | 0.9201 | 0.9357 | 0.8096 | 0.9997 | 0.7516 |
| MiniLM | Transformer-AE | 0.8805 | 0.9071 | 0.8979 | 0.9979 | 0.7893 |
| MiniLM | xLSTM-AE | 0.6557 | 0.5344 | 0.8516 | 0.9352 | 0.5169 |
| MiniLM | Mamba-AE | 0.9997 | 0.9990 | 0.5721 | 0.9992 | 0.5719 |
| BGE-M3 | BiLSTM-AE | 0.9724 | 0.9692 | 0.7544 | 1.0000 | 0.7310 |
| BGE-M3 | Simple-AE | 0.9763 | 0.9692 | 0.7940 | 0.9990 | 0.7729 |
| BGE-M3 | Stacked-AE | 0.8740 | 0.8528 | 0.8398 | 0.9990 | 0.7312 |
| BGE-M3 | Transformer-AE | 0.8807 | 0.9127 | 0.8323 | 0.9948 | 0.7299 |
| BGE-M3 | xLSTM-AE | 0.6625 | 0.5289 | 0.8008 | 0.8445 | 0.4792 |
| BGE-M3 | Mamba-AE | 0.9995 | 0.9980 | 0.8674 | 0.9990 | 0.8659 |
| Qwen-Emb | BiLSTM-AE | 0.9641 | 0.9598 | 0.9167 | 1.0000 | 0.8831 |
| Qwen-Emb | Simple-AE | 0.9641 | 0.9637 | 0.9167 | 0.9997 | 0.8833 |
| Qwen-Emb | Stacked-AE | 0.8823 | 0.8596 | 0.9167 | 0.9987 | 0.8078 |
| Qwen-Emb | Transformer-AE | 0.8339 | 0.7573 | 0.9167 | 0.9984 | 0.7625 |
| Qwen-Emb | xLSTM-AE | 0.6594 | 0.5284 | 0.9167 | 0.8854 | 0.5357 |
| Qwen-Emb | Mamba-AE | 0.9997 | 0.9990 | 0.9167 | 0.9997 | 0.9161 |
Table 6.
Test classification performance with uncertainty-based task weighting. F1 denotes trajectory macro F1; FM denotes Full Match Accuracy. Bold values identify the highest reported score in each column.
Table 6.
Test classification performance with uncertainty-based task weighting. F1 denotes trajectory macro F1; FM denotes Full Match Accuracy. Bold values identify the highest reported score in each column.
| Text Encoder | Trajectory AE | Traj. Acc. | F1 | Perc. Acc. | Dec. Acc. | FM |
|---|
| EmbeddingGemma | BiLSTM-AE | 0.8510 | 0.8745 | 0.9167 | 0.9997 | 0.7784 |
| EmbeddingGemma | Simple-AE | 0.9372 | 0.8802 | 0.9167 | 1.0000 | 0.8594 |
| EmbeddingGemma | Stacked-AE | 0.7672 | 0.7762 | 0.9167 | 0.9990 | 0.7008 |
| EmbeddingGemma | Transformer-AE | 0.8135 | 0.8115 | 0.9164 | 0.9792 | 0.7263 |
| EmbeddingGemma | xLSTM-AE | 0.6440 | 0.5258 | 0.9159 | 0.7602 | 0.4315 |
| EmbeddingGemma | Mamba-AE | 0.9995 | 0.9980 | 0.9167 | 0.9992 | 0.9154 |
| MiniLM | BiLSTM-AE | 0.8622 | 0.8913 | 0.2690 | 0.9997 | 0.2195 |
| MiniLM | Simple-AE | 0.9818 | 0.9787 | 0.6000 | 0.9997 | 0.5888 |
| MiniLM | Stacked-AE | 0.8919 | 0.8957 | 0.4076 | 0.9997 | 0.3773 |
| MiniLM | Transformer-AE | 0.8620 | 0.8773 | 0.8930 | 0.9961 | 0.7693 |
| MiniLM | xLSTM-AE | 0.6628 | 0.5478 | 0.8286 | 0.8729 | 0.5255 |
| MiniLM | Mamba-AE | 0.9995 | 0.9980 | 0.3052 | 0.9990 | 0.3036 |
| BGE-M3 | BiLSTM-AE | 0.9146 | 0.9287 | 0.8159 | 1.0000 | 0.7531 |
| BGE-M3 | Simple-AE | 0.9792 | 0.9760 | 0.8422 | 0.9997 | 0.8224 |
| BGE-M3 | Stacked-AE | 0.8901 | 0.9090 | 0.8609 | 0.9990 | 0.7643 |
| BGE-M3 | Transformer-AE | 0.8771 | 0.9081 | 0.7073 | 0.9958 | 0.6174 |
| BGE-M3 | xLSTM-AE | 0.6924 | 0.6421 | 0.7370 | 0.8375 | 0.4401 |
| BGE-M3 | Mamba-AE | 0.9995 | 0.9980 | 0.8578 | 0.9990 | 0.8562 |
| Qwen-Emb | BiLSTM-AE | 0.9682 | 0.9664 | 0.9167 | 1.0000 | 0.8867 |
| Qwen-Emb | Simple-AE | 0.9680 | 0.9660 | 0.9167 | 0.9997 | 0.8870 |
| Qwen-Emb | Stacked-AE | 0.8240 | 0.8172 | 0.9167 | 0.9987 | 0.7570 |
| Qwen-Emb | Transformer-AE | 0.8643 | 0.8948 | 0.9167 | 0.9802 | 0.7732 |
| Qwen-Emb | xLSTM-AE | 0.6667 | 0.5556 | 0.9167 | 0.7185 | 0.4190 |
| Qwen-Emb | Mamba-AE | 0.9995 | 0.9980 | 0.9167 | 1.0000 | 0.9161 |
Table 7.
Mean classification performance by trajectory encoder and text encoder under fixed and uncertainty-based weighting. Each trajectory encoder mean is computed over four text encoders; each text encoder mean is computed over six trajectory encoders. Trajectory F1 is macro F1 over ST/LT/RT; perception and decision F1 are positive-class F1 scores. Bold values identify the highest reported score in each column.
Table 7.
Mean classification performance by trajectory encoder and text encoder under fixed and uncertainty-based weighting. Each trajectory encoder mean is computed over four text encoders; each text encoder mean is computed over six trajectory encoders. Trajectory F1 is macro F1 over ST/LT/RT; perception and decision F1 are positive-class F1 scores. Bold values identify the highest reported score in each column.
| Encoder | Weighting | Trajectory | Perception | Decision | FM |
|---|
| Acc. | F1 | Acc. | F1 | Acc. | F1 |
|---|
| BiLSTM-AE | Fixed | 0.9108 | 0.8898 | 0.7991 | 0.8737 | 0.9997 | 0.9996 | 0.7299 |
| Uncertainty | 0.8990 | 0.9152 | 0.7296 | 0.7847 | 0.9999 | 0.9998 | 0.6594 |
| Simple-AE | Fixed | 0.9695 | 0.9582 | 0.7307 | 0.7906 | 0.9993 | 0.9991 | 0.7063 |
| Uncertainty | 0.9665 | 0.9502 | 0.8189 | 0.8868 | 0.9998 | 0.9997 | 0.7894 |
| Stacked-AE | Fixed | 0.8559 | 0.8509 | 0.8707 | 0.9269 | 0.9993 | 0.9990 | 0.7439 |
| Uncertainty | 0.8433 | 0.8495 | 0.7755 | 0.8403 | 0.9991 | 0.9987 | 0.6499 |
| Transformer-AE | Fixed | 0.8557 | 0.8496 | 0.8909 | 0.9396 | 0.9939 | 0.9914 | 0.7563 |
| Uncertainty | 0.8542 | 0.8729 | 0.8583 | 0.9162 | 0.9878 | 0.9831 | 0.7215 |
| xLSTM-AE | Fixed | 0.6492 | 0.4847 | 0.8714 | 0.9282 | 0.8504 | 0.7243 | 0.4850 |
| Uncertainty | 0.6665 | 0.5678 | 0.8495 | 0.9133 | 0.7973 | 0.6780 | 0.4540 |
| Mamba-AE | Fixed | 0.9997 | 0.9988 | 0.8182 | 0.8840 | 0.9993 | 0.9990 | 0.8174 |
| Uncertainty | 0.9995 | 0.9980 | 0.7491 | 0.8048 | 0.9993 | 0.9990 | 0.7479 |
| EmbeddingGemma | Fixed | 0.8313 | 0.7853 | 0.9167 | 0.9565 | 0.9530 | 0.9005 | 0.7322 |
| Uncertainty | 0.8354 | 0.8110 | 0.9165 | 0.9564 | 0.9562 | 0.9099 | 0.7353 |
| MiniLM | Fixed | 0.8844 | 0.8528 | 0.6725 | 0.7609 | 0.9886 | 0.9825 | 0.5771 |
| Uncertainty | 0.8767 | 0.8648 | 0.5506 | 0.6386 | 0.9779 | 0.9698 | 0.4640 |
| BGE-M3 | Fixed | 0.8942 | 0.8718 | 0.8148 | 0.8880 | 0.9727 | 0.9554 | 0.7184 |
| Uncertainty | 0.8921 | 0.8936 | 0.8035 | 0.8792 | 0.9718 | 0.9494 | 0.7089 |
| Qwen-Emb | Fixed | 0.8839 | 0.8447 | 0.9167 | 0.9565 | 0.9803 | 0.9699 | 0.7981 |
| Uncertainty | 0.8818 | 0.8663 | 0.9167 | 0.9565 | 0.9495 | 0.9432 | 0.7732 |
Table 8.
Reported validation and test performance averaged over six trajectory encoders for each text encoder. All denotes the unweighted mean over all 24 encoder combinations. Perc. denotes perception accuracy and FM denotes Full Match Accuracy.
Table 8.
Reported validation and test performance averaged over six trajectory encoders for each text encoder. All denotes the unweighted mean over all 24 encoder combinations. Perc. denotes perception accuracy and FM denotes Full Match Accuracy.
| Text Encoder | Weighting | Perc. Val. | Perc. Test | FM Val. | FM Test |
|---|
| EmbeddingGemma | Fixed | 1.0000 | 0.9167 | 0.9757 | 0.7322 |
| EmbeddingGemma | Uncertainty | 0.9996 | 0.9165 | 0.9416 | 0.7353 |
| MiniLM | Fixed | 0.9938 | 0.6725 | 0.9727 | 0.5771 |
| MiniLM | Uncertainty | 0.9471 | 0.5506 | 0.8986 | 0.4640 |
| BGE-M3 | Fixed | 1.0000 | 0.8148 | 0.9758 | 0.7184 |
| BGE-M3 | Uncertainty | 0.9732 | 0.8035 | 0.9217 | 0.7089 |
| Qwen-Emb | Fixed | 0.9574 | 0.9167 | 0.9393 | 0.7981 |
| Qwen-Emb | Uncertainty | 0.9088 | 0.9167 | 0.8543 | 0.7732 |
| All | Fixed | 0.9878 | 0.8302 | 0.9659 | 0.7064 |
| All | Uncertainty | 0.9572 | 0.7968 | 0.9041 | 0.6704 |
Table 9.
Latent-vector ablation results averaged over four text encoder pairings. M: full multimodal input; ZT: text latent vector set to zero; ZR: trajectory latent vector set to zero. Bold values identify the highest reported score in each column. All reported values are accuracies.
Table 9.
Latent-vector ablation results averaged over four text encoder pairings. M: full multimodal input; ZT: text latent vector set to zero; ZR: trajectory latent vector set to zero. Bold values identify the highest reported score in each column. All reported values are accuracies.
| Weighting | Trajectory AE | Traj. M | Traj. ZT | Perc. M | Perc. ZR | Dec. M | Dec. ZR |
|---|
| Fixed | BiLSTM-AE | 0.9108 | 0.5033 | 0.7991 | 0.8906 | 0.9997 | 0.3952 |
| Fixed | Simple-AE | 0.9695 | 0.7201 | 0.7307 | 0.8411 | 0.9993 | 0.3555 |
| Fixed | Stacked-AE | 0.8559 | 0.3831 | 0.8707 | 0.9167 | 0.9993 | 0.3555 |
| Fixed | Transformer-AE | 0.8557 | 0.4198 | 0.8909 | 0.9193 | 0.9939 | 0.3555 |
| Fixed | xLSTM-AE | 0.6492 | 0.6001 | 0.8714 | 0.9167 | 0.8504 | 0.3555 |
| Fixed | Mamba-AE | 0.9997 | 0.9997 | 0.8182 | 0.7552 | 0.9993 | 0.6445 |
| Uncertainty | BiLSTM-AE | 0.8990 | 0.3900 | 0.7296 | 0.8086 | 0.9999 | 0.3555 |
| Uncertainty | Simple-AE | 0.9665 | 0.6565 | 0.8189 | 0.9232 | 0.9998 | 0.3555 |
| Uncertainty | Stacked-AE | 0.8433 | 0.3762 | 0.7755 | 0.9167 | 0.9991 | 0.3555 |
| Uncertainty | Transformer-AE | 0.8542 | 0.4422 | 0.8583 | 0.8737 | 0.9878 | 0.3555 |
| Uncertainty | xLSTM-AE | 0.6665 | 0.6133 | 0.8495 | 0.8737 | 0.7973 | 0.3555 |
| Uncertainty | Mamba-AE | 0.9995 | 0.9994 | 0.7491 | 0.7005 | 0.9993 | 0.6445 |