PRISM-MTL: Inter-Modal Selective Multi-Task Learning for Assistive Driving Perception
Abstract
1. Introduction
2. Related Works
2.1. Single-Task Learning for Driver State Recognition Categories
2.2. Single-Task Learning for Traffic Context Perception Categories
2.3. Multi-Task Learning for ADAS
| Authors | Year | Dataset | Modalities | Method | DER | DBR | TCR | VBR |
|---|---|---|---|---|---|---|---|---|
| Saadi et al. [20] | 2023 | KMU-FED + FER2013 | Driver face images | DFER-GCViT | ✓ | - | - | - |
| Saadi et al. [23] | 2024 | KMU-FED + KDEF | Driver face images | ShuffViT-DFER | ✓ | - | - | - |
| Duan et al. [3] | 2023 | SFD + AUCDD-V1 | Face + hands video | FRNet | - | ✓ | - | - |
| Kim et al. [15] | 2025 | NTHU-DDD | Face + body video | STFTransNet | - | ✓ | - | - |
| Doshi [19] | 2025 | StateFarm | Driver images | Anchor-ViT | - | ✓ | - | - |
| Qu et al. [36] | 2021 | TuSimple + CULane | Front camera images | FOLOLane | - | - | ✓ | - |
| Ko et al. [35] | 2022 | TuSimple + CULane | Front camera images | PINet | - | - | ✓ | - |
| Yan et al. [37] | 2025 | TuSimple + CULane | Front camera images | MHFS-FORMER | - | - | ✓ | - |
| Wasi et al. [6] | 2024 | DAAD | Multi-view video + gaze | M2MVT | - | - | - | ✓ |
| Liu et al. [51] | 2025 | AIDE | Multi-view video + joints | TEM3-Learning | ✓ | ✓ | ✓ | ✓ |
| Liu et al. [52] | 2025 | AIDE | Multi-view video + joints | UMD-Net | ✓ | ✓ | ✓ | ✓ |
| Liu et al. [53] | 2025 | AIDE | Multi-view video + joints | MMTL-UniAD | ✓ | ✓ | ✓ | ✓ |
3. Perception and Recognition with Inter-Modal Selective Multi-Task Learning
3.1. Problem Formulation
3.2. Multimodal Data Preprocessing
3.3. Hierarchical Stage-Wise Attention Network
3.4. Joint Modality Encoder with Token-SE
3.5. Task-Specific Modality Fusion
3.6. Multi-Head Task-to-Modality Cross-Attention
3.7. Clip-Level Prediction and Multi-Task Objective
4. Experimental Results
4.1. Dataset
4.2. Evaluation Metrics
4.3. Implementation Details
4.4. Comparison with the State-of-the-Art
4.5. Ablation Studies
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- World Health Organization. Global Status Report on Road Safety 2023; World Health Organization: Geneva, Switzerland, 2023. [Google Scholar]
- Centers for Disease Control and Prevention. Global Road Safety: Road Traffic Injuries and Economic Cost; U.S. Department of Health & Human Services: Washington, DC, USA, 2025. Available online: https://www.cdc.gov/transportation-safety/global/ (accessed on 5 August 2025).
- Duan, C.; Gong, Y.; Liao, J.; Zhang, M.; Cao, L. FRNet: DCNN for real-time distracted driving detection toward embedded deployment. IEEE Trans. Intell. Transp. Syst. 2023, 24, 9835–9848. [Google Scholar] [CrossRef]
- Saadi, I.; Cunningham, D.W.; Taleb-Ahmed, A.; Hadid, A.; El Hillali, Y. Driver’s facial expression recognition: A comprehensive survey. Expert Syst. Appl. 2024, 242, 122784. [Google Scholar] [CrossRef]
- Rocky, A.; Wu, Q.M.J.; Zhang, W. Review of accident detection methods using dashcam videos for autonomous driving vehicles. IEEE Trans. Intell. Transp. Syst. 2024, 25, 6885–6900. [Google Scholar] [CrossRef]
- Wasi, A.; Gangisetty, S.; Rai, S.N.; Jawahar, C.V. Early anticipation of driving maneuvers. In Proceedings of the European Conference on Computer Vision (ECCV); Lecture Notes in Computer Science; Springer Nature: Cham, Switzerland, 2024; Volume 15128, pp. 152–169. [Google Scholar]
- Yang, D.; Huang, S.; Xu, Z.; Li, Z.; Wang, S.; Li, M.; Wang, Y.; Liu, Y.; Yang, K.; Chen, Z.; et al. AIDE: A vision-driven multi-view, multi-modal, multi-tasking dataset for assistive driving perception. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 20459–20470. [Google Scholar]
- Kendall, A.; Gal, Y.; Cipolla, R. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 7482–7491. [Google Scholar]
- Misra, I.; Shrivastava, A.; Gupta, A.; Hebert, M. Cross-stitch networks for multi-task learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016. [Google Scholar]
- Xu, Y.; Qian, Y.; Jie, Z.; Ma, L. Multi-view action recognition for distracted driver behavior localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vancouver, BC, Canada, 17–24 June 2023; pp. 7172–7181. [Google Scholar]
- Abosaq, H.A.; Ramzan, M.; Althobiani, F.; Abid, A.; Aamir, K.M.; Abdushkour, H.; Irfan, M.; Gommosani, M.E.; Ghonaim, S.M.; Shamji, V.R.; et al. Unusual driver behavior detection in videos using deep learning models. Sensors 2023, 23, 311. [Google Scholar]
- Althabhawee, A.F.Y.; Ibrahim, R.M.; Oleiwi, B.K. Enhancing road safety using deep learning-based driver behavior detection system. Int. J. Transp. Dev. Integr. 2025, 9, 317–324. [Google Scholar] [CrossRef]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 770–778. [Google Scholar]
- Huang, C.; Wang, X.; Cao, J.; Wang, S.; Zhang, Y. HCF: A hybrid CNN framework for behavior detection of distracted drivers. IEEE Access 2020, 8, 109335–109349. [Google Scholar] [CrossRef]
- Kim, M.; Choi, G. STFTransNet: A transformer-based spatial–temporal fusion network for enhanced multimodal driver inattention state recognition system. Sensors 2025, 25, 5819. [Google Scholar] [CrossRef] [PubMed]
- Khan, T.; Choi, G.; Lee, S. EFFNet-CA: An efficient driver distraction detection based on multiscale features extractions and channel attention mechanism. Sensors 2023, 23, 3835. [Google Scholar] [CrossRef] [PubMed]
- Walizad, M.E.; Hurroo, M.; Sethia, D. Driver drowsiness detection system using convolutional neural network. In Proceedings of the 2022 6th International Conference on Trends in Electronics and Informatics (ICOEI), Tirunelveli, India, 28–30 April 2022; pp. 1073–1080. [Google Scholar]
- Chaudhari, A.; Bhatt, C.; Krishna, A.; Mazzeo, P.L. ViTFER: Facial emotion recognition with vision transformers. Appl. Syst. Innov. 2022, 5, 80. [Google Scholar] [CrossRef]
- Doshi, V. Anchor-ViT: Spatially focused vision transformer for distracted driving detection. In Proceedings of the IEEE International Conference on Image Processing (ICIP), Anchorage, AK, USA, 21–24 September 2025. [Google Scholar]
- Saadi, I.; Cunningham, D.W.; Taleb-Ahmed, A.; Hadid, A.; El Hillali, Y. Driver’s facial expression recognition using global context vision transformer. In Proceedings of the IEEE International Conference on Computer Vision and Machine Intelligence (CVMI), Gwalior, India, 10–11 November 2023; pp. 1–8. [Google Scholar]
- Mohammed, A.A.Q.; Geng, X.; Wang, J.; Ali, Z. Driver distraction detection using semi-supervised lightweight vision transformer. Eng. Appl. Artif. Intell. 2024, 129, 107618. [Google Scholar] [CrossRef]
- Liang, J.; Zhu, H.; Zhang, E.; Zhang, J. Stargazer: A transformer-based driver action detection system for intelligent transportation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vancouver, BC, Canada, 17–24 June 2022; pp. 3159–3166. [Google Scholar]
- Saadi, I.; Cunningham, D.W.; Taleb-Ahmed, A.; Hadid, A.; El Hillali, Y. Shuffle Vision Transformer: Lightweight, fast and efficient recognition of driver’s facial expression. arXiv 2024, arXiv:2409.03438. [Google Scholar]
- Cui, J.; Chen, Y.; Wu, Z.; Wu, H.; Wu, W. A driver behavior detection model for human–machine co-driving systems based on an improved Swin Transformer. World Electr. Veh. J. 2025, 16, 7. [Google Scholar]
- Ying, N.; Jiang, Y.; Guo, C.; Zhou, D.; Zhao, J. A multimodal driver emotion recognition algorithm based on the audio and video signals in Internet of Vehicles platform. IEEE Internet Things J. 2024, 11, 35812–35824. [Google Scholar] [CrossRef]
- Manavand, M.R.; Salarifar, M.H.; Ghavami, M.; Taghipour-Gorjikolaie, M. Driver’s facial expression recognition by using deep local and global features. Inf. Sci. 2025, 692, 121658. [Google Scholar] [CrossRef]
- Hasan, M.Z.; Joshi, A.; Rahman, M.S.; Venkatachalapathy, A.; Sharma, A.; Hegde, C.; Sarkar, S. DriveCLIP: Zero-shot transfer for distracted driving activity understanding using CLIP. In Proceedings of the NeurIPS Workshop on Machine Learning for Autonomous Driving (ML4AD), New Orleans, LA, USA, 3 December 2022. [Google Scholar]
- Zhao, J.; Wu, Y.; Deng, R.; Xu, S.; Gao, J.; Burke, A. A survey of autonomous driving from a deep learning perspective. ACM Comput. Surv. 2025, 57, 263. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 779–788. [Google Scholar]
- Zeng, G.; Wu, Z.; Xu, L.; Liang, Y. Efficient Vision Transformer YOLOv5 for accurate and fast traffic sign detection. Electronics 2024, 13, 880. [Google Scholar] [CrossRef]
- Liu, L.; Su, B.; Jiang, J.; Wu, G.; Guo, C.; Xu, C.; Yang, H.F. Towards accurate and efficient 3D object detection for autonomous driving: A mixture of experts computing system on edge. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–3 November 2025; pp. 25903–25913. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. Adv. Neural Inf. Process. Syst. 2015, 28, 91–99. [Google Scholar] [CrossRef]
- Gao, X.; Liu, Y.; Chen, J.; Li, H. Improved traffic sign detection algorithm based on Faster R-CNN. Appl. Sci. 2022, 12, 8948. [Google Scholar] [CrossRef]
- Pan, X.; Shi, J.; Luo, P.; Wang, X.; Tang, X. Spatial as deep: Spatial CNN for traffic scene understanding. Proc. AAAI Conf. Artif. Intell. 2018, 32, 7276–7283. [Google Scholar] [CrossRef]
- Ko, Y.; Lee, Y.; Azam, S.; Munir, F.; Jeon, M.; Pedrycz, W. Key points estimation and point instance segmentation approach for lane detection. IEEE Trans. Intell. Transp. Syst. 2022, 23, 8949–8958. [Google Scholar] [CrossRef]
- Qu, Z.; Jin, H.; Zhou, Y.; Yang, Z.; Zhang, W. Focus on local: Detecting lane marker from bottom up via key point. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 19–25 June 2021; pp. 14117–14125. [Google Scholar]
- Yan, D.; Zhang, T. MHFS-FORMER: Multiple-scale hybrid features transformer for lane detection. Sensors 2025, 25, 2876. [Google Scholar] [CrossRef] [PubMed]
- Xue, J.-R.; Fang, J.-W.; Zhang, P. A survey of scene understanding by event reasoning in autonomous driving. Int. J. Autom. Comput. 2018, 15, 249–266. [Google Scholar] [CrossRef]
- Lv, C.; Qi, M.; Liu, L.; Ma, H. T2SG: Traffic topology scene graph for topology reasoning in autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 15–20 June 2025; pp. 17197–17206. [Google Scholar]
- Rong, F.; Peng, W.; Lan, M.; Zhang, Q.; Zhang, L. Driving scene understanding with traffic scene-assisted topology graph transformer. In Proceedings of the ACM International Conference on Multimedia (ACM MM), Melbourne, Australia, 28 October–1 November 2024; pp. 10075–10084. [Google Scholar]
- Zhang, Y.; Qian, D.; Li, D.; Pan, Y.; Chen, Y.; Liang, Z.; Zhang, Z.; Liu, Y.; Mei, J.; Fu, M.; et al. GraphAD: Interaction scene graph for end-to-end autonomous driving. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), Montreal, QC, Canada, 3–9 August 2025; pp. 2422–2430. [Google Scholar]
- Hu, A.; Murez, Z.; Mohan, N.; Dudas, S.; Hawke, J.; Badrinarayanan, V.; Cipolla, R.; Kendall, A. FIERY: Future instance prediction in bird’s-eye view from surround monocular cameras. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 15253–15262. [Google Scholar]
- Hu, S.; Chen, L.; Wu, P.; Li, H.; Yan, J.; Tao, D. ST-P3: End-to-end vision-based autonomous driving via spatial–temporal feature learning. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 533–549. [Google Scholar]
- Prakash, A.; Chitta, K.; Geiger, A. Multi-modal fusion transformer for end-to-end autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 19–25 June 2021; pp. 7077–7087. [Google Scholar]
- Li, Z.; Wang, W.; Li, H.; Xie, E.; Sima, C.; Lu, T.; Qiao, Y.; Dai, J. BEVFormer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 1–18. [Google Scholar]
- Hu, Y.; Yang, J.; Chen, L.; Li, K.; Sima, C.; Zhu, X.; Chai, S.; Du, S.; Lin, T.; Wang, W.; et al. Planning-oriented autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 17853–17862. [Google Scholar]
- Sima, C.; Renz, K.; Chitta, K.; Chen, L.; Zhang, H.; Xie, C.; Beißwenger, J.; Luo, P.; Geiger, A.; Li, H. DriveLM: Driving with graph visual question answering. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 256–274. [Google Scholar]
- Shao, H.; Hu, Y.; Wang, L.; Song, G.; Waslander, S.L.; Liu, Y.; Li, H. LMDrive: Closed-loop end-to-end driving with large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 15120–15130. [Google Scholar]
- Ruder, S. An overview of multi-task learning in deep neural networks. arXiv 2017, arXiv:1706.05098. [Google Scholar]
- Xu, X.; Zhao, H.; Vineet, V.; Lim, S.-N.; Torralba, A. MTFormer: Multi-task learning via transformer and cross-task reasoning. In Proceedings of the European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 304–321. [Google Scholar]
- Liu, W.; Qiao, Y.; Wang, Z.; Guo, Q.; Chen, Z.; Zhou, M.; Li, X.; Wang, L.; Li, Z.; Liu, H.; et al. TEM3-Learning: Time-efficient multimodal multi-task learning for advanced assistive driving. arXiv 2025, arXiv:2506.18084. [Google Scholar]
- Liu, W.; Qiao, Y.; Li, Z.; Wang, W.; Zhang, W.; Zhu, J.; Jiang, Y.; Wang, L.; Wang, H.; Liu, H.; et al. UMD-Net: A unified multi-task assistive driving network based on multimodal fusion. IEEE Trans. Intell. Transp. Syst. 2025, 26, 12315–12328. [Google Scholar] [CrossRef]
- Liu, W.; Wang, W.; Qiao, Y.; Guo, Q.; Zhu, J.; Li, P.; Chen, Z.; Yang, H.; Li, Z.; Wang, L.; et al. MMTL-UniAD: A unified framework for multimodal and multi-task learning in assistive driving perception. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 15–20 June 2025; pp. 6864–6874. [Google Scholar]
- State Farm. State Farm Distracted Driver Detection; Kaggle: San Francisco, CA, USA, 2016; Available online: https://www.kaggle.com/c/state-farm-distracted-driver-detection (accessed on 28 May 2026).
- Weng, C.-H.; Lai, Y.-H.; Lai, S.-H. Driver drowsiness detection via a hierarchical temporal deep belief network. In Proceedings of the Asian Conference on Computer Vision Workshops (ACCVW); Springer International Publishing: Cham, Switzerland, 2017; pp. 117–133. [Google Scholar]
- Ramanishka, V.; Chen, Y.-T.; Misu, T.; Saenko, K. Toward driving scene understanding: A dataset for learning driver behavior and causal reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 7699–7707. [Google Scholar]
- Li, Z.; Zhao, X.; Wu, F.; Chen, D.; Wang, C. A lightweight and efficient distracted driver detection model fusing convolutional neural network and vision transformer. IEEE Trans. Intell. Transp. Syst. 2024, 25, 19962–19978. [Google Scholar] [CrossRef]
- Bai, J.; Yu, W.; Xiao, Z.; Havyarimana, V.; Regan, A.C.; Jiang, H.; Jiao, L. Two-stream spatial–temporal graph convolutional networks for driver drowsiness detection. IEEE Trans. Cybern. 2022, 52, 13821–13833. [Google Scholar] [CrossRef] [PubMed]
- Yang, L.; Yang, H.; Wei, H.; Hu, Z.; Lv, C. Video-based driver drowsiness detection with optimised utilization of key facial features. IEEE Trans. Intell. Transp. Syst. 2024, 25, 6938–6950. [Google Scholar] [CrossRef]
- Huang, Y.; Liu, C.; Chang, F.; Lu, Y. Self-supervised multi-granularity graph attention network for vision-based driver fatigue detection. IEEE Trans. Emerg. Top. Comput. Intell. 2024, 8, 3067–3080. [Google Scholar] [CrossRef]
- Chen, J.; Mittal, G.; Yu, Y.; Kong, Y.; Chen, M. GateHUB: Gated history unit with background suppression for online action detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 19925–19934. [Google Scholar]
- Wang, J.; Chen, G.; Huang, Y.; Wang, L.; Lu, T. Memory-and-anticipation transformer for online action understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 13824–13835. [Google Scholar]
- Cao, S.; Luo, W.; Wang, B.; Zhang, W.; Ma, L. E2E-LOAD: End-to-end long-form online action detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 10422–10432. [Google Scholar]
- Guo, H.; Wang, H.; Ji, Q. Bayesian evidential deep learning for online action detection. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; Volume 15074, pp. 283–301. [Google Scholar]








| Symbol | Description |
|---|---|
| Task set (DER, DBR, TCR, VBR) | |
| Modality set (face, body, scene, posture, gesture) | |
| E | Modality embedding dimension (E = 128) |
| D | Shared attention dimension of TSMF (D = 256) |
| H | Number of attention heads in TSMF (H = 2) |
| k | Kernel size of the spatial attention (k = 7) |
| tk | Learnable task token used as the query for the k-th task |
| Q, K, V | Query, key, and value in the cross-attention |
| αt,m | Task-conditioned modality relevance of task t for modality m |
| FFN | Feed-forward network refining task-specific features |
| Task | Count | Class |
|---|---|---|
| driver emotion | 5 | Anxiety, Peace, Weariness, Happiness, Anger |
| driver behavior | 7 | Smoking, Making Phone, Looking Around, Dozing off, Normal Driving, Talking, Body Movement |
| traffic context | 3 | Traffic Jam, Waiting, Smooth Traffic |
| vehicle behavior | 5 | Parking, Turning, Backward Moving, Changing Lane, Forward Moving |
| Author | Modality | Method | Acc. (%) | F1-Score |
|---|---|---|---|---|
| Khan et al. [16] (2023) | Driver area | EFFNet-CA | 99.58 | 1.000 |
| Li et al. [57] (2024) | Driver area | CoViT | 97.89 | - |
| Doshi [19] (2025) | Driver area | Anchor-ViT | 92.30 ± 0.30 | 0.924 |
| Kim et al. [15] (2025) | Driver area | STFTransNet | 99.65 ± 0.13 | 0.996 ± 0.001 |
| Ours | Face/Body | PRISM-MTL | 99.55 ± 0.09 | 0.996 ± 0.001 |
| Author | Method | Model | Acc. (%) | F1-Score |
|---|---|---|---|---|
| Bai et al. [58] (2022) | Face landmark | 2s-STGCN | 92.70 | 0.881 |
| Yang et al. [59] (2024) | Face area | VBFLLFA | 91.30 | - |
| Huang et al. [60] (2024) | Face area | SMGA-Net | 81.00 | 0.811 |
| Kim et al. [15] (2025) | Driver area | STFTransNet | 95.86 ± 0.17 | 0.957 ± 0.002 |
| Ours | Face/Body | PRISM-MTL | 96.47 ± 0.21 | 0.964 ± 0.002 |
| Author | Method | Model | mAP (%) |
|---|---|---|---|
| Chen et al. [61] (2022) | Sensors | GateHUB | 32.1 |
| Wang et al. [62] (2023) | Sensors | MAT | 32.7 |
| Cao et al. [63] (2023) | Scene | E2E-LOAD | 48.1 |
| Guo et al. [64] (2024) | Sensors | BEDL | 33.0 |
| Ours | Scene/Sensors | PRISM-MTL | 64.45 ± 0.01 |
| HSA-Net Stage | DER | DBR | TCR | VBR | mAcc. (%) | ||
|---|---|---|---|---|---|---|---|
| Stage 2 | Stage 3 | Stage 4 | Acc. | Acc. | Acc. | Acc. | |
| √ | 83.68 | 79.14 | 93.96 | 85.55 | 85.58 | ||
| √ | 84.69 | 78.79 | 93.45 | 86.03 | 85.74 | ||
| √ | 85.17 | 78.79 | 93.45 | 84.83 | 85.56 | ||
| √ | √ | 82.59 | 76.72 | 94.48 | 85.34 | 84.78 | |
| √ | √ | 84.31 | 79.31 | 94.66 | 85.69 | 85.99 | |
| √ | √ | 84.48 | 79.14 | 94.83 | 85.00 | 85.86 | |
| √ | √ | √ | 83.79 | 79.48 | 95.34 | 86.38 | 86.25 |
| Multimodal Data | DER | DBR | TCR | VBR | mAcc. (%) | ||
|---|---|---|---|---|---|---|---|
| Face + Body | Scene | Post + Gest | Acc. | Acc. | Acc. | Acc. | |
| √ | 83.28 | 79.31 | 90.52 | 82.24 | 83.84 | ||
| √ | 83.28 | 78.28 | 94.48 | 85.86 | 85.86 | ||
| √ | 70.00 | 71.90 | 84.48 | 74.31 | 75.17 | ||
| √ | √ | 85.34 | 77.07 | 93.10 | 85.86 | 85.34 | |
| √ | √ | 84.31 | 79.48 | 93.62 | 85.86 | 85.82 | |
| √ | √ | 82.07 | 76.38 | 90.17 | 82.24 | 82.72 | |
| √ | √ | √ | 83.79 | 79.48 | 95.34 | 86.38 | 86.25 |
| Multi-Task | DER | DBR | TCR | VBR | |||
|---|---|---|---|---|---|---|---|
| DER | DBR | TCR | VBR | Acc. | Acc. | Acc. | Acc. |
| √ | 81.14 | - | - | - | |||
| √ | - | 79.14 | - | - | |||
| √ | - | - | 94.48 | - | |||
| √ | - | - | - | 86.21 | |||
| √ | √ | √ | √ | 83.79 | 79.48 | 95.34 | 86.38 |
| Multi-Task | DER | DBR | TCR | VBR | |||
|---|---|---|---|---|---|---|---|
| DER | DBR | TCR | VBR | Acc. | Acc. | Acc. | Acc. |
| √ | √ | 82.76 | 78.97 | - | - | ||
| √ | √ | - | - | 92.93 | 86.03 | ||
| √ | √ | √ | 84.83 | 79.14 | 92.93 | - | |
| √ | √ | √ | 84.31 | 78.62 | - | 85.17 | |
| √ | √ | √ | 80.86 | - | 92.93 | 86.27 | |
| √ | √ | √ | - | 77.07 | 93.79 | 85.69 | |
| √ | √ | √ | √ | 83.79 | 79.48 | 95.34 | 86.38 |
| Method | DER | DBR | TCR | VBR | mAcc. (%) | Params (M) | FLOPs (G) | |
|---|---|---|---|---|---|---|---|---|
| HSA-Net | TSMF | Acc. | Acc. | Acc. | Acc. | |||
| 82.41 | 81.21 | 93.28 | 81.38 | 82.41 | 34.51 | 175.07 | ||
| √ | 82.59 | 75.34 | 91.72 | 82.93 | 83.15 | 34.61 | 175.09 | |
| √ | 82.93 | 79.45 | 93.62 | 85.17 | 85.29 | 34.68 | 175.12 | |
| √ | √ | 83.79 | 79.48 | 95.34 | 86.38 | 86.25 | 34.77 | 175.13 |
| Method | DER | DBR | TCR | VBR | mAcc. (%) |
|---|---|---|---|---|---|
| Acc. | Acc. | Acc. | Acc. | ||
| task-agnostic attention | 83.62 | 76.55 | 93.79 | 85.52 | 84.87 |
| TSMF | 83.79 | 79.48 | 95.34 | 86.38 | 86.25 |
| TSMF Head Rate | DER | DBR | TCR | VBR | mAcc. (%) |
|---|---|---|---|---|---|
| Acc. | Acc. | Acc. | Acc. | ||
| 1 Head | 84.31 | 78.45 | 94.66 | 86.03 | 85.86 |
| 2 Head | 86.38 | 78.62 | 94.66 | 85.34 | 86.25 |
| 4 Head | 84.66 | 78.10 | 94.48 | 86.72 | 85.99 |
| 8 Head | 84.31 | 79.66 | 93.28 | 86.90 | 86.03 |
| Embedding Dimension (E, D) | DER | DBR | TCR | VBR | mAcc. (%) |
|---|---|---|---|---|---|
| Acc. | Acc. | Acc. | Acc. | ||
| (64, 128) | 84.31 | 79.48 | 93.28 | 85.52 | 85.65 |
| (128, 256) default | 83.79 | 79.48 | 95.34 | 86.38 | 86.25 |
| (256, 512) | 84.83 | 79.48 | 93.97 | 84.48 | 85.69 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kim, M.; Choi, G. PRISM-MTL: Inter-Modal Selective Multi-Task Learning for Assistive Driving Perception. Mathematics 2026, 14, 2812. https://doi.org/10.3390/math14152812
Kim M, Choi G. PRISM-MTL: Inter-Modal Selective Multi-Task Learning for Assistive Driving Perception. Mathematics. 2026; 14(15):2812. https://doi.org/10.3390/math14152812
Chicago/Turabian StyleKim, Minjun, and Gyuho Choi. 2026. "PRISM-MTL: Inter-Modal Selective Multi-Task Learning for Assistive Driving Perception" Mathematics 14, no. 15: 2812. https://doi.org/10.3390/math14152812
APA StyleKim, M., & Choi, G. (2026). PRISM-MTL: Inter-Modal Selective Multi-Task Learning for Assistive Driving Perception. Mathematics, 14(15), 2812. https://doi.org/10.3390/math14152812

