Road Traffic Anomaly Detection by Human-Attention-Assisted Text–Vision Learning
Abstract
1. Introduction
2. Related Work
2.1. Road Traffic Anomaly Detection Methods
2.2. Visual Text Multimodal Learning for Video Anomaly Detection
3. Text Annotation on the TADS Dataset
- (1)
- Traffic Anomaly Description: “A traffic anomaly occurred in the scene”. For aligning anomalous video representations with text representations during the training process, all features of traffic anomaly videos are aligned with the text features of this sentence.
- (2)
- Normal Scene Description: “The traffic in this scenario is normal”. For aligning normal video representations with text representations during the training process, all features of videos prior to the occurrence of any traffic anomalies are aligned with the text features of this sentence.
4. The Detection Approach
4.1. Multimodal Feature Extraction
4.2. The Introduction of the Eye-Gaze Map
4.3. Loss Function and Traffic Anomaly Discrimination
4.3.1. Loss Between Visual and Coarse-Grained Text Features
4.3.2. Loss Between Fine-Grained Text and Coarse-Grained Text Features
4.3.3. Loss Between Visual and Fine-Grained Text Features
4.3.4. Loss Between Attention and Eye-Gaze Maps
| Algorithm 1 The pseudocode of training process | |
| Require: frames_group; Require: gaze_map; Require: fine_grained_texts; Require: coarse_grained_texts; for each input_data in dataLoader do frame_target = frames_group[:, -1, :, :, :] frame_emb_f, _ = CLIP_visual(frames_group) frame_emb = temporal_fusion(frame_emb_f) _, frame_emb_withpn = CLIP_visual(frame_target) coarse_emb = CLIP_text(coarse_grained_texts) fine_emb = CLIP_text(fine_grained_texts) fine_grained_out, the_atten = cross_attn_i2t(frame_emb_withpn, fine_emb) atten_calcul = attn_transfer(the_atten) loss_coarse = loss-v2c(frame_emb, coarse_emb) loss_fine = loss-v2f((frame_emb, fine_emb)) loss_coarse2fine = loss-f2c(fine_emb, coarse_emb) loss_gaze = KL-loss(gaze_map, atten_calcul) loss = weighted_sum(loss_coarse, loss_fine, loss_coarse2fine, loss_gaze) loss.train() end for | (b,t,c,h,w) (b,1,h,w) (b,77) (b,77) take the last frame as the target |
4.3.5. Traffic Anomaly Score and Discrimination
5. Experiments and Analysis
5.1. Evaluation Metrics
5.2. Experimental Setup
5.3. Performance Comparison and Analysis
5.3.1. Competitors
5.3.2. Overall Performance Comparison
5.3.3. Performance Comparison on Various Categories of Traffic Anomaly
5.3.4. Qualitative Results
5.4. Ablation Experiment
5.4.1. Overall Performance for Ablation
5.4.2. Performance on Various Categories of Traffic Anomaly for Ablation
5.5. Zero-Shot and Inference Speed
5.5.1. Zero-Shot and Fine-Tuning
5.5.2. Inference Speed
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Fang, J.; Qiao, J.; Bai, J.; Yu, H.; Xue, J. Traffic accident detection via self-supervised consistency learning in driving scenarios. IEEE Trans. Intell. Transp. Syst. 2022, 23, 9601–9614. [Google Scholar] [CrossRef]
- Fang, J.; Yan, D.; Qiao, J.; Xue, J.; Wang, H.; Li, S. Dada-2000: Can driving accident be predicted by driver attentionƒ analyzed by a benchmark. In Proceedings of the 2019 IEEE Intelligent Transportation Systems Conference (ITSC); IEEE: Piscataway, NJ, USA, 2019; pp. 4303–4309. [Google Scholar]
- Cui, X.; Li, L.L.; Chai, Y.; Fang, J.; Silamu, W. Self-supervised Traffic Accident Detection by Motion-Conditioned Diffusive Frame Prediction. In Proceedings of the International Conference on Autonomous Unmanned Systems; Springer: Berlin/Heidelberg, Germany, 2023; pp. 292–303. [Google Scholar]
- Chai, Y.; Fang, J.; Liang, H.; Silamu, W. TADS: A novel dataset for road traffic accident detection from a surveillance perspective. J. Supercomput. 2024, 80, 26226–26249. [Google Scholar] [CrossRef]
- Fang, J.; Qiao, J.; Xue, J.; Li, Z. Vision-based traffic accident detection and anticipation: A survey. IEEE Trans. Circuits Syst. Video Technol. 2023, 34, 1983–1999. [Google Scholar] [CrossRef]
- Sultani, W.; Chen, C.; Shah, M. Real-world anomaly detection in surveillance videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 6479–6488. [Google Scholar]
- Li, L.; Chen, Y.C.; Cheng, Y.; Gan, Z.; Yu, L.; Liu, J. Hero: Hierarchical encoder for video+ language omni-representation pre-training. arXiv 2020, arXiv:2005.00200. [Google Scholar] [CrossRef]
- Xu, H.; Ghosh, G.; Huang, P.Y.; Arora, P.; Aminzadeh, M.; Feichtenhofer, C.; Metze, F.; Zettlemoyer, L. Vlm: Task-agnostic video-language model pre-training for video understanding. arXiv 2021, arXiv:2105.09996. [Google Scholar]
- Fu, T.J.; Li, L.; Gan, Z.; Lin, K.; Wang, W.Y.; Wang, L.; Liu, Z. Violet: End-to-end video-language transformers with masked visual-token modeling. arXiv 2021, arXiv:2111.12681. [Google Scholar]
- Xu, P.; Zhu, X.; Clifton, D.A. Multimodal learning with transformers: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12113–12132. [Google Scholar] [CrossRef] [PubMed]
- Nag, S.; Zhu, X.; Song, Y.Z.; Xiang, T. Zero-shot temporal action detection via vision-language prompting. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 681–697. [Google Scholar]
- Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.H.; Li, Z.; Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 4904–4916. [Google Scholar]
- Gu, X.; Lin, T.Y.; Kuo, W.; Cui, Y. Open-vocabulary object detection via vision and language knowledge distillation. arXiv 2021, arXiv:2104.13921. [Google Scholar]
- Yao, L.; Han, J.; Liang, X.; Xu, D.; Zhang, W.; Li, Z.; Xu, H. Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 23497–23506. [Google Scholar]
- Xu, J.; De Mello, S.; Liu, S.; Byeon, W.; Breuel, T.; Kautz, J.; Wang, X. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2022; pp. 18134–18144. [Google Scholar]
- Zhou, Z.; Lei, Y.; Zhang, B.; Liu, L.; Liu, Y. Zegclip: Towards adapting clip for zero-shot semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 11175–11185. [Google Scholar]
- Baldrati, A.; Agnolucci, L.; Bertini, M.; Del Bimbo, A. Zero-shot composed image retrieval with textual inversion. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2023; pp. 15338–15347. [Google Scholar]
- Tschannen, M.; Mustafa, B.; Houlsby, N. Clippo: Image-and-language understanding from pixels only. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 11006–11017. [Google Scholar]
- Wu, W.; Sun, Z.; Ouyang, W. Revisiting classifier: Transferring vision-language models for video recognition. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Palo Alto, CA, USA, 2023; Volume 37, pp. 2847–2855. [Google Scholar]
- Ni, B.; Peng, H.; Chen, M.; Zhang, S.; Meng, G.; Fu, J.; Xiang, S.; Ling, H. Expanding language-image pretrained models for general video recognition. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 1–18. [Google Scholar]
- Chen, S.; Xu, Q.; Ma, Y.; Qiao, Y.; Wang, Y. Attentive snippet prompting for video retrieval. IEEE Trans. Multimed. 2023, 26, 4348–4359. [Google Scholar] [CrossRef]
- Ma, Y.; Xu, G.; Sun, X.; Yan, M.; Zhang, J.; Ji, R. X-clip: End-to-end multi-grained contrastive learning for video-text retrieval. In Proceedings of the 30th ACM International Conference on Multimedia; ACM: New York, NY, USA, 2022; pp. 638–647. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning; PmLR: Cambridge, MA, USA, 2021; pp. 8748–8763. [Google Scholar]
- Wu, P.; Zhou, X.; Pang, G.; Sun, Y.; Liu, J.; Wang, P.; Zhang, Y. Open-vocabulary video anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2024; pp. 18297–18307. [Google Scholar]
- Zanella, L.; Menapace, W.; Mancini, M.; Wang, Y.; Ricci, E. Harnessing large language models for training-free video anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2024; pp. 18527–18536. [Google Scholar]
- Fang, J.; Li, L.l.; Zhou, J.; Xiao, J.; Yu, H.; Lv, C.; Xue, J.; Chua, T.S. Abductive ego-view accident video understanding for safe driving perception. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2024; pp. 22030–22040. [Google Scholar]
- Liang, R.; Li, Y.; Zhou, J.; Li, X. Text-driven traffic anomaly detection with temporal high-frequency modeling in driving videos. IEEE Trans. Circuits Syst. Video Technol. 2024, 34, 8684–8697. [Google Scholar] [CrossRef]
- Dogru, N.; Subasi, A. Traffic accident detection using random forest classifier. In Proceedings of the 2018 15th Learning and Technology Conference (L&T); IEEE: Piscataway, NJ, USA, 2018; pp. 40–45. [Google Scholar]
- Wang, W.; Chen, S.; Qu, G. Comparison between partial least squares regression and support vector machine for freeway incident detection. In Proceedings of the 2007 IEEE Intelligent Transportation Systems Conference; IEEE: Piscataway, NJ, USA, 2007; pp. 190–195. [Google Scholar]
- Liu, W.; Luo, W.; Lian, D.; Gao, S. Future frame prediction for anomaly detection—A new baseline. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 6536–6545. [Google Scholar]
- Chong, Y.S.; Tay, Y.H. Abnormal event detection in videos using spatiotemporal autoencoder. In Proceedings of the Advances in Neural Networks-ISNN 2017: 14th International Symposium, ISNN 2017, Sapporo, Hakodate, and Muroran, Hokkaido, Japan, 21–26 June 2017; Proceedings, Part II 14; Springer: Berlin/Heidelberg, Germany, 2017; pp. 189–196. [Google Scholar]
- Wang, M.; Xing, J.; Liu, Y. Actionclip: A new paradigm for video action recognition. arXiv 2021, arXiv:2109.08472. [Google Scholar] [CrossRef]
- Luo, H.; Ji, L.; Zhong, M.; Chen, Y.; Lei, W.; Duan, N.; Li, T. Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning. Neurocomputing 2022, 508, 293–304. [Google Scholar] [CrossRef]
- Wu, P.; Zhou, X.; Pang, G.; Zhou, L.; Yan, Q.; Wang, P.; Zhang, Y. Vadclip: Adapting vision-language models for weakly supervised video anomaly detection. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Palo Alto, CA, USA, 2024; Volume 38, pp. 6074–6082. [Google Scholar]
- Fang, J.; Yan, D.; Qiao, J.; Xue, J.; Yu, H. DADA: Driver attention prediction in driving accident scenarios. IEEE Trans. Intell. Transp. Syst. 2021, 23, 4959–4971. [Google Scholar] [CrossRef]
- Ristea, N.C.; Croitoru, F.A.; Ionescu, R.T.; Popescu, M.; Khan, F.S.; Shah, M. Self-distilled masked auto-encoders are efficient video anomaly detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2024; pp. 15984–15995. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2016; pp. 770–778. [Google Scholar]
- Park, H.; Noh, J.; Ham, B. Learning memory-guided normality for anomaly detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2020; pp. 14372–14381. [Google Scholar]
- Liu, Z.; Nie, Y.; Long, C.; Zhang, Q.; Li, G. A hybrid video anomaly detection framework via memory-augmented flow reconstruction and flow-guided frame prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2021; pp. 13588–13597. [Google Scholar]







| Models | Evaluation Metrics (%) | ||
|---|---|---|---|
| Micro AUC | Macro AUC | Accuracy | |
| Resnet50 [37] | 52.7 | 50.3 | 55.4 |
| LS-Resnet50 | 59.0 | 47.2 | 61.8 |
| MNAD(recon) [38] | 54.1 | 58.0 | 53.0 |
| MNAD(pred) [38] | 66.8 | 66.6 | 63.7 |
| OLP-TAD [30] | 61.3 | 63.7 | 63.2 |
| RF-RG [4] | 68.6 | 69.1 | 62.8 |
| Ours | 74.3 | 76.0 | 66.4 |
| Models | Evaluation Metrics (%) | ||
|---|---|---|---|
| Micro AUC | Macro AUC | Accuracy | |
| base | 69.3 | 67.8 | 62.3 |
| base-t | 71.0 | 73.0 | 64.6 |
| base-t-f | 72.4 | 71.4 | 67.2 |
| base-t-f-g | 74.3 | 76.0 | 66.4 |
| Models | Evaluation Metrics (%) | Inference Speed (FPS) | ||
|---|---|---|---|---|
| Micro AUC | Macro AUC | Accuracy | ||
| CLIP(zero_shot) | 38.6 | 39.6 | 39.4 | 83 |
| Ours | 74.3 | 76.0 | 66.4 | 54 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chai, Y.; Silamu, W. Road Traffic Anomaly Detection by Human-Attention-Assisted Text–Vision Learning. Sensors 2026, 26, 2638. https://doi.org/10.3390/s26092638
Chai Y, Silamu W. Road Traffic Anomaly Detection by Human-Attention-Assisted Text–Vision Learning. Sensors. 2026; 26(9):2638. https://doi.org/10.3390/s26092638
Chicago/Turabian StyleChai, Yachuang, and Wushouer Silamu. 2026. "Road Traffic Anomaly Detection by Human-Attention-Assisted Text–Vision Learning" Sensors 26, no. 9: 2638. https://doi.org/10.3390/s26092638
APA StyleChai, Y., & Silamu, W. (2026). Road Traffic Anomaly Detection by Human-Attention-Assisted Text–Vision Learning. Sensors, 26(9), 2638. https://doi.org/10.3390/s26092638

