CB-OWL-ViT: A Multimodal Cost-Effective Framework for Contagious Disease Monitoring
Abstract
1. Introduction
1.1. Research Background and Motivations
- RQ 1: How can a mask detection model be developed to robustly adapt to new scenarios without requiring additional retraining?
- RQ 2: How can a social distance estimation model be developed to cost-effectively operate with various types of cameras, including but not limited to stereo cameras?
1.2. Problem Statements and Objectives
1.3. Research Contributions
- Cluster-Based Integration of Open-World Vision-Language Detection: This study introduces Cluster-Based OWL-ViT (CB-OWL-ViT), a novel integration of OWL-ViT with a clustering-based decision mechanism that aggregates evidence from multiple semantically related textual queries. This strategy enhances robustness in mask-wearing assessment under ambiguous visual conditions, without modifying the underlying detection architecture or requiring retraining.
- Practical Use of Established Distance Estimation Techniques: Homography transformation and metric depth estimation, two established geometric and learning-based methods, are incorporated to enable cost-effective social distance estimation using monocular cameras.
- Flexible and Cost-Effective Monitoring Framework: The proposed framework enables rapid adaptation to new monitoring scenarios through query redefinition and supports deployment with a wide range of camera types, addressing practical constraints in large-scale public health monitoring.
- New Dataset and Empirical Evaluation: This study introduces a newly curated dataset collected at the University of Waterloo to empirically evaluate the proposed framework under controlled indoor conditions.
- Comprehensive Comparison with State-of-the-Art (SOTA) Closed-Set Detectors: Extensive experiments compare the proposed framework with recent YOLO variants, highlighting the trade-offs between closed-set supervised detectors and open-world Vision-Language Models (VLMs).
2. Related Work
2.1. Object Detection
2.1.1. Unimodal Object Detection Models
2.1.2. Multimodal Object Detection Models
2.2. Depth Estimation
2.3. Condition Monitoring During Contagious Disease Outbreaks
3. Methodology and Framework
3.1. Object Detection Module
- Text Encoder consists of a 12-layer Transformer architecture incorporating Multi-Head Self-Attention (MHSA) layers with eight attention heads, a Feed-Forward Neural Network (FFNN) with a hidden dimension of 2048, layer normalization, residual connections, and positional embeddings. It employs a model-wide hidden dimension of 512 and generates query embeddings as output, extracted from the final End-Of-Sequence (EOS) token, for use in open-vocabulary object detection tasks.
- Image Encoder employs the standard ViT-B/32 architecture to generate image embeddings from input image patches. It has a hidden dimension of 768 and produces token representations with a sequence length of 576.
- Linear Projection comprises a single fully connected layer without any nonlinear activation function. It maps each output token representation from the ViT-B/32 model, which has a hidden dimension of 768, to an image embedding. This mapping facilitates direct comparison with query embeddings for object classification.
- MLP Head comprises three fully connected layers configured for bounding box prediction. The input dimension is 768, corresponding to the hidden dimension of the ViT-B/32 model. The network includes two intermediate layers with a dimension of 2048, each followed by a ReLU activation function. The final output layer produces four numerical outputs corresponding to the bounding box coordinates—. Consistent with standard Transformer design conventions, the intermediate layers have a dimensionality of 2048. The output layer applies no nonlinear activation function, and the predictions are intentionally biased toward the centers of the respective image patches to enhance localization accuracy.
| Algorithm 1: Detection algorithm |
Require: Image, Queries, OWL-ViT Model Result: boxes, scores, labels
|
| Algorithm 2: Clustering algorithm |
![]() |
| Algorithm 3: Classification algorithm |
![]() |
3.2. Social Distance Estimation Module
3.2.1. Homography Transformation
3.2.2. Metric Depth Estimation
4. Experimental Results
4.1. Setup
4.2. Evaluation Metrics
4.3. Object Detection Module
4.4. Social Distance Estimation Module
5. Evaluation and Discussion
5.1. Comparative Analysis with Other Object Detection Models
5.2. Comparative Analysis with Other SOTA Models
5.3. Practical Deployment Analysis and Model Validation
5.3.1. Inference Speed and Computational Cost
5.3.2. Comparison with VLM-Based Detector (YOLO-World)
5.4. Discussion
6. Conclusions and Future Directions
- Dataset Expansion for OWL-ViT Fine-Tuning: Collecting a task-specific dataset that encompasses diverse mask-wearing scenarios and crowd conditions. This facilitates the effective fine-tuning of the OWL-ViT model, thereby enhancing its applicability to real-world public health monitoring.
- Advanced Depth Modeling for Social Distance Estimation: Developing a more sophisticated and accurate depth estimation model. This approach extends beyond conventional pinhole camera assumptions and has the potential to significantly improve the accuracy of social distance estimates, particularly in complex 3D environments.
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| ML | Machine Learning |
| OWL-ViT | Open-World Localization Vision Transformer |
| CB-OWL-ViT | Cluster-Based OWL-ViT |
| SOTA | State-of-the-Art |
| VLM | Vision-Language Model |
| DSFD | Dual Shot Face Detector |
| CNN | Convolutional Neural Network |
| ViT | Vision Transformer |
| MDE | Monocular Depth Estimation |
| SSD | Single Shot MultiBox Detector |
| RMFD | Real-World Masked Face Dataset |
| MHSA | Multi-Head Self-Attention |
| FFNN | Feed-Forward Neural Network |
| EOS | End-Of-Sequence |
| NMS | Non-Maximum Suppression |
| IoU | Intersection over Union |
| DLT | Direct Linear Transformation |
| AP | Average Precision |
| TP | True Positives |
| FP | False Positives |
| FN | False Negatives |
| mAP | mean AP |
| MAE | Mean Absolute Error |
| MSE | Mean Squared Error |
| RMSE | Root Mean Square Error |
Appendix A. Comparative Detection Results Across Models




Appendix B. Comparative Social Distance Estimation Results Across Methods

References
- Sadjadi, E.N. Challenges and Opportunities for Education Systems with the Current Movement toward Digitalization at the Time of COVID-19. Mathematics 2023, 11, 259. [Google Scholar] [CrossRef] [Scilit]
- Ajagbe, S.A.; Adigun, M.O. Deep learning techniques for detection and prediction of pandemic diseases: A systematic literature review. Multimed. Tools Appl. 2024, 83, 5893–5927. [Google Scholar] [CrossRef] [Scilit]
- Sadjadi, E.N. The recovery plans at the time of COVID-19 foster the journey toward smart city development and sustainability: A narrative review. Environ. Dev. Sustain. 2024, 27, 9743–9771. [Google Scholar] [CrossRef] [Scilit]
- Fatahi, M.; Alizadeh, M.; Moshiri, B. A Novel Model for Student’s Mental Health Monitoring Based on Hard and Soft Data Fusion. In Proceedings of the 2023 31st International Conference on Electrical Engineering (ICEE), Tehran, Iran, 9–11 May 2023; pp. 723–728. [Google Scholar] [CrossRef] [Scilit]
- Kwon, S.; Joshi, A.D.; Lo, C.H.; Drew, D.A.; Nguyen, L.H.; Guo, C.G.; Ma, W.; Mehta, R.S.; Shebl, F.M.; Warner, E.T.; et al. Association of social distancing and face mask use with risk of COVID-19. Nat. Commun. 2021, 12, 3737. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Motaharifar, M.; Norouzzadeh, A.; Abdi, P.; Iranfar, A.; Lotfi, F.; Moshiri, B.; Lashay, A.; Mohammadi, S.F.; Taghirad, H.D. Applications of Haptic Technology, Virtual Reality, and Artificial Intelligence in Medical Training During the COVID-19 Pandemic. Front. Robot. AI 2021, 8, 612949. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Davoodi, M.; Ghaffari, M. Learning-based systems for assessing hazard places of contagious diseases and diagnosing patient possibility. Expert Syst. Appl. 2023, 213, 119043. [Google Scholar] [CrossRef] [Scilit]
- Mokeddem, M.L.; Belahcene, M.; Bourennane, S. Real-time social distance monitoring and face mask detection based Social-Scaled-YOLOv4, DeepSORT and DSFD&MobileNetv2 for COVID-19. Multimed. Tools Appl. 2024, 83, 30613–30639. [Google Scholar] [CrossRef] [Scilit]
- Fatahi, M.; Sadrian Zadeh, D.; Moshiri, B.; Basir, O. Entropy-based genetic feature engineering and multi-classifier fusion for anomaly detection in vehicle controller area networks. Future Gener. Comput. Syst. 2025, 169, 107779. [Google Scholar] [CrossRef] [Scilit]
- Fatahi, M.; Sadrian Zadeh, D.; Ghojogh, B.; Moshiri, B.; Basir, O. An Optimal Cascade Feature-Level Spatiotemporal Fusion Strategy for Anomaly Detection in CAN Bus. arXiv 2025, arXiv:2501.18821. [Google Scholar] [CrossRef] [Scilit]
- Mostafa, S.A.; Ravi, S.; Asaad Zebari, D.; Asaad Zebari, N.; Abed Mohammed, M.; Nedoma, J.; Martinek, R.; Deveci, M.; Ding, W. A YOLO-based deep learning model for Real-Time face mask detection via drone surveillance in public spaces. Inf. Sci. 2024, 676, 120865. [Google Scholar] [CrossRef] [Scilit]
- Thai, C.; Tran, V.; Bui, M.; Nguyen, D.; Ninh, H.; Tran, H. Real-time masked face classification and head pose estimation for RGB facial image via knowledge distillation. Inf. Sci. 2022, 616, 330–347. [Google Scholar] [CrossRef] [Scilit]
- Zeng, D.; Liu, H.; Zhao, F.; Ge, S.; Shen, W.; Zhang, Z. Proposal pyramid networks for fast face detection. Inf. Sci. 2019, 495, 136–149. [Google Scholar] [CrossRef] [Scilit]
- Pagano, C.; Granger, E.; Sabourin, R.; Marcialis, G.; Roli, F. Adaptive ensembles for face recognition in changing video surveillance environments. Inf. Sci. 2014, 286, 75–101. [Google Scholar] [CrossRef] [Scilit]
- Gupta, N.; Mujumdar, S.; Patel, H.; Masuda, S.; Panwar, N.; Bandyopadhyay, S.; Mehta, S.; Guttula, S.; Afzal, S.; Sharma Mittal, R.; et al. Data Quality for Machine Learning Tasks. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, Virtual Event, 14–18 August 2021; pp. 4040–4041. [Google Scholar] [CrossRef] [Scilit]
- Minderer, M.; Gritsenko, A.; Stone, A.; Neumann, M.; Weissenborn, D.; Dosovitskiy, A.; Mahendran, A.; Arnab, A.; Dehghani, M.; Shen, Z.; et al. Simple Open-Vocabulary Object Detection. In Computer Vision—ECCV 2022; Series Title: Lecture Notes in Computer Science; Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T., Eds.; Springer Nature: Cham, Switzerland, 2022; Volume 13670, pp. 728–755. [Google Scholar] [CrossRef] [Scilit]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. arXiv 2024, arXiv:2405.14458v2. [Google Scholar] [CrossRef] [Scilit]
- Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-Centric Real-Time Object Detectors. arXiv 2025, arXiv:2502.12524v1. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs Beat YOLOs on Real-time Object Detection. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 16965–16974. [Google Scholar] [CrossRef] [Scilit]
- Lv, W.; Zhao, Y.; Chang, Q.; Huang, K.; Wang, G.; Liu, Y. RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer. arXiv 2024, arXiv:2407.17140v1. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Xia, C.; Lv, F.; Shi, Y. RT-DETRv3: Real-Time End-to-End Object Detection with Hierarchical Dense Positive Supervision. In Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Tucson, AZ, USA, 28 February–4 March 2025; pp. 1628–1636. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Wang, Y.; Wang, C.; Tai, Y.; Qian, J.; Yang, J.; Wang, C.; Li, J.; Huang, F. DSFD: Dual Shot Face Detector. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 5055–5064. [Google Scholar] [CrossRef] [Scilit]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, PMLR, Virtual, 18–24 July 2021; pp. 8748–8763. [Google Scholar]
- Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.H.; Li, Z.; Duerig, T. Scaling Up Visual and Vision-Language Representation Learning with Noisy Text Supervision. In Proceedings of the 38th International Conference on Machine Learning, PMLR, Virtual, 18–24 July 2021; Volume 139, pp. 4904–4916. [Google Scholar]
- Cheng, T.; Song, L.; Ge, Y.; Liu, W.; Wang, X.; Shan, Y. YOLO-World: Real-Time Open-Vocabulary Object Detection. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 16901–16911. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929v2. [Google Scholar] [CrossRef] [Scilit]
- Bao, H.; Dong, L.; Piao, S.; Wei, F. BEiT: BERT Pre-Training of Image Transformers. arXiv 2021, arXiv:2106.08254v2. [Google Scholar] [CrossRef] [Scilit]
- Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. DINOv2: Learning Robust Visual Features without Supervision. arXiv 2020, arXiv:2304.07193v2. [Google Scholar] [CrossRef] [Scilit]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 10674–10685. [Google Scholar] [CrossRef] [Scilit]
- Yang, L.; Kang, B.; Huang, Z.; Xu, X.; Feng, J.; Zhao, H. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 10371–10381. [Google Scholar] [CrossRef] [Scilit]
- Yang, L.; Kang, B.; Huang, Z.; Zhao, Z.; Xu, X.; Feng, J.; Zhao, H. Depth Anything V2. arXiv 2024, arXiv:2406.09414v2. [Google Scholar] [CrossRef] [Scilit]
- Bochkovskii, A.; Delaunoy, A.; Germain, H.; Santos, M.; Zhou, Y.; Richter, S.R.; Koltun, V. Depth Pro: Sharp Monocular Metric Depth in Less Than a Second. arXiv 2024, arXiv:2410.02073. [Google Scholar] [CrossRef] [Scilit]
- Yadav, S. Deep Learning based Safe Social Distancing and Face Mask Detection in Public Areas for COVID-19 Safety Guidelines Adherence. Int. J. Res. Appl. Sci. Eng. Technol. 2020, 8, 1368–1375. [Google Scholar] [CrossRef] [Scilit]
- Walia, I.S.; Kumar, D.; Sharma, K.; Hemanth, J.D.; Popescu, D.E. An Integrated Approach for Monitoring Social Distancing and Face Mask Detection Using Stacked ResNet-50 and YOLOv5. Electronics 2021, 10, 2996. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Huang, B.; Wang, G.; Yi, P.; Jiang, K. Masked Face Recognition Dataset and Application. IEEE Trans. Biom. Behav. Identity Sci. 2023, 5, 298–304. [Google Scholar] [CrossRef] [Scilit]
- Saponara, S.; Elhanashi, A.; Gagliardi, A. Implementing a real-time, AI-based, people detection and social distancing measuring system for Covid-19. J. Real-Time Image Process. 2021, 18, 1937–1947. [Google Scholar] [CrossRef] [Scilit]
- Razavi, M.; Alikhani, H.; Janfaza, V.; Sadeghi, B.; Alikhani, E. An Automatic System to Monitor the Physical Distance and Face Mask Wearing of Construction Workers in COVID-19 Pandemic. SN Comput. Sci. 2022, 3, 27. [Google Scholar] [CrossRef] [Scilit]
- MVD, A. Face Mask Detection. 2025. Available online: https://www.kaggle.com/datasets/andrewmvd/face-mask-detection (accessed on 9 July 2025).
- Meivel, S.; Sindhwani, N.; Anand, R.; Pandey, D.; Alnuaim, A.A.; Altheneyan, A.S.; Jabarulla, M.Y.; Lelisho, M.E. Mask Detection and Social Distance Identification Using Internet of Things and Faster R-CNN Algorithm. Comput. Intell. Neurosci. 2022, 2022, 2103975. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Elhanashi, A.; Saponara, S.; Dini, P.; Zheng, Q.; Morita, D.; Raytchev, B. An integrated and real-time social distancing, mask detection, and facial temperature video measurement system for pandemic monitoring. J. Real-Time Image Process. 2023, 20, 95. [Google Scholar] [CrossRef] [Scilit]
- Shao, Y.; Ning, J.; Shao, H.; Zhang, D.; Chu, H.; Ren, Z. Lightweight face mask detection algorithm with attention mechanism. Eng. Appl. Artif. Intell. 2024, 137, 109077. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Farhadi, A. YOLO9000: Better, Faster, Stronger. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 6517–6525. [Google Scholar] [CrossRef] [Scilit]
- Bochkovskiy, A.; Wang, C.Y.; Liao, H.Y.M. YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv 2004, arXiv:2004.10934. [Google Scholar] [CrossRef] [Scilit]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 4510–4520. [Google Scholar] [CrossRef] [Scilit]
- Wojke, N.; Bewley, A.; Paulus, D. Simple online and realtime tracking with a deep association metric. In Proceedings of the 2017 IEEE International Conference on Image Processing (ICIP), Beijing, China, 17–20 September 2017; pp. 3645–3649. [Google Scholar] [CrossRef] [Scilit]
- Kuhn, H.W. The Hungarian method for the assignment problem. Nav. Res. Logist. Q. 1955, 2, 83–97. [Google Scholar] [CrossRef] [Scilit]
- Yu, H.; Su, J.; Cai, G.; Piao, Y.; Liu, N.; Huang, M. 3DSAC: Size Adaptive Clustering for 3D object detection in point clouds. Int. J. Appl. Earth Obs. Geoinf. 2023, 118, 103231. [Google Scholar] [CrossRef] [Scilit]
- Hartley, R.; Zisserman, A. Multiple View Geometry in Computer Vision, 2nd ed.; Cambridge University Press: Cambridge, UK, 2004. [Google Scholar] [CrossRef] [Scilit]
- Szeliski, R. Computer Vision: Algorithms and Applications, 2nd ed.; Texts in Computer Science; Springer: Cham, Switzerland, 2022; ISSN 1868-0941, 1868-095X. [Google Scholar] [CrossRef] [Scilit]
- Nelson, J. Mask Wearing Dataset. 2022. Available online: https://universe.roboflow.com/joseph-nelson/mask-wearing (accessed on 9 July 2025).
- Zhang, Y.; Sun, P.; Jiang, Y.; Yu, D.; Weng, F.; Yuan, Z.; Luo, P.; Liu, W.; Wang, X. ByteTrack: Multi-object Tracking by Associating Every Detection Box. In Computer Vision—ECCV 2022; Series Title: Lecture Notes in Computer Science; Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T., Eds.; Springer Nature: Cham, Switzerland, 2022; Volume 13682, pp. 1–21. [Google Scholar] [CrossRef] [Scilit]






| Reference | Year | Mask Monitoring | Distance Monitoring | Dataset Availability |
|---|---|---|---|---|
| [33] | 2020 | SSD | Fixed assumed height (165 cm) | Custom (×) |
| [34] | 2021 | ResNet-50 + DSFD | Stereo camera | RMFD [35] (✓) |
| [36] | 2021 | YOLOv2 | Stereo camera | Custom (×) |
| [37] | 2022 | Faster R-CNN | Stereo camera | Kaggle [38] (✓) |
| [39] | 2022 | Faster R-CNN | Stereo camera | Custom (×) |
| [40] | 2023 | YOLOv4 | Stereo camera | Custom (×) |
| [8] | 2024 | MobileNetV2 + DSFD | DeepSORT with Hungarian algorithm | RMFD [35] (✓) |
| [41] | 2024 | Custom | - | Kaggle [38] (✓) |
| Dataset | Subset | Instances | Class Label | Objects |
|---|---|---|---|---|
| Kaggle [38] | Training | 641 | W/O Mask (0) | 462 |
| W/ Mask (1) | 2255 | |||
| Validation | 113 | W/O Mask (0) | 96 | |
| W/ Mask (1) | 346 | |||
| Roboflow [52] | Training | 1300 | W/O Mask (0) | 1047 |
| W/ Mask (1) | 1894 | |||
| Validation | 350 | W/O Mask (0) | 283 | |
| W/ Mask (1) | 438 |
| Dataset | Subset | Instances | Class Label | Objects |
|---|---|---|---|---|
| Smartphone | Test | 334 | W/O Mask (0) | 196 |
| W/ Mask (1) | 334 | |||
| Webcam | Test | 330 | W/O Mask (0) | 223 |
| W/ Mask (1) | 229 |
| Dataset | Model | Class | mAP@ | P | R | F1 | |
|---|---|---|---|---|---|---|---|
| 0.50 | 0.20 | ||||||
| Kaggle | CB-OWL-ViT | 1 | 0.3214 | 0.5317 | 0.88 | 0.58 | 0.70 |
| 0 | 0.2414 | 0.2792 | 0.39 | 0.67 | 0.49 | ||
| All | 0.2814 | 0.4054 | 0.64 | 0.62 | 0.59 | ||
| OWL-ViT | 1 | 0.3807 | 0.4693 | 0.66 | 0.56 | 0.61 | |
| 0 | 0.2180 | 0.2316 | 0.10 | 0.91 | 0.17 | ||
| All | 0.2994 | 0.3504 | 0.38 | 0.73 | 0.39 | ||
| Roboflow | CB-OWL-ViT | 1 | 0.0780 | 0.5198 | 0.83 | 0.60 | 0.70 |
| 0 | 0.3218 | 0.5570 | 0.62 | 0.81 | 0.71 | ||
| All | 0.1999 | 0.5384 | 0.73 | 0.71 | 0.70 | ||
| OWL-ViT | 1 | 0.0580 | 0.3976 | 0.53 | 0.53 | 0.53 | |
| 0 | 0.4131 | 0.6064 | 0.20 | 0.98 | 0.33 | ||
| All | 0.2356 | 0.5020 | 0.36 | 0.75 | 0.43 | ||
| Smartphone | CB-OWL-ViT | 1 | 0.2884 | 0.8703 | 1.00 | 0.87 | 0.93 |
| 0 | 0.5047 | 0.7886 | 0.79 | 0.99 | 0.88 | ||
| All | 0.3966 | 0.8295 | 0.89 | 0.93 | 0.91 | ||
| OWL-ViT | 1 | 0.5846 | 0.6407 | 0.94 | 0.65 | 0.77 | |
| 0 | 0.3775 | 0.5068 | 0.49 | 1.00 | 0.65 | ||
| All | 0.4811 | 0.5737 | 0.71 | 0.82 | 0.71 | ||
| Webcam | CB-OWL-ViT | 1 | 0.5478 | 0.6761 | 0.97 | 0.69 | 0.81 |
| 0 | 0.8271 | 0.8605 | 0.70 | 0.98 | 0.82 | ||
| All | 0.6875 | 0.7683 | 0.83 | 0.84 | 0.81 | ||
| OWL-ViT | 1 | 0.3281 | 0.3716 | 0.94 | 0.37 | 0.53 | |
| 0 | 0.5622 | 0.5932 | 0.55 | 1.00 | 0.71 | ||
| All | 0.4452 | 0.4824 | 0.75 | 0.68 | 0.62 | ||
| Dataset | Method | MAE | MSE | RMSE |
|---|---|---|---|---|
| Smartphone | Metric Depth Estimation | 0.5300 | 0.3799 | 0.6163 |
| Homography Transformation | 0.1116 | 0.0189 | 0.1376 | |
| Webcam | Metric Depth Estimation | 0.2429 | 0.1076 | 0.3280 |
| Homography Transformation | 0.1364 | 0.0450 | 0.2122 |
| Model | Class | mAP@ | P | R | F1 | |
|---|---|---|---|---|---|---|
| 0.50 | 0.20 | |||||
| YOLOv10m | 1 | 0.8951 | 0.9177 | 0.90 | 0.92 | 0.91 |
| 0 | 0.7514 | 0.7514 | 0.89 | 0.77 | 0.83 | |
| All | 0.8232 | 0.8346 | 0.89 | 0.85 | 0.87 | |
| YOLOv10l | 1 | 0.9042 | 0.9134 | 0.89 | 0.92 | 0.91 |
| 0 | 0.8216 | 0.8436 | 0.91 | 0.85 | 0.88 | |
| All | 0.8629 | 0.8785 | 0.90 | 0.89 | 0.89 | |
| YOLOv12m | 1 | 0.9028 | 0.9140 | 0.91 | 0.92 | 0.92 |
| 0 | 0.7894 | 0.7894 | 0.91 | 0.80 | 0.85 | |
| All | 0.8461 | 0.8517 | 0.91 | 0.86 | 0.88 | |
| YOLOv12l | 1 | 0.8912 | 0.8969 | 0.91 | 0.90 | 0.91 |
| 0 | 0.7361 | 0.7460 | 0.94 | 0.76 | 0.84 | |
| All | 0.8136 | 0.8215 | 0.92 | 0.83 | 0.87 | |
| RT-DETRl | 1 | 0.9153 | 0.9298 | 0.84 | 0.94 | 0.89 |
| 0 | 0.8137 | 0.8348 | 0.80 | 0.85 | 0.82 | |
| All | 0.8645 | 0.8823 | 0.82 | 0.90 | 0.85 | |
| CB-OWL-ViT | 1 | 0.3214 | 0.5317 | 0.88 | 0.58 | 0.70 |
| 0 | 0.2414 | 0.2792 | 0.39 | 0.67 | 0.49 | |
| All | 0.2814 | 0.4054 | 0.64 | 0.62 | 0.59 | |
| Model | Class | mAP@ | P | R | F1 | |
|---|---|---|---|---|---|---|
| 0.50 | 0.20 | |||||
| YOLOv10m | 1 | 0.0553 | 0.9114 | 0.81 | 0.93 | 0.87 |
| 0 | 0.3093 | 0.6825 | 0.91 | 0.69 | 0.78 | |
| All | 0.1823 | 0.7969 | 0.86 | 0.81 | 0.83 | |
| YOLOv10l | 1 | 0.0579 | 0.9007 | 0.84 | 0.91 | 0.87 |
| 0 | 0.3154 | 0.7383 | 0.89 | 0.75 | 0.81 | |
| All | 0.1866 | 0.8195 | 0.86 | 0.83 | 0.84 | |
| YOLOv12m | 1 | 0.0656 | 0.9267 | 0.81 | 0.94 | 0.87 |
| 0 | 0.2659 | 0.7325 | 0.95 | 0.73 | 0.83 | |
| All | 0.1657 | 0.8296 | 0.88 | 0.84 | 0.85 | |
| YOLOv12l | 1 | 0.0631 | 0.9110 | 0.79 | 0.94 | 0.85 |
| 0 | 0.2483 | 0.6839 | 0.97 | 0.69 | 0.80 | |
| All | 0.1557 | 0.7974 | 0.88 | 0.81 | 0.83 | |
| RT-DETRl | 1 | 0.0631 | 0.9110 | 0.79 | 0.94 | 0.85 |
| 0 | 0.2483 | 0.6839 | 0.97 | 0.69 | 0.80 | |
| All | 0.1557 | 0.7974 | 0.88 | 0.81 | 0.83 | |
| CB-OWL-ViT | 1 | 0.0780 | 0.5198 | 0.83 | 0.60 | 0.70 |
| 0 | 0.3218 | 0.5570 | 0.62 | 0.81 | 0.71 | |
| All | 0.1999 | 0.5384 | 0.73 | 0.71 | 0.70 | |
| Model | Class | mAP@ | P | R | F1 | |
|---|---|---|---|---|---|---|
| 0.50 | 0.20 | |||||
| YOLOv10m | 1 | 0.9278 | 0.9969 | 0.95 | 1.00 | 0.97 |
| 0 | 0.5497 | 0.8345 | 0.95 | 0.84 | 0.89 | |
| All | 0.7387 | 0.9157 | 0.95 | 0.92 | 0.93 | |
| YOLOv10l | 1 | 0.9137 | 0.9999 | 0.99 | 1.00 | 0.99 |
| 0 | 0.5770 | 0.8928 | 0.99 | 0.89 | 0.94 | |
| All | 0.7454 | 0.9464 | 0.99 | 0.95 | 0.97 | |
| YOLOv12m | 1 | 0.9339 | 1.0000 | 0.98 | 1.00 | 0.99 |
| 0 | 0.5097 | 0.8736 | 0.94 | 0.88 | 0.91 | |
| All | 0.7218 | 0.9368 | 0.96 | 0.94 | 0.95 | |
| YOLOv12l | 1 | 0.9054 | 1.0000 | 0.89 | 1.00 | 0.94 |
| 0 | 0.5363 | 0.7959 | 1.00 | 0.80 | 0.89 | |
| All | 0.7208 | 0.8980 | 0.94 | 0.90 | 0.91 | |
| RT-DETRl | 1 | 0.9219 | 0.9997 | 0.92 | 1.00 | 0.96 |
| 0 | 0.6169 | 0.8821 | 0.95 | 0.88 | 0.92 | |
| All | 0.7694 | 0.9409 | 0.94 | 0.94 | 0.94 | |
| CB-OWL-ViT | 1 | 0.2884 | 0.8703 | 1.00 | 0.87 | 0.93 |
| 0 | 0.5047 | 0.7886 | 0.79 | 0.99 | 0.88 | |
| All | 0.3966 | 0.8295 | 0.89 | 0.93 | 0.91 | |
| Model | Class | mAP@ | P | R | F1 | |
|---|---|---|---|---|---|---|
| 0.50 | 0.20 | |||||
| YOLOv10m | 1 | 0.5048 | 0.9909 | 0.97 | 0.99 | 0.98 |
| 0 | 0.6720 | 0.6720 | 0.97 | 0.67 | 0.80 | |
| All | 0.5884 | 0.8314 | 0.97 | 0.83 | 0.89 | |
| YOLOv10l | 1 | 0.5371 | 0.9910 | 0.99 | 0.99 | 0.99 |
| 0 | 0.7875 | 0.8201 | 0.99 | 0.82 | 0.90 | |
| All | 0.6623 | 0.9055 | 0.99 | 0.91 | 0.94 | |
| YOLOv12m | 1 | 0.5195 | 0.9865 | 0.99 | 0.99 | 0.99 |
| 0 | 0.7256 | 0.7941 | 0.98 | 0.79 | 0.88 | |
| All | 0.6225 | 0.8903 | 0.98 | 0.89 | 0.93 | |
| YOLOv12l | 1 | 0.5575 | 0.9910 | 0.97 | 0.99 | 0.98 |
| 0 | 0.6752 | 0.6846 | 0.95 | 0.69 | 0.80 | |
| All | 0.6164 | 0.8378 | 0.96 | 0.84 | 0.89 | |
| RT-DETRl | 1 | 0.5619 | 0.9952 | 0.94 | 1.00 | 0.97 |
| 0 | 0.7843 | 0.7936 | 0.94 | 0.79 | 0.86 | |
| All | 0.6731 | 0.8944 | 0.94 | 0.90 | 0.91 | |
| CB-OWL-ViT | 1 | 0.5478 | 0.6761 | 0.97 | 0.69 | 0.81 |
| 0 | 0.8271 | 0.8605 | 0.70 | 0.98 | 0.82 | |
| All | 0.6875 | 0.7683 | 0.83 | 0.84 | 0.81 | |
| Datest | Model | Class | mAP@ | P | R | F1 | |
|---|---|---|---|---|---|---|---|
| 0.50 | 0.20 | ||||||
| Kaggle | ResNet50 + DSFD [34] | 1 | 0.5839 | 0.7216 | 0.77 | 0.88 | 0.82 |
| 0 | 0.4270 | 0.4794 | 0.75 | 0.62 | 0.68 | ||
| All | 0.5055 | 0.6005 | 0.76 | 0.75 | 0.75 | ||
| MobileNetV2 + DSFD [8] | 1 | 0.5591 | 0.6843 | 0.74 | 0.86 | 0.79 | |
| 0 | 0.2885 | 0.3013 | 0.63 | 0.47 | 0.54 | ||
| All | 0.4238 | 0.4928 | 0.69 | 0.66 | 0.67 | ||
| CB-OWL-ViT | 1 | 0.3214 | 0.5317 | 0.88 | 0.58 | 0.70 | |
| 0 | 0.2414 | 0.2792 | 0.39 | 0.67 | 0.49 | ||
| All | 0.2814 | 0.4054 | 0.64 | 0.62 | 0.59 | ||
| Roboflow | ResNet50 + DSFD [34] | 1 | 0.0985 | 0.7170 | 0.74 | 0.92 | 0.82 |
| 0 | 0.3566 | 0.6965 | 0.86 | 0.78 | 0.82 | ||
| All | 0.2275 | 0.7067 | 0.80 | 0.85 | 0.82 | ||
| MobileNetV2 + DSFD [8] | 1 | 0.0953 | 0.6952 | 0.72 | 0.90 | 0.80 | |
| 0 | 0.3605 | 0.6555 | 0.83 | 0.76 | 0.80 | ||
| All | 0.2279 | 0.6754 | 0.78 | 0.83 | 0.80 | ||
| CB-OWL-ViT | 1 | 0.0780 | 0.5198 | 0.83 | 0.60 | 0.70 | |
| 0 | 0.3218 | 0.5570 | 0.62 | 0.81 | 0.71 | ||
| All | 0.1999 | 0.5384 | 0.73 | 0.71 | 0.70 | ||
| Smartphone | ResNet50 + DSFD [34] | 1 | 0.9578 | 0.9608 | 0.93 | 1.00 | 0.96 |
| 0 | 0.5510 | 0.7861 | 0.90 | 0.88 | 0.89 | ||
| All | 0.7544 | 0.8735 | 0.91 | 0.94 | 0.93 | ||
| MobileNetV2 + DSFD [8] | 1 | 0.9935 | 0.9965 | 0.99 | 1.00 | 1.00 | |
| 0 | 0.6306 | 0.9023 | 0.91 | 0.99 | 0.95 | ||
| All | 0.8121 | 0.9494 | 0.95 | 1.00 | 0.97 | ||
| CB-OWL-ViT | 1 | 0.2884 | 0.8703 | 1.00 | 0.87 | 0.93 | |
| 0 | 0.5047 | 0.7886 | 0.79 | 0.99 | 0.88 | ||
| All | 0.3966 | 0.8295 | 0.89 | 0.93 | 0.91 | ||
| Webcam | ResNet50 + DSFD [34] | 1 | 0.2630 | 0.5814 | 0.56 | 1.00 | 0.72 |
| 0 | 0.3050 | 0.3090 | 0.95 | 0.31 | 0.46 | ||
| All | 0.2840 | 0.4432 | 0.76 | 0.65 | 0.59 | ||
| MobileNetV2 + DSFD [8] | 1 | 0.4511 | 0.8867 | 0.95 | 0.89 | 0.92 | |
| 0 | 0.8809 | 0.9028 | 0.85 | 0.96 | 0.90 | ||
| All | 0.6660 | 0.8947 | 0.90 | 0.93 | 0.91 | ||
| CB-OWL-ViT | 1 | 0.5478 | 0.6761 | 0.97 | 0.69 | 0.81 | |
| 0 | 0.8271 | 0.8605 | 0.70 | 0.98 | 0.82 | ||
| All | 0.6875 | 0.7683 | 0.83 | 0.84 | 0.81 | ||
| Dataset | Method | MAE | MSE | RMSE |
|---|---|---|---|---|
| Smartphone | Metric Depth Estimation | 0.5300 | 0.3799 | 0.6163 |
| Homography Transformation | 0.1116 | 0.0189 | 0.1376 | |
| Fixed Human Height [33] | 0.7275 | 0.6555 | 0.8096 | |
| Webcam | Metric Depth Estimation-Based | 0.2429 | 0.1076 | 0.3280 |
| Homography Transformation | 0.1364 | 0.0450 | 0.2122 | |
| Fixed Human Height [33] | 0.9231 | 1.3125 | 1.1456 |
| Model | Number of Parameters | Model Size (MB) | Inference Time (ms/img) |
|---|---|---|---|
| MobileNetV2 + DSFD [8] | 132,803,175 | 608.455 | 45 |
| YOLOv8n | 3,011,238 | 5.94 | 25 |
| ResNet50 + DSFD [34] | 160,424,359 | 932.067 | 68 |
| CB-OWL-ViT | 153,231,879 | 584.53 | 80 |
| Datest | Model | Class | mAP@ | P | R | F1 | |
|---|---|---|---|---|---|---|---|
| 0.50 | 0.20 | ||||||
| Kaggle | CB-OWL-ViT | 1 | 0.3214 | 0.5317 | 0.88 | 0.58 | 0.70 |
| 0 | 0.2414 | 0.2792 | 0.39 | 0.67 | 0.49 | ||
| All | 0.2814 | 0.4054 | 0.64 | 0.62 | 0.59 | ||
| OWL-ViT | 1 | 0.3807 | 0.4693 | 0.66 | 0.56 | 0.61 | |
| 0 | 0.2180 | 0.2316 | 0.10 | 0.91 | 0.17 | ||
| All | 0.2994 | 0.3504 | 0.38 | 0.73 | 0.39 | ||
| YOLO-World | 1 | 0.335 | 0.413 | 0.58 | 0.49 | 0.53 | |
| 0 | 0.192 | 0.204 | 0.09 | 0.80 | 0.16 | ||
| All | 0.265 | 0.308 | 0.34 | 0.64 | 0.35 | ||
| Roboflow | CB-OWL-ViT | 1 | 0.0780 | 0.5198 | 0.83 | 0.60 | 0.70 |
| 0 | 0.3218 | 0.5570 | 0.62 | 0.81 | 0.71 | ||
| All | 0.1999 | 0.5384 | 0.73 | 0.71 | 0.70 | ||
| OWL-ViT | 1 | 0.0580 | 0.3976 | 0.53 | 0.53 | 0.53 | |
| 0 | 0.4131 | 0.6064 | 0.20 | 0.98 | 0.33 | ||
| All | 0.2356 | 0.5020 | 0.36 | 0.75 | 0.43 | ||
| YOLO-World | 1 | 0.051 | 0.350 | 0.47 | 0.47 | 0.47 | |
| 0 | 0.364 | 0.532 | 0.18 | 0.86 | 0.29 | ||
| All | 0.208 | 0.441 | 0.32 | 0.66 | 0.38 | ||
| Smartphone | CB-OWL-ViT | 1 | 0.2884 | 0.8703 | 1.00 | 0.87 | 0.93 |
| 0 | 0.5047 | 0.7886 | 0.79 | 0.99 | 0.88 | ||
| All | 0.3966 | 0.8295 | 0.89 | 0.93 | 0.91 | ||
| OWL-ViT | 1 | 0.5846 | 0.6407 | 0.94 | 0.65 | 0.77 | |
| 0 | 0.3775 | 0.5068 | 0.49 | 1.00 | 0.65 | ||
| All | 0.4811 | 0.5737 | 0.71 | 0.82 | 0.71 | ||
| YOLO-World | 1 | 0.514 | 0.565 | 0.83 | 0.57 | 0.67 | |
| 0 | 0.332 | 0.446 | 0.43 | 0.88 | 0.54 | ||
| All | 0.423 | 0.506 | 0.63 | 0.73 | 0.64 | ||
| Webcam | CB-OWL-ViT | 1 | 0.5478 | 0.6761 | 0.97 | 0.69 | 0.81 |
| 0 | 0.8271 | 0.8605 | 0.70 | 0.98 | 0.82 | ||
| All | 0.6875 | 0.7683 | 0.83 | 0.84 | 0.81 | ||
| OWL-ViT | 1 | 0.3281 | 0.3716 | 0.94 | 0.37 | 0.53 | |
| 0 | 0.5622 | 0.5932 | 0.55 | 1.00 | 0.71 | ||
| All | 0.4452 | 0.4824 | 0.75 | 0.68 | 0.62 | ||
| YOLO-World | 1 | 0.290 | 0.327 | 0.83 | 0.33 | 0.47 | |
| 0 | 0.495 | 0.522 | 0.48 | 0.88 | 0.60 | ||
| All | 0.392 | 0.425 | 0.65 | 0.61 | 0.54 | ||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Fatahi, M.; Sadrian Zadeh, D.; Noormohammadi-Asl, A.; Moshiri, B.; Basir, O.A.; Sadjadi, E.N.; García-Herrero, J.; Molina, J.M. CB-OWL-ViT: A Multimodal Cost-Effective Framework for Contagious Disease Monitoring. Mathematics 2026, 14, 647. https://doi.org/10.3390/math14040647
Fatahi M, Sadrian Zadeh D, Noormohammadi-Asl A, Moshiri B, Basir OA, Sadjadi EN, García-Herrero J, Molina JM. CB-OWL-ViT: A Multimodal Cost-Effective Framework for Contagious Disease Monitoring. Mathematics. 2026; 14(4):647. https://doi.org/10.3390/math14040647
Chicago/Turabian StyleFatahi, Mohammad, Danial Sadrian Zadeh, Ali Noormohammadi-Asl, Behzad Moshiri, Otman A. Basir, Ebrahim Navid Sadjadi, Jesús García-Herrero, and José M. Molina. 2026. "CB-OWL-ViT: A Multimodal Cost-Effective Framework for Contagious Disease Monitoring" Mathematics 14, no. 4: 647. https://doi.org/10.3390/math14040647
APA StyleFatahi, M., Sadrian Zadeh, D., Noormohammadi-Asl, A., Moshiri, B., Basir, O. A., Sadjadi, E. N., García-Herrero, J., & Molina, J. M. (2026). CB-OWL-ViT: A Multimodal Cost-Effective Framework for Contagious Disease Monitoring. Mathematics, 14(4), 647. https://doi.org/10.3390/math14040647



