Next Article in Journal
Service-Level Interoperability for Distributed Co-Simulation of Heterogeneous Building Performance Models
Next Article in Special Issue
Lightweight Gated Parallel Fusion of CoordAtt and ASPP for Degradation-Robust Monocular Parking-Line Segmentation
Previous Article in Journal
Experimental Approach to Intelligent Estimation of the State-of-Charge (SoC) of Batteries: Case of Electric Vehicles
Previous Article in Special Issue
SiStNet: A Single-Stage Convolutional Neural Network for Vehicle Detection
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Visible–Infrared Image Fusion for Computer Vision: A Review of Datasets and Fusion Strategies in Object Detection and Facial-Expression Recognition

by
Muhammad Tahir Naseem
1,
Chan-Su Lee
1,* and
Muhammad Adnan Khan
2,*
1
Department of Electronic Engineering, Yeungnam University, Gyeongsan-si 38541, Republic of Korea
2
Department of Software, Faculty of Artificial Intelligence and Software, Gachon University, Seongnam-si 13120, Republic of Korea
*
Authors to whom correspondence should be addressed.
Appl. Sci. 2026, 16(13), 6757; https://doi.org/10.3390/app16136757
Submission received: 22 May 2026 / Revised: 19 June 2026 / Accepted: 22 June 2026 / Published: 6 July 2026
(This article belongs to the Special Issue Applied Computer Vision and Deep Learning)

Abstract

Visible and infrared (IR) image fusion has become an important strategy for improving computer vision performance under low illumination, occlusion, and some poor-visibility conditions. By integrating complementary textural information from visible images with thermal or IR cues, VIR fusion can enhance object localization, detection robustness, and facial-expression recognition (FER). This review examines VIR fusion techniques and datasets for computer vision applications, with object detection (OD) considered as a relatively mature scene-level task and FER considered as an emerging human-centered application. It summarizes major multimodal datasets, compares early-fusion approaches, including sensor- and feature-level fusion, with late-fusion approaches, including score- and decision-level fusion, and discusses representative machine learning and deep learning methods. The review also evaluates commonly used performance metrics and identifies current limitations, including dataset imbalance, sensor misalignment, limited demographic diversity in facial-expression datasets, computational complexity, and weak real-time generalization. Finally, key application areas, including surveillance, healthcare, remote sensing, autonomous systems, and human–computer interaction, are discussed. This review highlights the need for better-aligned multimodal datasets, standardized evaluation protocols, lightweight fusion architectures, and robust models capable of operating in dynamic real-world environments.
Keywords: visible–infrared fusion; thermal imaging; object detection; facial-expression recognition; multimodal learning; deep learning; image-fusion datasets; sensor fusion visible–infrared fusion; thermal imaging; object detection; facial-expression recognition; multimodal learning; deep learning; image-fusion datasets; sensor fusion

Share and Cite

MDPI and ACS Style

Naseem, M.T.; Lee, C.-S.; Khan, M.A. Visible–Infrared Image Fusion for Computer Vision: A Review of Datasets and Fusion Strategies in Object Detection and Facial-Expression Recognition. Appl. Sci. 2026, 16, 6757. https://doi.org/10.3390/app16136757

AMA Style

Naseem MT, Lee C-S, Khan MA. Visible–Infrared Image Fusion for Computer Vision: A Review of Datasets and Fusion Strategies in Object Detection and Facial-Expression Recognition. Applied Sciences. 2026; 16(13):6757. https://doi.org/10.3390/app16136757

Chicago/Turabian Style

Naseem, Muhammad Tahir, Chan-Su Lee, and Muhammad Adnan Khan. 2026. "Visible–Infrared Image Fusion for Computer Vision: A Review of Datasets and Fusion Strategies in Object Detection and Facial-Expression Recognition" Applied Sciences 16, no. 13: 6757. https://doi.org/10.3390/app16136757

APA Style

Naseem, M. T., Lee, C.-S., & Khan, M. A. (2026). Visible–Infrared Image Fusion for Computer Vision: A Review of Datasets and Fusion Strategies in Object Detection and Facial-Expression Recognition. Applied Sciences, 16(13), 6757. https://doi.org/10.3390/app16136757

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop