Next Article in Journal
Decoupled Model-Free Adaptive Control with Prediction Features Experimentally Applied to a Three-Tank System Following Time-Varying Trajectories
Previous Article in Journal
Artificial Intelligence in Electric Vehicle Battery Disassembly: A Systematic Review
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems

by
Huthaifa I. Ashqar
1,2,*,
Taqwa I. Alhadidi
3,
Mohammed Elhenawy
4 and
Nour O. Khanfar
5
1
Civil Engineering Department, Arab American University, Jenin P.O. Box 240, Palestine
2
Artificial Intelligence Program, Fu Foundation School of Engineering and Applied Science, Columbia University, New York, NY 10027, USA
3
Civil Engineering Department, Al-Ahliyya Amman University, Amman 19328, Jordan
4
CARRS-Q, Queensland University of Technology, Brisbane, QLD 4001, Australia
5
Natural, Engineering and Technology Sciences Department, Arab American University, Jenin P.O. Box 240, Palestine
*
Author to whom correspondence should be addressed.
Automation 2024, 5(4), 508-526; https://doi.org/10.3390/automation5040029
Submission received: 19 August 2024 / Revised: 19 September 2024 / Accepted: 8 October 2024 / Published: 10 October 2024

Abstract

The integration of thermal imaging data with multimodal large language models (MLLMs) offers promising advancements for enhancing the safety and functionality of autonomous driving systems (ADS) and intelligent transportation systems (ITS). This study investigates the potential of MLLMs, specifically GPT-4 Vision Preview and Gemini 1.0 Pro Vision, for interpreting thermal images for applications in ADS and ITS. Two primary research questions are addressed: the capacity of these models to detect and enumerate objects within thermal images, and to determine whether pairs of image sources represent the same scene. Furthermore, we propose a framework for object detection and classification by integrating infrared (IR) and RGB images of the same scene without requiring localization data. This framework is particularly valuable for enhancing the detection and classification accuracy in environments where both IR and RGB cameras are essential. By employing zero-shot in-context learning for object detection and the chain-of-thought technique for scene discernment, this study demonstrates that MLLMs can recognize objects such as vehicles and individuals with promising results, even in the challenging domain of thermal imaging. The results indicate a high true positive rate for larger objects and moderate success in scene discernment, with a recall of 0.91 and a precision of 0.79 for similar scenes. The integration of IR and RGB images further enhances detection capabilities, achieving an average precision of 0.93 and an average recall of 0.56. This approach leverages the complementary strengths of each modality to compensate for individual limitations. This study highlights the potential of combining advanced AI methodologies with thermal imaging to enhance the accuracy and reliability of ADS, while identifying areas for improvement in model performance.
Keywords: multimodal large language models (MLLMs); thermal images; RGB; object detection; autonomous driving systems multimodal large language models (MLLMs); thermal images; RGB; object detection; autonomous driving systems

Share and Cite

MDPI and ACS Style

Ashqar, H.I.; Alhadidi, T.I.; Elhenawy, M.; Khanfar, N.O. Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems. Automation 2024, 5, 508-526. https://doi.org/10.3390/automation5040029

AMA Style

Ashqar HI, Alhadidi TI, Elhenawy M, Khanfar NO. Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems. Automation. 2024; 5(4):508-526. https://doi.org/10.3390/automation5040029

Chicago/Turabian Style

Ashqar, Huthaifa I., Taqwa I. Alhadidi, Mohammed Elhenawy, and Nour O. Khanfar. 2024. "Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems" Automation 5, no. 4: 508-526. https://doi.org/10.3390/automation5040029

APA Style

Ashqar, H. I., Alhadidi, T. I., Elhenawy, M., & Khanfar, N. O. (2024). Leveraging Multimodal Large Language Models (MLLMs) for Enhanced Object Detection and Scene Understanding in Thermal Images for Autonomous Driving Systems. Automation, 5(4), 508-526. https://doi.org/10.3390/automation5040029

Article Metrics

Back to TopTop