Next Article in Journal
Container-Based Electronic Control Unit Virtualisation: A Paradigm Shift Towards a Centralised Automotive E/E Architecture
Next Article in Special Issue
SentimentFormer: A Transformer-Based Multimodal Fusion Framework for Enhanced Sentiment Analysis of Memes in Under-Resourced Bangla Language
Previous Article in Journal
A Provably Secure and Lightweight Two-Factor Authentication Protocol for Wireless Sensor Network
Previous Article in Special Issue
Named Entity Recognition for Equipment Fault Diagnosis Based on RoBERTa-wwm-ext and Deep Learning Integration
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Picture May Be Worth a Hundred Words for Visual Question Answering †

1
Graduate School of Information Science and Technology, Osaka University, Osaka 565-0871, Japan
2
CyberAgent, Inc., Tokyo 150-0042, Japan
3
Department of Intelligence Science and Technology, Graduate School of Informatics, Kyoto University, Kyoto 606-8507, Japan
*
Author to whom correspondence should be addressed.
This paper is an extended version of our paper published in a paper entitled “Visual Question Answering with Textual Representations for Images”, which was presented at Proceedings of the IEEE/CVF International Conference on Computer Vision, Virtual, 11–17 October 2021.
Electronics 2024, 13(21), 4290; https://doi.org/10.3390/electronics13214290
Submission received: 3 October 2024 / Revised: 24 October 2024 / Accepted: 29 October 2024 / Published: 31 October 2024

Abstract

How far can textual representations go in understanding images? In image understanding, effective representations are essential. Deep visual features from object recognition models currently dominate various tasks, especially Visual Question Answering (VQA). However, these conventional features often struggle to capture image details in ways that match human understanding, and their decision processes lack interpretability. Meanwhile, the recent progress in language models suggests that descriptive text could offer a viable alternative. This paper investigated the use of descriptive text as an alternative to deep visual features in VQA. We propose to process description–question pairs rather than visual features, utilizing a language-only Transformer model. We also explored data augmentation strategies to enhance training set diversity and mitigate statistical bias. Extensive evaluation shows that textual representations using approximately a hundred words can effectively compete with deep visual features on both the VQA 2.0 and VQA-CP v2 datasets. Our qualitative experiments further reveal that these textual representations enable clearer investigation of VQA model decision processes, thereby improving interpretability.
Keywords: visual question answering; textual representations; data augmentation; interpretability; vision-and-language visual question answering; textual representations; data augmentation; interpretability; vision-and-language

Share and Cite

MDPI and ACS Style

Hirota, Y.; Garcia, N.; Otani, M.; Chu, C.; Nakashima, Y. A Picture May Be Worth a Hundred Words for Visual Question Answering. Electronics 2024, 13, 4290. https://doi.org/10.3390/electronics13214290

AMA Style

Hirota Y, Garcia N, Otani M, Chu C, Nakashima Y. A Picture May Be Worth a Hundred Words for Visual Question Answering. Electronics. 2024; 13(21):4290. https://doi.org/10.3390/electronics13214290

Chicago/Turabian Style

Hirota, Yusuke, Noa Garcia, Mayu Otani, Chenhui Chu, and Yuta Nakashima. 2024. "A Picture May Be Worth a Hundred Words for Visual Question Answering" Electronics 13, no. 21: 4290. https://doi.org/10.3390/electronics13214290

APA Style

Hirota, Y., Garcia, N., Otani, M., Chu, C., & Nakashima, Y. (2024). A Picture May Be Worth a Hundred Words for Visual Question Answering. Electronics, 13(21), 4290. https://doi.org/10.3390/electronics13214290

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop