Next Article in Journal
Do LLMs Offer a Robust Defense Mechanism Against Membership Inference Attacks on Graph Neural Networks?
Previous Article in Journal
DietQA: A Comprehensive Framework for Personalized Multi-Diet Recipe Retrieval Using Knowledge Graphs, Retrieval-Augmented Generation, and Large Language Models
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

SADAMB: Advancing Spatially-Aware Vision-Language Modeling Through Datasets, Metrics, and Benchmarks

by
Giorgos Papadopoulos
,
Petros Drakoulis
*,
Athanasios Ntovas
,
Alexandros Doumanoglou
and
Dimitris Zarpalas
*
Centre for Research and Technology HELLAS (CERTH), Information Technologies Institute (ITI), 57001 Thessaloniki, Greece
*
Authors to whom correspondence should be addressed.
Computers 2025, 14(10), 413; https://doi.org/10.3390/computers14100413
Submission received: 28 August 2025 / Revised: 16 September 2025 / Accepted: 22 September 2025 / Published: 29 September 2025

Abstract

Understanding spatial relationships between objects in images is crucial for robotic navigation, augmented reality systems, and autonomous driving applications, among others. However, existing vision-language benchmarks often overlook explicit spatial reasoning, limiting progress in this area. We attribute this limitation in part to existing open datasets and evaluation metrics, which tend to overlook spatial details. To address this gap, we make three contributions: First, we greatly extend the COCO dataset with annotations of spatial relations, providing a resource for spatially aware image captioning and visual question answering. Second, we propose a new evaluation framework encompassing metrics that assess image captions’ spatial accuracy at both the sentence and dataset levels. And third, we conduct a benchmark study of various vision encoder–text decoder transformer architectures for image captioning using the introduced dataset and metrics. Results reveal that current models capture spatial information only partially, underscoring the challenges of spatially grounded caption generation.
Keywords: vision-language modeling; spatial relations; spatial grounding; spatial image captioning; spatial visual question answering; dataset; metrics; benchmark vision-language modeling; spatial relations; spatial grounding; spatial image captioning; spatial visual question answering; dataset; metrics; benchmark

Share and Cite

MDPI and ACS Style

Papadopoulos, G.; Drakoulis, P.; Ntovas, A.; Doumanoglou, A.; Zarpalas, D. SADAMB: Advancing Spatially-Aware Vision-Language Modeling Through Datasets, Metrics, and Benchmarks. Computers 2025, 14, 413. https://doi.org/10.3390/computers14100413

AMA Style

Papadopoulos G, Drakoulis P, Ntovas A, Doumanoglou A, Zarpalas D. SADAMB: Advancing Spatially-Aware Vision-Language Modeling Through Datasets, Metrics, and Benchmarks. Computers. 2025; 14(10):413. https://doi.org/10.3390/computers14100413

Chicago/Turabian Style

Papadopoulos, Giorgos, Petros Drakoulis, Athanasios Ntovas, Alexandros Doumanoglou, and Dimitris Zarpalas. 2025. "SADAMB: Advancing Spatially-Aware Vision-Language Modeling Through Datasets, Metrics, and Benchmarks" Computers 14, no. 10: 413. https://doi.org/10.3390/computers14100413

APA Style

Papadopoulos, G., Drakoulis, P., Ntovas, A., Doumanoglou, A., & Zarpalas, D. (2025). SADAMB: Advancing Spatially-Aware Vision-Language Modeling Through Datasets, Metrics, and Benchmarks. Computers, 14(10), 413. https://doi.org/10.3390/computers14100413

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop