Next Article in Journal
Implementing Computer Vision in Android Apps and Presenting the Background Technology with Mathematical Demonstrations
Previous Article in Journal
A Reinforcement Learning-Based Dynamic Clustering of Sleep Scheduling Algorithm (RLDCSSA-CDG) for Compressive Data Gathering in Wireless Sensor Networks
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Cross-Attention Fusion of Visual and Geometric Features for Large-Vocabulary Arabic Lipreading

1
SMARTS Laboratory, Technopark of Sfax, Sakiet Ezzit, Sfax 3021, Tunisia
2
Digital Research Center of Sfax, Technopark of Sfax, Sakiet Ezzit, Sfax 3021, Tunisia
3
ISSAT Institut Supérieur des Sciences Appliquées et de Technologie, Gafsa University, Sidi Ahmed Zarrouk University Campus, Gafsa 2112, Tunisia
*
Author to whom correspondence should be addressed.
Technologies 2025, 13(1), 26; https://doi.org/10.3390/technologies13010026
Submission received: 15 November 2024 / Revised: 23 December 2024 / Accepted: 26 December 2024 / Published: 9 January 2025
(This article belongs to the Section Information and Communication Technologies)

Abstract

Lipreading involves recognizing spoken words by analyzing the movements of the lips and surrounding area using visual data. It is an emerging research topic with many potential applications, such as human–machine interaction and enhancing audio-based speech recognition. Recent deep learning approaches integrate visual features from the mouth region and lip contours. However, simple methods such as concatenation may not effectively optimize the feature vector. In this article, we propose extracting optimal visual features using 3D convolution blocks followed by a ResNet-18, while employing a graph neural network to extract geometric features from tracked lip landmarks. To fuse these complementary features, we introduce a cross-attention mechanism that combines visual and geometric information to obtain an optimal representation of lip movements for lipreading tasks. To validate our approach for Arabic, we introduce the first large-scale Lipreading in the Wild for Arabic (LRW-AR) dataset, consisting of 20,000 videos across 100 word classes, spoken by 36 speakers. Experimental results on both the LRW-AR and LRW datasets demonstrate the effectiveness of our approach, achieving accuracies of 85.85% and 89.41%, respectively.
Keywords: lipreading; deep learning; LRW-AR; graph neural networks; Transformer; Arabic language lipreading; deep learning; LRW-AR; graph neural networks; Transformer; Arabic language

Share and Cite

MDPI and ACS Style

Daou, S.; Ben-Hamadou, A.; Rekik, A.; Kallel, A. Cross-Attention Fusion of Visual and Geometric Features for Large-Vocabulary Arabic Lipreading. Technologies 2025, 13, 26. https://doi.org/10.3390/technologies13010026

AMA Style

Daou S, Ben-Hamadou A, Rekik A, Kallel A. Cross-Attention Fusion of Visual and Geometric Features for Large-Vocabulary Arabic Lipreading. Technologies. 2025; 13(1):26. https://doi.org/10.3390/technologies13010026

Chicago/Turabian Style

Daou, Samar, Achraf Ben-Hamadou, Ahmed Rekik, and Abdelaziz Kallel. 2025. "Cross-Attention Fusion of Visual and Geometric Features for Large-Vocabulary Arabic Lipreading" Technologies 13, no. 1: 26. https://doi.org/10.3390/technologies13010026

APA Style

Daou, S., Ben-Hamadou, A., Rekik, A., & Kallel, A. (2025). Cross-Attention Fusion of Visual and Geometric Features for Large-Vocabulary Arabic Lipreading. Technologies, 13(1), 26. https://doi.org/10.3390/technologies13010026

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop