1. Introduction
In the Philippines, around 7.5% of children and 14.7% of adults live with hearing impairments, which hamper their ability to communicate effectively, and where communication strategies such as lip-reading come into play [
1,
2]. Lip-reading systems rely solely on visual cues/movements made by the speaker’s lips, or visemes, to interpret and predict the speaker’s speech [
3]. While lip movements are central, other facial features, such as jaw motion, cheek muscle activity, and eye expressions [
4,
5], provide complementary cues that enhance speech recognition accuracy. However, most lip-reading systems focus on major languages, such as English [
6,
7]. Hence, lip-reading systems are limited in applicability in countries such as the Philippines, particularly given its diverse set of dialects [
8].
The dataset employed in this study was derived from the structure and parameters of the GRID Corpus, a widely used audiovisual resource for speech and lip-reading research. Each video clip in the GRID Corpus has a fixed duration of three seconds, recorded at 25 frames per second, yielding 75 frames per clip [
6]. A resolution of 360 × 288 pixels was consistently maintained to ensure standardized visibility of facial and lip movements.
The model architecture adopted in this study was based on LipNet, the first end-to-end lip-reading framework [
9]. LipNet introduced the integration of 3D-CNNs with Bidirectional Long Short-Term Memory (Bi-LSTM) layers, trained using Connectionist Temporal Classification (CTC) loss. This design enables the handling of variable-length input sequences without requiring explicit frame-to-character alignment. The 3D CNN extracts spatial and temporal features from consecutive frames, while Vi-LSTM captures contextual dependencies across time. The CTC loss subsequently aligns predicted outputs with target sequences. Subsequent studies have validated the effectiveness of this CNN–Bi-LSTM–CTC architecture for visual speech recognition [
10,
11].
By developing a lip-reading system for Tagalog, 3D-CNN with Bi-LSTM architecture is constructed to address the current absence of Tagalog-based lip-reading systems and datasets and enhance communication accessibility for the Filipino deaf community and contribute to the expansion of lip-reading research in the Philippines.
2. Materials and Methods
The system’s input is visual-only recordings of a person speaking. In addition to the study’s limitations, it focuses on common and basic Filipino words. 3D CNN pre-processes the video input (including video cropping and lip detection) and predicts the speech through the speaker’s lip movements. Bi-LSTM aids 3D CNN in training and tackles the vanishing gradient issue. After processing, the system outputs the predicted message of the speaker in text.
2.1. Model Architecture
The 3D-CNN layers process video frames as small 3D volumes, allowing the model to learn both spatial lip features and temporal motion. Unlike 2D CNNs that handle frames individually, 3D convolutions use 3 × 3 × 3 filters to capture subtle mouth movements such as opening and shifting between sounds. The extracted features are then passed to Bi-LSTM layers, which process the sequence forward and backward to understand the full context. Each Bi-LSTM layer has 128 hidden units with orthogonal initialization and a 0.5 dropout rate to ensure stable training and minimize overfitting. Finally, Dense and Softmax layers produce character probabilities for each frame, and the Connectionist Temporal Classification (CTC) loss aligns predictions with target sequences without exact frame-to-letter matching. The full architecture is summarized in
Table 1.
2.2. Dataset Creation
The dataset centered around 74 random yet common Tagalog words used in day-to-day communication. Each set consisted of 80 lines, wherein there were three words per line that were randomized. Each video was three seconds long, shot with 25 frames per second, amounting to 75 frames per video, and with a resolution of 360 × 288. Transcription was done manually by matching the lip movement to the corresponding word the speaker spoke and logging the frame ranges to create the alignment files, as illustrated in
Figure 1. A total of 35 sets were recorded, with 25 sets for training the system and 10 sets for testing. A total of four speakers were included in the dataset.
For the dataset, we applied video augmentation techniques to increase diversity and improve the model’s ability to generalize from limited data. Horizontal flipping was performed on all video samples, creating mirrored versions that helped the model recognize lip movements from both orientations. We adjusted brightness and contrast to generate lighter and darker variations, simulating real-world lighting differences. Each original video produced around five new augmented versions, expanding visual variability in lighting, contrast, and perspective.
Following this, we generated lighter and darker versions of every video to simulate real-world illumination differences that occur due to lighting direction, brightness levels, or camera exposure. Although normalization was later applied in preprocessing, brightness-based augmentation introduced meaningful pixel-level variability that normalization alone could not eliminate. This method made the model more robust to illumination changes and improved its ability to accurately predict speech under diverse visual conditions [
12]. As a result, each original video sample produced augmented variants, including flipped, brightened, and darkened versions, which substantially increased the diversity and realism of the dataset and improved model adaptability and overall recognition performance (
Figure 2).
2.3. Preprocessing and Feature Extraction
The preprocessing method adopts LipNet, focusing on precise mouth localization and consistent frame preparation for visual speech recognition [
9]. Facial detection was performed using the Dlib library, with landmarks 48 to 67 used to isolate the mouth region. Each frame was converted to grayscale to emphasize motion over color, then cropped, resized, or padded to fit the model’s input size of 75 × 40 × 100 × 1. Normalization was applied to standardize pixel values and improve feature extraction (
Figure 3).
2.4. System Performance Evaluation
We used character error rate (CER) and word error rate (WER) to evaluate the performance of the lip-reading system, both of which are standard metrics for assessing speech and visual speech recognition systems [
13]. Both metrics are used to measure the accuracy of the system’s predictions by determining the minimum number of corrections required to bring it to ground-truth [
9,
14]. WER enables performance analysis at the sentence level, while CER provides analysis at the character level. Both metrics are computed using the Levenshtein distance algorithm, as shown in Equation (1), where
N denotes the number of reference units (words or characters),
S represents the number of substitutions,
D the number of deletions, and
I is the number of insertions required to align the predicted output with the ground truth [
14].
Every prediction was compared to the ground-truth, wherein word-level performance was measured in terms of overall accuracy (
A), precision (
P), recall (
R), and F1-score (Equation (2)).
where
TP,
TN,
FP, and
FN represent true positive, true negative, false positive, and false negative, respectively.
3. Results and Discussion
The system achieved an overall average CER and WER of 10.09 and 24.08% (
Table 2). These results suggest that minimal corrections are required to align the predictions with the ground-truth transcriptions. The low CER indicates that the system effectively distinguishes between different visemes at the character level. Meanwhile, the relatively higher WER indicates that the system can still accurately predict some words, but its performance decreases slightly when predicting words in a phrase or sentence.
Across the 2398 individual word pair predictions from 800 test videos, the system achieved an overall word accuracy of 76.27%. However, a substantial discrepancy was observed between accuracy and other word-level performance metrics, with a precision of 19.28%, a recall of 15.02%, and an F1-score of 16.68%. This discrepancy suggests the presence of class imbalance, wherein words occurring with higher frequency are more accurately predicted, while those with lower frequency reduce overall system performance. The imbalance is reflected in the system’s CER and WER (
Table 2), as infrequent words contributed disproportionately to errors.
A limitation was also observed in prior lip-reading studies [
8,
15,
16,
17]. Misclassifications were particularly evident in word pairs with overlapping visemes, such as linggo–tao, tatlo–relo, and kailan–saan. Manual transcription confirmed these similarities, reinforcing that viseme overlap remains a fundamental challenge for lip-reading systems.
Despite these limitations, the system demonstrates the potential of the proposed Tagalog lip-reading framework. Its performance is consistent with results reported in comparable studies across other languages under similar constraints, where CER ranged from 30 to 45%, WER from 56 to 73%, and accuracy from 58 to 77% [
6,
15,
16,
17,
18,
19].
Training required four hours per epoch due to the large number of video frames and the computational complexity of the 3D CNN–Bi-LSTM model. The final three epochs were not recorded because of memory limitations. However, checkpointing preserved the latest weights and prevented data loss. The last recorded metrics were a training loss of 3.2946, validation loss of 0.7823, and a learning rate of 1.0 × 10−4. Although training ended prematurely, the model continued to converge effectively.
4. Conclusions and Recommendations
The presented Taglog lip-reading system achieved an overall accuracy of 76.27% with a CER and WER of 10.09 and 24.08%, respectively. The results provide a baseline for Tagalog lip-reading systems, as they are comparable to other studies in the same field. The system effectively captures fine-grained viseme patterns at the character level and maintains reasonable accuracy at the word level. Considering the limitations identified in the study, the dataset needs to be expanded by adding more words and speakers, and alternative methods need to be explored for building the model to address issues such as class imbalance, viseme similarity, long training times, and out-of-memory errors.
Author Contributions
Conceptualization, C.C.P.; methodology, T.J.G.A., A.D.V.P. and C.C.P.; software, A.D.V.P. and T.J.G.A.; validation, A.D.V.P., T.J.G.A. and C.C.P.; formal analysis, T.J.G.A. and A.D.V.P.; investigation, A.D.V.P. and T.J.G.A.; resources, T.J.G.A., A.D.V.P. and C.C.P.; data curation, T.J.G.A., A.D.V.P. and C.C.P.; writing—original draft preparation, T.J.G.A. and A.D.V.P.; writing—review and editing, C.C.P.; visualization, T.J.G.A. and A.D.V.P.; supervision, C.C.P.; project administration, A.D.V.P.; funding acquisition, T.J.G.A. and A.D.V.P. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Ethical review and approval were waived for this study because the data collected consisted of video recordings of the researchers and volunteer participants performing non-invasive, minimal-risk tasks (speaking words aloud) solely for the purpose of algorithm training, presenting no physical or psychological risk to the human subjects.
Informed Consent Statement
Informed consent was obtained from all subjects involved in the study. Written informed consent has been obtained from the participant(s) to publish this paper.
Data Availability Statement
The datasets presented in this article are not readily available due to privacy and ethical restrictions regarding identifiable human subjects. Requests to access the datasets should be directed to the corresponding author.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Newall, J.P.; Martinez, N.; Swanepoel, D.W.; McMahon, C.M. A national survey of hearing loss in the Philippines. Asia Pac. J. Public Health 2020, 32, 235–241. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Santana, R.S.; Menezes, E.C.; Ralin, V.; Givigi, S.; Batorowicz, B. Lipreading as a communication strategy to enhance speech recognition in individuals with hearing impairment: A scoping review. Disabil. Rehabil. Assist. Technol. 2025, 20, 1235–1246. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Deshmukh, N.; Ahire, A.; Bhandari, S.H.; Mali, A.; Warkari, K. Vision based lip reading system using deep learning. In Proceedings of the 2021 International Conference on Computing, Communication and Green Engineering (CCGE); IEEE: New York, NY, USA, 2021; pp. 1–6. [Google Scholar]
- Arceo, A.J.C.; Borejon, R.Y.N.; Hortinela, M.C.R.; Ballado, A.H.; Paglinawan, A.C. Design of an e-attendance checker through facial recognition using histogram of oriented gradients with support vector machine. In Proceedings of the 2020 IEEE 8th International Conference on Smart City and Informatization (iSCI); IEEE: New York, NY, USA, 2020; pp. 1–5. [Google Scholar]
- Balbin, J.R.; Paglinawan, C.C.; de Castro, M.J.A.; Llamas, J.K.C.; Medina, M.E.T.; Pangilinan, J.J.O.; Valiente, F.L. Augmented reality aided analysis of customer satisfaction based on taste-induced facial expression recognition using AFFDEX software developer’s kit. In Proceedings of the 2019 9th International Conference on Biomedical Engineering and Technology; ACM: New York, NY, USA, 2019; pp. 204–209. [Google Scholar]
- Cooke, M.; Barker, J.; Cunningham, S.; Shao, X. The Grid Audio-Visual Speech Corpus (1.0) [Data Set]; Zenodo: Genève, Switzerland, 2006. [Google Scholar] [CrossRef]
- Afouras, T.; Chung, J.S.; Senior, A.; Vinyals, O.; Zisserman, A. Deep audio-visual speech recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 44, 8717–8727. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abadiano, A.B.; Sangilan, F.O.; Cruz, J.P.T. Lip-Reading for Philippine Regional Language Classification of Cebuano and Ilocano. In Proceedings of the 2025 17th International Conference on Computer and Automation Engineering (ICCAE); IEEE: New York, NY, USA, 2025; pp. 1–5. [Google Scholar]
- Assael, Y.M.; Shillingford, B.; Whiteson, S.; De Freitas, N. Lipnet: End-to-end sentence-level lipreading. arXiv 2016, arXiv:1611.01599. [Google Scholar]
- Prashanth, B.S.; Kumar, M.M.; Puneetha, B.H.; Lohith, R.; Gowda, V.D.; Chandan, V.; Sneha, H.R. Lip reading with 3D convolutional and bidirectional LSTM networks on the GRID corpus. In Proceedings of the 2024 Second International Conference on Networks, Multimedia and Information Technology (NMITCON); IEEE: New York, NY, USA, 2024; pp. 1–8. [Google Scholar]
- Shillingford, B.; Assael, Y.; Hoffman, M.W.; Paine, T.; Hughes, C.; Prabhu, U.; Liao, H.; Sak, H.; Rao, K.; Bennett, L.; et al. Large-scale visual speech recognition. arXiv 2018, arXiv:1807.05162. [Google Scholar] [CrossRef] [Scilit]
- Shorten, C.; Khoshgoftaar, T.M. A survey on image data augmentation for deep learning. J. Big Data 2019, 6, 60. [Google Scholar] [CrossRef] [Scilit]
- James, J.; Gopinath, D.P. Advocating character error rate for multilingual asr evaluation. arXiv 2024, arXiv:2410.07400. [Google Scholar] [CrossRef] [Scilit]
- Baevski, A.; Zhou, Y.; Mohamed, A.; Auli, M. wav2vec 2.0: A framework for self-supervised learning of speech representations. Adv. Neural Inf. Process. Syst. 2020, 33, 12449–12460. [Google Scholar]
- Al-Ghanim, A.; Nourah, A.O.; Al-Haidary, R.; Al-Zeer, S.; Altammami, S.; Mahmoud, H.A. I see what you say (iswys): Arabic lip reading system. In Proceedings of the 2013 International Conference on Current Trends in Information Technology (CTIT); IEEE: New York, NY, USA, 2013; pp. 11–17. [Google Scholar]
- Gorman, B.M. Reducing viseme confusion in speech-reading. In ACM SIGACCESS Accessibility and Computing; ACM: New York, NY, USA, 2016; pp. 36–43. [Google Scholar]
- Peymanfard, J.; Saeedi, V.; Mohammadi, M.R.; Zeinali, H.; Mozayani, N. Leveraging Visemes for Better Visual Speech Representation and Lip Reading. arXiv 2023, arXiv:2307.10157. [Google Scholar] [CrossRef] [Scilit]
- Patil, M.S.; Chickerur, S.; Meti, A.; Nabapure, P.M.; Mahindrakar, S.; Naik, S.; Kanyal, S. LSTM based lip reading approach for devanagiri script. ADCAIJ Adv. Distrib. Comput. Artif. Intell. J. 2019, 8, 13–26. [Google Scholar] [CrossRef] [Scilit]
- Fernandez-Lopez, A.; Sukno, F.M. End-to-end lip-reading without large-scale data. IEEE/ACM Trans. Audio Speech Lang. Process. 2022, 30, 2076–2090. [Google Scholar] [CrossRef] [Scilit]
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |