Using Automatic Speech Recognition to Assess Thai Speech Language Fluency in the Montreal Cognitive Assessment (MoCA)
Abstract
1. Introduction
2. Related Work
3. Materials and Methods
3.1. Data Collection
3.2. Data Preprocessing
3.3. Speech Corpus and Data Augmentation
3.4. Model Training
3.5. GMM-HMM Acoustic Model
3.6. DNN-HMM Acoustic Model
3.7. Language Model and Lexicon
3.8. Evaluation Metrics
3.9. Decoding and Scoring
4. Results
4.1. Automatic Speech Recognition Results
4.2. Data Augmentation Analysis
4.3. Word Count and Recoginition Result Analysis
4.4. Language Fluency Assessment
5. Discussion
5.1. Automatic Speech Recognition
5.2. Data Augmentation
5.3. Word Detection Error Analysis
5.3.1. Substitution Error
5.3.2. Deletion Error
5.3.3. Insertion Error
5.4. Language Fluency Assessment
5.5. Limitations and Future Directions
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Acknowledgments
Conflicts of Interest
References
- Deary, I.J.; Corley, J.; Gow, A.; Harris, S.E.; Houlihan, L.M.; Marioni, R.; Penke, L.; Rafnsson, S.B.; Starr, J.M. Age-associated cognitive decline. Br. Med. Bull. 2009, 92, 135–152. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lo, R.Y. The borderland between normal aging and dementia. Tzu Chi Med. J. 2017, 29, 65–71. [Google Scholar] [CrossRef] [Scilit]
- Gale, S.A.; Acar, D.; Daffner, K.R. Dementia. Am. J. Med. 2018, 131, 1161–1169. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Duong, S.; Patel, T.; Chang, F. Dementia. Can. Pharm. J. 2017, 150, 118–129. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Geda, Y.E. Mild Cognitive Impairment in Older Adults. Curr. Psychiatry Rep. 2012, 14, 320–327. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Petersen, R.C.; Doody, R.; Kurz, A.; Mohs, R.C.; Morris, J.C.; Rabins, P.V.; Ritchie, K.; Rossor, M.; Thal, L.; Winblad, B. Current Concepts in Mild Cognitive Impairment. Arch. Neurol. 2001, 58, 1985–1992. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Folstein, M.F.; Folstein, S.E.; McHugh, P.R. Mini-mental state. A practical method for grading the cognitive state of patients for the clinician. J. Psychiatr. Res. 1975, 12, 189–198. [Google Scholar] [CrossRef] [Scilit]
- Nasreddine, Z.S.; Phillips, N.A.; Bédirian, V.; Charbonneau, S.; Whitehead, V.; Collin, I.; Cummings, J.L.; Chertkow, H. The Montreal Cognitive Assessment, MoCA: A Brief Screening Tool for Mild Cognitive Impairment. J. Am. Geriatr. Soc. 2005, 53, 695–699, Corrigendum in J. Am. Geriatr. Soc. 2019, 67, 1991. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ahmed, S.; De Jager, C.; Wilcock, G. A comparison of screening tools for the assessment of Mild Cognitive Impairment: Preliminary findings. Neurocase 2012, 18, 336–351. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wutiwiwatchai, C.; Furui, S. Thai speech processing technology: A review. Speech Commun. 2007, 49, 8–27. [Google Scholar] [CrossRef] [Scilit]
- Koanantakool, H.T.; Karoonboonyanan, T.; Wutiwiwatchai, C. Computers and the Thai Language. IEEE Ann. Hist. Comput. 2009, 31, 46–61. [Google Scholar] [CrossRef] [Scilit]
- Chaiwongsai, J.; Chiracharit, W.; Chamnongthai, K.; Miyanaga, Y. An architecture of HMM-based isolated-word speech recognition with tone detection function. In Proceedings of the 2008 International Symposium on Intelligent Signal Processing and Communications Systems, Bangkok, Thailand, 8–11 February 2009; pp. 11–14. [Google Scholar] [CrossRef] [Scilit]
- König, A.; Satt, A.; Sorin, A.; Hoory, R.; Toledo-Ronen, O.; Derreumaux, A.; Manera, V.; Verhey, F.R.J.; Aalten, P.; Robert, P.H.; et al. Automatic speech analysis for the assessment of patients with predementia and Alzheimer’s disease. Alzheimer’s Dementia: Diagn. Assess. Dis. Monit. 2015, 1, 112–124. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, L.; Fraser, K.C.; Rudzicz, F. Speech Recognition in Alzheimer’s Disease and in its Assessment. In Proceedings of the Interspeech 2016, San Francisco, CA, USA, 8–12 September 2016; pp. 1948–1952. [Google Scholar]
- Povey, D.; Boulianne, G.; Burget, L.; Motlicek, P.; Schwarz, P. The Kaldi Speech Recognition. In Proceedings of the IEEE 2011 Workshop on Automatic Speech Recognition and Understanding, Waikoloa, HI, USA, 11–15 December 2011; Available online: http://kaldi.sf.net/ (accessed on 1 January 2022).
- Eshao, Z.; Ejanse, E.; Evisser, K.; Meyer, A.S. What do verbal fluency tasks measure? Predictors of verbal fluency performance in older adults. Front. Psychol. 2014, 5, 772. [Google Scholar] [CrossRef] [Scilit]
- Pakhomov, S.V.; Marino, S.E.; Banks, S.; Bernick, C. Using automatic speech recognition to assess spoken responses to cognitive tests of semantic verbal fluency. Speech Commun. 2015, 75, 14–26. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tröger, J.; Linz, N.; König, A.; Robert, P.; Alexandersson, J. Telephone-based Dementia Screening I. In Proceedings of the 12th EAI International Conference on Pervasive Computing Technologies for Healthcare, New York, NY, USA, 21–24 May 2018; pp. 59–66. [Google Scholar] [CrossRef] [Scilit]
- Lauraitis, A.; Maskeliūnas, R.; Damaševičius, R.; Krilavičius, T. A Mobile Application for Smart Computer-Aided Self-Administered Testing of Cognition, Speech, and Motor Impairment. Sensors 2020, 20, 3236. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Theera-Umpon, N.; Chansareewittaya, S.; Auephanwiriyakul, S. Phoneme and tonal accent recognition for Thai speech. Expert Syst. Appl. 2011, 38, 13254–13259. [Google Scholar] [CrossRef] [Scilit]
- Hu, X.; Saiko, M.; Hori, C. Incorporating tone features to convolutional neural network to improve Mandarin/Thai speech recognition. In Proceedings of the Signal and Information Processing Association Annual Summit and Conference (APSIPA), Kuala Lumpur, Malaysia, 9–12 December 2014; pp. 1–5. [Google Scholar]
- RNNoise—Recurrent Neural Network for Audio. Available online: https://github.com/xiph/rnnoise (accessed on 1 January 2022).
- Kasuriya, S.; Sornlertlamvanich, V.; Cotsomrong, P.; Kanokphara, S.; Thatphithakkul, N. Thai Speech Corpus for Speech Recognition. In Proceedings of the Oriental COCOSDA, Singapore, 1–3 October 2003; pp. 54–61. [Google Scholar]
- Juang, B.H.; Rabiner, L.R. Automatic Speech Recognition—A Brief History of the Technology Development; Georgia Institute of Technology; Atlanta Rutgers University and the University of California: Santa Barbara, CA, USA, 2005; p. 67. [Google Scholar]
- Hinton, G.; Deng, L.; Yu, D.; Dahl, G.E.; Mohamed, A.-R.; Jaitly, N.; Senior, A.; Vanhoucke, V.; Nguyen, P.; Sainath, T.N.; et al. Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups. IEEE Signal Process. Mag. 2012, 29, 82–97. [Google Scholar] [CrossRef] [Scilit]
- Ghahremani, P.; Baba Ali, B.; Povey, D.; Riedhammer, K.; Trmal, J.; Khudanpur, S. A pitch extraction algorithm tuned for automatic speech recognition. In Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy, 4–9 May 2014; pp. 2494–2498. [Google Scholar] [CrossRef] [Scilit]
- Anastasakos, T.; McDonough, J.; Makhoul, J. Speaker adaptive training: A maximum likelihood approach to speaker normalization. In Proceedings of the 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing, Munich, Germany, 21–24 April 1997; pp. 6–9. [Google Scholar]
- Waibel, A.; Hanazawa, T.; Hinton, G.; Shikano, K.; Lang, K. Phoneme recognition using time-delay neural networks. IEEE Trans. Acoust. Speech Signal Process. 1989, 37, 328–339. [Google Scholar] [CrossRef] [Scilit]
- Hadian, H.; Sameti, H.; Povey, D.; Khudanpur, S. End-to-end Speech Recognition Using Lattice-free MMI. In Proceedings of the 19th Annual Conference of the International Speech Communication Association, Hyderabad, India, 2–6 September 2018; pp. 12–16. [Google Scholar] [CrossRef] [Scilit]


















| Thai Consonants | Initial Consonant Sound | Final Consonant Sound | Thai Consonants | Initial Consonant Sound | Final Consonant Sound |
|---|---|---|---|---|---|
| ก | g | k | ป | bp | p |
| ข, ฃ, ค, ฅ, ฆ, | k | k | ผ, พ, ภ | p | p |
| ง | ng | ng | ฝ, ฟ | f | p |
| จ | j | t | ม | m | m |
| ฉ, ช, ฌ | ch | t | ย | y | y |
| ญ | y | n | ร | r | n |
| ด, ฎ | d | t | ล, ฬ | l | n |
| ถ, ฐ, ท, ธ, ฑ, ฒ | t | t | ว | w | w |
| ต, ฏ | Dt | t | ซ, ศ, ษ, ส, ทร | s | t |
| น, ณ | N | n | ห, ฮ | h | - |
| บ | B | p |
| Thai Script | Phonemic Transcription | English Transcription | Tone | |
|---|---|---|---|---|
| 1 | คา | kh¯aa | a kind of grass | M |
| 2 | ข่า | kh`aa | galingale | L |
| 3 | ฆ่า | khˆaa | to kill | F |
| 4 | ค้า | kh´aa | to trade | H |
| 5 | ขา | khˇaa | a leg | R |
| Corpus | Number of Utterances | Duration (h) |
|---|---|---|
| LOTUS-PD | 3040 | 5:24 |
| Digital MoCA | 780 | 1:40 |
| Augmented Digital MoCA | 3900 | 8:24 |
| Augmented LOTUS + MoCA | 21,628 | 42:10 |
| Group | Train Dataset | Validation Dataset | K-Folds | % WER | |
|---|---|---|---|---|---|
| AVG | SD | ||||
| Baseline | MOCA | MOCA − Dev | 5 | 86.95 | 23.27 |
| LOTUS | LOTUS | 5 | 2.30 | 0.47 | |
| Augmentation | MOCA + LOTUS | MOCA + LOTUS | 5 | 7.91 | 1.46 |
| MOCA+Pitch Shift | MOCA + LOTUS | 5 | 87.11 | 5.21 | |
| MOCA+Pitch Shift+ Noise Reduce | MOCA + LOTUS | 5 | 80.80 | 10.89 | |
| MOCA + LOTUS + AUG (Pitch Shift + Noise Reduce) | MOCA + LOTUS | 5 | 21.20 | 2.90 | |
| Train Dataset | Test Dataset | Features | Model | %WER |
|---|---|---|---|---|
| MOCA + LOTUS + AUG | MOCA | MFCC + Pitch | GMM-HMM | 82.97 |
| MOCA + LOTUS + AUG | MOCA | MFCC + iVector | TDNN-HMM (LF-MMI) | 62.64 |
| MOCA + LOTUS + AUG | MOCA | MFCC | TDNN-HMM (EE LF-MMI) | 31.87 |
| Model | TDNN Layers | Epochs | Beam Search | Acoustic Weight | %WER |
|---|---|---|---|---|---|
| Baseline | 13 | 10 | 15 | 1 | 31.87 |
| EE-LF-MMI-1 | 13 | 5 | 15 | 1 | 31.87 |
| EE-LF-MMI-2 | 15 | 5 | 15 | 1 | 32.42 |
| EE-LF-MMI-3 | 9 | 5 | 15 | 1 | 34.89 |
| EE-LF-MMI-4 | 13 | 5 | 15 | 0.8 | 30.77 |
| EE-LF-MMI-5 | 13 | 5 | 10 | 0.8 | 32.69 |
| Train Dataset | Test Dataset | %WER |
|---|---|---|
| MOCA | MOCA | 95.05 |
| MOCA+Pitch Shift | MOCA | 89.94 |
| MOCA+Pitch Shift+Noise Reduce | MOCA | 86.54 |
| MOCA + LOTUS + AUG (Pitch Shift+Noise Reduce) | MOCA | 82.97 |
| Speaker | ID | #Word | Corr | Sub | Ins | Del | Err |
|---|---|---|---|---|---|---|---|
| F0028 | raw | 24 | 8 | 8 | 0 | 8 | 16 |
| F0028 | sys | 24 | 33.33 | 33.33 | 0 | 33.33 | 66.67 |
| F0031 | raw | 34 | 25 | 7 | 2 | 2 | 11 |
| F0031 | sys | 34 | 73.53 | 20.59 | 5.88 | 5.88 | 32.35 |
| F0034 | raw | 28 | 19 | 4 | 0 | 5 | 9 |
| F0034 | sys | 28 | 67.86 | 14.29 | 0 | 17.86 | 32.14 |
| F0036 | raw | 42 | 34 | 8 | 1 | 0 | 9 |
| F0036 | sys | 42 | 80.95 | 19.05 | 2.38 | 0 | 21.43 |
| F0045 | raw | 22 | 17 | 3 | 0 | 2 | 5 |
| F0045 | sys | 22 | 77.27 | 13.64 | 0 | 9.09 | 22.73 |
| F0049 | raw | 16 | 12 | 4 | 0 | 0 | 4 |
| F0049 | sys | 16 | 75 | 25 | 0 | 0 | 25 |
| F0050 | raw | 46 | 34 | 12 | 0 | 0 | 12 |
| F0050 | sys | 46 | 73.91 | 26.09 | 0 | 0 | 26.09 |
| F0052 | raw | 24 | 16 | 8 | 0 | 0 | 8 |
| F0052 | sys | 24 | 66.67 | 33.33 | 0 | 0 | 33.33 |
| M0005 | raw | 18 | 12 | 6 | 0 | 0 | 6 |
| M0005 | sys | 18 | 66.67 | 33.33 | 0 | 0 | 33.33 |
| M0008 | raw | 18 | 14 | 4 | 0 | 0 | 4 |
| M0008 | sys | 18 | 77.78 | 22.22 | 0 | 0 | 22.22 |
| M0015 | raw | 12 | 9 | 2 | 0 | 1 | 3 |
| M0015 | sys | 12 | 75 | 16.67 | 0 | 8.33 | 25 |
| M0017 | raw | 10 | 7 | 3 | 0 | 0 | 3 |
| M0017 | sys | 10 | 70 | 30 | 0 | 0 | 30 |
| M0020 | raw | 34 | 22 | 11 | 0 | 1 | 12 |
| M0020 | sys | 34 | 64.71 | 32.35 | 0 | 2.94 | 35.29 |
| M0025 | raw | 36 | 26 | 8 | 0 | 2 | 10 |
| M0025 | sys | 36 | 72.22 | 22.22 | 0 | 5.56 | 27.78 |
| SUM | raw | 364 | 255 | 88 | 3 | 21 | 112 |
| SUM | sys | 364 | 70.05 | 24.18 | 0.82 | 5.77 | 30.77 |
| Word | IPA | Count | Word Count in Training | Correctly Detected |
|---|---|---|---|---|
| ไก่ | kày | 12 | 0 | 11 |
| เก็บ | kèp | 10 | 0 | 10 |
| เกี่ยว | kìaw | 10 | 0 | 10 |
| โกรธ | kròot | 10 | 0 | 6 |
| กิน | kin | 10 | 245 | 8 |
| เกี้ยว | kîaw | 6 | 0 | 3 |
| แก้ว | kɛ̂ɛw | 6 | 0 | 6 |
| โกง | kooŋ | 6 | 0 | 3 |
| กด | kòt | 6 | 50 | 3 |
| กบ | kòp | 6 | 120 | 5 |
| Single Words | Long Utterances | |||
|---|---|---|---|---|
| Detection Result | Count | Percent | Count | Percent |
| Correct | 80 | 44% | 175 | 96% |
| Substitution | 82 | 45% | 3 | 2% |
| Deletion | 17 | 9% | 4 | 2% |
| Insertion | 3 | 2% | 0 | 0% |
| Error | Ref | IPA | Detect | IPA | Ref | IPA | Detect | IPA |
|---|---|---|---|---|---|---|---|---|
| Substitution | โกง | kooŋ | กง | koŋ | กอง | kɔɔŋ | ก้อง | kɔ̂ŋ |
| Substitution | เกี้ยว | kîaw | เกี่ยว | kìaw | กาง | kaaŋ | ก้าง | kâaŋ |
| Substitution | กราบ | kràap | กระดาษ | kràdàat | กิน | kin | เก่ง | kèŋ |
| Substitution | โกรธ | kròot | กอด | kɔ̀ɔt | แกง | kɛɛŋ | ก้อง | kɔ̂ŋ |
| Substitution | กก | kòk | กบ | kòp | แก่ | kɛ̀ɛ | ก้อง | kɔ̂ŋ |
| Substitution | กด | kòt | กิน | kin | แก่ | kɛ̀ɛ | แก้ว | kɛ̂ɛw |
| Substitution | กบ | kòp | กง | koŋ | ไกล | klay | ไก่ | kày |
| Substitution | แก | kɛɛ | แก่ | kɛ̀ɛ | ไกล | klay | กลาย | klaay |
| Substitution | กรง | kroŋ | ก้อน | kɔ̂ɔn | ไกว | kway | ไก่ | kày |
| Substitution | กระ | krà | กลับ | klàp | ไกว | kway | ใกล้ | klây |
| Error | Ref | IPA | Detect | IPA | Count |
|---|---|---|---|---|---|
| Deletion | กด | kòt | *** | 2 | |
| Substitution | โกง | kooŋ | กง | koŋ | 3 |
| Substitution | เกี้ยว | kîaw | เกี่ยว | kìaw | 3 |
| Substitution | กราบ | kràap | กระดาษ | kràdàat | 2 |
| Substitution | โกรธ | kròot | กอด | kɔ̀ɔt | 2 |
| Substitution | กด | kòt | กิน | kin | 1 |
| Substitution | กราบ | kràap | กา | kaa | 1 |
| Substitution | กราบ | kràap | กลาด | klâat | 1 |
| Substitution | โกรธ | kròot | กก | kòk | 1 |
| Substitution | โกรธ | kròot | กด | kòt | 1 |
| Ref | Prediction | Count | Ref | Prediction | Count |
|---|---|---|---|---|---|
| เกรี้ยวกราด | *** | 1 | กังขา | *** | 1 |
| เกียรติยศ | *** | 1 | กัด | *** | 1 |
| แก้ม | *** | 1 | ก้าว | *** | 1 |
| กด | *** | 1 | กิ๊บ | *** | 1 |
| กระดุม | *** | 1 | กึก | *** | 1 |
| กระรอก | *** | 1 | กึกก้อง | *** | 1 |
| กรุง | *** | 1 | กุ๊ก | *** | 1 |
| กล้ามเนื้อ | *** | 1 | กุ๊กกิ๊ก | *** | 1 |
| กะเสือกกะสน | *** | 1 |
Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. |
© 2022 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Kantithammakorn, P.; Punyabukkana, P.; Pratanwanich, P.N.; Hemrungrojn, S.; Chunharas, C.; Wanvarie, D. Using Automatic Speech Recognition to Assess Thai Speech Language Fluency in the Montreal Cognitive Assessment (MoCA). Sensors 2022, 22, 1583. https://doi.org/10.3390/s22041583
Kantithammakorn P, Punyabukkana P, Pratanwanich PN, Hemrungrojn S, Chunharas C, Wanvarie D. Using Automatic Speech Recognition to Assess Thai Speech Language Fluency in the Montreal Cognitive Assessment (MoCA). Sensors. 2022; 22(4):1583. https://doi.org/10.3390/s22041583
Chicago/Turabian StyleKantithammakorn, Pimarn, Proadpran Punyabukkana, Ploy N. Pratanwanich, Solaphat Hemrungrojn, Chaipat Chunharas, and Dittaya Wanvarie. 2022. "Using Automatic Speech Recognition to Assess Thai Speech Language Fluency in the Montreal Cognitive Assessment (MoCA)" Sensors 22, no. 4: 1583. https://doi.org/10.3390/s22041583
APA StyleKantithammakorn, P., Punyabukkana, P., Pratanwanich, P. N., Hemrungrojn, S., Chunharas, C., & Wanvarie, D. (2022). Using Automatic Speech Recognition to Assess Thai Speech Language Fluency in the Montreal Cognitive Assessment (MoCA). Sensors, 22(4), 1583. https://doi.org/10.3390/s22041583

