Regularized Urdu Speech Recognition with Semi-Supervised Deep Learning
AbstractAutomatic Speech Recognition, (ASR) has achieved the best results for English, with end-to-end neural network based supervised models. These supervised models need huge amounts of labeled speech data for good generalization, which can be quite a challenge to obtain for low-resource languages like Urdu. Most models proposed for Urdu ASR are based on Hidden Markov Models (HMMs). This paper proposes an end-to-end neural network model, for Urdu ASR, regularized with dropout, ensemble averaging and Maxout units. Dropout and ensembles are averaging techniques over multiple neural network models while Maxout are units in a neural network which adapt their activation functions. Due to limited labeled data, Semi Supervised Learning (SSL) techniques are also incorporated to improve model generalization. Speech features are transformed into a lower dimensional manifold using an unsupervised dimensionality-reduction technique called Locally Linear Embedding (LLE). Transformed data along with higher dimensional features is used to train neural networks. The proposed model also utilizes label propagation-based self-training of initially trained models and achieves a Word Error Rate (WER) of 4% less than that reported as the benchmark on the same Urdu corpus using HMM. The decrease in WER after incorporating SSL is more significant with an increased validation data size. View Full-Text
Share & Cite This Article
Ali Humayun, M.; Hameed, I.A.; Muslim Shah, S.; Hassan Khan, S.; Zafar, I.; Bin Ahmed, S.; Shuja, J. Regularized Urdu Speech Recognition with Semi-Supervised Deep Learning. Appl. Sci. 2019, 9, 1956.
Ali Humayun M, Hameed IA, Muslim Shah S, Hassan Khan S, Zafar I, Bin Ahmed S, Shuja J. Regularized Urdu Speech Recognition with Semi-Supervised Deep Learning. Applied Sciences. 2019; 9(9):1956.Chicago/Turabian Style
Ali Humayun, Mohammad; Hameed, Ibrahim A.; Muslim Shah, Syed; Hassan Khan, Sohaib; Zafar, Irfan; Bin Ahmed, Saad; Shuja, Junaid. 2019. "Regularized Urdu Speech Recognition with Semi-Supervised Deep Learning." Appl. Sci. 9, no. 9: 1956.
Note that from the first issue of 2016, MDPI journals use article numbers instead of page numbers. See further details here.