Next Article in Journal
A Power Law Reconstruction of Ultrasound Backscatter Images
Next Article in Special Issue
Vocal Directivity of the Greek Singing Voice on the First Three Formant Frequencies
Previous Article in Journal
The Historical Building and Room Acoustics of the Stockholm Public Library (1925–28, 1931–32)
Previous Article in Special Issue
Acoustic Analyses of L1 and L2 Vowel Interactions in Mandarin–Cantonese Late Bilinguals
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Technical Note

Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer

by
Noé Tits
*,
Prernna Bhatnagar
and
Thierry Dutoit
ISIA Lab, Faculty of Engineering, University of Mons, 7000 Mons, Belgium
*
Author to whom correspondence should be addressed.
Acoustics 2024, 6(3), 772-781; https://doi.org/10.3390/acoustics6030042
Submission received: 19 June 2024 / Revised: 20 August 2024 / Accepted: 27 August 2024 / Published: 29 August 2024
(This article belongs to the Special Issue Developments in Acoustic Phonetic Research)

Abstract

In this paper, we present a novel approach for text-independent phone-to-audio alignment based on phoneme recognition, representation learning and knowledge transfer. Our method leverages a self-supervised model (Wav2Vec2) fine-tuned for phoneme recognition using a Connectionist Temporal Classification (CTC) loss, a dimension reduction model and a frame-level phoneme classifier trained using forced-alignment labels (using Montreal Forced Aligner) to produce multi-lingual phonetic representations, thus requiring minimal additional training. We evaluate our model using synthetic native data from the TIMIT dataset and the SCRIBE dataset for American and British English, respectively. Our proposed model outperforms the state-of-the-art (charsiu) in statistical metrics and has applications in language learning and speech processing systems. We leave experiments on other languages for future work but the design of the system makes it easily adaptable to other languages.
Keywords: speech recognition; phoneme recognition; deep learning; transfer learning speech recognition; phoneme recognition; deep learning; transfer learning

Share and Cite

MDPI and ACS Style

Tits, N.; Bhatnagar, P.; Dutoit, T. Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer. Acoustics 2024, 6, 772-781. https://doi.org/10.3390/acoustics6030042

AMA Style

Tits N, Bhatnagar P, Dutoit T. Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer. Acoustics. 2024; 6(3):772-781. https://doi.org/10.3390/acoustics6030042

Chicago/Turabian Style

Tits, Noé, Prernna Bhatnagar, and Thierry Dutoit. 2024. "Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer" Acoustics 6, no. 3: 772-781. https://doi.org/10.3390/acoustics6030042

APA Style

Tits, N., Bhatnagar, P., & Dutoit, T. (2024). Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer. Acoustics, 6(3), 772-781. https://doi.org/10.3390/acoustics6030042

Article Metrics

Back to TopTop