Skip to Content
ElectronicsElectronics
  • This is an early access version, the complete PDF, HTML, and XML versions will be available soon.
  • Article
  • Open Access

10 September 2026

Acoustic-to-Articulatory Inversion Based on Externally Visible Articulator Trajectories and a Channel Attention Mechanism

,
,
,
,
and
1
School of Information Sciences, Beijing Language and Culture University, Beijing 100083, China
2
School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing 100876, China
*
Author to whom correspondence should be addressed.

Abstract

Acoustic-to-articulatory inversion (AAI) predicts articulatory movements from acoustic signals; however, current methods often overlook external, visible articulatory features and frequency domain information. Limited paired acoustic–articulatory data remain a practical constraint on robust AAI modeling and evaluation. This paper proposes a novel AAI framework integrating external visible articulatory features and a channel attention mechanism for feature fusion. We first construct a Mandarin EMA dataset for language learning, then explore the impact of visible features (e.g., lip movements) on predicting internal articulatory trajectories. We additionally investigate short-time Fourier magnitude representations computed from EMA-measured lip–jaw trajectories as auxiliary kinematic features. SENet- and ECA-based group-attention variants are evaluated as alternatives for reweighting input groups before concatenation. Experiments on SAIT-EMA and MOCHA show that visible articulatory features and frequency-domain information improve prediction accuracy, with channel attention fusion providing further gains in selected configurations. These findings demonstrate the value of attention-based multimodal data fusion for intelligent audio and articulatory information processing in pronunciation-learning contexts.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.