Next Article in Journal
PUF-Based Secure Authentication Protocol for Cloud-Assisted Wireless Medical Sensor Networks
Previous Article in Journal
Personalized Learning Path Recommendation Based on Knowledge Graphs: A Survey
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

AirSpeech: Lightweight Speech Synthesis Framework for Home Intelligent Space Service Robots

1
Beijing Research Institute of Automation for Machinery Industry Co., Ltd., Beijing 100120, China
2
Yanqi Lake Institute of Basic Manufacturing Technology Research Co., Ltd., Beijing 100044, China
*
Authors to whom correspondence should be addressed.
Electronics 2026, 15(1), 239; https://doi.org/10.3390/electronics15010239
Submission received: 12 November 2025 / Revised: 17 December 2025 / Accepted: 30 December 2025 / Published: 5 January 2026

Abstract

Text-to-Speech (TTS) methods typically employ a sequential approach with an Acoustic Model (AM) and a vocoder, using a Mel spectrogram as an intermediate representation. However, in home environments, TTS systems often struggle with issues such as inadequate robustness against environmental noise and limited adaptability to diverse speaker characteristics. The quality of the Mel spectrogram directly affects the performance of TTS systems, yet existing methods overlook the potential of enhancing Mel spectrogram quality through more comprehensive speech features. To address the complex acoustic characteristics of home environments, this paper introduces AirSpeech, a post-processing model for Mel-spectrogram synthesis. We adopt a Generative Adversarial Network (GAN) to improve the accuracy of Mel spectrogram prediction and enhance the expressiveness of synthesized speech. By incorporating additional conditioning extracted from synthesized audio using specified speech feature parameters, our method significantly enhances the expressiveness and emotional adaptability of synthesized speech in home environments. Furthermore, we propose a global normalization strategy to stabilize the GAN training process. Through extensive evaluations, we demonstrate that the proposed method significantly improves the signal quality and naturalness of synthesized speech, providing a more user-friendly speech interaction solution for smart home applications.
Keywords: text-to-speech; home service robot; acoustic model; Mel-spectrogram; generative adversarial networks text-to-speech; home service robot; acoustic model; Mel-spectrogram; generative adversarial networks

Share and Cite

MDPI and ACS Style

Qin, X.; Pan, F.; Gao, J.; Huang, S.; Sun, Y.; Zhong, X. AirSpeech: Lightweight Speech Synthesis Framework for Home Intelligent Space Service Robots. Electronics 2026, 15, 239. https://doi.org/10.3390/electronics15010239

AMA Style

Qin X, Pan F, Gao J, Huang S, Sun Y, Zhong X. AirSpeech: Lightweight Speech Synthesis Framework for Home Intelligent Space Service Robots. Electronics. 2026; 15(1):239. https://doi.org/10.3390/electronics15010239

Chicago/Turabian Style

Qin, Xiugong, Fenghu Pan, Jing Gao, Shilong Huang, Yichen Sun, and Xiao Zhong. 2026. "AirSpeech: Lightweight Speech Synthesis Framework for Home Intelligent Space Service Robots" Electronics 15, no. 1: 239. https://doi.org/10.3390/electronics15010239

APA Style

Qin, X., Pan, F., Gao, J., Huang, S., Sun, Y., & Zhong, X. (2026). AirSpeech: Lightweight Speech Synthesis Framework for Home Intelligent Space Service Robots. Electronics, 15(1), 239. https://doi.org/10.3390/electronics15010239

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop