Next Article in Journal
Filament Type Recognition for Additive Manufacturing Using a Spectroscopy Sensor and Machine Learning
Next Article in Special Issue
Sit-and-Reach Pose Detection Based on Self-Train Method and Ghost-ST-GCN
Previous Article in Journal
Event-Triggered State Filter Estimation for INS/DVL Integrated Navigation with Correlated Noise and Outliers
Previous Article in Special Issue
Robust Multi-Subtype Identification of Breast Cancer Pathological Images Based on a Dual-Branch Frequency Domain Fusion Network
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Simultaneous Speech and Eating Behavior Recognition Using Data Augmentation and Two-Stage Fine-Tuning

by
Toshihiro Tsukagoshi
1,
Masafumi Nishida
1 and
Masafumi Nishimura
1,2,*
1
Graduate School of Science and Technology, Shizuoka University, 3-5-1 Johoku, Chuo-ku, Hamamatsu 432-8011, Japan
2
Department of Smart Design, Faculty of Architecture and Design, Aichi Sangyo University, 12-5 Harayama, Oka-cho, Okazaki 444-0005, Japan
*
Author to whom correspondence should be addressed.
Sensors 2025, 25(5), 1544; https://doi.org/10.3390/s25051544
Submission received: 31 December 2024 / Revised: 16 February 2025 / Accepted: 27 February 2025 / Published: 2 March 2025
(This article belongs to the Special Issue AI-Based Automated Recognition and Detection in Healthcare)

Abstract

Speaking and eating are essential components of health management. To enable the daily monitoring of these behaviors, systems capable of simultaneously recognizing speech and eating behaviors are required. However, due to the distinct acoustic and contextual characteristics of these two domains, achieving high-precision integrated recognition remains underexplored. In this study, we propose a method that combines data augmentation through synthetic data creation with a two-stage fine-tuning approach tailored to the complexity of domain adaptation. By concatenating speech and eating sounds of varying lengths and sequences, we generated training data that mimic real-world environments where speech and eating behaviors co-exist. Additionally, efficient model adaptation was achieved through two-stage fine-tuning of the self-supervised learning model. The experimental evaluations demonstrate that the proposed method maintains speech recognition accuracy while achieving high detection performance for eating behaviors, with an F1 score of 0.918 for chewing detection and 0.926 for swallowing detection. These results underscore the potential of using voice recognition technology for daily health monitoring.
Keywords: speech recognition; health monitoring; eating behavior recognition; self-supervised learning; skin-contact microphones speech recognition; health monitoring; eating behavior recognition; self-supervised learning; skin-contact microphones

Share and Cite

MDPI and ACS Style

Tsukagoshi, T.; Nishida, M.; Nishimura, M. Simultaneous Speech and Eating Behavior Recognition Using Data Augmentation and Two-Stage Fine-Tuning. Sensors 2025, 25, 1544. https://doi.org/10.3390/s25051544

AMA Style

Tsukagoshi T, Nishida M, Nishimura M. Simultaneous Speech and Eating Behavior Recognition Using Data Augmentation and Two-Stage Fine-Tuning. Sensors. 2025; 25(5):1544. https://doi.org/10.3390/s25051544

Chicago/Turabian Style

Tsukagoshi, Toshihiro, Masafumi Nishida, and Masafumi Nishimura. 2025. "Simultaneous Speech and Eating Behavior Recognition Using Data Augmentation and Two-Stage Fine-Tuning" Sensors 25, no. 5: 1544. https://doi.org/10.3390/s25051544

APA Style

Tsukagoshi, T., Nishida, M., & Nishimura, M. (2025). Simultaneous Speech and Eating Behavior Recognition Using Data Augmentation and Two-Stage Fine-Tuning. Sensors, 25(5), 1544. https://doi.org/10.3390/s25051544

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop