Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (1,111)

Search Parameters:
Keywords = speech assessment

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
19 pages, 462 KB  
Article
Abstract Language Use in Gifted and Typically Developing Pre-School Children During Free Play: A Comparative Study
by Lingzhu Gong, Shiyu Hu, Jiangbo Hu, Yimin Fu, Hui Zhang and Yanfei Cao
Educ. Sci. 2026, 16(8), 1183; https://doi.org/10.3390/educsci16081183 - 23 Jul 2026
Abstract
This study examines the characteristics of abstract language use in gifted preschool children during natural free-play activities and compares them with those of typically developing children. The aim is to enhance our awareness of characteristics of gifted children language use and support them [...] Read more.
This study examines the characteristics of abstract language use in gifted preschool children during natural free-play activities and compares them with those of typically developing children. The aim is to enhance our awareness of characteristics of gifted children language use and support them in everyday conversations and activities. Abstract language refers to expressions that describe decontextualized situations or that convey causal reasoning and prediction. It is closely related to the development of children’s language, cognition, and thinking. A total of 16 gifted children (10 boys and 6 girls; M = 73.06 months) were identified using two standardized instruments, the Combined Raven’s Test (Chinese revised version) and the Test of Nonverbal Intelligence, Second Edition (TONI-2). All identified children completed both assessments and scored at or above the 99th percentile on the Combined Raven’s Test and the 98th percentile on the TONI-2. Another 16 typically developing children (10 boys and 6 girls; M = 72.81 months) from the same classes were selected as matched controls based on teacher reports and their typical developmental performance. Drawing on corpus analysis, the study analyzed 2778 utterances produced by the gifted children and 2879 utterances produced by the typically developing children during 30-min free-play sessions. The results showed that, as the level of abstract language increased, both groups displayed a similar declining trend in the frequency of language use. However, a cumulative logit mixed-effects model showed that gifted children had significantly higher odds of producing language at higher levels of abstraction than typically developing children. In addition, gifted children demonstrated higher levels of abstract language during educator–child and peer interactions, whereas their private speech was more concentrated at lower levels of abstraction than that of typically developing children. Based on these findings, the study proposes educational recommendations in three areas: recognising children’s advanced thinking, addressing their specific language needs, and leveraging their communicative strengths. These recommendations may help create a more supportive educational environment for gifted children and better foster their developmental potential. Full article
(This article belongs to the Section Early Childhood Education)
Show Figures

Figure 1

25 pages, 4200 KB  
Article
Challenges in Emotion Recognition Across Modalities: A Comparative Analysis
by Rafał Gasz
Appl. Sci. 2026, 16(14), 7239; https://doi.org/10.3390/app16147239 - 20 Jul 2026
Viewed by 170
Abstract
Emotion recognition remains a challenging task despite substantial progress in machine learning and affective computing. This study examines challenges in emotion recognition through a comparative analysis of two widely used modalities: facial images and speech signals. The analysis was conducted using FER-2013 for [...] Read more.
Emotion recognition remains a challenging task despite substantial progress in machine learning and affective computing. This study examines challenges in emotion recognition through a comparative analysis of two widely used modalities: facial images and speech signals. The analysis was conducted using FER-2013 for facial emotion recognition and the TESS and RAVDESS datasets for speech emotion recognition. A MobileNetV2-based approach was applied to visual data, while speech analysis employed MFCC-based representations and both classical and deep learning models. The study combines quantitative performance evaluation with qualitative analysis of classification behavior, focusing on emotion-specific recognition difficulties and recurring error patterns across modalities. Model performance was assessed using accuracy, precision, recall, F1-score, and confusion matrices. Across the analysed datasets, overall classification accuracy ranged from approximately 73% to 96%, while class-level F1-scores ranged from 0.48 to 0.89 depending on the emotion and modality. Happiness and surprise consistently achieved the highest recognition performance, whereas neutral emotion, fear, and disgust exhibited the lowest class-level F1-scores and generated the highest numbers of misclassifications. The experimental results confirmed that happiness and surprise achieved the highest classification performance across modalities, while neutral emotion, fear, and disgust showed reduced recognition accuracy due to weak expressive cues and overlapping feature representations. These difficulties are associated with weak or ambiguous expressive signals, overlap between emotional categories, and variability in emotional expression. The comparative findings suggest that recognition challenges arise from both modality-specific limitations and the inherent properties of emotional expression. The results highlight the importance of multimodal approaches and more flexible representations for improving emotion recognition systems. Full article
(This article belongs to the Special Issue Computational Models and Machine Learning for Biomedical Applications)
Show Figures

Figure 1

20 pages, 505 KB  
Review
AI-Enabled First-Response Support After Sexual and Gender-Based Violence: A PRISMA-ScR Scoping Review
by Paolo Bailo, Chiara Carsana, Maria Garreffa, Anna Carannante, Marco Giustini, Cecilia Fazio, Loredana Falzano, Andrea Piccinini and Simona Gaudi
Healthcare 2026, 14(14), 2174; https://doi.org/10.3390/healthcare14142174 - 18 Jul 2026
Viewed by 203
Abstract
Background: Artificial intelligence (AI) is increasingly proposed to augment early-stage assistance for survivors of sexual and gender-based violence (GBV), including intimate partner and domestic violence, across crisis hotlines, specialist services, digital reporting channels, legal support tools and healthcare pathways. However, the scope, maturity [...] Read more.
Background: Artificial intelligence (AI) is increasingly proposed to augment early-stage assistance for survivors of sexual and gender-based violence (GBV), including intimate partner and domestic violence, across crisis hotlines, specialist services, digital reporting channels, legal support tools and healthcare pathways. However, the scope, maturity and evaluative strength of the peer-reviewed evidence remain uncertain. We aimed to map the application domains, evaluative maturity, and implementation and safety gaps of this evidence base. Methods: We conducted a scoping review reported according to the PRISMA Extension for Scoping Reviews (PRISMA-ScR), using a Population–Concept–Context framework focused on AI-enabled first-response and early support. Searches in Scopus, Web of Science Core Collection and PubMed were supplemented by targeted searches of IEEE Xplore and ACM Digital Library. Records were screened against predefined criteria, charted using a structured form and synthesised descriptively. Results: Original searches yielded 187 records and 21 included sources of evidence. The supplementary search identified 539 records/candidates; 27 full texts were assessed and 6 additional sources met eligibility criteria, yielding 27 included sources of evidence. Evidence covered survivor-facing conversational support; screening and triage in emergency and specialist services; social-triage and online disclosure models; survivor-informed help-seeking and chatbot design; legal/support routing; and enabling modalities such as speech-based approaches. Most sources reported technical performance, usability, acceptability or systems-audit findings, while no workflow-integrated evaluation was identified and survivor-centred effectiveness outcomes, service uptake and adverse-event monitoring were rarely reported. Conclusions: The evidence remains heterogeneous and early-stage, with limited support for service-integrated effectiveness or safety. Included sources more often assessed models, interfaces or prototypes than downstream pathway outcomes. The findings support cautious, pathway-aware interpretation and identify recurring concerns regarding escalation, accountability, equity, digital trace safety and human handover. The proposed practice considerations and outcome domains are author-informed priorities for future pilot and implementation studies, not validated guidelines. Full article
Show Figures

Figure 1

21 pages, 2071 KB  
Review
Voice, Speech, and Large Language Models in Neurology: From Acoustic Biomarkers to Conversational AI
by Shahar Shelly
Computation 2026, 14(7), 160; https://doi.org/10.3390/computation14070160 - 16 Jul 2026
Viewed by 306
Abstract
Background: Speech models (wav2vec 2.0, HuBERT, Whisper), large language models (GPT, LLaMA), and conversational AI have expanded computational speech analysis from handcrafted acoustic features to dialogue-based neurological assessment. How well these approaches address clinical practice has not been evaluated. Methods: We conducted a [...] Read more.
Background: Speech models (wav2vec 2.0, HuBERT, Whisper), large language models (GPT, LLaMA), and conversational AI have expanded computational speech analysis from handcrafted acoustic features to dialogue-based neurological assessment. How well these approaches address clinical practice has not been evaluated. Methods: We conducted a narrative review searching PubMed, Google Scholar, and IEEE Xplore, supplemented by Interspeech and ICASSP proceedings. Findings are organized along three layers: acoustic-motor (voice quality, prosody, articulation), language-transcript (lexical, syntactic, semantic, and discourse analysis), and integrated multimodal-conversational (interactive dialogue systems). Traditional acoustic biomarkers provide background; the primary focus is on foundation models, LLMs, and conversational AI. Findings: Speech foundation models outperform handcrafted features on several classification tasks but degrade on severely impaired speech due to domain mismatch with healthy training data. LLMs classify transcripts and score cognitive tests, but operate on text alone and cannot access acoustic-motor information. Conversational AI can administer cognitive screening through naturalistic dialogue, but validation is limited to small single-centre feasibility studies. Prospective clinical validation remains limited. Cross-linguistic generalizability is untested for most methods. Interpretation: The field is moving toward integrated speech-language assessment, but the gap between technical capability and clinical utility remains wide. Closing it requires diverse multilingual datasets, standardized benchmarks, prospective validation, and ethical governance. Full article
Show Figures

Figure 1

23 pages, 3222 KB  
Article
Thrivaad: A Multilingual, Predictive Eye-Sign-Based AAC System Powered by Optimized Deep Learning
by Rajesh Kannan Megalingam, Sakthiprasad Kuttankulangara Manoharan, Dhilna Cheriyan Manjooran and Dhanaraj Kamble
Sensors 2026, 26(14), 4503; https://doi.org/10.3390/s26144503 - 15 Jul 2026
Viewed by 271
Abstract
Around 1.5% of the global population is suffering from speech impairments; the major causes for this are cerebral palsy and ALS, and the only way for these individuals to communicate is through Augmentative and Alternative Communication (AAC). These systems are either electronic or [...] Read more.
Around 1.5% of the global population is suffering from speech impairments; the major causes for this are cerebral palsy and ALS, and the only way for these individuals to communicate is through Augmentative and Alternative Communication (AAC). These systems are either electronic or non-electronic. Based on new study developments, electronic methods, such as Brain–Computer Interaction (BCI) and eye-gaze-based communication, are assessed as the best choices, but they have their own limitations, incorporating limited adaptability to changing conditions, such as setup variations and user fatigue, which reduces the system‘s robustness. Our previous study, Netravad, shows potential for addressing these gaps, but it lacks multilingual support and will not yield the same results under changing lighting conditions. This study, Thrivaad, provides multilingual support and text prediction and integrates optimized deep learning to accurately capture eye movements even in varying environmental lighting conditions. Thrivaad uses eye movements as input from a webcam, and the Optuna-optimized YOLOv5 model is used to detect the eye direction accurately. Then communication is established in English, Malayalam, and Hindi. The text-prediction feature of this system improves communication by reducing the number of eye gestures required to form a message. This study included a total of 60 participants across three age groups with 35,263 eye-sign images collected. With this data, the YOLOv5 model is trained and then optimized by Optuna. The proposal system provides accurate eye direction, text prediction, multilingual support, and improved adaptability to changing conditions for eye-based AAC. Full article
(This article belongs to the Section Intelligent Sensors)
Show Figures

Figure 1

29 pages, 3972 KB  
Article
The Impact of Anodal tDCS on Verbal and Nonverbal Functional Communication in Subacute Aphasia—An Observational Study
by Ilona Rubi-Fessen, Kathrin Gerbershagen, Prisca Stenneken and Klaus Willmes
Brain Sci. 2026, 16(7), 727; https://doi.org/10.3390/brainsci16070727 - 9 Jul 2026
Viewed by 327
Abstract
Background/Objectives: While a growing number of studies have demonstrated positive effects of adjuvant anodal transcranial direct current stimulation (tDCS) on linguistic performance in aphasia, evidence for corresponding effects on verbal and—in particular—nonverbal functional communication remains absent. This is a critical gap, given [...] Read more.
Background/Objectives: While a growing number of studies have demonstrated positive effects of adjuvant anodal transcranial direct current stimulation (tDCS) on linguistic performance in aphasia, evidence for corresponding effects on verbal and—in particular—nonverbal functional communication remains absent. This is a critical gap, given that restoring everyday communicative competence is the ultimate goal of speech and language therapy (SLT). In the present observational study, we investigated the add-on effect of anodal tDCS over language-relevant left-hemispheric areas on verbal and nonverbal functional communication in n = 34 individuals with subacute aphasia and examined the relationship between communicative and linguistic change. Methods: Participants underwent two consecutive two-week therapy phases (P1, P2), each consisting of 10 SLT sessions. Severity-specific linguistic and communicative assessments were administered before, between, and after both phases: the Aachen Aphasia Test (AAT; n = 26) combined with the Amsterdam–Nijmegen Everyday Language Test (ANELT, A-scale), and the Bielefeld Aphasia Screening Rehabilitation (BIAS-R; n = 8) combined with the Scenario Test. During P2, SLT was supplemented by online anodal tDCS. Results: Overall test performance improved significantly more in tDCS-supported P2 compared to P1 (AAT profile level, p < 0.001; BIAS-R mean percentage value (MPW), p = 0.027; ANELT A-scale raw score Version 1, p = 0.081, and version 2, p = 0.038; and Scenario Test total score, p = 0.003). Significant correlations between AAT profile level and ANELT total scores were found across all time points. Between the MPV and subtests of the BIAS-R and the Scenario total score, there was a tendency toward decreasing correlation levels from T1 to T3. Conclusions: To our knowledge, this is the first study demonstrating that adjuvant tDCS in subacute aphasia enhances not only linguistic performance but also verbal and nonverbal functional communication beyond SLT alone—assessed with standardized, performance-based, ecologically valid instruments. Full article
Show Figures

Figure 1

36 pages, 11266 KB  
Article
Evaluating Moral and Ethical Alignment in Large Language Models: A Dungeons & Dragons Benchmark
by Jan Sawicki, Maria Ganzha and Marcin Paprzycki
Appl. Sci. 2026, 16(13), 6769; https://doi.org/10.3390/app16136769 - 6 Jul 2026
Viewed by 184
Abstract
(1) Background: Large language models (LLMs) are increasingly used to generate morally distinct non-player characters in interactive fiction and tabletop role-playing games. However, prior work shows that reinforcement learning from human feedback (RLHF) safety training suppresses morally extreme expression. Hence, whether models can [...] Read more.
(1) Background: Large language models (LLMs) are increasingly used to generate morally distinct non-player characters in interactive fiction and tabletop role-playing games. However, prior work shows that reinforcement learning from human feedback (RLHF) safety training suppresses morally extreme expression. Hence, whether models can reliably generate and evaluate speech with a fixed moral–ethical alignment, within an assigned fictional persona, remains an open question. (2) Methods: To explore this issue, eight generator LLMs each produced three-turn role-play conversations for 37 Dungeons & Dragons 5th Edition deities spanning all nine cells of the 3 × 3 moral–ethical alignment grid (888 conversations by design, 885 obtained after one generator’s midexperiment deprecation). Next, eight evaluator LLMs classified each conversation on both alignment axes using structured output evaluation (6951 evaluation records), assessed across 17 analysis metrics. (3) Results: The obtained results can be summarised as follows. Moral-axis accuracy (Good/Neutral/Evil) exceeded ethical-axis accuracy (Lawful/Neutral/Chaotic) by approximately 10 percentage points on average across all 64 generator–evaluator pairs, with the pair-level gap ranging up to 32 points and reversing sign for one generator (nemotron-120b). Moreover, the alignment-cell hardness ranged 27.6-fold, from Lawful Good (0.910) to Neutral Neutral (0.033). Furthermore, a strong good-ward moral drift, consistent with RLHF suppression effects reported in prior work, was observed for nemotron-120b (0.931 of Evil-target outputs classified as Good) and was near-absent for grok-4.1-fast (0.024). Finally, deity wiki prominence did not predict per-deity accuracy after controlling for alignment cell (partial ρ < 0.20, p > 0.20). (4) Conclusions: In summary, good-ward drift, consistent with RLHF suppression of morally extreme expression, turned out to be graded and model-specific. Moreover, iconic characters override it (Cyric, Chaotic Evil: 0.947, nearly matching the top Lawful Good score). Finally, alignment-cell identity is the dominant accuracy predictor, not the deity familiarity, within the role-played persona setting studied here. Full article
(This article belongs to the Special Issue Natural Language Processing (NLP): Technologies and Applications)
Show Figures

Figure 1

24 pages, 2865 KB  
Article
A Comparative Investigation of Cepstral Feature Extraction Methods for Deepfake Speech Detection
by Nida Akıncı and Erdal Özbay
Appl. Sci. 2026, 16(13), 6707; https://doi.org/10.3390/app16136707 - 4 Jul 2026
Viewed by 219
Abstract
The widespread adoption of voice-based authentication systems has been accompanied by an escalating threat from deep learning-based synthetic speech generation techniques. This study presents a comparative and experimental investigation of cepstral feature extraction methods for deepfake speech detection. Specifically, Mel-Frequency Cepstral Coefficients (MFCC), [...] Read more.
The widespread adoption of voice-based authentication systems has been accompanied by an escalating threat from deep learning-based synthetic speech generation techniques. This study presents a comparative and experimental investigation of cepstral feature extraction methods for deepfake speech detection. Specifically, Mel-Frequency Cepstral Coefficients (MFCC), Linear-Frequency Cepstral Coefficients (LFCC), and Constant-Q Cepstral Coefficients (CQCC) are systematically evaluated with respect to their frequency scaling characteristics, spectral resolution properties, and capacity to capture artifacts specific to synthetic speech production. Experiments were conducted on 5571 audio samples drawn from the ASVspoof 2021 Logical Access evaluation partition, with all methods assessed under identical classification conditions using a linear Support Vector Machine. Results indicate that CQCC attains the highest numerical performance, achieving 83.59% accuracy, 89.15% ROC-AUC, and 15.83% Equal Error Rate (EER); however, the performance difference between MFCC and CQCC does not reach statistical significance (p = 0.202). Five-fold cross-validation corroborates this finding (CQCC: 87.89% ± 0.81%). McNemar’s test confirms that the performance difference between LFCC and CQCC is statistically significant (p = 0.036). A fine-grained attack-wise analysis across 13 spoofing systems reveals that no single feature representation consistently outperforms the others across all attack types; CQCC achieves the highest accuracy on 6 out of 13 systems, while MFCC remains competitive on several attack categories. The overall findings indicate that deepfake detection performance is highly sensitive not only to the classifier architecture but also to the choice of frequency scale, cepstral transformation design, and data conditions. Empirical motivation is provided that multi-feature strategies integrating complementary frequency representations may offer more robust and generalizable detection solutions. Full article
Show Figures

Figure 1

12 pages, 227 KB  
Article
A Multidimensional Comparison of Psychosocial Adjustment, Functional Hearing, and Device Benefit in Adult Hearing Device Users
by Bamini Gopinath, Alicia Zou, Jessica Turner, George Burlutsky and Diana Tang
Healthcare 2026, 14(13), 1980; https://doi.org/10.3390/healthcare14131980 - 3 Jul 2026
Viewed by 255
Abstract
Background: While previous research has examined individual aspects of hearing loss, such as perceived hearing handicap, functional hearing ability, or hearing device outcomes, these domains are rarely evaluated together within the same cohort. We aimed to address this gap by providing a [...] Read more.
Background: While previous research has examined individual aspects of hearing loss, such as perceived hearing handicap, functional hearing ability, or hearing device outcomes, these domains are rarely evaluated together within the same cohort. We aimed to address this gap by providing a multidimensional assessment of hearing-related outcomes and comparing measures of perceived hearing handicap, psychosocial adjustment to hearing loss, functional hearing ability, and device-related benefit between hearing aid, cochlear implant, and bimodal device users. Methods: Cross-sectional data were collected from 749 adults aged ≥40 years using hearing aids (n = 467), cochlear implants (n = 150), or bimodal devices (n = 132), during 2022–2024. Measures of hearing handicap (Hearing Handicap Inventory for the Elderly Screening, HHIE-S), psychosocial responses to hearing loss (Attitudes towards Loss of Hearing Questionnaire, ALHQ), real-world hearing ability (Speech, Spatial and Qualities of Hearing Scale, SSQ), and outcomes of hearing device use (International Outcome Inventory for Hearing Devices, IOI-HD) were collected. Results: After multivariate adjustment, hearing aid users versus bimodal and cochlear implant users had significantly higher ALHQ scores in the domain of negative associations (suggestive of negative attitudes or stigma attached to hearing aid use): 2.22 (0.05) versus 1.92 (0.09) (p = 0.0067) and 1.90 (0.09) (p = 0.0145), respectively. Hearing aid users had significantly higher SSQ spatial scores (better sound localization and spatial perception) compared to bimodal device users (5.13 versus 4.45, p = 0.0254). Total IOI-HD scores were significantly lower in hearing aid users (reflective of irregular, inconsistent daily usage of the device) compared to cochlear implant users (p < 0.0001) and bimodal device users (p = 0.0007). HHIE-S scores were not significantly different between device user groups after multivariate adjustment. Conclusions: This multidimensional approach provides new insight into the heterogeneity of hearing experiences and may help inform more person-centered models of hearing care. Full article
23 pages, 311 KB  
Article
Evaluating Adversarial Robustness of Deepfake Audio Detectors and Vocoder Fingerprint Detectors Against Universal Adversarial Perturbations
by Quang Minh Tran, Wei Zong, Yang-Wai Chow and Willy Susilo
Future Internet 2026, 18(7), 344; https://doi.org/10.3390/fi18070344 - 29 Jun 2026
Viewed by 343
Abstract
Audio deepfake and vocoder fingerprint detectors are increasingly used to identify synthetic speech and attribute it to its generating model. However, their robustness against adversarial perturbations remains unclear across attack algorithms, perturbation domains, detector representations, and vocoder types. This paper presents a focused, [...] Read more.
Audio deepfake and vocoder fingerprint detectors are increasingly used to identify synthetic speech and attribute it to its generating model. However, their robustness against adversarial perturbations remains unclear across attack algorithms, perturbation domains, detector representations, and vocoder types. This paper presents a focused, quality-aware evaluation of four representative adversarial attacks, namely the Fast Gradient Sign Method (FGSM), Basic Iterative Method (BIM), Projected Gradient Descent (PGD), and Carlini–Wagner (CW) attack, against audio deepfake and vocoder fingerprint detectors. Each attack is implemented in both the waveform domain and the short-time Fourier transform (STFT) magnitude domain. All attacks are optimized against Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks (AASIST) under a targeted fake-to-real objective and are evaluated on synthetic speech generated by HiFi-GAN, Fullband MelGAN, StyleMelGAN, and Parallel WaveGAN. Attack performance is first measured on the source AASIST detector, after which black-box transferability is assessed on three target detector families: ResNet with Linear Frequency Cepstral Coefficient (LFCC) features, LCNN with Constant-Q Cepstral Coefficient (CQCC) features, and a bidirectional long short-term memory (BiLSTM) detector. The results show that adversarial effectiveness depends strongly on perturbation domain and detector representation. STFT-magnitude PGD transfers strongly to LFCC-based ResNet detectors but has limited effect on CQCC-based and recurrent detectors. In contrast, waveform-domain attacks produce broader transferability across feature-based detectors, with different attacks showing distinct ASR–quality trade-offs. Under the chosen waveform-domain budget, FGSM and BIM preserve transcription-level intelligibility while retaining meaningful black-box transferability, whereas CW provides the strongest overall source-detector and black-box attack performance. To distinguish effective adversarial perturbations from destructive signal degradation, we evaluate audio quality and intelligibility using word error rate (WER) and signal-to-noise ratio (SNR). Overall, the findings show that robustness claims in audio deepfake and vocoder fingerprint detection are limited when adversarial perturbations, black-box transferability, and audio quality are jointly considered. Full article
(This article belongs to the Special Issue Adversarial Attacks and Cyber Security)
Show Figures

Graphical abstract

20 pages, 7862 KB  
Article
Identification of Parkinson’s Disease from Native Italian People: Machine Learning Voice Analysis
by Mohammad Amran Hossain, Enea Traini and Francesco Amenta
BioMed 2026, 6(3), 15; https://doi.org/10.3390/biomed6030015 - 29 Jun 2026
Viewed by 231
Abstract
Background: Parkinson’s Disease (PD) is a neurodegenerative disorder frequently accompanied by speech impairments, which could serve as non-invasive biomarkers for early detection. This study investigates the efficacy of machine learning models trained on voice and speech acoustic features for distinguishing PD patients from [...] Read more.
Background: Parkinson’s Disease (PD) is a neurodegenerative disorder frequently accompanied by speech impairments, which could serve as non-invasive biomarkers for early detection. This study investigates the efficacy of machine learning models trained on voice and speech acoustic features for distinguishing PD patients from healthy controls (HC) using the publicly available Italian Parkinson’s Voice and Speech (IPVS) dataset. Methods: A comprehensive set of acoustic features was extracted, including perturbation, prosodic and temporal features, Mel-Frequency Cepstral Coefficients (MFCCs), and Gammatone Cepstral Coefficients (GTCCs). These features were evaluated individually and in combination using six supervised classifiers: Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), K-Nearest Neighbors (KNN), XGBoost (XGB), and Multi-Layer Perceptron (MLP). Results: The best-performing configuration combination of GTCC and acoustic features with the SVM model achieved 94.68% accuracy, 94.37% sensitivity, 95.04% specificity, 95.71% precision, 96.04% F1-score, an MCC value of 0.89, and an ROC-AUC of 0.98. In the combination of all feature sets, the most impressive performance was observed with the MLP classifier. This achieved 93.08% test accuracy, 94.30% sensitivity, 91.18% specificity, 94.30% precision, 94.30 F1-score, an ROC-AUC of 0.97, and an MCC value of 0.85. Conclusions: The findings demonstrate that combining clinically relevant acoustic features with robust machine learning classifiers offers a reliable, interpretable, and computationally efficient solution for PD detection. The use of a publicly available dataset, open-source tools, and subject-wise validation contributes to the reproducibility and clinical relevance of the proposed approach. This study reinforces the potential of speech as a digital biomarker for early PD detection and supports the integration of voice-based assessments into diagnostic platforms. Full article
Show Figures

Figure 1

25 pages, 902 KB  
Article
Baseline Differences in Cochlear Implant Candidates: Bilateral Traditional vs. Expanded Indications
by Jack Y. Lin, Andrew L. S. Thornton, Margaret L. Wilson, Barak M. Spector, Terrin N. Tamati and Aaron C. Moberly
J. Clin. Med. 2026, 15(13), 5068; https://doi.org/10.3390/jcm15135068 - 29 Jun 2026
Viewed by 325
Abstract
Background/Objectives: Cochlear implant (CI) candidacy has expanded beyond traditional bilateral hearing loss (HL) to include single-sided deafness (SSD) and asymmetric hearing loss (AHL), yet baseline differences in speech recognition and patient-reported outcome measures (PROMs) between these groups—bilateral HL, SSD, and AHL—remain poorly characterized. [...] Read more.
Background/Objectives: Cochlear implant (CI) candidacy has expanded beyond traditional bilateral hearing loss (HL) to include single-sided deafness (SSD) and asymmetric hearing loss (AHL), yet baseline differences in speech recognition and patient-reported outcome measures (PROMs) between these groups—bilateral HL, SSD, and AHL—remain poorly characterized. The objective of this study was to characterize and compare preoperative speech recognition performance and PROMs between traditional bilateral HL and SSD/AHL CI candidates, and to examine associations between preoperative word recognition scores and PROMs across the full cohort. Methods: Sixty-eight adults (mean age 71.6 years, SD 7.4) undergoing preoperative CI evaluation were enrolled (31 bilateral HL, 12 SSD, and 25 AHL). Consonant–Nucleus–Consonant (CNC) word recognition and AzBio sentence recognition were assessed for both the ear-to-be-implanted (CI ear) and the contralateral ear. The following PROMs were evaluated: the Speech, Spatial and Qualities of Hearing Scale (SSQ-12); Cochlear Implant Quality of Life–35 (CIQOL-35); Patient Health Questionnaire-2 (PHQ-2); Tinnitus Handicap Inventory (THI); and the Instrumental Activities of Daily Living (IADL). Group comparisons used Mann–Whitney U tests and t-tests. CNC, SSQ-Mean, and CIQOL-Global associations were assessed using multivariable linear regression analysis. Results: Preoperative CI-ear speech recognition did not differ between the bilateral HL and SSD/AHL groups. SSD/AHL candidates had significantly higher contralateral-ear speech recognition performance, better SSQ-12 scores across all domains, and higher CIQOL-35 Global, Communication, Entertainment, and Environment scores compared to bilateral HL candidates. However, the CIQOL-35 Emotional, Listening Effort, and Social domains, PHQ-2, THI, and IADL did not differ significantly between the bilateral HL and SSD/AHL groups. Across our entire sample of candidates, CI-ear CNC scores were not significantly associated with preoperative SSQ-Mean or CIQOL-Global scores, while contralateral-ear CNC scores showed moderate, significant associations with both measures. Conclusions: Traditional bilateral and SSD/AHL CI candidates exhibit distinct preoperative PROM profiles (namely, the SSQ-12 and CIQOL-35) despite having no significant differences in CI-ear speech recognition. Contralateral-ear CNC scores—but not CI-ear scores—were significantly associated with the SSQ-Mean and CIQOL-Global, suggesting that contralateral-ear CNC scores may offer relevant insight into CI candidates’ functional hearing. These findings support population-specific counseling and highlight the complementary value of PROMs and audiometric data in CI candidacy evaluations. Full article
Show Figures

Figure 1

12 pages, 765 KB  
Article
Laryngostroboscopic Screening in Asymptomatic Adults Undergoing Prosthetic Rehabilitation: A Prospective Observational Study
by Desislava Atanasova Konstantinova, Kalina Stoyanova Georgieva-Bozhkova, Anna Kirilova Nenova-Nogalcheva and Stoyan Georgiev Katsarov
Diagnostics 2026, 16(13), 2004; https://doi.org/10.3390/diagnostics16132004 - 27 Jun 2026
Viewed by 257
Abstract
Background and Objectives: Laryngostroboscopy is considered the gold standard for the functional assessment of vocal fold vibration and enables the detection of subtle structural and vibratory abnormalities that may not be apparent during routine examination. In interdisciplinary research involving speech analysis and prosthetic [...] Read more.
Background and Objectives: Laryngostroboscopy is considered the gold standard for the functional assessment of vocal fold vibration and enables the detection of subtle structural and vibratory abnormalities that may not be apparent during routine examination. In interdisciplinary research involving speech analysis and prosthetic rehabilitation, exclusion of underlying laryngeal pathology is methodologically important. The aim of the present study was to evaluate the diagnostic findings obtained through laryngostroboscopic screening in asymptomatic Bulgarian adults examined within a broader research project on speech function and prosthetic rehabilitation. Materials and Methods: A prospective observational study was conducted between April 2022 and July 2023 at the Medical University–Varna, Bulgaria. Eighty adults without self-reported voice-related symptoms underwent laryngostroboscopic examination using an ATMOS Strobo 21 LED system (Advanced Technology Medical Systems GmbH, Lenzkirch, Germany). Participants were assessed for structural and functional laryngeal abnormalities, including alterations in movement frequency, oscillation amplitude, phase symmetry, and visible pathological changes. Descriptive statistics and chi-square tests and Fisher’s exact test analyses were used to evaluate possible associations between laryngeal pathology and demographic variables. Results: Normal laryngeal status was observed in 64 participants (80.0%), whereas 16 (20.0%) showed laryngostroboscopic findings. Isolated vibratory deviations were recorded separately and were not automatically classified as laryngeal pathology. Minor structural or functional variations were found in 5 participants (6.3%), functional laryngeal disorders in 6 (7.5%), benign lesions in 1 (1.3%), and diffuse inflammatory changes consistent with laryngitis in 4 (5.0%). Deviations in vibratory parameters were identified in 25 participants (31.3%) for movement frequency, 16 (20.0%) for oscillation amplitude, and 22 (27.5%) for phase synchronization. No statistically significant associations were found between laryngeal pathology and gender or age group (p > 0.05). Conclusions: Laryngostroboscopic examination identified structural and functional laryngeal findings in a proportion of asymptomatic adults recruited within a speech-function research framework. Functional vibratory deviations were observed more frequently than overt structural pathology. These findings demonstrate that previously unrecognized laryngeal abnormalities may be present even in individuals without apparent voice-related complaints. Further studies incorporating speech-function outcomes and larger cohorts are required to clarify the clinical significance of these observations. Full article
(This article belongs to the Section Biomedical Optics)
Show Figures

Figure 1

21 pages, 1298 KB  
Article
Recovery Phenotypes After Head-and-Neck Reconstructive Surgery: A Prospective Cohort Comparing Free-Flap and Pedicled-Flap Pathways
by Sonia Roxana Burtic, Bogdan Florin Capastraru, Panche Taskov, Daian Ionel Popa, Codrina Mihaela Levai, Livia Stanga, Melania Lavinia Bratu and Adelina Maria Jianu
Diseases 2026, 14(7), 226; https://doi.org/10.3390/diseases14070226 - 23 Jun 2026
Viewed by 235
Abstract
Background: Recovery after major head-and-neck reconstruction extends beyond flap survival and wound closure, involving swallowing, psychological adaptation, body image, and overall quality of life. Integrated multidimensional assessments remain limited in routine reconstructive outcomes research. Aim: The aim of this study was to characterize [...] Read more.
Background: Recovery after major head-and-neck reconstruction extends beyond flap survival and wound closure, involving swallowing, psychological adaptation, body image, and overall quality of life. Integrated multidimensional assessments remain limited in routine reconstructive outcomes research. Aim: The aim of this study was to characterize and compare six-month multidimensional recovery—clinical, functional, nutritional, psychological, and body-image outcomes—between microvascular free-flap and regional pedicled-flap reconstruction and to identify factors that stratify risk for persistent functional and psychosocial impairment. Methods: We conducted a single-center prospective cohort study at the “Victor Babeș” University of Medicine and Pharmacy, Timișoara, Romania, enrolling 87 adults undergoing major reconstructive surgery after ablative treatment of head-and-neck defects (52 microvascular free flaps; 35 regional pedicled flaps). Patients were assessed at baseline and 6 months using the SF-36, WHOQOL-BREF, Body Image Scale (BIS), HADS, PHQ-9, GAD-7, Functional Oral Intake Scale (FOIS), speech intelligibility, and PEG/tracheostomy dependence. Results: At 6 months, most SF-36 and WHOQOL-BREF domains improved with moderate effect sizes (d = 0.3–0.7; all p ≤ 0.009), and body image distress decreased significantly (ΔBIS −2.9 ± 4.6; p < 0.001), whereas social functioning showed no robust gain (p = 0.098; not surviving false-discovery-rate correction). Pedicled reconstruction was associated with higher PEG dependence (37.1% vs. 9.6%; p = 0.005) and worse FOIS (4.7 ± 1.4 vs. 5.6 ± 1.2; p = 0.003). Major complications were linked to blunted or worsening psychological trajectories and a threefold higher rate of clinically significant depression (HADS-D ≥ 11: 66.7% vs. 18.7%; p = 0.001). In a reduced four-predictor multivariable model, pedicled flap (aOR 4.6), adjuvant radiotherapy (aOR 2.8), major complication (aOR 3.3), and lower baseline FOIS (aOR 0.5 per point) were independently associated with PEG dependence (optimism-corrected AUC 0.79). Clustering identified three recovery phenotypes—functional/emotional responders, psychological/body-image responders, and global slow recovery—with significantly different PEG rates (5.9%, 21.4%, 40.0%; p = 0.006). Exploratory mediation analysis suggested that the association between reconstruction technique and mental quality-of-life recovery was partly statistically accounted for by swallowing and body-image improvement. Conclusions: Recovery after major head-and-neck reconstruction is multidimensional and heterogeneous. Baseline swallowing function, reconstruction technique, radiotherapy, and major complications jointly stratify risk for persistent functional and psychosocial impairment, supporting risk-adapted multidisciplinary rehabilitation and early psycho-oncologic screening. Full article
Show Figures

Figure 1

20 pages, 1348 KB  
Article
Auditory Brainstem Response Recorded with the NeuroAudio System in Children Under 3 Years of Age
by Milaine Dominici Sanfins, Diego Lourenço dos Santos Silva, Rhayane Vitória Lopes, Emilia Czaplicka and Piotr Henryk Skarzynski
Life 2026, 16(7), 1044; https://doi.org/10.3390/life16071044 - 23 Jun 2026
Viewed by 530
Abstract
Background: The click-evoked Auditory Brainstem Response (ABR) is the gold standard electrophysiological tool for assessing auditory pathway integrity in infants and young children. As normative data are inherently equipment-specific, the absence of pediatric reference values for the NeuroAudio system (Neurosoft, Ivanovo, Russia) represents [...] Read more.
Background: The click-evoked Auditory Brainstem Response (ABR) is the gold standard electrophysiological tool for assessing auditory pathway integrity in infants and young children. As normative data are inherently equipment-specific, the absence of pediatric reference values for the NeuroAudio system (Neurosoft, Ivanovo, Russia) represents a significant gap in clinical practice, given that existing normative datasets for this system are restricted to adult populations. Objective: To establish normative data for click ABR recorded with the NeuroAudio system in children under three years of age, stratified by age group according to auditory maturation patterns. Methods: A prospective, cross-sectional study was conducted at the Electrophysiology Laboratory of the Department of Speech Therapy, Paulista School of Medicine, Federal University of São Paulo (UNIFESP/EPM), under the approval of the Research Ethics Committee (protocol 7.939.564). A total of 203 children (121 males, 82 females; age range: 2 weeks to 36 months) with confirmed normal peripheral auditory function were included. Click stimuli (0.1 ms, rarefaction polarity) were delivered monaurally via ER-3A insert earphones at 80 dB nHL and a repetition rate of 19.3/s. Two average runs of 2000 artifact-free sweeps were recorded per ear. Absolute latencies of waves I, III, and V, interpeak intervals I–III, III–V, and I–V, and amplitudes of waves I and V were analyzed. Results: Statistical modeling supported the consolidation of 12 initial age bins into three clinically and statistically validated categories: 0–3, 4–12, and 13–36 months. Wave I latency remained stable across age groups, whereas waves III and V and all interpeak intervals showed progressive shortening with increasing age. Wave V amplitude increased progressively with age, while wave I amplitude remained unchanged. Females presented shorter latencies than males for waves III and V and for all interpeak intervals. The right ear exhibited a shorter III–V interpeak interval than the left ear, with a significant ear × age interaction indicating that this asymmetry is modulated during early maturation. Age, sex, and ear-stratified normative values (two SD and three SD reference limits) are reported. Conclusion: This study provides the first pediatric normative dataset for click-evoked ABR acquired with the NeuroAudio system in children under three years of age. The proposed three age stratifications, together with sex- and ear-specific reference values for the III–V interpeak interval, offer a clinically actionable framework for the accurate interpretation of pediatric ABR recordings and for the early identification of auditory pathway abnormalities. Full article
(This article belongs to the Section Physiology and Pathology)
Show Figures

Figure 1

Back to TopTop