Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (54)

Search Parameters:
Keywords = prosody analysis

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
18 pages, 563 KB  
Article
The Topicalization Comparative Construction—Evidence from Sinitic and Beyond
by Wen Lu
Languages 2026, 11(9), 192; https://doi.org/10.3390/languages11090192 - 17 Sep 2026
Viewed by 90
Abstract
In this paper, we discuss topicalization comparatives, an underrecognized subtype within the Heinean event-schema framework. Drawing on data from Nyanja, Malay, and Sinitic languages, the analysis corroborates the existence of a construction marked by a fronted topic with an obligatory standard of comparison [...] Read more.
In this paper, we discuss topicalization comparatives, an underrecognized subtype within the Heinean event-schema framework. Drawing on data from Nyanja, Malay, and Sinitic languages, the analysis corroborates the existence of a construction marked by a fronted topic with an obligatory standard of comparison and an optional comparee. Topicalization creates a cognitive domain that facilitates contrast between the standard of comparison and the comparee, with prosodic cues or an overt topic and/or comment marker. Furthermore, two formal subtypes of topicalization comparative construction are distinguished: a fuller form aligning with Heine’s prototype and a reduced variant observed in Wu and Hui Sinitic, wherein topic–comment ordering reinforced by prosody enables a contrastive reading between the standard and the comparee. Despite the centrality of topicalization, its cross-linguistic distribution remains limited; as a matter of fact, not all topic-prominent languages employ this schema, due to internal competition among cognitive strategies and external pressure from standard languages. These findings underscore the marginal status of topicalization comparatives in a global typological perspective. The study contributes empirical evidence from under-documented Sinitic varieties and points to the need for broader cross-linguistic investigation into the competition among alternative cognitive schemas in shaping comparative constructions in topic-prominent languages. Full article
Show Figures

Figure 1

24 pages, 457 KB  
Article
Collocational Patterns and Semantic Prosodies of “Saudization” in Modern Standard Arabic Discourse: A Corpus-Based Analysis
by Basant S. M. Moustafa
Languages 2026, 11(9), 188; https://doi.org/10.3390/languages11090188 - 15 Sep 2026
Viewed by 258
Abstract
Since its intensification in the early 2000s, Saudization, perceived as Saudi Arabia’s labor nationalization policy, has profoundly shaped the socio-economic and socio-political discourse in the Arab world. In spite of the extensive policy and economic research on Saudization, it remains linguistically underexplored, particularly [...] Read more.
Since its intensification in the early 2000s, Saudization, perceived as Saudi Arabia’s labor nationalization policy, has profoundly shaped the socio-economic and socio-political discourse in the Arab world. In spite of the extensive policy and economic research on Saudization, it remains linguistically underexplored, particularly through corpus-based analysis of Arabic language discourse. The current study, hence, investigates the collocational patterns of the term “Sa’wadah” (the transliterated Arabic language equivalent of “Saudization”) in a mega-corpus of Modern Standard Arabic (MSA) with the aim of sketching its semantic prosodies and working towards the ideological representations permeating through its collocates. To this end, using the methodological synergy of Critical Discourse Analysis and Corpus Linguistics, the Arabic corpus (arTenTen24) integrated within the online interface of Sketch Engine is scrutinized to extract a statistically significant collocational profile for Saudization and work out its semantic prosodies and their ideological representations. The findings show that the collocational profile of Saudization can be categorized into institutional, economic and evaluative topic domains. Second, the semantic prosodies associated with the term are mainly mixed, with a slight tendency towards negative and critical evaluation. Third, the ideological representations of Saudizations in MSA discourse are marked by ambivalence: the policy is constructed as a national duty whose implementation is a contested practice. Full article
22 pages, 6996 KB  
Article
Multimodal Dimensions of Hungarian Infant-Directed Communication During Storytelling
by Anna Kohári, Uwe D. Reichel, Sarolta Murányi, Rebeka Dóra Keserű, Katalin Pirsel and Katalin Mády
Languages 2026, 11(8), 163; https://doi.org/10.3390/languages11080163 - 4 Aug 2026
Viewed by 942
Abstract
Adults, when talking to infants, adjust their communication across several dimensions, including acoustic, linguistic, and gestural characteristics. However, our knowledge of how these dimensions covary in this register adjustment remains limited. Our objective was to analyze the linguistic elements, gestures and acoustic features [...] Read more.
Adults, when talking to infants, adjust their communication across several dimensions, including acoustic, linguistic, and gestural characteristics. However, our knowledge of how these dimensions covary in this register adjustment remains limited. Our objective was to analyze the linguistic elements, gestures and acoustic features (speech rate, vowel space, prosody) of storytelling directed to 6-month-old infants in a multidimensional way. The storytelling of 92 mothers to their own children was investigated and compared to their communication with other adults. We analyzed the acoustic features, gestures, and the linguistic construction of their storytelling. Our results have shown that pointing and showing gestures were more common in infant-directed storytelling, just as interactive expressions unrelated to the story were. The correlation tests across dimensions revealed that infant-directed communication is multidimensional. The cross-dimensional analyses demonstrated that infant-directed communication, compared to adult-directed communication, exhibited more systematic multidimensional coordination. Mothers who used more gestures and interactive elements in their infant-directed communication also tended to modify the acoustic characteristics of their speech to a greater extent relative to their adult-directed speech. Full article
Show Figures

Figure 1

21 pages, 2071 KB  
Review
Voice, Speech, and Large Language Models in Neurology: From Acoustic Biomarkers to Conversational AI
by Shahar Shelly
Computation 2026, 14(7), 160; https://doi.org/10.3390/computation14070160 - 16 Jul 2026
Viewed by 1230
Abstract
Background: Speech models (wav2vec 2.0, HuBERT, Whisper), large language models (GPT, LLaMA), and conversational AI have expanded computational speech analysis from handcrafted acoustic features to dialogue-based neurological assessment. How well these approaches address clinical practice has not been evaluated. Methods: We conducted a [...] Read more.
Background: Speech models (wav2vec 2.0, HuBERT, Whisper), large language models (GPT, LLaMA), and conversational AI have expanded computational speech analysis from handcrafted acoustic features to dialogue-based neurological assessment. How well these approaches address clinical practice has not been evaluated. Methods: We conducted a narrative review searching PubMed, Google Scholar, and IEEE Xplore, supplemented by Interspeech and ICASSP proceedings. Findings are organized along three layers: acoustic-motor (voice quality, prosody, articulation), language-transcript (lexical, syntactic, semantic, and discourse analysis), and integrated multimodal-conversational (interactive dialogue systems). Traditional acoustic biomarkers provide background; the primary focus is on foundation models, LLMs, and conversational AI. Findings: Speech foundation models outperform handcrafted features on several classification tasks but degrade on severely impaired speech due to domain mismatch with healthy training data. LLMs classify transcripts and score cognitive tests, but operate on text alone and cannot access acoustic-motor information. Conversational AI can administer cognitive screening through naturalistic dialogue, but validation is limited to small single-centre feasibility studies. Prospective clinical validation remains limited. Cross-linguistic generalizability is untested for most methods. Interpretation: The field is moving toward integrated speech-language assessment, but the gap between technical capability and clinical utility remains wide. Closing it requires diverse multilingual datasets, standardized benchmarks, prospective validation, and ethical governance. Full article
Show Figures

Figure 1

22 pages, 3077 KB  
Article
AI-Driven Detection of Neurodevelopmental Disorder from Emotional Speech Using a Hybrid CNN–BiLSTM–Attention Framework
by Nayarah Shabir, Parveen Kumar Lehana and Sheema Khan
Appl. Sci. 2026, 16(13), 6647; https://doi.org/10.3390/app16136647 - 3 Jul 2026
Viewed by 478
Abstract
Neurodevelopmental disorders (NDDs) are associated with impairments in communication, behavior, and social interaction, making accurate diagnosis clinically challenging. Autism Spectrum Disorder (ASD), a major NDD, often exhibits atypical speech patterns characterized by altered prosody and reduced emotional expressiveness. The study proposes a hybrid [...] Read more.
Neurodevelopmental disorders (NDDs) are associated with impairments in communication, behavior, and social interaction, making accurate diagnosis clinically challenging. Autism Spectrum Disorder (ASD), a major NDD, often exhibits atypical speech patterns characterized by altered prosody and reduced emotional expressiveness. The study proposes a hybrid dual-path framework for ASD detection from emotional speech using two strategies: PCA–GMM-based acoustic modeling and a CNN–BiLSTM–Attention architecture for spectral–temporal feature learning. The proposed framework captures probabilistic, spectral, and temporal speech characteristics for robust ASD classification. Acoustic analysis demonstrated clear separability between ASD and non-ASD speech, while the deep learning framework achieved stable and reliable performance across multiple emotional conditions. Experimental evaluation achieved 98.3% accuracy, AUC values ranging from 0.9699 to 0.9864, and F1-scores up to 0.9891. The findings highlight the potential of AI-driven speech analysis as a scalable and non-invasive tool for early ASD screening and predictive healthcare applications. Full article
Show Figures

Figure 1

24 pages, 4206 KB  
Article
Development and Preliminary Validation of the Turkish Prosodic Comprehension Test (PCT)
by Merve Savaş, Göknur Miray Ceyhan Tasin, Senanur Kahraman Beğen, Melis Buse Arslan, Ayşe Nur Koçak and Tutku Altıntaş
Audiol. Res. 2026, 16(4), 99; https://doi.org/10.3390/audiolres16040099 - 30 Jun 2026
Viewed by 433
Abstract
Background: Prosodic cues play a critical role in marking syntactic boundaries and guiding sentence interpretation. However, Turkish clinical language batteries lack dedicated measures targeting linguistic prosody and the syntax–prosody interface. Consequently, subtle auditory–prosodic comprehension difficulties may go undetected in stroke populations who perform [...] Read more.
Background: Prosodic cues play a critical role in marking syntactic boundaries and guiding sentence interpretation. However, Turkish clinical language batteries lack dedicated measures targeting linguistic prosody and the syntax–prosody interface. Consequently, subtle auditory–prosodic comprehension difficulties may go undetected in stroke populations who perform within normal limits on standard aphasia assessments. This study presents the development and initial psychometric evaluation of the Prosodic Comprehension Test (PCT), a Turkish sentence–picture matching tool designed to isolate prosodic contributions to meaning under controlled syntactic conditions. Methods: A total of 440 neurologically healthy native Turkish-speaking adults participated. An initial pool of 80 sentences (40 minimal pairs), identical in segmental and syntactic structure but differing in interpretation through prosodic boundary placement, was created. Audio stimuli were recorded by a professional actor, and corresponding visual stimuli represented alternative interpretations. Following expert review, 32 sentences (16 pairs) were retained and organized into three subcomponents: In situ Prosody, Focus–Topic Marking, and Pragmatic Disambiguation. Administration included fixed-intensity auditory presentation and a structured learning phase. Results: Internal consistency was acceptable (Cronbach’s α = 0.73). Principal component analysis was consistent with the theoretically proposed three-component structure (KMO = 0.74; Bartlett’s test significant, p < 0.001), with the three components collectively accounting for 28.2% of the total variance. Convergent validity was supported by a significant positive correlation with MoCA-TR (r = 0.23, p < 0.001, 95% CI [0.14, 0.32]). Conclusions: The PCT appears to be a linguistically grounded and psychometrically promising tool for assessing prosodic comprehension in Turkish. The present findings are based on a healthy adult sample and should be interpreted as preliminary normative evidence. Further research should address test–retest reliability, confirmatory factor analyses, and validation in clinical populations. Full article
(This article belongs to the Section Speech and Language)
Show Figures

Figure 1

30 pages, 1349 KB  
Article
A Lightweight Multimodal Architecture for Punctuation Restoration in Kazakh ASR
by Aidana Karibayeva, Oleg Myssov, Balzhan Abduali, Dina Amirova and Adina Karybayeva
Computers 2026, 15(6), 345; https://doi.org/10.3390/computers15060345 - 28 May 2026
Viewed by 1299
Abstract
In this paper, we first present a multimodal architecture called CrossAttn-v1. This model is designed to recover punctuation marks in Kazakh and combines contextual XLM-RoBERTa-large text embeddings with the Whisper large-v3 encoder states via a cross-attention mechanism. In addition, a 4-dimensional prosodic vector [...] Read more.
In this paper, we first present a multimodal architecture called CrossAttn-v1. This model is designed to recover punctuation marks in Kazakh and combines contextual XLM-RoBERTa-large text embeddings with the Whisper large-v3 encoder states via a cross-attention mechanism. In addition, a 4-dimensional prosodic vector and a CRF output layer are used. The model was trained using an adapted Whisper ASR model on 33,332 utterances from the KazakhTTS2 corpus. After adaptation, the word error rate decreased from 45.7% to 4.25%. On the in-domain test set (56,396 tokens), CrossAttn-v1 achieved F1-macro = 0.8485 for recovering five-class punctuation marks. Furthermore, CrossAttn-v1 outperformed the GPT-4o zero-shot model by +0.294 F1 and the M3 Hybrid model based on prosody alone by +0.070 F1. The class analysis showed that the Whisper encoder states were particularly useful for prosody-dependent punctuation. For example, it outperformed M3 Hybrid by +9.5 percentage points on the QUESTION mark and by +20.2 percentage points on the EXCLAIM mark. On 883 out-of-domain natural speech recordings, the model performed similarly to the text-only baseline model (Δ = −0.041, not significant), suggesting that domain mismatch in the Whisper training corpus was a major factor limiting generalization. Full article
(This article belongs to the Special Issue Advances in Multimodal Learning and Representation)
Show Figures

Figure 1

24 pages, 1038 KB  
Article
Avant-Garde Poetry and the Tékhnē of Traditional Versification
by Evgenii Kazartsev and Nikita Kirichenko
Arts 2026, 15(5), 97; https://doi.org/10.3390/arts15050097 - 2 May 2026
Cited by 1 | Viewed by 685
Abstract
This article offers a theoretically nuanced and empirically grounded investigation into the paradoxical afterlife of classical versification within the poetic practices of the Russian and Soviet avant-garde. Challenging the persistent historiographic narrative that equates avant-garde poetics with an unequivocal rupture from tradition, the [...] Read more.
This article offers a theoretically nuanced and empirically grounded investigation into the paradoxical afterlife of classical versification within the poetic practices of the Russian and Soviet avant-garde. Challenging the persistent historiographic narrative that equates avant-garde poetics with an unequivocal rupture from tradition, the study demonstrates that canonical metrical forms—most notably iambic tetrameter—continued to operate as structurally productive, albeit critically reconfigured, elements within experimental verse. Drawing on a broad corpus encompassing poetic manifestos, verse texts, and prose writings by Vladimir Maiakovskii, Ilia Sel’vinskii, Semen Kirsanov, and Nikolai Aseev, the authors combine close formal analysis with quantitative prosodic modeling, including linguistic and speech models derived from Kolmogorov–Taranovsky verse theory. The article argues that avant-garde poets did not simply negate inherited metrics but subjected them to a process of internal recomposition, shifting attention from meter as a fixed scheme to rhythm as a dynamic, semantically charged construct. While rhythmic innovation is shown to be consciously engineered in verse, the analysis of verse-like fragments in prose reveals persistent, unconscious attachments to “classical” rhythmic patterns, particularly the Pushkinian alternating rhythm. This tension between declarative rejection and latent continuity illuminates the avant-garde’s distinctive mode of negotiating tradition: not abolishing it, but instrumentalizing it within a broader project of total artistic reorganization. The study thus reframes avant-garde prosody as a site where innovation and inheritance coexist in a state of productive contradiction, reshaping our understanding of modernist poetic technique. Full article
Show Figures

Figure 1

18 pages, 676 KB  
Article
Iambic Production Advantage and Unbiased Recognition in Word Learning by Mandarin-Speaking Children with Cochlear Implants
by Xinyuan Shi, Jingjing Yang and Dandan Liang
Behav. Sci. 2026, 16(4), 491; https://doi.org/10.3390/bs16040491 - 26 Mar 2026
Viewed by 540
Abstract
This study examines whether Mandarin-speaking children with cochlear implants (CIs) exhibit challenges or advantages in learning novel words with overt trochaic versus iambic patterns. Mandarin full-tone words lack salient stress cues, whereas neutral-tone words exhibit a clear trochaic pattern. Given the unique prosody [...] Read more.
This study examines whether Mandarin-speaking children with cochlear implants (CIs) exhibit challenges or advantages in learning novel words with overt trochaic versus iambic patterns. Mandarin full-tone words lack salient stress cues, whereas neutral-tone words exhibit a clear trochaic pattern. Given the unique prosody of Mandarin and CI users’ difficulty in acquiring the neutral tone, we predicted distinct effects of lexical stress on word production and recognition. Fifteen Mandarin-speaking preschoolers with CIs and 15 age-matched children with normal hearing (NH) learned 16 pairs of stress-contrasted novel words. A referent-naming task assessed stress production through pattern proportion, accuracy, and acoustic analysis, while a referent-matching task evaluated stress identification and word-referent mapping accuracy. In the naming task, children with CIs showed a preference for iambic words in both frequency and accuracy. They also produced longer second syllables in trochaic words than their NH peers. In the matching task, the CI group performed worse overall, although neither group showed a stress-specific effect. These results indicate that CI users struggle with syllable duration control in trochees. This difficulty reflects both an inability to shorten the unstressed syllable and the potential adoption of a final-syllable lengthening strategy linked to higher prosodic domains. The insensitivity to stress contrast in recognition may stem from the generally weak word-level stress cues in Mandarin. Full article
(This article belongs to the Section Developmental Psychology)
Show Figures

Figure 1

38 pages, 6181 KB  
Article
An AIoT-Based Framework for Automated English-Speaking Assessment: Architecture, Benchmarking, and Reliability Analysis of Open-Source ASR
by Paniti Netinant, Rerkchai Fooprateepsiri, Ajjima Rukhiran and Meennapa Rukhiran
Informatics 2026, 13(2), 19; https://doi.org/10.3390/informatics13020019 - 26 Jan 2026
Viewed by 3065
Abstract
The emergence of low-cost edge devices has enabled the integration of automatic speech recognition (ASR) into IoT environments, creating new opportunities for real-time language assessment. However, achieving reliable performance on resource-constrained hardware remains a significant challenge, especially on the Artificial Internet of Things [...] Read more.
The emergence of low-cost edge devices has enabled the integration of automatic speech recognition (ASR) into IoT environments, creating new opportunities for real-time language assessment. However, achieving reliable performance on resource-constrained hardware remains a significant challenge, especially on the Artificial Internet of Things (AIoT). This study presents an AIoT-based framework for automated English-speaking assessment that integrates architecture and system design, ASR benchmarking, and reliability analysis on edge devices. The proposed AIoT-oriented architecture incorporates a lightweight scoring framework capable of analyzing pronunciation, fluency, prosody, and CEFR-aligned speaking proficiency within an automated assessment system. Seven open-source ASR models—four Whisper variants (tiny, base, small, and medium) and three Vosk models—were systematically benchmarked in terms of recognition accuracy, inference latency, and computational efficiency. Experimental results indicate that Whisper-medium deployed on the Raspberry Pi 5 achieved the strongest overall performance, reducing inference latency by 42–48% compared with the Raspberry Pi 4 and attaining the lowest Word Error Rate (WER) of 6.8%. In contrast, smaller models such as Whisper-tiny, with a WER of 26.7%, exhibited two- to threefold higher scoring variability, demonstrating how recognition errors propagate into automated assessment reliability. System-level testing revealed that the Raspberry Pi 5 can sustain near real-time processing with approximately 58% CPU utilization and around 1.2 GB of memory, whereas the Raspberry Pi 4 frequently approaches practical operational limits under comparable workloads. Validation using real learner speech data (approximately 100 sessions) confirmed that the proposed system delivers accurate, portable, and privacy-preserving speaking assessment using low-power edge hardware. Overall, this work introduces a practical AIoT-based assessment framework, provides a comprehensive benchmark of open-source ASR models on edge platforms, and offers empirical insights into the trade-offs among recognition accuracy, inference latency, and scoring stability in edge-based ASR deployments. Full article
Show Figures

Figure 1

33 pages, 3147 KB  
Review
Perception–Production of Second-Language Mandarin Tones Based on Interpretable Computational Methods: A Review
by Yujiao Huang, Zhaohong Xu, Xianming Bei and Huakun Huang
Mathematics 2026, 14(1), 145; https://doi.org/10.3390/math14010145 - 30 Dec 2025
Cited by 1 | Viewed by 2936
Abstract
We survey recent advances in second-language (L2) Mandarin lexical tones research and show how an interpretable computational approach can deliver parameter-aligned feedback across perception–production (P ↔ P). We synthesize four strands: (A) conventional evaluations and tasks (identification, same–different, imitation/read-aloud) that reveal robust tone-pair [...] Read more.
We survey recent advances in second-language (L2) Mandarin lexical tones research and show how an interpretable computational approach can deliver parameter-aligned feedback across perception–production (P ↔ P). We synthesize four strands: (A) conventional evaluations and tasks (identification, same–different, imitation/read-aloud) that reveal robust tone-pair asymmetries and early P ↔ P decoupling; (B) physiological and behavioral instrumentation (e.g., EEG, eye-tracking) that clarifies cue weighting and time course; (C) audio-only speech analysis, from classic F0 tracking and MFCC–prosody fusion to CNN/RNN/CTC and self-supervised pipelines; and (D) interpretable learning, including attention and relational models (e.g., graph neural networks, GNNs) opened with explainable AI (XAI). Across strands, evidence converges on tones as time-evolving F0 trajectories, so movement, turning-point timing, and local F0 range are more diagnostic than height alone, and the contrast between Tone 2 (rising) and Tone 3 (dipping/low) remains the persistent difficulty; learners with tonal vs. non-tonal language backgrounds weight these cues differently. Guided by this synthesis, we outline a tool-oriented framework that pairs perception and production on the same items, jointly predicts tone labels and parameter targets, and uses XAI to generate local attributions and counterfactual edits, making feedback classroom-ready. Full article
(This article belongs to the Section E1: Mathematics and Computer Science)
Show Figures

Figure 1

14 pages, 2974 KB  
Data Descriptor
Articulatory Data on Preboundary Lengthening Across Prominence Conditions in American English
by Jiyoung Jang, Sahyang Kim and Taehong Cho
Data 2025, 10(12), 197; https://doi.org/10.3390/data10120197 - 1 Dec 2025
Viewed by 691
Abstract
This article presents articulatory–kinematic data on preboundary lengthening (Intonational Phrase-final lengthening) from the productions of ten native speakers of American English—a relatively rare class of phonetic data compared with the more widely available acoustic data. The dataset includes three trisyllabic nonce words (bábaba, [...] Read more.
This article presents articulatory–kinematic data on preboundary lengthening (Intonational Phrase-final lengthening) from the productions of ten native speakers of American English—a relatively rare class of phonetic data compared with the more widely available acoustic data. The dataset includes three trisyllabic nonce words (bábaba, babába, bababá), each designed to manipulate the location of lexical stress. These were produced under prosodic conditions that varied in boundary position and focus-induced phrasal prominence, enabling analysis of how preboundary lengthening is distributed across words with different lexical stress locations and how it interacts with prosodic prominence. Articulatory data were collected using electromagnetic articulography (EMA, Carstens AG200), providing kinematic measurements such as movement duration, peak velocity, and displacement of articulatory gestures. The accompanying files allow examination of individual speaker variation in these measures as modulated by prosodic structure, including boundary and prominence effects. While theoretical findings have been reported in a previous study, the full dataset, including detailed descriptions of individual speaker patterns, is made available here. By making these less commonly available articulatory data publicly available, we aim to promote broad reuse and support further research in prosody, articulatory phonetics, and speech production. Full article
Show Figures

Figure 1

26 pages, 4013 KB  
Article
Music Genre Classification Using Prosodic, Stylistic, Syntactic and Sentiment-Based Features
by Erik-Robert Kovacs and Stefan Baghiu
Big Data Cogn. Comput. 2025, 9(11), 296; https://doi.org/10.3390/bdcc9110296 - 19 Nov 2025
Viewed by 4704
Abstract
Romanian popular music has had a storied history across the last century and a half. Incorporating different influences at different times, today it boasts a wide range of both autochthonous and imported genres, such as traditional folk music, rock, rap, pop, and manele, [...] Read more.
Romanian popular music has had a storied history across the last century and a half. Incorporating different influences at different times, today it boasts a wide range of both autochthonous and imported genres, such as traditional folk music, rock, rap, pop, and manele, to name a few. We aim to trace the linguistic differences between the lyrics of these genres using natural language processing and a computational linguistics approach by studying the prosodic, stylistic, syntactic, and sentiment-based features of each genre. For this purpose, we have crawled a dataset of ~14,000 Romanian songs from publicly available websites along with the user-provided genre labels, and characterized each song and each genre, respectively, with regard to these features, discussing similarities and differences. We improve on existing tools for Romanian language natural language processing by building a lexical analysis library well suited to song lyrics or poetry which encodes a set of 17 linguistic features. In addition, we build lexical analysis tools for profanity-based features and improve the SentiLex sentiment analysis library by manually rebalancing its lexemes to overcome the limitations introduced by it having been machine translated into Romanian. We estimate the accuracy gain using a benchmark Romanian sentiment analysis dataset and register a 25% increase in accuracy over the SentiLex baseline. The contribution is meant to describe the characteristics of the Romanian expression of autochthonous as well as international genres and provide technical support to researchers in natural language processing, musicology or the digital humanities in studying the lyrical content of Romanian music. We have released our data and code for research use. Full article
(This article belongs to the Special Issue Artificial Intelligence (AI) and Natural Language Processing (NLP))
Show Figures

Figure 1

11 pages, 703 KB  
Article
Distinguishing Between Healthy and Unhealthy Newborns Based on Acoustic Features and Deep Learning Neural Networks Tuned by Bayesian Optimization and Random Search Algorithm
by Salim Lahmiri, Chakib Tadj and Christian Gargour
Entropy 2025, 27(11), 1109; https://doi.org/10.3390/e27111109 - 27 Oct 2025
Cited by 1 | Viewed by 821
Abstract
Voice analysis and classification for biomedical diagnosis purpose is receiving a growing attention to assist physicians in the decision-making process in clinical milieu. In this study, we develop and test deep feedforward neural networks (DFFNN) to distinguish between healthy and unhealthy newborns. The [...] Read more.
Voice analysis and classification for biomedical diagnosis purpose is receiving a growing attention to assist physicians in the decision-making process in clinical milieu. In this study, we develop and test deep feedforward neural networks (DFFNN) to distinguish between healthy and unhealthy newborns. The DFFNN are trained with acoustic features measured from newborn cries, including auditory-inspired amplitude modulation (AAM), Mel Frequency Cepstral Coefficients (MFCC), and prosody. The configuration of the DFFNN is optimized by using Bayesian optimization (BO) and random search (RS) algorithm. Under both optimization techniques, the experimental results show that the DFFNN yielded to the highest classification rate when trained with all acoustic features. Specifically, the DFFNN-BO and DFFNN-RS achieved 87.80% ± 0.23 and 86.12% ± 0.33 accuracy, respectively, under ten-fold cross-validation protocol. Both DFFNN-BO and DFFNN-RS outperformed existing approaches tested on the same database. Full article
(This article belongs to the Section Signal and Data Analysis)
Show Figures

Figure 1

19 pages, 7222 KB  
Article
Multi-Channel Spectro-Temporal Representations for Speech-Based Parkinson’s Disease Detection
by Hadi Sedigh Malekroodi, Nuwan Madusanka, Byeong-il Lee and Myunggi Yi
J. Imaging 2025, 11(10), 341; https://doi.org/10.3390/jimaging11100341 - 1 Oct 2025
Cited by 1 | Viewed by 2082
Abstract
Early, non-invasive detection of Parkinson’s Disease (PD) using speech analysis offers promise for scalable screening. In this work, we propose a multi-channel spectro-temporal deep-learning approach for PD detection from sentence-level speech, a clinically relevant yet underexplored modality. We extract and fuse three complementary [...] Read more.
Early, non-invasive detection of Parkinson’s Disease (PD) using speech analysis offers promise for scalable screening. In this work, we propose a multi-channel spectro-temporal deep-learning approach for PD detection from sentence-level speech, a clinically relevant yet underexplored modality. We extract and fuse three complementary time–frequency representations—mel spectrogram, constant-Q transform (CQT), and gammatone spectrogram—into a three-channel input analogous to an RGB image. This fused representation is evaluated across CNNs (ResNet, DenseNet, and EfficientNet) and Vision Transformer using the PC-GITA dataset, under 10-fold subject-independent cross-validation for robust assessment. Results showed that fusion consistently improves performance over single representations across architectures. EfficientNet-B2 achieves the highest accuracy (84.39% ± 5.19%) and F1-score (84.35% ± 5.52%), outperforming recent methods using handcrafted features or pretrained models (e.g., Wav2Vec2.0, HuBERT) on the same task and dataset. Performance varies with sentence type, with emotionally salient and prosodically emphasized utterances yielding higher AUC, suggesting that richer prosody enhances discriminability. Our findings indicate that multi-channel fusion enhances sensitivity to subtle speech impairments in PD by integrating complementary spectral information. Our approach implies that multi-channel fusion could enhance the detection of discriminative acoustic biomarkers, potentially offering a more robust and effective framework for speech-based PD screening, though further validation is needed before clinical application. Full article
(This article belongs to the Special Issue Celebrating the 10th Anniversary of the Journal of Imaging)
Show Figures

Figure 1

Back to TopTop