Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (33)

Search Parameters:
Keywords = speech envelope

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
17 pages, 1536 KB  
Article
The Temporal Resolution Needed for Speech Intelligibility Assessed with Mosaic Speech: Effects of Block Duration, Age, and Word Familiarity
by Gerard B. Remijn, Yuna Uzuhashi, Emi Hasuo, Kazuo Ueda and Yoshitaka Nakajima
Audiol. Res. 2026, 16(5), 122; https://doi.org/10.3390/audiolres16050122 - 24 Aug 2026
Viewed by 154
Abstract
Background/Objectives: Mosaic speech was used to further investigate the auditory system’s temporal resolution needed for speech intelligibility. Mosaic speech is a form of degraded speech segmented in frequency × time blocks with no discernible temporal fine structure and with degraded amplitude envelope [...] Read more.
Background/Objectives: Mosaic speech was used to further investigate the auditory system’s temporal resolution needed for speech intelligibility. Mosaic speech is a form of degraded speech segmented in frequency × time blocks with no discernible temporal fine structure and with degraded amplitude envelope cues. Methods: We performed a listening experiment with mosaic speech consisting of 20 frequency bands segmented into 20, 40, 80, 160, and 320 ms. Younger listeners (<25 years; n = 20), with self-reported normal hearing and having passed a limited screening test, and elderly listeners (>65 years; n = 19) with hearing thresholds ranging from normal hearing to moderate hearing loss listened to Japanese low- and high-familiarity mosaic words and wrote down what they heard. The original words were included as control stimuli. Results: Although younger listeners had significantly higher intelligibility scores, elderly listeners could integrate and parse coarse blocks of mosaic speech of 20 ms and 40 ms with intelligibility scores of 84% or higher for high-familiarity words. For longer block durations, however, the elderly listeners’ intelligibility dropped rapidly for both high- and low-familiarity words. In both the young and the elderly listener groups, intelligibility reached the floor for block durations of 160 and 320 ms. Word familiarity strongly affected intelligibility scores. For blocks up to 80 ms, intelligibility was significantly higher for high-familiarity words than for low-familiarity words in both age groups, with a 10–30% difference. Conclusions: Elderly listeners (n = 19) with normal hearing to moderate hearing loss could maintain relatively high intelligibility for mosaic words segmented in blocks of 20 or 40 ms, provided the words had a high familiarity level. Full article
(This article belongs to the Section Speech and Language)
Show Figures

Figure 1

20 pages, 985 KB  
Article
Spanish + Portuguese = Mixing(?): Forcing the Issue
by John Lipski
Languages 2026, 11(7), 143; https://doi.org/10.3390/languages11070143 - 3 Jul 2026
Cited by 1 | Viewed by 484
Abstract
In sustained Spanish–Portuguese contact zones, there is often a mismatch between observable linguistic data and speakers’ views on the degree of resulting hybridity. Terms like “Portuñol/Portunhol” are suggestive of a “third” language, but little empirical evidence has been offered in support of such [...] Read more.
In sustained Spanish–Portuguese contact zones, there is often a mismatch between observable linguistic data and speakers’ views on the degree of resulting hybridity. Terms like “Portuñol/Portunhol” are suggestive of a “third” language, but little empirical evidence has been offered in support of such outcomes, and spontaneous speech does not provide the complete envelope of variation. The present study draws on data collected in Misiones province in northeastern Argentina, where vernacular Portuguese is spoken natively in rural households, and Spanish is first acquired informally. To empirically test popular notions of “Portuñol” hybridity, an array of experimental tasks has been employed to present various monolingual and mixed configurations to bilingual participants. In the sociolinguistic free-fall environment of Misiones, the potential for the emergence of a stable hybrid is at its greatest, but the increasingly fine-grained array of experimental techniques designed to force-feed putative “Portuñol” shows that even with priming, lexical borrowing is still the only consistent contact-induced phenomenon. Full article
(This article belongs to the Special Issue Shifting Borders: Spanish Morphosyntax in Contact Zones)
Show Figures

Figure 1

42 pages, 6690 KB  
Article
MS-SENet: A Multi-Scale Squeeze–Excitation Network for Deep-Learning-Based Automatic Modulation Classification in Cognitive Radio Systems
by Evelio Astaiza Hoyos, Héctor Fabio Bermúdez-Orozco and Nasly Cristina Rodriguez-Idrobo
Future Internet 2026, 18(7), 343; https://doi.org/10.3390/fi18070343 - 29 Jun 2026
Viewed by 319
Abstract
Automatic modulation classification (AMC) is a critical enabler of cognitive radio (CR) systems, allowing secondary users to identify primary user modulation schemes and adapt transmission parameters in real time. Traditional AMC approaches, based on likelihood functions or hand-crafted features, suffer from degraded performance [...] Read more.
Automatic modulation classification (AMC) is a critical enabler of cognitive radio (CR) systems, allowing secondary users to identify primary user modulation schemes and adapt transmission parameters in real time. Traditional AMC approaches, based on likelihood functions or hand-crafted features, suffer from degraded performance under low signal-to-noise ratio (SNR) conditions and realistic channel impairments. In this paper, we propose MS-SENet (Multi-Scale Squeeze–Excitation Network), a novel deep-learning architecture that integrates multi-scale convolutional feature extraction, squeeze-and-excitation channel attention, residual learning, bidirectional long short-term memory (BiLSTM) temporal modelling, and global attention pooling into a unified framework for robust AMC. The multi-scale convolution module employs parallel branches with kernel sizes of 3, 5, and 7 to capture both fine-grained phase transitions and coarse envelope patterns from raw in-phase/quadrature (I/Q) signal samples. Squeeze–excitation residual blocks perform channel-wise feature recalibration, enabling the network to emphasize informative feature maps while suppressing less relevant ones. A bidirectional LSTM layer models temporal dependencies across the signal sequence, and a global attention pooling mechanism performs weighted temporal aggregation prior to classification. We present a comprehensive taxonomy of deep-learning architectures for AMC organised along five axes—input representation, feature extraction, temporal modelling, regularization strategy, and architectural complexity—and conduct a rigorous comparative evaluation against ten baseline architectures on a RadioML-style synthetic dataset (110,000 samples, 11 modulation classes, and 20 SNR levels from −20 to +18 dB). The experimental results demonstrate that MS-SENet achieves a mean classification accuracy of 87.9% at SNR ≥ 0 dB (the average of the medium and high SNR regime averages: 86.06% for 0 ≤ SNR < 10 dB and 89.68% for SNR ≥ 10 dB) while maintaining a compact footprint of approximately 406 K parameters, making it suitable for deployment on resource-constrained edge devices. We further analyze the robustness of the proposed architecture to multipath fading, carrier frequency offset, and sample rate offset, confirming its resilience under practical operating conditions. MS-SENet is an architecture designed for automatic modulation classification of I/Q signals and is not related to the homonymous architecture for speech emotion recognition. Full article
Show Figures

Figure 1

28 pages, 8935 KB  
Article
Wind-Sound Synergy and Fractal Design: Intelligent, Adaptive Acoustic Façades for High-Performance, Climate-Responsive Buildings
by Lingge Tan, Xinyue Zhang, Donghui Cui and Stephen Jia Wang
Buildings 2026, 16(8), 1615; https://doi.org/10.3390/buildings16081615 - 20 Apr 2026
Viewed by 608
Abstract
The building façade serves as the primary interface between the built environment and external climate, marking the transition from static regulation to dynamic response in climate-adaptive design. While existing research predominantly addresses periodic climatic elements such as temperature and solar radiation, the highly [...] Read more.
The building façade serves as the primary interface between the built environment and external climate, marking the transition from static regulation to dynamic response in climate-adaptive design. While existing research predominantly addresses periodic climatic elements such as temperature and solar radiation, the highly stochastic wind environment and its potential for internal acoustic problems remain systematically unexplored. This study investigates the acoustic modulation mechanism of building façades under dynamic wind conditions through a simulation-based methodology. The primary aim is to demonstrate the use of active control to mitigate the influence of fluctuating wind on the internal acoustic environment of buildings with open windows or semi-open boundaries, focusing on the coupling between stochastic wind fields and architectural acoustics in humid subtropical climates. We propose a wind-responsive adaptive acoustic façade system employing fractal geometry and configurable delay strategies, and develop a high-fidelity simulation framework to quantify how façade geometry and activation logic regulate acoustic parameters under varying wind conditions (1–8 m/s). Results indicate that: (1) support vector regression-based mapping of wind speed to delay strategies maintains key sound-field parameters (Lateral Fraction (LF), Speech Clarity (C50), and Early Decay Time to Reverberation Time ratio (EDT/RT30)) within 10% fluctuation across wind regimes; (2) fractal configurations achieve balanced wide-band (125 Hz–8 kHz) performance, with SPL fluctuation <3 dB, spectral tilt (+0.3 dB), and reverberation time slope <0.3; (3) configurational switching between column (high LF) and row (high C50) arrangements enables dynamic trade-off between spatial impression and speech clarity. This work establishes an integrated framework coupling wind dynamics, façade morphology, and acoustic modulation to regulate objective indoor acoustic parameters. Based on the simulated omnidirectional point-source model, the results show that key acoustic indicators remain stable across varying wind conditions, providing a theoretical and quantifiable basis for climate-responsive acoustic envelope design. Future work will include empirical prototype testing and listening tests to determine whether these simulated acoustic parameters translate into improved comfort and well-being for occupants. Full article
(This article belongs to the Special Issue Advanced Research on Improvement of the Indoor Acoustic Environment)
Show Figures

Figure 1

21 pages, 518 KB  
Communication
Ordering and Quantifying Textual Cohesion via Semantic, Geometric and Statistical Structure
by Stelios Arvanitis
Stats 2026, 9(2), 25; https://doi.org/10.3390/stats9020025 - 3 Mar 2026
Viewed by 774
Abstract
We propose a semantic, geometric, and statistical framework for quantifying and ordering textual cohesion in long-form discourse. Sentences are embedded into a semantic similarity graph and Ollivier–Ricci curvature is used to extract sentence- and document-level structural profiles, represented as step functions on a [...] Read more.
We propose a semantic, geometric, and statistical framework for quantifying and ordering textual cohesion in long-form discourse. Sentences are embedded into a semantic similarity graph and Ollivier–Ricci curvature is used to extract sentence- and document-level structural profiles, represented as step functions on a normalized rhetorical-time axis. On this functional space we define the Weighted Utopia Index (wUI), a corpus-relative measure of weighted shortfall from an upper-envelope profile under a dominance-type ordering. The rhetorical-time weighting function is learned self-supervised: we generate controlled sentence-order perturbations with known ordinal coherence degradation and estimate the weight parameters via an ordered probit model on a training split. We evaluate ordering recovery on held-out State of the Union speeches using rank correlations, pairwise and adjacent ordering accuracy, and violation-localization diagnostics with bootstrap uncertainty. Across these criteria, wUI systematically outperforms embedding-only adjacent-similarity baselines, while a Nash-type aggregation provides an interpretable semantic–structural trade-off score. An application to later-period speeches illustrates how the method yields interpretable cohesion rankings and curvature-profile diagnostics without requiring external annotations. Full article
(This article belongs to the Section Applied Statistics and Machine Learning Methods)
Show Figures

Figure 1

14 pages, 1658 KB  
Article
The Effect of Modulation Enhancement Scheme on Speech Recognition in Spatial Noise Among Young Adults with Normal Hearing
by Vibha Kanagokar, M. A. Yashu, Jayashree S. Bhat and Arivudai Nambi Pitchaimuthu
Audiol. Res. 2026, 16(1), 26; https://doi.org/10.3390/audiolres16010026 - 14 Feb 2026
Viewed by 778
Abstract
Background/Objectives: Speech understanding in noise relies on both temporal fine structure (TFS) and temporal envelope (ENV) cues. While TFS primarily conveys interaural time differences (ITDs) at low frequencies, ENV cues can also support ITD processing, especially when TFS is unavailable or degraded. [...] Read more.
Background/Objectives: Speech understanding in noise relies on both temporal fine structure (TFS) and temporal envelope (ENV) cues. While TFS primarily conveys interaural time differences (ITDs) at low frequencies, ENV cues can also support ITD processing, especially when TFS is unavailable or degraded. Expanding the ENV by increasing modulation depth has been proposed to improve speech perception, but its effects on spatial release from masking (SRM) and binaural temporal processing in normal-hearing listeners remain unclear. The goal of this study was to evaluate the effect of ENV enhancement on SRM in young adults with normal hearing and its influence on ITD sensitivity and interaural coherence (IC). Method: Thirty normal-hearing native Kannada speakers (19–34 years) participated. Speech stimuli consisted of Kannada sentences embedded in four-talker babble at −5, 0, and +5 dB signal to noise ratio (SNR). Target and masker were spatialized using head-related transfer functions at 0°, 15°, and 37.5° azimuths. Stimuli were presented with and without ENV enhancement (compression–expansion algorithm). Speech recognition scores were analyzed using generalized linear mixed models, and SRM was calculated as performance differences between co-located and spatially separated conditions. Cross-correlation analyses were performed to estimate ITDs and IC across SNRs. Result: ENV enhancement yielded significantly higher SRM values across all SNRs and spatial separations. Benefits were greatest at lower SNRs and wider target–masker separations. Cross-correlation analysis showed enhanced IC and more reliable ITD estimates under the expanded condition, particularly at moderate SNRs. Conclusions: Temporal ENV enhancement strengthens spatial unmasking and binaural timing cues in normal-hearing adults, especially under adverse listening conditions. These findings highlight its potential application in auditory rehabilitation and hearing technologies where ENV cues are critical. Full article
(This article belongs to the Section Hearing)
Show Figures

Figure 1

29 pages, 4560 KB  
Article
Graph Fractional Hilbert Transform: Theory and Application
by Daxiang Li and Zhichao Zhang
Fractal Fract. 2026, 10(2), 74; https://doi.org/10.3390/fractalfract10020074 - 23 Jan 2026
Cited by 2 | Viewed by 713
Abstract
The graph Hilbert transform (GHT) is a key tool in constructing analytic signals and extracting envelope and phase information in graph signal processing. However, its utility is limited by confinement to the graph Fourier domain, a fixed phase shift, information loss for real-valued [...] Read more.
The graph Hilbert transform (GHT) is a key tool in constructing analytic signals and extracting envelope and phase information in graph signal processing. However, its utility is limited by confinement to the graph Fourier domain, a fixed phase shift, information loss for real-valued spectral components, and the absence of tunable parameters. The graph fractional Fourier transform introduces domain flexibility through a fractional order parameter α but does not resolve the issues of phase rigidity and information loss. Inspired by the dual-parameter fractional Hilbert transform (FRHT) in classical signal processing, we propose the graph FRHT (GFRHT). The GFRHT incorporates a dual-parameter framework: the fractional order α enables analysis across arbitrary fractional domains, interpolating between vertex and spectral spaces, while the angle parameter β provides adjustable phase shifts and a non-zero real-valued response (cosβ) for real eigenvalues, thereby eliminating information loss. We formally define the GFRHT, establish its core properties, and design a method for graph analytic signal construction, enabling precise envelope extraction and demodulation. Experiments on anomaly identification, speech classification and edge detection demonstrate that GFRHT outperforms GHT, offering greater flexibility and superior performance in graph signal processing. Full article
Show Figures

Figure 1

27 pages, 1695 KB  
Review
Overcoming the Challenge of Singing Among Cochlear Implant Users: An Analysis of the Disrupted Feedback Loop and Strategies for Improvement
by Stephanie M. Younan, Emmeline Y. Lin, Brooke Barry, Arjun Kurup, Karen C. Barrett and Nicole T. Jiam
Brain Sci. 2025, 15(11), 1192; https://doi.org/10.3390/brainsci15111192 - 4 Nov 2025
Cited by 1 | Viewed by 2717
Abstract
Background: Cochlear implants (CIs) are transformative neuroprosthetics that restore speech perception for individuals with severe-to-profound hearing loss. However, temporal envelope cues are well-represented within the signal processing, while spectral envelope cues are poorly accessed by CI users, resulting in substantial deficits compared to [...] Read more.
Background: Cochlear implants (CIs) are transformative neuroprosthetics that restore speech perception for individuals with severe-to-profound hearing loss. However, temporal envelope cues are well-represented within the signal processing, while spectral envelope cues are poorly accessed by CI users, resulting in substantial deficits compared to normal-hearing individuals. This profoundly impairs the perception of complex auditory stimuli like music and vocal prosody, significantly impacting users’ quality of life, social engagement, and artistic expression. Methods: This narrative review synthesizes research on CI signal-processing limitations, perceptual and production challenges in music and singing, the role of the auditory–motor feedback loop, and strategies for improvement, including rehabilitation, technology, and the influence of neuroplasticity and sensitive developmental periods. Results: The degraded signal causes marked deficits in pitch, timbre, and vocal emotion perception. Critically, this impoverished input functionally breaks the high-fidelity auditory–motor feedback loop essential for vocal control, transforming it from a precise fine-tuner into a gross error detector sensitive only to massive pitch shifts (~6 semitones). This neurophysiological breakdown directly causes pervasive pitch inaccuracies and melodic distortion in singing. Despite these challenges, improvements are possible through advanced sound-processing strategies, targeted auditory–motor training that leverages neuroplasticity, and capitalizing on sensitive periods for auditory development. Conclusions: The standard CI signal creates a fundamental neurophysiological barrier to singing. Overcoming this requires a paradigm shift toward holistic, patient-centered care that moves beyond speech-centric goals. Integrating personalized, music-based rehabilitation with advanced CI programming is essential for improving vocal production, fostering musical engagement, and ultimately enhancing the overall quality of life for CI users. Full article
(This article belongs to the Special Issue Language, Communication and the Brain—2nd Edition)
Show Figures

Figure 1

18 pages, 3164 KB  
Article
Cough Detection Using Acceleration Signals and Deep Learning Techniques
by Daniel Sanchez-Morillo, Diego Sales-Lerida, Blanca Priego-Torres and Antonio León-Jiménez
Electronics 2024, 13(12), 2410; https://doi.org/10.3390/electronics13122410 - 20 Jun 2024
Cited by 9 | Viewed by 5304
Abstract
Cough is a frequent symptom in many common respiratory diseases and is considered a predictor of early exacerbation or even disease progression. Continuous cough monitoring offers valuable insights into treatment effectiveness, aiding healthcare providers in timely intervention to prevent exacerbations and hospitalizations. Objective [...] Read more.
Cough is a frequent symptom in many common respiratory diseases and is considered a predictor of early exacerbation or even disease progression. Continuous cough monitoring offers valuable insights into treatment effectiveness, aiding healthcare providers in timely intervention to prevent exacerbations and hospitalizations. Objective cough monitoring methods have emerged as superior alternatives to subjective methods like questionnaires. In recent years, cough has been monitored using wearable devices equipped with microphones. However, the discrimination of cough sounds from background noise has been shown a particular challenge. This study aimed to demonstrate the effectiveness of single-axis acceleration signals combined with state-of-the-art deep learning (DL) algorithms to distinguish intentional coughing from sounds like speech, laugh, or throat noises. Various DL methods (recurrent, convolutional, and deep convolutional neural networks) combined with one- and two-dimensional time and time–frequency representations, such as the signal envelope, kurtogram, wavelet scalogram, mel, Bark, and the equivalent rectangular bandwidth spectrum (ERB) spectrograms, were employed to identify the most effective approach. The optimal strategy, which involved the SqueezeNet model in conjunction with wavelet scalograms, yielded an accuracy and precision of 92.21% and 95.59%, respectively. The proposed method demonstrated its potential for cough monitoring. Future research will focus on validating the system in spontaneous coughing of subjects with respiratory diseases under natural ambulatory conditions. Full article
Show Figures

Figure 1

32 pages, 7815 KB  
Article
Neural Adaptation at Stimulus Onset and Speed of Neural Processing as Critical Contributors to Speech Comprehension Independent of Hearing Threshold or Age
by Jakob Schirmer, Stephan Wolpert, Konrad Dapper, Moritz Rühle, Jakob Wertz, Marjoleen Wouters, Therese Eldh, Katharina Bader, Wibke Singer, Etienne Gaudrain, Deniz Başkent, Sarah Verhulst, Christoph Braun, Lukas Rüttiger, Matthias H. J. Munk, Ernst Dalhoff and Marlies Knipper
J. Clin. Med. 2024, 13(9), 2725; https://doi.org/10.3390/jcm13092725 - 6 May 2024
Cited by 10 | Viewed by 3595
Abstract
Background: It is assumed that speech comprehension deficits in background noise are caused by age-related or acquired hearing loss. Methods: We examined young, middle-aged, and older individuals with and without hearing threshold loss using pure-tone (PT) audiometry, short-pulsed distortion-product otoacoustic emissions [...] Read more.
Background: It is assumed that speech comprehension deficits in background noise are caused by age-related or acquired hearing loss. Methods: We examined young, middle-aged, and older individuals with and without hearing threshold loss using pure-tone (PT) audiometry, short-pulsed distortion-product otoacoustic emissions (pDPOAEs), auditory brainstem responses (ABRs), auditory steady-state responses (ASSRs), speech comprehension (OLSA), and syllable discrimination in quiet and noise. Results: A noticeable decline of hearing sensitivity in extended high-frequency regions and its influence on low-frequency-induced ABRs was striking. When testing for differences in OLSA thresholds normalized for PT thresholds (PTTs), marked differences in speech comprehension ability exist not only in noise, but also in quiet, and they exist throughout the whole age range investigated. Listeners with poor speech comprehension in quiet exhibited a relatively lower pDPOAE and, thus, cochlear amplifier performance independent of PTT, smaller and delayed ABRs, and lower performance in vowel-phoneme discrimination below phase-locking limits (/o/-/u/). When OLSA was tested in noise, listeners with poor speech comprehension independent of PTT had larger pDPOAEs and, thus, cochlear amplifier performance, larger ASSR amplitudes, and higher uncomfortable loudness levels, all linked with lower performance of vowel-phoneme discrimination above the phase-locking limit (/i/-/y/). Conslusions: This study indicates that listening in noise in humans has a sizable disadvantage in envelope coding when basilar-membrane compression is compromised. Clearly, and in contrast to previous assumptions, both good and poor speech comprehension can exist independently of differences in PTTs and age, a phenomenon that urgently requires improved techniques to diagnose sound processing at stimulus onset in the clinical routine. Full article
Show Figures

Graphical abstract

17 pages, 3162 KB  
Article
A Mixed-Rate Strategy on a Bilaterally-Synchronized Cochlear Implant Processor Offering the Opportunity to Provide Both Speech Understanding and Interaural Time Difference Cues
by Stephen R. Dennison, Tanvi Thakkar, Alan Kan, Mario A. Svirsky, Mahan Azadpour and Ruth Y. Litovsky
J. Clin. Med. 2024, 13(7), 1917; https://doi.org/10.3390/jcm13071917 - 26 Mar 2024
Cited by 4 | Viewed by 2243
Abstract
Background/Objective: Bilaterally implanted cochlear implant (CI) users do not consistently have access to interaural time differences (ITDs). ITDs are crucial for restoring the ability to localize sounds and understand speech in noisy environments. Lack of access to ITDs is partly due to lack [...] Read more.
Background/Objective: Bilaterally implanted cochlear implant (CI) users do not consistently have access to interaural time differences (ITDs). ITDs are crucial for restoring the ability to localize sounds and understand speech in noisy environments. Lack of access to ITDs is partly due to lack of communication between clinical processors across the ears and partly because processors must use relatively high rates of stimulation to encode envelope information. Speech understanding is best at higher stimulation rates, but sensitivity to ITDs in the timing of pulses is best at low stimulation rates. Methods: We implemented a practical “mixed rate” strategy that encodes ITD information using a low stimulation rate on some channels and speech information using high rates on the remaining channels. The strategy was tested using a bilaterally synchronized research processor, the CCi-MOBILE. Nine bilaterally implanted CI users were tested on speech understanding and were asked to judge the location of a sound based on ITDs encoded using this strategy. Results: Performance was similar in both tasks between the control strategy and the new strategy. Conclusions: We discuss the benefits and drawbacks of the sound coding strategy and provide guidelines for utilizing synchronized processors for developing strategies. Full article
Show Figures

Figure 1

20 pages, 19399 KB  
Article
Speech Inpainting Based on Multi-Layer Long Short-Term Memory Networks
by Haohan Shi, Xiyu Shi and Safak Dogan
Future Internet 2024, 16(2), 63; https://doi.org/10.3390/fi16020063 - 17 Feb 2024
Cited by 6 | Viewed by 4091
Abstract
Audio inpainting plays an important role in addressing incomplete, damaged, or missing audio signals, contributing to improved quality of service and overall user experience in multimedia communications over the Internet and mobile networks. This paper presents an innovative solution for speech inpainting using [...] Read more.
Audio inpainting plays an important role in addressing incomplete, damaged, or missing audio signals, contributing to improved quality of service and overall user experience in multimedia communications over the Internet and mobile networks. This paper presents an innovative solution for speech inpainting using Long Short-Term Memory (LSTM) networks, i.e., a restoring task where the missing parts of speech signals are recovered from the previous information in the time domain. The lost or corrupted speech signals are also referred to as gaps. We regard the speech inpainting task as a time-series prediction problem in this research work. To address this problem, we designed multi-layer LSTM networks and trained them on different speech datasets. Our study aims to investigate the inpainting performance of the proposed models on different datasets and with varying LSTM layers and explore the effect of multi-layer LSTM networks on the prediction of speech samples in terms of perceived audio quality. The inpainted speech quality is evaluated through the Mean Opinion Score (MOS) and a frequency analysis of the spectrogram. Our proposed multi-layer LSTM models are able to restore up to 1 s of gaps with high perceptual audio quality using the features captured from the time domain only. Specifically, for gap lengths under 500 ms, the MOS can reach up to 3~4, and for gap lengths ranging between 500 ms and 1 s, the MOS can reach up to 2~3. In the time domain, the proposed models can proficiently restore the envelope and trend of lost speech signals. In the frequency domain, the proposed models can restore spectrogram blocks with higher similarity to the original signals at frequencies less than 2.0 kHz and comparatively lower similarity at frequencies in the range of 2.0 kHz~8.0 kHz. Full article
(This article belongs to the Special Issue Deep Learning and Natural Language Processing II)
Show Figures

Figure 1

12 pages, 2154 KB  
Article
A Novel Computationally Efficient Approach for Exploring Neural Entrainment to Continuous Speech Stimuli Incorporating Cross-Correlation
by Luong Do Anh Quan, Le Thi Trang, Hyosung Joo, Dongseok Kim and Jihwan Woo
Appl. Sci. 2023, 13(17), 9839; https://doi.org/10.3390/app13179839 - 31 Aug 2023
Viewed by 2512
Abstract
A linear system identification technique has been widely used to track neural entrainment in response to continuous speech stimuli. Although the approach of the standard regularization method using ridge regression provides a straightforward solution to estimate and interpret neural responses to continuous speech [...] Read more.
A linear system identification technique has been widely used to track neural entrainment in response to continuous speech stimuli. Although the approach of the standard regularization method using ridge regression provides a straightforward solution to estimate and interpret neural responses to continuous speech stimuli, inconsistent results and costly computational processes can arise due to the need for parameter tuning. We developed a novel approach to the system identification method called the detrended cross-correlation function, which aims to map stimulus features to neural responses using the reverse correlation and derivative of convolution. This non-parametric (i.e., no need for parametric tuning) approach can maintain consistent results. Moreover, it provides a computationally efficient training process compared to the conventional method of ridge regression. The detrended cross-correlation function correctly captures the temporal response function to speech envelope and the spectral–temporal receptive field to speech spectrogram in univariate and multivariate forward models, respectively. The suggested model also provides more efficient computation compared to the ridge regression to process electroencephalography (EEG) signals. In conclusion, we suggest that the detrended cross-correlation function can be comparably used to investigate continuous speech- (or sound-) evoked EEG signals. Full article
(This article belongs to the Special Issue Modern Advances in Neurolinguistics and EEG Language Processing)
Show Figures

Figure 1

17 pages, 1125 KB  
Article
Investigations on the Optimal Estimation of Speech Envelopes for the Two-Stage Speech Enhancement
by Yanjue Song and Nilesh Madhu
Sensors 2023, 23(14), 6438; https://doi.org/10.3390/s23146438 - 16 Jul 2023
Cited by 2 | Viewed by 2501
Abstract
Using the source-filter model of speech production, clean speech signals can be decomposed into an excitation component and an envelope component that is related to the phoneme being uttered. Therefore, restoring the envelope of degraded speech during speech enhancement can improve the intelligibility [...] Read more.
Using the source-filter model of speech production, clean speech signals can be decomposed into an excitation component and an envelope component that is related to the phoneme being uttered. Therefore, restoring the envelope of degraded speech during speech enhancement can improve the intelligibility and quality of output. As the number of phonemes in spoken speech is limited, they can be adequately represented by a correspondingly limited number of envelopes. This can be exploited to improve the estimation of speech envelopes from a degraded signal in a data-driven manner. The improved envelopes are then used in a second stage to refine the final speech estimate. Envelopes are typically derived from the linear prediction coefficients (LPCs) or from the cepstral coefficients (CCs). The improved envelope is obtained either by mapping the degraded envelope onto pre-trained codebooks (classification approach) or by directly estimating it from the degraded envelope (regression approach). In this work, we first investigate the optimal features for envelope representation and codebook generation by a series of oracle tests. We demonstrate that CCs provide better envelope representation compared to using the LPCs. Further, we demonstrate that a unified speech codebook is advantageous compared to the typical codebook that manually splits speech and silence as separate entries. Next, we investigate low-complexity neural network architectures to map degraded envelopes to the optimal codebook entry in practical systems. We confirm that simple recurrent neural networks yield good performance with a low complexity and number of parameters. We also demonstrate that with a careful choice of the feature and architecture, a regression approach can further improve the performance at a lower computational cost. However, as also seen from the oracle tests, the benefit of the two-stage framework is now chiefly limited by the statistical noise floor estimate, leading to only a limited improvement in extremely adverse conditions. This highlights the need for further research on joint estimation of speech and noise for optimum enhancement. Full article
(This article belongs to the Special Issue Machine Learning and Signal Processing Based Acoustic Sensors)
Show Figures

Figure 1

12 pages, 2813 KB  
Article
Contributions of Temporal Modulation Cues in Temporal Amplitude Envelope of Speech to Urgency Perception
by Masashi Unoki, Miho Kawamura, Maori Kobayashi, Shunsuke Kidani, Junfeng Li and Masato Akagi
Appl. Sci. 2023, 13(10), 6239; https://doi.org/10.3390/app13106239 - 19 May 2023
Cited by 1 | Viewed by 2410
Abstract
We previously investigated the perception of noise-vocoded speech to determine whether the temporal amplitude envelope (TAE) of speech plays an important role in the perception of linguistic information as well as non-linguistic information. However, it remains unclear if these TAEs also play a [...] Read more.
We previously investigated the perception of noise-vocoded speech to determine whether the temporal amplitude envelope (TAE) of speech plays an important role in the perception of linguistic information as well as non-linguistic information. However, it remains unclear if these TAEs also play a role in the urgency perception of non-linguistic information. In this paper, we comprehensively investigated whether the TAE of speech contributes to urgency perception. To this end, we compared noise-vocoded stimuli containing TAEs identical to those of original speech with those containing TAEs controlled by low-pass or high-pass filtering. We derived degrees of urgency from a paired comparison of the results and then used them as a basis to clarify the relationship between the temporal modulation components in TAEs of speech and urgency perception. Our findings revealed that (1) the perceived degrees of urgency of noise-vocoded stimuli are similar to those of the original, (2) significant cues for urgency perception are temporal modulation components of the noise-vocoded stimuli higher than the modulation frequency of 6 Hz, (3) additional significant cues for urgency perception are temporal modulation components of the noise-vocoded stimuli lower than the modulation frequency of 8 Hz, and (4) the TAE of the time-reversal speech is not likely to contain important cues for the perception of urgency. We therefore conclude that temporal modulation cues in the TAE of speech are a significant component in the perception of urgency. Full article
(This article belongs to the Special Issue Audio, Speech and Language Processing)
Show Figures

Figure 1

Back to TopTop