Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (23)

Search Parameters:
Keywords = acoustic salience

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
21 pages, 14719 KB  
Article
Respiratory Disease Classification Using NMF-Enhanced Log-Mel Spectrograms and Convolutional Recurrent Neural Networks
by Bowen Han, Wei Quan, Bogdan Matuszewski and Dennis Corbett
Sensors 2026, 26(13), 4268; https://doi.org/10.3390/s26134268 - 4 Jul 2026
Viewed by 538
Abstract
Respiratory disease classification using lung sound recordings remains challenging due to signal interference, heterogeneous acquisition conditions, and substantial overlap among clinically related acoustic patterns. This study presents a framework for respiratory disease classification using NMF-enhanced log-mel spectrograms and deep neural classifiers. Respiratory sound [...] Read more.
Respiratory disease classification using lung sound recordings remains challenging due to signal interference, heterogeneous acquisition conditions, and substantial overlap among clinically related acoustic patterns. This study presents a framework for respiratory disease classification using NMF-enhanced log-mel spectrograms and deep neural classifiers. Respiratory sound recordings from two publicly available datasets were harmonized into a unified label space comprising Asthma, Bronchiectasis, Bronchiolitis, COPD, Healthy, Pneumonia and URTI. Following signal standardization and fixed-length segmentation, a non-negative matrix factorization (NMF)-based enhancement stage was applied to increase the salience of respiratory components prior to log-mel spectrogram generation. The proposed classifier was a convolutional recurrent neural network (CRNN) that combined convolutional feature extraction, bidirectional recurrent modelling, and attention-based temporal aggregation. For comparison, RDLINet, a conventional CNN, ResNet, and a YOLO-style backbone were implemented under the same preprocessing and training framework. Experimental results demonstrated that the proposed CRNN achieved the best overall performance, attaining 96.14 ± 0.50% accuracy and 94.05 ± 1.21% Macro-F1 on the unified seven-class cohort. Class-wise analysis, confusion-matrix evaluation, and output-space visualization further showed that the CRNN provided more balanced recognition across disease categories and clearer class separation than competing architectures. These findings indicate that NMF-enhanced spectro-temporal modelling combined with convolutional recurrent learning offers an effective approach for automated multi-class respiratory disease classification. Full article
Show Figures

Figure 1

26 pages, 3508 KB  
Article
Dual-Track Residual Framework for Residual Strength-Controlled Emotional Speech Synthesis
by Youdong Ding, Yafan Geng, Wenjing Yu and Feifan Cai
Appl. Sci. 2026, 16(13), 6613; https://doi.org/10.3390/app16136613 - 2 Jul 2026
Viewed by 501
Abstract
Recent text-to-speech (TTS) systems can synthesize natural and intelligible speech, but adding controllable emotional expression to a pretrained model while preserving target-speaker identity remains challenging. This setting is especially constrained when the acoustic backbone is kept frozen and emotional adaptation relies on additional [...] Read more.
Recent text-to-speech (TTS) systems can synthesize natural and intelligible speech, but adding controllable emotional expression to a pretrained model while preserving target-speaker identity remains challenging. This setting is especially constrained when the acoustic backbone is kept frozen and emotional adaptation relies on additional trainable modules. We study emotional adaptation for a frozen flow-matching TTS backbone and propose the dual-track residual framework (DTRF). The DTRF keeps the neutral-adapted Matcha-Base backbone unchanged, represents emotion as a neutral-anchor residual, and introduces two residual control paths: an asymmetric zero-initialized acoustic control branch for spectral vector field modulation and an emotional duration adapter (EDA) for duration-level prosody control. Rather than injecting emotion only into the acoustic path, the DTRF applies emotion control to both acoustic vector field prediction and phoneme duration prediction, jointly adjusting spectral realization and temporal prosody. A global neutral anchor converts absolute emotion embeddings into relative residuals so that the control signal describes the deviation from neutral speech toward the target emotion rather than an absolute style vector. During inference, a shared scalar factor α scales both residual paths, providing a practical residual strength interface for controllable emotion rendering. Moderate α values tend to increase emotional salience, whereas larger extrapolative values introduce trade-offs in naturalness, speaker similarity, and intelligibility. Experiments on the English subset of the Emotional Speech Dataset (ESD) show that the DTRF improves emotion-related metrics relative to the internal full-parameter updating and style token conditioning baselines, while maintaining a practical balance among speaker similarity, naturalness, and intelligibility. The emotion control modules contain approximately 27 M trainable parameters, corresponding to 29.61% of the full model parameters and a 70.39% reduction compared with full-parameter updating. These results suggest that jointly modeling acoustic and duration residuals can be an effective strategy for adding residual strength-controlled emotional rendering to a frozen flow-matching TTS model without full-backbone updating. Full article
(This article belongs to the Special Issue Deep Learning for Speech, Image and Language Processing)
Show Figures

Figure 1

22 pages, 1398 KB  
Article
Phoneme Monitoring in Developmental Dyslexia: Pupillometric Evidence for Cognitive Rather than Acoustic Origins of Phonological Deficits
by Marina Rossi, Massimiliano Canzi and Tamara V. Rathcke
Brain Sci. 2026, 16(7), 697; https://doi.org/10.3390/brainsci16070697 - 30 Jun 2026
Viewed by 271
Abstract
Background. Developmental dyslexia (DD) is a neurodevelopmental disorder characterized by persistent difficulties acquiring fluent reading. A core feature is impaired phonological processing, though its etiology remains debated. Two competing accounts attribute phonological deficits either to reduced acoustic sensitivity to lexical stress cues or [...] Read more.
Background. Developmental dyslexia (DD) is a neurodevelopmental disorder characterized by persistent difficulties acquiring fluent reading. A core feature is impaired phonological processing, though its etiology remains debated. Two competing accounts attribute phonological deficits either to reduced acoustic sensitivity to lexical stress cues or to insufficient cognitive support during phoneme processing. Methods. To test these accounts, 57 Italian children (28 with DD, 29 typically developing) completed a phoneme monitoring task in which targets appeared in strong or weak syllables, varying in acoustic salience. A composite acoustic salience factor was derived from target cue properties, and an individual cognitive factor was computed from IQ, working memory, and shifting attention. Pupillometry was used to assess auditory sensitivity to acoustic salience and cognitive effort in real time. Results. The results showed that, behaviourally, children with DD showed significantly lower target identification accuracy and d’-sensitivity. However, pupil dilation during target processing did not differ between the two groups, while children with DD showed reduced pupillary responses on trials involving distractor rejection and missed targets. These physiological patterns correlated primarily with individual cognitive scores rather than acoustic salience. Conclusions. Taken together, these findings point toward an important role of individually varying cognitive resources during phonological processing and highlight the value of pupillometry as a sensitive, real-time index of cognitive engagement during a phonological task. Full article
(This article belongs to the Special Issue Exploring Neurophysiology Aspect in Dyslexia)
Show Figures

Figure 1

20 pages, 3786 KB  
Article
A New Condition Diagnosis Method for Ball Bearings Using Ultrasonic Visualization and Light CNN
by Hangyeol Jo, Sung-Ho Hong, Choon-Su Park, Moonsuk Kim, Miao Dai and Sang-Woo Ban
Lubricants 2026, 14(7), 249; https://doi.org/10.3390/lubricants14070249 - 23 Jun 2026
Viewed by 337
Abstract
Early fault diagnosis of ball bearings is essential for maintaining the reliability of rotating machinery and preventing unexpected downtime. This study proposes a fault diagnosis framework that combines non-contact ultrasonic visualization with a lightweight convolutional neural network (Light CNN). Seven bearing conditions, including [...] Read more.
Early fault diagnosis of ball bearings is essential for maintaining the reliability of rotating machinery and preventing unexpected downtime. This study proposes a fault diagnosis framework that combines non-contact ultrasonic visualization with a lightweight convolutional neural network (Light CNN). Seven bearing conditions, including ferrous-particle contamination and grease starvation, were investigated using ultrasonic, vibration, and acoustic emission (AE) sensors under identical experimental conditions. Saliency-map extraction and two-dimensional histogram analysis were applied to ultrasonic RGB images to generate compact feature representations, which were compressed into 20 × 20 feature maps and used as inputs to a three-layer Light CNN. The proposed method achieved an average classification accuracy of 99.98% and an F1-score of 99.98%. In addition, an average inference throughput of 11.47 IPS was obtained, representing approximately ten times higher computational efficiency than vibration- and AE-based approach-es. Stable diagnostic performance was also maintained under a low-speed operating condition of 500 rpm. These results demonstrate the effectiveness of combining ultrasonic visualization and a lightweight CNN for accurate and computationally efficient bearing fault diagnosis. Full article
(This article belongs to the Special Issue Multiphysics Modelling in Bearing Lubrication)
Show Figures

Figure 1

26 pages, 9596 KB  
Article
Enhanced Neural Responses to Self-Name Stimuli Relative to Tone and Reversed Speech Deviants in the Auditory Oddball Paradigm
by Fang Duan, Xiongping Cao, Zheng Yan and Jianming Chen
Brain Sci. 2026, 16(6), 608; https://doi.org/10.3390/brainsci16060608 - 2 Jun 2026
Viewed by 484
Abstract
Background: Auditory oddball paradigms are widely used to investigate neural responses to deviant stimuli and attentional processing. However, different paradigms involve deviant stimuli with varying levels of stimulus relevance, and the corresponding neural responses have rarely been directly compared within a unified [...] Read more.
Background: Auditory oddball paradigms are widely used to investigate neural responses to deviant stimuli and attentional processing. However, different paradigms involve deviant stimuli with varying levels of stimulus relevance, and the corresponding neural responses have rarely been directly compared within a unified experimental framework. The aim of this study was to compare neural responses elicited by three variants of the auditory oddball paradigm that differ in the type of deviant stimuli: tone, reversed speech, and self-name deviants. Methods: Electroencephalography (EEG) data were recorded from 38 healthy participants while they performed three paradigm variants. Event-related potentials (ERPs) were analyzed to examine neural responses to deviant stimuli. In addition, cortical activation patterns were identified via source reconstruction, and classification analyses were conducted to assess the discriminability of neural responses across the three variants. Results: ERP results revealed that the self-name paradigm elicited the largest ERP responses, characterized by a significant P300 amplitude (3.95 μV) and prominent MMN (−6.39 μV). Crucially, source-space analysis revealed a graded expansion of cortical recruitment: acoustic deviance (tone) and structural reanalysis (reversed speech) were associated with 7 and 6 significant clusters, respectively, primarily in the auditory and fronto-cingulate cortices, whereas the self-name paradigm engaged 12 significant clusters spanning a distributed network encompassing salience-processing regions and cortical midline structures associated with self-referential processing (including the insula and posterior cingulate cortex). Classification analyses mirrored these findings, with the self-name paradigm consistently yielding the highest neural separability (~80% accuracy) and greater robustness to interindividual variability, demonstrating the superior discriminability of self-referential neural patterns. Conclusions: These findings demonstrate that self-referential auditory stimuli elicit stronger and more discriminable neural responses than other auditory deviant stimuli in the oddball paradigm. These results provide a comparative perspective on how different dimensions of auditory relevance modulate neural processing and may inform the design of effective auditory paradigms for cognitive neuroscience and related translational applications. Full article
(This article belongs to the Section Sensory and Motor Neuroscience)
Show Figures

Figure 1

18 pages, 1083 KB  
Article
The Perception of Novel Consonant Clusters: A Comparison of Salience and Sonority
by Marina Oganyan, Matthew C. Kelley, Yuan Chai, Akira Omaki and Richard A. Wright
Brain Sci. 2026, 16(6), 583; https://doi.org/10.3390/brainsci16060583 - 29 May 2026
Viewed by 392
Abstract
Background/Objectives: In human language, speech sound units, referred to as segments, are rigidly ordered. Certain orderings are typologically common, while others are typologically rare. This pattern has led linguists to posit a scalar segment-intrinsic feature, referred to as “sonority,” and a hierarchy relating [...] Read more.
Background/Objectives: In human language, speech sound units, referred to as segments, are rigidly ordered. Certain orderings are typologically common, while others are typologically rare. This pattern has led linguists to posit a scalar segment-intrinsic feature, referred to as “sonority,” and a hierarchy relating to ordering referred to as the “sonority hierarchy.” In generative phonology, it has been proposed that the sonority hierarchy is present in the innate underlying grammar as the Sonority Sequencing Principle (SSP) and variations thereof. In grammar-based approaches it is common for speech sounds to be abstractly classified based on manner with obstruents being the least sonorous and vowels being most sonorous and with three levels of adherence to the SSP: 1 strict adherence with a sonority rise as in the sequence /pla/, 2 slight adherence/slight violation with a sonority plateau as in /pta/ or /mna/, and 3 violation with a sonority fall before the vowel as in /lpa/. Some linguists have proposed that segmental ordering is the result of generalized pressures relating to speech perception, cognitive processing, and gestural coordination, and therefore not part of the underlying grammar. This study examines the relative contribution of perceptual salience (defined as loudness and segmental recoverability) to segmental ordering. Methods: We used eye-tracking in the visual world paradigm to track real-time processing of perception of novel consonant clusters. We compared whether salience or sonority had a stronger association with how participants simplify clusters in perception, where salience and sonority had different predictions. Results: We found that participants looked more towards salience-predicted competitors than sonority-predicted ones prior to focusing on the target. Conclusions: This finding is in line with theories based on acoustic and auditory salience being the strongest predictors of perceptual ease. Full article
(This article belongs to the Special Issue Language Perception and Processing)
Show Figures

Figure 1

19 pages, 5566 KB  
Article
Noise Characteristics and Multi-Dimensional Sound Quality Evaluation of High-Frequency Transformers Under Non-Sinusoidal Excitation
by Cai Zeng, Li Li, Yexin Zhu, Xing Du, Jie Zhang, Xiaoqiong He and Xinbiao Xiao
Acoustics 2026, 8(2), 28; https://doi.org/10.3390/acoustics8020028 - 26 Apr 2026
Viewed by 856
Abstract
High-frequency transformer (HFT) noise is a pivotal indicator of equipment performance. To conduct a comprehensive evaluation, this study systematically performed testing and evaluation on the noise generated by a 70 kW HFT under no-load conditions. Acoustic data were collected using acoustic sensors and [...] Read more.
High-frequency transformer (HFT) noise is a pivotal indicator of equipment performance. To conduct a comprehensive evaluation, this study systematically performed testing and evaluation on the noise generated by a 70 kW HFT under no-load conditions. Acoustic data were collected using acoustic sensors and a head-and-torso simulator, followed by an analysis of noise characteristics focusing on the impacts of voltage levels and operating frequencies. A multi-dimensional evaluation of HFT noise was carried out using sound quality parameters to unravel its intrinsic attributes under electrical parameter excitation. The key findings are as follows: HFT noise exhibits steady-state time-domain behavior and distinct tonal frequency-domain features; the dominant frequency is twice the operating frequency, with prominent harmonics. The noise intensity increases with the voltage levels (~47.0 dB (A) at 200 V to ~72.0 dB (A) at 750 V at 5 kHz) but decreases with the operating frequencies (~82.0 dB (A) at 4 kHz to ~47.0 dB (A) at 10 kHz at 750 V). This study establishes correlations between the electrical parameters and sound quality metrics; the loudness, sharpness, tone-to-noise ratio and prominence ratio are sensitive to the electrical parameters of HFT. Single-frequency noise from HFT exhibits remarkable perceptual salience, exacerbating the perceived annoyance. Thus, HFT design should prioritize reducing single-frequency noise to alleviate such issues. Full article
Show Figures

Figure 1

22 pages, 3040 KB  
Article
Prefabricated Co-Working Spaces’ Window Design: Emotional Salience Scale-Based Optimisation
by Antonio Ciervo, Massimiliano Masullo, Luigi Maffei, Roxana Adina Toma, Maria Dolores Morelli and Michelangelo Scorpio
Buildings 2026, 16(4), 875; https://doi.org/10.3390/buildings16040875 - 22 Feb 2026
Viewed by 705
Abstract
Windows are key elements of the building’s system; they connect workers with the outdoor environment, influence daylight penetration, sound insulation, and thermal exchanges of façades, but they also moderate the workers’ well-being and productivity. This research investigates how the window-to-wall ratio, as well [...] Read more.
Windows are key elements of the building’s system; they connect workers with the outdoor environment, influence daylight penetration, sound insulation, and thermal exchanges of façades, but they also moderate the workers’ well-being and productivity. This research investigates how the window-to-wall ratio, as well as the position and orientation of mullions, in movable offices affect the combination of workers’ perceptual and emotional responses. A smart co-working prefabricated movable office was modelled in virtual reality to include dynamic visual elements and acoustic stimuli. Experiments were performed in a laboratory under controlled thermal conditions involving 32 volunteers. The Igroup Presence and Emotional Salience Questionnaires were used to collect subjective responses. ANOVA analysis and post hoc test with the Bonferroni correction were used for data elaboration. Results revealed that window design affects emotional salience. High window-to-wall ratio and no mullions achieved the highest scores. Increasing the number of mullions, particularly when they obstruct key visual elements, reduced the positive emotional salience rating. Horizontal mullions diminish the outdoors’ spatial perception, interrupting visual continuity and restricting users’ capacity to recognise variations in the views. Finally, the results suggest some valuable insights and suggestions that can help designers improve window design and people’s well-being and satisfaction. Full article
(This article belongs to the Section Building Energy, Physics, Environment, and Systems)
Show Figures

Figure 1

36 pages, 5431 KB  
Article
Explainable AI-Driven Quality and Condition Monitoring in Smart Manufacturing
by M. Nadeem Ahangar, Z. A. Farhat, Aparajithan Sivanathan, N. Ketheesram and S. Kaur
Sensors 2026, 26(3), 911; https://doi.org/10.3390/s26030911 - 30 Jan 2026
Cited by 9 | Viewed by 2753
Abstract
Artificial intelligence (AI) is increasingly adopted in manufacturing for tasks such as automated inspection, predictive maintenance, and condition monitoring. However, the opaque, black-box nature of many AI models remains a major barrier to industrial trust, acceptance, and regulatory compliance. This study investigates how [...] Read more.
Artificial intelligence (AI) is increasingly adopted in manufacturing for tasks such as automated inspection, predictive maintenance, and condition monitoring. However, the opaque, black-box nature of many AI models remains a major barrier to industrial trust, acceptance, and regulatory compliance. This study investigates how explainable artificial intelligence (XAI) techniques can be used to systematically open and interpret the internal reasoning of AI systems commonly deployed in manufacturing, rather than to optimise or compare model performance. A unified explainability-centred framework is proposed and applied across three representative manufacturing use cases encompassing heterogeneous data modalities and learning paradigms: vision-based classification of casting defects, vision-based localisation of metal surface defects, and unsupervised acoustic anomaly detection for machine condition monitoring. Diverse models are intentionally employed as representative black-box decision-makers to evaluate whether XAI methods can provide consistent, physically meaningful explanations independent of model architecture, task formulation, or supervision strategy. A range of established XAI techniques, including Grad-CAM, Integrated Gradients, Saliency Maps, Occlusion Sensitivity, and SHAP, are applied to expose model attention, feature relevance, and decision drivers across visual and acoustic domains. The results demonstrate that XAI enables alignment between model behaviour and physically interpretable defect and fault mechanisms, supporting transparent, auditable, and human-interpretable decision-making. By positioning explainability as a core operational requirement rather than a post hoc visual aid, this work contributes a cross-modal framework for trustworthy AI in manufacturing, aligned with Industry 5.0 principles, human-in-the-loop oversight, and emerging expectations for transparent and accountable industrial AI systems. Full article
(This article belongs to the Section Intelligent Sensors)
Show Figures

Figure 1

16 pages, 1157 KB  
Article
User-Centered Redesign of Monitoring Alarms: A Pre–Post Study on Perception, Functionality, and Recognizability Following Real-Life Clinical Implementation
by Cynthia Hunn, Christoph B. Nöthiger, Julia Braun, Yoko Sen, Avery Sen, Samira Akbas, Matthias Hoffmann, Elena Neumann, Greta Gasciauskaite, David W. Tscholl and Tadzio R. Roche
Healthcare 2025, 13(23), 3033; https://doi.org/10.3390/healthcare13233033 - 24 Nov 2025
Cited by 1 | Viewed by 784
Abstract
Background: Auditory alarms in patient monitoring are vital for clinical safety, but their harsh acoustic properties and high frequency contribute to stress, alarm fatigue, and reduced acceptance among healthcare staff. In collaboration with Sen Sound, Philips redesigned its alarm sounds to reduce auditory [...] Read more.
Background: Auditory alarms in patient monitoring are vital for clinical safety, but their harsh acoustic properties and high frequency contribute to stress, alarm fatigue, and reduced acceptance among healthcare staff. In collaboration with Sen Sound, Philips redesigned its alarm sounds to reduce auditory harshness, particularly for low- and medium-priority alarms, while preserving the salience of high-priority alerts. This study evaluated the impact of these refined alarm sounds in a real-world clinical setting. Objective: The goal was to determine whether anesthesia professionals perceive the refined Philips alarm sounds as more pleasant, clinically appropriate, and reliably recognizable compared with the traditional sounds. Methods: We conducted a single-center, pre–post intervention study at the University Hospital Zurich, Switzerland. Anesthesia providers assessed traditional and refined Philips alarm sounds with respect to perceived sound appeal, perceived functionality, and recognition accuracy. The primary outcome (sound appeal) was tested for superiority; using mixed-effects regression models. Results: Seventy-seven participants completed both study phases. Refined alarm sounds significantly improved perceived sound appeal (mean difference +0.51; 95% CI, 0.37–0.64; p < 0.001), while perceived functionality showed a small decrease (mean difference −0.15; 95% CI, −0.27 to −0.03). Recognition accuracy for low- and medium-priority alarms was higher with traditional sounds (low: 95.2% vs. 87.5%, p = 0.002; medium: 81.1% vs. 62.0%, p < 0.001), while high-priority alarms were more accurately identified with refined sounds (89.0% vs. 81.4%, p = 0.002). Overall, 71% of participants preferred the refined sounds, and 92% supported further development. Conclusions: Refined alarm sounds reduced perceived harshness and improved auditory comfort for anesthesia providers, but were associated with slightly lower perceived functionality and mixed recognition accuracy. High-priority alarms were identified more reliably, whereas low- and medium-priority alarms were less distinctly recognized, indicating a limited trade-off between sound appeal and clarity that primarily affected lower-priority signals. These findings suggest that while refinement can enhance the auditory environment, further development, potentially incorporating auditory icons or voice-based alerts, will be needed to optimize both user experience and patient safety in clinical practice. Full article
(This article belongs to the Section Clinical Care)
Show Figures

Figure 1

24 pages, 1784 KB  
Article
Indoor Soundscape Perception and Soundscape Appropriateness Assessment While Working at Home: A Comparative Study with Relaxing Activities
by Jiaxin Li, Yong Huang, Rumei Han, Yuan Zhang and Jian Kang
Buildings 2025, 15(15), 2642; https://doi.org/10.3390/buildings15152642 - 26 Jul 2025
Cited by 5 | Viewed by 2364
Abstract
The COVID-19 pandemic’s rapid shift to working from home has fundamentally challenged residential acoustic design, which traditionally prioritises rest and relaxation rather than sustained concentration. However, a clear gap exists in understanding how acoustic needs and the subjective evaluation of soundscape appropriateness ( [...] Read more.
The COVID-19 pandemic’s rapid shift to working from home has fundamentally challenged residential acoustic design, which traditionally prioritises rest and relaxation rather than sustained concentration. However, a clear gap exists in understanding how acoustic needs and the subjective evaluation of soundscape appropriateness (SA) differ between these conflicting activities within the same domestic space. Addressing this gap, this study reveals critical differences in how people experience and evaluate home soundscapes during work versus relaxation activities in the same residential spaces. Through an online survey of 247 Chinese participants during lockdown, we assessed soundscape perception attributes, the perceived saliencies of various sound types, and soundscape appropriateness (SA) ratings while working and relaxing at home. Our findings demonstrate that working at home creates a more demanding acoustic context: participants perceived indoor soundscapes as significantly less comfortable and less full of content when working compared to relaxing (p < 0.001), with natural sounds becoming less noticeable (−13.3%) and distracting household sounds more prominent (+7.5%). Structural equation modelling revealed distinct influence mechanisms: while comfort significantly mediates SA enhancement in both activities, the effect is stronger during relaxation (R2 = 0.18). Critically, outdoor man-made noise, building-service noise, and neighbour sounds all negatively impact SA during work, with neighbour sounds showing the largest detrimental effect (total effect size = −0.17), whereas only neighbour sounds and outdoor man-made noise significantly disrupt relaxation activities. Additionally, natural sounds act as a positive factor during relaxation. These results expose a fundamental mismatch: existing residential acoustic environments, designed primarily for rest, fail to support the cognitive demands of work activities. This study provides evidence-based insights for acoustic design interventions, emphasising the need for activity-specific soundscape considerations in residential spaces. As hybrid work arrangements become the norm post-pandemic, our findings highlight the urgency of reimagining residential acoustic design to accommodate both focused work and restorative relaxation within the same home. Full article
(This article belongs to the Section Architectural Design, Urban Science, and Real Estate)
Show Figures

Figure 1

16 pages, 1589 KB  
Article
Evaluating the Relative Perceptual Salience of Linguistic and Emotional Prosody in Quiet and Noisy Contexts
by Minyue Zhang, Hui Zhang, Enze Tang, Hongwei Ding and Yang Zhang
Behav. Sci. 2023, 13(10), 800; https://doi.org/10.3390/bs13100800 - 26 Sep 2023
Cited by 10 | Viewed by 6078
Abstract
How people recognize linguistic and emotional prosody in different listening conditions is essential for understanding the complex interplay between social context, cognition, and communication. The perception of both lexical tones and emotional prosody depends on prosodic features including pitch, intensity, duration, and voice [...] Read more.
How people recognize linguistic and emotional prosody in different listening conditions is essential for understanding the complex interplay between social context, cognition, and communication. The perception of both lexical tones and emotional prosody depends on prosodic features including pitch, intensity, duration, and voice quality. However, it is unclear which aspect of prosody is perceptually more salient and resistant to noise. This study aimed to investigate the relative perceptual salience of emotional prosody and lexical tone recognition in quiet and in the presence of multi-talker babble noise. Forty young adults randomly sampled from a pool of native Mandarin Chinese with normal hearing listened to monosyllables either with or without background babble noise and completed two identification tasks, one for emotion recognition and the other for lexical tone recognition. Accuracy and speed were recorded and analyzed using generalized linear mixed-effects models. Compared with emotional prosody, lexical tones were more perceptually salient in multi-talker babble noise. Native Mandarin Chinese participants identified lexical tones more accurately and quickly than vocal emotions at the same signal-to-noise ratio. Acoustic and cognitive dissimilarities between linguistic prosody and emotional prosody may have led to the phenomenon, which calls for further explorations into the underlying psychobiological and neurophysiological mechanisms. Full article
Show Figures

Figure 1

19 pages, 7952 KB  
Article
Analysis and Hardware Architecture on FPGA of a Robust Audio Fingerprinting Method Using SSM
by Ignacio Algredo-Badillo, Brenda Sánchez-Juárez, Kelsey A. Ramírez-Gutiérrez, Claudia Feregrino-Uribe, Francisco López-Huerta and Johan J. Estrada-López
Technologies 2022, 10(4), 86; https://doi.org/10.3390/technologies10040086 - 19 Jul 2022
Cited by 4 | Viewed by 4259
Abstract
The significant volume of sharing of digital media has recently increased due to the pandemic, raising the number of unauthorized uses of these media, such as emerging unauthorized copies, forgery, the lack of copyright, and electronic fraud, among others. In particular, several applications [...] Read more.
The significant volume of sharing of digital media has recently increased due to the pandemic, raising the number of unauthorized uses of these media, such as emerging unauthorized copies, forgery, the lack of copyright, and electronic fraud, among others. In particular, several applications integrate services or products such as music distribution, content management, audiobooks, streaming, and so on, which require users to demonstrate and guarantee their audio ownership. The use of acoustic fingerprint technology has emerged as a solution that is widely used to secure audio applications. This technique extracts and analyzes certain information that identifies the inherent properties of a partial or complete audio file. In this paper, we introduce two audio fingerprinting hardware architectures with a feature extraction system based on spectrogram saliency maps (SSM) and a brute-force search. The first of these conducts a search in 33 saliency maps of 32 × 32 pixels in size. After analyzing the first algorithm, a second architecture is proposed, in which the saliency map is reduced to 27 × 25 pixels, requiring 75.67% fewer hardware resources, lowering the power consumption by 64.58%, and improving the efficiency by 3.19 times via a throughput reduction of 22.29%. Full article
(This article belongs to the Special Issue 10th Anniversary of Technologies—Recent Advances and Perspectives)
Show Figures

Figure 1

16 pages, 1256 KB  
Article
Possible Event-Related Potential Correlates of Voluntary Attention and Reflexive Attention in the Emei Music Frog
by Wenjun Niu, Di Shen, Ruolei Sun, Yanzhu Fan, Jing Yang, Baowei Zhang and Guangzhan Fang
Biology 2022, 11(6), 879; https://doi.org/10.3390/biology11060879 - 8 Jun 2022
Cited by 1 | Viewed by 11047
Abstract
Attention, referring to selective processing of task-related information, is central to cognition. It has been proposed that voluntary attention (driven by current goals or tasks and under top-down control) and reflexive attention (driven by stimulus salience and under bottom-up control) struggle to control [...] Read more.
Attention, referring to selective processing of task-related information, is central to cognition. It has been proposed that voluntary attention (driven by current goals or tasks and under top-down control) and reflexive attention (driven by stimulus salience and under bottom-up control) struggle to control the focus of attention with interaction in a push–pull fashion for everyday perception in higher vertebrates. However, how auditory attention engages in auditory perception in lower vertebrates remains unclear. In this study, each component of auditory event-related potentials (ERP) related to attention was measured for the telencephalon, diencephalon and mesencephalon in the Emei music frog (Nidirana daunchina), during the broadcasting of acoustic stimuli invoking voluntary attention (using binary playback paradigm with silence replacement) and reflexive attention (using equiprobably random playback paradigm), respectively. Results showed that (1) when the sequence of acoustic stimuli could be predicted, the amplitudes of stimulus preceding negativity (SPN) evoked by silence replacement in the forebrain were significantly greater than that in the mesencephalon, suggesting voluntary attention may engage in auditory perception in this species because of the correlation between the SPN component and top-down control such as expectation and/or prediction; (2) alternately, when the sequence of acoustic stimuli could not be predicted, the N1 amplitudes evoked in the mesencephalon were significantly greater than those in other brain areas, implying that reflexive attention may be involved in auditory signal processing because the N1 components relate to selective attention; and (3) both SPN and N1 components could be evoked by the predicted stimuli, suggesting auditory perception of the music frogs might invoke the two kind of attention resources simultaneously. The present results show that human-like ERP components related to voluntary attention and reflexive attention exist in the lower vertebrates also. Full article
(This article belongs to the Section Zoology)
Show Figures

Figure 1

12 pages, 904 KB  
Article
Music with Concurrent Saliences of Musical Features Elicits Stronger Brain Responses
by Lorenzo J. Tardón, Ignacio Rodríguez-Rodríguez, Niels T. Haumann, Elvira Brattico and Isabel Barbancho
Appl. Sci. 2021, 11(19), 9158; https://doi.org/10.3390/app11199158 - 1 Oct 2021
Cited by 5 | Viewed by 3329
Abstract
Brain responses are often studied under strictly experimental conditions in which electroencephalograms (EEGs) are recorded to reflect reactions to short and repetitive stimuli. However, in real life, aural stimuli are continuously mixed and cannot be found isolated, such as when listening to music. [...] Read more.
Brain responses are often studied under strictly experimental conditions in which electroencephalograms (EEGs) are recorded to reflect reactions to short and repetitive stimuli. However, in real life, aural stimuli are continuously mixed and cannot be found isolated, such as when listening to music. In this audio context, the acoustic features in music related to brightness, loudness, noise, and spectral flux, among others, change continuously; thus, significant values of these features can occur nearly simultaneously. Such situations are expected to give rise to increased brain reaction with respect to a case in which they would appear in isolation. In order to assert this, EEG signals recorded while listening to a tango piece were considered. The focus was on the amplitude and time of the negative deflation (N100) and positive deflation (P200) after the stimuli, which was defined on the basis of the selected music feature saliences, in order to perform a statistical analysis intended to test the initial hypothesis. Differences in brain reactions can be identified depending on the concurrence (or not) of such significant values of different features, proving that coterminous increments in several qualities of music influence and modulate the strength of brain responses. Full article
(This article belongs to the Special Issue Processing Techniques Applied to Audio, Image and Brain Signals)
Show Figures

Figure 1

Back to TopTop