Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (125)

Search Parameters:
Keywords = room acoustic modelling

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
19 pages, 1730 KB  
Article
Developing a Kazakh Audio–Visual Multimodal Speech Recognition Model Based on Hierarchical and Cross-Modal Attention
by Turdybek Kurmetkan, Orken Mamyrbayev, Adem Tekerek and Ainur Toleu
Information 2026, 17(8), 756; https://doi.org/10.3390/info17080756 - 6 Aug 2026
Abstract
This study presents an audio–visual speech recognition (AVSR) model for Kazakh that jointly exploits audio and visual channels. The study introduces QazAVSR, a 57 h dataset collected from 271 speakers, and extracts synchronized audio signals and lip-region video sequences using FFmpeg 7.0, Dlib [...] Read more.
This study presents an audio–visual speech recognition (AVSR) model for Kazakh that jointly exploits audio and visual channels. The study introduces QazAVSR, a 57 h dataset collected from 271 speakers, and extracts synchronized audio signals and lip-region video sequences using FFmpeg 7.0, Dlib 19.24, and OpenCV 4.9.0. The proposed architecture uses the self-supervised HuBERT_BASE model in the audio branch and an ImageNet-pretrained ViT-B/16 model in the visual branch. Audio and visual representations are fused by a three-layer BiModalHformer block, where intra- and cross-attention operations are performed at each level. Extensive experimental validation, supplemented by rigorous paired bootstrap resampling significance tests, demonstrates that the full multimodal BiModalHformer model achieves a highly robust average character error rate (CER) of 31.2% and a Word Error Rate (WER) of 43.1%. These results significantly outperform traditional audio-only, video-only, and standard representation-level fusion baselines. Furthermore, comparisons against powerful external baseline architectures—including Whisper-Small and AV-HuBERT configurations rigorously adapted for the Kazakh language—statistically validate the architectural efficacy of the BiModalHformer framework. Additional systematic evaluations utilizing extended metrics such as the Match Error Rate (MER), word information preserved (WIP), and the Multimodal Synergy Index (MSI) confirm that the full audio–visual configuration preserves lexical information significantly more effectively. Finally, extensive noise perturbation experiments confirm that the multimodal architecture exhibits superior structural robustness to complex acoustic distortions, including environmental noise, synthetic room reverberation, and overlapping speech topologies. Full article
Show Figures

Figure 1

16 pages, 9568 KB  
Article
Equivalent Circuit Extraction of SAW Resonator with Spurious Modes Interference over a −55 °C to 85 °C Temperature Range
by Xianli Tang, Yonghao Jia and Yuandong Gu
Micromachines 2026, 17(8), 893; https://doi.org/10.3390/mi17080893 - 25 Jul 2026
Viewed by 441
Abstract
Surface acoustic wave (SAW) resonators are widely employed to design radio-frequency (RF) filters in wireless communication. To adapt to various application scenarios, the article proposes an equivalent circuit model based on the Butterworth–Van Dyke (BVD) model for SAW devices operating with spurious modes [...] Read more.
Surface acoustic wave (SAW) resonators are widely employed to design radio-frequency (RF) filters in wireless communication. To adapt to various application scenarios, the article proposes an equivalent circuit model based on the Butterworth–Van Dyke (BVD) model for SAW devices operating with spurious modes and at extreme ambient temperatures. In cases where the resonance frequencies of the spurious modes are close to those of the main mode, isolation capacitances (IC) are proposed in the modeling process. With the IC, the different resonance frequencies produced by the proposed equivalent circuit model can be flexibly adjusted. The extreme temperature influence on the SAW resonators is investigated using the proposed model. In the temperature-dependent test environment, the performance of the SAW devices changes, and these changes are captured by the proposed model. Especially for the resonators’ spurious-mode frequencies, which are less influenced by temperature near room temperature unless extreme temperatures are applied. The parameters motional resistance Rm and motional inductance Lm are considered temperature-dependent and are used to describe the influence of the ambient temperature. The RF characteristics of the SAW devices are modeled with the proposed model and verified with measurement data. The consistent results between the measured data and simulated data indicate that the proposed model is accurate and the modeling work is effective. Full article
(This article belongs to the Special Issue MEMS/NEMS Devices and Applications, 4th Edition)
Show Figures

Figure 1

22 pages, 34093 KB  
Article
The Teatro Nuovo Acoustic Heritage in Turin
by Silvana Sukaj, Amelia Trematerra and Ilaria Lombardi
Appl. Sci. 2026, 16(15), 7372; https://doi.org/10.3390/app16157372 - 23 Jul 2026
Viewed by 295
Abstract
Growing attention has recently been devoted to the role of sound environments in the preservation of cultural heritage. Within this framework, the acoustic history of the Teatro Nuovo in Turin is investigated through virtual reconstructions of its principal architectural configurations, and buildings referring [...] Read more.
Growing attention has recently been devoted to the role of sound environments in the preservation of cultural heritage. Within this framework, the acoustic history of the Teatro Nuovo in Turin is investigated through virtual reconstructions of its principal architectural configurations, and buildings referring to architecture, built for acoustic purposes and used for sound-related activities, are considered places where acoustic heritage exists. For this reason, this paper aims to find out the acoustic heritage of the Teatro Nuovo in Turin, Italy, through a simulated reconstruction of its sound field evolution, before the last theatre’s recent closing. Acoustics of both the first realization in 1936 and the altered one in 1950 are analyzed, considering geometrical and material characteristics over time. Reverberation Time, RT, Clarity, C80, Strength, G, Definition, D50, and lateral fraction, JFL, are considered, comparing the two scenarios, according to the Just Noticeable Difference, JND. The evaluation to use the space as a multipurpose auditorium follows. Conclusions underline that the second configuration was much more suitable both for speech activities and opera, with the Reverberation Time reducing from 1.78 to 1.24 s, and C80 and D50 increasing, but the multipurpose capability of the hall fair remaining. Acoustic suggestions for further restoration processes are indicated. Full article
(This article belongs to the Special Issue Acoustics Analysis and Noise Control for Buildings)
Show Figures

Figure 1

29 pages, 3549 KB  
Article
Exploratory Room-Level Acoustic Soundscape Monitoring of Cough-like Events Under Standard and Ventilation-Restricted Pig-Housing Conditions Using Audio Spectrogram Transformer
by Md Sharifuzzaman, Hong-Seok Mun, Md Kamrul Hasan, Jin-Gu Kang, Eddiemar B. Lagua, Hae-Rang Park, Keiven Mark B. Ampode, Young-Hwa Kim, Ahsan Mehtab and Chul-Ju Yang
Animals 2026, 16(14), 2275; https://doi.org/10.3390/ani16142275 - 22 Jul 2026
Viewed by 281
Abstract
Respiratory sound monitoring is a promising non-invasive tool for precision pig farming, but practical evidence from calibrated room-level deployment under degraded air-quality conditions remains limited. This study reports a 28-day exploratory room-level case study in which 52 growing pigs were housed in two [...] Read more.
Respiratory sound monitoring is a promising non-invasive tool for precision pig farming, but practical evidence from calibrated room-level deployment under degraded air-quality conditions remains limited. This study reports a 28-day exploratory room-level case study in which 52 growing pigs were housed in two rooms: one standard-ventilation room and one ventilation-restricted room, and monitored with one microphone per room emphasizing mixed room-level soundscape monitoring rather than individual pig cough counts or replicated treatment inference. Because the design lacked independent room-level replication, all room contrasts and p-values were interpreted as exploratory descriptive screening summaries rather than causal treatment effects. Airflow verification, playback calibration at multiple pen positions, and background-noise spectral analysis were performed to address measurement bias. Signal inspection showed that biologically relevant vocal energy was retained after 16 kHz resampling, while class imbalance was handled by inverse-frequency weighting and macro-F1-based model selection. The Audio Spectrogram Transformer (AST) pipeline was subjected to five-fold group-blocked cross-validation, and temporal validation. The model achieved a test macro-F1 of 0.937, five-fold macro-F1 of 0.928 ± 0.019, and three-day deployment validation macro-F1 of 0.914. In this two-room dataset, the ventilation-restricted room displayed higher room-level cough-like detections, aggressive vocalizations, normal vocalizations, lower silence, reduced growth, and poorer air quality. Cough-like detections showed recurring clock-time clustering, with the most sustained elevation during 19:00–22:00 and a smaller peak around 10:00 with the highest occurrences at 20.00 (2.84 room-level cough-like detections standardized to group size). Audio-only early-warning analysis flagged deteriorated air-quality windows with AUROC = 0.91 and AUPRC = 0.88 and provided a median 34 min lead time before environmental threshold exceedance, highlighting practical utility as an early inspection cue for farmers before air-quality deterioration becomes more pronounced. Cough-like events descriptively co-varied positively with NH3, temperature, and CO2. Overall, calibrated AST-based monitoring can summarize group-level acoustic changes associated with degraded room environments, while multi-room and multi-farm replication remains necessary for causal inference and generalization. Full article
(This article belongs to the Special Issue Application of Precision Farming in Pig Systems)
Show Figures

Figure 1

16 pages, 1346 KB  
Article
Occupant-Centred Acoustic Assessment of Teachers’ Responses to Sound Sources in a Secondary School
by Hang Fu, Jiayi Zhou, Jie Zhang, Rumei Han and Yuan Zhang
Buildings 2026, 16(14), 2894; https://doi.org/10.3390/buildings16142894 - 21 Jul 2026
Viewed by 368
Abstract
The acoustic quality of a school’s indoor environment affects the health, performance and comfort of the teachers who work in it, yet assessment often reduces that environment to a single overall level. In a cross-sectional case study of one secondary school in China, [...] Read more.
The acoustic quality of a school’s indoor environment affects the health, performance and comfort of the teachers who work in it, yet assessment often reduces that environment to a single overall level. In a cross-sectional case study of one secondary school in China, all 148 teachers rated eleven school sound-source categories on audibility, annoyance and interference with teaching and office concentration, and reported perceived noisiness, voice raising, voice strain, fatigue, emotional distress and noise sensitivity, while selected-room acoustic measurements provided site context not matched to individual respondents. Ratings across the eleven sources were dominated by one component, but the resulting composite score was not associated with all teacher responses alike. Source load was associated with emotional distress, voice raising and perceived noisiness but not fatigue, which instead tracked voice strain, whereas emotional distress was associated most strongly with noise sensitivity. In exploratory source-specific models, traffic and other external sources had the largest associations with perceived noisiness, whereas playground activity and student talking had the largest associations with voice raising. No tested subgroup interaction survived false-discovery-rate correction. The selected-room measurements documented conditions compatible with source intrusion and raised vocal effort. The findings support an occupant-centred, source-resolved approach to school indoor acoustics and provide hypotheses for future multi-school testing. Full article
(This article belongs to the Section Building Energy, Physics, Environment, and Systems)
Show Figures

Figure 1

24 pages, 16216 KB  
Article
A COMSOL–MATLAB Coupled Optimization Framework for Cost-Effective Acoustic Renovation of Educational Buildings Using Wood-Based Materials
by Shuang Yan, Liutao Zhang, Zhenbo Liu and Yuanyuan Miao
Buildings 2026, 16(13), 2676; https://doi.org/10.3390/buildings16132676 - 6 Jul 2026
Viewed by 307
Abstract
Poor classroom acoustic conditions can impair speech intelligibility, increase cognitive load, and reduce the quality of learning environments. Simulation-guided optimization provides a promising approach for improving building acoustic performance and indoor environmental quality while reducing trial-and-error material selection in renovation practice. The field-validated [...] Read more.
Poor classroom acoustic conditions can impair speech intelligibility, increase cognitive load, and reduce the quality of learning environments. Simulation-guided optimization provides a promising approach for improving building acoustic performance and indoor environmental quality while reducing trial-and-error material selection in renovation practice. The field-validated optimized configuration combined slotted wood sound-absorbing panels and mineral wool panels, with a total material cost of 1512 RMB. Field measurements showed that this configuration reduced RT from 2.42–1.76 s to 0.78–0.43 s. In addition, the simulation-based STI evaluation increased from 0.19 to 0.75, indicating a potential improvement in speech intelligibility. Since STI was not directly measured in the field, this result is interpreted as a model-based prediction supported by the calibrated RT validation. The average relative error between simulated and measured RT values was 6.83%, demonstrating the predictive reliability of the calibrated model for RT prediction. The proposed framework provides a practical decision-support method for cost-controlled acoustic renovation of educational buildings and was validated using an existing classroom case. The framework was demonstrated using one existing classroom and provides a methodological basis for adaptation to other educational spaces by updating room-specific inputs. Its external transferability requires validation in additional classrooms. Full article
(This article belongs to the Section Building Energy, Physics, Environment, and Systems)
Show Figures

Figure 1

20 pages, 2639 KB  
Article
Model-Informed Speech Enhancement Using Virtual Room Acoustics and Acoustic Descriptor Optimization
by Samuel Yaw Mensah, Tao Zhang, Xin Zhao and Nahid-Al Mahmud
Sensors 2026, 26(12), 3630; https://doi.org/10.3390/s26123630 - 6 Jun 2026
Viewed by 522
Abstract
Reverberation and background noise remain persistent obstacles to achieving clear and intelligible speech in enclosed environments. Conventional data-driven or purely empirical dereverberation systems often perform well only under training conditions but lack robustness and physical interpretability when exposed to new acoustic spaces. To [...] Read more.
Reverberation and background noise remain persistent obstacles to achieving clear and intelligible speech in enclosed environments. Conventional data-driven or purely empirical dereverberation systems often perform well only under training conditions but lack robustness and physical interpretability when exposed to new acoustic spaces. To address these limitations, this paper proposes a physics-informed speech enhancement algorithm that integrates analytical room acoustics modeling with a descriptor-guided optimization framework. The method employs virtual field simulations based on the Helmholtz equation to estimate key acoustic descriptors, reverberation time (RT60), direct-to-reverberant ratio (DRR), and clarity index (C50), which are then used to adaptively control a model-informed dereverberation filter. This hybrid formulation bridges physical modeling and signal processing, allowing the algorithm to minimize late reverberation energy while maintaining spectral fidelity. Experimental results across multiple simulated and real-room conditions demonstrate measurable improvements over baseline methods, achieving average gains of +6.4 dB in SNR, +1.2 in PESQ, and +0.13 in STOI, along with reduced RT60 and enhanced clarity. The proposed approach offers both computational efficiency and interpretability, making it suitable for real-time deployment in teleconferencing, hearing-assistive, and smart audio applications. Full article
Show Figures

Figure 1

29 pages, 8416 KB  
Article
Pilot Room-Level Acoustic and Physiological Monitoring of Respiratory Disturbance in Pigs Following Experimental Klebsiella pneumoniae Challenge
by Md Sharifuzzaman, Hong-Seok Mun, Eddiemar B. Lagua, Md Kamrul Hasan, Ahsan Mehtab, Jin-Gu Kang, Hae-Rang Park, Young-Hwa Kim and Chul-Ju Yang
Vet. Sci. 2026, 13(6), 550; https://doi.org/10.3390/vetsci13060550 - 3 Jun 2026
Viewed by 1010
Abstract
Respiratory disease remains a major challenge in pig production. This two-room pilot study evaluated whether room-level acoustic monitoring combined with physiological measurements could provide an early warning after an experimental Klebsiella pneumoniae challenge. Forty growing pigs balanced by sex and body weight were [...] Read more.
Respiratory disease remains a major challenge in pig production. This two-room pilot study evaluated whether room-level acoustic monitoring combined with physiological measurements could provide an early warning after an experimental Klebsiella pneumoniae challenge. Forty growing pigs balanced by sex and body weight were housed for 28 days in one control room and one challenged room (20 pigs/room; four pens/room). Challenged pigs were intranasally inoculated on days 8, 12, 16, and 20 with a culture whose dose was retrospectively verified by serial-dilution plating. Nasal and fecal samples were cultured on Klebsiella ChromoSelect agar, and colonies with expected morphology were enumerated as presumptive Klebsiella/K. pneumoniae colonies. A fine-tuned Audio Spectrogram Transformer (AST) classified five sound classes from facility-specific audio and was evaluated by group-blocked hold-out testing, five-fold group-blocked cross-validation, temporal deployment validation, and window-threshold sensitivity analysis. The model achieved hold-out macro-F1 of 0.947, five-fold macro-F1 of 0.928 ± 0.019, and 24 h deployment macro-F1 of 0.914. Presumptive nasal bacterial load was higher in challenged pigs at 1-week post-inoculation (log10 4.03 vs. 0.67). Group-size-standardized cough detections were also higher in the challenged room (54.84 vs. 36.80 detections/day), and daily coughing first exceeded the baseline threshold on day 8. Thresholds of 0.764 (control) and 1.115 (treatment) were obtained from an integrated score that included coughing, sneezing, ear temperatures, rectal temperature, and respiration rate; the treatment score and treatment–control contrast score first surpassed the threshold on day 8, and daily multimodal scores varied between groups (t = −6.636, p < 0.001). Integrated score improved discrimination of post-inoculation disturbance compared with cough detections alone (leave-one-day-out AUROC: 0.94 vs. 0.88). Because each condition was represented by one room, findings are exploratory temporal contrasts, not replicated treatment effects or a stand-alone diagnostic test. Full article
Show Figures

Graphical abstract

21 pages, 14302 KB  
Article
Audio-Based Device for Automated Surgical Counting, ToolSafe
by Michael R. Gardner, Latifa A. Aladdal, Lama Alshammari, Fatima Aldalgan, Maram A. Alomair, Shahad Alomair and Amani Alrashed
Appl. Sci. 2026, 16(11), 5181; https://doi.org/10.3390/app16115181 - 22 May 2026
Viewed by 416
Abstract
Manual counting of surgical tools, known as surgical counting, is a time-consuming and error-prone task that increases the risk of retained surgical instruments and extends operating room (OR) time. Presently, in hospitals around the world, surgical counting is often performed manually with paper [...] Read more.
Manual counting of surgical tools, known as surgical counting, is a time-consuming and error-prone task that increases the risk of retained surgical instruments and extends operating room (OR) time. Presently, in hospitals around the world, surgical counting is often performed manually with paper or tablet checklists, often leading to delays, increased infection risk, and financial cost. RFID, barcode-based, and computer vision solutions exist but are expensive and have challenges with sterilization and signal interference. This paper presents ToolSafe, a low-cost, portable system that classifies surgical tools by their acoustic signatures when dropped into a detection box. A pilot dataset of 4004 audio samples from four tool types (n = 996, tissue forceps; n = 1005, iris scissors; n = 1006, scalpel handle; n = 997, testing needle) was collected using ToolSafe. A convolutional neural network (CNN) was evaluated using stratified five-fold cross-validation on the laboratory dataset, with a k-nearest neighbors (KNN) classifier implemented as a control model. In each fold, both models were trained on 80% of the data and tested on the remaining 20%, ensuring that all samples were used for both training and evaluation. The CNN achieved a mean (±standard deviation) classification accuracy of 99.55% (±0.19%) across the validation folds, outperforming the KNN model, which achieved a mean accuracy of 97.28% (±0.50%). The difference was statistically significant according to a paired t-test across folds (p = 0.0003), indicating CNN’s superior performance on the dataset. For a run of 100 additional samples using the Raspberry Pi-based system, spectrogram generation averaged 0.121 s (±0.025 s), CNN inference averaged 0.180 s (±0.033 s), and total end-to-end latency averaged 1.851 s (±0.253 s) per tool. This pilot study proposes a possible technological solution for surgical counting that reduces human error and enhances patient safety. ToolSafe may be subsequently improved by increasing the number of surgical tools used in the training dataset, testing under more robust OR-like environments, and comparing to other classification algorithms. Further refinement and incorporation of ToolSafe in operating room workflows have the potential to reduce patient risks from extended surgical times and retained surgical instruments. Full article
Show Figures

Figure 1

22 pages, 4019 KB  
Article
DGSNA: Dynamic Generative Scene-Based Noise Addition Method
by Zihao Chen, Zhentao Lin, Bi Zeng, Linyi Huang and Jia Cai
Computation 2026, 14(5), 109; https://doi.org/10.3390/computation14050109 - 9 May 2026
Viewed by 388
Abstract
To ensure the reliable operation of speech systems across diverse environments, noise addition methods have emerged as the standard solution. However, existing methods offer limited coverage of real-world scenes and depend on pre-existing noise libraries and scene metadata. This paper presents prompt-based Dynamic [...] Read more.
To ensure the reliable operation of speech systems across diverse environments, noise addition methods have emerged as the standard solution. However, existing methods offer limited coverage of real-world scenes and depend on pre-existing noise libraries and scene metadata. This paper presents prompt-based Dynamic Generative Scene-based Noise Addition (DGSNA), a novel approach driven by generative language models that integrates Dynamic Generation of Scene-based Information (DGSI) with Scene-based Noise Addition for Speech (SNAS). The DGSI module, with a BET (Background, Examples, Task) prompt framework, dynamically generates logic-compliant scene-based information, including scene dimensions, sound sources, and microphone positions, thereby addressing the challenges of scene enumeration and detailed description. Complementing this, the SNAS module employs a Time–Frequency Diffusion-based (TFD) Text-to-Audio model to synthesize scene-specific noise. By integrating this noise with clean speech via Room Impulse Response (RIR) filters, the module streamlines the traditionally labor-intensive process of replicating diverse acoustic environments. Experimental results show that DGSNA significantly enhances the robustness of speech recognition and keyword spotting models, achieving relative improvements of up to 11.32%. Furthermore, DGSNA is highly compatible with existing noise addition techniques. Full article
(This article belongs to the Section Computational Engineering)
Show Figures

Figure 1

22 pages, 33241 KB  
Article
Eigenbeam–vMF-Based Room Acoustic Analyzer: A Comparative Study with First-Order and Higher-Order Ambisonic Recordings
by Amy Bastine, Thushara D. Abhayapala and Jihui (Aimee) Zhang
Appl. Sci. 2026, 16(9), 4470; https://doi.org/10.3390/app16094470 - 2 May 2026
Viewed by 637
Abstract
Comprehensive room acoustic characterization requires resolving reflection behavior across time, frequency, and space. The recently proposed eigenbeam–vMF-based analyzer provides a framework for this by modeling the reflection field as a time–frequency-dependent directional power distribution, estimated via spatial correlation of eigenbeams (ambisonics) and parameterized [...] Read more.
Comprehensive room acoustic characterization requires resolving reflection behavior across time, frequency, and space. The recently proposed eigenbeam–vMF-based analyzer provides a framework for this by modeling the reflection field as a time–frequency-dependent directional power distribution, estimated via spatial correlation of eigenbeams (ambisonics) and parameterized using von Mises–Fisher clustering. This formulation enables a unified and interpretable description of anisotropic early reflections, their transition into diffuse reverberation, and frequency-dependent acoustic behavior. Prior work showed that the analyzer reliably captures these features using higher-order ambisonics from a 32-channel spherical microphone array (SMA) and that constraining the same array to the first order still led to retaining the dominant features. This paper investigates whether this capability extends to first-order microphone arrays with sparser spatial sampling for more economical and practical deployment. A comparative study is conducted in a recording studio with variable wall panels (wood and felt), evaluating a four-channel first-order array against a 32-channel SMA. The results reveal distinct acoustic differences between panel settings, which are consistent across both arrays. While the SMA captures finer spatial detail and prolonged anisotropic reflections more effectively, the first-order array demonstrates potential for preliminary room acoustic assessments by identifying room mode frequencies, dominant reflection directions, and highly reflective surfaces. Full article
(This article belongs to the Special Issue Architectural Acoustics: From Theory to Application—2nd Edition)
Show Figures

Figure 1

19 pages, 7601 KB  
Article
On the Reflection of a Spherical Sound Wave from a Finite Size Surface
by Jens Holger Rindel
Appl. Sci. 2026, 16(9), 4243; https://doi.org/10.3390/app16094243 - 26 Apr 2026
Viewed by 438
Abstract
Room acoustics computer models based on geometrical acoustics usually handle the sound reflections by the assumption of plane waves. However, if the sound source is a point source, which is usually the case, the spherical wave reflection would be more correct. An approximate [...] Read more.
Room acoustics computer models based on geometrical acoustics usually handle the sound reflections by the assumption of plane waves. However, if the sound source is a point source, which is usually the case, the spherical wave reflection would be more correct. An approximate model for the spherical wave reflection is presented, starting with the assumption of an infinite plane. It was found that the errors caused due to the simplified plane wave assumption can be significant, especially for hard surfaces and near grazing incidence. As something new, the gradual transition from a spherical wave to a plane wave approximation was addressed. For sound propagation exceeding 50 times the wavelength, the plane wave approximation was found to be fully justified, but for shorter distances the spherical wave reflection model should be applied. In contrast to previous work on spherical wave reflection, the reflection from a finite-sized surface was studied. For the first time, the spherical wave reflection model was combined with the complex radiation impedance of a finite-sized surface. One interesting application example of the spherical reflection model is the attenuation of sound propagation above the audience area in a performance space. Finally, the extension of the spherical wave reflection model to higher order reflections was addressed. Full article
(This article belongs to the Special Issue Architectural Acoustics: From Theory to Application—2nd Edition)
Show Figures

Figure 1

24 pages, 2467 KB  
Article
Comparative Development of Machine Learning Models for Short-Term Indoor CO2 Forecasting Using Low-Cost IoT Sensors: A Case Study in a University Smart Laboratory
by Zhanel Baigarayeva, Assiya Boltaboyeva, Zhuldyz Kalpeyeva, Raissa Uskenbayeva, Maksat Turmakhan, Adilet Kakharov, Aizhan Anartayeva and Aiman Moldagulova
Algorithms 2026, 19(5), 328; https://doi.org/10.3390/a19050328 - 24 Apr 2026
Viewed by 745
Abstract
Unlike reactive systems, mechanical ventilation controlled by CO2 concentration operates at a target efficiency that dynamically increases whenever the target CO2 level is exceeded. This approach eliminates the typical ‘dead-time’ and prevents air quality degradation by ensuring the system adjusts its [...] Read more.
Unlike reactive systems, mechanical ventilation controlled by CO2 concentration operates at a target efficiency that dynamically increases whenever the target CO2 level is exceeded. This approach eliminates the typical ‘dead-time’ and prevents air quality degradation by ensuring the system adjusts its performance immediately in response to concentration changes. In this work, the study focuses on the development and evaluation of data-driven predictive models for near-term indoor CO2 forecasting that can be integrated into pre-occupancy ventilation strategies, rather than designing a complete control scheme. Experimental data were collected over four months in a 48 m2 smart laboratory configured as an open-plan office, where a heterogeneous IoT sensing architecture logged synchronized time-series measurements of CO2 and microclimate variables (temperature, relative humidity, PM2.5, TVOCs), together with acoustic noise levels and appliance-level energy consumption used as indirect occupancy-related signals. Raw telemetry was transformed into a 22-feature state vector using a structured feature engineering method incorporating z-score standardization, cyclic time encodings, multi-horizon CO2 lags, rolling statistics, momentum features, and non-linear interactions to represent temporal autocorrelation and daily periodicity. The study benchmarks multiple regression paradigms, including simple baselines and ensemble methods, and found that an automated multi-level stacked ensemble achieved the highest predictive fidelity for short-term forecasting, with an Mean Absolute Error (MAE) of 32.97 ppm across an observed CO2 range of 403–2305 ppm, representing improvements of approximately 24% and 43% over Linear Regression and K-Nearest Neighbors (KNN), respectively. Temporal diagnostics showed strong phase alignment with observed CO2 rises during occupancy transitions and statistically reliable prediction intervals. Five-fold walk-forward cross-validation confirmed the temporal stability of these results, with top models achieving consistent R2 values of 0.93–0.95 across Folds 2–5. These results demonstrate that, within a single-room university laboratory setting, historical sensor data from low-cost IoT devices can support accurate short-term CO2 forecasting, providing a predictive layer that could support future proactive ventilation scheduling aimed at reducing CO2 lag at the start of occupancy while avoiding unnecessary ventilation runtime. Generalization to other building types and occupancy profiles requires further validation. Full article
(This article belongs to the Special Issue Emerging Trends in Distributed AI for Smart Environments)
Show Figures

Figure 1

36 pages, 6746 KB  
Article
An Archaeoacoustic Analysis of a Single-Nave Hall in the Cellars of Diocletian’s Palace in Split, Croatia
by Mateja Nosil Mešić, Marko Horvat and Zoran Veršić
Acoustics 2026, 8(2), 26; https://doi.org/10.3390/acoustics8020026 - 20 Apr 2026
Viewed by 833
Abstract
Diocletian’s palace with its cellars represents one of the most important cultural heritage sites of the ancient Roman civilisation on the present-day Croatian territory. The cellar complex has been rediscovered only recently and has been preserved remarkably well due to its centuries-long concealment [...] Read more.
Diocletian’s palace with its cellars represents one of the most important cultural heritage sites of the ancient Roman civilisation on the present-day Croatian territory. The cellar complex has been rediscovered only recently and has been preserved remarkably well due to its centuries-long concealment beneath mediaeval urban matrices. An archaeoacoustic analysis was performed on a selected single-nave hall as a small part of this complex. A model of the hall was developed in room acoustics simulation software and calibrated based on the results of field measurements. Acoustic suitability of the hall for speech-based events and music performances was then evaluated according to contemporary objective criteria, and the findings were compared with the results of similar studies performed on other heritage sites. The hall was found to be very well suited for speech in terms of intelligibility and mid-frequency reverberation, thus showing potential for revitalisation, with excessive low-frequency reverberation in the hall and reduced audibility in the farthest part of the audience as potential issues. With a feasible audience size, the hall is not reverberant enough for music performances but provides high clarity. In terms of sound strength, the hall is suitable for solo performers or small ensembles. Excessive perceptive broadening of the sound source is expected due to strong early lateral energy. In terms of traditional Dalmatian a cappella singing, the acoustics of the hall are likely to support and enhance such performances. Full article
(This article belongs to the Collection Historical Acoustics)
Show Figures

Figure 1

20 pages, 3276 KB  
Article
Reaction Time to Amplitude-Modulated Tones Under Spectral Masking: Implications for Architectural Acoustic Design
by Ryota Shimokura and Yoshiharu Soeta
Appl. Sci. 2026, 16(8), 3814; https://doi.org/10.3390/app16083814 - 14 Apr 2026
Viewed by 648
Abstract
Detectability of auditory signals in built environments is a critical issue in architectural acoustics, particularly in public spaces where notification sounds must be perceived reliably under background noise. This study investigated reaction times (RTs) to amplitude-modulated pure tones under silent, white noise, and [...] Read more.
Detectability of auditory signals in built environments is a critical issue in architectural acoustics, particularly in public spaces where notification sounds must be perceived reliably under background noise. This study investigated reaction times (RTs) to amplitude-modulated pure tones under silent, white noise, and bandpass-noise conditions. Twenty young and twenty elderly participants responded to 1 and 2 kHz tones with flat, gentle, and steep onset envelopes. To describe perceptual detection in physically interpretable terms, a time-integrated sound-exposure level model, LAE(t), was applied. RT was defined as the moment when cumulative acoustic energy exceeded a criterion value relative to the hearing threshold. In silent conditions, RTs were accurately predicted by LAE(t), with onset-envelope shape influencing early energy accumulation. In noise conditions, RTs increased systematically with spectral proximity between target and masker, consistent with auditory filter theory. When spectral separation exceeded approximately four ERB numbers, masking effects were minimal, and RT approached silent-condition values. These findings demonstrate that perceptual detection timing is governed by cumulative acoustic energy and spectral masking rather than instantaneous sound pressure level. The LAE(t) model provides a detection-oriented metric that complements conventional room-acoustic parameters and may support evidence-based design of perceptually robust auditory signals in architectural environments. Full article
(This article belongs to the Special Issue Architectural Acoustics: From Theory to Application—2nd Edition)
Show Figures

Figure 1

Back to TopTop