Physics-Aware Generative Demasking: Spatially Conditioned Diffusion for Robust Transient Detection in Industrial Noise
Abstract
1. Introduction
- We propose a physics-aware acquisition framework utilizing a wearable dual-sensor set-up. By contrasting audio characteristics captured at the source (glove) and the periphery (chest/desk), we construct a hard spatial constraint that effectively disentangles the target transient signatures from the complex background noise manifold, utilizing the ambient signal as a dynamic reference for the environmental noise floor.
- We develop a spatially conditioned diffusion probabilistic model (SC-DPM) for signal enhancement. By injecting the ambient noise reference as a dynamic guidance condition into the reverse diffusion process, the model learns to physically demask the target signal. Unlike traditional subtractive methods, this approach leverages generative modeling to restore high-fidelity morphological details of the connector insertion sound from heavily corrupted mixtures.
- We introduce an efficient feature extraction mechanism based on causal random convolutional kernels. By employing causal dilations and local proportion of positive values (LPPV) pooling, this method captures the transient morphological features of the enhanced “click” sound without computational overhead, facilitating rapid and robust detection suitable for real-time industrial cycles.
2. Related Work
2.1. Time Series Classification with Random Convolution Kernels
2.2. Denoising Diffusion Models and Physics Awareness
3. Methods
3.1. System Overview
- Physics-Aware Acquisition: A dual-sensor configuration is employed to simultaneously capture a source-proximal acoustic signal and an ambient noise reference, thereby imposing a physical spatial constraint on the observed mixture.
- Generative Demasking (SC-DPM): A spatially conditioned diffusion probabilistic model (SC-DPM) leverages the ambient reference as an explicit conditioning variable to disentangle and reconstruct the clean insertion transient from the noisy observation.
- Causal Feature Extraction: The reconstructed waveforms are transformed into discriminative representations using causal random convolutional kernels, enabling temporally precise and computationally efficient classification.
3.2. Physics-Aware Acquisition Strategy
- Source-Proximal Sensor (): This is mounted on the operator’s glove at a distance of approximately from the connector. This sensor captures the high-energy insertion transient while remaining contaminated by local environmental noise.
- Ambient Reference Sensor (): This is positioned on the operator’s chest or a nearby workbench at a distance exceeding . Due to distance attenuation and partial body shadowing, this sensor predominantly records environmental noise, with negligible contribution from the insertion transient.
3.3. Spatially Conditioned Diffusion Probabilistic Model (SC-DPM)
3.3.1. Forward Diffusion Process
3.3.2. Conditional Reverse Process
3.3.3. Training Objective
3.4. Efficient Feature Extraction via Causal Random Kernels
3.4.1. Causal Dilated Convolution
3.4.2. Discrete Kernel Weights
3.4.3. Local Proportion of Positive Values
4. Data Preparation and Experiments
4.1. Data Acquisition System
- Source-Proximal Sensor (): A high-fidelity electret microphone is integrated into an industrial safety glove, positioned approximately from the operator’s fingertips (Figure 2a). This sensor is configured to capture the high-intensity near-field transient signature of connector insertions.
- Ambient Reference Sensor (): An auxiliary microphone is deployed in the far field to capture the environmental noise profile correlated with the proximal sensor, providing the necessary spatial condition for the SC-DPM.
- Controlled Baseline (Lab): An anechoic setting was used to acquire high-quality ground truth signals of connector insertions, serving as the clean reference.
- Unseen Non-Stationary Environments (OOD): Recordings were conducted in outdoor construction sites and windy areas. These scenarios introduced irregular, impulsive interference (e.g., impact sounds and wind gusts) to test the model’s generalization capability against out-of-distribution (OOD) noise.
- Target Operational Environment (Workshop): Field recordings were acquired from an active automotive assembly line. This environment featured continuous, broadband background noise generated by pneumatic tools, automated conveyors, and overlapping human speech.
4.2. LLM-Assisted Automated Annotation Pipeline
- Stage I: Voice-Anchored Temporal Localization. To overcome the limitations of traditional energy-based detection in high-noise environments, we employed the pretrained Qwen-Audio model [10]. As a state-of-the-art large audio model (LAM), Qwen-Audio demonstrates exceptional zero-shot generalization in speech understanding amidst noise. We fine-tuned the model for keyword spotting (KWS) to robustly detect specific vocal triggers (e.g., shouting “One”) and pinpoint the timestamp within the proximal stream.
- Stage II: Synchronized Retroactive Extraction. Leveraging the causal relationship between the physical action and the subsequent vocalization, the system extracts a fixed temporal window across all synchronized channels. This window is empirically calibrated to encompass the complete action sequence—approach, insertion click, and release—while excluding the vocal cue itself.
4.3. Dataset Organization and Spectral Characteristics
- OOD-Driven Training Set (Training): This dataset comprised recordings acquired in diverse non-workshop environments (e.g., construction sites). By training on these harsh, irregular noise conditions, we forced the SC-DPM to learn generalized physical demasking rules rather than overfitting to specific workshop frequency patterns. This strategy enhanced robustness against out-of-distribution (OOD) acoustic shifts.
- In-Domain Evaluation Set (Testing): This dataset consisted exclusively of recordings collected from the actual automotive assembly line. It represents the target domain, characterized by stationary mechanical hum and pneumatic tool transients and serving as the rigorous standard for evaluating the system’s practical performance.
5. Results and Discussion
5.1. Signal Reconstruction Analysis
5.1.1. Proximal Signal Characteristics
5.1.2. Ambient Reference Signal Analysis
5.1.3. Enhanced Signal via Spatially Conditioned Diffusion
- Input Noisy Signals (Figure 6a,d): Both examples demonstrate that the target transient was almost entirely submerged in non-stationary industrial noise, rendering discriminative detection infeasible.
- Ambient Reference Signals (Figure 6b,e): The reference channel captured consistent background noise patterns without introducing target-related artifacts, thereby serving as a reliable “negative” template.
- Enhanced Outputs (Figure 6c,f): The proposed method successfully reconstructed the transient insertion signatures in both cases, yielding clean and temporally localized events that closely resemble ideal baseline signals.
5.2. Classification Accuracy
- Susceptibility to Domain Shift: In relatively controlled noise conditions, all learning-based methods achieved reasonable performance. However, in the challenging real factory environment, the accuracy of the baseline CRNN model degraded significantly to 69.4%, highlighting the severe impact of non-stationary industrial noise on standard deep learning architectures.
- Baseline Performance: The SIMPF method [30], which utilizes simple pooling front-ends, demonstrated moderate robustness with an accuracy of 83.4%. Notably, our framework’s classification back-end alone (ablation) achieved only 76.8%, indicating that lightweight classifiers are insufficient to handle heavy industrial interference without proper signal enhancement.
- Decisive Impact of Spatial Conditioning: The integration of the SC-DPM mechanism yielded a substantial performance leap. Specifically, when comparing the full model with the ablation version, the accuracy improved by approximately 16.5 percentage points (93.3–76.8%) in the factory environment. Furthermore, the proposed method outperformed the CRNN baseline by 23.9 percentage points, validating that explicitly modeling spatial noise correlation significantly enhances the discriminability between true insertion events and background disturbances.
5.3. Computational Cost and Efficiency Analysis
- Marginal Overhead: Introducing the SC-DPM module increased the training time only slightly from 115.2 min (ablation) to 120.4 min (full), corresponding to a negligible 4.5% overhead.
- Significant Gain: This small additional cost yielded a massive return in performance, achieving the state-of-the-art accuracy of 93.3%. Compared with the ablated model, the accuracy gain exceeded 16 percentage points while maintaining nearly identical computational cost.
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Rubin, H.-D.; Pascucci, V.C.; Toran, J.; Druckenmiller, R.; Lipschutz, M.; Conde, P.; Vasudevan, V.; Gupta, J.; Oon, Y.-H.; Han, C. A standardized reliability evaluation framework for connectors—Stress levels and test recommendations. In Proceedings of the 2019 IEEE Holm Conference on Electrical Contacts, Milwaukee, WI, USA, 14–18 September 2019; pp. 324–334. [Google Scholar]
- Hershey, S.; Chaudhuri, S.; Ellis, D.P.W.; Gemmeke, J.F.; Jansen, A.; Moore, R.C.; Plakal, M.; Platt, D.; Saurous, R.A.; Seybold, B.; et al. CNN architectures for large-scale audio classification. In Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA, 5–9 March 2017; pp. 131–135. [Google Scholar]
- Cakir, E.; Parascandolo, G.; Heittola, T.; Huttunen, H.; Virtanen, T. Convolutional recurrent neural networks for polyphonic sound event detection. IEEE/ACM Trans. Audio Speech Lang. Process. 2017, 25, 1291–1303. [Google Scholar] [CrossRef] [Scilit]
- Kong, Q.; Cao, Y.; Iqbal, T.; Wang, Y.; Wang, W.; Plumbley, M.D. PANNs: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Trans. Audio Speech Lang. Process. 2020, 28, 2880–2894. [Google Scholar] [CrossRef] [Scilit]
- Chis, T.V.; Cioca, L.-I.; Badea, D.O.; Cristea, I.; Darabont, D.C.; Iordache, R.M.; Platon, S.N.; Trifu, A.; Barsan, V.-A. Integrated Noise Management Strategies in Industrial Environments: A Framework for Occupational Safety, Health, and Productivity. Sustainability 2025, 17, 1181. [Google Scholar] [CrossRef] [Scilit]
- Nam, H.; Kim, S.-H.; Ko, B.-Y.; Park, Y.-H. Frequency dynamic convolution: Frequency-adaptive pattern recognition for sound event detection. arXiv 2022, arXiv:2203.15296. [Google Scholar] [CrossRef] [Scilit]
- Yue, H.; Zhang, Z.; Mu, D.; Dang, Y.; Yin, J.; Tang, J. Full-frequency dynamic convolution: A physical frequency-dependent convolution for sound event detection. arXiv 2024, arXiv:2401.04976. [Google Scholar]
- Nam, H.; Kim, S.-H.; Park, Y.-H. FilterAugment: An acoustic environmental data augmentation method. In Proceedings of the 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, 23–27 May 2022; pp. 4308–4312. [Google Scholar]
- Xiao, Y.; Das, R.K. WildDESED: An LLM-powered dataset for wild domestic environment sound event detection system. arXiv 2024, arXiv:2407.03656. [Google Scholar]
- Chu, Y.; Xu, J.; Zhou, X.; Yang, C.; Zhang, S.; Yan, Z.; Chang, C.; Zhou, J. Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models. arXiv 2023, arXiv:2311.07919. [Google Scholar]
- Liu, X.; Liu, H.; Kong, Q.; Mei, X.; Zhao, J.; Huang, Q.; Plumbley, M.D.; Wang, W. Separate What You Describe: Language-Queried Audio Source Separation. arXiv 2022, arXiv:2203.15147. [Google Scholar] [CrossRef] [Scilit]
- Yin, H.; Xiao, Y.; Bai, J.; Das, R.K. Leveraging LLM and text-queried separation for noise-robust sound event detection. arXiv 2024, arXiv:2411.01174. [Google Scholar]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, 33, 6840–6851. [Google Scholar]
- Nichol, A.Q.; Dhariwal, P. Improved denoising diffusion probabilistic models. In Proceedings of the 38th International Conference on Machine Learning (ICML), Virtual Event, 18–24 July 2021; pp. 8162–8171. [Google Scholar]
- Turner, R.E.; Diaconu, C.-D.; Markou, S.; Shysheya, A.; Foong, A.Y.K.; Mlodozeniec, B. Denoising diffusion probabilistic models in six simple steps. arXiv 2024, arXiv:2402.04384. [Google Scholar] [CrossRef] [Scilit]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. Adv. Neural Inf. Process. Syst. 2014, 27, 2672–2680. [Google Scholar]
- Pascual, S.; Bonafonte, A.; Serrà, J. SEGAN: Speech enhancement generative adversarial network. arXiv 2017, arXiv:1703.09452. [Google Scholar] [CrossRef] [Scilit]
- Kong, Z.; Ping, W.; Huang, J.; Zhao, K.; Catanzaro, B. DiffWave: A versatile diffusion model for audio synthesis. arXiv 2021, arXiv:2009.09761. [Google Scholar] [CrossRef] [Scilit]
- Chen, N.; Zhang, Y.; Zen, H.; Weiss, R.J.; Norouzi, M.; Chan, W. WaveGrad: Estimating Gradients for Waveform Generation. arXiv 2021, arXiv:2009.00713. [Google Scholar]
- Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; Plumbley, M.D. AudioLDM: Text-to-Audio Generation with Latent Diffusion Models. arXiv 2023, arXiv:2301.12503. [Google Scholar]
- Ye, T.; Dong, L.; Xia, Y.; Sun, Y.; Zhu, Y.; Huang, G.; Wei, F. Differential transformer. arXiv 2024, arXiv:2410.05258. [Google Scholar]
- Ni, W.; Zhang, C.; Liu, T.; Zeng, Q.; Xu, L.; Wang, H. An efficient astronomical seeing forecasting method by random convolutional kernel transformation. Eng. Appl. Artif. Intell. 2024, 127, 107259. [Google Scholar] [CrossRef] [Scilit]
- Dempster, A.; Petitjean, F.; Webb, G.I. ROCKET: Exceptionally fast and accurate time series classification using random convolutional kernels. Data Min. Knowl. Discov. 2020, 34, 1454–1495. [Google Scholar] [CrossRef] [Scilit]
- Dempster, A.; Schmidt, D.F.; Webb, G.I. MiniRocket: A very fast (almost) deterministic transform for time series classification. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Virtual Event, 14–18 August 2021; pp. 248–257. [Google Scholar]
- Tan, C.W.; Dempster, A.; Bergmeir, C.; Webb, G.I. MultiRocket: Multiple pooling operators and transformations for fast and effective time series classification. Data Min. Knowl. Discov. 2022, 36, 1623–1646. [Google Scholar] [CrossRef] [Scilit]
- Salehinejad, H.; Wang, Y.; Yu, Y.; Jin, T.; Valaee, S. S-ROCKET: Selective random convolution kernels for time series classification. arXiv 2022, arXiv:2203.03445. [Google Scholar]
- Chen, S.; Sun, W.; Huang, L.; Li, X.; Wang, Q.; John, D. P-ROCKET: Pruning random convolution kernels for time series classification. arXiv 2023, arXiv:2309.08499. [Google Scholar]
- Uribarri, G.; Barone, F.; Ansuini, A.; Fransén, E. Detach-ROCKET: Sequential feature selection for time series classification with random convolutional kernels. Data Min. Knowl. Discov. 2024, 38, 3922–3947. [Google Scholar] [CrossRef] [Scilit]
- Mansour Lo, M.; Morvan, G.; Rossi, M.; Morganti, F.; Mercier, D. Time series classification with random convolution kernel-based transforms: Pooling operators and input representations matter. arXiv 2024, arXiv:2409.08137. [Google Scholar]
- Liu, X.; Liu, H.; Kong, Q.; Mei, X.; Plumbley, M.D.; Wang, W. Simple Pooling Front-Ends for Efficient Audio Classification. arXiv 2022, arXiv:2210.00943. [Google Scholar]
- Wu, D.; Cao, H.; Lv, N.; Fan, J.; Tan, X.; Yang, S. Feature Matching Conditional GAN for Fast Radio Burst Localization with Cluster-fed Telescope. Astrophys. J. Lett. 2019, 887, L10. [Google Scholar] [CrossRef] [Scilit]
- Lu, Y.J.; Wang, Z.Q.; Watanabe, S.; Richard, A.; Yu, C.; Tsao, Y. Conditional diffusion probabilistic model for speech enhancement. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, 23–27 May 2022; pp. 7402–7406. [Google Scholar]
- Zhang, Z.; Cao, H.; Fan, J.; Peng, J.; Liu, S. Active RIS-Assisted Sup-Degree of Freedom Interference Suppression for a Large Antenna Array: A Deep-Learning Approach With Location Awareness. IEEE Trans. Antennas Propag. 2024, 72, 628–641. [Google Scholar] [CrossRef] [Scilit]
- Peng, J.; Cao, H.; Fan, J.; Zhang, Z.; Wu, D. Active Reconfigurable Intelligent Surface-Assisted Mainlobe Wideband RFI Mitigation With Deep Reinforcement Learning for a Large Reflector Antenna. IEEE Trans. Geosci. Remote. Sens. 2024, 62, 2003415. [Google Scholar] [CrossRef] [Scilit]
- Shi, Z.; Zheng, H.; Xu, C.; Dong, C.; Pan, B.; Xie, X.; He, A.; Li, T.; Fu, H. Resfusion: Denoising diffusion probabilistic models for image restoration based on prior residual noise. Adv. Neural Inf. Process. Syst. (NeurIPS) 2024, 37, 130664–130693. [Google Scholar]
- Su, K.; Cui, K.; Patel, R.; Rumbo, R.; Wang, X. Physics-driven diffusion models for impact sound synthesis from videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 9749–9759. [Google Scholar]
- Wang, Y.; Liu, Z.; Zhang, L. DiffPhysiNet: A bearing diagnostic framework based on physics-driven diffusion network for unseen working conditions. PHM Soc. Eur. Conf. 2024, 8, 1–10. [Google Scholar]
- Ding, Y.; Zhang, W.; Zhao, X. Denoising diffusion implicit model for bearing fault diagnosis under different working loads. ITM Web Conf. 2024, 63, 01025. [Google Scholar] [CrossRef] [Scilit]
- Xia, Z.; Luo, Z.; Chen, C.H.; Shen, X. An Effective Photoplethysmography Denosing Method Based on Diffusion Probabilistic Model. IEEE J. Biomed. Health Inform. 2025, 29, 4071–4080. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Liu, X.; Plumbley, M.D.; Wang, W. SoloAudio: Target sound extraction with language-oriented audio diffusion transformer. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 6–11 April 2025. [Google Scholar]








| Model | OOD-Synthesized Noise | Target Factory Domain |
|---|---|---|
| CRNN Baseline | 81.9% | 69.4% |
| SIMPF [30] | 88.1% | 83.4% |
| Ours (Ablation: No SC-DPM) | 85.3% | 76.8% |
| Ours (Full: with SC-DPM) | 94.8% | 93.3% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cao, H.; Lv, Z.; Hu, J.; Wang, H.; Yang, L.; Zhang, G. Physics-Aware Generative Demasking: Spatially Conditioned Diffusion for Robust Transient Detection in Industrial Noise. Entropy 2026, 28, 364. https://doi.org/10.3390/e28040364
Cao H, Lv Z, Hu J, Wang H, Yang L, Zhang G. Physics-Aware Generative Demasking: Spatially Conditioned Diffusion for Robust Transient Detection in Industrial Noise. Entropy. 2026; 28(4):364. https://doi.org/10.3390/e28040364
Chicago/Turabian StyleCao, Hailin, Zixi Lv, Jinjie Hu, Hui Wang, Lisheng Yang, and Guoxin Zhang. 2026. "Physics-Aware Generative Demasking: Spatially Conditioned Diffusion for Robust Transient Detection in Industrial Noise" Entropy 28, no. 4: 364. https://doi.org/10.3390/e28040364
APA StyleCao, H., Lv, Z., Hu, J., Wang, H., Yang, L., & Zhang, G. (2026). Physics-Aware Generative Demasking: Spatially Conditioned Diffusion for Robust Transient Detection in Industrial Noise. Entropy, 28(4), 364. https://doi.org/10.3390/e28040364

