AMUSE++: A Mamba-Enhanced Speech Enhancement Framework with Bi-Directional and Advanced Front-End Modeling
Abstract
1. Introduction
- A transition from magnitude-only, frequency-domain approaches to deep learning models operating in time-frequency and time-domain representations.
- Increased focus on explicit phase modeling to achieve perceptual quality and intelligibility improvements.
- Development of lightweight architectures optimized for edge devices and mobile deployment.
- Growing use of multi-modal integration and cross-sensor fusion to handle real-world acoustic challenges.
- We propose a 2D bi-directional Mamba module that explicitly models spectro-temporal dependencies along both time and frequency axes and show that it significantly outperforms 1D uni-directional Mamba module in speech enhancement.
- We design a preliminary denoising module that applies the proposed 2D bi-directional Mamba in a front-end stage and introduce an auxiliary enhancement branch that provides direct supervision to this module via the MUSE++ multi-objective loss.
- Through extensive experiments and ablation studies on the VoiceBank-DEMAND corpus, we demonstrate that AMUSE++ achieves superior perceptual quality and intelligibility compared with MUSE, MUSE++, and several state-of-the-art lightweight SE models, while maintaining a compact parameter footprint
2. Backbone Model: MUSE++ and Its 1D Uni-Directional Mamba Block
- 1.
- Efficient Sequence Modeling with Mamba-2: The MET Transformer in MUSE is replaced by the 1D Mamba-2 state space model, which maintains linear complexity with respect to sequence length. This facilitates both a significant reduction in parameter count and computational cost, while enabling effective long-range dependency modeling in sequential audio data.
- 2.
- Dynamic SNR-based Data Augmentation: Instead of fixed-SNR mixing, MUSE++ performs dynamic, on-the-fly mixing for every training utterance by sampling SNR values uniformly from a broad range, generating noisy-clean pairs on demand. This dramatically increases data diversity and improves noise robustness.
- 3.
- Augmented Multi-Objective Loss: The objective is expanded to include STFT consistency, time-domain, and multi-resolution STFT losses in addition to standard spectrogram-based terms, yielding richer and more comprehensive supervision for optimization.
2.1. Input Representation and Data Flow
2.2. 1D Mamba (Mamba-2): The Backbone Sequence Model
2.3. Dense Convolution Codec
2.4. Employed Loss Function
2.5. Limitations of the MUSE++ Backbone and Motivation for Further Development
3. Proposed Method: AMUSE++
3.1. Advanced Mamba Module
3.1.1. Transition from Uni-Directional to Bi-Directional Mamba
3.1.2. Extension from 1D to 2D Mamba for Time–Frequency Modeling
- Temporal branch (along time):where the reshape operation converts all time trajectories into a batch of 1D sequences of length T, so that a single Mambatime block can be shared across all frequency bins while keeping each trajectory independent.
- Frequency branch (along frequency):where each of the frequency trajectories becomes a 1D sequence of length F in the batch, processed by a single shared Mambafreq block.
3.1.3. Integration: 2D Bi-Directional Mamba Module
3.2. The AMUSE++ Framework
- Replacement of 1D uni-directional Mamba blocks: All 1D uni-directional Mamba (Mamba-2) modules in the original MUSE++ backbone are replaced by the advanced Mamba module. This allows each sequence modeling block to capture dependencies along both the time and frequency axes, in both forward and backward directions, providing richer contextual representations than purely 1D, uni-directional processing.
- Addition of a front-end Preliminary Denoising Module: A stack of advanced Mamba modules is introduced as a dedicated Preliminary Denoising Module (PDM) at the front end of AMUSE++. This module performs an initial enhancement of the input speech features before they are passed to the main backbone for further refinement.
- 1.
- Input projection: Given an input tensorwhere B is the batch size, C the input channel dimension, and T and F the numbers of time frames and frequency bins, respectively, a Conv2D layer projects into the internal model with a new channel dimension while preserving the time–frequency resolution:
- 2.
- Stack of 2D bi-directional Mamba blocks: The projected features are then refined by a stack of N Mamba blocks:where each implements the proposed 2D bi-directional Mamba over the time–frequency plane.
- 3.
- Output projection: The output of the last Mamba block is mapped back to the original channel dimension C using another Conv2D layer:where is configured such that .
- 4.
- Residual connection and preliminary enhancement. Finally, a residual connection is applied by adding the original input:The resulting tensor is interpreted as a preliminarily denoised time–frequency representation and is fed into the subsequent AMUSE++ backbone for further enhancement and decoding.
4. Experimental Setup
- Data preprocessing and training protocol: All waveforms are normalized to zero mean and unit variance before being fed into the model. Signals are uniformly segmented into chunks of 30,700 samples. The STFT uses an FFT size of 510, window length 510, hop size 100, and a sampling rate of 16 kHz. Models are trained for up to 100 epochs using the AdamW optimizer with an initial learning rate of 0.0005, an exponential decay factor of 0.99, weight decay of , and a batch size of 2. Early stopping is applied if the validation loss does not improve for 10 consecutive epochs.
- Model architecture configuration: AMUSE++ adopts a three-level U-Net encoder–decoder architecture with the dense channel dimension initialized at 16 and doubled at each downsampling stage. Each Mamba module in the U-Net is configured with state size , convolution width , and expansion factor . For the PDM module, both the input channel dimension C and the Conv2D-processed channel dimension are set to 16, the number of Mamba blocks , and the Conv2D layers use a kernel size of (i.e., ).
- Loss function configuration: For the multi-resolution STFT loss () in Equation (7), we employ three STFT settings: [FFT size, window length, hop size] = [510, 510, 100], [800, 800, 200], [320, 320, 80]. The loss function weights in Equations (7) and (8) are set as: , , , , , , and . These hyperparameters follow prior work on MUSE++ [20], MUSE [7] and MP-SENet [13,14], the latter serving as the main baseline during the development of MUSE. Finally, the weight for the auxiliary loss term in Equation (41) is set to .
- Implementation details: The front-end dense encoder and the back-end mask and phase decoders follow the MP-SENet design, using dilated convolutions with dilation rates and dense skip connections. The magnitude mask is predicted using a learnable sigmoid activation with the initial parameter . Our implementation is based on the official MUSE repository, with additional modifications for integrating Mamba, dynamic SNR augmentation, and the augmented loss terms. Our implementation is built upon a modified version of the official MUSE repository. A refactored and documented AMUSE++ codebase, along with pre-trained models, is planned to be released publicly to support reproducibility and future extensions.To quantify computational efficiency, we further measured the runtime of AMUSE++ on an Intel(R) Xeon(R) CPU E5-1620 and an NVIDIA RTX 3060 (12 GB). On the VoiceBank+DEMAND test set, AMUSE++ achieves a real-time factor (RTF) of 0.032 and reduces GPU memory usage from 7.5 GB (MUSE) to 3.6 GB. The RTF details will be clarified in the subsequent section.
- Perceptual Evaluation of Speech Quality (PESQ) [28]: Scores range from to 4.5, with higher values indicating improved perceived quality, based on a predictive model of human mean opinion scores (MOS).
- Short-Time Objective Intelligibility (STOI) [29]: Measures speech intelligibility on a 0 to 1 scale; higher scores signify clearer, more understandable speech.
- Segmental Signal-to-Noise Ratio (SSNR) [30]: Assesses segmental SNR throughout an utterance, where increased values reflect more effective noise attenuation.
- Composite Overall Quality (COVL) [30]: Reports a MOS-like score from 0 to 5 for holistic speech quality, with higher numbers denoting better quality.
- Composite Signal Distortion (CSIG) [30]: Rates signal distortion on a MOS scale (0 to 5), with higher results indicating less distortion in the output.
- Composite Background Noise Intrusiveness (CBAK) [30]: Evaluates the intrusiveness of background noise, again on a 0–5 MOS scale; higher values indicate greater background noise suppression.
5. Results and Discussions
5.1. Overall Performance Evaluation
- AMUSE++ vs. MUSE++ (our backbone): Relative to MUSE++, AMUSE++ consistently improves all objective metrics. The gains in PESQ, CSIG, CBAK, and COVL indicate that the proposed architectural changes lead to better perceived quality, less speech distortion, and less intrusive background noise. SSNR and STOI are also higher, showing that AMUSE++ achieves stronger noise reduction and slightly better intelligibility while still inheriting the compact Mamba-based backbone of MUSE++.
- AMUSE++ vs. MUSE: Compared with the original MUSE model, AMUSE++ not only achieves better scores on every metric (PESQ, CSIG, CBAK, COVL, SSNR, STOI) but also uses fewer parameters (0.31 M vs. 0.51 M). This demonstrates that starting from the lightweight MUSE++ backbone and enhancing it with 2D bi-directional Mamba leads to a model that is both more efficient and more effective than the heavier Transformer-based baseline.
- Trade-off between quality and complexity: While AMUSE++ increases the parameter count from 0.17 M (MUSE++) to 0.31 M, this is still substantially smaller than MUSE, and the additional capacity is reflected in consistent improvements across all evaluation indices. Thus, AMUSE++ can be viewed as a strengthened successor to MUSE++, offering a more favorable quality–complexity trade-off than either of the two baselines.
- Behavior of quality- and noise-related indices: The simultaneous improvement in CSIG and CBAK shows that AMUSE++ is able to reduce background noise without introducing additional speech distortion, which is a common failure mode for overly aggressive denoisers. At the same time, the higher SSNR together with the gains in PESQ and COVL suggest that the model not only removes more noise energy but also reconstructs cleaner and more natural spectral details.
- Intelligibility and robustness: The STOI gains over both MUSE and MUSE++ indicate that the proposed 2D bi-directional Mamba modules help preserve critical linguistic cues, rather than merely optimizing for signal-level metrics. This is particularly important for downstream ASR and human listening, and it confirms that AMUSE++ improves robustness in challenging noisy conditions while maintaining high intelligibility.
5.2. Ablation Study of AMUSE++
- From 1D uni-directional to 1D bi-directional Mamba: The two rightmost columns (both excluding PDM and dynamic SNR/loss) isolate the effect of bi-directionality in the 1D configuration. Transitioning from 1D-uni to 1D-Bi yields modest gains in some metrics—PESQ improves from 3.2860 to 3.3386 and COVL from 4.0334 to 4.0793—but at the cost of reduced SSNR and STOI performance. These mixed results suggest that bi-directional temporal context may offer benefits in perceptual quality metrics despite trade-offs in signal fidelity and intelligibility measures, with only a minor parameter increase from 0.17 M to 0.27 M.
- Benefit of 2D bi-directional modeling (1D-Bi vs. 2D-Bi): The fourth column (2D-Bi, no PDM, no dynamic SNR/loss) replaces 1D-Bi with the proposed 2D bi-directional Mamba while keeping the rest of the architecture unchanged. This further improves PESQ (3.3386 → 3.3845), CSIG (4.6389 → 4.6603), CBAK (3.7342 → 3.8349), COVL (4.0793 → 4.1225), and STOI (0.9494 → 0.9538), together with a substantial SSNR gain (8.4190 → 9.6307). These results indicate that jointly modeling temporal and frequency dependencies in a bi-directional manner is more effective than restricting Mamba to a single temporal dimension, as it better exploits the inherent 2D structure of spectrograms and yields cleaner, more speech-like reconstructions.
- Effect of the preliminary denoising module (2D-Bi vs. 2D-Bi + PDM): The third column introduces the PDM on top of the 2D-Bi configuration, still without dynamic SNR or augmented loss. Adding the PDM yields consistent improvements across almost all metrics: PESQ increases from 3.3845 to 3.4425, CSIG from 4.6603 to 4.6842, COVL from 4.1225 to 4.1733, and SSNR from 9.6307 to 9.5079 (comparable), with STOI remaining at a similarly high level. The parameter count only slightly increases from 0.29 M to 0.31 M. These observations suggest that a dedicated front-end denoising stage helps the backbone operate on cleaner and more structured time–frequency representations, enabling it to focus on finer-grained refinement and thereby improving perceived quality.
- Dynamic SNR and augmented loss (2D-Bi + PDM vs. full AMUSE++): The second column adds dynamic SNR augmentation together with the augmented loss on top of the 2D-Bi + PDM architecture. This full configuration achieves the best performance across all metrics: PESQ, CSIG, CBAK, COVL, SSNR, and STOI all reach their maximum values (e.g., PESQ 3.5305, COVL 4.2753, SSNR 10.8090, STOI 0.9592), with no change in the number of parameters. The improvements over the second column therefore stem purely from a stronger and more diverse training strategy, indicating that dynamic SNR and the augmented loss provide more effective supervision and improve robustness across different noise levels.
- Global trend and interaction of components: Overall, the sequenceexhibits a nearly monotonic improvement across all objective metrics. This pattern shows that the proposed design choices are complementary: (i) moving from 1D uni-directional to 1D bi-directional Mamba improves temporal context aggregation, (ii) upgrading to 2D bi-directional Mamba exploits the full time–frequency structure, (iii) adding the PDM provides a useful front-end enhancement stage, and (iv) dynamic SNR with the augmented loss further aligns optimization with perceptual and robustness objectives. The final AMUSE++ configuration therefore represents a well-balanced combination of architectural and training improvements that jointly contribute to its strong overall performance.
| AMUSE++ Variants | |||||
|---|---|---|---|---|---|
| Mamba Type | 2D-Bi | 2D-Bi | 2D-Bi | 1D-Bi | 1D-uni |
| PDM | + | + | - | - | - |
| Dynamic SNR + Augmented Loss | + | - | - | - | - |
| PESQ | 3.5305 | 3.4425 | 3.3845 | 3.3386 | 3.2860 |
| CSIG | 4.7660 | 4.6842 | 4.6603 | 4.6389 | 4.6082 |
| CBAK | 3.9791 | 3.8539 | 3.8349 | 3.7342 | 3.7610 |
| COVL | 4.2753 | 4.1733 | 4.1225 | 4.0793 | 4.0334 |
| SSNR | 10.8090 | 9.5079 | 9.6307 | 8.4190 | 9.2603 |
| STOI | 0.9592 | 0.9535 | 0.9538 | 0.9494 | 0.9504 |
| Params (M) | 0.31 | 0.31 | 0.29 | 0.27 | 0.17 |
5.3. Qualitative Evaluation Using Spectrograms
5.4. Comparison with Some State-of-the-Art SE Methods
- Performance of classical baselines. The two conventional non-neural SE methods, Wiener filtering [31] and LogMMSE [5], provide only limited enhancement compared with modern deep models. Specifically, Wiener filtering improves PESQ from 1.97 (Noisy) to 2.22 but yields a slightly lower CSIG (3.23 vs. 3.35) and essentially no gain in COVL (2.63), suggesting that it mainly performs mild noise suppression without substantially improving perceived quality or signal distortion. In contrast, LogMMSE attains higher scores across all reported metrics (PESQ 2.34, CSIG 3.67, CBAK 3.12, COVL 3.04, STOI 0.91) than both the Noisy and Wiener conditions, indicating that it can better preserve speech structure while reducing background noise. However, both methods remain clearly inferior to recent lightweight neural approaches such as TSTNN, DPT-FSNet, MUSE/MUSE++, and AMUSE++, whose PESQ, CSIG, CBAK, and COVL scores are substantially higher, highlighting the advantage of data-driven architectures over traditional spectral-domain estimators in this task.
- Overall enhancement quality. AMUSE++ attains the best or on-par-with-the-best scores across all objective metrics, achieving a PESQ of 3.53, CSIG of 4.77, CBAK of 3.98, COVL of 4.28, and STOI of 0.96. These scores are consistently higher than those of most competing methods, indicating that AMUSE++ not only improves perceived speech quality (PESQ, COVL) and signal distortion (CSIG), but also yields competitive background noise suppression (CBAK) and intelligibility (STOI). In particular, the gains in PESQ and COVL suggest that the proposed architecture is especially effective at producing natural-sounding enhanced speech.
- Comparison within the MUSE family. Within the MUSE family, both MUSE++ and AMUSE++ clearly outperform the original MUSE baseline across all reported metrics, confirming the effectiveness of the Mamba-based backbone introduced in MUSE++. Building on this stronger backbone, AMUSE++ further improves performance over MUSE++ itself, with noticeable gains in PESQ, CSIG, CBAK, and COVL. This indicates that the proposed preliminary denoising module, the 2D bi-directional Mamba extension, and the auxiliary enhancement branch provide additional benefits on top of the MUSE++ architecture, rather than merely reproducing its behavior with a different parametrization.
- Parameter efficiency and trade-offs. Despite its strong performance, AMUSE++ remains highly compact, with only 0.31 M parameters. This is one order of magnitude smaller than DB-AIAT (2.81 M), while still achieving clearly superior PESQ and COVL scores. Compared with other competitive lightweight models such as TSTNN (0.92 M), DPT-FSNet (0.88 M), and MetricGAN-OKDv2 (0.82 M), AMUSE++ uses fewer parameters yet consistently matches or outperforms them across all metrics, particularly in PESQ and COVL. Although AMUSE++ has slightly more parameters than its backbone model MUSE++ (0.17 M), the increase in model size is modest and is accompanied by consistent performance gains, indicating a favorable trade-off between complexity and enhancement quality.
- Impact of the proposed design. Taken together, these results suggest that the combination of the preliminary denoising module, the 2D bi-directional Mamba extension of the MUSE++ backbone, and the auxiliary enhancement branch allows AMUSE++ to exploit richer spectro-temporal context than conventional 1D or purely convolutional designs. The improvements over MUSE/MUSE++ and other lightweight baselines indicate that the proposed architecture can effectively strengthen speech enhancement performance without incurring a prohibitive parameter cost, making it attractive for deployment in resource-constrained scenarios.
6. Conclusions and Future Work
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Leglaive, S.; Fraticelli, M.; ElGhazaly, H.; Borne, L.; Sadeghi, M.; Wisdom, S.; Pariente, M.; Hershey, J.R.; Pressnitzer, D.; Barker, J.P. Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge. Comput. Speech Lang. 2025, 89, 101685. [Google Scholar] [CrossRef] [Scilit]
- Zheng, C.; Zhang, H.; Liu, W.; Luo, X.; Li, A.; Li, X.; Moore, B.C.J. Sixty Years of Frequency-Domain Monaural Speech Enhancement: From Traditional to Deep Learning Methods. Trends Hear. 2023, 27, 23312165231209913. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Natarajan, S.; Rahman Al-Haddad, S.A.; Ahmad, F.A.; Kamil, R.; Hassan, M.K.; Azrad, S.; Macleans, J.F.; Abdulhussain, S.H.; Mahmmod, B.M.; Saparkhojayev, N.; et al. Deep neural networks for speech enhancement and speech recognition: A systematic review. Ain Shams Eng. J. 2025, 16, 103405. [Google Scholar] [CrossRef] [Scilit]
- Boll, S.F. Suppression of acoustic noise in speech using spectral subtraction. IEEE Trans. Acoust. Speech Signal Process. 1979, 27, 113–120. [Google Scholar] [CrossRef] [Scilit]
- Ephraim, Y.; Malah, D. Speech enhancement using a minimum mean-square error log-spectral amplitude estimator. IEEE Trans. Acoust. Speech Signal Process. 1985, 33, 443–445. [Google Scholar] [CrossRef] [Scilit]
- Paliwal, K.K.; Wojcicki, K.; Rao, B.P. The importance of phase in speech enhancement. Speech Commun. 2010, 53, 465–494. [Google Scholar] [CrossRef] [Scilit]
- Lin, Z.; Chen, X.; Wang, J. MUSE: Flexible Voiceprint Receptive Fields and Multi-Path Fusion Enhanced Taylor Transformer for U-Net-based Speech Enhancement. In Proceedings of the INTERSPEECH 2024, Kos, Greece, 1–5 September 2024; pp. 672–676. [Google Scholar]
- Wahab, F.E.; Ye, Z.; Saleem, N.; Ullah, R. Compact deep neural networks for real-time speech enhancement on resource-limited devices. Speech Commun. 2024, 156, 103008. [Google Scholar] [CrossRef] [Scilit]
- Saleem, N.; Bourouis, S.; Elmannai, H.; Algarni, A.D. CTSE-Net: Resource-efficient convolutional and TF-transformer network for speech enhancement. Knowl.-Based Syst. 2024, 290, 110597. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Zhuang, X.; Qian, Y.; Wang, M. Lightweight Dynamic Sparse Transformer for Monaural Speech Enhancement. In Proceedings of the Interspeech 2024, Kos, Greece, 1–5 September 2024; pp. 3816–3820. [Google Scholar]
- Mattursun, A.; Wang, L.; Yu, Y.; Ma, C. Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), Nantes, France, 30 June–4 July 2025. [Google Scholar]
- Yin, D.; Huang, J.; Wu, Y.; Zou, Y.; Xue, W.; Jin, Z.Y.; Zhang, S.; Wu, J.; Yu, D. PHASEN: A self-supervised phase-and-harmonics-aware speech enhancement network. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; Volume 34, pp. 9458–9465. [Google Scholar]
- Lu, Y.X.; Ai, Y.; Ling, Z.H. MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra. In Proceedings of the Interspeech 2023, Dublin, Ireland, 20–24 August 2023; pp. 3834–3838. [Google Scholar]
- Lu, Y.X.; Ai, Y.; Ling, Z.H. Explicit estimation of magnitude and phase spectra in speech enhancement. Neural Netw. 2025, 189, 107562. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, L.; Liu, W.; Meng, R.; Lee, G.; Baek, S.; Moon, H.G. Fspen: An Ultra-Lightweight Network for Real Time Speech Enahncment. In Proceedings of the ICASSP 2024—2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Republic of Korea, 14–19 April 2024; pp. 10671–10675. [Google Scholar] [CrossRef] [Scilit]
- Michelsanti, D.; Tan, Z.H.; Xu, Y.; Richter, S.R.; Ma, M.; Sørensen, J.; Jensen, J.; Gerkmann, T.; Jensen, S.; Virtanen, T.; et al. An Overview of Deep-Learning-Based Audio-Visual Speech Enhancement and Separation. IEEE/ACM Trans. Audio Speech Lang. Process. 2021, 29, 1368–1396. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Du, Z.; Lei, S.; Huoyijun, H.; Yang, M.; Zhang, Z.; Shen, L. Lip landmark-based audio-visual speech enhancement with cross-modality attention. Neurocomputing 2023, 545, 127409. [Google Scholar]
- Kuang, K.; Yang, F.; Yang, J. A lightweight speech enhancement network fusing bone- and air-conducted speech. J. Acoust. Soc. Am. 2024, 156, 1355–1366. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, X.; Zhang, Y.; Wang, S. Mixed T-domain and TF-domain Magnitude and Phase Representations for GAN-based Speech Enhancement. Sci. Rep. 2024, 14, 17698. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, T.J.; Hung, J.W. Enhancing the MUSE Speech Enhancement Framework with Mamba-Based Architecture and Extended Loss Functions. Mathematics 2025, 13, 3481. [Google Scholar] [CrossRef] [Scilit]
- Gu, A.; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. In Proceedings of the Conference on Language Modeling (COLM), Philadelphia, PA, USA, 7–9 October 2024. [Google Scholar]
- Dao, T.; Gu, A. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. In Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, 21–27 July 2024. ICML’24. [Google Scholar]
- Choromanski, K.; Likhosherstov, V.; Dohan, D.; Song, X.; Gane, A.; Sarlós, T.; Hawkins, P.; Davis, J.; Mohiuddin, A.; Kaiser, Ł; et al. Rethinking Attention with Performers. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2021. [Google Scholar]
- Zhang, X.; Zhang, Q.; Liu, H.; Xiao, T.; Qian, X.; Ahmed, B.; Ambikairajah, E.; Li, H.; Epps, J. Mamba in Speech: Towards an Alternative to Self-Attention. IEEE Trans. Audio Speech Lang. Process. 2025, 33, 1933–1948. [Google Scholar] [CrossRef] [Scilit]
- Kim, S.H.; Kim, T.G.; Chun, C.J. Mamba-based Hybrid Model for Speech Enhancement. In Proceedings of the Interspeech, Rotterdam, The Netherlands, 17–21 August 2025. [Google Scholar]
- Valentini-Botinhao, C.; Wang, X.; Takaki, S.; Yamagishi, J. Investigating RNN-based speech enhancement methods for noise-robust Text-to-Speech. In Proceedings of the 9th ISCA Workshop on Speech Synthesis Workshop (SSW 9), Sunnyvale, CA, USA, 13–15 September 2016; pp. 146–152. [Google Scholar] [CrossRef] [Scilit]
- Thiemann, J.; Ito, N.; Vincent, E. The Diverse Environments Multi-channel Acoustic Noise Database (DEMAND): A database of multichannel environmental noise recordings. In Proceedings of the 21st International Congress on Acoustics, Montreal, QC, Canada, 2–7 June 2013; pp. 1–6. [Google Scholar]
- ITU-T. Perceptual Evaluation of Speech Quality (PESQ), an Objective Method for End-to-End Speech Quality Assessment of Narrowband Telephone Networks and Speech Codecs; Technical Report P.862; International Telecommunication Union: Geneva, Switzerland, 2001. [Google Scholar]
- Taal, C.H.; Hendriks, R.C.; Heusdens, R.; Jensen, J. An Algorithm for Intelligibility Prediction of Time–Frequency Weighted Noisy Speech. IEEE Trans. Audio Speech Lang. Process. 2011, 19, 2125–2136. [Google Scholar] [CrossRef] [Scilit]
- Hu, Y.; Loizou, P.C. Evaluation of Objective Quality Measures for Speech Enhancement. IEEE Trans. Audio Speech Lang. Process. 2008, 16, 229–238. [Google Scholar] [CrossRef] [Scilit]
- Scalart, P.; Vieira Filho, J.V. Speech enhancement based on a priori signal to noise estimation. In Proceedings of the 1996 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Atlanta, GA, USA, 7–10 May 1996; Volume 2, pp. 629–632. [Google Scholar]
- Logmmse: A Python Implementation of the LogMMSE Speech Enhancement/Noise Reduction Algorithm. Version 1.5.3. Available online: https://pypi.org/project/logmmse/ (accessed on 2 January 2026).
- Wang, K.; He, B.; Zhu, W.P. TSTNN: Two-stage Transformer Based Neural Network for Speech Enhancement in the Time Domain. In Proceedings of the ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, 6–11 June 2021; IEEE: New York, NY, USA, 2021; pp. 7098–7102. [Google Scholar]
- Yu, G.; Li, A.; Zheng, C.; Guo, Y.; Wang, Y.; Wang, H. Dual-Branch Attention-In-Attention Transformer for Single-Channel Speech Enhancement. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, 23–27 May 2022; pp. 7847–7851. [Google Scholar]
- Dang, F.; Chen, H.; Zhang, P. DPT-FSNet: Dual-Path Transformer Based Full-Band and Sub-Band Fusion Network for Speech Enhancement. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, 23–27 May 2022; pp. 6857–6861. [Google Scholar]
- Shin, W.; Lee, B.H.; Kim, J.S.; Park, H.J.; Han, S.W. MetricGAN-OKD: Multi-Metric Optimization of MetricGAN via Online Knowledge Distillation for Speech Enhancement. In Proceedings of the 40th International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; pp. 31521–31538. [Google Scholar]
- Shin, W.; Park, H.J.; Kim, J.S.; Lee, B.H.; Han, S.W. Multi-View Attention Transfer for Efficient Speech Enhancement. In Proceedings of the Interspeech, Incheon, Republic of Korea, 18–22 September 2022; pp. 1196–1200. [Google Scholar] [CrossRef] [Scilit]
- Pascual, S.; Bonafonte, A.; Serrà, J. SEGAN: Speech Enhancement Generative Adversarial Network. In Proceedings of the Interspeech 2017, Stockholm, Sweden, 20–24 August 2017; pp. 3642–3646. [Google Scholar]
- Watanabe, S.; Mandel, M.I.; Barker, J.; Vincent, E.; Arora, A.; Chang, X.; Khudanpur, S.; Manohar, V.; Povey, D.; Raj, D.; et al. CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings. In Proceedings of the 6th International Workshop on Speech Processing in Everyday Environments (CHiME 2020), Online Virtual Workshop, 4 May 2020. [Google Scholar]
- Reddy, C.K.A.; Dubey, H.; Noufal, A.A.; Gopal, V.; Cutler, R.; Braun, S.; Gamper, H.; Aichner, R.; Srinivasan, S.; Tashev, I.; et al. The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results. In Proceedings of the Interspeech, Shanghai, China, 25–29 October 2020; pp. 2472–2476. [Google Scholar] [CrossRef] [Scilit]






| PESQ | CSIG | CBAK | COVL | SSNR | STOI | #Para. (M) | |
|---|---|---|---|---|---|---|---|
| MUSE | 3.3475 | 4.6163 | 3.7965 | 4.0827 | 9.3309 | 0.9506 | 0.51 |
| MUSE++ | 3.3636 | 4.6619 | 3.8584 | 4.1209 | 10.1838 | 0.9538 | 0.17 |
| AMUSE++ | 3.5305 | 4.7660 | 3.9791 | 4.2753 | 10.8090 | 0.9592 | 0.31 |
| SNR | Method | PESQ | CSIG | CBAK | COVL | SSNR | STOI |
|---|---|---|---|---|---|---|---|
| 17.5 dB | MUSE++ | 3.8163 | 4.9355 | 4.2455 | 4.5522 | 12.7526 | 0.9722 |
| AMUSE++ | 3.9414 | 4.9734 | 4.3472 | 4.6607 | 13.3976 | 0.9748 | |
| 12.5 dB | MUSE++ | 3.5472 | 4.8180 | 3.9963 | 4.3007 | 10.9278 | 0.9652 |
| AMUSE++ | 3.6922 | 4.8943 | 4.1033 | 4.4318 | 11.5146 | 0.9680 | |
| 7.5 dB | MUSE++ | 3.2848 | 4.6320 | 3.7749 | 4.0440 | 9.4844 | 0.9545 |
| AMUSE++ | 3.4650 | 4.7511 | 3.9012 | 4.2115 | 10.1093 | 0.9596 | |
| 2.5 dB | MUSE++ | 2.8122 | 4.2659 | 3.4221 | 3.5924 | 7.6042 | 0.9237 |
| AMUSE++ | 3.0291 | 4.4482 | 3.5695 | 3.8023 | 8.2486 | 0.9345 |
| SNR | Method | PESQ | CSIG | CBAK | COVL | SSNR | STOI |
|---|---|---|---|---|---|---|---|
| −5 dB | MUSE++ | 2.3215 | 3.8013 | 3.0254 | 3.0981 | 5.4021 | 0.8865 |
| AMUSE++ | 2.5450 | 4.0268 | 3.2047 | 3.3282 | 6.3655 | 0.9032 | |
| −10 dB | MUSE++ | 1.8938 | 3.3068 | 2.6001 | 2.6142 | 2.8387 | 0.8131 |
| AMUSE++ | 2.0571 | 3.5192 | 2.7770 | 2.8133 | 3.9378 | 0.8381 |
| #Para. (M) ↓ | RTF ↓ | IFT (s) ↓ | THP ↑ | Peak VRAM (GB) ↓ | |
|---|---|---|---|---|---|
| MUSE | 0.51 | 0.1038 | 218 | 9.64 | 7.58 |
| MUSE++ | 0.17 | 0.0116 | 27 | 85.98 | 1.20 |
| AMUSE++ | 0.31 | 0.0320 | 75 | 31.23 | 3.64 |
| Method | Parameters | PESQ | CSIG | CBAK | COVL | STOI |
|---|---|---|---|---|---|---|
| Noisy | - | 1.97 | 3.35 | 2.44 | 2.63 | 0.91 |
| Wiener | - | 2.22 | 3.23 | 2.68 | 2.63 | - |
| logMMSE | - | 2.34 | 3.67 | 3.12 | 3.04 | 0.91 |
| TSTNN | 0.92 M | 2.96 | 4.33 | 3.53 | 3.67 | 0.95 |
| DB-AIAT | 2.81 M | 3.31 | 4.61 | 3.75 | 3.96 | - |
| DPT-FSNet | 0.88 M | 3.33 | 4.58 | 3.72 | 4.00 | 0.96 |
| MetricGAN-OKDv2 | 0.82 M | 3.12 | 4.27 | 3.16 | 3.71 | 0.95 |
| MANNER-S-5.3GF | 0.90 M | 3.06 | 4.42 | 3.58 | 3.77 | 0.95 |
| MUSE | 0.51 M | 3.35 | 4.62 | 3.80 | 4.08 | 0.95 |
| MUSE++ | 0.17 M | 3.36 | 4.66 | 3.86 | 4.12 | 0.95 |
| AMUSE++ (ours) | 0.31 M | 3.53 | 4.77 | 3.98 | 4.28 | 0.96 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, T.-J.; Chen, B.; Hung, J.-W. AMUSE++: A Mamba-Enhanced Speech Enhancement Framework with Bi-Directional and Advanced Front-End Modeling. Electronics 2026, 15, 282. https://doi.org/10.3390/electronics15020282
Li T-J, Chen B, Hung J-W. AMUSE++: A Mamba-Enhanced Speech Enhancement Framework with Bi-Directional and Advanced Front-End Modeling. Electronics. 2026; 15(2):282. https://doi.org/10.3390/electronics15020282
Chicago/Turabian StyleLi, Tsung-Jung, Berlin Chen, and Jeih-Weih Hung. 2026. "AMUSE++: A Mamba-Enhanced Speech Enhancement Framework with Bi-Directional and Advanced Front-End Modeling" Electronics 15, no. 2: 282. https://doi.org/10.3390/electronics15020282
APA StyleLi, T.-J., Chen, B., & Hung, J.-W. (2026). AMUSE++: A Mamba-Enhanced Speech Enhancement Framework with Bi-Directional and Advanced Front-End Modeling. Electronics, 15(2), 282. https://doi.org/10.3390/electronics15020282

