Efficient Speech Enhancement via Flow Matching with Gated Bidirectional Mamba2
Abstract
1. Introduction
- (1)
- We are among the first to integrate flow matching with selective state-space models for SE, achieving a real-time factor of —more than five times faster than representative diffusion-based baselines while maintaining comparable perceptual quality.
- (2)
- We design DiMamba, a parameter-efficient bidirectional Mamba2 that shares weights across the forward and backward directions and fuses them through a multiplicative agreement gate, reducing the parameter overhead of bidirectionalization by approximately relative to the concatenation-based alternative.
- (3)
- We show through experiments that the proposed framework provides a favorable trade-off among perceptual quality, speaker similarity, intelligibility, and runtime efficiency when compared with representative discriminative, diffusion-based, LM-based, and flow-matching baselines.
2. Materials and Methods
2.1. Overall Architecture
2.2. Flow Matching for Speech Enhancement
2.3. Flow Matching with Mamba2
2.3.1. Encoder
2.3.2. Input Embedding
2.3.3. DiMamba Block
2.3.4. Training Objective
2.3.5. Decoder
3. Experimental Setup
3.1. Datasets
3.1.1. Dataset Configuration
3.1.2. Training, Validation, and Test Splits
- (1)
- LibriSpeech. We used the train-other-500 subset for training and the official dev-clean and test-clean subsets for validation and testing, respectively. These official splits are speaker-disjoint by construction.
- (2)
- VCTK. Speakers were partitioned into train/validation/test splits with a ratio of , ensuring that no speaker appears in more than one split. Utterances within each split were used in full.
- (3)
- GigaSpeech. A ∼100 h subset was drawn from the GigaSpeech training partition for training only; GigaSpeech was not used for evaluation.
3.1.3. Online Degradation Pipeline
- (1)
- Reverberation. With probability , an RIR was randomly sampled (uniformly) from the OpenSLR-26/28 pool and convolved with the clean utterance; otherwise, the utterance was kept anechoic.
- (2)
- Additive noise. A noise segment was randomly cropped from the combined DEMAND + WHAM! noise pool and added to every training sample. The SNR was drawn from a four-bin categorical distribution that emphasizes moderately noisy conditions: bins , , , and dB were selected with probabilities , , , and , respectively, and the SNR was sampled uniformly within the chosen bin. This allocates more probability mass to the moderate-SNR regime, which is closer to typical real-world conditions while still exposing the model to both severe and mild cases.
- (3)
- Bandwidth limitation. With probability , the mixture was further low-pass filtered with a cutoff frequency drawn uniformly from kHz, simulating bandwidth-limited recording conditions.
- (4)
- Loudness normalization. The noisy mixture and its clean reference were jointly scaled to a target dBFS level drawn from dBFS, with the same gain applied to both signals to preserve their relative amplitude and the target SNR.
3.1.4. Test Sets
3.2. Evaluation Metrics
3.3. Training Setup
3.4. Baselines
4. Results and Discussion
4.1. Overall Comparison with Baselines
4.2. Speaker Preservation and Intelligibility Analysis
4.3. Runtime Efficiency Analysis
4.4. Ablation Study on the Backbone Design
4.5. Speaker Reconstruction Experiment
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Westhausen, N.L.; Meyer, B.T. Dual-Signal Transformation LSTM Network for Real-Time Noise Suppression. In Proceedings of the Interspeech 2020, Shanghai, China, 25–29 October 2020; pp. 2477–2481. [Google Scholar] [CrossRef]
- Le, X.; Chen, H.; Chen, K.; Lu, J. DPCRN: Dual-Path Convolution Recurrent Network for Single Channel Speech Enhancement. In Proceedings of the Interspeech 2021, Brno, Czechia, 30 August–3 September 2021; pp. 2811–2815. [Google Scholar] [CrossRef]
- Lu, Y.X.; Ai, Y.; Ling, Z.H. MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra. In Proceedings of the Interspeech 2023, Dublin, Ireland, 20–24 August 2023; pp. 3834–3838. [Google Scholar]
- Li, X.; Wang, Q.; Liu, X. MaskSR: Masked Language Model for Full-band Speech Restoration. In Proceedings of the Interspeech 2024, Kos, Greece, 1–5 September 2024. [Google Scholar] [CrossRef]
- Wang, Z.; Zhu, X.; Zhang, Z.; Lv, Y.; Jiang, N.; Zhao, G.; Xie, L. Selm: Speech enhancement using discrete tokens and language models. In Proceedings of the ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Republic of Korea, 14–19 April 2024; IEEE: New York, NY, USA, 2024; pp. 11561–11565. [Google Scholar]
- Yang, H.; Su, J.; Kim, M.; Jin, Z. Genhancer: High-fidelity speech enhancement via generative modeling on discrete codec tokens. In Proceedings of the Interspeech, Kos, Greece, 1–5 September 2024; Volume 2024, pp. 1170–1174. [Google Scholar]
- Zhang, J.; Yang, J.; Fang, Z.; Wang, Y.; Zhang, Z.; Wang, Z.; Fan, F.; Wu, Z. Anyenhance: A unified generative model with prompt-guidance and self-critic for voice enhancement. IEEE Trans. Audio Speech Lang. Process. 2025, 33, 3085–3098. [Google Scholar]
- Kang, B.; Zhu, X.; Zhang, Z.; Ye, Z.; Liu, M.; Wang, Z.; Zhu, Y.; Ma, G.; Chen, J.; Xiao, L.; et al. LLaSE-G1: Incentivizing generalization capability for llama-based speech enhancement. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Vienna, Austria, 2025; pp. 13292–13305. [Google Scholar]
- Lu, Y.J.; Wang, Z.Q.; Watanabe, S.; Richard, A.; Yu, C.; Tsao, Y. Conditional diffusion probabilistic model for speech enhancement. In Proceedings of the ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, 23–27 May 2022; IEEE: New York, NY, USA, 2022; pp. 7402–7406. [Google Scholar]
- Lemercier, J.M.; Richter, J.; Welker, S.; Gerkmann, T. Storm: A diffusion-based stochastic regeneration model for speech enhancement and dereverberation. IEEE/ACM Trans. Audio Speech Lang. Process. 2023, 31, 2724–2737. [Google Scholar]
- Welker, S.; Richter, J.; Gerkmann, T. Speech Enhancement with Score-Based Generative Models in the Complex STFT Domain. In Proceedings of the Interspeech 2022, Incheon, Republic of Korea, 18–22 September 2022; pp. 2928–2932. [Google Scholar] [CrossRef]
- Lipman, Y.; Chen, R.T.; Ben-Hamu, H.; Nickel, M.; Le, M. Flow Matching for Generative Modeling. In Proceedings of the Eleventh International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Wang, Z.; Liu, Z.; Zhu, X.; Zhu, Y.; Liu, M.; Chen, J.; Xiao, L.; Weng, C.; Xie, L. FlowSE: Efficient and high-quality speech enhancement via flow matching. arXiv 2025, arXiv:2505.19476. [Google Scholar]
- Lee, S.; Cheong, S.; Han, S.; Shin, J.W. Flowse: Flow matching-based speech enhancement. In Proceedings of the ICASSP 2025—2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 6–11 April 2025; IEEE: New York, NY, USA, 2025; pp. 1–5. [Google Scholar]
- Rong, X.; Gao, J.; Wang, Z.; Yesilbursa, M.; Wojcicki, K.; Lu, J. StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement. arXiv 2026, arXiv:2603.09234. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. In Proceedings of the First Conference on Language Modeling, Philadelphia, PA, USA, 7–9 October 2024. [Google Scholar]
- Fan, C.; Liu, E.; Li, A.; Tao, J.; Zhou, J.; Li, J.; Zheng, C.; Lv, Z. BSDB-net: Band-split dual-branch network with selective state spaces mechanism for monaural speech enhancement. Proc. AAAI Conf. Artif. Intell. 2025, 39, 23850–23858. [Google Scholar] [CrossRef]
- Dao, T.; Gu, A. Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality. In Proceedings of the International Conference on Machine Learning, Vienna, Austria, 21–27 July 2024; PMLR: Cambridge, MA, USA, 2024. [Google Scholar]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv 2024, arXiv:2401.09417. [Google Scholar] [CrossRef]
- Wang, J.; Lin, Z.; Wang, T.; Ge, M.; Wang, L.; Dang, J. Mamba-SEUNet: Mamba UNet for monaural speech enhancement. In Proceedings of the ICASSP 2025—2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 6–11 April 2025; IEEE: New York, NY, USA, 2025; pp. 1–5. [Google Scholar]
- Zhang, X.; Zhang, Q.; Liu, H.; Xiao, T.; Qian, X.; Ahmed, B.; Ambikairajah, E.; Li, H.; Epps, J. Mamba in speech: Towards an alternative to self-attention. IEEE Trans. Audio Speech Lang. Process. 2025, 33, 1933–1948. [Google Scholar]
- Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; Soricut, R. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
- Dehghani, M.; Gouws, S.; Vinyals, O.; Uszkoreit, J.; Kaiser, Ł. Universal transformers. arXiv 2018, arXiv:1807.03819. [Google Scholar]
- Dauphin, Y.N.; Fan, A.; Auli, M.; Grangier, D. Language modeling with gated convolutional networks. In Proceedings of the International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; PMLR: Cambridge, MA, USA, 2017; pp. 933–941. [Google Scholar]
- Shazeer, N. Glu variants improve transformer. arXiv 2020, arXiv:2002.05202. [Google Scholar] [CrossRef]
- Peebles, W.; Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; IEEE: New York, NY, USA, 2023; pp. 4195–4205. [Google Scholar]
- Ho, J.; Salimans, T. Classifier-free diffusion guidance. arXiv 2022, arXiv:2207.12598. [Google Scholar] [CrossRef]
- Siuzdak, H. Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis. arXiv 2023, arXiv:2306.00814. [Google Scholar]
- Panayotov, V.; Chen, G.; Povey, D.; Khudanpur, S. Librispeech: An asr corpus based on public domain audio books. In Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia, 19–24 April 2015; IEEE: New York, NY, USA, 2015; pp. 5206–5210. [Google Scholar]
- Veaux, C.; Yamagishi, J.; King, S. The voice bank corpus: Design, collection and data analysis of a large regional accent speech database. In Proceedings of the 2013 International Conference Oriental COCOSDA Held Jointly with 2013 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE), Gurgaon, India, 25–27 November 2013; pp. 1–4. [Google Scholar]
- Chen, G.; Chai, S.; Wang, G.; Du, J.; Zhang, W.; Weng, C.; Su, D.; Povey, D.; Trmal, J.; Zhang, J.; et al. GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio. In Proceedings of the Interspeech 2021, Brno, Czechia, 30 August–3 September 2021. [Google Scholar]
- Thiemann, J.; Ito, N.; Vincent, E. The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings. In Proceedings of the Meetings on Acoustics, Acoustical Society of America, San Francisco, CA, USA, 2–6 December 2013; Volume 19, p. 035081. [Google Scholar]
- Wichern, G.; Antognini, J.; Flynn, M.; Zhu, L.R.; McQuinn, E.; Crow, D.; Manilow, E.; Le Roux, J. WHAM!: Extending Speech Separation to Noisy Environments. In Proceedings of the Interspeech 2019, Graz, Austria, 15–19 September 2019. [Google Scholar]
- Ko, T.; Peddinti, V.; Povey, D.; Seltzer, M.L.; Khudanpur, S. A study on data augmentation of reverberant speech for robust speech recognition. In Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA, 5–9 March 2017; IEEE: New York, NY, USA, 2017; pp. 5220–5224. [Google Scholar]
- Reddy, C.K.; Dubey, H.; Koishida, K.; Nair, A.; Gopal, V.; Cutler, R.; Braun, S.; Gamper, H.; Aichner, R.; Srinivasan, S. INTERSPEECH 2021 Deep Noise Suppression Challenge. In Proceedings of the Interspeech 2021, Brno, Czechia, 30 August–3 September 2021. [Google Scholar]
- Reddy, C.K.; Gopal, V.; Cutler, R. DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors. In Proceedings of the ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, 6–11 June 2021; IEEE: New York, NY, USA, 2021; pp. 6493–6497. [Google Scholar]
- Wang, H.; Liang, C.; Wang, S.; Chen, Z.; Zhang, B.; Xiang, X.; Deng, Y.; Qian, Y. Wespeaker: A research and production oriented speaker embedding learning toolkit. In Proceedings of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; IEEE: New York, NY, USA, 2023; pp. 1–5. [Google Scholar]
- Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust speech recognition via large-scale weak supervision. In Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; PMLR: Cambridge, MA, USA, 2023; pp. 28492–28518. [Google Scholar]
- Luo, Y.; Mesgarani, N. Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation. IEEE/ACM Trans. Audio Speech Lang. Process. 2019, 27, 1256–1266. [Google Scholar] [CrossRef]
- Defossez, A.; Synnaeve, G.; Adi, Y. Real Time Speech Enhancement in the Waveform Domain. In Proceedings of the Interspeech 2020, Shanghai, China, 25–29 October 2020. [Google Scholar]




| Model | Type | With Reverb | Without Reverb | Real Recording | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DNSMOS ↑ | Spk Sim ↑ | DNSMOS ↑ | Spk Sim ↑ | DNSMOS ↑ | ||||||||
| SIG | BAK | OVL | SIG | BAK | OVL | SIG | BAK | OVL | ||||
| Noisy | - | 1.760 | 1.497 | 1.392 | 0.942 | 3.392 | 2.618 | 2.483 | 0.970 | 3.053 | 2.509 | 2.255 |
| Conv-TasNet | D | 2.415 | 2.710 | 2.010 | 0.815 | 3.092 | 3.341 | 3.001 | 0.802 | 3.102 | 2.975 | 2.410 |
| Demucs | D | 2.510 | 2.641 | 2.215 | 0.804 | 3.124 | 3.257 | 3.010 | 0.802 | 2.974 | 2.870 | 2.291 |
| CDiffuSE | Diff | 2.541 | 2.300 | 2.190 | 0.761 | 3.294 | 3.641 | 3.047 | 0.765 | 3.201 | 3.104 | 2.781 |
| SGMSE | Diff | 2.730 | 2.741 | 2.430 | 0.764 | 3.501 | 3.710 | 3.137 | 0.782 | 3.297 | 2.894 | 2.793 |
| StoRM | Diff | 2.947 | 3.141 | 2.516 | 0.790 | 3.514 | 3.941 | 3.205 | 0.798 | 3.410 | 3.379 | 2.940 |
| FlowSE | Flow | 3.391 | 3.957 | 3.109 | 0.861 | 3.556 | 4.032 | 3.291 | 0.876 | 3.481 | 4.004 | 3.119 |
| FMM(Mamba2) | Flow | 3.443 | 3.803 | 3.049 | 0.868 | 3.581 | 4.091 | 3.323 | 0.886 | 3.502 | 4.061 | 3.227 |
| Model | Type | WER ↓ | QMOS ↑ |
|---|---|---|---|
| Noisy | - | 25.20 | - |
| Conv-TasNet | D | 18.6 | 3.00 ± 0.07 |
| Demucs | D | 17.4 | 3.07 ± 0.10 |
| CDiffuSE | Diff | 16.8 | 3.13 ± 0.08 |
| SGMSE | Diff | 15.3 | 3.27 ± 0.07 |
| StoRM | Diff | 15.1 | 3.37 ± 0.10 |
| FMM | Flow | 4.7 | 3.58 ± 0.10 |
| Model | Type | NFE | Reverb Set ↑ | No Reverb Set ↑ | ||||
|---|---|---|---|---|---|---|---|---|
| SIG | BAK | OVL | SIG | BAK | OVL | |||
| CDiffuSE | Diff | >50 | 2.451 | 2.310 | 2.228 | 3.225 | 3.214 | 2.687 |
| SGMSE | Diff | >50 | 2.853 | 2.812 | 2.510 | 3.267 | 2.956 | 2.834 |
| StoRM | Diff | >50 | 2.967 | 3.241 | 2.627 | 3.401 | 3.476 | 2.914 |
| FMM | Flow | 10 | 3.380 | 3.944 | 3.100 | 3.477 | 4.046 | 3.198 |
| FMM | Flow | 15 | 3.445 | 4.012 | 3.165 | 3.472 | 4.052 | 3.200 |
| FMM | Flow | 20 | 3.473 | 4.036 | 3.189 | 3.474 | 4.065 | 3.202 |
| Model | Type | Parameters | Reverb Set ↑ | No Reverb Set ↑ | ||||
|---|---|---|---|---|---|---|---|---|
| SIG | BAK | OVL | SIG | BAK | OVL | |||
| Noisy | - | - | 1.870 | 1.678 | 1.428 | 3.021 | 2.543 | 2.168 |
| Mamba | Flow | 261 M | 2.875 | 3.330 | 2.502 | 3.288 | 3.700 | 2.877 |
| Concat Mamba2 | Flow | 364 M | 3.269 | 3.913 | 2.950 | 3.451 | 4.021 | 3.166 |
| FMM(our) | Flow | 286 M | 3.380 | 3.944 | 3.100 | 3.477 | 4.046 | 3.198 |
| Model | Type | Spk-Sim | Reverb Set ↑ | No Reverb Set ↑ | ||||
|---|---|---|---|---|---|---|---|---|
| SIG | BAK | OVL | SIG | BAK | OVL | |||
| Noisy | - | - | 1.760 | 1.497 | 1.392 | 3.392 | 2.618 | 2.483 |
| SELM | LM | 0.793 | 3.160 | 3.577 | 2.695 | 3.508 | 4.086 | 3.258 |
| Masksr | LM | 0.816 | 3.542 | 4.016 | 3.223 | 3.575 | 4.082 | 3.307 |
| FMM (our) | Flow | 0.868 | 3.443 | 3.803 | 3.049 | 3.581 | 4.091 | 3.323 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yuan, J.; Zhou, R.; Fan, C. Efficient Speech Enhancement via Flow Matching with Gated Bidirectional Mamba2. Appl. Sci. 2026, 16, 4757. https://doi.org/10.3390/app16104757
Yuan J, Zhou R, Fan C. Efficient Speech Enhancement via Flow Matching with Gated Bidirectional Mamba2. Applied Sciences. 2026; 16(10):4757. https://doi.org/10.3390/app16104757
Chicago/Turabian StyleYuan, Jiajun, Ruohua Zhou, and Cunhang Fan. 2026. "Efficient Speech Enhancement via Flow Matching with Gated Bidirectional Mamba2" Applied Sciences 16, no. 10: 4757. https://doi.org/10.3390/app16104757
APA StyleYuan, J., Zhou, R., & Fan, C. (2026). Efficient Speech Enhancement via Flow Matching with Gated Bidirectional Mamba2. Applied Sciences, 16(10), 4757. https://doi.org/10.3390/app16104757

