DualMambaFormer: A Parallel Hybrid Transformer–Mamba Network for Hyperspectral Image Classification
Highlights
- A novel parallel hybrid network, DualMambaFormer, was developed, demonstrating superior classification performance over state-of-the-art CNN, Transformer, and Mamba models across four benchmark datasets.
- The proposed dual-stream encoder successfully integrates Multi-Head Self-Attention and a Local Enhanced Mamba (LEM) module, enabling the synchronous extraction of global spatial correlations and local dynamic sequences without incurring quadratic computational complexity.
- The empirical success of the proposed “global–sequence–local” modeling paradigm demonstrates that the proposed dual-stream encoder combines MHSA and LEM to capture complementary global correlations, sequential dependencies, and local spatial context within a unified framework. This architectural shift provides a robust, scalable foundation for future multi-modal fusion and real-time edge deployment tasks in Earth observation.
- The study establishes a critical design principle for adapting sequence models to non-causal 2D spatial data: while State Space Models ensure linear complexity spectral processing, their integration with explicit local spatial operations is indispensable for preserving fine-grained morphological boundaries and complex land-cover textures.
Abstract
1. Introduction
- We propose a parallel hybrid architecture, termed DualMambaFormer, for hyperspectral image classification. The framework adopts a dual-stream paradigm to jointly model global spatial dependencies and dynamic sequential characteristics, thereby enhancing the representation of complex spectral–spatial information.
- A Spectral–spatial Residual Network and a Local Enhanced Mamba (LEM) branch are further introduced to strengthen feature extraction. Specifically, the SS-ResNet reduces spectral redundancy while enhancing local feature embedding, and the LEM branch integrates state space modeling with depthwise convolution to capture both long-range dependencies and fine-grained local textures, effectively alleviating the loss of spatial continuity caused by flattening 2D images into 1D sequences.
- Extensive experiments conducted on four benchmark hyperspectral datasets demonstrate the effectiveness of the proposed method. The results show that DualMambaFormer consistently outperforms representative convolutional, self-attention-based, and state-space-based methods, highlighting its strong classification accuracy, robustness, and generalization capability.
2. Related Work
2.1. CNN-Based Hyperspectral Classification
2.2. Transformer-Based Global Modeling
2.3. State Space Models (SSMs)
2.4. Prototype Learning and Few-Shot HSI Classification
3. Methodology
3.1. Overall Framework
3.2. Spectral–Spatial Residual Network
3.3. Parallel Dual-Branch Encoder
3.4. Dual-Token Fusion and Classification
4. Experiments
4.1. Experimental Datasets
- Indian Pines (IP): Acquired by the AVIRIS sensor over northwestern Indiana, USA, this dataset consists of pixels with a spatial resolution of 20 m. After removing bands associated with water absorption, 200 spectral bands remain for analysis. The scene comprises 16 agricultural land-cover classes. Characterized by relatively low spatial resolution and a significant presence of mixed pixels, the IP dataset serves as a benchmark for evaluating the model’s noise robustness and context extraction capabilities.
- Pavia University (PU): Collected by the ROSIS-03 sensor over the University of Pavia, Italy, this image has dimensions of pixels and a high spatial resolution of 1.3 m. The dataset provides 103 spectral bands and covers 9 urban categories, such as asphalt, bricks, and shadows. Given its rich textural details and the highly fragmented spatial distribution of urban structures, this dataset challenges the model’s ability to preserve class boundaries.
- Salinas (SA): Also acquired by the AVIRIS sensor, this dataset covers the Salinas Valley in California. It consists of pixels with a spatial resolution of 3.7 m and contains 204 spectral bands covering 16 crop classes. Unlike the IP dataset, SA is characterized by large, continuous spatial regions with high homogeneity, serving as an ideal testbed for evaluating classification consistency over homogeneous regions.
- WHU-Hi-HongHu (WHUHH): Acquired in Honghu City, Hubei Province, using a UAV platform equipped with a Nano-Hyperspec-VNIR imaging spectrometer, this dataset features an ultra-high spatial resolution of 0.043 m. The image measures pixels and includes 270 spectral bands. Due to the extremely high resolution, individual objects often span hundreds of pixels with fine-grained textures, requiring the model to effectively capture both long-range dependencies and local features.
4.2. Experimental Setup and Evaluation Metrics
4.3. Comparison with State-of-the-Art Methods
4.3.1. Qualitative Analysis
4.3.2. Quantitative Analysis
5. Discussion
5.1. Robustness and Generalization Under Limited Training Samples
5.2. Accuracy-Efficiency Trade-Off and Deployment Potential
5.3. Class-Wise Error Patterns and Failure Cases
5.4. Contribution of Key Architectural Components
5.4.1. Role of the Spectral–Spatial Residual Network
5.4.2. Role of the Local Enhanced Mamba Branch
5.5. Limitations and Future Work
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Hong, D.F.; He, W.; Yokoya, N.; Yao, J.; Gao, L.R.; Zhang, L.P.; Chanussot, J.; Zhu, X.X. Interpretable Hyperspectral Artificial Intelligence: When nonconvex modeling meets hyperspectral remote sensing. IEEE Geosci. Remote Sens. Mag. 2021, 9, 52–87. [Google Scholar] [CrossRef]
- Wang, D.; Hu, M.Q.; Jin, Y.; Miao, Y.C.; Yang, J.Q.; Xu, Y.C.; Qin, X.L.; Ma, J.Q.; Sun, L.Y.; Li, C.X.; et al. HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 6427–6444. [Google Scholar] [CrossRef]
- Li, S.; Wang, M.; Cheng, C.; Gao, X.; Ye, Z.; Liu, W. Spectral-Spatial-Sensorial Attention Network with Controllable Factors for Hyperspectral Image Classification. Remote Sens. 2024, 16, 1253. [Google Scholar] [CrossRef]
- Duan, P.H.; Shan, T.C.; Kang, X.D.; Li, S.T. Spectral Super-Resolution in Frequency Domain. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 12338–12348. [Google Scholar] [CrossRef]
- Obermeier, W.A.; Lehnert, L.W.; Pohl, M.; Gianonni, S.M.; Silva, B.; Seibert, R.; Laser, H.; Moser, G.; Müller, C.; Luterbacher, J. Grassland ecosystem services in a changing environment: The potential of hyperspectral monitoring. Remote Sens. Environ. 2019, 232, 111273. [Google Scholar] [CrossRef]
- Li, Z.; Chen, B.; Wu, S.; Su, M.; Chen, J.M.; Xu, B. Deep learning for urban land use category classification: A review and experimental assessment. Remote Sens. Environ. 2024, 311, 114290. [Google Scholar] [CrossRef]
- Siebels, K.; Goïta, K.; Germain, M. Estimation of mineral abundance from hyperspectral data using a new supervised neighbor-band ratio unmixing approach. IEEE Trans. Geosci. Remote Sens. 2020, 58, 6754–6766. [Google Scholar] [CrossRef]
- Liu, Y.; Fan, Y.; Feng, H.; Chen, R.; Bian, M.; Ma, Y.; Yue, J.; Yang, G. Estimating potato above-ground biomass based on vegetation indices and texture features constructed from sensitive bands of UAV hyperspectral imagery. Comput. Electron. Agric. 2024, 220, 108918. [Google Scholar] [CrossRef]
- Sahadevan, A.S. Extraction of spatial-spectral homogeneous patches and fractional abundances for field-scale agriculture monitoring using airborne hyperspectral images. Comput. Electron. Agric. 2021, 188, 106325. [Google Scholar] [CrossRef]
- Wang, C.; Liu, B.; Liu, L.; Zhu, Y.; Hou, J.; Liu, P.; Li, X. A review of deep learning used in the hyperspectral image analysis for agriculture. Artif. Intell. Rev. 2021, 54, 5205–5253. [Google Scholar] [CrossRef]
- Lu, B.; Dao, P.D.; Liu, J.; He, Y.; Shang, J. Recent advances of hyperspectral imaging technology and applications in agriculture. Remote Sens. 2020, 12, 2659. [Google Scholar] [CrossRef]
- Li, S.; Song, W.; Fang, L.; Chen, Y.; Ghamisi, P.; Benediktsson, J.A. Deep learning for hyperspectral image classification: An overview. IEEE Trans. Geosci. Remote Sens. 2019, 57, 6690–6709. [Google Scholar] [CrossRef]
- Melgani, F.; Bruzzone, L. Classification of hyperspectral remote sensing images with support vector machines. IEEE Trans. Geosci. Remote Sens. 2004, 42, 1778–1790. [Google Scholar] [CrossRef]
- Joelsson, S.R.; Benediktsson, J.A.; Sveinsson, J.R. Random forest classifiers for hyperspectral data. In Proceedings of the 2005 IEEE International Geoscience and Remote Sensing Symposium, 2005, IGARSS’05, Seoul, Republic of Korea, 29 July 2005; p. 4. [Google Scholar]
- Li, J.; Bioucas-Dias, J.M.; Plaza, A. Semisupervised hyperspectral image segmentation using multinomial logistic regression with active learning. IEEE Trans. Geosci. Remote Sens. 2010, 48, 4085–4098. [Google Scholar] [CrossRef]
- Farrell, M.D.; Mersereau, R.M. On the impact of PCA dimension reduction for hyperspectral detection of difficult targets. IEEE Geosci. Remote Sens. Lett. 2005, 2, 192–195. [Google Scholar] [CrossRef]
- Bandos, T.V.; Bruzzone, L.; Camps-Valls, G. Classification of hyperspectral images with regularized linear discriminant analysis. IEEE Trans. Geosci. Remote Sens. 2009, 47, 862–873. [Google Scholar] [CrossRef]
- Zhao, W.; Du, S. Spectral–spatial feature extraction for hyperspectral image classification: A dimension reduction and deep learning approach. IEEE Trans. Geosci. Remote Sens. 2016, 54, 4544–4554. [Google Scholar] [CrossRef]
- Mou, L.; Ghamisi, P.; Zhu, X.X. Unsupervised spectral–spatial feature learning via deep residual Conv–Deconv network for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2017, 56, 391–406. [Google Scholar] [CrossRef]
- Hu, W.; Huang, Y.; Wei, L.; Zhang, F.; Li, H. Deep convolutional neural networks for hyperspectral image classification. J. Sens. 2015, 2015, 258619. [Google Scholar] [CrossRef]
- Chen, Y.; Zhu, L.; Ghamisi, P.; Jia, X.; Li, G.; Tang, L. Hyperspectral images classification with Gabor filtering and convolutional neural network. IEEE Geosci. Remote Sens. Lett. 2017, 14, 2355–2359. [Google Scholar] [CrossRef]
- Zhong, Z.; Li, J.; Luo, Z.; Chapman, M. Spectral–spatial residual network for hyperspectral image classification: A 3-D deep learning framework. IEEE Trans. Geosci. Remote Sens. 2017, 56, 847–858. [Google Scholar] [CrossRef]
- Roy, S.K.; Krishna, G.; Dubey, S.R.; Chaudhuri, B.B. HybridSN: Exploring 3-D–2-D CNN feature hierarchy for hyperspectral image classification. IEEE Geosci. Remote Sens. Lett. 2019, 17, 277–281. [Google Scholar] [CrossRef]
- Gong, Z.; Zhong, P.; Yu, Y.; Hu, W.; Li, S. A CNN with multiscale convolution and diversified metric for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2019, 57, 3599–3618. [Google Scholar] [CrossRef]
- Zhu, M.Z.; Fan, J.Y.; Yang, Q.H.; Chen, T. SC-EADNet: A Self-Supervised Contrastive Efficient Asymmetric Dilated Network for Hyperspectral Image Classification. Ieee Trans. Geosci. Remote Sens. 2022, 60, 17. [Google Scholar] [CrossRef]
- Yang, J.; Wu, C.; Du, B.; Zhang, L. Enhanced multiscale feature fusion network for HSI classification. IEEE Trans. Geosci. Remote Sens. 2021, 59, 10328–10347. [Google Scholar] [CrossRef]
- Yang, J.; Du, B.; Xu, Y.; Zhang, L. Can spectral information work while extracting spatial distribution?—An online spectral information compensation network for HSI classification. IEEE Trans. Image Process. 2023, 32, 2360–2373. [Google Scholar] [CrossRef] [PubMed]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Hong, D.; Han, Z.; Yao, J.; Gao, L.; Zhang, B.; Plaza, A.; Chanussot, J. SpectralFormer: Rethinking hyperspectral image classification with transformers. IEEE Trans. Geosci. Remote Sens. 2021, 60, 5518615. [Google Scholar] [CrossRef]
- Sun, L.; Zhao, G.; Zheng, Y.; Wu, Z. Spectral–spatial feature tokenization transformer for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5522214. [Google Scholar] [CrossRef]
- Roy, S.K.; Deria, A.; Shah, C.; Haut, J.M.; Du, Q.; Plaza, A. Spectral–spatial morphological attention transformer for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5503615. [Google Scholar] [CrossRef]
- Mei, S.; Song, C.; Ma, M.; Xu, F. Hyperspectral image classification using group-aware hierarchical transformer. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5539014. [Google Scholar] [CrossRef]
- Zhao, Z.; Xu, X.; Li, S.; Plaza, A. Hyperspectral image classification using groupwise separable convolutional vision transformer network. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5511817. [Google Scholar] [CrossRef]
- Xu, Y.; Wang, D.; Zhang, L.; Zhang, L. Dual selective fusion transformer network for hyperspectral image classification. Neural Netw. 2025, 187, 107311. [Google Scholar] [CrossRef] [PubMed]
- Ouyang, E.; Li, B.; Hu, W.; Zhang, G.; Zhao, L.; Wu, J. When Multigranularity Meets Spatial-Spectral Attention: A Hybrid Transformer for Hyperspectral Image Classification. IEEE Trans. Geosci. Remote Sens. 2023, 61, 4401118. [Google Scholar] [CrossRef]
- Fu, C.; Zhou, T.; Guo, T.; Zhu, Q.; Luo, F.; Du, B. CNN-Transformer and Channel-Spatial Attention based network for hyperspectral image classification with few samples. Neural Netw. 2025, 186, 107283. [Google Scholar] [CrossRef]
- Yang, J.; Du, B.; Zhang, L. From center to surrounding: An interactive learning framework for hyperspectral image classification. ISPRS-J. Photogramm. Remote Sens. 2023, 197, 145–166. [Google Scholar] [CrossRef]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv 2024, arXiv:2401.09417. [Google Scholar] [CrossRef]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. Vmamba: Visual state space model. Adv. Neural Inf. Process. Syst. 2024, 37, 103031–103063. [Google Scholar]
- Li, Y.P.; Luo, Y.; Zhang, L.F.; Wang, Z.M.; Du, B. MambaHSI: Spatial-Spectral Mamba for Hyperspectral Image Classification. IEEE Trans. Geosci. Remote Sens. 2024, 62, 16. [Google Scholar] [CrossRef]
- He, Y.; Tu, B.; Liu, B.; Li, J.; Plaza, A. 3DSS-Mamba: 3D-spectral-spatial mamba for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5534216. [Google Scholar] [CrossRef]
- Wang, G.; Zhang, X.; Peng, Z.; Zhang, T.; Jiao, L. S 2 Mamba: A spatial–spectral state space model for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5511413. [Google Scholar]
- Wang, H.; Zhuang, P.; Zhang, X.; Li, J. DBMGNet: A dual-branch mamba-GCN network for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4410517. [Google Scholar] [CrossRef]
- Yang, A.; Li, M.; Ding, Y.; Fang, L.; Cai, Y.; He, Y. GraphMamba: An efficient graph structure learning vision mamba for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5537414. [Google Scholar] [CrossRef]
- Tang, H.; Huang, Z.; Li, Y.; Zhang, L.; Xie, W. A multiscale spatial–spectral prototypical network for hyperspectral image few-shot classification. IEEE Geosci. Remote Sens. Lett. 2022, 19, 6011205. [Google Scholar] [CrossRef]
- Tang, H.; Wu, Y.; Li, H.; Tang, D.; Yang, X.; Xie, W. Global–local prototype-based few-shot learning for cross-domain hyperspectral image classification. Knowl.-Based Syst. 2025, 314, 113199. [Google Scholar] [CrossRef]
- Tang, H.; Zhang, C.; Tang, D.; Lin, X.; Yang, X.; Xie, W. Few-shot hyperspectral image classification with deep fuzzy metric learning. IEEE Geosci. Remote Sens. Lett. 2025, 22, 5502205. [Google Scholar] [CrossRef]
- Yang, J.; Du, B.; Wang, D.; Zhang, L. ITER: Image-to-pixel representation for weakly supervised HSI classification. IEEE Trans. Image Process. 2023, 33, 257–272. [Google Scholar] [CrossRef]









| Class | SpectralFormer | GAHT | SSFTT | 3DSS-Mamba | GSCVIT | DualMambaFormer (Ours) |
|---|---|---|---|---|---|---|
| 1 | 66.00 ± 11.47 | 88.39 ± 9.29 | 77.11 ± 17.13 | 86.27 ± 3.88 | 93.48 ± 3.62 | 98.76 ± 0.80 |
| 2 | 77.25 ± 8.09 | 89.27 ± 4.43 | 69.92 ± 20.30 | 89.97 ± 6.17 | 91.69 ± 6.45 | 99.08 ± 1.04 |
| 3 | 80.32 ± 11.09 | 87.25 ± 6.23 | 60.63 ± 14.57 | 86.32 ± 5.04 | 94.09 ± 4.18 | 99.60 ± 0.43 |
| 4 | 91.10 ± 2.12 | 96.87 ± 1.65 | 97.26 ± 2.03 | 90.23 ± 8.46 | 93.61 ± 4.88 | 95.63 ± 0.88 |
| 5 | 99.86 ± 0.21 | 99.85 ± 0.19 | 99.85 ± 0.22 | 99.10 ± 0.87 | 99.86 ± 0.15 | 99.84 ± 0.14 |
| 6 | 77.09 ± 11.96 | 92.96 ± 5.98 | 77.10 ± 16.86 | 93.34 ± 3.42 | 93.86 ± 4.74 | 99.93 ± 0.18 |
| 7 | 75.66 ± 10.84 | 97.07 ± 3.82 | 78.88 ± 20.94 | 94.68 ± 3.94 | 98.53 ± 1.17 | 99.95 ± 0.11 |
| 8 | 62.86 ± 20.38 | 88.08 ± 5.43 | 76.74 ± 20.03 | 88.42 ± 6.92 | 95.12 ± 1.94 | 99.12 ± 0.67 |
| 9 | 96.80 ± 1.23 | 99.35 ± 0.49 | 98.83 ± 1.95 | 98.27 ± 1.26 | 99.64 ± 0.44 | 98.52 ± 0.77 |
| OA (%) | 76.46 ± 2.36 | 90.69 ± 2.44 | 75.79 ± 9.10 | 90.10 ± 2.31 | 93.40 ± 2.43 | 98.95 ± 0.55 |
| (%) | 69.94 ± 2.64 | 87.91 ± 3.09 | 69.92 ± 9.98 | 87.10 ± 2.84 | 91.39 ± 3.05 | 98.61 ± 0.73 |
| AA (%) | 80.77 ± 1.91 | 93.23 ± 1.38 | 81.81 ± 4.10 | 91.84 ± 1.75 | 95.54 ± 1.07 | 98.94 ± 0.23 |
| Class | SpectralFormer | GAHT | SSFTT | 3DSS-Mamba | GSCVIT | DualMambaFormer (Ours) |
|---|---|---|---|---|---|---|
| 1 | 90.00 ± 7.83 | 100.00 ± 0.00 | 96.77 ± 4.33 | 93.23 ± 7.28 | 99.03 ± 1.48 | 96.56 ± 0.98 |
| 2 | 68.72 ± 3.71 | 88.08 ± 3.69 | 84.45 ± 7.24 | 65.83 ± 6.84 | 89.47 ± 2.85 | 99.68 ± 0.97 |
| 3 | 82.14 ± 6.69 | 93.76 ± 3.85 | 79.05 ± 26.67 | 76.72 ± 11.39 | 96.53 ± 2.67 | 94.38 ± 4.20 |
| 4 | 94.60 ± 2.40 | 99.63 ± 0.79 | 97.65 ± 2.74 | 97.59 ± 3.79 | 99.89 ± 0.21 | 97.59 ± 1.80 |
| 5 | 90.81 ± 2.73 | 96.56 ± 2.07 | 95.38 ± 2.33 | 88.29 ± 4.64 | 97.23 ± 1.69 | 99.52 ± 0.65 |
| 6 | 95.32 ± 1.80 | 99.34 ± 0.42 | 96.82 ± 5.89 | 94.93 ± 2.89 | 99.43 ± 0.61 | 97.46 ± 1.73 |
| 7 | 98.46 ± 4.62 | 100.00 ± 0.00 | 100.00 ± 0.00 | 99.23 ± 2.31 | 100.00 ± 0.00 | 99.38 ± 0.41 |
| 8 | 97.27 ± 2.36 | 99.86 ± 0.28 | 99.58 ± 0.55 | 97.76 ± 2.84 | 99.93 ± 0.21 | 100.00 ± 0.00 |
| 9 | 98.00 ± 6.00 | 100.00 ± 0.00 | 100.00 ± 0.00 | 100.00 ± 0.00 | 100.00 ± 0.00 | 100.00 ± 0.00 |
| 10 | 81.40 ± 3.86 | 91.91 ± 2.68 | 88.37 ± 5.57 | 74.06 ± 14.17 | 94.03 ± 2.51 | 100.00 ± 0.00 |
| 11 | 73.53 ± 4.42 | 82.13 ± 4.48 | 60.80 ± 25.84 | 66.80 ± 10.69 | 89.50 ± 3.21 | 93.96 ± 2.70 |
| 12 | 74.84 ± 6.65 | 94.97 ± 2.23 | 91.14 ± 3.34 | 83.41 ± 10.83 | 95.82 ± 1.88 | 94.36 ± 2.35 |
| 13 | 99.42 ± 0.54 | 100.00 ± 0.00 | 99.94 ± 0.19 | 99.10 ± 1.36 | 99.94 ± 0.19 | 96.15 ± 1.65 |
| 14 | 92.10 ± 2.82 | 95.62 ± 1.83 | 96.29 ± 2.00 | 93.83 ± 3.15 | 97.77 ± 0.88 | 99.87 ± 0.39 |
| 15 | 93.45 ± 4.24 | 98.72 ± 1.20 | 97.11 ± 1.79 | 94.64 ± 4.24 | 99.38 ± 0.80 | 99.80 ± 0.16 |
| 16 | 99.53 ± 0.93 | 100.00 ± 0.00 | 99.77 ± 0.70 | 99.07 ± 1.14 | 100.00 ± 0.00 | 99.76 ± 0.26 |
| OA (%) | 81.88 ± 1.36 | 91.39 ± 1.32 | 83.47 ± 9.07 | 79.31 ± 3.65 | 94.26 ± 0.89 | 96.56 ± 0.98 |
| (%) | 79.43 ± 1.51 | 90.17 ± 1.48 | 81.43 ± 9.94 | 76.62 ± 4.01 | 93.43 ± 1.01 | 96.06 ± 1.12 |
| AA (%) | 89.35 ± 1.10 | 96.29 ± 0.45 | 92.70 ± 3.71 | 89.03 ± 2.36 | 97.37 ± 0.39 | 98.17 ± 0.51 |
| Class | SpectralFormer | GAHT | SSFTT | 3DSS-Mamba | GSCVIT | DualMambaFormer (Ours) |
|---|---|---|---|---|---|---|
| 1 | 95.02 ± 1.06 | 100.00 ± 0.00 | 99.96 ± 0.12 | 99.91 ± 0.09 | 99.80 ± 0.45 | 100.00 ± 0.00 |
| 2 | 99.35 ± 0.48 | 99.99 ± 0.02 | 99.68 ± 0.61 | 99.89 ± 0.15 | 99.76 ± 0.31 | 100.00 ± 0.00 |
| 3 | 96.08 ± 2.06 | 99.63 ± 0.32 | 99.00 ± 0.65 | 98.52 ± 2.75 | 99.84 ± 0.35 | 98.53 ± 4.39 |
| 4 | 97.97 ± 0.79 | 99.83 ± 0.24 | 99.93 ± 0.09 | 99.32 ± 0.60 | 99.80 ± 0.24 | 99.78 ± 0.38 |
| 5 | 92.68 ± 2.63 | 99.17 ± 0.91 | 99.15 ± 0.64 | 98.42 ± 1.98 | 98.73 ± 1.52 | 99.13 ± 0.67 |
| 6 | 99.74 ± 0.41 | 100.00 ± 0.00 | 99.99 ± 0.01 | 99.41 ± 0.75 | 99.93 ± 0.19 | 99.98 ± 0.07 |
| 7 | 97.95 ± 1.27 | 99.94 ± 0.08 | 99.87 ± 0.22 | 98.88 ± 1.20 | 99.97 ± 0.05 | 100.00 ± 0.00 |
| 8 | 80.26 ± 2.21 | 88.40 ± 1.12 | 85.86 ± 2.11 | 85.72 ± 5.83 | 87.33 ± 1.83 | 93.66 ± 3.11 |
| 9 | 97.69 ± 1.01 | 99.80 ± 0.21 | 99.44 ± 0.74 | 98.37 ± 1.34 | 100.00 ± 0.01 | 100.00 ± 0.00 |
| 10 | 92.55 ± 1.81 | 97.92 ± 1.18 | 96.82 ± 1.17 | 95.64 ± 2.75 | 98.17 ± 1.27 | 99.31 ± 0.85 |
| 11 | 93.97 ± 2.86 | 99.78 ± 0.23 | 99.44 ± 0.35 | 99.72 ± 0.32 | 99.43 ± 0.46 | 99.99 ± 0.03 |
| 12 | 97.91 ± 1.55 | 99.99 ± 0.02 | 99.88 ± 0.11 | 98.49 ± 1.93 | 99.79 ± 0.25 | 99.97 ± 0.03 |
| 13 | 99.97 ± 0.10 | 100.00 ± 0.00 | 99.70 ± 0.31 | 99.28 ± 0.98 | 99.99 ± 0.03 | 100.00 ± 0.00 |
| 14 | 99.16 ± 0.56 | 99.74 ± 0.20 | 99.29 ± 0.84 | 98.38 ± 1.89 | 99.39 ± 1.17 | 99.96 ± 0.09 |
| 15 | 80.27 ± 3.50 | 89.70 ± 1.50 | 82.71 ± 4.23 | 88.66 ± 8.16 | 91.25 ± 2.57 | 93.23 ± 9.05 |
| 16 | 97.34 ± 1.61 | 99.32 ± 0.60 | 98.28 ± 1.19 | 98.96 ± 1.72 | 99.04 ± 0.64 | 99.99 ± 0.02 |
| OA (%) | 91.23 ± 0.50 | 95.92 ± 0.26 | 94.22 ± 0.59 | 94.60 ± 1.26 | 95.87 ± 0.38 | 97.60 ± 0.95 |
| (%) | 90.25 ± 0.55 | 95.46 ± 0.29 | 93.57 ± 0.66 | 93.99 ± 1.41 | 95.41 ± 0.42 | 97.33 ± 1.06 |
| AA (%) | 94.87 ± 0.34 | 98.33 ± 0.13 | 97.44 ± 0.22 | 97.35 ± 0.90 | 98.26 ± 0.18 | 98.97 ± 0.53 |
| Class | SpectralFormer | GAHT | SSFTT | 3DSS-Mamba | GSCVIT | DualMambaFormer (Ours) |
|---|---|---|---|---|---|---|
| 1 | 95.28 ± 1.04 | 96.13 ± 0.81 | 82.37 ± 15.01 | 88.38 ± 5.95 | 93.46 ± 5.86 | 96.11 ± 0.58 |
| 2 | 90.01 ± 2.12 | 96.23 ± 1.47 | 69.44 ± 31.34 | 91.15 ± 4.82 | 95.07 ± 4.31 | 97.67 ± 1.41 |
| 3 | 78.32 ± 1.68 | 84.00 ± 8.94 | 80.07 ± 11.96 | 82.10 ± 5.97 | 89.40 ± 2.69 | 92.83 ± 2.64 |
| 4 | 93.20 ± 2.40 | 95.46 ± 0.70 | 84.41 ± 11.17 | 86.76 ± 6.24 | 95.56 ± 5.62 | 98.10 ± 0.68 |
| 5 | 85.74 ± 5.17 | 93.45 ± 2.38 | 72.78 ± 29.87 | 85.04 ± 7.13 | 90.68 ± 0.77 | 98.34 ± 1.27 |
| 6 | 90.41 ± 2.92 | 94.12 ± 2.52 | 66.43 ± 29.85 | 90.97 ± 2.07 | 93.32 ± 2.90 | 96.88 ± 1.20 |
| 7 | 70.49 ± 3.32 | 78.60 ± 3.92 | 69.35 ± 14.28 | 62.44 ± 8.83 | 80.01 ± 5.05 | 88.78 ± 2.19 |
| 8 | 63.43 ± 3.11 | 80.02 ± 4.75 | 64.45 ± 12.07 | 57.92 ± 9.37 | 83.12 ± 3.24 | 95.95 ± 1.42 |
| 9 | 96.94 ± 0.86 | 97.27 ± 1.74 | 95.38 ± 3.40 | 94.25 ± 2.73 | 97.55 ± 1.37 | 97.85 ± 1.74 |
| 10 | 61.54 ± 5.74 | 89.57 ± 2.24 | 81.11 ± 8.85 | 72.52 ± 6.10 | 82.34 ± 11.69 | 93.23 ± 5.29 |
| 11 | 71.51 ± 3.36 | 85.85 ± 3.57 | 55.64 ± 19.38 | 78.03 ± 7.08 | 87.65 ± 3.80 | 96.25 ± 1.01 |
| 12 | 70.96 ± 4.17 | 79.35 ± 5.21 | 56.33 ± 20.32 | 69.13 ± 5.68 | 81.50 ± 9.10 | 92.50 ± 2.77 |
| 13 | 67.83 ± 5.20 | 78.49 ± 4.74 | 67.83 ± 5.20 | 64.14 ± 8.39 | 78.77 ± 4.33 | 88.07 ± 2.81 |
| 14 | 85.54 ± 3.17 | 94.55 ± 2.26 | 85.43 ± 8.74 | 85.41 ± 3.68 | 92.92 ± 5.11 | 98.54 ± 1.26 |
| 15 | 94.02 ± 2.90 | 98.26 ± 1.01 | 97.13 ± 2.83 | 94.59 ± 3.02 | 98.89 ± 0.66 | 99.76 ± 0.51 |
| 16 | 89.36 ± 3.51 | 94.82 ± 1.69 | 91.33 ± 3.20 | 84.57 ± 9.00 | 98.31 ± 1.48 | 99.06 ± 1.47 |
| 17 | 89.21 ± 2.37 | 96.65 ± 1.96 | 92.08 ± 3.87 | 88.33 ± 4.74 | 95.30 ± 5.49 | 98.51 ± 3.06 |
| 18 | 90.66 ± 3.82 | 95.68 ± 2.52 | 64.21 ± 25.62 | 92.29 ± 4.39 | 96.14 ± 2.73 | 99.10 ± 0.57 |
| 19 | 90.03 ± 1.94 | 94.89 ± 0.94 | 67.38 ± 24.16 | 85.48 ± 5.92 | 94.21 ± 2.12 | 96.56 ± 1.07 |
| 20 | 94.45 ± 2.00 | 95.69 ± 2.37 | 75.31 ± 22.05 | 89.94 ± 6.66 | 97.42 ± 1.69 | 99.57 ± 0.30 |
| 21 | 87.86 ± 5.67 | 93.64 ± 12.16 | 55.38 ± 36.65 | 85.45 ± 10.67 | 89.48 ± 11.80 | 99.84 ± 0.49 |
| 22 | 91.54 ± 3.50 | 95.71 ± 2.00 | 88.84 ± 9.76 | 89.14 ± 7.58 | 97.14 ± 1.80 | 98.99 ± 0.73 |
| OA (%) | 86.36 ± 1.05 | 91.65 ± 0.85 | 77.17 ± 9.41 | 83.05 ± 3.54 | 91.79 ± 2.32 | 96.09 ± 0.38 |
| (%) | 83.03 ± 1.24 | 89.55 ± 1.05 | 72.46 ± 10.69 | 79.23 ± 4.06 | 89.76 ± 2.73 | 95.08 ± 0.47 |
| AA (%) | 84.01 ± 0.83 | 91.29 ± 0.90 | 75.03 ± 9.24 | 82.64 ± 2.20 | 91.28 ± 1.23 | 96.48 ± 0.43 |
| Dataset | Method | Params (M) | FLOPs (M) | Latency (ms/Sample) | Throughput (Samples/s) |
|---|---|---|---|---|---|
| PU | SpectralFormer | 0.184 | 15.73 | 0.217 | 4613 |
| SSFTT | 0.153 | 11.4 | 0.083 | 11,981 | |
| GAHT | 0.927 | 45.41 | 0.194 | 5158 | |
| GSCVIT | 0.153 | 4.96 | 0.26 | 3841 | |
| 3DSS-Mamba | 0.024 | 13.96 | 0.137 | 7283 | |
| DualMambaFormer | 1.087 | 428.2 | 0.148 | 6771 | |
| IP | SpectralFormer | 0.356 | 35.41 | 0.216 | 4622 |
| SSFTT | 0.153 | 11.4 | 0.083 | 12,045 | |
| GAHT | 0.831 | 40.67 | 0.202 | 4955 | |
| GSCVIT | 0.638 | 20.96 | 0.28 | 3568 | |
| 3DSS-Mamba | 0.024 | 13.96 | 0.137 | 7297 | |
| DualMambaFormer | 1.101 | 435.97 | 0.203 | 4925 | |
| SA | SpectralFormer | 0.366 | 36.43 | 0.999 | 1001 |
| SSFTT | 0.153 | 11.4 | 0.073 | 13,739 | |
| GAHT | 0.973 | 47.61 | 0.204 | 4893 | |
| GSCVIT | 0.179 | 6.62 | 0.294 | 3407 | |
| 3DSS-Mamba | 0.024 | 13.96 | 0.134 | 7446 | |
| DualMambaFormer | 1.101 | 436.29 | 0.169 | 5927 | |
| WHUHH | SpectralFormer | 0.559 | 55.02 | 0.213 | 4694 |
| SSFTT | 0.154 | 11.4 | 0.08 | 12,428 | |
| GAHT | 1.515 | 74.15 | 0.299 | 3342 | |
| GSCVIT | 0.249 | 10.25 | 0.295 | 3389 | |
| 3DSS-Mamba | 0.025 | 13.96 | 0.139 | 7212 | |
| DualMambaFormer | 1.111 | 441.57 | 0.198 | 5059 |
| MHSA | LEM | SS-ResNet | PU | IP | SA | WHUHH |
|---|---|---|---|---|---|---|
| √ | √ | 93.89 ± 1.18 | 85.04 ± 0.96 | 93.42 ± 1.78 | 88.36 ± 1.08 | |
| √ | √ | 97.84 ± 0.39 | 95.11 ± 0.52 | 95.95 ± 0.67 | 92.92 ± 0.79 | |
| √ | √ | √ | 98.95 ± 0.57 | 96.56 ± 0.55 | 97.60 ± 0.95 | 96.09 ± 0.38 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yu, J.; Li, J.; Sun, G.; Lu, J.; Cheng, X.; Zhou, R.; Sun, W.; Gao, X. DualMambaFormer: A Parallel Hybrid Transformer–Mamba Network for Hyperspectral Image Classification. Remote Sens. 2026, 18, 1516. https://doi.org/10.3390/rs18101516
Yu J, Li J, Sun G, Lu J, Cheng X, Zhou R, Sun W, Gao X. DualMambaFormer: A Parallel Hybrid Transformer–Mamba Network for Hyperspectral Image Classification. Remote Sensing. 2026; 18(10):1516. https://doi.org/10.3390/rs18101516
Chicago/Turabian StyleYu, Jiang, Jingwei Li, Gan Sun, Jingying Lu, Xuejun Cheng, Ruimeng Zhou, Wei Sun, and Xianjun Gao. 2026. "DualMambaFormer: A Parallel Hybrid Transformer–Mamba Network for Hyperspectral Image Classification" Remote Sensing 18, no. 10: 1516. https://doi.org/10.3390/rs18101516
APA StyleYu, J., Li, J., Sun, G., Lu, J., Cheng, X., Zhou, R., Sun, W., & Gao, X. (2026). DualMambaFormer: A Parallel Hybrid Transformer–Mamba Network for Hyperspectral Image Classification. Remote Sensing, 18(10), 1516. https://doi.org/10.3390/rs18101516

