DPMSCH-Net: A Dual Prior-Guidance Multi-Scale SSM-CNN Hybrid Network with Manifold Constraints for Hyperspectral Image Classification Under Small-Sample Conditions
Highlights
- DPMSCH-Net with prior-guided DMB captures spatial–spectral information under data scarcity, and SupCon regularizationmitigates feature collapse through manifold constraints.
- It achieves state-of-the-art five-shot classification performance across HanChuan, HongHu, and XiongAn datasets.
- It offers a robust framework for real-world agricultural HSI monitoring with minimal annotation costs.
- It demonstrates the potential of injecting physical priors into sequence modeling for few-shot tasks.
Abstract
1. Introduction
- (1)
- We propose DPMSCH-Net, an asymmetric SSM–CNN hybrid backbone designed for few-shot learning for hyperspectral image classification. By progressively integrating dual Mamba blocks (DMBs), hierarchical context modules (HCMs), and CNNs across different decreasing resolutions, the network effectively captures local-to-global spatial–spectral contexts.
- (2)
- We design a dual Mamba block (DMB) guided by domain-specific priors to significantly improve feature discriminability under extreme data scarcity. Aided by base-weight constraints, the constructed Gaussian spatial prior and PCA-inspired spectral decay prior guide the model to precisely focus on central targets while emphasizing informative spectral representations, without excessively suppressing useful peripheral spatial information or informative spectral cues.
- (3)
- To mitigate feature collapse under extreme label scarcity, we introduce a feature-level supervised contrastive (SupCon) regularization to impose explicit manifold constraints. Via implicit hard sample mining, it avoids tedious pair construction and explicitly enforces intra-class compactness and inter-class dispersion.
2. Related Work
2.1. CNN-Based Methods in HSIC
2.2. Mamba-Based Methods in HSIC
2.3. Spectral–Spatial Methods in HSIC
2.4. Contrastive Learning in HSIC
3. Proposed Method
3.1. Overall Architecture
3.2. Dual Mamba Block
3.3. Feature-Level Contrastive Regularization
4. Experimental Section
4.1. Datasets
- (A)
- HongHu: The HongHu dataset is an HS scene collected over the Honghu area and is frequently adopted for HSI analysis tasks. The image covers an area of 940 × 475 pixels at a fine spatial resolution of 0.043 m. It contains 270 spectral channels ranging from 400 to 1000 nm and provides labeled samples for 22 different land-cover classes. A visual summary of the dataset is provided in Figure 3a, showing the original HSI and the corresponding ground truth.
- (B)
- HanChuan: HanChuan is a representative HS benchmark dataset acquired over Hanchuan City. The dataset provides an HSI of size 1217 × 303 pixels with a spatial resolution of 0.109 m. It consists of 274 spectral bands spanning the wavelength range from 400 to 1000 nm and includes ground-truth annotations for 16 land-cover categories. Figure 3b offers a visual overview of the dataset.
- (C)
- XiongAn: The XiongAn dataset comprises an HSI with a spatial size of 3750 × 1580 pixels. The image contains 250 spectral bands covering the 400–1000 nm wavelength range, with a spatial resolution of 0.5 m. The dataset includes 19 annotated land-cover classes, which are mainly composed of economic crops. A visual summary of the dataset, paired with its ground truth, is provided in Figure 3c.
4.2. Experiment Setting
- (A)
- Evaluation metrics: In the experiments, all performance metrics are assessed on the test set. Three widely used evaluation indicators for HSI classification are adopted, including Overall Accuracy (OA), Average Accuracy (AA), and the Kappa coefficient [54]. For clarity of presentation in the subsequent tables, the values of OA and AA are expressed as percentages (%), and the Kappa coefficient is multiplied by 100.
- (B)
- Configuration: All experiments were conducted on a workstation equipped with four NVIDIA Tesla V100-SXM2-32GB GPUs0 and an Intel(R) Xeon(R) Gold 6230 CPU. The operating system was Ubuntu 22.04.4 LTS, and the deep learning framework was implemented using PyTorch 2.7.1 with CUDA 11.8.
- (C)
- Implementation details: To simulate an FSL scenario, exactly five labeled samples per class were randomly selected to construct the training set, while all remaining samples were reserved for testing. All experiments were independently conducted ten times using different random seeds (0–9) to mitigate the bias caused by this random selection, with the final results reported as mean ± standard deviation. The training process used the Adam optimizer with a learning rate of and was run for 100 epochs. Furthermore, a batch size of 64 was adopted for the effective optimization of Equation (16) under the 5-shot setting. In the HanChuan, XiongAn, and HongHu datasets, each class contains only 5 labeled samples, with the total training pools strictly limited to 80, 95, and 110 samples, respectively. With random sampling at a batch size of 64, the expected number of samples per class in a single mini-batch is approximately 4.00, 3.37, and 2.91, ensuring that most classes obtain 2 to 4 positive anchors per batch. Meanwhile, since each class contains a maximum of 5 samples, this batch size guarantees that any given sample will have at least 59 negative samples within the mini-batch. Furthermore, the SupCon loss effectively accommodates rare edge cases—where only a single sample from a class is drawn due to sampling fluctuations—enabling smooth training without triggering division-by-zero errors. Therefore, setting the batch size to 64 not only effectively prevents feature-space collapse but also avoids the irrationality of using a batch size larger than the total training set.
4.3. Parameter Analysis
- (A)
- The retained dimensions of PCA (c)PCA dimensionality reduction aims to preserve the main discriminative spectral features in HS data, meanwhile eliminating redundant information and noise. With the size of input patch cubes fixed at , we evaluated the impact of c (ranging from 20 to 80 with a step of 5) on model performance, and the experimental results are shown in Figure 4. It can be observed that on the HongHu and HanChuan datasets, the model performance improves steadily and approaches the optimum at 40 dimensions, while further increasing dimensions leads to a drop in accuracy. On the XiongAn dataset, the effect of PCA dimensions on performance is relatively slight and accuracy stays stable.This observation suggests that the sensitivity of model performance to the number of retained PCA dimensions is highly data-dependent. When spectral bands exhibit redundancy, the choice of dimensionality becomes critical; conversely, when the spectral structure is inherently compact, the model exhibits stronger robustness to dimensional variation. In practice, for datasets like HongHu and HanChuan, where significant inter-band redundancy exists and discriminative information is distributed across numerous spectral channels, retaining too few dimensions forfeits critical information, while excessive retention introduces noise—hence the pronounced sensitivity to the PCA dimension. In contrast, for spectrally compact data such as XiongAn, where the majority of discriminative energy is concentrated in the first few principal components, performance remains stable as long as the retained dimension exceeds a certain minimum threshold; further increasing the dimensionality yields marginal additional information. This phenomenon is essentially attributable to the eigenvalue decay characteristics of the HSI covariance matrix.
- (B)
- The input patch size (w)We investigated the effect of the size of input patch cubes w with the number of retained PCA dimensions c fixed at 40. A series of input patch sizes from to with a step of 2 are tested, and the results are shown in Figure 5. All patch sizes are odd to ensure that the central pixel is the target pixel. On the HongHu and HanChuan datasets, the classification performance exhibits an overall upward trend as the spatial context expands, peaking at . Beyond this optimal size, the accuracy fluctuates and even degrades. This is because an excessively large receptive field introduces irrelevant background noise and heterogeneous pixels, which mislead the model, alongside a sharp rise in computational overhead. On the XiongAn dataset, accuracy rises continuously but much more slowly beyond , along with excessive computational cost.The spatial resolutions of HongHu and HanChuan are relatively high, at 0.043 m and 0.109 m, respectively. When w = 19, the receptive field fits neatly around the central pixel, capturing a purely homogeneous local context. But because agricultural plots in these scenes are often small and fragmented, enlarging the patch further tends to cross object boundaries and pull in a lot of heterogeneous pixels—background or other crop types. This inter-class interference dilutes the spectral purity of the target and severely misleads the model. In contrast, XiongAn has a coarser resolution of 0.5 m, and its crops are mostly distributed in large continuous blocks. Even when the patch size exceeds , the receptive field still largely stays within the same class, so the accuracy continues to rise, albeit slowly. Still, this improvement comes with a sharp increase in computation, making the marginal return from further enlarging the patch rather limited.
- (C)
- The parameter for the PCA spectral prior ()The baseline bias in the PCA spectral prior serves to preserve valid discriminative information in low-energy channels, thereby preventing feature collapse under extreme five training samples condition. To evaluate its influence on model performance, we tested over a range of candidate values—0.01, 0.05, 0.1, 0.2, and 0.3. Figure 6 presents the resulting OA, AA, and Kappa trends on the three datasets as varies. On the three datasets, classification performance improves steadily as increases from 0.01 to 0.1, peaking at the latter value. Further raising it to 0.2 or 0.3, however, leads to a noticeable decline in accuracy.When is set too small, like 0.01, the soft constraint imposed by the baseline bias becomes overly weak, which causes the network to over-suppress low-energy spectral channels. Although these tail channels account for only a small fraction of the overall variance, they may still carry fine-grained spectral cues that help distinguish similar crop types. Over-suppressing them therefore discards critical detail. Conversely, when is too large—0.3, for instance—the baseline bias becomes excessively strong and substantially weakens the guiding effect of the PCA spectral decay prior. The network then loses its ability to focus on the most informative principal bands, and instead assigns excessive weight to redundant bands and spectral noise, which ultimately compromises the discriminability of the extracted features.
- (D)
- The parameter for the Gaussian spatial prior ()The baseline bias in the Gaussian spatial prior is a soft constraint that prevents boundary features from being completely attenuated. To assess its impact on model performance, we tested over candidate values ranging from 0.1 to 0.9 at intervals of 0.2. Figure 7 illustrates the resulting classification performance on the three datasets. The optimal choice of varies slightly across datasets and evaluation metrics. On the HongHu dataset, the model attains its best performance across all metrics at = 0.5. For HanChuan, OA reaches a local peak of 83.02% at 0.7, whereas AA is highest at 0.3. A similar pattern emerges on the XiongAn dataset: although OA and Kappa peak at 0.5, reaching 67.35% and 63.73, respectively, AA shows a slight preference for a smaller bias of 0.3, where it reaches 76.17%.This performance fluctuation suggests that a smaller enables the network to better capture fine-grained features of minority classes—which benefits Average Accuracy—yet it may over-suppress global spatial context. On the other hand, a larger weakens the prior’s ability to focus on central targets because the baseline bias becomes too strong, introducing excessive heterogeneous background noise. Given that spatial heterogeneity and class imbalance vary across different agricultural scenes, the optimal trade-off point shifts accordingly.
4.4. Quantitative and Qualitative Analysis
4.4.1. Quantitative Analysis
4.4.2. Qualitative Analysis
4.5. Ablation Study
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Wu, C.; Feng, R.; Zhang, P.; Mao, H.; Wang, D.; Bai, Z.; Li, Y. MRC-Net: A Multistage Network with Rank-Reduced Multihead Self-Attention and Cascade Learning for Hyperspectral Pansharpening. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 22466–22485. [Google Scholar] [CrossRef]
- Lebedev, L.I.; Yasakov, Y.V.; Utesheva, T.H.; Gromov, V.P.; Borusjak, A.V.; Turlapov, V. Complex analysis and monitoring of the environment based on Earth sensing data. Comput. Opt. 2019, 43, 282–295. [Google Scholar] [CrossRef]
- Chen, X.; Zheng, X.; Zhang, Y.; Lu, X. Remote sensing scene classification by local–global mutual learning. IEEE Geosci. Remote Sens. Lett. 2022, 19, 6506405. [Google Scholar] [CrossRef]
- Pasolli, E.; Melgani, F.; Tuia, D.; Pacifici, F.; Emery, W.J. SVM Active Learning Approach for Image Classification Using Spatial Information. IEEE Trans. Geosci. Remote Sens. 2014, 52, 2217–2233. [Google Scholar] [CrossRef]
- Ma, L.; Crawford, M.M.; Tian, J. Local Manifold Learning-Based k-Nearest-Neighbor for Hyperspectral Image Classification. IEEE Trans. Geosci. Remote Sens. 2010, 48, 4099–4109. [Google Scholar]
- Ham, J.; Chen, Y.; Crawford, M.M.; Ghosh, J. Investigation of the random forest framework for classification of hyperspectral data. IEEE Trans. Geosci. Remote Sens. 2005, 43, 492–501. [Google Scholar] [CrossRef]
- Wu, C.; Wang, D.; Bai, Y.; Mao, H.; Li, Y.; Shen, Q. HSR-Diff: Hyperspectral image super-resolution via conditional diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2023; pp. 7060–7070. [Google Scholar]
- Zhong, Z.; Li, J.; Luo, Z.; Chapman, M. Spectral–spatial residual network for hyperspectral image classification: A 3-D deep learning framework. IEEE Trans. Geosci. Remote Sens. 2018, 56, 847–858. [Google Scholar] [CrossRef]
- Song, W.; Li, S.; Fang, L.; Lu, T. Hyperspectral image classification with deep feature fusion network. IEEE Trans. Geosci. Remote Sens. 2018, 56, 3173–3184. [Google Scholar] [CrossRef]
- Zhao, Z.; Xu, X.; Li, S.; Plaza, A. Hyperspectral Image Classification Using Groupwise Separable Convolutional Vision Transformer Network. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5511817. [Google Scholar] [CrossRef]
- Hong, D.; Han, Z.; Yao, J.; Gao, L.; Zhang, B.; Plaza, A.; Chanussot, J. SpectralFormer: Rethinking Hyperspectral Image Classification with Transformers. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5518615. [Google Scholar] [CrossRef]
- Zhang, J.; Liu, L.; Zhao, R.; Shi, Z. A Bayesian meta-learning-based method for few-shot hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5500613. [Google Scholar] [CrossRef]
- Li, W.; Liu, Q.; Zhang, Y.; Wang, Y.; Yuan, Y.; Jia, Y.; He, Y. Few-shot hyperspectral image classification using meta learning and regularized finetuning. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5529514. [Google Scholar] [CrossRef]
- Wang, W.; Liu, F.; Xiao, L. I2MEP-Net: Inter- and intra-modality enhancing prototypical network for few-shot hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5516815. [Google Scholar] [CrossRef]
- Gao, K.; Liu, B.; Yu, X.; Qin, J.; Zhang, P.; Tan, X. Deep relation network for hyperspectral image few-shot classification. Remote Sens. 2020, 12, 923. [Google Scholar] [CrossRef]
- Zeng, J.; Xue, Z.; Zhang, L.; Lan, Q.; Zhang, M. Multistage relation network with dual-metric for few-shot hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5510017. [Google Scholar] [CrossRef]
- Deng, B.; Jia, S.; Shi, D. Deep metric learning-based feature embedding for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2020, 58, 1422–1435. [Google Scholar] [CrossRef]
- Zhu, L.; Chen, Y.; Ghamisi, P.; Benediktsson, J.A. Generative adversarial networks for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2018, 56, 5046–5063. [Google Scholar] [CrossRef]
- Li, W.; Wu, G.; Zhang, F.; Du, Q. Hyperspectral image classification using deep pixel-pair features. IEEE Trans. Geosci. Remote Sens. 2017, 55, 844–853. [Google Scholar] [CrossRef]
- Nalepa, J.; Myller, M.; Kawulok, M. Training- and test-time data augmentation for hyperspectral image segmentation. IEEE Geosci. Remote Sens. Lett. 2020, 17, 292–296. [Google Scholar] [CrossRef]
- Zhang, H.; Cisse, M.; Dauphin, Y.N.; Lopez-Paz, D. mixup: Beyond empirical risk minimization. In Proceedings of the 6th International Conference on Learning Representations (ICLR); OpenReview.net: Vancouver, BC, Canada, 2018. [Google Scholar]
- Wang, D.; Bai, Y.; Wu, C.; Li, Y.; Shang, C.; Shen, Q. Convolutional LSTM-based hierarchical feature fusion for multispectral pan-sharpening. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5404016. [Google Scholar] [CrossRef]
- Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; Krishnan, D. Supervised contrastive learning. Adv. Neural Inf. Process. Syst. 2020, 33, 18661–18673. [Google Scholar]
- Chen, Y.; Lin, Z.; Zhao, X.; Wang, G.; Gu, Y. Deep learning-based classification of hyperspectral data. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2014, 7, 2094–2107. [Google Scholar] [CrossRef]
- Zhao, W.; Du, S. Spectral–spatial feature extraction for hyperspectral image classification: A dimension reduction and deep learning approach. IEEE Trans. Geosci. Remote Sens. 2016, 54, 4544–4554. [Google Scholar] [CrossRef]
- Hu, W.; Huang, Y.; Wei, L.; Zhang, F.; Li, H. Deep convolutional neural networks for hyperspectral image classification. J. Sens. 2015, 2015, 258619. [Google Scholar] [CrossRef]
- Li, Y.; Zhang, H.; Shen, Q. Spectral–spatial classification of hyperspectral imagery with 3D convolutional neural network. Remote Sens. 2017, 9, 67. [Google Scholar] [CrossRef]
- Roy, S.K.; Krishna, G.; Dubey, S.R.; Chaudhuri, B.B. HybridSN: Exploring 3-D–2-D CNN feature hierarchy for hyperspectral image classification. IEEE Geosci. Remote Sens. Lett. 2020, 17, 277–281. [Google Scholar] [CrossRef]
- Gong, Z.; Zhong, P.; Yu, Y.; Hu, W.; Li, S. A CNN with multiscale convolution and diversified metric for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2019, 57, 3599–3618. [Google Scholar] [CrossRef]
- Paoletti, M.E.; Haut, J.M.; Fernandez-Beltran, R.; Plaza, J.; Plaza, A.J.; Pla, F. Deep pyramidal residual networks for spectral–spatial hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2019, 57, 740–754. [Google Scholar] [CrossRef]
- He, M.; Li, B.; Chen, H. Multi-scale 3D deep convolutional neural network for hyperspectral image classification. In 2017 IEEE International Conference on Image Processing (ICIP); IEEE: Piscataway, NJ, USA, 2017; pp. 3904–3908. [Google Scholar]
- Xu, F.; Mei, S.; Zhang, G.; Wang, N.; Du, Q. Bridging CNN and transformer with cross-attention fusion network for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5522214. [Google Scholar] [CrossRef]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2024, arXiv:2312.00752. [Google Scholar]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision Mamba: Efficient visual representation learning with bidirectional state space model. arXiv 2024, arXiv:2401.09417. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. VMamba: Visual state space model. Adv. Neural Inf. Process. Syst. 2024, 37, 103031–103063. [Google Scholar] [CrossRef]
- Li, Y.; Luo, Y.; Zhang, L.; Wang, Z.; Du, B. MambaHSI: Spatial–spectral Mamba for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5524216. [Google Scholar] [CrossRef]
- Wang, Y.; Liu, L.; Xiao, J.; Yu, D.; Tao, Y.; Zhang, W. MambaHSI+: Multidirectional state propagation for efficient hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4411414. [Google Scholar] [CrossRef]
- Mao, J.; Ma, H.; Liang, Y. BiMambaHSI: Bidirectional spectral–spatial state space model for hyperspectral image classification. Remote Sens. 2025, 17, 3676. [Google Scholar] [CrossRef]
- He, Y.; Tu, B.; Liu, B.; Li, J.; Plaza, A. 3DSS-Mamba: 3D-spectral-spatial Mamba for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5534216. [Google Scholar] [CrossRef]
- Xu, Y.; Han, C.; Chen, S.; Jin, Y.; Miao, Y.; Guo, H.; Wang, D. PHDMamba: Progressive hybrid Mamba for hyperspectral image classification. IEEE Geosci. Remote Sens. Lett. 2025, 22, 5510605. [Google Scholar] [CrossRef]
- Yang, X.; Wei, Y.; Tan, J.; Li, S.; Tang, H.; Liu, W. LMamba: Local-guided Mamba with multi-scale filtering for hyperspectral image classification. Remote Sens. 2026, 18, 1629. [Google Scholar] [CrossRef]
- Fang, Y.; Sun, L.; Zheng, Y.; Wu, Z. Deformable convolution-enhanced hierarchical transformer with spectral–spatial cluster attention for hyperspectral image classification. IEEE Trans. Image Process. 2025, 34, 701–716. [Google Scholar] [CrossRef]
- Xi, B.; Zhang, Y.; Li, J.; Zheng, T.; Zhao, X.; Xu, H.; Xue, C.; Li, Y.; Chanussot, J. MCTGCL: Mixed CNN–Transformer for Mars hyperspectral image classification with graph contrastive learning. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5503214. [Google Scholar] [CrossRef]
- Zhang, J.; Li, L.; Jiao, L.; Liu, X.; Liu, F.; Ma, W.; Yang, S.; Chen, P. PIMamba: Physics-Informed Mamba for Few-Shot Hyperspectral Image Classification. IEEE Trans. Geosci. Remote Sens. 2026, 64, 5511116. [Google Scholar] [CrossRef]
- Li, J.; Wu, H.; Song, R.; Xu, H.; Li, Y.; Du, Q. Physics-Guided Time-Interactive-Frequency Network for Cross-Domain Few-Shot Hyperspectral Image Classification. IEEE Trans. Neural Netw. Learn. Syst. 2026, 37, 438–452. [Google Scholar] [CrossRef] [PubMed]
- He, Y.; Tu, B.; Liu, B.; Li, J.; Plaza, A. HSI-MFormer: Integrating Mamba and Transformer experts for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5621916. [Google Scholar] [CrossRef]
- Sheng, J.; Zhou, J.; Wang, J.; Ye, P.; Fan, J. DualMamba: A Lightweight Spectral–Spatial Mamba-Convolution Network for Hyperspectral Image Classification. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5501415. [Google Scholar] [CrossRef]
- Hou, S.; Shi, H.; Cao, X.; Zhang, X.; Jiao, L. Hyperspectral Imagery Classification Based on Contrastive Learning. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5521213. [Google Scholar] [CrossRef]
- Lee, H.; Kwon, H. Self-Supervised Contrastive Learning for Cross-Domain Hyperspectral Image Representation. In ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing; IEEE: Piscataway, NJ, USA, 2022; pp. 3239–3243. [Google Scholar]
- Ayuba, D.L.A.; Guillemaut, J.Y.; Marti-Cardona, B.; Mendez, O. HyperKon: A Self-Supervised Contrastive Network for Hyperspectral Image Analysis. Remote Sens. 2024, 16, 3399. [Google Scholar] [CrossRef]
- Huang, L.; Chen, Y.; He, X.; Ghamisi, P. Supervised Contrastive Learning-Based Classification for Hyperspectral Image. Remote Sens. 2022, 14, 5530. [Google Scholar] [CrossRef]
- Schroff, F.; Kalenichenko, D.; Philbin, J. FaceNet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2015; pp. 815–823. [Google Scholar]
- Wen, Y.; Zhang, K.; Li, Z.; Qiao, Y. A discriminative feature learning approach for deep face recognition. In Computer Vision—ECCV 2016; Springer International Publishing: Cham, Switzerland, 2016; pp. 499–515. [Google Scholar]
- Li, S.; Song, W.; Fang, L.; Chen, Y.; Ghamisi, P.; Benediktsson, J.A. Deep Learning for Hyperspectral Image Classification: An Overview. IEEE Trans. Geosci. Remote Sens. 2019, 57, 6690–6709. [Google Scholar] [CrossRef]
- Zhang, B.; Chen, Y.; Xiong, S.; Lu, X. Hyperspectral image classification via cascaded spatial cross-attention network. IEEE Trans. Image Process. 2025, 34, 899–913. [Google Scholar] [CrossRef]
- Liu, S.; Fu, C.; Duan, Y.; Wang, X.; Luo, F. Spatial–spectral enhancement and fusion network for hyperspectral image classification with few labeled samples. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5502414. [Google Scholar] [CrossRef]
- Xu, Y.; Wang, D.; Zhang, L.; Zhang, L. Dual selective fusion transformer network for hyperspectral image classification. Neural Netw. 2025, 187, 107311. [Google Scholar] [CrossRef] [PubMed]
- He, Y.; Tu, B.; Jiang, P.; Liu, B.; Li, J.; Plaza, A. IGroupSS-Mamba: Interval group spatial–spectral Mamba for hyperspectral image classification. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5538817. [Google Scholar] [CrossRef]











| No. | Color | HongHu | XiongAn | HanChuan | |||
|---|---|---|---|---|---|---|---|
| Class | Numbers | Class | Numbers | Class | Numbers | ||
| 1 | ![]() | Red roof | 14,041 | Swamp rice | 26,138 | Strawberry | 44,735 |
| 2 | ![]() | Road | 3512 | Rice stubble | 187,425 | Cowpea | 22,753 |
| 3 | ![]() | Bare soil | 21,821 | Water | 124,862 | Soybean | 10,287 |
| 4 | ![]() | Cotton | 163,285 | Grassland | 91,518 | Sorghum | 5353 |
| 5 | ![]() | Cotton firewood | 6218 | Willow | 197,218 | Water spinach | 1200 |
| 6 | ![]() | Rape | 44,557 | Elm | 19,663 | Watermelon | 4533 |
| 7 | ![]() | Chinese cabbage | 24,103 | Acer negundo | 296,538 | Greens | 5903 |
| 8 | ![]() | Pakchoi | 4054 | Ash tree | 276,755 | Trees | 17,978 |
| 9 | ![]() | Cabbage | 10,819 | Goldenrain tree | 44,232 | Grass | 9469 |
| 10 | ![]() | Tuber mustard | 12,394 | Chinese scholar tree | 372,708 | Red roof | 10,516 |
| 11 | ![]() | Brassica parachinensis | 11,015 | Peach tree | 67,210 | Gray roof | 16,911 |
| 12 | ![]() | Brassica chinensis | 8954 | Vegetable garden | 29,763 | Plastic | 3679 |
| 13 | ![]() | Small Brassica chinensis | 22,507 | Corn | 85,547 | Bare soil | 9116 |
| 14 | ![]() | Lactuca sativa | 7356 | Poplar | 68,885 | Road | 18,560 |
| 15 | ![]() | Celtuce | 1002 | Pear tree | 986,139 | Bright object | 1136 |
| 16 | ![]() | Film covered lettuce | 7262 | Soybeans | 7456 | Water | 75,401 |
| 17 | ![]() | Romaine lettuce | 3010 | Lotus leaf | 27,178 | / | / |
| 18 | ![]() | Carrot | 3217 | Locust | 6506 | / | / |
| 19 | ![]() | White radish | 8712 | Buildings | 26,140 | / | / |
| 20 | ![]() | Garlic sprout | 3486 | / | / | / | / |
| 21 | ![]() | Broad bean | 1328 | / | / | / | / |
| 22 | ![]() | Tree | 4040 | / | / | / | / |
| Total | 386,693 | 2,941,881 | 257,530 | ||||
| Class | CNN-Based Methods | Transformer-Based Methods | Mamba-Based Methods | Hybrid Architecture Methods | Ours | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| CSCANet | SSEFN | GSC-VIT | DSFormer | MambaHSI+ | MambaHSI | IGroupSS-Mamba | SCluster-Former | MCTGCL | HSI-MFormer | ||
| 1 | |||||||||||
| 2 | |||||||||||
| 3 | |||||||||||
| 4 | |||||||||||
| 5 | |||||||||||
| 6 | |||||||||||
| 7 | |||||||||||
| 8 | |||||||||||
| 9 | |||||||||||
| 10 | |||||||||||
| 11 | |||||||||||
| 12 | |||||||||||
| 13 | |||||||||||
| 14 | |||||||||||
| 15 | |||||||||||
| 16 | |||||||||||
| 17 | |||||||||||
| 18 | |||||||||||
| 19 | |||||||||||
| 20 | |||||||||||
| 21 | |||||||||||
| 22 | |||||||||||
| OA | |||||||||||
| AA | |||||||||||
| Kappa | |||||||||||
| Class | CNN-Based Methods | Transformer-Based Methods | Mamba-Based Methods | Hybrid Architecture Methods | Ours | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| CSCANet | SSEFN | GSC-VIT | DSFormer | MambaHSI+ | MambaHSI | IGroupSS-Mamba | SCluster-Former | MCTGCL | HSI-MFormer | ||
| 1 | |||||||||||
| 2 | |||||||||||
| 3 | |||||||||||
| 4 | |||||||||||
| 5 | |||||||||||
| 6 | |||||||||||
| 7 | |||||||||||
| 8 | |||||||||||
| 9 | |||||||||||
| 10 | |||||||||||
| 11 | |||||||||||
| 12 | |||||||||||
| 13 | |||||||||||
| 14 | |||||||||||
| 15 | |||||||||||
| 16 | |||||||||||
| OA | |||||||||||
| AA | |||||||||||
| Kappa | |||||||||||
| Class | CNN-Based Methods | Transformer-Based Methods | Mamba-Based Methods | Hybrid Architecture Methods | Ours | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| CSCANet | SSEFN | GSC-VIT | DSFormer | MambaHSI+ | MambaHSI | IGroupSS-Mamba | SCluster-Former | MCTGCL | HSI-MFormer | ||
| 1 | |||||||||||
| 2 | |||||||||||
| 3 | |||||||||||
| 4 | |||||||||||
| 5 | |||||||||||
| 6 | |||||||||||
| 7 | |||||||||||
| 8 | |||||||||||
| 9 | |||||||||||
| 10 | |||||||||||
| 11 | |||||||||||
| 12 | |||||||||||
| 13 | |||||||||||
| 14 | |||||||||||
| 15 | |||||||||||
| 16 | |||||||||||
| 17 | |||||||||||
| 18 | |||||||||||
| 19 | |||||||||||
| OA | |||||||||||
| AA | |||||||||||
| Kappa | |||||||||||
| Methods | Params (M) | FLOPs (M) | Train Time (s) | Test Time (s) | |
|---|---|---|---|---|---|
| CNN-Based Methods | CSCANet | 0.45 | 103.52 | 10.08 | 64.11 |
| SSEFN | 10.62 | 359.30 | 310.88 | 977.48 | |
| Transformer-Based Methods | GSC-VIT | 0.25 | 10.70 | 29.06 | 473.61 |
| DSFormer | 0.69 | 142.50 | 35.62 | 800.85 | |
| Mamba-Based Methods | MambaHSI+ | 0.46 | 80.31 | 409.38 | 467.24 |
| MambaHSI | 0.44 | 89.08 | 242.03 | 420.13 | |
| IGroupSS-Mamba | 0.37 | 145.48 | 576.12 | 1418.30 | |
| Hybrid Architecture Methods | SCluster-Former | 2.03 | 604.57 | 80.34 | 495.25 |
| MCTGCL | 0.29 | 88.85 | 81.17 | 48.03 | |
| HSI-MFormer | 0.17 | 74.35 | 23.14 | 65.32 | |
| Ours | Ours | 0.21 | 75.55 | 79.67 | 420.49 |
| Multi-Scale | DMB | SupCon | HanChuan | HongHu | XiongAn | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| OA | AA | Kappa | OA | AA | Kappa | OA | AA | Kappa | |||
| × | × | × | |||||||||
| × | ✓ | ✓ | |||||||||
| × | × | ✓ | |||||||||
| × | ✓ | × | |||||||||
| ✓ | × | ✓ | |||||||||
| ✓ | × | × | |||||||||
| ✓ | ✓ | × | |||||||||
| ✓ | ✓ | ✓ | |||||||||
| Spatial Prior | Spectral Prior | HanChuan | HongHu | XiongAn | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| OA | AA | Kappa | OA | AA | Kappa | OA | AA | Kappa | ||
| ✓ | × | 82.45 ± 3.28 | 81.33 ± 1.81 | 79.71 ± 3.64 | 88.20 ± 1.21 | 85.41 ± 1.40 | 85.21 ± 1.49 | 67.08 ± 2.21 | 76.03 ± 1.08 | 63.62 ± 2.30 |
| × | ✓ | 82.56 ± 2.27 | 81.27 ± 1.78 | 80.11 ± 2.54 | 87.95 ± 1.62 | 85.23 ± 1.49 | 84.93 ± 1.93 | 67.16 ± 1.82 | 76.05 ± 0.84 | 63.57 ± 1.84 |
| ✓ | ✓ | |||||||||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Feng, R.; Wu, C.; Zhou, M.; Wang, D.; Jia, D.; Bai, Z. DPMSCH-Net: A Dual Prior-Guidance Multi-Scale SSM-CNN Hybrid Network with Manifold Constraints for Hyperspectral Image Classification Under Small-Sample Conditions. Remote Sens. 2026, 18, 2452. https://doi.org/10.3390/rs18152452
Feng R, Wu C, Zhou M, Wang D, Jia D, Bai Z. DPMSCH-Net: A Dual Prior-Guidance Multi-Scale SSM-CNN Hybrid Network with Manifold Constraints for Hyperspectral Image Classification Under Small-Sample Conditions. Remote Sensing. 2026; 18(15):2452. https://doi.org/10.3390/rs18152452
Chicago/Turabian StyleFeng, Rui, Chanyue Wu, Meili Zhou, Dong Wang, Dongdong Jia, and Zongwen Bai. 2026. "DPMSCH-Net: A Dual Prior-Guidance Multi-Scale SSM-CNN Hybrid Network with Manifold Constraints for Hyperspectral Image Classification Under Small-Sample Conditions" Remote Sensing 18, no. 15: 2452. https://doi.org/10.3390/rs18152452
APA StyleFeng, R., Wu, C., Zhou, M., Wang, D., Jia, D., & Bai, Z. (2026). DPMSCH-Net: A Dual Prior-Guidance Multi-Scale SSM-CNN Hybrid Network with Manifold Constraints for Hyperspectral Image Classification Under Small-Sample Conditions. Remote Sensing, 18(15), 2452. https://doi.org/10.3390/rs18152452























