4.3. Performance Comparison
To validate the effectiveness of the proposed method, comprehensive experiments were conducted on the five mentioned benchmark datasets. Several state-of-the-art hyperspectral image classification methods were selected as comparison models, including Lite-HCNet [
31], LSSAN [
32], MSDAN [
33], SimPoolFormer [
34], SpectralFormer [
6], CacfNet [
35], GSCViT [
36], and SSFTT [
7].
(1) Results on QHUP Dataset:
As shown in
Table 1, FDM-Net achieves the best overall performance on the QHUP dataset, with an OA of 97.58%, AA of 92.35% and Kappa coefficient of 96.39%, all ranking first among the compared methods. Specifically, compared with the second-best method Lite-HCNet in terms of OA, FDM-Net yields an improvement of 0.54%. For AA and Kappa metrics, it outperforms the runner-up Lite-HCNet by 0.52% and 0.80% respectively, verifying its strong classification capability for complex land-cover scenes. The gains are particularly evident on challenging categories: for spectrally heterogeneous classes with large intra-class variance (e.g., Class 1 and Class 4), FDM-Net attains 91.50% and 96.93% respectively, benefiting from its frequency-adaptive feature decomposition that enhances robustness to intra-class spectral variations. On the most challenging Class 9, which suffers from limited training samples and severe spectral confusion, FDM-Net achieves a competitive 75.12% with higher stability.
Figure 3 presents the classification maps of all methods on the QHUP dataset for visual verification. Most competing methods produce scattered salt-and-pepper misclassification noise in homogeneous ground object regions, with blurred and jagged boundaries between different categories. Especially for methods like LSSAN and CacfNet, misclassification noise is more prominent in fine-grained local regions, and the structural morphology of small ground objects deviates obviously from the ground truth. In contrast, FDM-Net generates classification maps with significantly higher spatial consistency: it effectively suppresses random noise, retains the complete structure of ground objects, and renders smoother and more accurate category boundaries. Its overall visual effect is highly consistent with the ground truth map, which matches the quantitative evaluation conclusions.
(2) Results on QHUQ Dataset: As shown in
Table 2, FDM-Net achieves the best OA (95.07%) and Kappa (93.47%) on the QHUQ dataset, with a competitive AA of 88.57%. Compared with CacfNet, it yields a 0.73% OA improvement and a 0.97% Kappa gain. Even for challenging categories with limited training samples and severe spectral confusion (e.g., Class 3), FDM-Net maintains competitive accuracy, verifying its robust classification capability and balanced recognition performance across all land-cover classes. These results demonstrate the effectiveness of the proposed Frequency Decomposition Fusion Enhancement (FDE) strategy in extracting discriminative spectral–spatial representations through explicit Frequency Decomposition and complementary integration of low-frequency and high-frequency components. At the class-wise level, FDM-Net achieves optimal accuracy on three of the six categories (Class 1, Class 2, and Class 6), attaining 96.16%, 97.16%, and 94.44% respectively, which demonstrates the effectiveness of the proposed FDE strategy in extracting discriminative representations for spectrally homogeneous classes.
Figure 4 presents the classification maps of all methods. Most methods produce obvious scattered misclassifications in homogeneous regions, with blurred object boundaries and degraded spatial consistency; LSSAN and SpectralFormer suffer from more misclassified pixels in fine-grained local areas. In contrast, FDM-Net generates classification maps closest to the ground truth: it effectively suppresses isolated noise, preserves the structural integrity of ground objects, and outputs clearer category boundaries. These visual findings align with the quantitative results, which benefit from the collaborative modeling of low-frequency semantic information and high-frequency structural details, achieving both global consistency and local discriminability. The superior preservation of fine-grained spatial structures and sharp boundaries further validates the effectiveness of the MLFA module in capturing multi-level spatial details within the high-frequency branch.
(3) Results on QHUT Dataset: The QHUT dataset poses greater classification challenges due to its abundant land-cover categories and serious spectral overlap between similar ground objects. As recorded in
Table 3, FDM-Net ranks first on OA (97.58%), AA (93.86%) and Kappa (97.25%). Relative to CacfNet, our method gains 0.76% OA and 0.87% Kappa improvements. Many existing algorithms suffer from unstable per-class performance here—for example, SpectralFormer delivers an extremely low AA of 76.72% with a large deviation. In contrast, our frequency-decoupled framework achieves balanced, stable classification across all categories, demonstrating strong robustness against complex spectral interference. Notably, competing methods collapse on the most challenging categories where training samples are extremely limited. For instance, SpectralFormer drops to 0.08% and 10.35% on Class 12 and Class 14, as its Transformer-based architecture suffers from insufficient data for learning discriminative token representations. In contrast, FDM-Net still retains 72.44% and 79.70% on these same categories. Equally telling is the stability gap: competing methods show deviations up to 36.90%, while FDM-Net delivers consistent performance across all categories.
As displayed in
Figure 5, fragmented land parcels and numerous linear features such as roads lead to severe structural distortion and widespread misclassification noise for most compared models, and several methods even generate large-scale wrong predictions for spectrally similar land covers. By separately modeling low-frequency global context and high-frequency details, FDM-Net effectively suppresses intra-region noise, maintains the morphological continuity of linear terrain features, and reconstructs precise, smooth classification boundaries, thus delivering the most visually faithful prediction against the ground truth. The notable fidelity in reconstructing fine linear structures and intricate land parcel boundaries further underscores the effectiveness of the MLFA module in preserving detailed spatial information through multi-level feature aggregation within the high-frequency branch.
(4) Results on Houston Dataset:
As a typical urban scene with finely distributed diverse ground objects, the Houston dataset demands strong fine-grained feature discrimination capability. As shown in
Table 4, FDM-Net ranks first in OA and Kappa (94.79% and 94.39%), and second in AA (94.24%, slightly behind Lite-HCNet at 94.28%). Compared with the runner-up Lite-HCNet, it yields an OA improvement of 0.28% and a Kappa gain of 0.33%. Many compared methods suffer from severe performance degradation in this complex scene—CacfNet, for instance, only achieves an OA of 79.20%. FDM-Net additionally avoids the severe collapses seen in competing methods, while CacfNet and SimPoolFormer fall below 40% on Class 13. For fine-grained urban classes, FDM-Net surpasses competitors by up to 16.91 %, with consistently lower per-class variance, confirming the robustness of the frequency-decoupled design.
The classification maps for visual verification are shown in
Figure 6. It is evident that most compared methods suffer from degraded spatial consistency in homogeneous regions, as well as structural fracture and misclassification on narrow linear features such as roads. Some methods even exhibit large-area classification errors on spectrally similar urban objects, which is consistent with their poor quantitative performance. In contrast, FDM-Net produces classification maps that are significantly closer to the ground truth, with markedly clearer boundaries and higher spatial consistency. It effectively preserves the integrity of small urban targets and maintains the morphological continuity of linear features. These visual results confirm the effectiveness of FDM-Net. This advantage stems from the collaborative modeling of high-frequency edge details through the MLFA module and low-frequency global semantics through the FDE strategy, which well accommodates the fine-grained classification demands of complex urban scenarios.
(5) Results on Indian Pines Dataset: The Indian Pines dataset poses severe few-shot classification challenges due to extreme sample imbalance and heavy spectral overlap across various crop types. As summarized in
Table 5, our FDM-Net achieves SOTA results with 98.78% OA, 94.55% AA and 98.60% Kappa, surpassing the second-best GSCViT by 1.01% in OA. Compared algorithms suffer from dramatic accuracy degradation on minority crop classes with scarce training samples, whereas our frequency-domain feature mining framework achieves steady high recognition accuracy across all categories, demonstrating strong robustness against imbalanced data distribution. On minority classes, the advantage is pronounced: FDM-Net reaches 96.58% on Class 1 where MSDAN drops to 49.78%, and attains 80.00% on Class 9 where competing methods exhibit deviations exceeding 43%, confirming the robustness of the frequency-decoupled design under extreme sample scarcity.
The classification maps of various methods are shown in
Figure 7, their classification maps exhibit noticeable misclassified regions within large farmland areas, along with blurred field edges and distorted outlines of tiny cropland parcels. These errors arise because compared methods focus primarily on local pixel relationships and struggle to capture both global context and fine boundary details simultaneously. In contrast, FDM-Net produces classification maps that are significantly closer to the ground truth. Benefiting from the decoupled modeling of low-frequency global contextual information and high-frequency boundary features, FDM-Net retains the complete geometric shape of scattered small farm plots and maintains high spatial consistency across homogeneous regions. These visual results confirm that FDM-Net effectively accommodates the fine-grained classification demands of complex agricultural landscapes.
4.4. Visual Radar Chart Analysis
As shown in
Figure 8a, FDM-Net achieves the optimal OA across all five datasets, obtaining 97.58% (QHUP), 95.07% (QHUQ), 97.58% (QHUT), 94.79% (Houston), and 98.78% (IP). It outperforms the best method on each dataset by 0.54%, 0.57%, 0.76%, 0.28%, and 1.01%, respectively, demonstrating outstanding overall classification generality under diverse scene distributions. The AA radar chart in
Figure 8b further validates the balanced class-level performance of FDM-Net. Our method achieves leading AA values on QHUP (92.35%) and QHUT (93.86%), along with competitive AA on QHUQ (88.57%), Houston (94.24%), and IP (94.55%). It surpasses Lite-HCNet by up to 2.57%, showing superior capability in handling hard and minority classes. Consistent performance gains are also observed in the Kappa metric in
Figure 8c. FDM-Net achieves the highest Kappa scores on all datasets (96.39%, 93.47%, 97.25%, 94.39%, 98.60%), with consistent improvements over Lite-HCNet. Its fuller and more balanced radar coverage demonstrates stronger cross-dataset generalization. Collectively, the consistent OA, AA and Kappa superiority verifies that the collaborative design of FDE for frequency-aware decomposition and MLFA for multi-level spatial aggregation effectively enhances spectral–spatial feature discrimination and overall classification robustness.
4.5. Sensitivity Analysis
A sensitivity analysis evaluated the impact of multi-scale fusion weights (
,
,
for fine to coarse scales, summing to 1) on classification accuracy, as reported in
Table 6. Results show that fine-scale information dominates performance: the optimal configuration R3 (
,
,
) achieved the highest OA of 94.12%, confirming that a descending weight distribution (
) yields the most discriminative representation. Overemphasizing fine details R1 (
) reduced OA to 91.42% due to noise amplification, while moderate coarse-scale emphasis (R7,
,
,
) gave the second-best OA of 93.68%, but further increasing
to 0.60 in R8 lowered it to 92.91%. The uniform allocation R5 (
,
,
) performed poorly (84.23%), about 9.9% below the optimum, indicating asymmetric contributions and feature dilution. Excluding R5, the other seven configurations showed OA within a narrow 91.42–94.12% range (2.70% spread), demonstrating robustness to weight variations across a broad operational range.
4.6. Ablation Study
4.6.1. Effectiveness of FDE and MLFA
The proposed FDM-Net contains two core modules: the Frequency Decomposition Enhancement (FDE) module and the Multi-Level Feature Aggregation (MLFA) module. To quantitatively verify the individual effectiveness of the two designs, we conduct ablation experiments by removing each module separately while keeping other network settings unchanged. Experimental results across five datasets consistently validate that both FDE and MLFA positively boost classification performance and model stability. Specifically, the FDE module addresses the frequency-specific bias problem in hybrid frameworks by explicitly decoupling input features into low-frequency global semantic information and high-frequency local detail features, thereby eliminating feature interference and enabling discriminative spectral–spatial representation learning. Meanwhile, the MLFA module tackles the insufficient exploitation of multi-level feature interactions by adaptively aggregating multi-level spatial details within the high-frequency branch, effectively preserving fine-grained structures that are critical for distinguishing spectrally similar and sample-scarce categories. Moreover, the two modules are mutually reinforcing: FDE supplies MLFA with spectrally purified high-frequency features via frequency decoupling, while MLFA preserves fine-grained spatial structures that complement the low-frequency global semantics, together yielding consistent gains over each module in isolation.
Taking the QHUP dataset as an example, as shown in
Table 7, removing MLFA reduces OA, AA, and Kappa to 96.52%, 91.40%, and 95.41% (drops of 1.06%, 0.95%, and 0.98%), while removing FDE yields 96.49%, 91.29%, and 95.20% (drops of 1.09%, 1.06%, and 1.19%). Similar trends are observed on QHUQ and QHUT, where both variants consistently underperform the full model in OA and Kappa, confirming the effectiveness of MLFA and FDE in improving discriminative feature learning. On Houston and Indian Pines, the full model also achieves the best OA, with clear margins over both ablated variants, highlighting their importance for complex and imbalanced scenes.
4.6.2. Ablation Study on the Gate Path in MLFA
The MLFA module adopts a dual-path design comprising a Gate Path and a Value Path. To verify the effectiveness of this gating mechanism, we construct a variant that removes the Gate Path entirely while retaining the Value Path and the FDE module.
Table 8 reports the comparison across all five datasets.
Removing the Gate Path leads to consistent performance degradation across all five datasets. On IP, OA drops sharply from 98.78% to 93.64% (−5.14%); on Houston and QHUT, the declines are 1.82% and 1.37%; on QHUP, OA falls by 0.90% with AA decreasing by 10.19%. Even on QHUQ (6 classes), the full model achieves slightly higher OA (95.07% vs. 95.06%). These results conclusively demonstrate that the Gate Path, despite its simple form, proves essential for adaptively modulating multi-scale spatial features.
Notably, its contribution is more pronounced on datasets with higher intra-class variability and more categories (e.g., IP with 16 classes, Houston with 15 classes), consistent with the expectation that gating-based feature selection becomes increasingly critical as feature space complexity grows.
4.6.3. Superiority of Mamba over Alternative Architectures
The low-frequency branch of FDM-Net employs a Mamba-based state space model for long-range dependency modeling. To justify this design choice, we replace the Mamba module with a Transformer Encoder of comparable parameter count while keeping all other components (FDE, MLFA with Gate Path) unchanged.
Table 9 reports the comparison.
Across all five datasets, the Mamba-based branch consistently outperforms the Transformer-based alternative. On QHUT and IP, Mamba achieves OA improvements of 1.67% and 2.54%, respectively; on QHUP, the AA gap reaches 13.86%, demonstrating that Mamba’s state-space formulation is more effective than self-attention at capturing long-range spectral–spatial dependencies in hyperspectral data. This advantage stems from Mamba’s linear complexity and selective scan mechanism, which efficiently models global context without the quadratic overhead of Transformer self-attention, making it particularly suitable for the low-frequency branch where global semantic reasoning is required.
4.6.4. t-SNE Visualization
For qualitative analysis, t-SNE visualization is conducted on the Houston dataset (
Figure 9) to compare the full FDM-Net, the FDE-only branch, and the MLFA-only branch. The FDE-only variant, which relies primarily on frequency-decomposed global context without multi-level spatial aggregation, exhibits noticeable inter-class overlap, particularly among spectrally similar categories. The MLFA-only variant improves local compactness through multi-level spatial feature aggregation but still suffers from scattered intra-class distributions due to the absence of explicit Frequency Decomposition. In contrast, the full FDM-Net produces the most compact intra-class clusters and the clearest inter-class separation, demonstrating that the FDE strategy and the MLFA module are complementary and jointly enhance feature discriminability.
Overall, both quantitative ablation results and qualitative t-SNE visualizations demonstrate that the Gate Path, Mamba branch, FDE, and MLFA modules are indispensable and mutually complementary. The cooperative integration of explicit Frequency Decomposition, gated multi-level spatial aggregation, and Mamba-based global context modeling enables FDM-Net to fully exploit spectral–spatial cues, thereby achieving superior and stable classification performance in complex hyperspectral scenarios.
4.7. Analysis of Model Complexity and Inference Efficiency
Experiments on the Houston dataset use an 11 × 11 input patch size, batch size 256, and PCA-reduced spectral dimension from 144 to 16. All models are evaluated under identical settings in terms of parameters, FLOPs, and inference time (
Table 10).
MSDAN exhibits the highest complexity (12.75 MB, 442.135 M FLOPs, 43.611 ms) due to its multi-scale convolutional branches. SimPoolFormer also incurs relatively high cost (5.45 MB, 55.967 M FLOPs) owing to its Transformer-based architecture. In contrast, Lite-HCNet, CacfNet, SpectralFormer, and SSFTT are extremely lightweight (≤0.44 MB, ≤2.04 M FLOPs), with SSFTT achieving the fastest inference (2.386 ms) thanks to its efficient tokenization and compact Transformer design. GSCViT maintains moderate complexity (0.50 MB, 4.625 M FLOPs) with competitive inference speed (4.098 ms).
Our FDM-Net achieves the best overall classification performance while maintaining low model complexity (0.19 MB, 2.302 M FLOPs), using only 1.49% of MSDAN and 3.49% of SimPoolFormer parameters. Although its inference time (15.322 ms) is higher than some extremely lightweight models due to multi-branch frequency-domain fusion, it remains 64.9% faster than MSDAN. Overall, FDM-Net strikes a favorable balance between state-of-the-art accuracy and acceptable computational cost by combining frequency-domain enhancement with multi-level spatial aggregation, making it suitable for resource-constrained hyperspectral applications.