LMFusion: Breaking the Computational Barrier for Multimodal Classification in Remote Sensing
Highlights
- We propose LMFusion, an efficient multimodal remote sensing classification framework that integrates linear-complexity cross-attention for bidirectional HSI-LiDAR feature interaction and Mamba-based state space modeling for spatial-spectral representation learning.
- Experimental results on multiple multimodal remote sensing datasets demonstrate that LMFusion achieves competitive classification performance while providing an effective feature fusion strategy for hyperspectral and LiDAR data.
- The proposed framework shows that effective multimodal feature interaction can be achieved with reduced computational burden, making multimodal remote sensing classification more suitable for resource-constrained scenarios.
- The introduced selective quantization-aware optimization further improves the compactness of the model, providing potential support for future deployment on low-bit or edge-computing hardware platforms.
Abstract
1. Introduction
- We propose LMFusion, an efficient multimodal classification framework that integrates linear-complexity cross-modal interaction with Mamba-based long-range spatial–spectral modeling, enabling effective HSI–LiDAR feature fusion under limited computational budgets.
- We develop a quantization-aware optimization scheme for LMFusion, supporting ultra-low-bit deployment while preserving feature learning and improving model compactness and inference efficiency.
- Extensive experiments on multiple multimodal benchmark datasets demonstrate that the proposed method consistently outperforms representative state-of-the-art approaches, while ablation studies further verify the effectiveness of Linear Cross Attention, SS2D-based state-space modeling, and quantization-aware training.
2. Related Works
2.1. Multi-Modal Learning in Remote Sensing
2.2. Efficient and Lightweight Optimization
3. Methodology
3.1. Problem Definition
3.2. Framework Overview
3.3. Linear Cross Attention
3.4. SS2D-Based Spatial State Space Modeling
3.5. Quantization-Aware Training
| Algorithm 1 Quantization-Aware Training Procedure |
|
4. Experiments and Analysis
4.1. Data Description
4.1.1. Houston2013 Dataset
4.1.2. Augsburg Dataset
4.1.3. MUUFL Gulfport Scene Dataset
4.2. Evaluation Metrics and Parameter Setting
4.2.1. Evaluation Metrics
4.2.2. Parameter Setting
4.3. Ablation Study
- A.
- Modeling Paradigm Analysis:
- B.
- Cross-Attention Design Analysis:
- C.
- Quantization Analysis:
5. Results
5.1. Comparisons with Previous Methods
5.2. Result Visualization
6. Discussion
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Chen, Y.; Lin, Z.; Zhao, X.; Wang, G.; Gu, Y. Deep Learning-Based Classification of Hyperspectral Data. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2014, 7, 2094–2107. [Google Scholar] [CrossRef]
- Ghamisi, P.; Yokoya, N.; Li, J.; Liao, W.; Liu, S.; Plaza, J.; Rasti, B.; Plaza, A. Advances in Hyperspectral Image and Signal Processing: A Comprehensive Overview of the State of the Art. IEEE Geosci. Remote Sens. Mag. 2017, 5, 37–78. [Google Scholar] [CrossRef]
- Matsuki, T.; Yokoya, N.; Iwasaki, A. Hyperspectral Tree Species Classification of Japanese Complex Mixed Forest with the Aid of Lidar Data. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2015, 8, 2177–2187. [Google Scholar] [CrossRef]
- Ghamisi, P.; Höfle, B.; Zhu, X.X. Hyperspectral and LiDAR Data Fusion Using Extinction Profiles and Deep Convolutional Neural Network. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2017, 10, 3011–3024. [Google Scholar] [CrossRef]
- Hang, R.; Li, Z.; Ghamisi, P.; Hong, D.; Xia, G.; Liu, Q. Classification of Hyperspectral and LiDAR Data Using Coupled CNNs. IEEE Trans. Geosci. Remote Sens. 2020, 58, 4939–4950. [Google Scholar] [CrossRef]
- Zhao, X.; Tao, R.; Li, W.; Li, H.C.; Du, Q.; Liao, W.; Philips, W. Joint Classification of Hyperspectral and LiDAR Data Using Hierarchical Random Walk and Deep CNN Architecture. IEEE Trans. Geosci. Remote Sens. 2020, 58, 7355–7370. [Google Scholar] [CrossRef]
- Du, X.; Zheng, X.; Lu, X.; Doudkin, A.A. Multisource Remote Sensing Data Classification with Graph Fusion Network. IEEE Trans. Geosci. Remote Sens. 2021, 59, 10062–10072. [Google Scholar] [CrossRef]
- Xue, Z.; Yu, X.; Tan, X.; Liu, B.; Yu, A.; Wei, X. Multiscale Deep Learning Network with Self-Calibrated Convolution for Hyperspectral and LiDAR Data Collaborative Classification. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5514116. [Google Scholar] [CrossRef]
- Luo, R.; Liao, W.; Zhang, H.; Zhang, L.; Scheunders, P.; Pi, Y.; Philips, W. Fusion of Hyperspectral and LiDAR Data for Classification of Cloud-Shadow Mixed Remote Sensed Scene. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2017, 10, 3768–3781. [Google Scholar] [CrossRef]
- Gbodjo, Y.J.E.; Montet, O.; Ienco, D.; Gaetano, R.; Dupuy, S. Multisensor Land Cover Classification with Sparsely Annotated Data Based on Convolutional Neural Networks and Self-Distillation. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2021, 14, 11485–11499. [Google Scholar] [CrossRef]
- Zhang, M.; Li, W.; Tao, R.; Li, H.; Du, Q. Information Fusion for Classification of Hyperspectral and LiDAR Data Using IP-CNN. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5506812. [Google Scholar] [CrossRef]
- Li, J.; Ma, Y.; Song, R.; Xi, B.; Hong, D.; Du, Q. A Triplet Semisupervised Deep Network for Fusion Classification of Hyperspectral and LiDAR Data. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5540513. [Google Scholar] [CrossRef]
- Hong, D.; Gao, L.; Yokoya, N.; Yao, J.; Chanussot, J.; Du, Q.; Zhang, B. More Diverse Means Better: Multimodal Deep Learning Meets Remote-Sensing Imagery Classification. IEEE Trans. Geosci. Remote Sens. 2021, 59, 4340–4354. [Google Scholar] [CrossRef]
- Xiu, D.; Pan, Z.; Wu, Y.; Hu, Y. MAGE: Multisource Attention Network with Discriminative Graph and Informative Entities for Classification of Hyperspectral and LiDAR Data. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5539714. [Google Scholar] [CrossRef]
- Šimundić, V.; Mihelčić, D.; Svirac, D.; Đurović, P.; Cupec, R. Safety System for Industrial Robots Based on Human Detection Using an RGB-D Camera. In Proceedings of the 2021 44th International Convention on Information, Communication and Electronic Technology (MIPRO), Opatija, Croatia, 27 September–1 October 2021; pp. 1178–1184. [Google Scholar] [CrossRef]
- Roy, S.K.; Krishna, G.; Dubey, S.R.; Chaudhuri, B.B. HybridSN: Exploring 3-D–2-D CNN Feature Hierarchy for Hyperspectral Image Classification. IEEE Geosci. Remote Sens. Lett. 2020, 17, 277–281. [Google Scholar] [CrossRef]
- Nugraheni, D.M.K.; de Vries, D. The effectiveness of SMS as verification of flood early warning messages from users’ perception. In Proceedings of the 2017 1st International Conference on Informatics and Computational Sciences (ICICoS), Semarang, Indonesia, 15–16 November 2017; pp. 77–82. [Google Scholar] [CrossRef]
- Zhu, X.X.; Tuia, D.; Mou, L.; Xia, G.S.; Zhang, L.; Xu, F.; Fraundorfer, F. Deep Learning in Remote Sensing: A Comprehensive Review and List of Resources. IEEE Geosci. Remote Sens. Mag. 2017, 5, 8–36. [Google Scholar] [CrossRef]
- Baltrušaitis, T.; Ahuja, C.; Morency, L.P. Multimodal Machine Learning: A Survey and Taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 423–443. [Google Scholar] [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual Event, 3–7 May 2021. [Google Scholar]
- Katharopoulos, A.; Vyas, A.; Pappas, N.; Fleuret, F. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. In Proceedings of the 37th International Conference on Machine Learning (ICML), Virtual Event, 13–18 July 2020; Daumé, H., III, Singh, A., Eds.; PMLR: Cambridge, MA, USA, 2020; Volume 119, pp. 5156–5165. [Google Scholar]
- Yang, W.; Ouyang, W.; Wang, X.; Ren, J.; Li, H.; Wang, X. 3D Human Pose Estimation in the Wild by Adversarial Learning. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 5255–5264. [Google Scholar] [CrossRef]
- Banner, R.; Nahshan, Y.; Hoffer, E.; Soudry, D. Post Training 4-bit Quantization of Convolutional Networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar]
- Krishnamoorthi, R. Quantizing Deep Convolutional Networks for Efficient Inference: A Whitepaper. arXiv 2018, arXiv:1806.08342. [Google Scholar] [CrossRef]
- Esser, S.K.; McKinstry, J.L.; Bablani, D.; Appuswamy, R.; Modha, D.S. Learned Step Size Quantization. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual Event, 26 April–1 May 2020. [Google Scholar]
- Ruan, Y.; Wu, S.; Yang, C.; Xie, K.; Zhi, H. Simulation and Experiments of the Eruption of Deep-sea Hydrothermal Plume. In Proceedings of the Global Oceans 2020: Singapore–U.S. Gulf Coast, Virtual Event, 5–30 October 2020; pp. 1–6. [Google Scholar] [CrossRef]
- Rehmat, M.; Ansari, A.; ur Rehman, M. Modeling and Analysis of 300 MW Photovoltaic System Using ETAP and Harmonic Filter Design. In Proceedings of the 2023 Third International Symposium on Instrumentation, Control, Artificial Intelligence, and Robotics (ICA-SYMP), Bangkok, Thailand, 18–20 January 2023; pp. 140–144. [Google Scholar] [CrossRef]
- Nagel, M.; Amjad, R.A.; van Baalen, M.; Louizos, C.; Blankevoort, T. A Data-Free Quantization Method for Deep Neural Networks. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual Event, 3–7 May 2021. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Qin, Z.; Sun, W.; Deng, H.; Jiang, X.; Sun, Y.; Zhao, X. Bridging the Divide: Reconsidering Softmax and Linear Attention. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), Abu Dhabi, United Arab Emirates, 7–11 December 2022; pp. 515–527. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar]
- Gu, A.; Goel, K.; Re, C. Efficiently Modeling Long Sequences with Structured State Spaces. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual Event, 25–29 April 2022. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. VMamba: Visual State Space Model. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 9–15 December 2024. [Google Scholar]
- Jacob, B.; Kligys, S.; Chen, B.; Zhu, M.; Tang, M.; Howard, A.; Adam, H.; Kalenichenko, D. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018. [Google Scholar]
- Zhou, S.; Wu, Y.; Ni, Z.; Zhou, X.; Wen, H.; Zou, Y. DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients. arXiv 2016, arXiv:1606.06160. [Google Scholar]
- Wu, X.; Hong, D.; Chanussot, J. Convolutional Neural Networks for Multimodal Remote Sensing Data Classification. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5517010. [Google Scholar] [CrossRef]
- Feng, M.; Gao, F.; Fang, J.; Dong, J. Hyperspectral and Lidar Data Classification Based on Linear Self-Attention. In Proceedings of the 2021 IEEE International Geoscience and Remote Sensing Symposium IGARSS, Virtual Event, 11–16 July 2021; pp. 2401–2404. [Google Scholar] [CrossRef]
- Zhao, G.; Ye, Q.; Sun, L.; Wu, Z.; Pan, C.; Jeon, B. Joint Classification of Hyperspectral and LiDAR Data Using a Hierarchical CNN and Transformer. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5500716. [Google Scholar] [CrossRef]
- Feng, Y.; Zhu, J.; Song, R.; Wang, X. S2EFT: Spectral-spatial-elevation fusion transformer for hyperspectral image and LiDAR classification. Knowl.-Based Syst. 2024, 283, 111190. [Google Scholar] [CrossRef]
- Yao, J.; Zhang, B.; Li, C.; Hong, D.; Chanussot, J. Extended Vision Transformer (ExViT) for Land Use and Land Cover Classification: A Multimodal Deep Learning Framework. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5514415. [Google Scholar] [CrossRef]
- Feng, Y.; Song, L.; Wang, L.; Wang, X. DSHFNet: Dynamic Scale Hierarchical Fusion Network Based on Multiattention for Hyperspectral Image and LiDAR Data Classification. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5522514. [Google Scholar] [CrossRef]
- Wang, A.; Dai, S.; Wu, H.; Lv, H.; Yan, S.; Wang, M. CTPMSN: Enhancing Multimodal Remote Sensing Classification with Composite Text Prompts. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 27960–27978. [Google Scholar] [CrossRef]




| Land Cover | Train | Test | Land Cover | Train | Test |
|---|---|---|---|---|---|
| Background | 662,013 | 652,648 | Grass-healthy | 198 | 1053 |
| Grass-stressed | 190 | 1064 | Grass-synthetic | 192 | 505 |
| Tree | 188 | 1056 | Soil | 186 | 1056 |
| Water | 182 | 143 | Residential | 196 | 1072 |
| Commercial | 191 | 1053 | Road | 193 | 1059 |
| Highway | 191 | 1036 | Railway | 181 | 1054 |
| Parking-lot1 | 192 | 1041 | Parking-lot2 | 184 | 285 |
| Tennis-court | 181 | 247 | Running-track | 187 | 473 |
| Land Cover | Train | Test | Land Cover | Train | Test |
|---|---|---|---|---|---|
| Background | 160,259 | 83,487 | Forest | 146 | 13,361 |
| Commercial Area | 264 | 30,065 | Residential Area | 21 | 3830 |
| Industrial Area | 248 | 26,609 | Low Plants | 52 | 523 |
| Allotment | 7 | 1638 | Water | 23 | 1507 |
| Land Cover | Train | Test | Land Cover | Train | Test |
|---|---|---|---|---|---|
| Background | 68,817 | 20,496 | Trees | 1162 | 22,084 |
| Grass-Pure | 214 | 4056 | Grass-Groundsurface | 344 | 6538 |
| Dirt-And-Sand | 91 | 1735 | Road-Materials | 334 | 6353 |
| Water | 23 | 443 | Buildings’-Shadow | 112 | 2121 |
| Buildings | 312 | 5928 | Sidewalk | 69 | 1316 |
| Yellow-Curb | 9 | 174 | ClothPanels | 13 | 256 |
| CNN | Trans | Mamba | Linear CA | Quant | Bit | OA | AA | Params | FLOPs | Size | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| A. Modeling Paradigm | |||||||||||
| ✓ | 95.82 | 89.55 | 94.56 | 4.08 M | 14.16 G | 15.55 MB | |||||
| ✓ | 96.45 | 91.91 | 95.37 | 4.05 M | 85.75 G | 15.46 MB | |||||
| ✓ | 97.14 | 95.53 | 96.40 | 7.58 M | 19.86 G | 28.92 MB | |||||
| B. Cross-Attention Design | |||||||||||
| ✓ | 97.14 | 95.53 | 96.40 | 7.58 M | 19.86 G | 28.92 MB | |||||
| ✓ | ✓ | FP32 | 96.61 | 94.97 | 95.82 | 7.59 M | 18.30 G | 28.96 MB | |||
| C. Quantization (based on Mamba + Linear CA) | |||||||||||
| ✓ | ✓ | ✓ | 1 | 97.19 | 95.65 | 96.44 | 7.59 M | 18.30 G | 0.91 MB | ||
| ✓ | ✓ | ✓ | 2 | 96.61 | 95.02 | 95.81 | 7.59 M | 18.30 G | 1.81 MB | ||
| ✓ | ✓ | ✓ | 4 | 96.69 | 95.09 | 95.91 | 7.59 M | 18.30 G | 3.62 MB | ||
| ✓ | ✓ | ✓ | 8 | 96.25 | 94.63 | 95.43 | 7.59 M | 18.30 G | 7.24 MB | ||
| ✓ | ✓ | ✓ | 16 | 96.73 | 95.18 | 95.95 | 7.59 M | 18.30 G | 14.48 MB | ||
| No. | Class | Coupled CNN | CCR-Net | LSAF | HCT | MHST | SSEFT | ExViT | DSHF | Ours |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Healthy grass | 82.81 | 88.41 | 83.10 | 97.34 | 92.18 | 83.38 | 82.91 | 87.56 | 98.34 |
| 2 | Stressed grass | 99.81 | 99.91 | 83.08 | 96.62 | 96.80 | 97.09 | 98.68 | 99.34 | 99.51 |
| 3 | Synthetic grass | 97.23 | 100.00 | 100.00 | 84.75 | 97.82 | 97.23 | 99.60 | 98.81 | 100.00 |
| 4 | Tree | 100.00 | 99.15 | 89.49 | 96.78 | 99.81 | 98.86 | 99.15 | 99.81 | 98.11 |
| 5 | Soil | 100.00 | 99.34 | 100.00 | 100.00 | 100.00 | 99.72 | 99.91 | 99.91 | 100.00 |
| 6 | Water | 100.00 | 95.80 | 100.00 | 96.50 | 96.50 | 91.61 | 99.30 | 93.71 | 100.00 |
| 7 | Residential | 92.35 | 94.96 | 92.72 | 82.09 | 95.52 | 91.32 | 96.08 | 83.96 | 98.69 |
| 8 | Commercial | 94.30 | 91.26 | 92.69 | 95.54 | 96.39 | 90.03 | 90.03 | 80.63 | 85.39 |
| 9 | Road | 92.73 | 92.63 | 97.07 | 90.84 | 87.72 | 80.36 | 86.12 | 64.68 | 91.06 |
| 10 | Highway | 85.14 | 79.63 | 68.44 | 58.88 | 82.24 | 55.31 | 72.97 | 96.72 | 76.7 |
| 11 | Railway | 98.86 | 94.21 | 89.94 | 97.53 | 98.01 | 84.82 | 88.99 | 78.75 | 100.00 |
| 12 | Park lot 1 | 93.28 | 88.38 | 96.25 | 90.11 | 85.49 | 77.71 | 90.39 | 86.36 | 91.5 |
| 13 | Park lot 2 | 89.47 | 77.19 | 89.12 | 97.19 | 92.98 | 61.75 | 90.18 | 88.77 | 99.18 |
| 14 | Tennis court | 100.00 | 96.76 | 100.00 | 100.00 | 94.33 | 99.76 | 99.60 | 98.79 | 100.00 |
| 15 | Running track | 98.94 | 99.79 | 100.00 | 100.00 | 97.04 | 92.18 | 95.14 | 100.00 | 100.00 |
| OA(%) | 94.37 | 93.15 | 90.51 | 91.15 | 93.81 | 86.33 | 91.40 | 89.01 | 95.84 | |
| AA(%) | 94.99 | 93.16 | 92.13 | 92.28 | 94.19 | 86.41 | 92.60 | 90.52 | 95.9 | |
| Kappa(%) | 93.88 | 92.56 | 89.69 | 90.40 | 93.28 | 85.18 | 90.66 | 88.06 | 95.54 |
| No. | Class | Coupled_CNN | CCR-Net | LSAF | HCT | MHST | SSEFT | ExViT | DSHF | Ours |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Forest | 92.96 | 92.06 | 97.21 | 94.23 | 98.92 | 91.83 | 93.51 | 97.06 | 99.85 |
| 2 | Commercial Area | 1.77 | 6.78 | 0.31 | 4.82 | 1.40 | 9.34 | 15.20 | 8.91 | 99.71 |
| 3 | Residential Area | 97.04 | 97.13 | 98.97 | 98.54 | 94.55 | 90.19 | 97.31 | 96.88 | 93.88 |
| 4 | Industrial Area | 76.66 | 62.09 | 30.31 | 43.79 | 64.33 | 52.22 | 64.26 | 14.83 | 99.65 |
| 5 | Low Plants | 96.10 | 84.20 | 95.97 | 95.33 | 86.04 | 88.29 | 87.57 | 98.01 | 92.25 |
| 6 | Allotment | 50.10 | 46.27 | 49.14 | 67.88 | 52.01 | 55.64 | 52.77 | 5.54 | 90.03 |
| 7 | Water | 30.66 | 26.94 | 11.94 | 52.09 | 34.37 | 8.63 | 27.01 | 47.51 | 94.78 |
| OA(%) | 91.39 | 86.47 | 90.13 | 90.90 | 87.47 | 84.42 | 88.28 | 89.81 | 99.05 | |
| AA(%) | 63.61 | 59.35 | 54.84 | 65.24 | 61.66 | 56.59 | 62.52 | 52.68 | 95.74 | |
| Kappa(%) | 87.56 | 80.63 | 85.45 | 86.83 | 82.17 | 77.74 | 83.28 | 84.93 | 98.64 |
| No. | Class | Coupled_CNN | CCR-Net | LSAF | HCT | MHST | SSEFT | ExViT | DSHF | Ours |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Trees | 97.67 | 97.44 | 97.00 | 97.03 | 98.86 | 97.37 | 98.58 | 95.06 | 98.51 |
| 2 | Grass-Pure | 88.09 | 77.93 | 83.36 | 90.29 | 79.83 | 77.76 | 87.70 | 81.01 | 91.55 |
| 3 | Grass-Groundsurface | 88.99 | 82.90 | 90.79 | 90.07 | 79.63 | 83.13 | 90.96 | 72.90 | 90.14 |
| 4 | Dirt-And-Sand | 95.22 | 88.13 | 94.18 | 94.18 | 93.95 | 82.54 | 90.61 | 94.87 | 95.8 |
| 5 | Road-Materials | 97.32 | 96.49 | 95.88 | 93.86 | 94.84 | 95.17 | 94.73 | 90.41 | 96.47 |
| 6 | Water | 99.10 | 94.36 | 94.58 | 95.71 | 92.78 | 92.10 | 93.68 | 1.81 | 98.36 |
| 7 | Buildings’Shadow | 85.38 | 83.36 | 87.36 | 87.09 | 87.13 | 78.97 | 90.05 | 97.97 | 89.51 |
| 8 | Buildings | 98.08 | 97.93 | 98.14 | 96.61 | 97.12 | 96.78 | 97.76 | 92.70 | 98.04 |
| 9 | Sidewalk | 51.60 | 52.89 | 75.91 | 46.35 | 68.77 | 63.22 | 68.54 | 56.08 | 87.77 |
| 10 | Yellow-Curb | 10.34 | 2.30 | 13.22 | 18.97 | 28.16 | 26.44 | 23.56 | 0 | 92.73 |
| 11 | ClothPanels | 81.64 | 85.16 | 84.77 | 75.39 | 93.75 | 91.02 | 80.86 | 0 | 97.75 |
| OA(%) | 93.65 | 91.50 | 93.70 | 92.95 | 92.43 | 91.17 | 94.37 | 88.65 | 94.95 | |
| AA(%) | 81.22 | 78.08 | 83.20 | 80.50 | 83.17 | 80.41 | 83.37 | 63.11 | 93.28 | |
| Kappa(%) | 91.59 | 88.74 | 91.68 | 90.69 | 89.92 | 88.32 | 92.54 | 85.15 | 93.28 |
| Model | Params (M) | Model Size (MB) |
|---|---|---|
| MAHiDFNet | 77.0 | 308.0 |
| FusAtNet | 36.9 | 147.6 |
| SepG-ResNet50 | 14.7 | 58.8 |
| Ours (FP32) | 7.60 | 29.00 |
| Ours (1-bit) | 7.60 | 0.91 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhou, S.; He, S.; Li, D.; Xie, W.; Li, Y. LMFusion: Breaking the Computational Barrier for Multimodal Classification in Remote Sensing. Remote Sens. 2026, 18, 1972. https://doi.org/10.3390/rs18121972
Zhou S, He S, Li D, Xie W, Li Y. LMFusion: Breaking the Computational Barrier for Multimodal Classification in Remote Sensing. Remote Sensing. 2026; 18(12):1972. https://doi.org/10.3390/rs18121972
Chicago/Turabian StyleZhou, Shenbo, Sibo He, Daixun Li, Weiying Xie, and Yunsong Li. 2026. "LMFusion: Breaking the Computational Barrier for Multimodal Classification in Remote Sensing" Remote Sensing 18, no. 12: 1972. https://doi.org/10.3390/rs18121972
APA StyleZhou, S., He, S., Li, D., Xie, W., & Li, Y. (2026). LMFusion: Breaking the Computational Barrier for Multimodal Classification in Remote Sensing. Remote Sensing, 18(12), 1972. https://doi.org/10.3390/rs18121972

