remotesensing-logo

Journal Browser

Journal Browser

Advancing Remote Sensing Through Large Multimodal Foundation Models: Toward Intelligent Earth Observation

A special issue of Remote Sensing (ISSN 2072-4292). This special issue belongs to the section "Remote Sensing Image Processing".

Deadline for manuscript submissions: 15 October 2026 | Viewed by 1994

Editors

Center of AI for Science, Shanghai Artificial Intelligence Laboratory, Shanghai, China
Interests: remote sensing image understanding; change detection; foundation models
Department of Geography, The University of Hong Kong, Hong Kong, China
Interests: remote sensing; deep learning; crop mapping; agricultural remote sensing; spatio-temporal modeling
Special Issues, Collections and Topics in MDPI journals

E-Mail Website
Guest Editor
Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences, Beijing, China
Interests: building height; weakly-supervised learning; multi-view imagery; high-resolution; change detection
Special Issues, Collections and Topics in MDPI journals
College of Computing & Data Science, Nanyang Technological University (NTU), Singapore, 639798, Singapore
Interests: remote sensing; earth observation; computer vision; foundation models; multimodal learning
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues,

Global warming, rapid urbanization and intensifying anthropogenic pressures have ushered in the Anthropocene, marked by interconnected environmental crises—biodiversity loss, extreme weather, land degradation, water scarcity and pollution—that threaten progress toward the UN Sustainable Development Goals (SDGs). Addressing these challenges requires not only comprehensive Earth observation data but also intelligent systems capable of transforming it into actionable, interpretable knowledge across spatiotemporal scales. Remote sensing has long provided synoptic, multi-spectral monitoring of Earth’s systems, yet the current explosion of multimodal data—from optical, SAR, LiDAR, hyperspectral, thermal and in situ sensors—exposes the limitations of traditional, task-specific workflows that lack semantic depth and cross-modal reasoning. Emerging foundation models, including generative large language models (LLMs), multimodal foundation models (MFMs), and AI agents, offer a paradigm shift. Pre-trained on vast, diverse datasets, they enable zero- or few-shot generalization, natural-language interaction, cross-sensor fusion, anomaly detection and even autonomous analysis pipelines. When tailored to Earth observation, these models unlock semantic understanding of landscapes, simulate environmental scenarios and support human-aligned decision-making—critical capabilities for monitoring and managing the coupled human–natural systems of the Anthropocene.

This Special Issue aims to foster interdisciplinary research that bridges cutting-edge AI—particularly generative models, multimodal reasoning, visual-language models and agentic frameworks—with remote sensing. It aligns closely with the journal’s scope by promoting innovative methodologies for RS data processing, interpretation and decision support. We invite original contributions that develop, adapt or benchmark foundation models for RS tasks, design intelligent agents for autonomous analysis or create standardized multimodal datasets and evaluation protocols.

Topics of interest include but are not limited to:

  • Multimodal foundation models for remote sensing (e.g., vision-language, vision-SAR, time-series fusion)
  • Generative AI for synthetic RS data generation, augmentation and simulation
  • LLM- and MFM-powered semantic interpretation and captioning of RS imagery
  • AI agents for autonomous RS data acquisition, processing and analysis workflows
  • Benchmark datasets and evaluation metrics for multimodal RS reasoning
  • Cross-modal alignment and transfer learning in Earth observation
  • Applications in multi-sphere Earth system monitoring (e.g., land cover change, water resources, urban dynamics)

Dr. Hao Chen
Dr. Wenyuan Li
Dr. Sen Lei
Dr. Yinxia Cao
Dr. Keyan Chen
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Remote Sensing is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2700 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • remote sensing
  • multimodal foundation models
  • generative AI
  • AI agents
  • multimodal fusion
  • earth system monitoring
  • intelligent
  • earth observation

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (3 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

33 pages, 7479 KB  
Article
Contribution Disparity and Key Factor Screening of Lightning Identification and Nowcasting with Multi-Source Data
by Xinjue Wang, Zirui Xu, Yujian Zhang, Yuanpeng Han, Lin Song, Qilin Zhang and Yi Liu
Remote Sens. 2026, 18(14), 2383; https://doi.org/10.3390/rs18142383 - 17 Jul 2026
Viewed by 286
Abstract
Radar and geostationary satellite observations provide essential information for lightning identification and nowcasting. However, multi-source channels commonly exhibit correlation and redundancy, and directly using all channels increases computational cost and may degrade model performance. This paper focuses on lightning identification and 0–1 h [...] Read more.
Radar and geostationary satellite observations provide essential information for lightning identification and nowcasting. However, multi-source channels commonly exhibit correlation and redundancy, and directly using all channels increases computational cost and may degrade model performance. This paper focuses on lightning identification and 0–1 h nowcasting, integrating radar composite reflectivity, Himawari-8/9 AHI multi-channel observations, and VLF-LLN lightning location data to construct a multi-source spatiotemporally aligned dataset. Using baseline deep-learning-based identification and nowcasting models, this paper proposes a workflow for analyzing multi-source channel contribution differences and selecting key factors by integrating permutation feature importance (PFI), SHAP attribution, channel collinearity grouping, and validation-set wrapper search. The results show that radar composite reflectivity has a stable dominant contribution in both tasks. The final selected channels are radar composite reflectivity, ΔBand 13, and Band 14 + Band 09 for the identification task, while they are radar composite reflectivity, Band 14, (Band 11 − Band 13) − (Band 13 − Band 15), and Δ(Band 08 − Band 13) for the nowcasting task. Independent test results show that under fixed thresholds, the three-channel identification model achieves a CSI of 0.433, outperforming the full 28-channel model with a CSI of 0.415, while the four-channel nowcasting model achieves a CSI of 0.284, outperforming the full-channel model with a CSI of 0.274. These results indicate that the proposed method can reduce input dimensionality while maintaining or improving model performance, and it can support multi-source predictor selection, lightweight deployment, and operational interpretation in lightning nowcasting. Full article
Show Figures

Figure 1

31 pages, 7054 KB  
Article
Fed-RSAdapter: Federated Fine-Tuning of Remote Sensing Images via Multi-Scale Adapter Modules
by Yuelei Wang, Liangkui Lin, Yirui Zhou, Shaolin Wang, Zhirui Wang and Xiyu Qi
Remote Sens. 2026, 18(14), 2352; https://doi.org/10.3390/rs18142352 - 14 Jul 2026
Viewed by 243
Abstract
In edge computing scenarios for remote sensing image interpretation, two fundamental challenges constrain the effectiveness of federated fine-tuning: the limited computational capacity of edge devices restricts the multi-scale feature learning capability of lightweight deployed models, while the highly heterogeneous and imbalanced data distributions [...] Read more.
In edge computing scenarios for remote sensing image interpretation, two fundamental challenges constrain the effectiveness of federated fine-tuning: the limited computational capacity of edge devices restricts the multi-scale feature learning capability of lightweight deployed models, while the highly heterogeneous and imbalanced data distributions (Non-IID settings) across clients render conventional parameter aggregation strategies ineffective. To address these challenges jointly, this paper proposes Fed-RSAdapter, a federated fine-tuning framework for remote sensing imagery based on multi-scale adapter modules. On the edge side, a multi-scale adapter architecture with cross-layer dense connections and feature map concatenation alignment is introduced. By inserting lightweight adapter modules in parallel within a frozen backbone, the proposed design captures multi-scale spatial characteristics inherent in high-resolution remote sensing imagery while confining trainable parameters to less than 5% of the full model, thereby satisfying strict on-device computational and communication constraints. On the server side, a parameter similarity-aware aggregation strategy is designed to handle client heterogeneity. Client adapter parameters are normalized and mapped into a similarity matrix via Gaussian kernel distance, enabling the server to perform personalized weighted aggregation that balances global generalization with client-specific adaptation. Extensive experiments on scene classification, object detection, and semantic segmentation benchmarks demonstrate that Fed-RSAdapter achieves an average improvement of 3.75% in overall accuracy for scene classification, 0.74% in mAP for object detection, and 0.25% in overall accuracy for semantic segmentation over federated baselines, while reducing the volume of transmitted parameters to approximately 4.93% of that required by conventional full-parameter federated learning. These results demonstrate the communication-efficient fine-tuning capability of the proposed framework under federated remote sensing settings. We further clarify that actual deployment on embedded edge hardware is not directly evaluated in this work and is therefore treated as an important direction for future validation rather than an experimentally proven conclusion. Full article
Show Figures

Figure 1

29 pages, 6909 KB  
Article
MDE-UNet: A Physically Guided Asymmetric Fusion Network for Multi-Source Meteorological Data Lightning Identification
by Yihua Chen, Yuanpeng Han, Yujian Zhang, Yi Liu, Lin Song, Jialei Wang, Xinjue Wang and Qilin Zhang
Remote Sens. 2026, 18(7), 1027; https://doi.org/10.3390/rs18071027 - 29 Mar 2026
Cited by 1 | Viewed by 531
Abstract
Utilizing multi-source meteorological data for lightning identification is crucial for monitoring severe convective weather. However, several key challenges persist in this field: dimensional imbalance and modal competition among multi-source heterogeneous data, model training bias caused by the extreme sparsity of lightning samples, and [...] Read more.
Utilizing multi-source meteorological data for lightning identification is crucial for monitoring severe convective weather. However, several key challenges persist in this field: dimensional imbalance and modal competition among multi-source heterogeneous data, model training bias caused by the extreme sparsity of lightning samples, and an imbalance between false alarms and missed detections resulting from complex background noise. To address these challenges, this paper proposes a lightning identification network guided by physical priors and constrained by supervision. First, to tackle the issue of modal competition in fusing satellite (high-dimensional) and radar (low-dimensional) data, a physical prior-guided asymmetric radar information enhancement mechanism is introduced. This mechanism uses radar physical features as contextual guidance to selectively enhance the latent weak radar signatures. Second, at the architectural level, a multi-source multi-scale feature fusion module and a weighted sliding window–multilayer perceptron (MLP) enhanced decoding unit are constructed. The former achieves the coupling of multi-scale physical features at a 2 km grid scale through cross-level semantic alignment, building a highly consistent feature field that effectively improves the model’s ability to detect lightning signals. The latter leverages adaptive receptive fields and the nonlinear modeling capability of MLPs to effectively smooth spatially discrete noise, ensuring spatial continuity in the reconstructed results. Finally, to address the model bias caused by severe class imbalance between positive and negative samples—resulting from the extreme sparsity of lightning events—an asymmetrically weighted BCE-DICE loss function is designed. Its “asymmetric” characteristic is implemented by assigning different penalty weights to false-positive and false-negative predictions. This loss function balances pixel-level accuracy and inter-class equilibrium while imposing high-weight penalties on false-positive predictions, achieving synergistic optimization of feature enhancement and directional suppression. Experimental results show that the proposed method effectively increases the hit rate while substantially reducing the false alarm rate, enabling efficient utilization of multi-source data and high-precision identification of lightning strike areas. Full article
Show Figures

Figure 1

Back to TopTop