Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (44)

Search Parameters:
Keywords = remote sensing images captioning

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
19 pages, 23860 KB  
Article
GeoGATE: Geo-Sensor-Guided Adaptive Token and Evidence Reasoning for High-Resolution Remote Sensing Image Understanding
by Jingnan Zhang and Fengjun Zhang
Appl. Sci. 2026, 16(17), 8616; https://doi.org/10.3390/app16178616 - 29 Aug 2026
Viewed by 233
Abstract
High-resolution remote sensing understanding requires models to preserve small spatial evidence, account for acquisition-dependent appearance, and separate genuine geographic change from nuisance variation. We introduce GeoGATE, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware [...] Read more.
High-resolution remote sensing understanding requires models to preserve small spatial evidence, account for acquisition-dependent appearance, and separate genuine geographic change from nuisance variation. We introduce GeoGATE, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware temporal reasoning. LoRA adaptation and NF4 quantization support efficient training and deployment. On the VRSBench test split, GeoGATE reaches 53.4 BLEU-1, 36.8 BLEU-2, 18.2 BLEU-4, 56.4 Acc@0.5, 82.3 VQA, 25.1 METEOR, and 42.6 ROUGE-L, outperforming the controlled GeoGATE (Base) configuration across captioning, question answering, and grounding. Component ablations associate adaptive slicing most strongly with localization, retrieval with language and VQA, and language model adaptation with all reported tasks. NF4 reduces measured video memory from 24.5 GiB to 7.2 GiB with only minor metric changes. These experiments support the single-image language and grounding components. Dedicated cross-sensor and bi-temporal benchmarks are not reported; the corresponding modules are therefore presented as architectural extensions rather than validated performance claims. Full article
Show Figures

Figure 1

33 pages, 6352 KB  
Article
ChangePixel: Pixel-Level Evidence-Grounded Disaster Change Narration via Single-Backbone Transfer
by Qinyu Zhou, Ben Yang, Xinyan Wei, Ding Qin, Tingting Leng and Xiaojing Liu
Remote Sens. 2026, 18(15), 2480; https://doi.org/10.3390/rs18152480 - 29 Jul 2026
Viewed by 416
Abstract
Remote sensing change captioning aims to describe disaster-related changes from bi-temporal imagery, yet existing methods typically produce image-level captions without explicit regional evidence, limiting interpretability and weakening the link between generated language and actual changed areas. We present ChangePixel, a single-backbone framework that [...] Read more.
Remote sensing change captioning aims to describe disaster-related changes from bi-temporal imagery, yet existing methods typically produce image-level captions without explicit regional evidence, limiting interpretability and weakening the link between generated language and actual changed areas. We present ChangePixel, a single-backbone framework that upgrades remote sensing change captioning into pixel-level, evidence-grounded change narration without introducing new manual grounding labels. ChangePixel incorporates three lightweight modules: a Bi-Temporal Change-Aware Transfer Adapter (BCTA) that converts shared pre- and post-event visual features into change-aware grounding representations, a Change Region Grounding Planner (CRGP) that localizes a compact set of informative changed regions before narration begins, and a Weak Evidence Alignment Bridge (WAB) that converts released change captions into phrase-to-region weak supervision. Through this design, the model jointly produces a global change caption and region-level evidence in the form of pixel masks paired with corresponding local change phrases. Experiments on the Remote Sensing Change Caption (RSCC) dataset and LEVIR-CC demonstrate that ChangePixel provides caption quality (ROUGE 19.52/ST5-SCS 76.91 on RSCC; CIDEr-D 56.82 on LEVIR-CC under zero-shot transfer) that is competitive with general-purpose vision–language models (VLMs) while adding pixel-level spatial evidence to change narration; additionally, evidence localization is quantified on LEVIR-MCI through semantic change-mask metrics, reaching Change mIoU 33.8 (15.3 points higher than a non-learned pixel-difference floor of 18.5), whereas phrase-to-region alignment is assessed qualitatively pending a dedicated grounding benchmark. The proposed framework offers a practical path from coarse image-level captioning to evidence-grounded disaster understanding. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Figure 1

22 pages, 3711 KB  
Article
Category-Aware Global–Local Semantic Alignment for Remote Sensing Image–Text Retrieval
by Da Ha and Haisu Zhang
Remote Sens. 2026, 18(14), 2335; https://doi.org/10.3390/rs18142335 - 13 Jul 2026
Viewed by 701
Abstract
In remote sensing image–text retrieval (RSITR), precise cross-modal retrieval is often hindered by centroid drift and ambiguous decision boundaries caused by high inter-class visual similarity. To address these bottlenecks, this study proposes a Category-aware Global–Local Semantic Alignment (CGLSA) framework fine-tuned on the CLIP [...] Read more.
In remote sensing image–text retrieval (RSITR), precise cross-modal retrieval is often hindered by centroid drift and ambiguous decision boundaries caused by high inter-class visual similarity. To address these bottlenecks, this study proposes a Category-aware Global–Local Semantic Alignment (CGLSA) framework fine-tuned on the CLIP (ViT-B/16) backbone. The architecture orchestrates two complementary mechanisms: Global Semantic Collaborative Alignment that regularizes macro-level category centroids using momentum updates and Bayesian prior calibration, and Local Fine-grained Feature Alignment that refines instance-level matching via dynamic scale adjustment and category-aware topological masks. Extensive evaluations on three major benchmark datasets (RSICD, RSITMD, and UCM-Captions) validate the model’s efficacy. Compared to strictly controlled CLIP-family baselines under equivalent supervised conditions, CGLSA achieves new state-of-the-art performance across all R@K metrics and mean recall. Extensions adapting this robust centroid formation to semi-supervised and open-vocabulary scenarios are identified for future work. Full article
Show Figures

Figure 1

21 pages, 32395 KB  
Article
OSM-CLIP: Enhancing Remote Sensing Image–Text Representation Learning with OpenStreetMap Data
by Alessio Pierdominici, Riccardo Ricci, Mohammed Alruqimi and Farid Melgani
Appl. Sci. 2026, 16(14), 7002; https://doi.org/10.3390/app16147002 - 13 Jul 2026
Viewed by 550
Abstract
Remote sensing vision–language models, such as RemoteCLIP and GeoRSCLIP, have advanced image–text representation learning. However, they rely on manually curated caption datasets that are expensive to scale and provide only global image-level supervision. In this paper, we introduce OSM-CLIP, a framework that exploits [...] Read more.
Remote sensing vision–language models, such as RemoteCLIP and GeoRSCLIP, have advanced image–text representation learning. However, they rely on manually curated caption datasets that are expensive to scale and provide only global image-level supervision. In this paper, we introduce OSM-CLIP, a framework that exploits the freely available, continuously growing annotations of OpenStreetMap (OSM) to provide regionally scalable, patch-level supervision for remote sensing image-text learning. We construct a large-scale dataset of over 265,000 satellite images covering the contiguous United States, each automatically paired with fine-grained geographic annotations scraped from OSM and mapped to individual image patches. A contrastive loss operating at the patch level associates each image region with its corresponding OSM textual description, enabling the model to learn spatially grounded representations without any manual labeling effort. After fine-tuning on standard remote sensing captioning datasets, OSM-CLIP achieves an average improvement of 10.81% in zero-shot classification, 5.06% in text-to-image retrieval (R@1), and 3.87% in image-to-text retrieval (R@1) over existing methods across 13 classification and 4 retrieval benchmarks. Our results demonstrate that freely available geographic annotations can serve as a powerful source of supervision for remote sensing vision–language models in regions with high-quality OSM coverage. Full article
Show Figures

Figure 1

21 pages, 26709 KB  
Article
From Landslide Detection to Multi-Source LLM-Based Reporting: A Complete Framework for Rapid Assessment of Post-Disaster Scenarios
by Mohammed Alruqimi, Abdelkader Riche, Pierluigi Confuorto, Mawloud Guermoui, Silvia Bianchini and Farid Melgani
Remote Sens. 2026, 18(11), 1821; https://doi.org/10.3390/rs18111821 - 2 Jun 2026
Viewed by 756
Abstract
Timely landslide detection and rapid qualitative assessment are fundamental to effective warning systems, hazard management, and risk mitigation. Yet, current practices that rely on on-site surveys and manual expert assessment remain risky, costly, and time-consuming. These limitations result in substantial delays between the [...] Read more.
Timely landslide detection and rapid qualitative assessment are fundamental to effective warning systems, hazard management, and risk mitigation. Yet, current practices that rely on on-site surveys and manual expert assessment remain risky, costly, and time-consuming. These limitations result in substantial delays between the event and the availability of actionable information. This study proposes a hybrid, multi-model framework that fuses RGB remote-sensing imagery with geospatial layers to enable timely landslide detection and actionable reporting. The pipeline couples an enhanced SegFormer (denoted as SDF-SegFormer-B2) model for landslide localization, a feature extraction technique for per-slide geo-attribute computation, and a lightweight instruction-tuned LLM (Mistral-7B-Instruct-v0.3) for structured, expert-style reporting. Although a few previous studies have explored landslide captioning, to our knowledge this is the first framework designed to generate structured technical reports enriched with terrain-context interpretation and qualitative intervention-priority indicators. Experiments use 26,758 georeferenced RGB tiles (64 × 64) with 3 m of spatial resolution from PlanetScope satellite imagery over Emilia–Romagna, Italy, with 68,592 annotated landslide boxes collected after the May 2023 rainfall events (~200 mm in 48 h on 1–3 May; 200–250 mm in 48 h on 16–17 May). The proposed SDF-SegFormer-B2 segmentation model achieved a precision of 85.54%, recall of 72.31%, and an F1-score of 78.39% on the unseen test dataset. To evaluate the quality of the generated landslide reports, 100 images were selected for domain-expert assessment. Among these, 58% of the reports were rated as “Very Good,” 30% as “Good,” 8% as “Acceptable,” and 4% as “Poor.” When considering only reports with complete and accurate inputs, 81.48% were rated “Very Good,” and 96.30% were rated either “Good” or “Very Good.” By integrating complementary models and modalities, the proposed approach automates localization-to-reporting and enables the generation of terrain-aware landslide summaries that may support preliminary decision-making and rapid post-disaster screening. Full article
(This article belongs to the Special Issue Artificial Intelligence and Remote Sensing for Geohazards)
Show Figures

Figure 1

30 pages, 13991 KB  
Article
Distribution-Aware CLIP-Adapter with Fine-Grained Text for Few-Shot Fine-Grained Classification
by Jingming Chen, Zhaoyang Huang, Feng Wang, Zixiao Wen, Jingxing Zhu and Guangyao Zhou
Remote Sens. 2026, 18(11), 1813; https://doi.org/10.3390/rs18111813 - 2 Jun 2026
Cited by 1 | Viewed by 545
Abstract
Fine-Grained Few-Shot Classification (FG-FSC) in remote sensing has become a critical task, as the scarcity of high-quality annotated data severely restricts the performance of deep learning models in fine-grained classification. Although Contrastive Language-Image Pre-Training (CLIP) exhibits strong generalization ability in few-shot learning, it [...] Read more.
Fine-Grained Few-Shot Classification (FG-FSC) in remote sensing has become a critical task, as the scarcity of high-quality annotated data severely restricts the performance of deep learning models in fine-grained classification. Although Contrastive Language-Image Pre-Training (CLIP) exhibits strong generalization ability in few-shot learning, it fails to generate discriminative text and image features when adapted to remote sensing tasks. In this paper, a framework is proposed to adapt CLIP to remote sensing FG-FSC from both visual and text aspects. First, we introduce a Distribution-AWare Adapter (DAWA) that adaptively fuses instance-level visual knowledge from few-shot samples with distribution-aware representations derived from Gaussian Discriminant Analysis based on the original CLIP zero-shot knowledge, leading to stable visual feature representations under various few-shot settings. A hybrid loss function that incorporates transductive and contrastive regularization is employed to further prevent overfitting and improve the discriminability of features. Furthermore, we generate category-level fine-grained text captions, optimizing the image–text alignment when extremely few training images are available. Experiments on multiple remote sensing and natural image datasets verify that the proposed framework achieves state-of-the-art few-shot fine-grained classification performance with a modest training cost, providing a practical solution for few-shot remote sensing image analysis. Full article
(This article belongs to the Special Issue Advancements of Vision-Language Models (VLMs) in Remote Sensing)
Show Figures

Figure 1

29 pages, 2266 KB  
Article
Test-Time Candidate-Aware Dual Refinement for Remote Sensing Image–Text Retrieval
by Bofan Zhang and Hao Wu
Remote Sens. 2026, 18(9), 1389; https://doi.org/10.3390/rs18091389 - 30 Apr 2026
Viewed by 816
Abstract
Remote sensing image–text retrieval (RSITR) is a pivotal task aimed at achieving efficient bidirectional matching between visual content and textual descriptions in large-scale remote sensing databases. Nevertheless, it faces a fundamental challenge: the severe information asymmetry between sparse, abstract captions and dense, multi-scale [...] Read more.
Remote sensing image–text retrieval (RSITR) is a pivotal task aimed at achieving efficient bidirectional matching between visual content and textual descriptions in large-scale remote sensing databases. Nevertheless, it faces a fundamental challenge: the severe information asymmetry between sparse, abstract captions and dense, multi-scale overhead imagery. Prior works predominantly focus on learning static cross-modal representations during training; however, this frozen inference process is fundamentally limited in bridging the asymmetry due to its inability to dynamically compensate for missing details or resolve visual ambiguities in heterogeneous scenes. To overcome this limitation, we propose CADRE (Test-Time Candidate-Aware Dual Refinement), a retrieval-backbone-agnostic framework exploiting retrieved candidates as feedback for bidirectional alignment. Operating on a novel Inject-and-Suppress paradigm, CADRE comprises two complementary modules. First, the Visual-Context Injection (VCI) module addresses textual sparsity by incorporating an adaptive filtering mechanism to efficiently mine hierarchical visual evidence from high-confidence candidates and inject it into the query via a domain-adapted Multimodal Large Language Model (MLLM). Second, the Query-Guided Disambiguation (QGD) module targets visual ambiguity by generating multi-view visual hypotheses and utilizing the query as a semantic probe to suppress background noise. Extensive experiments on three standard benchmarks (RSICD, RSITMD, and UCM) demonstrate good transferability across several strong RSITR backbones. Full article
Show Figures

Figure 1

23 pages, 1395 KB  
Article
A Mask-Guided Multigranular Mamba Network for Remote Sensing Change Captioning
by Yifan Qu and Huaidong Zhang
Remote Sens. 2026, 18(7), 1048; https://doi.org/10.3390/rs18071048 - 31 Mar 2026
Viewed by 831
Abstract
Remote sensing image change captioning (RSICC) aims to generate semantic textual descriptions characterizing changes between bi-temporal remote sensing images, with wide applications in disaster assessment and urban planning. However, existing methods face specific drawbacks: CNN-based models have limited ability to capture long-range spatial [...] Read more.
Remote sensing image change captioning (RSICC) aims to generate semantic textual descriptions characterizing changes between bi-temporal remote sensing images, with wide applications in disaster assessment and urban planning. However, existing methods face specific drawbacks: CNN-based models have limited ability to capture long-range spatial correlations due to local receptive fields, and Transformer-based models suffer from quadratic complexity while distributing attention uniformly across all spatial positions, resulting in weak perception of salient changes in background-dominated scenes. In this paper, we present PM3Net (Progressive Mask-guided Multigranular Mamba Network), which leverages Mamba state space models with linear complexity for efficient spatiotemporal change modeling. The Progressive Mask-guided Encoder (PME) creates dual-source change masks combining L2 norm spatial differences with cosine distance semantic differences for progressive change feature extraction from detailed structures to high-level semantics. The Mask-guided Feature Enhancement (MFE) module applies mask-weighted refinement and cross-layer fusion to emphasize salient change regions while suppressing background interference, producing multigranular visual representations. Experiments on LEVIR-MCI and WHU-CDC datasets show PM3Net achieves superior results compared to existing methods, with BLEU-4 scores of 66.89 and 73.05, respectively. The results confirm PM3Net’s ability to solve the RSICC task while demonstrating how Mamba models can succeed in this specific field. Full article
Show Figures

Figure 1

25 pages, 3799 KB  
Article
DR-CLIP: A Deformable Vision–Language Model for Scale-Invariant Object Counting in Remote Sensing Images
by Jingzhe Nie, Qun Liu, Tianze Li, Xu Lu and Liang Zhang
Sensors 2026, 26(6), 1863; https://doi.org/10.3390/s26061863 - 16 Mar 2026
Cited by 1 | Viewed by 854
Abstract
Object counting in remote sensing images is valuable for applications such as urban planning and environmental monitoring. However, it remains challenging due to heterogeneous annotations, semantic ambiguity in open-vocabulary queries, and performance degradation of small targets. To address these limitations, we propose DR-CLIP [...] Read more.
Object counting in remote sensing images is valuable for applications such as urban planning and environmental monitoring. However, it remains challenging due to heterogeneous annotations, semantic ambiguity in open-vocabulary queries, and performance degradation of small targets. To address these limitations, we propose DR-CLIP (Deformable Remote CLIP), a vision–language model for remote sensing image counting that incorporates deformable visual feature extraction with text-guided prediction. DR-CLIP includes a (1) Region-to-Instruction (R2I) mechanism to convert points, bounding boxes, and polygons into a unified image–text training representation, a (2) Multi-scale Deformable Attention (MSDA) to enhance discriminative feature extraction across extreme scale variations and cluttered backgrounds, and a (3) Text-Guided Counting Head that establishes robust cross-modal alignment through contrastive learning, achieving open-vocabulary counting capability without category-specific retraining. On DOTA-v2.0, DR-CLIP achieves a Mean Absolute Error (MAE) of 2.34 and a Root Mean Squared Error (RMSE) of 3.89, outperforming baselines by 19.0% in MAE. The MSDA module significantly increases Small-Object Recall (SOR) to 0.824, which is especially effective in situations involving dense and small object counting. In cross-modal retrieval, DR-CLIP attains R@1 scores of 68.3% (image-to-text) and 72.1% (text-to-image) on the Remote Sensing Image Captioning Dataset (RSICD). The framework generalizes robustly, with only 8.7% performance degradation in cross-domain tests, which is significantly lower than the 23.4% drop observed in baseline methods. Full article
(This article belongs to the Section Remote Sensors)
Show Figures

Figure 1

25 pages, 11205 KB  
Article
Remote Sensing Image Captioning via Self-Supervised DINOv3 and Transformer Fusion
by Maryam Mehmood, Ahsan Shahzad, Farhan Hussain, Lismer Andres Caceres-Najarro and Muhammad Usman
Remote Sens. 2026, 18(6), 846; https://doi.org/10.3390/rs18060846 - 10 Mar 2026
Cited by 1 | Viewed by 1812
Abstract
Effective interpretation of coherent and usable information from aerial images (e.g., satellite imagery or high-altitude drone photography) can greatly reduce human effort in many situations, both natural (e.g., earthquakes, forest fires, tsunamis) and man-made (e.g., highway pile-ups, traffic congestion), particularly in disaster management. [...] Read more.
Effective interpretation of coherent and usable information from aerial images (e.g., satellite imagery or high-altitude drone photography) can greatly reduce human effort in many situations, both natural (e.g., earthquakes, forest fires, tsunamis) and man-made (e.g., highway pile-ups, traffic congestion), particularly in disaster management. This research proposes a novel encoder–decoder framework for captioning of remote sensing images that integrates self-supervised DINOv3 visual features with a hybrid Transformer–LSTM decoder. Unlike existing approaches that rely on supervised CNN-based encoders (e.g., ResNet, VGG), the proposed method leverages DINOv3’s self-supervised learning capabilities to extract dense, semantically rich features from aerial images without requiring domain-specific labeled pretraining. The proposed hybrid decoder combines Transformer layers for global context modeling with LSTM layers for sequential caption generation, producing coherent and context-aware descriptions. Feature extraction is performed using the DINOv3 model, which employs the gram-anchoring technique to stabilize dense feature maps. Captions are generated through a hybrid of Transformer with Long Short-Term Memory (LSTM) layers, which adds contextual meaning to captions through sequential hidden layer modeling with gated memory. The model is first evaluated on two traditional remote sensing image captioning datasets: RSICD and UCM-Captions. Multiple evaluation metrics like Bilingual Evaluation Understudy (BLEU), Consensus-based Image Description Evaluation (CIDEr), Recall-Oriented Understudy for Gisting Evaluation (ROUGE-L), and Metric for Evaluation of Translation with Explicit Ordering (METEOR), are used to quantify the performance and robustness of the proposed DINOv3 hybrid model. The proposed model outperforms conventional Convolutional Neural Network (CNN) and Vision Transformers (ViT)-based models by approximately 9–12% across most evaluation metrics. Attention heatmaps are also employed to qualitatively validate the proposed model when identifying and describing key spatial elements. In addition, the proposed model is evaluated on advanced remote sensing datasets, including RSITMD, DisasterM3, and GeoChat. The results demonstrate that self-supervised vision transformers are robust encoders for multi-modal understanding in remote sensing image analysis and captioning. Full article
Show Figures

Figure 1

22 pages, 1012 KB  
Article
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-Guided Difference Perception
by Pei Deng, Wenqian Zhou and Hanlin Wu
Remote Sens. 2026, 18(4), 541; https://doi.org/10.3390/rs18040541 - 8 Feb 2026
Cited by 3 | Viewed by 1193
Abstract
The accurate interpretation of land cover changes in multi-temporal satellite imagery is critical for Earth observation. However, existing methods typically yield static outputs—such as binary masks or fixed captions—lacking interactivity and user guidance. To address this limitation, we introduce remote sensing image change [...] Read more.
The accurate interpretation of land cover changes in multi-temporal satellite imagery is critical for Earth observation. However, existing methods typically yield static outputs—such as binary masks or fixed captions—lacking interactivity and user guidance. To address this limitation, we introduce remote sensing image change analysis (RSICA), a novel paradigm that enables the instruction-guided, multi-turn exploration of temporal differences in bi-temporal images through visual question answering. To realize RSICA, we propose DeltaVLM, a vision language model specifically designed for interactive change understanding. DeltaVLM comprises three key components: (1) a fine-tuned bi-temporal vision encoder that independently extracts semantic features from each image in the input pair; (2) a visual difference perception module with a cross-semantic relation measuring (CSRM) mechanism to interpret changes; and (3) an instruction-guided Q-former that selects query-relevant change features and aligns them with a frozen large language model to generate context-aware responses. We also present ChangeChat-105k, a large-scale instruction-following dataset containing over 105k diverse samples. Extensive experiments show that DeltaVLM achieves state-of-the-art performance in both single-turn captioning and multi-turn interactive change analysis, surpassing both general multimodal models and specialized remote sensing vision language models. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Graphical abstract

22 pages, 7754 KB  
Article
CSSA: A Cross-Modal Spatial–Semantic Alignment Framework for Remote Sensing Image Captioning
by Xiao Han, Zhaoji Wu, Yunpeng Li, Xiangrong Zhang, Guanchun Wang and Biao Hou
Remote Sens. 2026, 18(3), 522; https://doi.org/10.3390/rs18030522 - 5 Feb 2026
Viewed by 1266
Abstract
Remote sensing image captioning (RSIC) aims to generate natural language descriptions for the given remote sensing image, which requires a comprehensive and in-depth understanding of image content and summarizes it with sentences. Most RSIC methods have successful vision feature extraction, but the representation [...] Read more.
Remote sensing image captioning (RSIC) aims to generate natural language descriptions for the given remote sensing image, which requires a comprehensive and in-depth understanding of image content and summarizes it with sentences. Most RSIC methods have successful vision feature extraction, but the representation of spatial features or fusion features fails to fully consider cross-modal differences between remote sensing images and texts, resulting in unsatisfactory performance. Thus, we propose a novel cross-modal spatial–semantic alignment (CSSA) framework for an RSIC task, which consists of a multi-branch cross-modal contrastive learning (MCCL) mechanism and a dynamic geometry Transformer (DG-former) module. Specifically, compared to discrete text, remote sensing images present a noisy property, interfering with the extraction of valid vision features. Therefore, we present an MCCL mechanism to learn consistent representation between image and text, achieving cross-modal semantic alignment. In addition, most objects are scattered in remote sensing images and exhibit a sparsity property due to the overhead view. However, the Transformer structure mines the objects’ relationships without considering the geometry information of the objects, leading to suboptimal capture of the spatial structure. To address this, a DG-former is designed to realize spatial alignment by introducing geometry information. We conduct experiments on three publicly available datasets (Sydney-Captions, UCM-Captions and RSICD), and the superior results demonstrate its effectiveness. Full article
Show Figures

Figure 1

23 pages, 1579 KB  
Article
Exploring Difference Semantic Prior Guidance for Remote Sensing Image Change Captioning
by Yunpeng Li, Xiangrong Zhang, Guanchun Wang and Tianyang Zhang
Remote Sens. 2026, 18(2), 232; https://doi.org/10.3390/rs18020232 - 11 Jan 2026
Cited by 1 | Viewed by 1660
Abstract
Understanding complex change scenes is a crucial challenge in remote sensing field. Remote sensing image change captioning (RSICC) task has emerged as a promising approach to translate appeared changes between bi-temporal remote sensing images into textual descriptions, enabling users to make accurate decisions. [...] Read more.
Understanding complex change scenes is a crucial challenge in remote sensing field. Remote sensing image change captioning (RSICC) task has emerged as a promising approach to translate appeared changes between bi-temporal remote sensing images into textual descriptions, enabling users to make accurate decisions. Current RSICC methods frequently encounter difficulties in consistency for contextual awareness and semantic prior guidance. Therefore, this study explores difference semantic prior guidance network to reason context-rich sentence for capturing appeared vision changes. Specifically, the context-aware difference module is introduced to guarantee the consistency of unchanged/changed context features, strengthening multi-level changed information to improve the ability of semantic change feature representation. Moreover, to effectively mine higher-level cognition ability to reason salient/weak changes, we employ difference comprehending with shallow change information to realize semantic change knowledge learning. In addition, the designed parallel cross refined attention in Transformer decoder can balance vision difference and semantic knowledge for implicit knowledge distilling, enabling fine-grained perception changes of semantic details and reducing pseudochanges. Compared with advanced algorithms on the LEVIR-CC and Dubai-CC datasets, experimental results validate the outstanding performance of the designed model in RSICC tasks. Notably, on the LEVIR-CC dataset, it reaches a CIDEr score of 143.34%, representing a 3.11% improvement over the most competitive SAT-cap. Full article
Show Figures

Figure 1

31 pages, 6416 KB  
Article
FireMM-IR: An Infrared-Enhanced Multi-Modal Large Language Model for Comprehensive Scene Understanding in Remote Sensing Forest Fire Monitoring
by Jinghao Cao, Xiajun Liu and Rui Xue
Sensors 2026, 26(2), 390; https://doi.org/10.3390/s26020390 - 7 Jan 2026
Cited by 2 | Viewed by 1713
Abstract
Forest fire monitoring in remote sensing imagery has long relied on traditional perception models that primarily focus on detection or segmentation. However, such approaches fall short in understanding complex fire dynamics, including contextual reasoning, fire evolution description, and cross-modal interpretation. With the rise [...] Read more.
Forest fire monitoring in remote sensing imagery has long relied on traditional perception models that primarily focus on detection or segmentation. However, such approaches fall short in understanding complex fire dynamics, including contextual reasoning, fire evolution description, and cross-modal interpretation. With the rise of multi-modal large language models (MLLMs), it becomes possible to move beyond low-level perception toward holistic scene understanding that jointly reasons about semantics, spatial distribution, and descriptive language. To address this gap, we introduce FireMM-IR, a multi-modal large language model tailored for pixel-level scene understanding in remote-sensing forest-fire imagery. FireMM-IR incorporates an infrared-enhanced classification module that fuses infrared and visual modalities, enabling the model to capture fire intensity and hidden ignition areas under dense smoke. Furthermore, we design a mask-generation module guided by language-conditioned segmentation tokens to produce accurate instance masks from natural-language queries. To effectively learn multi-scale fire features, a class-aware memory mechanism is introduced to maintain contextual consistency across diverse fire scenes. We also construct FireMM-Instruct, a unified corpus of 83,000 geometrically aligned RGB–IR pairs with instruction-aligned descriptions, bounding boxes, and pixel-level annotations. Extensive experiments show that FireMM-IR achieves superior performance on pixel-level segmentation and strong results on instruction-driven captioning and reasoning, while maintaining competitive performance on image-level benchmarks. These results indicate that infrared–optical fusion and instruction-aligned learning are key to physically grounded understanding of wildfire scenes. Full article
(This article belongs to the Special Issue Remote Sensing and UAV Technologies for Environmental Monitoring)
Show Figures

Figure 1

25 pages, 8526 KB  
Article
Describing Land Cover Changes via Multi-Temporal Remote Sensing Image Captioning Using LLM, ViT, and LoRA
by Javier Lamar León, Vitor Nogueira, Pedro Salgueiro and Paulo Quaresma
Remote Sens. 2026, 18(1), 166; https://doi.org/10.3390/rs18010166 - 4 Jan 2026
Cited by 5 | Viewed by 2196
Abstract
Describing land cover changes from multi-temporal remote sensing imagery requires capturing both visual transformations and their semantic meaning in natural language. Existing methods often struggle to balance visual accuracy with descriptive coherence. We propose MVLT-LoRA-CC (Multi-modal Vision Language Transformer with Low-Rank Adaptation for [...] Read more.
Describing land cover changes from multi-temporal remote sensing imagery requires capturing both visual transformations and their semantic meaning in natural language. Existing methods often struggle to balance visual accuracy with descriptive coherence. We propose MVLT-LoRA-CC (Multi-modal Vision Language Transformer with Low-Rank Adaptation for Change Captioning), a framework that integrates a Vision Transformer (ViT), a Large Language Model (LLM), and Low-Rank Adaptation (LoRA) for efficient multi-modal learning. The model processes paired temporal images through patch embeddings and transformer blocks, aligning visual and textual representations via a multi-modal adapter. To improve efficiency and avoid unnecessary parameter growth, LoRA modules are selectively inserted only into the attention projection layers and cross-modal adapter blocks rather than being uniformly applied to all linear layers. This targeted design preserves general linguistic knowledge while enabling effective adaptation to remote sensing change description. To assess performance, we introduce the Complementary Consistency Score (CCS) framework, which evaluates both descriptive fidelity for change instances and classification accuracy for no change cases. Experiments on the LEVIR-CC test set demonstrate that MVLT-LoRA-CC generates semantically accurate captions, surpassing prior methods in both descriptive richness and temporal change recognition. The approach establishes a scalable solution for multi-modal land cover change description in remote sensing applications. Full article
Show Figures

Figure 1

Back to TopTop