Next Article in Journal
Spatial and Temporal Patterns of Mangrove Forest Change in the Mekong Region over Four Decades Based on a Remote Sensing Data-Driven Approach
Previous Article in Journal
High-Resolution Remote Sensing and People-to-Pixel Integration for Mapping Farmland Abandonment in Central Himalayan Villages
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Hierarchical Prompt Engineering for Remote Sensing Scene Understanding with Large Vision–Language Models

College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai 200433, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2025, 17(22), 3727; https://doi.org/10.3390/rs17223727
Submission received: 29 September 2025 / Revised: 31 October 2025 / Accepted: 13 November 2025 / Published: 16 November 2025

Abstract

Vision–language models (VLMs) show strong potential for remote-sensing scene classification but still struggle with fine-grained categories and distribution shifts. We introduce a hierarchical prompting framework that decomposes recognition into a coarse-to-fine decision process with structured outputs, combined with parameter-efficient adaptation using LoRA/QLoRA. To evaluate robustness without depending on external benchmarks, we construct five protocol variants of the AID (V0–V4) that systematically vary label granularity, class consolidation, and augmentation settings. Each variant is designed to align with a specific prompting style and hierarchy. The data pipeline follows a strict split-before-augment strategy, in which augmentation is applied only to the training split to avoid train-test leakage. We further audit leakage using rotation/flip–invariant perceptual hashing across splits to ensure reproducibility. Experiments on all five AID variants show that hierarchical prompting consistently outperforms non-hierarchical prompts and matches or exceeds full fine-tuning, while requiring substantially less compute. Ablation studies on prompt design, adaptation strategy, and model capacity—together with confusion matrices and class-wise metrics—indicate improved recognition at both coarse and fine levels, as well as robustness to rotations and flips. The proposed framework provides a strong, reproducible baseline for remote-sensing scene classification under constrained compute and includes complete prompt templates and processing scripts to support replication.
Keywords: remote sensing; scene classification; hierarchical prompting; LoRA/QLoRA; data leakage prevention; split-before-augment; AID remote sensing; scene classification; hierarchical prompting; LoRA/QLoRA; data leakage prevention; split-before-augment; AID

Share and Cite

MDPI and ACS Style

Chen, T.; Ai, J. Hierarchical Prompt Engineering for Remote Sensing Scene Understanding with Large Vision–Language Models. Remote Sens. 2025, 17, 3727. https://doi.org/10.3390/rs17223727

AMA Style

Chen T, Ai J. Hierarchical Prompt Engineering for Remote Sensing Scene Understanding with Large Vision–Language Models. Remote Sensing. 2025; 17(22):3727. https://doi.org/10.3390/rs17223727

Chicago/Turabian Style

Chen, Tianyang, and Jianliang Ai. 2025. "Hierarchical Prompt Engineering for Remote Sensing Scene Understanding with Large Vision–Language Models" Remote Sensing 17, no. 22: 3727. https://doi.org/10.3390/rs17223727

APA Style

Chen, T., & Ai, J. (2025). Hierarchical Prompt Engineering for Remote Sensing Scene Understanding with Large Vision–Language Models. Remote Sensing, 17(22), 3727. https://doi.org/10.3390/rs17223727

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop