SAREval: A Multi-Dimensional and Multi-Task Benchmark for Evaluating Visual Language Models on SAR Image Understanding
Highlights
- We establish SAREval, the first comprehensive benchmark specifically designed for evaluating Vision-Language Models on SAR image understanding, comprising 20 tasks across perception, reasoning, and robustness dimensions with 14,950 expert-verified image–text pairs.
- Current VLMs demonstrate significant limitations in SAR image interpretation, particularly in fine-grained target discrimination and physical-attribute reasoning tasks, while exhibiting unexpected performance improvements under certain noise conditions that challenge conventional robustness understanding.
- The benchmark provides a standardized platform for developing and evaluating SAR-specific VLMs, addressing the critical gap in specialized evaluation frameworks for radar remote sensing and enabling fair comparison across different model architectures.
- Our findings reveal fundamental domain adaptation challenges in VLMs and establish a new paradigm for constructing multimodal benchmarks in specialized remote sensing domains, with direct implications for maritime surveillance, infrastructure monitoring, and environmental observation applications.
Abstract
1. Introduction
- Limited Support for SAR-VLM Alignment Needs: Existing benchmarks do not adequately meet the alignment requirements for SAR-VLM evaluation. For instance, they lack image–text pairs that reflect SAR-specific characteristics such as microwave scattering mechanisms (e.g., metallic targets appearing as bright spots due to strong reflection, while water bodies appear dark due to weak reflection). Textual descriptions also often fail to balance technical accuracy with semantic readability, which makes conventional optical image–text alignment paradigms unsuitable for SAR data.
- High Cost of Professional SAR Annotation: Interpreting SAR imagery requires deep domain expertise. Manual annotation involves not only delineating object boundaries but also providing specialized information such as scattering mechanisms and imaging parameters, leading to low annotation efficiency and difficulty in maintaining consistency.
- Lack of SAR-Adapted Task Paradigms: General-purpose benchmarks designed for optical imagery cannot fully address the interpretation needs of the SAR domain. Meanwhile, SAR-specific tasks—such as polarization mode inference, speckle noise robustness, and geometric distortion correction—have not yet established mature evaluation paradigms or quantifiable metrics, further hindering the development of SAR benchmarks.
- We construct a multi-level evaluation benchmark centered on perception, reasoning and robustness dimensions, deeply integrating SAR-specific characteristics. Unlike existing benchmarks that focus primarily on optical data or include only limited SAR tasks, SAREval is designed specifically for SAR data, with tasks closely aligned with SAR imaging mechanisms and scattering characteristics. It covers 20 specialized tasks ranging from image classification and object detection to physical interpretation and sensor parameter inversion. This offers a standardized platform for fine-grained assessment of model adaptability to SAR’s unique properties.
- We propose a task-adaptive hybrid annotation framework and introduce an innovative distractor generation strategy based on “physical constraints + gradient divergence.” To tackle the challenges of SAR image–text alignment and high annotation cost, we combine a label-driven method, LLM-assisted method, and expert verification into a unified data generation pipeline. Specifically, for numerical reasoning tasks, we design a three-tier distractor system based on physical constraints and gradient divergence, ensuring that distractors are both physically plausible and capable of discriminating between different levels of reasoning ability. Under this framework, we produce over ten thousand high-quality test samples across diverse scenarios, balancing professionalism and scale.
- We conduct systematic benchmarking of multiple state-of-the-art VLMs and provide critical performance insights and analysis. We comprehensively evaluate 11 mainstream models on SAREval, quantifying their performance differences across various tasks and deeply analyzing their limitations in specialized tasks such as geometric measurement and physical state reasoning. This offers empirical evidence to guide future development and optimization of VLMs for SAR data.
2. Related Works
2.1. Remote Sensing Multimodal Benchmarks
2.2. Strategies for Constructing Visual Language Benchmark Data
3. SAREval
3.1. Hierarchical Capability Taxonomy
3.1.1. Perception Capability Dimension
3.1.2. Reasoning Capability Dimension
3.1.3. Robustness Capability Dimension
3.2. Dataset Construction and Quality Control
3.2.1. Data Source
3.2.2. Task-Adaptive Hybrid Annotation Framework
4. Evaluation Based on SAREval
4.1. Experimental Setup
4.2. Main Results
4.2.1. Perception Capability Assessment
4.2.2. Reasoning Capability Assessment
4.2.3. Robustness Capability Assessment
5. Discussion and Analysis
5.1. Performance Disparity Analysis
5.2. Critical Bottleneck
5.3. Future Directions
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| VLMs | Vision Language Models |
| SAR | Synthetic Aperture Radar |
| REC | Referring Expression Comprehension |
| CLIP | Contrastive Language–Image Pre-training |
| ALBEF | Align Before Fuse |
| RNNs | Recurrent Neural Networks |
| FCNs | Fully Convolutional Networks |
| LLaVA | Large Language and Vision Assistant |
| MLLMs | Multimodal Large Language Models |
| CNNs | Convolutional Neural Networks |
| LSTM | Long Short-term Memory |
| LLM | Large Language Model |
| ID | Image Description |
| SC | Scene Classification |
| MCOD | Multi-Class Object Detection |
| ACD | Aircraft Classification Detection |
| SCD | Ship Classification Detection |
| VCD | Vehicle Classification Detection |
| ObjC | Object Counting |
| VG | Visual Grounding |
| LUTS | Land Use Type Segmentation |
| MOSD | Marine Oil Spill Detection |
| IBC | Interference Caused by Background Clutter |
| ICM | Interference Caused by Coherent Mechanisms |
| IRI | Imaging Resolution Inference |
| IAI | Incidence Angle Inference |
| POI | Polarization Orientation Inference |
| TAI | Target Azimuth Inference |
| OST-VE | Oil Storage Tank Volume Estimation |
| SMDE | Ship Motion Direction Estimation |
| BHE | Building Height Estimation |
| TDM | Target Dimension Measurement |
| Img-U | Image-Level Understanding |
| Obj-R | Object-Level Recognition |
| Pix-S | Pixel-Level Segmentation |
| RI | Robustness Against Interference |
| IPE | Imaging Parameter Estimation |
| LCAI | Land Cover Attribute Inference |
| TGM | Terrain Geometry Measurement |
| ROUGE-L | Recall-Oriented Understudy for Gisting Evaluation—Longest common subsequence |
| BLEU | Bilingual Evaluation Understudy |
| NLP | Natural Language Processing |
| OBBs | Oriented Bounding Boxes |
| ViT | Vision Transformer |
Appendix A
Appendix A.1. Dataset Sources
| Dataset | Imaging Platform | Resolution | Band | Polarization | Task Scenarios |
|---|---|---|---|---|---|
| MSAR-1.0 [39] | Haisi-1, Gaofen-3 | 1 m | C | HH, VV, HV, VH | Multi-class Object Detection, Object Counting, Visual Grounding, Scene Classification, Polarization Orientation Inference |
| SAR-AIRcraft-1.0 [40] | Gaofen-3 | 1 m | C | VV, HH | Multi-class Object Detection, Aircraft Classification Detection |
| FUSAR-Ship 1.0 [54] | Gaofen-3 | 1 m | C | VV, HH | Ship Classification Detection |
| SAR-Airport [55] | Sentinel-1B | 5 m | C | VH | Scene Classification |
| SIVED [56] | Airborne SAR | 0.1–0.3 m | X, Ku, Ka | VV, HH | Multi-class Object Detection, Object Counting |
| ATRNet-STAR [45] | Airborne SAR | 0.1–0.3 m | X, Ku | VV, HH, HV, VH | Vehicle Classification Detection, Target Azimuth Inference, Imaging Resolution Inference |
| OpenEarthMap-SAR [41] | - | 0.15–0.5 m | - | VV, HH | Scene Classification, Land Use Type Segmentation |
| SOS [57] | ALOS-PALSAR, Sentinel-1A | 5 m, 10 m | L, C | - | Marine Oil Spill Detection |
| OpenSARWake [44] | ALOS-PALSAR, Sentinel-1A, TerraSAR-X | 1.25–12.5 m | L, C, X | VV, HH | Ship Motion Direction Estimation |
| FAIR-CSAR [42] | Gaofen-3 | SL: 1 m, FSI: 5 m | C | VV, HH | Multi-class Object Detection, Aircraft Classification Detection, Object Counting, Oil Storage Tank Volume Estimation, Incidence Angle Inference, Target Dimension Measurement |
| QXS-SARPOT [58] | Gaofen-3 | 1 m | C | VV, HH | Image Description, Imaging Resolution Inference |
| BHdataset [43] | Sentinel-1A | 2.5 m | C | VV, VH | Building Height Estimation |
| SARBuD [59] | Gaofen-3 | 10 m | C | VV, HH | Imaging Resolution Inference |
Appendix A.1.1. Perception Dimension
Appendix A.1.2. Reasoning Dimension
Appendix A.1.3. Robustness Dimension
Appendix A.2. Data Structure Analysis
Appendix A.2.1. Overall Dataset Architecture
| L1 Level | L2 Level | L3 Level | Number of Tasks |
|---|---|---|---|
| Perception | Image-level Understanding | Scene Classification | 632 |
| Image Description | 395 | ||
| Object-level Recognition | Multi-class Object Detection | 856 | |
| Ship Classification Detection | 284 | ||
| Aircraft Classification Detection | 222 | ||
| Vehicle Classification Detection | 530 | ||
| Object Counting | 760 | ||
| Visual Grounding | 3605 | ||
| Pixel-level Segmentation | Land Use Type Segmentation | 3788 | |
| Marine Oil Spill Detection | 237 | ||
| Reasoning | Terrain Geometry Measurement | Target Dimension Measurement | 408 |
| Building Height Estimation | 697 | ||
| Object Attribute Inference | Ship Motion Direction Estimation | 404 | |
| Oil Storage Tank Volume Estimation | 386 | ||
| Imaging Parameter Estimation | Target Azimuth Inference | 200 | |
| Polarization Orientation Inference | 204 | ||
| Incidence Angle Inference | 117 | ||
| Imaging Resolution Inference | 227 | ||
| Robustness | Robustness against Interference | Speckle Noise Interference | 856 |
| Clutter Interference | 142 |

Appendix A.2.2. Quantitative Analysis of Dataset Composition


Appendix A.3. VLMs Selection
| Model Type | Model Name | Visual Encoder | Language Model | Release Date | Parameters | Training Scale |
|---|---|---|---|---|---|---|
| Intern Series | InternVL2–4B | InternViT-6B | InternLM2-7B | June 2024 | 4B | 1.25 B image–text pairs + 1.25 M instructions |
| Mono-InternVL-2B | InternViT-6B | InternLM2-7B | October 2024 | 2B | 980 M image–text pairs + 1.15 M instructions | |
| InternVideo2.5-4B | InternViT-300M | InternLM2.5-4B | January 2025 | 4B | 28 M videos + 1.25 B images + 950 K instructions | |
| InternVL2.5-8B | InternViT-300M | InternLM2.5-8B | January 2025 | 8B | 28 M videos + 1.25 B images + 960 K instructions | |
| InternVL3-9B | InternViT-6B | InternLM3-8B | April 2025 | 9B | 1.5 B multimodal pairs + 1.4 M instructions | |
| LLaVA Series | LLaVA-1.5-7B | CLIP-ViT-L/14 | Vicuna-7B | October 2023 | 7B | 558 K image–text pairs + 665 K instructions |
| LLaVA-Next | CLIP-ViT-L/14 | Vicuna-13B | January 2024 | 13B | 558 K image–text pairs + 1.32 M instructions | |
| LLaVA--onevision-qwen2-7b | CLIP-ViT-L/14 | Vicuna-7B | October 2024 | 7B | 250 M multimodal pairs + 9.39 M instructions | |
| Others | Qwen2.5-VL-7B | ViT | Qwen2.5-7B | February 2024 | 7B | 1.35 B multimodal pairs + 3.2 M instructions |
| Phi-3.5-vision | CLIP-ViT-L/14 | Phi-3 Mini | August 2024 | 4.2B | 420 M image–text pairs + 1.8 M instructions | |
| DeepSeek-VL2-tiny | SigLIP-L | DeepSeek-MoE 3B | March 2025 | 1B | 850 M image–text pairs + 2.2 M instructions |
Appendix A.4. Expert Annotation and Quality Control
Appendix A.4.1. Expert Annotation Protocol
| Task Type | Initial Annotated Sample Count | Directly Matched Sample Size | Direct Concordance Rate | Cohens κ | Valid Data After Review |
|---|---|---|---|---|---|
| Ship Motion Direction Estimation | 450 | 352 | 78.2% | 0.76 | 404 |
| Oil Storage Tank Volume Estimation | 420 | 298 | 71.0% | 0.68 | 386 |
Appendix A.4.2. Quality Control for LLM-Assisted Generation
Appendix A.5. Large Model Generation Details
| Prompt: |
|---|
|
You are a senior expert in SAR remote sensing and image description evaluation. Please act as an automatic scorer to comprehensively evaluate the quality of the predicted SAR image description based on three core dimensions: accuracy, professionalism, and completeness. Provide a quantitative score between 0 and 100 (integer), where a higher score indicates better description quality.
[Evaluation Dimensions & Weighting] 1. Accuracy (Weight: 0.4): - Whether the predicted description is completely consistent with the reference content (e.g., target category, attributes, spatial relationships, SAR image characteristics such as polarization and scattering properties); - No hallucinated information (e.g., non-existent targets, incorrect attributes) or missing key factual content. 2. Professionalism (Weight: 0.3): - Whether the description uses standardized SAR remote sensing terminology (e.g., “backscattering coefficient”, “polarization mode (HH/HV/VV/VH)”, “speckle noise”, “ship target”) instead of colloquial or ambiguous expressions; - Whether the technical expression conforms to academic norms in the SAR remote sensing field. 3. Completeness (Weight: 0.3): - Whether the description covers both global scene information (e.g., image coverage, main ground object types, overall environment) and local detail information (e.g., target quantity, position, size, morphological features, SAR-specific characteristics); - Whether all key elements in the reference description are fully reflected without omitting critical content. [Reference Materials] Reference Description (Ground Truth): {ref_caption} Predicted Description (to be evaluated): {pred_caption} [Scoring Requirements] 1. Calculate the sub-score for each dimension first (0–100 points), then compute the final score using the weighted sum formula: Final Score = (Accuracy Score × 0.4) + (Professionalism Score × 0.3) + (Completeness Score × 0.3); 2. The final score must be an integer between 0 and 100 (no decimals); 3. Only return the final quantitative score, without any additional explanations, comments, or formatting (e.g., no “Score:” prefix). |
| Prompt: |
|---|
|
You are an expert in SAR remote sensing and cross-modal image interpretation. Based on the reference attribute information of visible light-SAR paired data, you need to generate a structured description for the given SAR image that is accurate, comprehensive, and professional. The SAR image is sourced from OpenEarthMap-SAR, with the satellite platform: Gaofen-3, imaging parameters: C-band spotlight imaging mode, and spatial resolution of 1 m; the paired visible light image provides cross-modal reference support. Please strictly follow the following generation principles and requirements:
1. Specifications for Image Attributes: - Must clearly describe core metadata: image source, satellite platform, sensor type (SAR image), imaging parameters (C-band, spotlight imaging mode), and spatial resolution (1 m); - Do not mention undefined information such as image color mode (e.g., grayscale image/color image) or acquisition time (e.g., morning/evening), and avoid subjective speculative statements. 2. Core Requirements for Ground Object Description: - Fully cover the target categories specified in the reference attributes, and detailedly label key ground object attributes: quantity, material, morphological characteristics, actual size, and spatial location (including pixel-level/geographic coordinate references and relative positional relationships between ground objects); - Utilize the cross-modal complementary information from visible light images to calibrate ambiguous or easily confused ground object features in SAR images (e.g., verifying target boundaries and supplementing structural details that are difficult to identify via SAR with visible light) to improve description accuracy; - Determine target uniqueness based on reference attributes: if the target is unique, clearly label its distinctive identification features; if there are multiple targets of the same type, distinguish them by differences in location, size, and morphology. 3. Logic for Description Structure: - Strictly follow the progressive structure of “global scene overview → local detail focus”: first outline the overall coverage of the image, composition of core ground objects, and scene type (e.g., offshore area, suburban area); then conduct refined descriptions of individual target ground objects to ensure clear hierarchy and coherent logic. 4. Refined Standards for Typical Ground Objects: - Roads: Clarify shape (straight/curved/polyline), width magnitude, extended length characteristics, and direction; label road surface roughness-related features based on SAR scattering properties; - Buildings: Describe distribution density, single-building size, and roof structural features (color cannot be identified by SAR images, so no need to describe); if auxiliary facilities (e.g., courtyards, parking lots) are mentioned in the reference attributes, they can be supplemented; - Airports: Cover core facilities such as terminals, aprons, jet bridges, and boarding gates, and refine facility layout, relative positions, and scale characteristics; - Typical remote sensing targets such as ships/bridges: Supplement descriptions of strong/weak scattering features based on polarization mode and scattering properties, and calibrate target boundaries and size parameters. 5. Accuracy and Output Specifications: - Strictly generate descriptions based on reference attribute information and visible features of bimodal images; verify the authenticity of each piece of content one by one to eliminate hallucinations, attribute errors, and unfounded speculations; - Adopt coherent paragraph-style expression and avoid using lists; focus on objective feature descriptions of ground objects, reduce irrelevant expressions such as aesthetic evaluations and subjective feelings, and use concise and professional language; - The output must be compatible with JSON structured organization requirements to ensure that the description content can be directly extracted into fields without format confusion or redundant information. |
| Prompt: |
|---|
|
You are a professional expert in SAR remote sensing and target counting. Based on the following task definition and generation requirements, generate problem variants that comply with academic standards and are suitable for the oil tank target counting scenario in single-modal SAR images.
[Core Task Definition] - Task Type: Oil tank target counting in single-modal SAR images - Core Objective: Guide the model under test to accurately identify and count the total number of oil tank targets in SAR images, adapting to the characteristics of SAR images and the remote sensing recognition features of oil tank facilities - Input Data: Single-modal SAR images (SAR sensor characteristics such as polarization mode and spatial resolution do not need to be explicitly reflected in the questions, but the question expression must adapt to the target presentation logic of SAR images) - Output Requirement: A clear integer representing the number of oil tank targets, without additional attribute descriptions (e.g., size, location, material, etc.) [Detailed Generation Requirements] 1. Domain Adaptability: -Use professional and diverse terminology related to oil tanks, alternating between “oil tanks” (core term), “oil storage tanks” (more precise expression), and “oil storage facilities” (general term for oil tank-type storage facilities) to conform to remote sensing field expression habits; -Use standardized expressions for images, such as “SAR image” (core expression), “satellite image” (associating with satellite remote sensing scenarios), and “the scene” (scenario-based expression), avoiding colloquial expressions unrelated to the remote sensing field; -Focus on the core task of “counting” and do not introduce judgments on other attributes of oil tank targets (e.g., type, status), ensuring the task is single and clear. 2. Problem Diversity: -Quantity: Generate 10–15 problem variants covering 3 types: Direct Question Type (explicitly requiring “count” or asking about “number”), Scenario-based Question Type (combining image/scene expressions), and Generalized Facility Type (using generalized terms such as “oil storage facilities”); -Sentence Structure Variation: Flexibly adjust sentence structures, adopting different interrogative forms such as “How many…?”, “What is the total count of…?”, and “Please count the number of…?” to avoid repetition and redundancy; -Vocabulary Variation: Alternate core verbs (count), core nouns (oil tanks), and image expressions (SAR image/satellite image) to enhance the richness of problem variants. 3. Logical Rigor: -Each question must clearly present the logical chain of “Input (SAR image) → Task (counting) → Output (quantity)” without ambiguity; -Avoid vague expressions, and use words such as “visible”, “shown”, “present”, and “identify” to limit the counting scope to “oil tank targets identifiable in the image”; -Appropriately add words like “exact” to emphasize the requirement for counting accuracy without increasing task complexity. |
Appendix A.6. Qualitative Failure Cases in SAR Image Description

Appendix B. SAREval Dataset Documentation
Appendix B.1. Dataset Structure
| Text |
|---|
![]() |
- Image Storage: Each subfolder under ‘images/’ contains the image files associated with its specific task (e.g., ‘aircraft_1.png’ is located in ‘images/AircraftClassificationDetection/’).
- Dual Annotation Mapping:
- JSON Files: Provide detailed annotations with multiple prompt templates and structured fields, directly corresponding to the task folders under ‘images/’.
- TSV Files: Located in the ‘LMUData/’ folder, these files offer condensed, tab-separated summaries of the same task’s annotations, facilitating efficient batch processing.
Appendix B.2. Annotation Format Details
Appendix B.2.1. JSON Format
| JSON |
|---|
|
{
"image_path": "aircraft_1.png", # Relative path to the image "ground_truth": "Airbus_A220", # Ground-truth label "ground_truth_option": "C", # Correct option index "options_list": [ # List of options "Boeing737", "Other", "Airbus_A220", "Boeing747" ], "options": "A. Boeing737 B. Other C. Airbus_A220 D. Boeing747", "prompts": [ # Multiple prompt templates for the task "What type of aircraft is visible in this image?", "Which model does the identified aircraft belong to?" ], "task": "Aircraft Type Classification", # Task name "image_name": "aircraft_1.png", # Image filename "question_id": 0, # Unique ID for the question "cls_description": "High Difficulty" # Difficulty level of the sample } |
Appendix B.2.2. TSV Format
| image_path | ground_truth | Answer | Question | Index | A | B | C | D | E |
|---|---|---|---|---|---|---|---|---|---|
| 01080.jpg | 2 | A | How many bridges are visible in this SAR image? | 0 | 2 | 1 | 0 | 3 | 4 |
| 014417.jpg | 2 | E | What is the total count of bridges in the scene? | 1 | 0 | 3 | 1 | 4 | 2 |
| Field | Description |
|---|---|
| image_path | Relative path to the image (e.g., ‘ aircraft_1.png ’ maps to ‘ images/AircraftClassificationDetection/aircraft_1.png ’ ) |
| ground_truth | Task-specific ground truth (category label for classification, coordinates for detection, numerical value for reasoning tasks) |
| answer/ground_truth_option | Correct option letter (A/B/C/D/E) for multiple-choice tasks (randomized to avoid positional bias) |
| options_list | Complete list of options for multiple-choice tasks (JSON: direct list; TSV: serialized string list) |
| A/B/C/D/E | Individual option values (TSV-only, for quick access without parsing lists) |
| question/prompts | Multiple question templates to test model generalization (JSON: list; TSV: serialized string list) |
| task | Task name (consistent across JSON/TSV formats) |
| index/question_id | Unique identifier for the sample (ensures traceability) |
| cls_description | Difficulty label (High/Medium/Low) for stratified evaluation |
Appendix B.3. Evaluation Scripts
| yaml |
|---|
|
eval_backend: VLMEvalKit
eval_config: model: - type: llava-1.5-7b-hf # Model type (supports mainstream VLMs like LLaVA, Qwen-VL2) name: CustomAPIModel # Model name (for result logging) api_base: http://localhost:8000/v1/chat/completions # Local API endpoint key: EMPTY # API key (set to "EMPTY" for local deployment) temperature: 0.0 # Set to 0 for deterministic results img_size: −1 # Auto-adapt to image size data: - Image_Captioning # Target task (match task name in dataset) mode: all # Evaluate all samples in the task reuse: false # Disable result reuse work_dir: outputs # Directory to save evaluation results nproc: 1 # Number of parallel processes |
- Replace ‘api_base’ with the actual API endpoint of your deployed VLM.
- Modify the ’data’ field to switch tasks (e.g., ’Aircraft_Classification_Detection’ for the aircraft classification task).
- Remove the ‘limit’ field to run a full evaluation on the task.
References
- Chen, Y.; Cong, Y.; Zhang, L. Deformable Scattering Feature Correlation Network for Aircraft Detection in SAR Images. IEEE Geosci. Remote Sens. Lett. 2023, 20, 4007205. [Google Scholar] [CrossRef] [Scilit]
- Fan, Q.; Chen, F.; Cheng, M.; Lou, S.; Xiao, R.; Zhang, B.; Wang, C.; Li, J. Ship Detection Using a Fully Convolutional Network with Compact Polarimetric SAR Images. Remote Sens. 2019, 11, 2171. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Zhao, L.; Li, C.; Kuang, G. Pyramid Attention Dilated Network for Aircraft Detection in SAR Images. IEEE Geosci. Remote Sens. Lett. 2021, 18, 662–666. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Luo, R.; Xing, J.; Li, Z.; Yuan, Z.; Cai, X. Geospatial Transformer Is What You Need for Aircraft Detection in SAR Imagery. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5225715. [Google Scholar] [CrossRef] [Scilit]
- Chen, S.-W.; Cui, X.-C.; Wang, X.-S.; Xiao, S.-P. Speckle-Free SAR Image Ship Detection. IEEE Trans. Image Process. 2021, 30, 5969–5983. [Google Scholar] [CrossRef] [Scilit]
- Ai, J.; Xue, W.; Zhu, Y.; Zhuang, S.; Xu, C.; Yan, H.; Chen, L.; Wang, Z. AIS-PVT: Long-Time AIS Data Assisted Pyramid Vision Transformer for Sea-Land Segmentation in Dual-Polarization SAR Imagery. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5220712. [Google Scholar] [CrossRef] [Scilit]
- Mohamud, S.A.M.; Jalali, A.; Lee, M. Hierarchical Reasoning Based on Perception Action Cycle for Visual Question Answering. Expert Syst. Appl. 2024, 241, 122698. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Zhang, X.; Cheng, X.; Tang, X.; Jiao, L. Learning Consensus-Aware Semantic Knowledge for Remote Sensing Image Captioning. Pattern Recognit. 2024, 145, 109893. [Google Scholar] [CrossRef] [Scilit]
- Ke, X.; Liu, H.; Xu, P.; Lin, X.; Guo, W. Text-Based Person Search via Cross-Modal Alignment Learning. Pattern Recognit. 2024, 152, 110481. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Wen, C.; Hu, Y.; Yuan, Z.; Zhu, X.X. Vision-Language Models in Remote Sensing: Current Progress and Future Trends. IEEE Geosci. Remote Sens. Mag. 2024, 12, 32–66. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Li, X.; An, J.; Gao, L.; Hou, B.; Li, C. Natural Language Description of Remote Sensing Images Based on Deep Learning. In Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Fort Worth, TX, USA, 23–28 July 2017; IEEE: Fort Worth, TX, USA, 2017; pp. 4798–4801. [Google Scholar]
- Shi, Z.; Zou, Z. Can a Machine Generate Humanlike Language Descriptions for a Remote Sensing Image? IEEE Trans. Geosci. Remote Sens. 2017, 55, 3623–3634. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Wang, X.; Tang, X.; Zhou, H.; Li, C. Description Generation for Remote Sensing Images Using Attribute Attention Mechanism. Remote Sens. 2019, 11, 612. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Zhang, W.; Diao, W.; Yan, M.; Gao, X.; Sun, X. VAA: Visual Aligning Attention Model for Remote Sensing Image Captioning. IEEE Access 2019, 7, 137355–137364. [Google Scholar] [CrossRef] [Scilit]
- Zhan, Y.; Xiong, Z.; Yuan, Y. RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5604513. [Google Scholar] [CrossRef] [Scilit]
- Lu, X.; Wang, B.; Zheng, X.; Li, X. Exploring Models and Data for Remote Sensing Image Caption Generation. IEEE Trans. Geosci. Remote Sens. 2018, 56, 2183–2195. [Google Scholar] [CrossRef] [Scilit]
- Lobry, S.; Marcos, D.; Murray, J.; Tuia, D. RSVQA: Visual Question Answering for Remote Sensing Data. IEEE Trans. Geosci. Remote Sens. 2020, 58, 8555–8566. [Google Scholar] [CrossRef] [Scilit]
- Papadopoulos, S.; Ioannidis, K.; Vrochidis, S.; Kompatsiaris, I.; Patras, I. Vision-Language Pretraining for Variable-Shot Image Classification. In MultiMedia Modeling; Ide, I., Kompatsiaris, I., Xu, C., Yanai, K., Chu, W.-T., Nitta, N., Riegler, M., Yamasaki, T., Eds.; Springer Nature: Singapore, 2025; pp. 283–297. [Google Scholar]
- Ak, K.E.; Mohta, J.; Dimitriadis, D.; Manchanda, S.; Xu, Y.; Shen, M. Aligning Vision Language Models with Contrastive Learning. In Computer Vision—ECCV 2024 Workshops; Lecture Notes in Computer Science; Del Bue, A., Canton, C., Pont-Tuset, J., Tommasi, T., Eds.; Springer Nature: Cham, Switzerland, 2025; Volume 15640, pp. 32–45. ISBN 978-3-031-91671-7. [Google Scholar]
- Yang, J.-H.; Lin, J. Toward Automatic Relevance Judgment Using Vision–Language Models for Image–Text Retrieval Evaluation. arXiv 2024, arXiv:2408.01363. [Google Scholar]
- Li, J.; Li, D.; Savarese, S.; Hoi, S. BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models. PMLR 2023, 202, 19730–19742. [Google Scholar]
- Kuckreja, K.; Danish, M.S.; Naseer, M.; Das, A.; Khan, S.; Khan, F.S. GeoChat:Grounded Large Vision-Language Model for Remote Sensing. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 27831–27840. [Google Scholar]
- Zhang, W.; Cai, M.; Zhang, T.; Li, J.; Zhuang, Y.; Mao, X. EarthMarker: Visual Prompt Learning for Region-Level and Point-Level Remote Sensing Imagery Comprehension. arXiv 2024, arXiv:2407.13596. [Google Scholar]
- Luo, J.; Pang, Z.; Zhang, Y.; Wang, T.; Wang, L.; Dang, B.; Lao, J.; Wang, J.; Chen, J.; Tan, Y.; et al. SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding. arXiv 2024, arXiv:2406.10100. [Google Scholar]
- Soni, S.; Dudhane, A.; Debary, H.; Fiaz, M.; Munir, M.A.; Danish, M.S.; Fraccaro, P.; Watson, C.D.; Klein, L.J.; Khan, F.S.; et al. EarthDial: Turning Multi-Sensory Earth Observations to Interactive Dialogues. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 15 June 2025; pp. 14303–14313. [Google Scholar]
- Irvin, J.A.; Liu, E.R.; Chen, J.C.; Dormoy, I.; Kim, J.; Khanna, S.; Zheng, Z.; Ermon, S. TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data. arXiv 2024, arXiv:2410.06234. [Google Scholar] [CrossRef] [Scilit]
- Argenti, F.; Lapini, A.; Bianchi, T.; Alparone, L. A Tutorial on Speckle Reduction in Synthetic Aperture Radar Images. IEEE Geosci. Remote Sens. Mag. 2013, 1, 6–35. [Google Scholar] [CrossRef] [Scilit]
- Li, T. TACMT: Text-Aware Cross-Modal Transformer for Visual Grounding on High-Resolution SAR Images. ISPRS J. Photogramm. Remote Sens. 2025, 222, 152–166. [Google Scholar] [CrossRef] [Scilit]
- Zhang, W.; Cai, M.; Zhang, T.; Zhuang, Y.; Mao, X. EarthGPT: A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing Domain. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5917820. [Google Scholar] [CrossRef] [Scilit]
- Hu, Y.; Yuan, J.; Wen, C.; Lu, X.; Liu, Y.; Li, X. RSGPT: A Remote Sensing Vision Language Model and Benchmark. ISPRS J. Photogramm. Remote Sens. 2025, 224, 272–286. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Ding, J.; Elhoseiny, M. VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding. Adv. Neural Inf. Process. Syst. 2024, 37, 3229–3242. [Google Scholar]
- Danish, M.S.; Munir, M.A.; Shah, S.R.A.; Kuckreja, K.; Khan, F.S.; Fraccaro, P.; Lacoste, A.; Khan, S. GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 19–25 October 2025. [Google Scholar]
- Ma, Z.; Xiao, X.; Dong, S.; Wang, P.; Wang, H.; Pan, Q. SARChat-Bench-2M: A Multi-Task Vision-Language Benchmark for SAR Image Interpretation. arXiv 2025, arXiv:2502.08168. [Google Scholar]
- Muhtar, D.; Li, Z.; Gu, F.; Zhang, X.; Xiao, P. LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model. In European Conference on Computer Vision; Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics); 15132 LNCS; Springer: Cham, Switzerland, 2025; pp. 440–457. [Google Scholar]
- Wang, F.; Wang, H.; Chen, M.; Wang, D.; Wang, Y.; Guo, Z.; Ma, Q.; Lan, L.; Yang, W.; Zhang, J.; et al. XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery? In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025. [Google Scholar]
- An, X.; Sun, J.; Gui, Z.; He, W. CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models. arXiv 2024, arXiv:2411.18145. [Google Scholar]
- Zhang, C.; Wang, S. Good at Captioning, Bad at Counting: Benchmarking GPT-4V on Earth Observation Data. arXiv 2024, arXiv:2401.17600. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Tao, Y.; Zhang, S.; Liu, S.; Xiong, Z.; Luo, C.; Liu, L.; Pechenizkiy, M.; Zhu, X.X.; Huang, T. REOBench: Benchmarking Robustness of Earth Observation Foundation Models. arXiv 2025, arXiv:2505.16793. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Huang, Z.; Xia, R.; Wu, B.; Sheng, L.; Sun, L.; Yao, B. Large-Scale Multi-Class SAR Image Target Detection Dataset-1.0. J. Radars 2022, 14, 1488. [Google Scholar]
- Wang, Z.; Kang, Y.; Zeng, X.; Wang, Y.; Zhang, T.; Sun, X. SAR-AIRcraft-1.0: High-Resolution SAR Aircraft Detection and Recognition Dataset. J. Radars 2023, 12, 906–922. [Google Scholar] [CrossRef]
- Xia, J.; Chen, H.; Broni-Bediako, C.; Wei, Y.; Song, J.; Yokoya, N. OpenEarthMap-SAR: A Benchmark Synthetic Aperture Radar Dataset for Global High-Resolution Land Cover Mapping. arXiv 2025, arXiv:2501.10891. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Suo, Y.; Meng, Q.; Dai, W.; Miao, T.; Zhao, W.; Yan, Z.; Diao, W.; Xie, G.; Ke, Q.; et al. FAIR-CSAR: A Benchmark Dataset for Fine-Grained Object Detection and Recognition Based on Single-Look Complex SAR Images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5201022. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Weng, Q. A Deep Learning-Based Super-Resolution Method for Building Height Estimation at 2.5 m Spatial Resolution in the Northern Hemisphere. Remote Sens. Environ. 2024, 310, 114241. [Google Scholar] [CrossRef] [Scilit]
- Xu, C.; Wang, X. OpenSARWake: A Large-Scale SAR Dataset for Ship Wake Recognition with a Feature Refinement Oriented Detector. IEEE Geosci. Remote Sens. Lett. 2024, 21, 4010105. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.X.; Li, W.J.; Liu, L.; Zhou, J.; Peng, B.W.; Song, Y.F.; Xiong, X.Y.; Yang, W.; Liu, T.P.; Liu, Z.; et al. ATRNet-STAR: A Large Dataset and Benchmark towards Remote Sensing Object Recognition in the Wild. arXiv 2025, arXiv:2501.13354. [Google Scholar]
- Papineni, K.; Roukos, S.; Ward, T.; Zhu, W.-J. BLEU: A Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics—ACL ’02, Philadelphia, PA, USA, 7–12 July 2002; Association for Computational Linguistics: Philadelphia, PA, USA, 2001; p. 311. [Google Scholar]
- Lin, C.-Y. ROUGE: A Package for Automatic Evaluation of Summaries. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Barcelona, Spain, 21–26 July 2004. [Google Scholar]
- Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K.Q.; Artzi, Y. BERTScore: Evaluating Text Generation with BERT. arXiv 2019, arXiv:1904.09675. [Google Scholar]
- Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Adv. Neural Inf. Process. Syst. 2023, 36, 46595–46623. [Google Scholar]
- Callison-Burch, C.; Osborne, M.; Koehn, P. Re-Evaluating the Role of Bleu in Machine Translation Research. In Proceedings of the 11th Conference of the European Chapter of the Association for Computational Linguistics, Trento, Italy, 3–7 April 2006; McCarthy, D., Wintner, S., Eds.; Association for Computational Linguistics: Trento, Italy, 2006; pp. 249–256. [Google Scholar]
- Zeng, Z.; Sun, J.; Zhang, H.; Wen, T.; Su, Y.; Xie, Y.; Wang, Z.; Chen, B. HICEScore: A Hierarchical Metric for Image Captioning Evaluation. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, Australia, 28 October–1 November 2024; ACM: New York, NY, USA, 2024; pp. 866–875. [Google Scholar]
- Cui, Y.; Yang, G.; Veit, A.; Huang, X.; Belongie, S. Learning to Evaluate Image Captioning. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018. [Google Scholar]
- Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R.L.; Choi, Y. CLIPScore: A Reference-Free Evaluation Metric for Image Captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, 7–11 November 2021. [Google Scholar]
- Hou, X.; Ao, W.; Song, Q.; Lai, J.; Wang, H.; Xu, F. FUSAR-Ship: Building a High-Resolution SAR-AIS Matchup Dataset of Gaofen-3 for Ship Detection and Recognition. Sci. China Inf. Sci. 2020, 63, 140303. [Google Scholar] [CrossRef] [Scilit]
- Wang, D.; Zhang, F.; Ma, F.; Hu, W.; Tang, Y.; Zhou, Y. A Benchmark Sentinel-1 SAR Dataset for Airport Detection. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 6671–6686. [Google Scholar] [CrossRef] [Scilit]
- Lin, X.; Zhang, B.; Wu, F.; Wang, C.; Yang, Y.; Chen, H. SIVED: A SAR Image Dataset for Vehicle Detection Based on Rotatable Bounding Box. Remote Sens. 2023, 15, 2825. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Q.; Zhang, Y.; Li, Z.; Yan, X.; Guan, Q.; Zhong, Y.; Zhang, L.; Li, D. Oil Spill Contextual and Boundary-Supervised Detection Network Based on Marine SAR Images. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5213910. [Google Scholar] [CrossRef] [Scilit]
- Huang, M.; Xu, Y.; Qian, L.; Shi, W.; Zhang, Y.; Bao, W.; Wang, N.; Liu, X.; Xiang, X. The QXS-SAROPT Dataset for Deep Learning in SAR-Optical Data Fusion. arXiv 2021, arXiv:2103.08259. [Google Scholar]
- Wu, F.; Zhang, H.; Wang, C.; Li, L.; Li, J.J.; Chen, W.R.; Zhang, B. SARBuD1.0: A SAR Building Dataset Based on GF-3 FSII Imageries for Built-up Area Extraction with Deep Learning Method. Natl. Remote Sens. Bull. 2022, 26, 620–631. [Google Scholar] [CrossRef] [Scilit]








| Dataset | SAR | Robustness Test | Answer Type | Dimension | Annotation Method |
|---|---|---|---|---|---|
| RSIEval [30] | × | × | FF | 6 | M |
| LHRS-Bench [34] | × | × | MCQ | 11 | M |
| FIT-RSFG [24] | × | × | MCQ, BBox, FF | 11 | A + M |
| GeoChat-Bench [22] | × | × | FF, BBox | 6 | A + M |
| SARChat-Bench [33] | √ | × | BBox, FF | 6 | A + M |
| VRSBench [31] | × | × | BBox, FF | 3 | A + M |
| XLRS-Bench [35] | × | × | MCQ, BBox, FF | 16 | A + M |
| CHOICE [36] | √ | × | MCQ, BBox, Seg | 23 | A + M |
| VLEO-Bench [37] | × | × | MCQ, BBox, FF | 6 | A + M |
| GeoBench-VLM [32] | √ | × | MCQ, BBox, Seg | 31 | A + M |
| REOBench [38] | × | √ | BBox, Seg, FF | 6 | A + M |
| SAREval | √ | √ | MCQ, BBox, Seg, FF | 20 | A + M |
| Model Name | AC | ShipC | OTC | BC | TCC | VC | ACD | SCD-M | SCD-H | VCD-M | VCD-H | MCOD | SC |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Phi-3.5-vision | 22.68 | 24.03 | 23.66 | 14.59 | 16.67 | 17.19 | 27.48 | 36.62 | 25.35 | 23.40 | 20.38 | 24.42 | 44.46 |
| Qwen2.5-VL-7B | 14.43 | 20.16 | 16.79 | 40.00 | 25.56 | 17.97 | 23.87 | 47.89 | 22.54 | 18.87 | 17.74 | 34.81 | 45.09 |
| Deepseek-VL2-tiny | 14.43 | 10.85 | 9.92 | 41.62 | 20.00 | 11.72 | 17.12 | 16.20 | 19.72 | 12.45 | 11.32 | 23.48 | 20.89 |
| Mono-InternVL-2B | 28.87 | 31.78 | 21.37 | 34.59 | 26.67 | 23.44 | 27.48 | 38.03 | 21.83 | 19.62 | 20.00 | 32.83 | 50.79 |
| InternVL2-4B | 20.62 | 20.93 | 19.85 | 15.14 | 6.67 | 17.19 | 7.66 | 30.99 | 9.15 | 7.17 | 2.64 | 28.27 | 46.99 |
| InternVL2.5-4B | 29.90 | 22.48 | 27.48 | 35.14 | 22.22 | 24.22 | 27.03 | 31.69 | 19.01 | 24.91 | 19.25 | 38.55 | 58.70 |
| InternVideo2.5-8B | 20.62 | 19.38 | 23.66 | 25.41 | 6.67 | 2.34 | 12.16 | 18.31 | 6.34 | 8.30 | 2.64 | 36.57 | 44.78 |
| InternVL3-9B | 20.62 | 21.71 | 30.53 | 16.76 | 11.11 | 25.00 | 22.97 | 23.94 | 16.20 | 18.87 | 16.23 | 38.43 | 19.46 |
| LLaVA-1.5-7B | 14.43 | 20.93 | 24.43 | 16.76 | 18.89 | 18.75 | 27.48 | 39.44 | 21.83 | 23.77 | 18.87 | 43.34 | 47.78 |
| LLaVA-Next | 9.28 | 2.33 | 0.76 | 13.51 | 13.33 | 8.59 | 16.67 | 38.73 | 22.54 | 22.26 | 17.36 | 38.32 | 38.61 |
| LLaVA-onevision-qwen2-7b | 26.80 | 20.93 | 14.50 | 24.86 | 12.22 | 21.09 | 24.32 | 38.03 | 24.65 | 27.92 | 23.02 | 46.03 | 61.71 |
| Model Name | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | ROUGE-L | BERTScore | Subjective Score |
|---|---|---|---|---|---|---|---|
| Phi-3.5-vision | 0.2462 | 0.1037 | 0.0445 | 0.0244 | 0.1745 | 0.583 | 48.0739 |
| Qwen2.5-VL-7B | 0.2386 | 0.114 | 0.0523 | 0.0291 | 0.1759 | 0.5948 | 50.9989 |
| Deepseek-VL2-tiny | 0.1832 | 0.0861 | 0.0376 | 0.0206 | 0.1519 | 0.5563 | 51.393 |
| Mono-InternVL-2B | 0.0868 | 0.0429 | 0.0202 | 0.0114 | 0.0962 | 0.5062 | 34.9366 |
| InternVL2-4B | 0.1759 | 0.0844 | 0.0372 | 0.0204 | 0.1552 | 0.5775 | 47.4545 |
| InternVL2.5-4B | 0.2286 | 0.1111 | 0.0526 | 0.0297 | 0.1787 | 0.5936 | 48.7487 |
| InternVideo2.5-8B | 0.2772 | 0.1292 | 0.0582 | 0.0335 | 0.1955 | 0.5883 | 50.7744 |
| InternVL3-9B | 0.2093 | 0.101 | 0.0481 | 0.0254 | 0.1671 | 0.5947 | 58.0316 |
| LLaVA-1.5-7B | 0.2409 | 0.1179 | 0.0543 | 0.0316 | 0.1908 | 0.5555 | 54.8532 |
| LLaVA-Next | 0.187 | 0.0839 | 0.0344 | 0.0185 | 0.1589 | 0.5689 | 56.9065 |
| LLaVA-onevision-qwen2-7b | 0.1925 | 0.0905 | 0.0412 | 0.0226 | 0.1624 | 0.5829 | 46.0764 |
| Model Name | Acc@0.25 | Acc@0.5 |
|---|---|---|
| Phi-3.5-vision | 0.88 | 0.05 |
| Qwen2.5-VL-7B | 0.44 | 0.15 |
| Deepseek-VL2-tiny | 1.42 | 0.20 |
| Mono-InternVL-2B | 0.06 | 0.00 |
| InternVL2-4B | 1.72 | 0.39 |
| InternVL2.5-4B | 0.83 | 0.10 |
| InternVideo2.5-8B | 1.96 | 0.59 |
| InternVL3-9B | 2.01 | 0.44 |
| LLaVA-1.5-7B | 5.10 | 1.37 |
| LLaVA-Next | 6.08 | 2.89 |
| LLaVA-onevision-qwen2-7b | 0.64 | 0.15 |
| Model Name | IAI | TAI | POI | IRI | SMDE | BHE | TDM |
|---|---|---|---|---|---|---|---|
| Phi-3.5-vision | 7.32 | 29.56 | 18.00 | 29.52 | 26.73 | 29.56 | 20.10 |
| Qwen2.5-VL-7B | 13.66 | 30.77 | 11.50 | 20.26 | 37.62 | 0.02 | 19.66 |
| Deepseek-VL2-tiny | 18.05 | 27.35 | 31.00 | 29.07 | 26.73 | 20.59 | 31.56 |
| Mono-InternVL-2B | 19.51 | 23.93 | 16.50 | 17.62 | 16.83 | 23.39 | 23.28 |
| InternVL2-4B | 25.37 | 25.64 | 22.00 | 21.15 | 6.93 | 29.84 | 7.84 |
| InternVL2.5-4B | 13.66 | 9.40 | 18.50 | 22.47 | 30.69 | 11.05 | 0.74 |
| InternVideo2.5-8B | 39.02 | 33.33 | 18.00 | 25.55 | 33.66 | 29.56 | 22.06 |
| InternVL3-9B | 22.44 | 24.79 | 16.00 | 23.35 | 29.70 | 28.84 | 32.35 |
| LLaVA-1.5-7B | 22.44 | 33.33 | 19.00 | 25.11 | 16.83 | 23.04 | 28.26 |
| LLaVA-Next | 46.34 | 38.46 | 18.00 | 23.35 | 9.90 | 12.58 | 1.72 |
| LLaVA-onevision-qwen2-7b | 43.41 | 28.21 | 20.50 | 17.62 | 19.80 | 19.61 | 31.28 |
| Model Name | No Speckle Noise | Add Speckle Noise | ΔAccuracy |
|---|---|---|---|
| Phi-3.5-vision | 24.42 | 25.47 | 1.05 |
| Qwen2.5-VL-7B | 34.81 | 36.33 | 1.52 |
| Deepseek-VL2-tiny | 23.48 | 38.67 | 15.19 |
| Mono-InternVL-2B | 32.83 | 23.28 | −9.54 |
| InternVL2-4B | 28.27 | 35.63 | 7.36 |
| InternVL2.5-4B | 38.55 | 26.05 | −12.50 |
| InternVideo2.5-8B | 36.57 | 28.86 | −7.71 |
| InternVL3-9B | 38.43 | 35.63 | −2.80 |
| LLaVA-1.5-7B | 43.34 | 43.69 | 0.35 |
| LLaVA-Next | 38.32 | 33.64 | −4.67 |
| LLaVA-onevision-qwen2-7b | 46.03 | 43.81 | −2.22 |
| Model Name | SCD-M | SCD-H | ||||
|---|---|---|---|---|---|---|
| No Background Clutter | Add Background Clutter | ΔAccuracy | No Background Clutter | Add Background Clutter | ΔAccuracy | |
| Phi-3.5-vision | 36.62 | 35.21 | −1.41 | 25.35 | 23.94 | −1.41 |
| Qwen2.5-VL-7B | 47.89 | 19.01 | −28.87 | 22.54 | 4.93 | −17.61 |
| Deepseek-VL2-tiny | 16.20 | 23.24 | 7.04 | 19.72 | 16.90 | −2.82 |
| Mono-InternVL-2B | 38.03 | 19.01 | −19.01 | 21.83 | 22.54 | 0.70 |
| InternVL2-4B | 30.99 | 50.00 | 19.01 | 9.15 | 50.70 | 41.55 |
| InternVL2.5-4B | 31.69 | 28.87 | −2.82 | 19.01 | 11.97 | −7.04 |
| InternVideo2.5-8B | 18.31 | 38.73 | 20.42 | 6.34 | 19.01 | 12.68 |
| InternVL3-9B | 23.94 | 35.92 | 11.97 | 16.20 | 20.42 | 4.23 |
| LLaVA-1.5-7B | 39.44 | 42.25 | 2.82 | 21.83 | 23.24 | 1.41 |
| LLaVA-Next | 38.73 | 40.85 | 2.11 | 22.54 | 22.54 | 0.00 |
| LLaVA-onevision-qwen2-7b | 38.03 | 42.25 | 4.23 | 24.65 | 21.83 | −2.82 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, Z.; Liu, L.; Wan, G.; Lu, Y.; Zheng, F.; Sun, G.; Huang, Y.; Guo, S.; Li, X.; Yuan, L. SAREval: A Multi-Dimensional and Multi-Task Benchmark for Evaluating Visual Language Models on SAR Image Understanding. Remote Sens. 2026, 18, 82. https://doi.org/10.3390/rs18010082
Wang Z, Liu L, Wan G, Lu Y, Zheng F, Sun G, Huang Y, Guo S, Li X, Yuan L. SAREval: A Multi-Dimensional and Multi-Task Benchmark for Evaluating Visual Language Models on SAR Image Understanding. Remote Sensing. 2026; 18(1):82. https://doi.org/10.3390/rs18010082
Chicago/Turabian StyleWang, Ziyan, Lei Liu, Gang Wan, Yuchen Lu, Fengjie Zheng, Guangde Sun, Yixiang Huang, Shihao Guo, Xinyi Li, and Liang Yuan. 2026. "SAREval: A Multi-Dimensional and Multi-Task Benchmark for Evaluating Visual Language Models on SAR Image Understanding" Remote Sensing 18, no. 1: 82. https://doi.org/10.3390/rs18010082
APA StyleWang, Z., Liu, L., Wan, G., Lu, Y., Zheng, F., Sun, G., Huang, Y., Guo, S., Li, X., & Yuan, L. (2026). SAREval: A Multi-Dimensional and Multi-Task Benchmark for Evaluating Visual Language Models on SAR Image Understanding. Remote Sensing, 18(1), 82. https://doi.org/10.3390/rs18010082


