OATS-RS: Ontology-Aware Adaptive and Selective Zero-Shot Scene Classification for Remote Sensing
Abstract
1. Introduction
- 1.
- We define OATS-RS, a unified zero-shot inference pipeline that integrates ontology-aware prompt banks, hierarchical scoring, confuser penalties, adaptive multi-view aggregation, target-distribution alignment, transductive refinement, local ambiguity experts, candidate ensembling, and selective prediction.
- 2.
- We provide formal definitions of ontology-aware inference, frozen vision–language backbones, zero-shot recognition, coverage, selective risk, and top-K shortlist utility, making the method easier to reproduce and critique.
- 3.
- We add a reproducibility-oriented parameter and ablation protocol that specifies what must be reported for group weights, temperature parameters, negative-confuser weights, transductive refinements, comparison baselines, and runtime or throughput.
- 4.
- We analyze the supplied GeoRSCLIP ViT-B-32 EuroSAT RGB final run using overall metrics, bootstrap intervals, per-class precision/recall/F1, semantic family aggregation, risk–coverage behavior, and annual crop diagnostics.
- 5.
- We explicitly identify the limitations of the current evidence base. The attached artifacts do not include stage-by-stage ablations, additional datasets such as AID or NWPU-RESISC45, image-level qualitative examples, or hardware-specific runtime measurements, which are are specified as required future empirical additions rather than being fabricated in this revision.
2. Related Work
2.1. Vision Language Pretraining and Prompt Adaptation
2.2. Remote Sensing Foundation Models, Calibration, and Open-Vocabulary Earth Vision
3. Proposed Method, Materials, and Experimental Protocol
3.1. Method Overview and Formal Definitions
3.1.1. Zero-Shot Recognition
3.1.2. Frozen Vision–Language Model
3.1.3. Ontology-Aware Inference
3.1.4. Selective Prediction
3.2. Problem Setup and Notation
3.3. Ontology-Aware Prompt Bank
3.4. Base Zero-Shot Scoring
3.5. Hierarchical Scoring
3.6. Adaptive Multi-View Inference
3.7. Target-Distribution Feature Alignment
3.8. Balanced Transductive Refinement
3.9. Prompt Adaptation from High-Confidence Support
3.10. Ambiguity-Group Experts
3.11. Candidate Ensemble
3.12. Parameter Selection and Reproducibility Protocol
3.13. Selective Prediction
3.14. Shortlist Utility, Family Aggregation, and Ambiguity Budgets
3.15. Benchmark Materials and Experimental Design
3.16. Ablation and Fair Comparison Protocol
3.17. Statistical Analysis
4. Theoretical Analysis
4.1. Pairwise Margin Decomposition
- 1.
- Positive semantic attraction: Does the image align with the positive prompt prototype of class c more than with class d?
- 2.
- Confuser suppression: Does the image align more strongly with the negative or confuser prompts of one class than the other?
- 3.
- Relational contrast: Do pairwise prompts explicitly favor c over nearby alternatives?
- 4.
- Hierarchical consistency: Is the parent category of c more plausible than the parent category of d?
- 5.
- Local re-ranking: When the decision has narrowed to a few hard classes, do specialized experts change the ordering?
4.2. Balanced Transductive Refinement as a Regularized Latent-Variable Problem
4.3. Adaptive Multi-View Aggregation and Stability
4.4. Selective Prediction, Ambiguity Recovery, and Confidence Ranking
4.5. Class–Group Disparity and Semantic Anisotropy
4.6. Complexity and Scaling Considerations
4.7. Runtime and Throughput Reporting
5. Results
5.1. A Single Fixed Zero-Shot Evaluation Produces Moderate Top-1 Performance and Strong Top-3 Retrieval
5.2. Ablation and Comparison Evidence Required for Full Component Attribution
5.3. Performance Is Highly Class-Dependent
5.4. Qualitative Visual Results Panel
5.5. Selective Prediction Improves Reliability Only Modestly at High Coverage
5.6. The Top-1/Top-3 Gap Indicates Unresolved but Structured Ambiguity
5.7. Semantic-Family Asymmetry Is Large and Systematic
5.8. Annual-Crop Diagnostics Reveal a Specific Local Confusion Mode
5.9. Ranking Quality Is Stronger than Exact Top-1 Assignment
6. Discussion
6.1. What the Final EuroSAT Run Shows
6.2. Why the Moderate Top-1 Accuracy Is Informative but Not Sufficient
6.3. Limitations Requiring Additional Experiments
6.4. Future Directions
7. Broader Context and Forward-Looking Research Agenda
7.1. From Single-Benchmark Zero-Shot Recognition to Cross-Taxonomy Earth Observation
7.2. Trustworthy Deployment, Abstention, and Analyst Workflows
7.3. Towards Open-Vocabulary and Multimodal Earth Vision Systems
7.4. Reproducibility, Benchmarking, and Inferential Restraint
8. Conclusions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Notation and Symbol Table
| Symbol | Meaning |
|---|---|
| Set of scene classes. In the reported benchmark, . | |
| Input RGB remote sensing image patch. | |
| Ground-truth class label. | |
| Frozen image encoder of the remote sensing vision–language model. | |
| Frozen text encoder of the remote-sensing vision–language model. | |
| -normalized image embedding of x. | |
| -normalized text embedding of prompt p. | |
| Positive prompt family for class c. | |
| Negative or confuser prompt family for class c. | |
| Set of prompt groups such as literal, scene, context, geometry, signature, and pairwise groups. | |
| Aggregated positive text prototype for class c. | |
| Aggregated confuser prototype for class c. | |
| Weight of the confuser penalty. | |
| Pairwise discriminative margin between classes c and d. | |
| Weight of pairwise relational corrections. | |
| Weight of hierarchical parent gating. | |
| Set of test-time views associated with image x. | |
| Adaptive weight of view m for image x. | |
| Aggregated multi-view image embedding. | |
| Target-distribution alignment operator derived from unlabeled target covariance. | |
| Aligned image embedding after feature adaptation. | |
| Soft assignment of target image i to class c during transductive refinement. | |
| Refined class prototype for class c. | |
| Vector of class bias terms enforcing the target prior. | |
| Final class posterior assigned to class c. | |
| Maximum posterior confidence used for selective prediction. | |
| Confidence threshold for the selective classifier. | |
| Coverage of the selective classifier. | |
| Accuracy on accepted examples at threshold . | |
| Effective accuracy equal to coverage times selective accuracy. | |
| Top-K accuracy. | |
| Ambiguity recovery ratio derived from top-K accuracy. | |
| Mean class-wise F1 across a semantic family . | |
| Semantic disparity index over family-level F1 scores. |
Appendix B. Additional Derivations
Appendix B.1. Softmax Posterior and Temperature-Scaled Logits
Appendix B.2. Derivation of the Prototype Update
Appendix B.3. Selective Prediction and Effective Accuracy
Appendix B.4. Macro-F1 and Balanced Accuracy
Appendix B.5. Expected Calibration Error
Appendix B.6. Top-K Monotonicity and Shortlist Utility
Appendix B.7. Bootstrap Percentile Intervals
Appendix B.8. Family-Conditioned Disparity Indices
Appendix B.9. Set-Valued Prediction and Abstention Utility
Appendix C. Extended Result Tables
| Metric | Mean | Std. Dev. | 95% CI Low | 95% CI High |
|---|---|---|---|---|
| Accuracy | 0.5221 | 0.0106 | 0.5015 | 0.5443 |
| Balanced accuracy | 0.5219 | 0.0085 | 0.5050 | 0.5389 |
| Macro-F1 | 0.5344 | 0.0078 | 0.5196 | 0.5500 |
| Expected calibration error | 0.3843 | 0.0106 | 0.3639 | 0.4063 |
| Top-3 accuracy | 0.8870 | 0.0065 | 0.8742 | 0.8995 |
| Derived Quantity | Formula | Value |
|---|---|---|
| Top-3 gap | 0.365 | |
| Ambiguity-recovery ratio | 0.764 | |
| Selective gain at reported threshold | 0.016 | |
| Urban/agriculture F1 ratio | 2.932 | |
| Hydro/agriculture F1 ratio | 1.909 |
Appendix D. Algorithmic Summary
| Step | Operation |
|---|---|
| 1 | For each class, construct ontology-aware prompt families including literal prompts, scene descriptions, contextual prompts, signature prompts, and pairwise contrastive prompts. |
| 2 | Encode all prompts with the frozen text encoder and aggregate them into positive and negative class prototypes. |
| 3 | For each image, generate multiple test-time views, encode them with the frozen image encoder, and compute adaptive view weights from confidence, margin, entropy, and agreement. |
| 4 | Aggregate the view embeddings into a single image representation and apply target-distribution feature alignment. |
| 5 | Compute base class scores from positive similarity minus confuser penalty, then add pairwise relational and hierarchical corrections. |
| 6 | Perform balanced transductive refinement on the unlabeled target set by iterating soft assignments, prototype updates, and class-bias corrections. |
| 7 | Apply local ambiguity-group experts, graph smoothing, density re-ranking, and candidate-ensemble fusion where available. |
| 8 | Produce final class posteriors, top-K rankings, and the selective decision rule obtained by thresholding confidence. |
| 9 | Evaluate accuracy, balanced accuracy, macro-F1, calibration, top-K performance, and risk–coverage behavior on the labeled test set. |
References
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations, Virtual, 3–7 May 2021. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models From Natural Language Supervision. arXiv 2021, arXiv:2103.00020. [Google Scholar] [CrossRef] [Scilit]
- Zhai, X.; Wang, X.; Mustafa, B.; Steiner, A.; Keysers, D.; Kolesnikov, A.; Beyer, L. LiT: Zero-Shot Transfer With Locked-Image Text Tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18123–18133. [Google Scholar] [CrossRef] [Scilit]
- Cherti, M.; Beaumont, R.; Wightman, R.; Wortsman, M.; Ilharco, G.; Gordon, C.; Schuhmann, C.; Schmidt, L.; Jitsev, J. Reproducible Scaling Laws for Contrastive Language-Image Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 2818–2829. [Google Scholar] [CrossRef] [Scilit]
- Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. DINOv2: Learning Robust Visual Features without Supervision. arXiv 2023, arXiv:2304.07193. [Google Scholar] [CrossRef] [Scilit]
- Cong, Y.; Khanna, S.; Meng, C.; Liu, P.; Rozi, E.; He, Y.; Burke, M.; Lobell, D.; Ermon, S. SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery. Adv. Neural Inf. Process. Syst. 2022, 35, 197–211. [Google Scholar] [CrossRef] [Scilit]
- Gupta, R.; Reed, C.; Li, S.; Brockman, S.; Funk, C.; Clipp, B.; Keutzer, K.; Candido, S.; Uyttendaele, M.; Darrell, T. Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023. [Google Scholar] [CrossRef] [Scilit]
- Fuller, A.; Millard, K.; Green, J.R. CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders. Adv. Neural Inf. Process. Syst. 2023, 36, 5506–5538. [Google Scholar] [CrossRef] [Scilit]
- Helber, P.; Bischke, B.; Dengel, A.; Borth, D. EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2019, 12, 2217–2226. [Google Scholar] [CrossRef] [Scilit]
- Xia, G.S.; Hu, J.; Hu, F.; Shi, B.; Bai, X.; Zhong, Y.; Zhang, L.; Lu, X. AID: A Benchmark Data Set for Performance Evaluation of Aerial Scene Classification. IEEE Trans. Geosci. Remote Sens. 2017, 55, 3965–3981. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Han, J.; Lu, X. Remote Sensing Image Scene Classification: Benchmark and State of the Art. Proc. IEEE 2017, 105, 1865–1883. [Google Scholar] [CrossRef] [Scilit]
- Sumbul, G.; Charfuelan, M.; Demir, B.; Markl, V. BigEarthNet: A Large-Scale Benchmark Archive for Remote Sensing Image Understanding. In Proceedings of the IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Yokohama, Japan, 28 July–2 August 2019; pp. 5901–5904. [Google Scholar] [CrossRef] [Scilit]
- Schmitt, M.; Hughes, L.H.; Qiu, C.; Zhu, X.X. SEN12MS – A Curated Dataset of Georeferenced Multi-Spectral Sentinel-1/2 Imagery for Deep Learning and Data Fusion. Isprs Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2019, IV-2/W7, 153–160. [Google Scholar] [CrossRef] [Scilit]
- Zhou, K.; Yang, J.; Loy, C.C.; Liu, Z. Learning to Prompt for Vision-Language Models. Int. J. Comput. Vis. 2022, 130, 2337–2348. [Google Scholar] [CrossRef] [Scilit]
- Zhou, K.; Yang, J.; Loy, C.C.; Liu, Z. Conditional Prompt Learning for Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16816–16825. [Google Scholar] [CrossRef] [Scilit]
- Khattak, M.U.; Rasheed, H.; Maaz, M.; Khan, S.; Khan, F.S. MaPLe: Multi-Modal Prompt Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 19113–19122. [Google Scholar] [CrossRef] [Scilit]
- Gu, Y.; Wang, Y.; Li, Y. A Survey on Deep Learning-Driven Remote Sensing Image Scene Understanding: Scene Classification, Scene Retrieval and Scene-Guided Object Detection. Appl. Sci. 2019, 9, 2110. [Google Scholar] [CrossRef] [Scilit]
- Thapa, A.; Horanont, T.; Neupane, B.; Aryal, J. Deep Learning for Remote Sensing Image Scene Classification: A Review and Meta-Analysis. Remote Sens. 2023, 15, 4804. [Google Scholar] [CrossRef] [Scilit]
- Zhou, W.; Newsam, S.; Li, C.; Shao, Z. PatternNet: A Benchmark Dataset for Performance Evaluation of Remote Sensing Image Retrieval. ISPRS J. Photogramm. Remote Sens. 2018, 145, 197–209. [Google Scholar] [CrossRef] [Scilit]
- Zhu, X.X.; Hu, J.; Qiu, C.; Shi, Y.; Kang, J.; Mou, L.; Bagheri, H.; Häberle, M.; Hua, Y.; Huang, R.; et al. So2Sat LCZ42: A Benchmark Dataset for Global Local Climate Zones Classification. arXiv 2019, arXiv:1912.12171. [Google Scholar] [CrossRef] [Scilit]
- Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.; Parekh, Z.; Pham, H.; Le, Q.V.; Sung, Y.; Li, Z.; Duerig, T. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. In Proceedings of the 38th International Conference on Machine Learning, Virtual, 18–24 July 2021; Volume 139, pp. 4904–4916. [Google Scholar]
- Khattak, M.U.; Wasim, S.T.; Naseer, M.; Khan, S.; Yang, M.H.; Khan, F.S. Self-Regulating Prompts: Foundational Model Adaptation without Forgetting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 15190–15200. [Google Scholar]
- Zhang, R.; Zhang, W.; Fang, R.; Gao, P.; Li, K.; Dai, J.; Qiao, Y.; Li, H. Tip-Adapter: Training-Free Adaption of CLIP for Few-Shot Classification. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 493–510. [Google Scholar]
- Shu, M.; Nie, W.; Huang, D.A.; Yu, Z.; Goldstein, T.; Anandkumar, A.; Xiao, C. Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language Models. In Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; Volume 35. [Google Scholar] [CrossRef] [Scilit]
- Gao, P.; Geng, S.; Zhang, R.; Ma, T.; Fang, R.; Zhang, Y.; Li, H.; Qiao, Y. CLIP-Adapter: Better Vision-Language Models with Feature Adapters. arXiv 2021, arXiv:2110.04544. [Google Scholar] [CrossRef] [Scilit]
- Jia, M.; Tang, L.; Chen, B.C.; Cardie, C.; Belongie, S.; Hariharan, B.; Lim, S.N. Visual Prompt Tuning. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 709–727. [Google Scholar] [CrossRef] [Scilit]
- Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. LAION-5B: An Open Large-Scale Dataset for Training Next Generation Image-Text Models. Adv. Neural Inf. Process. Syst. 2022, 35, 25278–25294. [Google Scholar] [CrossRef] [Scilit]
- Zhai, X.; Mustafa, B.; Kolesnikov, A.; Beyer, L. Sigmoid Loss for Language Image Pre-Training. arXiv 2023, arXiv:2303.15343. [Google Scholar] [CrossRef] [Scilit]
- Drusch, M.; Del Bello, U.; Carlier, S.; Colin, O.; Fernandez, V.; Gascon, F.; Hoersch, B.; Isola, C.; Laberinti, P.; Martimort, P.; et al. Sentinel-2: ESA’s Optical High-Resolution Mission for GMES Operational Services. Remote Sens. Environ. 2012, 120, 25–36. [Google Scholar] [CrossRef] [Scilit]
- Sumbul, G.; de Wall, A.; Kreuziger, T.; Marcelino, F.; Costa, H.; Benevides, P.; Caetano, M.; Demir, B.; Markl, V. BigEarthNet-MM: A Large Scale Multi-Modal Multi-Label Benchmark Archive for Remote Sensing Image Classification and Retrieval. arXiv 2021, arXiv:2105.07921. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Zheng, Z.; Ma, A.; Lu, X.; Zhong, Y. LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation. arXiv 2021, arXiv:2110.08733. [Google Scholar] [CrossRef] [Scilit]
- Brown, C.F.; Brumby, S.P.; Guzder-Williams, B.; Birch, T.; Hyde, S.B.; Mazzariello, J.; Czerwinski, W.; Pasquarella, V.J.; Haertel, R.; Ilyushchenko, S.; et al. Dynamic World, Near Real-Time Global 10 m Land Use Land Cover Mapping. Sci. Data 2022, 9, 251. [Google Scholar] [CrossRef] [Scilit]
- Zanaga, D.; Van De Kerchove, R.; Daems, D.; De Keersmaecker, W.; Brockmann, C.; Kirches, G.; Wevers, J.; Cartus, O.; Santoro, M.; Fritz, S.; et al. ESA WorldCover 10 m 2021 v200. Dataset Release. 2022. Available online: https://zenodo.org/records/7254221 (accessed on 28 May 2026).
- Liu, F.; Chen, D.; Guan, Z.; Zhou, X.; Zhu, J.; Ye, Q.; Fu, L.; Zhou, J. RemoteCLIP: A Vision Language Foundation Model for Remote Sensing. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5622216. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Zhao, T.; Guo, Y.; Yin, J. RS5M and GeoRSCLIP: A Large-Scale Vision-Language Dataset and a Large Vision-Language Model for Remote Sensing. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5642123. [Google Scholar] [CrossRef] [Scilit]
- Lacoste, A.; Lehmann, N.; Rodriguez, O.; Sherwin, E.; Kerner, H.; Lutjens, B.; Zhu, X. GEO-Bench: Toward Foundation Models for Earth Monitoring. arXiv 2023, arXiv:2306.03831. [Google Scholar] [CrossRef] [Scilit]
- Sun, X.; Wang, P.; Lu, W.; Zhu, Z.; Lu, X.; He, Q.; Li, J.; Rong, X.; Yang, Z.; Chang, H.; et al. RingMo: A Remote Sensing Foundation Model with Masked Image Modeling. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5612822. [Google Scholar] [CrossRef] [Scilit]
- Bastani, F.; Wolters, P.; Gupta, R.; Ferdinando, J.; Kembhavi, A. SatlasPretrain: A Large-Scale Dataset for Remote Sensing Image Understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 16772–16782. [Google Scholar]
- Nedungadi, V.; Kariryaa, A.; Oehmcke, S.; Belongie, S.; Igel, C.; Lang, N. MMEarth: Exploring Multi-Modal Pretext Tasks for Geospatial Representation Learning. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Xiong, Z.; Wang, Y.; Zhang, F.; Stewart, A.J.; Hanna, J.; Borth, D.; Papoutsis, I.; Le Saux, B.; Camps-Valls, G.; Zhu, X.X. Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation. arXiv 2024, arXiv:2403.15356. [Google Scholar] [CrossRef] [Scilit]
- Astruc, G.; Gonthier, N.; Mallet, C.; Landrieu, L. AnySat: An Earth Observation Model for Any Resolutions, Scales, and Modalities. arXiv 2024, arXiv:2412.14123. [Google Scholar] [CrossRef] [Scilit]
- Szwarcman, D.; Roy, S.; Fraccaro, P.; Gislason, O.E.; Blumenstiel, B.; Ghosal, R.; Moreno, J.B. Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications. arXiv 2024, arXiv:2412.02732. [Google Scholar] [CrossRef] [Scilit]
- Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On Calibration of Modern Neural Networks. In Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; Volume 70, pp. 1321–1330. [Google Scholar]
- Minderer, M.; Djolonga, J.; Romijnders, R.; Hubis, F.; Zhai, X.; Houlsby, N.; Tran, D.; Lucic, M. Revisiting the Calibration of Modern Neural Networks. In Proceedings of the Advances in Neural Information Processing Systems, Online, 6–14 December 2021; Volume 34. [Google Scholar]
- Lakshminarayanan, B.; Pritzel, A.; Blundell, C. Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]
- Hendrycks, D.; Gimpel, K. A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks. In Proceedings of the International Conference on Learning Representations, Toulon, France, 24–26 April 2017. [Google Scholar]
- Geifman, Y.; El-Yaniv, R. Selective Classification for Deep Neural Networks. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]
- Geifman, Y.; El-Yaniv, R. SelectiveNet: A Deep Neural Network with an Integrated Reject Option. In Proceedings of the 36th International Conference on Machine Learning, Long Beach, CA, USA, 9-15 June 2019; Volume 97, pp. 2151–2159. [Google Scholar]
- Dempster, A.P.; Laird, N.M.; Rubin, D.B. Maximum Likelihood from Incomplete Data via the EM Algorithm. J. R. Stat. Soc. Ser. B 1977, 39, 1–38. [Google Scholar] [CrossRef] [Scilit]
- Grandvalet, Y.; Bengio, Y. Semi-Supervised Learning by Entropy Minimization. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, Canada, 13–18 December 2004; Volume 17. [Google Scholar]
- Zhou, D.; Bousquet, O.; Lal, T.N.; Weston, J.; Scholkopf, B. Learning with Local and Global Consistency. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, Canada, 13–18 December 2004; Volume 16, pp. 321–328. [Google Scholar]
- Wang, D.; Shelhamer, E.; Liu, S.; Olshausen, B.; Darrell, T. Tent: Fully Test-Time Adaptation by Entropy Minimization. In Proceedings of the International Conference on Learning Representations, Virtual Event, 3–7 May 2021. [Google Scholar]
- Zhang, M.; Levine, S.; Finn, C. MEMO: Test Time Robustness via Adaptation and Augmentation. Adv. Neural Inf. Process. Syst. 2022, 35, 38629–38642. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Fink, O.; Van Gool, L.; Dai, D. Continual Test-Time Domain Adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 7191–7201. [Google Scholar] [CrossRef] [Scilit]
- Niu, S.; Wu, J.; Zhang, Y.; Chen, Y.; Zheng, S.; Zhao, P.; Tan, M. Efficient Test-Time Model Adaptation without Forgetting. In Proceedings of the 39th International Conference on Machine Learning, Baltimore, MA, USA, 17–23 July 2022; Volume 62, pp. 16888–16905. [Google Scholar]
- Platt, J.C. Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods. Adv. Large Margin Classif. 1999, 10, 61–74. [Google Scholar]
- Lin, H.T.; Lin, C.J.; Weng, R.C. A Note on Platt’s Probabilistic Outputs for Support Vector Machines. Mach. Learn. 2007, 68, 267–276. [Google Scholar] [CrossRef] [Scilit]
- Niculescu-Mizil, A.; Caruana, R. Predicting Good Probabilities with Supervised Learning. In Proceedings of the 22nd International Conference on Machine Learning, Bonn, Germany, 7–11 August 2005; pp. 625–632. [Google Scholar] [CrossRef] [Scilit]
- Gal, Y.; Ghahramani, Z. Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. In Proceedings of the 33rd International Conference on Machine Learning, New York City, NY, USA, 19–24 June 2016; Volume 48, pp. 1050–1059. [Google Scholar]
- Kumar, A.; Liang, P.S.; Ma, T. Verified Uncertainty Calibration. Adv. Neural Inf. Process. Syst. 2019, 32. [Google Scholar]
- Chow, C.K. On Optimum Recognition Error and Reject Tradeoff. IEEE Trans. Inf. Theory 1970, 16, 41–46. [Google Scholar] [CrossRef] [Scilit]
- Sadinle, M.; Lei, J.; Wasserman, L. Least Ambiguous Set-Valued Classifiers with Bounded Error Levels. J. Am. Stat. Assoc. 2019, 114, 223–234. [Google Scholar] [CrossRef] [Scilit]
- Angelopoulos, A.N.; Bates, S. A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification. arXiv 2021, arXiv:2107.07511. [Google Scholar] [CrossRef] [Scilit]
- Cao, Q.; Chen, Y.; Ma, C.; Yang, X. Open-Vocabulary Remote Sensing Image Semantic Segmentation. arXiv 2024, arXiv:2409.07683. [Google Scholar] [CrossRef] [Scilit]
- Li, K.; Liu, R.; Cao, X.; Meng, D.; Wang, Z. SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images. arXiv 2024, arXiv:2410.01768. [Google Scholar] [CrossRef] [Scilit]
- Dutta, S.; Vasim, A.; Gole, S.; Rezatofighi, H.; Banerjee, B. AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Nashville, TN, USA, 11–15 June 2025; pp. 2279–2289. [Google Scholar]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.Y.; et al. Segment Anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 4015–4026. [Google Scholar] [CrossRef] [Scilit]
- Ledoit, O.; Wolf, M. A Well-Conditioned Estimator for Large-Dimensional Covariance Matrices. J. Multivar. Anal. 2004, 88, 365–411. [Google Scholar] [CrossRef] [Scilit]
- Efron, B. Bootstrap Methods: Another Look at the Jackknife. Ann. Stat. 1979, 7, 1–26. [Google Scholar] [CrossRef] [Scilit]









| Parameter Family | Setting Rule in a Strict Zero-Shot Run | Purpose and Failure Mode Controlled |
|---|---|---|
| Prompt-group weights | Fix before evaluation; normalize over groups; report literal, scene, definition, context, geometry, signature, contrastive, pairwise, and parent weights separately. | Controls how much the classifier trusts class names, semantic descriptions, geometric cues, and confuser prompts. Poor reporting means that prompt gains cannot be reproduced. |
| Prompt aggregation temperature | Use one value per group or a single global value; do not tune on test labels; report whether aggregation is mean-like or max-like. | Controls whether many prompts contribute smoothly or whether one high-scoring prompt dominates. |
| Negative-confuser weight | Increase only if label-free diagnostics show improved margins without severe class-prior collapse; report the final scalar. | Penalizes classes with confuser prompts that are highly compatible with the image. Excessive values can suppress genuinely ambiguous classes. |
| Hierarchy weight | Use a small fixed value when parent categories are reliable; set to zero if parent ontology is uncertain. | Prevents fine-grained labels from outranking implausible coarse families. |
| View temperature and weights | Fix across datasets or select by view-agreement stability; report number and type of views. | Controls adaptive multi-view fusion; unstable view weighting can amplify augmentations that accidentally raise confidence. |
| Alignment parameters | Report shrinkage strength, power exponent, and blending coefficient; use conservative shrinkage when target set is small. | Controls target–covariance alignment while limiting numerical instability and over-whitening. |
| Transductive parameters | Report assignment temperature, confidence exponent, prototype blend, and prior-correction step size; monitor class-prior entropy. | Controls unlabeled prototype refinement; aggressive settings can collapse predictions into easy classes. |
| Selective threshold | Report both target coverage and actual coverage; never report selective accuracy without coverage. | Determines abstention behavior and analyst workload. |
| Dataset | Task Type | Why It Matters for Transferability |
|---|---|---|
| EuroSAT RGB [9] | Ten-class Sentinel-2 scene classification | Current measured benchmark; tests compact land use and land cover semantics. |
| AID [10] | High-resolution aerial scene classification | Tests transfer to aerial imagery with different scale, appearance, and urban-object detail. |
| NWPU-RESISC45 [11] | Remote sensing scene classification with 45 classes | Tests a larger and more fine-grained scene taxonomy with many visually adjacent classes. |
| BigEarthNet/ BigEarthNet-MM [12,30] | Multi-label land cover tagging | Tests whether ontology-aware prompts extend beyond single-label scene recognition. |
| SEN12MS [13] | Multimodal Sentinel-1/2 representation learning | Tests whether the method can be adapted to multimodal optical-SAR evidence. |
| Variant | Components Enabled | Metrics to Report |
|---|---|---|
| GeoRSCLIP baseline | Single or standard prompt; frozen backbone; no ontology, no refinement. | Top-1, top-3, ECE |
| + ontology prompts | Prompt groups for name, definition, context, geometry, and signature. | Same plus class F1 |
| + contrastive/confuser scoring | Adds negative prompts and pairwise confuser margins. | Same plus confuser errors |
| + adaptive multi-view | Adds view scoring and weighted view aggregation. | Same plus throughput |
| + target alignment | Adds covariance-shrinkage feature alignment on unlabeled target embeddings. | Same plus stability |
| + balanced transductive refinement | Adds soft assignments, prototype updates, and class-prior correction. | Same plus prior entropy |
| + prompt support adaptation | Adds high-confidence support centroids and drift control. | Same plus support size |
| + ambiguity experts | Adds local re-ranking for vegetation, water, or other confuser groups. | Same plus local errors |
| Full OATS-RS | Adds candidate ensembling and selective prediction. | Same plus coverage/risk |
| Method Family | Fair-Use Requirement | Notes for Interpretation |
|---|---|---|
| CLIP/OpenCLIP [2,4] | Same RGB preprocessing and same prompt set when possible. | Tests general VLM transfer without remote-sensing adaptation. |
| RemoteCLIP [34] | Same split and image resolution; no target-label tuning. | Tests remote sensing VLM pretraining under matched inference rules. |
| GeoRSCLIP/RS5M [35] | Same checkpoint, preprocessing, and class names. | Natural backbone baseline for the current method. |
| TPT/test-time prompt tuning [24] | Report unlabeled-batch assumption and optimization steps. | Useful but may differ from frozen-prompt inference. |
| Tip-Adapter/CLIP-Adapter [23,25] | Clearly state whether labeled support examples are used. | Few-shot variants are not strict zero-shot baselines. |
| SAM or SAM-assisted open-vocabulary pipelines [64,65,66,67] | Convert masks to image-level labels or report dense metrics separately. | Relevant for visual grounding, but not a direct scene-classification comparator without an aggregation rule. |
| Full OATS-RS | Same backbone and target split as baselines; no target labels. | Reports top-1, top-3, calibration, coverage, and class-wise diagnostics. |
| Metric | What to Report |
|---|---|
| Image encoder throughput | Images per second for one view and for the full multi-view setting. |
| Text encoder cost | Number of prompts, total prompt encoding time, and whether text embeddings are cached. |
| End-to-end throughput | Images per second from image loading through final selective prediction. |
| Transductive overhead | Target-batch size, number of refinement iterations, nearest-neighbor cost, and additional seconds per batch. |
| Memory footprint | Peak GPU memory or CPU memory for feature storage, covariance estimation, and neighbor search. |
| Hardware and software | GPU/CPU model, precision, batch size, framework version, and backbone checkpoint. |
| Item | Setting |
|---|---|
| Benchmark | EuroSAT RGB |
| Label space | Ten scene classes |
| Model family | GeoRSCLIP |
| Backbone architecture | ViT-B-32 |
| Training on target labels | None (zero-shot) |
| Inference mode | Prompt-based zero-shot with adaptive selective inference |
| Input modality | RGB remote-sensing scene patches |
| Primary summary metrics | Accuracy, balanced accuracy, macro-F1, micro-F1, top-3 accuracy, expected calibration error, coverage, selective accuracy |
| Empirical basis used in this paper | Final attached performance artifacts only |
| Metric | Value | 95% Bootstrap CI |
|---|---|---|
| Accuracy | 0.522 | 0.501–0.544 |
| Balanced accuracy | 0.522 | 0.505–0.539 |
| Macro-F1 | 0.535 | 0.520–0.550 |
| Micro-F1 | 0.522 | – |
| Expected calibration error | 0.384 | 0.364–0.406 |
| Top-3 accuracy | 0.887 | 0.874–0.899 |
| Coverage | 0.934 | – |
| Selective accuracy | 0.538 | – |
| Selective Macro-F1 | 0.548 | – |
| Effective accuracy | 0.502 | – |
| Class | Precision | Recall | F1 |
|---|---|---|---|
| Industrial Area | 0.924 | 0.910 | 0.917 |
| Residential Area | 0.876 | 0.955 | 0.914 |
| River | 0.966 | 0.845 | 0.901 |
| Highway | 0.922 | 0.770 | 0.839 |
| Pasture | 0.825 | 0.800 | 0.812 |
| Annual Crop | 0.267 | 0.250 | 0.258 |
| Sea or Lake | 0.347 | 0.205 | 0.258 |
| Forest | 0.197 | 0.265 | 0.226 |
| Permanent Crop | 0.206 | 0.195 | 0.201 |
| Herbaceous Vegetation | 0.017 | 0.025 | 0.021 |
| Target Coverage | Actual Coverage | Risk | Selective Accuracy | Threshold |
|---|---|---|---|---|
| 0.250 | 0.250 | 0.470 | 0.530 | 0.147 |
| 0.500 | 0.500 | 0.406 | 0.594 | 0.138 |
| 0.750 | 0.750 | 0.437 | 0.563 | 0.130 |
| 0.900 | 0.900 | 0.453 | 0.547 | 0.124 |
| 0.934 | 0.934 | 0.462 | 0.538 | 0.122 |
| 0.950 | 0.950 | 0.466 | 0.534 | 0.120 |
| Semantic Family | Mean Precision | Mean Recall | Mean F1 |
|---|---|---|---|
| Agriculture/vegetation | 0.303 | 0.307 | 0.304 |
| Urban/transport | 0.907 | 0.878 | 0.890 |
| Hydrographic | 0.657 | 0.525 | 0.580 |
| Quantity | Value | Notes |
|---|---|---|
| Subset accuracy | 0.250 | Matches Annual Crop recall |
| Prediction share: Permanent Crop | 0.650 | Dominant confusion mode |
| Prediction share: Sea or Lake | 0.055 | Secondary error mode |
| Prediction share: Herbaceous Vegetation | 0.025 | Rare error mode |
| Prediction share: Pasture | 0.020 | Rare error mode |
| Mean confidence (correct) | 0.143 | Std. dev. 0.008 |
| Mean confidence (incorrect) | 0.131 | Std. dev. 0.006 |
| Mean margin (correct) | 0.010 | Top-1 minus top-2 |
| Mean margin (incorrect) | 0.012 | Top-1 minus top-2 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Horváth, J. OATS-RS: Ontology-Aware Adaptive and Selective Zero-Shot Scene Classification for Remote Sensing. Remote Sens. 2026, 18, 2038. https://doi.org/10.3390/rs18122038
Horváth J. OATS-RS: Ontology-Aware Adaptive and Selective Zero-Shot Scene Classification for Remote Sensing. Remote Sensing. 2026; 18(12):2038. https://doi.org/10.3390/rs18122038
Chicago/Turabian StyleHorváth, János. 2026. "OATS-RS: Ontology-Aware Adaptive and Selective Zero-Shot Scene Classification for Remote Sensing" Remote Sensing 18, no. 12: 2038. https://doi.org/10.3390/rs18122038
APA StyleHorváth, J. (2026). OATS-RS: Ontology-Aware Adaptive and Selective Zero-Shot Scene Classification for Remote Sensing. Remote Sensing, 18(12), 2038. https://doi.org/10.3390/rs18122038

