Parameter-Efficient Adaptation of Generative-Foundation (Flux, Qwen) vs. Zero-Shot (Gemini, SAM3) Models for Aerial Image Segmentation
Abstract
1. Introduction
2. Related Work
2.1. Conventional and Machine Learning Approaches for Roof Segmentation
2.2. Foundation and Editing Models in Computer Vision
2.3. Parameter-Efficient Fine-Tuning Techniques
3. Methodology
3.1. Dataset Development
3.1.1. Primary Aerial-Imagery Dataset
3.1.2. Ground-Truth Dataset Generation
3.1.3. Mask Polarity and Preprocessing Considerations
3.2. Editing-Based Models and LoRA Adaptation
3.2.1. Zero-Shot Editing-Based Models
Foundation Models
- FLUX.1 Kontext: A rectified flow transformer model known for high semantic adherence, serving as the primary baseline for stable LoRA adaptation (accessed 18 December 2025) [29].
- FLUX.2: A newer iteration of the FLUX family, selected to investigate whether aggressive optimisation for aesthetic generation compromises geometric stability in discriminative tasks (accessed 17 December 2025) [37].
- Qwen Image Edit 2509: A diffusion-based image-editing model included as a representative editing-based baseline within the comparative evaluation framework (accessed 17 December 2025) [31].
- Gemini 3 Pro image preview: A multimodal large language model (LMM) utilised as a zero-shot baseline to benchmark the performance of fine-tuned adapters against state-of-the-art promptable systems (accessed 31 December 2025).
Segmentation-First Baseline
3.2.2. LoRA Fine-Tuning Configuration
3.3. Validation Metrics and Evaluation Protocol
4. Results
4.1. Zero-Shot Performance of Editing-Based Models
4.1.1. Foundation Models’ Performance
4.1.2. Segmentation-First Baseline Performance (SAM3)
4.2. Effect of LoRA Fine-Tuning on Editing-Based Models
4.2.1. Performance of FLUX Kontext LoRA Pipeline
4.2.2. Performance of FLUX.2 Pipeline
4.2.3. Performance of Qwen Image Edit 2509 Pipeline
4.3. Cross-Model Performance Comparison
4.3.1. The Impact of Mask Polarity on Metric Reliability
4.3.2. Quantitative Cross-Model Comparison
4.3.3. Qualitative Visual Comparison
5. Discussion
5.1. The Zero-Shot Editing Models
5.2. Analysis of Mask Polarity and Metric Sensitivity
5.3. SAM3 as a Segmentation-First Baseline
5.4. Impact of LoRA Fine-Tuning
5.5. Cross-Model Trade-Offs
5.6. Analysis of Research Uncertainty and Reliability
6. Conclusions and Future Work
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Saadaoui, H.; Farah, S.; Lechgar, H.; Ghennioui, A.; Rhinane, H. Advancing Urban Roof Segmentation: Transformative Deep Learning Models from CNNs to Transformers for Scalable and Accurate Urban Imaging Solutions—A Case Study in Ben Guerir City, Morocco. Technologies 2025, 13, 452. [Google Scholar] [CrossRef]
- Kim, J.; Bae, H.; Kang, H.; Lee, S.G. CNN Algorithm for Roof Detection and Material Classification in Satellite Images. Electronics 2021, 10, 1592. [Google Scholar] [CrossRef]
- Cai, Y.; He, H.; Yang, K.; Fatholahi, S.N.; Ma, L.; Xu, L.; Li, J. A Comparative Study of Deep Learning Approaches to Rooftop Detection in Aerial Images. Can. J. Remote Sens. 2021, 47, 413–431. [Google Scholar] [CrossRef]
- Awais, M.; Naseer, M.; Khan, S.; Anwer, R.M.; Cholakkal, H.; Shah, M.; Yang, M.H.; Khan, F.S. Foundation Models Defining a New Era in Vision: A Survey and Outlook. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 2245–2264. [Google Scholar] [CrossRef] [PubMed]
- Asif, M.; Sharieff, R.; Olawale, M.; Khan, M.I. Unlocking the potential of unregulated rooftops for solar PV on residential buildings: Identifying and addressing key challenges. Energy Nexus 2025, 18, 100447. [Google Scholar] [CrossRef]
- Hussain, Z.k.; Congshi, J.; Adrees, M.; Chaudhary, H.; Shafqat, R. A Novel Architecture for building rooftop extraction using remote sensing and deep learning. Remote Sens. Appl. Soc. Environ. 2025, 38, 101551. [Google Scholar] [CrossRef]
- Li, Y.; Wu, Y.; Lai, Y.; Hu, M.; Yang, X. MedDINOv3: How to adapt vision foundation models for medical image segmentation? arXiv 2025, arXiv:2509.02379. [Google Scholar] [CrossRef]
- Pei, J.; Zhou, Z.; Zhang, T. Evaluation study on SAM 2 for class-agnostic instance-level segmentation. CAAI Artif. Intell. Res. 2025, 4, 9150055. [Google Scholar] [CrossRef]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. Lora: Low-rank adaptation of large language models. ICLR 2022, 1, 3. [Google Scholar]
- Yang, M.; Chen, J.; Zhang, Y.; Liu, J.; Zhang, J.; Ma, Q.; Verma, H.; Zhang, Q.; Zhou, M.; King, I. Low-rank adaptation for foundation models: A comprehensive review. arXiv 2025, arXiv:2501.00365. [Google Scholar]
- Shata, D.; Omrani, S.; Drogemuller, R.; Denman, S.; Wagdy, A. Unlocking solar potential through machine learning techniques for roof geometry prediction—A review. Sol. Energy 2025, 302, 113994. [Google Scholar] [CrossRef]
- Song, S.; Tang, Y.; Qin, R. Synthetic Data Matters: Retraining With Geo-Typical Synthetic Labels for Building Detection. IEEE Trans. Geosci. Remote Sens. 2025, 63, 1–13. [Google Scholar] [CrossRef]
- Neupane, B.; Aryal, J.; Rajabifard, A. Building Footprint Segmentation Using Transfer Learning: A Case Study of the City of Melbourne. ISPRS Ann. Photogramm. Remote Sens. Spatial Inf. Sci. 2022, X-4/W3-2022, 173–179. [Google Scholar] [CrossRef]
- Demir, I.; Koperski, K.; Lindenbaum, D.; Pang, G.; Huang, J.; Basu, S.; Hughes, F.; Tuia, D.; Raskar, R. Deepglobe 2018: A challenge to parse the earth through satellite images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA, 18–22 June 2018. [Google Scholar]
- Maggiori, E.; Tarabalka, Y.; Charpiat, G.; Alliez, P. Can semantic labeling methods generalize to any city? The inria aerial image labeling benchmark. In Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Fort Worth, TX, USA, 23–28 July 2017. [Google Scholar]
- Van Etten, A.; Lindenbaum, D.; Bacastow, T.M. Spacenet: A remote sensing dataset and challenge series. arXiv 2018, arXiv:1807.01232. [Google Scholar]
- Rottensteiner, F.; Sohn, G.; Gerke, M.; Wegner, J.D.; Breitkopf, U.; Jung, J. Results of the ISPRS benchmark on urban object detection and 3D building reconstruction. ISPRS J. Photogramm. Remote Sens. 2014, 93, 256–271. [Google Scholar] [CrossRef]
- Chen, T.; Zhu, L.; Ding, C.; Cao, R.; Zhang, S.; Wang, Y.; Li, Z.; Sun, L.; Mao, P.; Zang, Y. Sam fails to segment anything?–Sam-adapter: Adapting sam in underperformed scenes: Camouflage, shadow, and more. arXiv 2023, arXiv:2304.09148. [Google Scholar]
- Zhou, T.; Xia, W.; Zhang, F.; Chang, B.; Wang, W.; Yuan, Y.; Konukoglu, E.; Cremers, D. Image segmentation in foundation model era: A survey. arXiv 2024, arXiv:2408.12957. [Google Scholar] [CrossRef]
- Carion, N.; Gustafson, L.; Hu, Y.-T.; Debnath, S.; Hu, R.; Suris, D.; Ryali, C.; Alwala, K.V.; Khedr, H.; Huang, A. SAM 3: Segment Anything with Concepts. arXiv 2025, arXiv:2511.16719. [Google Scholar] [CrossRef]
- Wang, H.; Vasu, P.K.A.; Faghri, F.; Vemulapalli, R.; Farajtabar, M.; Mehta, S.; Rastegari, M.; Tuzel, O.; Pouransari, H. Sam-clip: Merging vision foundation models towards semantic and spatial understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024. [Google Scholar]
- Ji, W.; Li, J.; Bi, Q.; Liu, T.; Li, W.; Cheng, L. Segment Anything Is Not Always Perfect: An Investigation of Sam on Different Real-World Applications; Springer: Berlin/Heidelberg, Germany, 2024. [Google Scholar]
- Isensee, F.; Rokuss, M.; Krämer, L.; Dinkelacker, S.; Ravindran, A.; Stritzke, F.; Hamm, B.; Wald, T.; Langenberg, M.; Ulrich, C. nninteractive: Redefining 3d promptable segmentation. arXiv 2025, arXiv:2503.08373. [Google Scholar] [CrossRef]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.-Y. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023. [Google Scholar]
- Khatua, A.; Bhattacharya, A.; Goswami, A.K.; Aithal, B.H. Developing approaches in building classification and extraction with synergy of YOLOV8 and SAM models. Spat. Inf. Res. 2024, 32, 511–530. [Google Scholar] [CrossRef]
- Li, J.; Feng, Y.; Guo, Y.; Huang, J.; Piao, Y.; Bi, Q.; Zhang, M.; Zhao, X.; Chen, Q.; Zou, S. SAM3-I: Segment Anything with Instructions. arXiv 2025, arXiv:2512.04585. [Google Scholar] [CrossRef]
- Li, J.; Wang, Z.; Sun, X.; Xu, N.; You, Z.; Huang, D. AutoSAM: Auto-Prompting Mamba-Based Vision Foundation Model for Multimodal Remote Sensing Semantic Segmentation. IEEE Trans. Geosci. Remote Sens. 2026, 64, 5612421. [Google Scholar] [CrossRef]
- Du, X.; Kolkin, N.; Shakhnarovich, G.; Bhattad, A. Generative models: What do they know? Do they know things? Let’s find out! arXiv 2023, arXiv:2311.17137. [Google Scholar]
- Labs, B.F.; Batifol, S.; Blattmann, A.; Boesel, F.; Consul, S.; Diagne, C.; Dockhorn, T.; English, J.; English, Z.; Esser, P. FLUX. 1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space. arXiv 2025, arXiv:2506.15742. [Google Scholar]
- Yang, C.; Liu, C.; Deng, X.; Kim, D.; Mei, X.; Shen, X.; Chen, L.-C. 1.58-bit FLUX. arXiv 2024, arXiv:2412.18653. [Google Scholar]
- Wu, C.; Li, J.; Zhou, J.; Lin, J.; Gao, K.; Yan, K.; Yin, S.-m.; Bai, S.; Xu, X.; Chen, Y. Qwen-image technical report. arXiv 2025, arXiv:2508.02324. [Google Scholar] [CrossRef]
- Hayou, S.; Ghosh, N.; Yu, B. Lora+: Efficient low rank adaptation of large models. arXiv 2024, arXiv:2402.12354. [Google Scholar]
- Wang, Z.; Ma, H.; Zhai, J. Low-rank adaptation for edge AI. Sci. Rep. 2025, 15, 33109. [Google Scholar] [CrossRef]
- Akbulut, Z.; Özdemir, S.; Karslı, F. Comparative Analysis of Vision Foundation Models for Building Segmentation in Aerial Imagery. Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci. 2025, XLVIII-M-6-2025, 23–29. [Google Scholar] [CrossRef]
- Council, B.C. Our City’s Housing and Homelessness Strategy; Brisbane City Council: Brisbane, QLD, Australia, 2023; p. 32.
- Şimşek, E.; Negin, F.; Özyer, G.T.; Özyer, B. Leveraging foreground–background cues for semantically-driven, training-free moving object detection. Eng. Appl. Artif. Intell. 2024, 136, 108873. [Google Scholar] [CrossRef]
- Labs, B.F. FLUX.2: Frontier Visual Intelligence; Black Forest Labs: Freiburg, Germany, 2025; Available online: https://bfl.ai/blog/flux-2 (accessed on 31 December 2025).
- Chen, T.; Cao, R.; Yu, X.; Zhu, L.; Ding, C.; Ji, D.; Chen, C.; Zhu, Q.; Xu, C.; Mao, P. SAM3-Adapter: Efficient Adaptation of Segment Anything 3 for Camouflage Object Segmentation, Shadow Detection, and Medical Image Segmentation. arXiv 2025, arXiv:2511.19425. [Google Scholar] [CrossRef]
- Xiong, X.; Wu, Z.; Lu, L.; Xia, Y. SAM3-UNet: Simplified Adaptation of Segment Anything Model 3. arXiv 2025, arXiv:2512.01789. [Google Scholar] [CrossRef]
- Dong, W.; Yu, J.; Huang, Y.; Wang, H.; Zhu, L.; Chung, A.; Ren, H.; Bai, L. More than Segmentation: Benchmarking SAM 3 for Segmentation, 3D Perception, and Reconstruction in Robotic Surgery. arXiv 2025, arXiv:2512.07596. [Google Scholar] [CrossRef]
- Han, J.; Yoon, S.; Kang, M.; Kim, T. Approach to Enhancing Panoramic Segmentation in Indoor Construction Sites Based on a Perspective Image Segmentation Foundation Model. Appl. Sci. 2025, 15, 4875. [Google Scholar] [CrossRef]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. ImageNet classification with deep convolutional neural networks. Commun. ACM 2017, 60, 84–90. [Google Scholar] [CrossRef]
- Tournaire, O.; Brédif, M.; Boldo, D.; Durupt, M. An efficient stochastic approach for building footprint extraction from digital elevation models. ISPRS J. Photogramm. Remote Sens. 2010, 65, 317–327. [Google Scholar] [CrossRef]
- Sirmacek, B.; Unsalan, C. Urban-Area and Building Detection Using SIFT Keypoints and Graph Theory. IEEE Trans. Geosci. Remote Sens. 2009, 47, 1156–1167. [Google Scholar] [CrossRef]
- Argyridis, A.; Argialas, D.P. Building change detection through multi-scale GEOBIA approach by integrating deep belief networks with fuzzy ontologies. Int. J. Image Data Fusion 2016, 7, 148–171. [Google Scholar] [CrossRef]
- Powers, D.M. Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. arXiv 2020, arXiv:2010.16061. [Google Scholar] [CrossRef]
- Osco, L.P.; Wu, Q.; de Lemos, E.L.; Gonçalves, W.N.; Ramos, A.P.M.; Li, J.; Marcato, J. The Segment Anything Model (SAM) for remote sensing applications: From zero to one shot. Int. J. Appl. Earth Obs. Geoinf. 2023, 124, 103540. [Google Scholar] [CrossRef]
- Benchabana, A.; Kholladi, M.-K.; Bensaci, R.; Khaldi, B. Building detection in high-resolution remote sensing images by enhancing superpixel segmentation and classification using deep learning approaches. Buildings 2023, 13, 1649. [Google Scholar] [CrossRef]
- Yin, J.; Wu, F.; Qiu, Y.; Li, A.; Liu, C.; Gong, X. A Multiscale and Multitask Deep Learning Framework for Automatic Building Extraction. Remote Sens. 2022, 14, 4744. [Google Scholar] [CrossRef]
- Shata, D.; Denman, S.; Omrani, S.; Drogemuller, R.; Ali, H.; Wagdy, A. Parameter-Efficient Adaptation of Generative-Foundation (Flux, Qwen) vs. Zero-Shot (Gemini, SAM3) Models for Aerial Image Segmentation [Data Set]. Zenodo. 2026. Available online: https://zenodo.org/records/18571111 (accessed on 31 December 2025).
- Art Pixels. Pixels-Art.com. Available online: https://www.pixels-art.com (accessed on 31 December 2025).





























| Data Type | Description | Quantity | Format/Resolution |
|---|---|---|---|
| Primary aerial imagery | High-resolution tiles | 9797 | PNG, 1024 × 1024 px |
| Ground-truth rooftop masks | Binary masks | 121 | PNG, 1024 × 1024 px |
| LoRA training | Aerial images, corresponding binary masks | 97 training tiles and 24 test tiles | PNG mask |
| Parameter | Description | Value |
|---|---|---|
| Confidence threshold | Minimum confidence score to keep detection (range from 0 to 1) | 0.50 |
| Pipeline mode | Segmentation mode | Positive only |
| Max detections | Maximum number of instances per tile | 100 |
| Instances | Keeping only detections whose boxes overlap a positive box or contain a positive point | All instances |
| Crop factor | The bounding box crop factor for building combined SEGS | 1.5 |
| Min size | Minimum region size for keeping segments (pixels) | 25 |
| Fill holes | Post-processing to fill internal holes in segments | Enabled |
| Component | Setting |
|---|---|
| Dataset size | 97 images |
| Resolutions bucketing | 512 × 512, 768 × 768, 1024 × 1024 px |
| Checkpoint interval | Every 250 steps |
| Training steps | 5000 |
| Parameter | Values | Description |
|---|---|---|
| Network architecture | FLUX.1 Kontext, FLUX.2, and Qwen Image Edit 2509 | LoRA |
| LoRA linear rank/α | 32/32 | Rank and scaling for linear layers |
| LoRA convolutional rank/α | 16/16 | Rank and scaling for convolutional layers |
| Optimiser | AdamW (8-bit) | Memory-efficient optimiser |
| Learning rate | 1 × 10−4 | Fixed learning rate |
| Weight decay | 1 × 10−4 | Regularisation term |
| Loss function | MSE (latent residual) | Mean squared error on diffusion residuals |
| Noise scheduler | FlowMatch | Dynamic weighting across diffusion timesteps |
| Batch size | 1 | Per-step image batch |
| Training steps | 5000 | Total optimisation iterations |
| Precision | bf16 (mixed) | Mixed precision for stability and speed |
| Model Architecture | “Black-Building” IoU | “White-Building” IoU | Performance Difference (Δ) |
|---|---|---|---|
| FLUX.1 Kontext | 95.8% | 89% | 6.8% |
| Qwen Image Edit | 95.02% | 86.74% | 8.28% |
| Gemini 3 Pro image preview | 93.27% | 85.14% | 8.13% |
| SAM3 | 94.07% | 84% | 10.07% |
| FLUX.2 | 93.15% | 82% | 11.15% |
| Model Category | Model | Paradigm | IoU | Dice | Accuracy | Precision | Recall |
|---|---|---|---|---|---|---|---|
| Zero-Shot | FLUX.1 Kontext | Editing (Zero-shot) | 45 | 55 | 79 | 71 | 50 |
| FLUX.2 | Editing (Zero-shot) | 67 | 78 | 90 | 91 | 70 | |
| Qwen Image Edit 2509 | Editing (Zero-shot) | 33 | 46 | 63 | 41 | 62 | |
| Gemini Pro | Editing (Zero-shot) | 85 | 91 | 95 | 94 | 91 | |
| SAM3 | Segmentation-First | 84 | 91 | 96 | 95 | 88 | |
| LoRA-Adapted | FLUX.1 Kontext | Step 4250 | 89 | 94 | 97 | 96 | 93 |
| FLUX.2 | Step 750 | 82 | 90 | 95 | 93 | 87 | |
| Qwen Image Edit 2509 | Step 3000 | 87 | 93 | 96 | 96 | 90 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Shata, D.; Denman, S.; Omrani, S.; Drogemuller, R.; Ali, H.; Wagdy, A. Parameter-Efficient Adaptation of Generative-Foundation (Flux, Qwen) vs. Zero-Shot (Gemini, SAM3) Models for Aerial Image Segmentation. Buildings 2026, 16, 1369. https://doi.org/10.3390/buildings16071369
Shata D, Denman S, Omrani S, Drogemuller R, Ali H, Wagdy A. Parameter-Efficient Adaptation of Generative-Foundation (Flux, Qwen) vs. Zero-Shot (Gemini, SAM3) Models for Aerial Image Segmentation. Buildings. 2026; 16(7):1369. https://doi.org/10.3390/buildings16071369
Chicago/Turabian StyleShata, Dina, Simon Denman, Sara Omrani, Robin Drogemuller, Hend Ali, and Ayman Wagdy. 2026. "Parameter-Efficient Adaptation of Generative-Foundation (Flux, Qwen) vs. Zero-Shot (Gemini, SAM3) Models for Aerial Image Segmentation" Buildings 16, no. 7: 1369. https://doi.org/10.3390/buildings16071369
APA StyleShata, D., Denman, S., Omrani, S., Drogemuller, R., Ali, H., & Wagdy, A. (2026). Parameter-Efficient Adaptation of Generative-Foundation (Flux, Qwen) vs. Zero-Shot (Gemini, SAM3) Models for Aerial Image Segmentation. Buildings, 16(7), 1369. https://doi.org/10.3390/buildings16071369

