Energy-Aware Multilingual Vision–Language Models for Drone Smart Sensing
Highlights
- Inference energy and task accuracy are statistically independent (Spearman , ) across five open-source VLMs and thirteen languages, enabling energy-aware model selection without perception penalty; Phi-3-V, LLaVA-1.5, and LLaVA-1.6 form a Pareto-efficient frontier spanning a energy range (– Wh per 1000 queries).
- Low-resource languages (Arabic, Basque, and Luxembourgish) incur a double penalty: they simultaneously lower task accuracy and, for Arabic, result in substantially higher inference energy costs (up to in score and Wh/1K), while Basque triggers inference collapse rather than genuine efficiency gains.
- A formal UAV query budget model identifies the Pareto-optimal VLM for any platform; on the DJI Matrice 300 RTK LLaVA-1.6 is preferred, while the energy-constrained Matrice 30 calls for LLaVA-1.5, providing actionable guidelines for energy-aware drone smart sensing deployment.
- The documented double penalty for low-resource languages has direct regulatory relevance under the EU AI Act’s non-discrimination requirements, underscoring the need for targeted multilingual fine-tuning before deploying VLM-based UAV perception systems in non-English-dominant operational regions.
Abstract
1. Introduction
- Methodological: Formal UAV Inference Energy Budget. We derive a closed-form constraint (Equations (1)–(5)) that maps a per-query inference energy figure to a platform-conditioned admissibility set, and reduces VLM model selection to a constrained optimisation over the AI Energy Score. We instantiate the constraint for four commercial DJI platforms (Section 4.6), producing platform-specific Pareto-optimal model recommendations that are not directly derivable from accuracy benchmarks alone.
- Methodological: Per-Language Energy Decomposition and Multi-Metric Ranking Framework. We extend the AI Energy Score protocol with a language-block decomposition (Equation (12)) that attributes per-language inference energy within a single contiguous measurement window, addressing a gap in the specification, which is monolingual by design. We additionally introduce a multi-metric composite ranking framework (Section 4.5) that combines accuracy, energy, efficiency, four task-type rankings, and consistency into a single ranking lattice, motivated by the absence of pairwise statistical separation in the raw accuracy comparison.
- Empirical: Joint Energy–Accuracy Benchmark and Orthogonality Finding. We provide one of the first joint energy–accuracy benchmarks of five leading open-source VLMs across thirteen languages on drone-relevant visual sensing tasks, establishing a reproducible baseline for energy-aware VLM selection in UAV deployments, and demonstrating that inference energy and task accuracy are statistically uncorrelated (Spearman , ).
- Empirical: Multimodal Extension of the Double-Penalty Effect. We document and quantify the double-penalty effect in VLMs, previously theorised for text-only LLMs on tokeniser fertility grounds [28] but not, to our knowledge, established in the multimodal setting: Arabic, Basque, and Luxembourgish simultaneously incur higher inference energy costs and lower task accuracy relative to high-resource languages, with direct implications for equitable multinational UAV deployments under the EU AI Act [36]. We further decompose the effect into Arabic over-generation and Basque collapse, two structurally distinct failure modes with different operational implications.
2. Related Work
2.1. Vision–Language Model Architectures
2.2. UAV Smart Sensing and Aerial Scene Understanding
2.3. Multilingual Evaluation of Foundation Models
2.4. AI Energy Efficiency Measurement
2.5. Benchmarking Datasets for Autonomous Perception
2.6. Summary and Positioning
3. Materials and Methods
3.1. UAV Inference Energy Budget
3.2. Models Under Evaluation
3.3. Dataset and Image Pre-Processing
3.4. Multilingual Task Design
3.5. Inference Configuration
3.6. Scoring Protocol
3.6.1. Scene Understanding—Semantic Similarity
3.6.2. Vehicle Detection—Keyword Recall
3.6.3. Road Condition—Feature Recall
3.6.4. Navigation Reasoning—Hybrid Score
3.6.5. Aggregation
Scoring-Function Design Rationale
Edge-Case Interpretation
3.7. Energy Measurement
3.8. Statistical Analysis
3.9. Reproducibility and Software Stack
4. Results
4.1. Overall Model Performance
4.2. Task-Type Performance Profiles
4.3. Per-Language Performance Patterns
4.4. Energy Consumption on the NVIDIA H200 GPU
Language-Dependent Energy Variation
4.5. Efficiency Frontier and Composite Ranking
4.6. Mission-Level Query Budget on Commercial UAV Platforms
- DJI Matrice 300 RTK [15]: Heavy-lift inspection platform; two TB60 batteries ( ); W; and min.
- DJI Matrice 350 RTK [17]: Next-generation heavy-lift successor; two TB65 batteries ( ; each); W; and min.
- DJI Matrice 30 [16]: Compact enterprise UAV; TB30 battery ( ); W; and min.
- DJI Mavic 3 Enterprise [18]: Portable reconnaissance UAV; ; W; and min.
4.7. Double-Penalty Effect for Low-Resource Languages
- Arabic (ar) is the clearest and most severe manifestation: its accuracy drops by to across all five models, combined with an energy premia of to /1K in four of the five models. Arabic triangles cluster firmly in the lower-right quadrant. Qwen2-VL is the sole exception ( /1K), reflecting more efficient Arabic tokenisation from its broader multilingual pretraining [42].
- Basque (eu) incurs steep accuracy penalties ( from to ) and shows a predominantly negative energy signature. Four of the five models exhibit negative Basque energy deltas: InternVL2 ( /1K), LLaVA-1.6 ( /1K), LLaVA-1.5 ( /1K), and Phi-3-V ( /1K). These models return very short or null responses to Basque prompts, terminating inference early. The corresponding Basque triangles appear in the lower-left quadrant of Figure 20 and represent a collapse penalty, inference latency reduced through model failure, rather than a genuine efficiency gain. Qwen2-VL is the exception, showing a positive delta ( /1K), suggesting that it generates more substantive Basque responses at the cost of higher energy.
- Luxembourgish (lb) shows milder accuracy penalties due to lexical proximity to German and French. LLaVA-1.5 achieves , the shallowest accuracy gap among all low-resource languages in the study, though its energy premium reaches /1K relative to English.
5. Discussion
5.1. Energy and Accuracy Are Orthogonal Design Axes
5.2. The Double Penalty Is Real and Language-Family Structured
5.3. Task Architecture Matters More than Model Size
5.4. Limitations
5.5. Future Directions
6. Mitigation Strategies, Deployment Guidance, and Implications for VLM Design
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| AGL | Above Ground Level |
| API | Application Programming Interface |
| BDD | Berkeley DeepDrive (dataset) |
| BLEU | Bilingual Evaluation Understudy |
| CLIP | Contrastive Language–Image Pretraining |
| EU | European Union |
| GQA | Grouped-Query Attention |
| GPU | Graphics Processing Unit |
| HBM3 | High Bandwidth Memory (3rd generation) |
| HELM | Holistic Evaluation of Language Models |
| LLM | Large Language Model |
| LoRA | Low-Rank Adaptation |
| MEGA | Massively Multilingual Evaluation of Generative AI |
| M-RoPE | Multimodal Rotary Position Embedding |
| NLP | Natural Language Processing |
| NVML | NVIDIA Management Library |
| PUE | Power Usage Effectiveness |
| ROUGE | Recall-Oriented Understudy for Gisting Evaluation |
| TDP | Thermal Design Power |
| UAV | Unmanned Aerial Vehicle |
| VLM | Vision–Language Model |
| vLLM | High-Throughput LLM Inference Engine |
| VQA | Visual Question Answering |
| VRAM | Video Random Access Memory |
| Wh | Watt-hour |
| XTREME | Cross-lingual TRansfer Evaluation of Multilingual Encoders |
References
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. Adv. Neural Inf. Process. Syst. 2017, 30, 1–11. [Google Scholar]
- de Curtò, J.; de Zarzà, I.; Calafate, C.T. Semantic Scene Understanding with Large Language Models on Unmanned Aerial Vehicles. Drones 2023, 7, 114. [Google Scholar] [CrossRef] [Scilit]
- OpenAI. GPT-4 Technical Report. arXiv 2023, arXiv:2303.08774. [Google Scholar] [CrossRef] [Scilit]
- Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv 2023, arXiv:2307.09288. [Google Scholar] [CrossRef] [Scilit]
- Zhang, L.; Hao, X.; Xu, Q.; Zhang, Q.; Zhang, X.; Wang, P.; Zhang, J.; Wang, Z.; Zhang, S.; Xu, R. Mapnav: A novel memory representation via annotated semantic maps for vlm-based vision-and-language navigation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Vienna, Austria, 2025; pp. 13032–13056. [Google Scholar]
- Song, D.; Liang, J.; Payandeh, A.; Raj, A.H.; Xiao, X.; Manocha, D. Vlm-social-nav: Socially aware robot navigation through scoring using vision-language models. IEEE Robot. Autom. Lett. 2024, 10, 508–515. [Google Scholar] [CrossRef] [Scilit]
- Elnoor, M.; Weerakoon, K.; Seneviratne, G.; Xian, R.; Guan, T.; Jaffar, M.K.M.; Rajagopal, V.; Manocha, D. VLM-GroNav: Robot Navigation Using Physically Grounded Vision-Language Models in Outdoor Environments. In Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2025; pp. 2391–2398. [Google Scholar]
- Strubell, E.; Ganesh, A.; McCallum, A. Energy and Policy Considerations for Deep Learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 3645–3650. [Google Scholar]
- Schwartz, R.; Dodge, J.; Smith, N.A.; Etzioni, O. Green AI. Commun. ACM 2020, 63, 54–63. [Google Scholar] [CrossRef] [Scilit]
- Luccioni, A.S.; Viguier, S.; Ligozat, A.L. Power Hungry Processing: Watts Driving the Cost of AI Deployment? In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, Rio de Janeiro, Brazil, 3–6 June 2024; pp. 85–99. [Google Scholar]
- de Vries, A. The Growing Energy Footprint of Artificial Intelligence. Joule 2023, 7, 2191–2194. [Google Scholar] [CrossRef] [Scilit]
- International Energy Agency. Electricity 2024: Analysis and Forecast to 2026, 2024. Available online: https://www.iea.org/reports/electricity-2024 (accessed on 7 April 2026).
- Patterson, D.; Gonzalez, J.; Le, Q.; Liang, C.; Munguia, L.M.; Rothchild, D.; So, D.; Texier, M.; Dean, J. Carbon Emissions and Large Neural Network Training. arXiv 2021, arXiv:2104.10350. [Google Scholar] [CrossRef] [Scilit]
- Henderson, P.; Hu, J.; Romoff, J.; Brunskill, E.; Jurafsky, D.; Pineau, J. Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning. J. Mach. Learn. Res. 2020, 21, 1–43. [Google Scholar]
- DJI. Matrice 300 RTK Specs, 2023. Available online: https://enterprise.dji.com/matrice-300/specs (accessed on 7 April 2026).
- DJI. Matrice 30 Series Specs, 2023. Available online: https://enterprise.dji.com/matrice-30/specs (accessed on 7 April 2026).
- DJI. DJI Matrice 350 RTK—Specifications. 2023. Available online: https://enterprise.dji.com/matrice-350-rtk/specs (accessed on 7 April 2026).
- DJI. DJI Mavic 3 Enterprise—Specifications. 2022. Available online: https://enterprise.dji.com/mavic-3-enterprise/specs (accessed on 7 April 2026).
- Hugging Face; Luccioni, S.; Jernite, Y.; Pierrard, R.; Moutawwakil, I.; Mitchell, M.; Gamazaychikov, B.; Chamberlin, S.; Hooker, S.; Wu, C.J.; et al. AI Energy Score: Standardized Energy Efficiency Ratings for AI Models, 2025. Available online: https://huggingface.co/AIEnergyScore (accessed on 7 April 2026).
- García-Martín, E.; Rodrigues, C.F.; Riley, G.; Grahn, H. Estimation of Energy Consumption in Machine Learning. J. Parallel Distrib. Comput. 2019, 134, 75–88. [Google Scholar] [CrossRef] [Scilit]
- Dodge, J.; Prewitt, T.; des Combes, R.T.; Odber, E.; Schwartz, R.; Strubell, E.; Luccioni, A.S.; Smith, N.A.; DeCario, N.; Buchanan, W. Measuring the Carbon Intensity of AI in Cloud Instances. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul, Republic of Korea, 21–24 June 2022; pp. 1877–1894. [Google Scholar]
- Hu, J.; Ruder, S.; Siddhant, A.; Neubig, G.; Firat, O.; Johnson, M. XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation. In Proceedings of the 37th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2020; pp. 4411–4421. [Google Scholar]
- Liang, Y.; Duan, N.; Gong, Y.; Wu, N.; Guo, F.; Qi, W.; Gong, M.; Shou, L.; Jiang, D.; Cao, G.; et al. XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Online, 16–20 November 2020; pp. 6008–6018. [Google Scholar]
- Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzmán, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; Stoyanov, V. Unsupervised Cross-lingual Representation Learning at Scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 5–10 July 2020; pp. 8440–8451. [Google Scholar]
- Ahuja, K.; Diddee, H.; Hada, R.; Ochieng, M.; Ramesh, K.; Jain, P.; Nambi, A.; Ganu, T.; Segal, S.; Ahmed, M.; et al. MEGA: Multilingual Evaluation of Generative AI. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023; pp. 4232–4267. [Google Scholar]
- Lai, V.D.; Ngo, N.T.; Veyseh, A.P.B.; Man, H.; Dernoncourt, F.; Bui, T.; Nguyen, T.H. ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning. In Findings of the Association for Computational Linguistics: EMNLP 2023; Association for Computational Linguistics: Singapore, 2023; pp. 13171–13189. [Google Scholar]
- Rust, P.; Pfeiffer, J.; Vulić, I.; Ruder, S.; Gurevych, I. How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, Virtual Event, 1–6 August 2021; pp. 3118–3135. [Google Scholar]
- Petrov, A.; La Malfa, E.; Torr, P.; Biber, A. Language Model Tokenizers Introduce Unfairness Between Languages. Adv. Neural Inf. Process. Syst. 2024, 36, 1608. [Google Scholar]
- de Curtò, J.; de Zarzà, I. Metamorphic Testing for Semantic Invariance in Large Language Models. IEEE Access 2025, 13, 214772–214791. [Google Scholar] [CrossRef] [Scilit]
- Yu, F.; Chen, H.; Wang, X.; Xian, W.; Chen, Y.; Liu, F.; Madhavan, V.; Darrell, T. BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 2636–2645. [Google Scholar]
- Reimers, N.; Gurevych, I. Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, Hong Kong, China, 3–7 November 2019; pp. 3982–3992. [Google Scholar]
- Friedman, M. The Use of Ranks to Avoid the Assumption of Normality Implicit in the Analysis of Variance. J. Am. Stat. Assoc. 1937, 32, 675–701. [Google Scholar] [CrossRef]
- Holm, S. A Simple Sequentially Rejective Multiple Test Procedure. Scand. J. Stat. 1979, 6, 65–70. [Google Scholar]
- Spearman, C. The Proof and Measurement of Association between Two Things. Am. J. Psychol. 1904, 15, 72–101. [Google Scholar] [CrossRef] [Scilit]
- Cohen, J. Statistical Power Analysis for the Behavioral Sciences, 2nd ed.; Lawrence Erlbaum Associates: Mahwah, NJ, USA, 1988. [Google Scholar]
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 Laying Down Harmonised Rules on Artificial Intelligence (AI Act). Off. J. Eur. Union 2024, L 2024/1689, 1–144. Available online: http://data.europa.eu/eli/reg/2024/1689/oj (accessed on 7 April 2026).
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning; Proceedings of Machine Learning Research; PMLR: Cambridge, MA, USA, 2021; Volume 139, pp. 8748–8763. [Google Scholar]
- Li, J.; Li, D.; Savarese, S.; Hoi, S. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In Proceedings of the 40th International Conference on Machine Learning; Proceedings of Machine Learning Research; PMLR: Cambridge, MA, USA, 2023; Volume 202, pp. 19730–19742. [Google Scholar]
- Liu, H.; Li, C.; Li, Y.; Lee, Y.J. Improved Baselines with Visual Instruction Tuning. arXiv 2023, arXiv:2310.03744. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Li, C.; Li, Y.; Li, B.; Zhang, Y.; Shen, S.; Lee, Y.J. LLaVA-NeXT: Improved Reasoning, OCR, and World Knowledge, 2024. Available online: https://llava-vl.github.io/blog/2024-01-30-llava-next/ (accessed on 7 April 2026).
- Chen, Z.; Wang, W.; Tian, H.; Ye, S.; Gao, Z.; Cui, E.; Tong, W.; Hu, K.; Luo, J.; Ma, Z.; et al. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites. arXiv 2024, arXiv:2404.16821. [Google Scholar] [CrossRef] [Scilit]
- Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. arXiv 2024, arXiv:2409.12191. [Google Scholar]
- Abdin, M.; Aneja, J.; Awadalla, H.; Awasthi, A.; Awan, A.A.; Bach, N.; Bahree, A.; Bakhtiari, A.; Bao, J.; Behl, H.; et al. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone. arXiv 2024, arXiv:2404.14219. [Google Scholar] [CrossRef] [Scilit]
- Ainslie, J.; Lee-Thorp, J.; de Jong, M.; Zemlyanskiy, Y.; Lebrón, F.; Sanghai, S. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023; pp. 4895–4901. [Google Scholar]
- Elsken, T.; Metzen, J.H.; Hutter, F. Neural Architecture Search: A Survey. J. Mach. Learn. Res. 2019, 20, 1–21. [Google Scholar]
- Cai, H.; Gan, C.; Wang, T.; Zhang, Z.; Han, S. Once-for-All: Train One Network and Specialize It for Efficient Deployment. In Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia, 30 April 2020. [Google Scholar]
- Tan, M.; Le, Q. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2019; pp. 6105–6114. [Google Scholar]
- Shi, J.; Huang, K.; Pan, H.; Xu, J.; Cheng, C.; Zhang, H. Autonomous Subtask Generation for Indoor Search and Rescue Mission via Large-Language-Model and Behavior-Tree Integration. In Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2025; pp. 16710–16716. [Google Scholar]
- Zhou, G.; Hong, Y.; Wu, Q. Navgpt: Explicit reasoning in vision-and-language navigation with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2024; Volume 38, pp. 7641–7649. [Google Scholar]
- Shinoda, R.; Inoue, N.; Kataoka, H.; Onishi, M.; Ushiku, Y. Agrobench: Vision-language model benchmark in agriculture. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 19–23 October 2025; pp. 7634–7644. [Google Scholar]
- Wang, Y.; Cui, J.; Zhai, C.; Tao, X.; Li, Y. Integrating segmentation and vision-language model for automated and interpretable building damage assessment from satellite imagery. Adv. Eng. Inform. 2026, 71, 104320. [Google Scholar] [CrossRef] [Scilit]
- Xiao, D.; Dianati, M.; Jennings, P.; Woodman, R. Hazardvlm: A video language model for real-time hazard description in automated driving systems. IEEE Trans. Intell. Veh. 2024, 10, 3331–3343. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Hu, X.; Hou, W.; Chen, H.; Zheng, R.; Wang, Y.; Yang, L.; Huang, H.; Ye, W.; Geng, X.; et al. On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective. arXiv 2023, arXiv:2302.12095. [Google Scholar]
- Lu, Y.; Bartolo, M.; Moore, A.; Riedel, S.; Stenetorp, P. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. arXiv 2022, arXiv:2104.08786. [Google Scholar] [CrossRef] [Scilit]
- Lacoste, A.; Luccioni, A.; Schmidt, V.; Dandres, T. Quantifying the Carbon Emissions of Machine Learning. arXiv 2019, arXiv:1910.09700. [Google Scholar] [CrossRef] [Scilit]
- Courty, B.; Schmidt, V.; Luccioni, S.; Goyal-Kamal; MarionCoutarel; Feld, B.; Lecourt, J.; LiamConnell; Saboni, A.; Inimaz; et al. CodeCarbon: Track and Reduce CO2 Emissions from Compute, 2020. Available online: https://github.com/mlco2/codecarbon (accessed on 7 April 2026).
- Bannour, N.; Ghannay, S.; Nevéol, A.; Ligozat, A.L. Evaluating the Carbon Footprint of NLP Methods: A Survey and Analysis of Existing Tools. In Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing, Virtual, 10 November 2021; pp. 11–21. [Google Scholar]
- de Zarzà, I.; Liz, M.; de Curtò, J.; Calafate, C.T. Energy-Aware Multilingual Evaluation of Large Language Models. Electronics 2026, 15, 1395. [Google Scholar] [CrossRef] [Scilit]
- Chitty-Venkata, K.T.; Emani, M.; Vishwanath, V.; Somani, A.K. Neural Architecture Search for Transformers: A Survey. IEEE Access 2023, 11, 108374–108412. [Google Scholar] [CrossRef] [Scilit]
- Papineni, K.; Roukos, S.; Ward, T.; Zhu, W.J. BLEU: A Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, PA, USA, 7–12 July 2002; pp. 311–318. [Google Scholar]
- Lin, C.Y. ROUGE: A Package for Automatic Evaluation of Summaries. In Proceedings of the Text Summarization Branches Out, Barcelona, Spain, 25–26 July 2004; pp. 74–81. [Google Scholar]
- Gao, T.; Yao, X.; Chen, D. SimCSE: Simple Contrastive Learning of Sentence Embeddings. arXiv 2021, arXiv:2104.08821. [Google Scholar]
- Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; et al. Holistic Evaluation of Language Models. arXiv 2022, arXiv:2211.09110. [Google Scholar] [CrossRef] [Scilit]
- Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A.A.M.; Abid, A.; Fisch, A.; Brown, A.R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models. Trans. Mach. Learn. Res. 2023, 2023, 1–95. Available online: https://openreview.net/forum?id=uyTL5Bvosj (accessed on 7 April 2026).
- de Curtò, J.; de Zarzà, I. Comparative Analysis of Reasoning Capabilities in Foundation Models. In Proceedings of the 2024 2nd International Conference on Foundation and Large Language Models (FLLM), Dubai, United Arab Emirates, 26–29 November 2024; pp. 141–149. [Google Scholar] [CrossRef] [Scilit]
- Kwon, W.; Li, Z.; Zhuang, S.; Sheng, Y.; Zheng, L.; Yu, C.H.; Gonzalez, J.E.; Zhang, H.; Stoica, I. Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP); ACM: New York, NY, USA, 2023; pp. 611–626. [Google Scholar] [CrossRef] [Scilit]
- Zhu, P.; Wen, L.; Du, D.; Bian, X.; Fan, H.; Hu, Q.; Ling, H. Detection and Tracking Meet Drones Challenge. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 7380–7399. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bozcan, I.; Kayacan, E. AU-AIR: A Multi-modal Unmanned Aerial Vehicle Dataset for Low Altitude Traffic Surveillance. In Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May–4 June 2020; pp. 8504–8510. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; Zhou, D. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv 2023, arXiv:2203.11171. [Google Scholar] [CrossRef] [Scilit]
- Elazar, Y.; Kassner, N.; Ravfogel, S.; Ravichander, A.; Hovy, E.; Schütze, H.; Goldberg, Y. Measuring and Improving Consistency in Pretrained Language Models. Trans. Assoc. Comput. Linguist. 2021, 9, 1012–1031. [Google Scholar] [CrossRef] [Scilit]
- NLLB Team; Costa-jussà, M.R.; Cross, J.; Çelebi, O.; Elbayad, M.; Heafield, K.; Heffernan, K.; Kalbassi, E.; Lam, J.; Licht, D.; et al. No Language Left Behind: Scaling Human-Centered Machine Translation. arXiv 2022, arXiv:2207.04672. [Google Scholar] [CrossRef] [Scilit]
- Virtanen, P.; Gommers, R.; Oliphant, T.E.; Haberland, M.; Reddy, T.; Cournapeau, D.; Burovski, E.; Peterson, P.; Weckesser, W.; Bright, J.; et al. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python. Nat. Methods 2020, 17, 261–272. [Google Scholar] [CrossRef] [Scilit]
- Mann, H.B.; Whitney, D.R. On a Test of Whether One of Two Random Variables Is Stochastically Larger than the Other. Ann. Math. Stat. 1947, 18, 50–60. [Google Scholar] [CrossRef] [Scilit]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
- Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Online, 16–20 November 2020; pp. 38–45. [Google Scholar]
- Wendler, C.; Veselovsky, V.; Monea, G.; West, R. Do Llamas Work in English? On the Latent Language of Multilingual Transformers. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Bangkok, Thailand, 11–16 August 2024; pp. 15366–15394. [Google Scholar]
- Feng, F.; Yang, Y.; Cer, D.; Arivazhagan, N.; Wang, W. Language-agnostic BERT sentence embedding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Dublin, Ireland, 2022; pp. 878–891. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the International Conference on Learning Representations, Online, 25–29 April 2022. [Google Scholar]




















| Study | Multimodal | Multilingual | Energy | UAV Context |
|---|---|---|---|---|
| Hu et al. (2020) [22] | – | ✔ | – | – |
| Ahuja et al. (2023) [25] | – | ✔ | – | – |
| Lai et al. (2023) [26] | – | ✔ | – | – |
| Luccioni et al. (2024) [10] | ✔ | – | ✔ | – |
| AI Energy Score (2025) [19] | ✔ | – | ✔ | – |
| Petrov et al. (2024) [28] | – | ✔ | ✔ | – |
| de Zarzà et al. (2026) [58] | – | ✔ | ✔ | – |
| This work | ✔ | ✔ | ✔ | ✔ |
| Short Name | HuggingFace Identifier | Params | Key Design Feature |
|---|---|---|---|
| InternVL2 | OpenGVLab/InternVL2-8B | 8B | InternViT-6B encoder; multilingual pretraining |
| Qwen2-VL | Qwen/Qwen2-VL-7B-Instruct | 7B | Naive Dynamic Resolution; M-RoPE |
| LLaVA-1.5 | llava-hf/llava-1.5-7b-hf | 7B | CLIP-ViT-L + MLP connector; Vicuna backbone |
| LLaVA-1.6 | llava-hf/llava-v1.6-mistral-7b-hf | 7B | 4× resolution tiling; Mistral backbone |
| Phi-3-V | microsoft/Phi-3-vision-128k-instruct | 4.2B | Azure AI CLIP encoder; 128k context window |
| Task | Type Key | Difficulty | Prompt Summary (English) |
|---|---|---|---|
| Scene understanding | scene_understanding | Easy | Describe the driving scene; what can you see on the road and around it? |
| Vehicle detection | vehicle_detection | Medium | Identify vehicles, pedestrians, and important objects in this driving scene. |
| Road condition | road_condition | Medium | Describe road conditions and environment: road type, weather, and lighting. |
| Navigation reasoning | navigation_reasoning | Hard | From an autonomous driving perspective, identify obstacles or hazards and their positions relative to the vehicle. |
| Component | Version/Reference |
|---|---|
| Python | 3.10 |
| PyTorch | ≥2.3 [74] |
| transformers | ≥4.44 [75] |
| vLLM | ≥0.5 |
| sentence-transformers | ≥2.6.1 [31] |
| pynvml | ≥11.5 |
| SciPy | ≥1.10 [72] |
| datasets (HF) | ≥2.18 |
| pandas/seaborn/matplotlib | standard scientific stack |
| Model | Scene | Veh. | Road | Nav. | |||||
|---|---|---|---|---|---|---|---|---|---|
| (Wh) | (Wh/1K) | (W) | |||||||
| LLaVA-1.6 | 0.160 | 0.032 | 4260.3 | 130.0 | 355.1 | 0.229 | 0.215 | 0.083 | 0.113 |
| LLaVA-1.5 | 0.137 | 0.028 | 2701.4 | 82.5 | 355.0 | 0.242 | 0.146 | 0.036 | 0.124 |
| InternVL2 | 0.137 | 0.026 | 5401.5 | 164.9 | 358.2 | 0.183 | 0.164 | 0.072 | 0.127 |
| Qwen2-VL | 0.126 | 0.005 | 3657.2 | 111.6 | 344.9 | 0.192 | 0.123 | 0.075 | 0.115 |
| Phi-3-V | 0.125 | 0.030 | 2173.2 | 66.3 | 323.2 | 0.217 | 0.109 | 0.072 | 0.101 |
| Platform | B | Admitted | Pareto | ||||
|---|---|---|---|---|---|---|---|
| (Wh) | (W) | (min) | Hz | Choice | Hz | ||
| DJI Matrice 300 RTK | 548.0 | 598 | 55 | 332 | All five | LLaVA-1.6 | 166 |
| DJI Matrice 350 RTK | 322.8 | 352 | 55 | 196 | All five | LLaVA-1.6 | 98 |
| DJI Matrice 30 | 131.6 | 193 | 41 | 107 | Phi-3-V, LLaVA-1.5 | LLaVA-1.5 | 54 |
| DJI Mavic 3 Ent. | 77.0 | 103 | 45 | 57 | None | — | 29 |
| Platform | Phi-3-V | LLaVA-1.5 | Qwen2-VL | LLaVA-1.6 | InternVL2 |
|---|---|---|---|---|---|
| DJI Matrice 300 RTK | 1.1 min (2.0%) | 1.3 min (2.4%) | 1.8 min (3.2%) | 2.1 min (3.8%) | 2.6 min (4.7%) |
| DJI Matrice 350 RTK | 1.8 min (3.3%) | 2.2 min (4.0%) | 3.0 min (5.4%) | 3.4 min (6.2%) | 4.3 min (7.8%) |
| DJI Matrice 30 | 2.4 min (5.8%) | 2.9 min (7.1%) | 3.9 min (9.4%) | 4.4 min (10.8%) | 5.5 min (13.3%) |
| DJI Mavic 3 Ent. | 4.7 min (10.4%) | 5.7 min (12.6%) | 7.4 min (16.4%) | 8.4 min (18.6%) | 10.1 min (22.4%) |
| Platform (Budget) | High-Resource Lang. | Low-Resource Lang. |
|---|---|---|
| M300/M350 RTK (≥322 ) | LLaVA-1.6-Mistral-7B [LLaVA-1.5, InternVL2] | LLaVA-1.6-Mistral-7B + T2P |
| Matrice 30 (∼132 ) | LLaVA-1.5 [Phi-3-V] | LLaVA-1.5 + T2P |
| Mavic 3 Ent. (∼77 ) | Phi-3-V (relaxed ) [no model strictly admissible] | Phi-3-V + T2P (relaxed ) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
de Curtò, J.; Liz, M.; de Zarzà, I.; Calafate, C.T. Energy-Aware Multilingual Vision–Language Models for Drone Smart Sensing. Drones 2026, 10, 361. https://doi.org/10.3390/drones10050361
de Curtò J, Liz M, de Zarzà I, Calafate CT. Energy-Aware Multilingual Vision–Language Models for Drone Smart Sensing. Drones. 2026; 10(5):361. https://doi.org/10.3390/drones10050361
Chicago/Turabian Stylede Curtò, J., Mauro Liz, I. de Zarzà, and Carlos T. Calafate. 2026. "Energy-Aware Multilingual Vision–Language Models for Drone Smart Sensing" Drones 10, no. 5: 361. https://doi.org/10.3390/drones10050361
APA Stylede Curtò, J., Liz, M., de Zarzà, I., & Calafate, C. T. (2026). Energy-Aware Multilingual Vision–Language Models for Drone Smart Sensing. Drones, 10(5), 361. https://doi.org/10.3390/drones10050361

