Joint 3D Trajectory and Power Optimization for UAV Swarms in Cell-Free Massive MIMO Networks: A CTDE-MAPPO Framework for Sensing-Aware Precision Agriculture
Highlights
- A joint sensing–communication utility function is introduced that balances communication energy efficiency and field coverage completeness in a single weighted objective that is embedded directly into a MAPPO-CTDE reward design with hard battery depletion and return-to-depot constraints—a formulation that, to our knowledge, is the first to jointly integrate these specific elements within a CF-mMIMO UAV setting, building on established CTDE, MAPPO, and reward-shaping techniques.
- The CF-mMIMO CPU serves as a global critic during training, resolving inter-agent non-stationarity while each UAV executes a lightweight local policy at runtime, achieving 93% field coverage and a 92% depot-return rate without any inter-UAV communication overhead at execution time.
- Jointly optimizing 3D rotary-wing trajectories and transmit power through a shared coverage map creates strong inter-agent coupling that makes the CF-mMIMO CPU a natural and architecturally motivated centralized critic—enabling scalable swarm coordination without dedicated inter-UAV communication links.
- CF-mMIMO macro-diversity, combined with sensing-aware trajectory adaptation, offers a practical and deployable pathway toward energy-autonomous UAV swarms for large-scale crop monitoring in 6G agricultural networks.
Abstract
1. Introduction
1.1. Motivation and Problem Statement
- A UAV that hovers exclusively at the altitude yielding the best channel quality will maximize communication efficiency but fail to sweep or scan the field.
- A UAV that ignores aerodynamic battery depletion will drain its reserves prematurely and fail to return safely to the depot.
- A swarm mission that skips partially surveyed zones will leave the end-user with an incomplete, unusable crop map.
1.2. Paper Contributions
- We propose a CF-mMIMO multi-UAV remote sensing (RS) architecture for precision agriculture. To accurately reflect real-world deployments, this architecture incorporates (i) a three-dimensional (3D) rotary-wing mobility model, (ii) a probabilistic air-to-ground fading channel, and (iii) a full-mission battery depletion model.
- We formulate a joint sensing–communication (SC) utility function that simultaneously rewards communication energy efficiency and field coverage completeness under a single weighted objective (Section 4). We cast the resulting optimization as Problem : a non-convex, long-horizon, battery-constrained trajectory and power allocation problem subject to hard battery depletion limits, mandatory return-to-depot constraints, and minimum uplink rate requirements. We demonstrate analytically that cannot be decomposed slot-by-slot, thereby ruling out classic SCA and supervised DNN approaches, and firmly motivating our MADRL reformulation. This formulation is absent from prior CF-mMIMO UAV studies, which either omit sensing coverage entirely or treat battery constraints as soft penalties.
- Our proposed methodology centers on a Multi-Agent Proximal Policy Optimization (MAPPO) approach, which is selected specifically to mitigate the non-stationary dynamics of collaborative aerial swarms. By implementing a Centralized Training with Decentralized Execution (CTDE) strategy, we eliminate the need for inter-UAV communication during live missions. The training phase utilizes the CF-mMIMO CPU as a shared global critic that processes the complete system state to refine the behavior of all agents. Once deployed, each UAV operates autonomously, executing a localized policy that continuously adapts spatial trajectories and uplink power based on real-time battery levels and fading channel conditions. Furthermore, we carefully engineer the reward mechanisms and state-action spaces to ensure that the fleet naturally converges on strategies that maximize agricultural sensing coverage, avoid collisions, and respect absolute energy constraints.
2. Related Work
2.1. Power Allocation in CF-mMIMO Networks
2.2. UAV Trajectory Optimization
2.3. Joint Trajectory and Power Optimization for UAV-Based Systems
2.4. Positioning Relative to Recent UAV Swarm Optimization Frameworks
- First, their UAVs act as flying access points serving ground users. In contrast, the UAVs in this paper are mobile remote sensors whose primary function is to collect crop and soil data. This sensing role introduces a coverage completeness requirement that is absent in [32].
- Second, their energy model targets fixed-wing UAVs, which have fundamentally different propulsion characteristics from the rotary-wing platforms prevalent in agricultural settings. This paper adopts the Zeng et al. [21] rotary-wing model with per-UAV battery depletion enforced as a hard mission constraint.
- Third, the formulation of [32] includes no return-to-depot requirement. Such a requirement is a practical necessity in agricultural deployments, where UAVs must land at a charging station upon mission completion.
- Fourth, their objective is purely communication-centric, maximizing energy efficiency defined as throughput per unit energy. This paper instead introduces a joint sensing–communication (SC) utility that simultaneously rewards field coverage and communication efficiency, reflecting the dual purpose of the UAV swarm.
3. CF-mMIMO-Assisted Precision Agriculture: System Model and Assumptions
3.1. CF-mMIMO Network Topology
3.2. UAV Mobility Model
3.3. Air-to-Ground Channel Model
3.4. CF-mMIMO Uplink Transmission
3.5. Energy Consumption Model
3.6. Sensing Coverage Model
4. Joint Optimization Problem
4.1. Joint Sensing–Communication Utility
4.2. Optimization Problem
- is the minimum required field coverage fraction.
- Constraint (23) enforces per-UAV transmit power limits.
- Constraints (24) and (25) enforce mobility and altitude bounds.
- Constraint (26) is the collision avoidance condition.
- Constraint (27) enforces that the battery state cannot fall below zero. This is a physical, state-space invariant enforced by construction in the environment dynamics (Section 5.1.4) rather than a constraint requiring active enforcement by the learned policy. Avoidance of the depletion event itself—i.e., discouraging a UAV from actually reaching zero battery mid-mission—is instead handled through reward shaping, as detailed in Section 5.1.
- Constraint (28) enforces the return-to-depot requirement, which is essential for battery-charged agricultural UAVs but is absent from the formulation of [32].
- Constraints (29) and (30) enforce minimum uplink rate and coverage completeness, respectively.
4.3. Why This Problem Cannot Be Solved by Existing Methods?
5. Proposed MAPPO-Based MADRL Framework
5.1. Markov Decision Process Formulation
5.1.1. Local Observation Space
5.1.2. Global State
5.1.3. Action Space
5.1.4. Reward Function
5.2. MAPPO Algorithm
- is the importance sampling ratio;
- is the Generalized Advantage Estimate (GAE), which is given by
- is the clipping coefficient.
| Algorithm 1 MAPPO Training for Joint Sensing–Communication Optimization in CF-mMIMO UAV Networks |
|
6. Simulation Results
6.1. Simulation Setup
6.2. Experimental Protocol and Hyperparameter Configuration
6.2.1. Network Architecture and Initialization
6.2.2. Optimization and Exploration Schedule
6.2.3. Reward Weight Configuration and Tuning
6.2.4. Training Budget and Checkpoint Selection
6.2.5. Computational Infrastructure
6.3. Impact Analysis of Proposed Framework Components
6.4. Convergence of MAPPO Training
6.5. Uplink SINR Coverage Analysis
- Static lawnmower (CF-mMIMO): UAVs follow fixed parallel sweep trajectories with equal power allocation while retaining the full cell-free AP cooperation. This baseline isolates the gain of the adaptive trajectory.
- MADRL (Cellular mMIMO): The same MAPPO policy is deployed over a non-cooperative cellular architecture, where each UAV is served by its single strongest AP. This baseline isolates the gain of cell-free cooperation.
- Cellular mMIMO: A non-cooperative single-AP network is used that has no learned trajectory optimization.
6.6. System Energy Efficiency
- Supervised DNN: a neural network trained via behavioral cloning using SCA-generated optimal power labels at .
- Static Lawnmower: a non-adaptive baseline where UAVs follow pre-computed parallel sweep trajectories with a fixed power allocation of .
- SCA: an iterative successive convex approximation algorithm with perfect instantaneous CSI and a fixed lawnmower trajectory. SCA assumes perfect instantaneous CSI and optimizes transmit power only along the same fixed lawnmower trajectory used by the static lawnmower baseline; it therefore bounds fixed-trajectory power-control methods and is not a bound on the joint trajectory-and-power problem addressed by MAPPO.
- SCMA-MADDPG: the MARL-based scheme proposed in [32], which optimizes for communication metrics without sensing awareness. This baseline was reimplemented (not taken from the original paper), under our own environment, sharing our network architecture and training procedure, differing only in the reward function, with hyperparameters matched to the proposed method for fairness.
- Parameter-Shared MAPPO: a parameter-shared MAPPO baseline, trained under the same procedure, budget, and hyperparameters as our proposed method, differing only in that its actor weights are shared across all UAVs rather than trained independently.
6.7. Sensitivity to the Communication–Sensing Trade-Off Coefficient
6.8. Battery-Coverage Dynamics
6.9. 3D Spatial Trajectories and Territory Partitioning
7. Discussion
7.1. Computational Complexity and Scalability Analysis
- The number of UAVs K;
- The number of APs M;
- The mission horizon T;
- The hidden-layer width H; and
- The number of PPO update epochs .
| Algorithm | Training Complexity | Execution Complexity | CSI Needed | Scales with K | Adaptive |
|---|---|---|---|---|---|
| Proposed MAPPO | Local only | Linear | Online | ||
| SCA (Perfect CSI) | — | Full inst. CSI | Cubic | Offline | |
| Supervised DNN | Local only | Fixed | Static | ||
| Static Lawnmower | precompute | lookup | None | Trivially | None |
| SCMA-MADDPG [32] | Local only | Linear | Online |
7.2. Battery-Depletion Avoidance: Reward Shaping Versus Hard Constraints
7.3. Agricultural Field Realism
7.4. Discussion on Modeling Assumptions and Practical Limitations
7.4.1. Wireless Channel Model
7.4.2. Sensing Model Fidelity and Practical Limitations
- Altitude-Dependent Resolution (GSD): Ground Sampling Distance scales directly with UAV altitude . Flying near maximizes geometric coverage but coarsens spatial resolution, so fine-grained tasks such as early-stage weed identification or localized crop-stress classification demand a lower operating altitude—a direct trade-off between coverage speed and imaging quality that our model does not currently arbitrate.
- Image Overlap Requirements: Photogrammetric pipelines for orthomosaic generation typically require 70–80% frontal and lateral image overlap. Our coverage metric treats sensing as an instantaneous geometric footprint, whereas a deployable system would need to coordinate flight speed, camera trigger rate, and inter-path spacing to guarantee this overlap—a constraint that limits how abruptly the trajectory can change direction.
- Environmental and Atmospheric Factors: We assume clear-sky, static imaging conditions. In practice, wind-induced attitude oscillations (yaw, pitch, roll) cause gimbal jitter and motion blur, while cloud cover, shadowing, and diurnal solar angle shift illumination and introduce radiometric noise into the sensing pipeline.
- Probabilistic Sensing Uncertainty: Rather than the binary disc model in Equation (18), real sensor detection probability decays with distance from the camera’s nadir point rather than dropping sharply at the footprint boundary. Incorporating a probabilistic sensing kernel, together with robust optimization under wind-induced trajectory uncertainty, is a promising direction for future work.
7.5. Coverage and Depot-Return Gains
7.6. Analysis of Simulation-to-Reality (Sim-to-Real) Gaps
- Stochastic Wind Disturbances: Real-world agricultural environments are subject to unpredictable wind vectors and localized thermal updrafts. To maintain a stable hover or track the target 3D trajectory under wind resistance, the UAV’s flight controller must constantly apply rapid, high-frequency motor adjustments. This continuous attitude correction drastically scales up the aerodynamic power consumption relative to the idealized, smooth propulsion equations, accelerating battery depletion and prematurely forcing a return-to-depot sequence. This would directly contract the available mission horizon, likely degrading the 93% field coverage completeness threshold.
- GPS Localized Noise and Sensor Drift: Standard onboard Global Positioning System (GPS) receivers exhibit inherent positioning errors and multi-path fading, resulting in a localized noise envelope (typically to 3 m). Because our decentralized policy maps continuous spatial observations directly to actions, positional drift introduces state estimation errors. This error forces the UAVs into jagged, suboptimal trajectory corrections, causing overlapping camera footprints and slight misalignments with the cell-free APs. This spatial jitter would introduce a minor performance penalty across both objectives, shaving percentage points off the peak 92% communication efficiency due to suboptimal beamforming and uplink channel matching.
- Non-Linear Battery Discharge and Internal Resistance: Our model utilizes an idealized, linear energy consumption model where remaining capacity is proportional to power drawn over time. In real lithium-polymer (LiPo) cells, the discharge curve is highly non-linear, and it is characterized by a rapid voltage drop near the end of the discharge cycle and severe internal resistance losses (I2R heating) under peak current loads (e.g., during aggressive high-speed climbing maneuvers). Neglecting these electrochemical dynamics means our simulator under-represents the risk of voltage sags under low-charge states, which would force the safety-aware policy to adopt a more conservative return-to-depot margin in reality, slightly reducing overall mission utility.
8. Conclusions
- Heterogeneous Fleet Optimization and Asymmetric Dynamics: While the proposed framework focuses on homogeneous rotary-wing UAVs, large-scale agricultural monitoring may benefit from heterogeneous UAV swarms. An interesting extension is to support mixed fleets consisting of fixed-wing UAVs for long-range coverage and rotary-wing UAVs for high-resolution sensing tasks. Such a setting introduces heterogeneous flight dynamics, distinct energy consumption models, and a more complex joint action space. Future work will investigate hierarchical coordination and reward design strategies to enable efficient cooperation among different UAV types.
- Privacy-Preserving Federated Multi-Farm Learning: Another promising direction is the integration of Federated Learning (FL) to enable collaborative training across multiple farms while keeping local operational data private. In this setting, CF-mMIMO CPUs deployed at different farms can serve as local FL clients, performing policy updates using local data and sharing only model parameters or gradients with a central aggregation server. Future research will also explore differential privacy techniques to enhance data protection while maintaining learning performance.
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CF-mMIMO | Cell-Free massive Multiple-Input Multiple-Output |
| MARL | Multi-Agent Reinforcement Learning |
| MADRL | Multi-Agent Deep Reinforcement Learning |
| MAPPO | Multi-Agent Proximal Policy Optimization |
| CTDE | Centralized Training with Decentralized Execution |
| UAV | Unmanned Aerial Vehicle |
| RS | Remote Sensing |
| EE | Energy Efficiency |
| SE | Spectral Efficiency |
| APs | Access Points |
| SCA | Successive Convex Approximation |
| DNN | Deep Neural Network |
| CSI | Channel State Information |
| MDP | Markov Decision Process |
References
- Sharma, S.; Popli, R.; Singh, S.; Chhabra, G.; Saini, G.S.; Singh, M.; Sandhu, A.; Sharma, A.; Kumar, R. The Role of 6G Technologies in Advancing Smart City Applications: Opportunities and Challenges. Sustainability 2024, 16, 7039. [Google Scholar] [CrossRef]
- Li, P.; Fan, J.; Wu, J. Exploring the key technologies and applications of 6G wireless communication network. iScience 2025, 28, 112281. [Google Scholar] [CrossRef] [PubMed]
- Ozgen, R.E.; Yazar, A.; Osmanca, M.S. 6G Beyond Radio: From Connecting Devices to Sensing the World. IEEE Access 2026, 14, 51783–51799. [Google Scholar] [CrossRef]
- Fountas, S.; Espejo-Garcia, B.; Kasimati, A.; Gemtou, M.; Panoutsopoulos, H.; Anastasiou, E. Agriculture 5.0: Cutting-Edge Technologies, Trends, and Challenges. IT Prof. 2024, 26, 40–47. [Google Scholar] [CrossRef]
- Massaoudi, A.; Berguiga, A.; Harchay, A. Secure Irrigation System for Olive Orchards Using Internet of Things. Comput. Mater. Contin. 2022, 72, 4664–4672. [Google Scholar] [CrossRef]
- Naseer, A.; Shmoon, M.; Shakeel, T.; Ur Rehman, S.; Ahmad, A.; Gruhn, V. A Systematic Literature Review of the IoT in Agriculture—Global Adoption, Innovations, Security, and Privacy Challenges. IEEE Access 2024, 12, 60986–61021. [Google Scholar] [CrossRef]
- Reddy Maddikunta, P.K.; Hakak, S.; Alazab, M.; Bhattacharya, S.; Gadekallu, T.R.; Khan, W.Z.; Pham, Q.V. Unmanned Aerial Vehicles in Smart Agriculture: Applications, Requirements, and Challenges. IEEE Sens. J. 2021, 21, 17608–17619. [Google Scholar] [CrossRef]
- Guan, S.; Zhu, Z.; Wang, G. A Review on UAV-Based Remote Sensing Technologies for Construction and Civil Applications. Drones 2022, 6, 117. [Google Scholar] [CrossRef]
- Phang, S.K.; Chiang, T.H.A.; Happonen, A.; Chang, M.M.L. From Satellite to UAV-Based Remote Sensing: A Review on Precision Agriculture. IEEE Access 2023, 11, 127057–127076. [Google Scholar] [CrossRef]
- Massaoudi, A.; Berguiga, A.; Harchay, A.; Ben Ayed, M.; Belmabrouk, H. Spectral and Energy Efficiency Trade-Off in UAV-Based Olive Irrigation Systems. Appl. Sci. 2023, 13, 739. [Google Scholar] [CrossRef]
- Massaoudi, A.; Aydi, W. Optimized Power Allocation Algorithms for UAV-Based Agricultural Systems. IEEE Access 2025, 13, 103292–103308. [Google Scholar] [CrossRef]
- Qu, C.; Boubin, J.; Gafurov, D.; Zhou, J.; Aloysius, N.; Nguyen, H.; Calyam, P. UAV Swarms in Smart Agriculture: Experiences and Opportunities. In Proceedings of the 2022 IEEE 18th International Conference on e-Science (e-Science), Salt Lake City, UT, USA, 11–14 October 2022; pp. 148–158. [Google Scholar] [CrossRef]
- Kassam, J.; Castanheira, D.; Silva, A.; Dinis, R.; Gameiro, A. Cell-Free Massive MIMO Technology and Applications in 6G. In Massive MIMO for Future Wireless Communication Systems: Technology and Applications; John Wiley & Sons, Inc.: Hoboken, NJ, USA, 2025; pp. 95–122. [Google Scholar] [CrossRef]
- Elhoushy, S.; Ibrahim, M.; Hamouda, W. Cell-Free Massive MIMO: A Survey. IEEE Comm. Surv. Tutor. 2022, 24, 492–523. [Google Scholar] [CrossRef]
- Björnson, E.; Sanguinetti, L. Scalable Cell-Free Massive MIMO Systems. IEEE Trans. Commun. 2020, 68, 4247–4261. [Google Scholar] [CrossRef]
- Pan, X.; Zheng, Z.; Huang, X.; Fei, Z. On the Uplink Distributed Detection in UAV-Enabled Aerial Cell-Free mMIMO Systems. IEEE Trans. Wirel. Comm. 2024, 23, 13812–13825. [Google Scholar] [CrossRef]
- Bjornson, E.; Sanguinetti, L.; Wymeersch, H.; Hoydis, J.; Marzetta, T.L. Massive MIMO is a reality—What is next?: Five promising research directions for antenna arrays. Digit. Signal Process. 2019, 94, 3–20. [Google Scholar] [CrossRef]
- Braga, I.M.; Antonioli, R.P.; Fodor, G.; Silva, Y.C.B.; Freitas, W.C. Decentralized Joint Pilot and Data Power Control Based on Deep Reinforcement Learning for the Uplink of Cell-Free Systems. IEEE Trans. Veh. Technol. 2023, 72, 957–972. [Google Scholar] [CrossRef]
- Zhang, Y.; Zhang, J.; Buzzi, S.; Xiao, H.; Ai, B. Unsupervised Deep Learning for Power Control of Cell-Free Massive MIMO Systems. IEEE Trans. Veh. Technol. 2023, 72, 9585–9590. [Google Scholar] [CrossRef]
- Zaher, M.; Demir, O.T.; Bjornson, E.; Petrova, M. Learning-Based Downlink Power Allocation in Cell-Free Massive MIMO Systems. IEEE Trans. Wirel. Commun. 2023, 22, 174–188. [Google Scholar] [CrossRef]
- Zeng, Y.; Xu, J.; Zhang, R. Energy Minimization for Wireless Communication with Rotary-Wing UAV. IEEE Trans. Wirel. Commun. 2019, 18, 2329–2345. [Google Scholar] [CrossRef]
- Wu, Q.; Zeng, Y.; Zhang, R. Joint Trajectory and Communication Design for Multi-UAV Enabled Wireless Networks. IEEE Trans. Wirel. Commun. 2018, 17, 2109–2121. [Google Scholar] [CrossRef]
- Luong, N.C.; Hoang, D.T.; Gong, S.; Niyato, D.; Wang, P.; Liang, Y.C.; Kim, D.I. Applications of Deep Reinforcement Learning in Communications and Networking: A Survey. IEEE Commun. Surv. Tutor. 2019, 21, 3133–3174. [Google Scholar] [CrossRef]
- Shah, S.A.A.; Fernando, X.; Kashef, R. Joint Trajectory and Pilot Assignment Optimization for UAV Enabled Cell-Free Massive MIMO. In Proceedings of the 2025 IEEE International Conference on Communications Workshops (ICC Workshops), Montreal, QC, Canada, 8–12 June 2025; pp. 1876–1881. [Google Scholar] [CrossRef]
- Kenneth, O.; Mwangi, E.; Konditi, D.B.O. Optimizing UAV Location for Deployment in Cell-Free Massive MIMO Networks Using a Soft Actor-Critic Reinforcement Learning. IEEE Access 2025, 13, 139680–139695. [Google Scholar] [CrossRef]
- Huang, T.; Pan, H.; Sun, W.; Gao, H. Sine Resistance Network-Based Motion Planning Approach for Autonomous Electric Vehicles in Dynamic Environments. IEEE Trans. Transp. Electrif. 2022, 8, 2862–2873. [Google Scholar] [CrossRef]
- Zhang, S.; Zhang, H.; He, Q.; Bian, K.; Song, L. Joint Trajectory and Power Optimization for UAV Relay Networks. IEEE Commun. Lett. 2018, 22, 161–164. [Google Scholar] [CrossRef]
- Hur, J.; Lee, S.H. Joint Trajectory and Power Optimization for Energy-Efficient UAV Redeployment Against an Eavesdropper Under Consistent Fairness Constraint. IEEE Commun. Lett. 2024, 28, 2347–2351. [Google Scholar] [CrossRef]
- Wu, X.; Hou, C.; Meng, G.; Zhou, Z.; Liu, Q. Joint Trajectory and Power Optimization for Loosely Coupled Tasks: A Decoupled-Critic MAPPO Approach. Drones 2026, 10, 116. [Google Scholar] [CrossRef]
- Shah, S.A.A.; Fernando, X.N.; Kashef, R. Joint Optimization of UAV Trajectory, Transmit Power, and User Association in Aerial-Terrestrial Cell-Free Massive MIMO Network. IEEE Trans. Wirel. Commun. 2026, 25, 15818–15832. [Google Scholar] [CrossRef]
- Ma, X.; Li, J.; Nie, J.; Li, D.; Feng, W.; Jiang, W. SSDNN: A Self-Supervised DNN for Energy Efficiency Optimization in UAV CF-mMIMO Under URLLC. In Proceedings of the IEEE International Conference on Communications, Montreal, QC, Canada, 8–12 June 2025; pp. 5047–5052. [Google Scholar] [CrossRef]
- Liu, Z.; Zhang, J.; Zeng, Y.; Ai, B. Energy-Efficient Multi-Agent Reinforcement Learning for UAV Trajectory Optimization in Cell-Free Massive MIMO Networks. IEEE Trans. Wirel. Commun. 2025, 24, 5917–5930. [Google Scholar] [CrossRef]
- Lee, J.; Ko, H. Joint Optimization on Trajectory, Data Relay, and Wireless Power Transfer in UAV-Based Environmental Monitoring System. Electronics 2024, 13, 828. [Google Scholar] [CrossRef]
- Saraiva, J.V.; Antonioli, R.P.; Fodor, G.; Freitas, W.C.; Silva, Y.C.B. Integrating Aerial and Ground Users: Deep Reinforcement Learning and Convex Optimization for Data Power Control in Cell-Free Networks. IEEE Trans. Veh. Technol. 2026, 75, 12904–12920. [Google Scholar] [CrossRef]
- Bianchi, D.; Borri, A.; Di Gennaro, S.; Preziuso, M. UAV trajectory control with rule-based minimum-energy reference generation. In Proceedings of the 2022 European Control Conference (ECC), London, UK, 12–15 July 2022; pp. 1497–1502. [Google Scholar] [CrossRef]
- Zhang, L.; Li, L.; Wei, W.; Song, H.; Yang, Y.; Liang, J. Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement Learning. In Proceedings of the Advances in Neural Information Processing Systems; Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2024; Volume 37, pp. 138698–138730. [Google Scholar] [CrossRef]
- Zhang, C.; Li, Y.; Liu, Y.; Huang, Y.; Zhang, Z. TTBO-FL: Joint Training, Trajectory, and Beamforming Optimization for Energy-Efficient Federated Learning in UAV Swarm. IEEE Internet Things J. 2026, 13, 2206–2220. [Google Scholar] [CrossRef]
- Mu, L.; Duan, T.; Wu, C.; Bai, B. Attention-Biased Reinforcement Learning Framework for Adaptive and Scalable Flocking of UAV Swarms. IEEE Trans. Autom. Sci. Eng. 2026, 23, 756–771. [Google Scholar] [CrossRef]
- Dong, L.; Kong, H.W.; Yuan, X. Reinforcement Learning-Based Spectral Performance Optimization for UAV-Assisted MIMO Communication System. IEEE/CAA J. Autom. Sin. 2025, 12, 1283–1285. [Google Scholar] [CrossRef]
- Ao, T.; Li, H.; Zhang, K.; Shi, H.; Shi, L.; Liu, F.; Zhou, Y. Heterogeneous UAVs Trajectory Optimization for Post-Disaster Target Search Based on MARL with Graph Attention Network. IEEE Trans. Veh. Technol. 2026, 75, 1412–1426. [Google Scholar] [CrossRef]
- Han, S.; Zhu, K.; Zhou, M.; Liu, X. Joint Deployment Optimization and Flight Trajectory Planning for UAV Assisted IoT Data Collection: A Bilevel Optimization Approach. IEEE Trans. Intell. Transp. Syst. 2022, 23, 21492–21504. [Google Scholar] [CrossRef]
- Al-Hourani, A.; Kandeepan, S.; Lardner, S. Optimal LAP Altitude for Maximum Coverage. IEEE Wirel. Commun. Lett. 2014, 3, 569–572. [Google Scholar] [CrossRef]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-agent actor-critic for mixed cooperative-competitive environments. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS’17); Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 6382–6393. [Google Scholar]








| Feature | [29] | [32] | [34] | This Work |
|---|---|---|---|---|
| CF-mMIMO | × | ✓ | ✓ | ✓ |
| UAV trajectory opt. | ✓ | ✓ | × | ✓ |
| Power control | ✓ | × | ✓ | ✓ |
| Rotary-wing model | × | × | × | ✓ |
| Battery hard constr. | × | × | × | ✓ |
| Return-to-depot | × | × | × | ✓ |
| Sensing coverage | × | × | × | ✓ |
| Joint SC utility | × | × | × | ✓ |
| CTDE | × | ✓ * | × | ✓ |
| Reference | Domain | Learning Paradigm | Joint Action Space | Energy Model | Scales to Larger Swarms? |
|---|---|---|---|---|---|
| Zhang et al. (TTBO-FL) [37] | Federated learning over UAV relay swarm | Soft Actor–Critic (single-agent) | Trajectory + local epochs + 3D beamforming | UAV propulsion (unspecified detail) | ✗ |
| Mu et al. (ABDDPG) [38] | Flocking/obstacle avoidance | MADDPG + attention-biased Transformer, LoRA fine-tuning | 3D trajectory (velocity) | Not modeled | ✓ * |
| Dong et al. [39] | Single UAV-assisted MIMO relay | RL-tuned PSO (hybrid, single-agent) | Trajectory + beamforming/power | Not modeled | N/A (single UAV) |
| Ao et al. (GATAC) [40] | Heterogeneous post-disaster target search | MARL + Graph Attention Network actor–critic | Trajectory (role-specific: fixed-wing leader/rotor follower) | Multi-rotor + fixed-wing propulsion | ✓ |
| Han et al. (bi-level) [41] | Static IoT data collection | Dandelion algorithm + iterated greedy (meta-heuristic, non-RL) | Deployment (footholds) + trajectory | Communication + flight energy | N/A (offline, re-solved per instance) |
| This paper | CF-mMIMO sensing–communication, precision agriculture | CTDE-MAPPO (on-policy, multi-agent) | 3D trajectory + transmit power | Rotary-wing propulsion [21] + Tx power | ✓ † |
| Parameter | Symbol | Value |
|---|---|---|
| Field dimensions | 500 m × 500 m | |
| Number of APs | M | 16 |
| Number of UAVs | K | 4 |
| Carrier frequency | 2.4 GHz | |
| Noise power | dBm | |
| Max UAV power | 200 mW | |
| UAV battery | 50 kJ | |
| Max UAV speed | 15 m/s | |
| Altitude range | m | |
| Min separation | 10 m | |
| Mission horizon | T | 200 s |
| Min rate | 1 bit/s/Hz | |
| MAPPO clip | 0.2 | |
| Discount factor | 0.99 | |
| GAE parameter | 0.95 | |
| Training episodes | 5000 |
| Variant | U | Depot | ||
|---|---|---|---|---|
| EE only | 1 | 0.418 ± 0.028 | 0.68 ± 0.042 | 0.88 ± 0.058 |
| Coverage only | 0 | 0.391 ± 0.024 | 0.95 ± 0.015 | 0.91 ± 0.061 |
| w/o battery term () | 0.6 | 0.431 ± 0.033 | 0.89 ± 0.026 | 0.38 ± 0.047 |
| w/o mission time | 0.6 | 0.428 ± 0.029 | 0.91 ± 0.022 | 0.32 ± 0.053 |
| w/o shared critic (indep. PPO) | 0.6 | 0.445 ± 0.035 | 0.90 ± 0.031 | 0.85 ± 0.049 |
| Proposed: Full model | 0.6 | 0.531 ± 0.021 | 0.93 ± 0.018 | 0.92 ± 0.031 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Massaoudi, A.; Aydi, W. Joint 3D Trajectory and Power Optimization for UAV Swarms in Cell-Free Massive MIMO Networks: A CTDE-MAPPO Framework for Sensing-Aware Precision Agriculture. Drones 2026, 10, 576. https://doi.org/10.3390/drones10080576
Massaoudi A, Aydi W. Joint 3D Trajectory and Power Optimization for UAV Swarms in Cell-Free Massive MIMO Networks: A CTDE-MAPPO Framework for Sensing-Aware Precision Agriculture. Drones. 2026; 10(8):576. https://doi.org/10.3390/drones10080576
Chicago/Turabian StyleMassaoudi, Ayman, and Walid Aydi. 2026. "Joint 3D Trajectory and Power Optimization for UAV Swarms in Cell-Free Massive MIMO Networks: A CTDE-MAPPO Framework for Sensing-Aware Precision Agriculture" Drones 10, no. 8: 576. https://doi.org/10.3390/drones10080576
APA StyleMassaoudi, A., & Aydi, W. (2026). Joint 3D Trajectory and Power Optimization for UAV Swarms in Cell-Free Massive MIMO Networks: A CTDE-MAPPO Framework for Sensing-Aware Precision Agriculture. Drones, 10(8), 576. https://doi.org/10.3390/drones10080576

