Energy-Efficient Access Point Switch On/Off in Cell-Free Massive MIMO Using Proximal Policy Optimization
Round 1
Reviewer 1 Report
Comments and Suggestions for Authors1. It would be beneficial to expand the Related Work section by incorporating additional studies on cell-free massive MIMO systems. This would improve the comprehensiveness of the manuscript and provide readers with a broader contextual understanding of the considered scenario.
2. Although the spatial correlation matrix is adopted from prior work without modification, it would enhance readability and clarity to explicitly present the corresponding mathematical expressions within the manuscript.
3. While the manuscript emphasizes energy efficiency and scalability, the power consumption parameters are fixed as shown in Table 2. Although these parameters are referenced from prior studies, concerns remain regarding the suitability and generalizability of the adopted power model for the considered system. In practical deployments, AP on/off decisions may interact with dynamic power control mechanisms to enhance overall network throughput. Furthermore, energy consumption and throughput trends may vary depending on network density and deployment scale. In this context, the authors are encouraged to clarify the practical meaning, scope, and insights of the manuscript under the considered modeling assumptions and scenario constraints.
4. In a cell-free network architecture, it remains unclear which network entity is responsible for solving the formulated optimization problem. The manuscript should clarify the architectural assumption regarding the decision-making authority and computational entity responsible for AP activation.
5. It would be beneficial for the authors to provide a clearer justification for the selection of the PPO-based approach in this study. Also, explicitly clarifying the performance differences compared to alternative RL-based baselines would be beneficial.
6. The reviewer has concerns regarding the practical relevance of the adopted simulation environment. It would strengthen the manuscript to provide additional analysis demonstrating how network performance varies under different deployment conditions, thereby offering clearer insights into the practical implications and generalizability of the proposed approach.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThe paper is written in a proper academic English, with accurate technical terminology in a formal tone. Some phrases are excessively repeated, such as those connected to energy efficiency, large-scale fading and combinatorial features. Some sentences need to be splitted into parts, as the one in the introduction,Section 1.2 from line 100, and line 395 in Section 4.1. These are really hard to follow.
The reviewed paper presents a study on energy-efficient access point switch-on/off strategies in cell-free massive MIMO networks using Proximal Policy Optimization (abbr. PPO), a reinforcement learning (abbr. RL) algorithm. Key algorithms and presented solutions, such as the PPO-based policy with per-AP state features seem to be novel in combination. PPO is also applied to ASO. A proper review of prior RL-related papers on ASO is given in the introductory part of the aper, stating different features. The paper is a methodological advancement in evaluation scale/power modeling over the prior research.
​
The major novelty lies in PPO policy learninig with attention to AP activation probabilities from compact state, what enables both scalabilty, and enables the authors to present a realistic downlink model, sweeping the configurations of various baselines. On the contrary,the authors present a downlink-only solution, with fixed pilots, thus with no attention to mobility issues. A methodological rigor over algorithmic invention is highly noticeable, with such contributions as EE-explicit formulation, scalable RL approach, in-depth benchmark presentation, and distribution analysis. A mature PPO approach is presented to a current problem.
However, the novelty in PPO is rather incremental, presenting downlink only, with simple baselines (why is e.g. RL-SAC omitted)? Why have you decided to include the greedy approach, since it is rather a heavy approach? Would it be possible to include sensitivity analysis where N/K varies, and pilot is placed eslwhere? In the future, I would see a joint up- and downlink ASO with ful duplex, including dynamic scenarios (mobility!). Add some episodic large-scale evolutions,use RL training of PPO with LSTMs, so much poplar now.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for AuthorsFirst of all, thank you for providing me a chance to review your paper. I have found some flaws in your paper. Those flaws are provided below in the form of pointers
- Please provide achieved results in the form of percentage in the end part of your abstract.
- Do not use abbreviations in the keywords section.
- The related works should be a separate section.
- Provide key contribution of your in the end paragraph of the introduction section in the form of pointers.
- After the key contributions provide the organization of your work.
- At the end of the related works provide a paragraph that should enclose the knowledge on the fact that what shortcoming have you found out in the present literature that gas motivated you to do this specific work.
- The way you have structured your article is not quite right, take guidance from some experienced person in your group so that he or she may help you in structuring your article.
- Provide the flow diagram of your work before the results.
- Provide the pseudo code of your work so that if some reader wants to recreate the work for their own analysis, then they may can do so.
- In your article the policy models that you have opted, each AP activation independently (Bernoulli outputs). This ignores potential correlation structures between APs, which may limit the ability to capture cooperative activation patterns.
- Please note that the absence of Graph Neural Networks (GNNs) or permutation-invariant architectures is a methodological limitation. How can only the cell-free massive MIMO naturally form a graph (AP–UE bipartite structure).
- In your study you have restricted the analysis to downlink operation, a little bit more discussion is required on this point because.
- This have the capacity to limit applicability in realistic TDD systems where uplink/downlink trade-offs influence AP activation.
- Your power consumption model is too simplified, as you are not accounting for the following.
- Dynamic hardware scaling.
- Load dependent circuit efficiencies.
- Hardware nonlinearities.
I am asking this because it has the capacity to definitely impacts the practical validly of your system.
- I would suggest to compare against practical heuristic policies that should be based on large-scale fading ranking or clustering, which could be a strong non-learning baseline.
- To access the scalability, I think you should test for ultra dense deployments, that may be 200+ Aps.
- Why have you avoided to choose for over densification and also sparse deployments.
- Using only means, max, and min of large-scale fading per AP compresses significant information, may possibly discard the user distribution structure.
- Please provide sensitivity analysis with respect to, power amplifier efficiency, shadow fading variance, traffic dependent co-efficient.
- You have chosen a 0.5 fixed threshold for the AP activation use, however no analysis had been provided weather this threshold is optimal or not.
- Please provide some discussion on the fact that why have you chosen to train the policy offline under the strategic large scale fading per episode.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for Authors1. Although the authors state that the adopted power consumption model follows a representative configuration used in the literature, the current formulation relies on fixed parameter values. In practical deployments, power consumption characteristics may vary depending on factors such as hardware implementation, traffic load, network density, and operational mechanisms, including adaptive power control. Therefore, the reviewer remains concerned about the extent to which the reported energy-efficiency results remain practically meaningful under these simplified and static modeling assumptions. It would be beneficial if the authors could further elaborate on the practical implications and applicability of the obtained results in realistic deployment environments.
2. Is there a particular reason why the proposed method is not directly compared with the reinforcement learning–based approaches summarized in Table 1?
3. It would also be beneficial for the authors to discuss how the network performance may vary under different traffic patterns or user distribution conditions.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Round 3
Reviewer 1 Report
Comments and Suggestions for AuthorsThe authors have carefully addressed the reviewer's comments in the response letter.
The reviewer has no further comments or questions.

