Adaptive Task Scheduling for Edge-Intelligent Systems: An Online Sleeping Restless Bandits Framework
Abstract
1. Introduction
- (1)
- Application 1: Task Offloading in Wireless Edge Networks. The decision maker (edge server) needs to select which end devices to offload tasks to or retrieve data from. Each device exhibits restless behavior: its internal state (e.g., buffer level, channel quality, energy reserve) autonomously evolves over time [18,19,20]. However, due to environmental interference or energy depletion, the availability of each device follows an unknown distribution, resulting in a dynamically changing set of selectable devices. Consequently, system transitions heavily depend on which devices are available in each round, fundamentally altering the policy optimization.
- (2)
- Application 2: Drone-Assisted Data Collection. In smart city or agricultural IoT scenarios, drones serve as mobile edge nodes to collect data from distributed sensors. Each sensor’s state (e.g., data freshness, storage capacity) changes continuously. Due to the drone’s mobility and varying weather conditions, a sensor’s availability to establish a reliable communication link shifts over time. This dynamic availability implies that state transitions and data collection efficiency depend not only on the chosen scheduling action but also on the real-time availability of the communication links, complicating the design of symmetric data distribution and scheduling policies.
- We are the first to investigate the open challenge of the online sleeping Restless Multi-Armed Bandit (RMAB) problem, which demands addressing the complex interplay between arm availability and restless Markovian reward optimization.
- Given the transition dynamics, reward distributions, and availability, we design a sleeping index policy (SIP) as an oracle for the online problem by constructing a fluid process.
- We propose the Minimum Exploration Guarantee (MEG) mechanism that enables efficient exploration for online algorithms in dynamic environments (In this paper, dynamic environments refer to the time-varying available set induced by stochastic arm availability, together with the restless evolution of arm states under stationary transition kernels and fixed availability probabilities, rather than non-stationary model parameters drifting over time ). We employ Laplace analysis to describe the characteristic function of the exploration rounds, and apply Taylor approximation to prove their convergence to the Gumbel distribution under the MEG mechanism. In conjunction with the online SIP derived from the modified LP solution, we present a low-complexity online algorithm OSILA with an regret guarantee.
- Based on real-world applications, we constructed problem instances for experimental validation. The results confirm that SIP is asymptotically optimal in the offline setting. In the online scenario, the regret performance of OSILA outperforms other baseline algorithms.
2. Related Work
3. Model and Problem Statement
3.1. Setup: Sleeping RMAB
3.2. Objective and Online Setting
4. Online Learning Algorithm Design
4.1. Oracle: Optimal Sleeping Index-Based Policy
4.2. Online Learning Algorithm: OSILA
4.2.1. Minimum Exploration Guarantee Mechanism
| Algorithm 1 Minimum Exploration Guarantee Mechanism |
|
4.2.2. Modified LP-Based Exploitation Mechanism
4.2.3. Main Idea of OSILA
| Algorithm 2 Online Sleeping Index-aware Learning Algorithm (OSILA) |
|
5. Performance Analysis
5.1. Main Results and Discussions
5.2. Proof Sketch of Theorem 2
6. Experiments
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Appendix A. Main Variables
| Variable | Description |
|---|---|
| K | Number of arms |
| T | Total number of time steps |
| State space | |
| Action space | |
| M | Budget constraint on the number of active arms |
| Availability probability of arm i | |
| Set of available arms at time t | |
| Expected reward for arm i in state with action | |
| Transition probability from to under action | |
| Policy | |
| Long-term average reward under policy | |
| Probability of arm i being in state and taking action at time t | |
| Number of exploration rounds per arm | |
| Estimated reward, transition probability, and availability probability |
Appendix B. Equivalent and Modified Linear Programming Problems
Appendix C. Details of the Numerical Case-Study

References
- Ji, C.; Wu, F.; Zhu, Z.; Chang, L.P.; Liu, H.; Zhai, W. Memory-efficient deep learning inference with incremental weight loading and data layout reorganization on edge systems. J. Syst. Archit. 2021, 118, 102183. [Google Scholar] [CrossRef] [Scilit]
- Ji, C.; Pan, R.; Chang, L.P.; Shi, L.; Zhu, Z.; Liang, Y.; Kuo, T.W.; Xue, C.J. Inspection and Characterization of App File Usage in Mobile Devices. ACM Trans. Storage 2020, 16, 1–25. [Google Scholar] [CrossRef] [Scilit]
- Raeisi-Varzaneh, M.; Dakkak, O.; Habbal, A.; Kim, B.S. Resource scheduling in edge computing: Architecture, taxonomy, open issues and future research directions. IEEE Access 2023, 11, 25329–25350. [Google Scholar] [CrossRef] [Scilit]
- Avan, A.; Azim, A.; Mahmoud, Q.H. A state-of-the-art review of task scheduling for edge computing: A delay-sensitive application perspective. Electronics 2023, 12, 2599. [Google Scholar] [CrossRef] [Scilit]
- Fan, W.; Liu, X.; Yuan, H.; Li, N.; Liu, Y. Time-slotted task offloading and resource allocation for cloud-edge-end cooperative computing networks. IEEE Trans. Mob. Comput. 2024, 23, 8225–8241. [Google Scholar] [CrossRef] [Scilit]
- Xu, C.; Guo, J.; Li, Y.; Zou, H.; Jia, W.; Wang, T. Dynamic parallel multi-server selection and allocation in collaborative edge computing. IEEE Trans. Mob. Comput. 2024, 23, 10523–10537. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Xiong, B.; Chen, X.; Min, G.; Li, J. Joint computation offloading and resource allocation in multi-edge smart communities with personalized federated deep reinforcement learning. IEEE Trans. Mob. Comput. 2024, 23, 11604–11619. [Google Scholar] [CrossRef] [Scilit]
- Whittle, P. Restless bandits: Activity allocation in a changing world. J. Appl. Probab. 1988, 25, 287–298. [Google Scholar] [CrossRef] [Scilit]
- Kadota, I.; Sinha, A.; Uysal-Biyikoglu, E.; Singh, R.; Modiano, E. Scheduling policies for minimizing age of information in broadcast wireless networks. IEEE/ACM Trans. Netw. 2018, 26, 2637–2650. [Google Scholar] [CrossRef] [Scilit]
- Maatouk, A.; Kriouile, S.; Assad, M.; Ephremides, A. On the optimality of the Whittle’s index policy for minimizing the age of information. IEEE Trans. Wirel. Commun. 2020, 20, 1263–1277. [Google Scholar] [CrossRef] [Scilit]
- Xiong, G.; Wang, S.; Yan, G.; Li, J. Reinforcement learning for dynamic dimensioning of cloud caches: A restless bandit approach. IEEE/ACM Trans. Netw. 2023, 31, 2147–2161. [Google Scholar] [CrossRef] [Scilit]
- Sun, S.; Wu, W.; Fu, C.; Qiu, X.; Luo, J.; Wang, J. AoI Optimization in Multi-source Update Network Systems under Stochastic Energy Harvesting Model. IEEE J. Sel. Areas Commun. 2024, 42, 3172–3187. [Google Scholar] [CrossRef] [Scilit]
- Mate, A.; Perrault, A.; Tambe, M. Risk-Aware Interventions in Public Health: Planning with Restless Multi-Armed Bandits. In Proceedings of the AAMAS, Virtual, 3–7 May 2021; pp. 880–888. [Google Scholar]
- Ortner, R.; Ryabko, D.; Auer, P.; Munos, R. Regret bounds for restless markov bandits. In International Conference on Algorithmic Learning Theory; Springer: Berlin/Heidelberg, Germany, 2012; pp. 214–228. [Google Scholar]
- Wang, S.; Huang, L.; Lui, J. Restless-UCB, an efficient and low-complexity algorithm for online restless bandits. Adv. Neural Inf. Process. Syst. 2020, 33, 11878–11889. [Google Scholar]
- Wang, K.; Xu, L.; Taneja, A.; Tambe, M. Optimistic whittle index policy: Online learning for restless bandits. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 10131–10139. [Google Scholar]
- Xiong, G.; Li, J. Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback. In Proceedings of the Forty-First International Conference on Machine Learning, Vienna, Austria, 21–27 July 2024. [Google Scholar]
- Kadota, I.; Modiano, E. Minimizing the Age of Information in Wireless Networks with Stochastic Arrivals. IEEE Trans. Mob. Comput. 2021, 20, 1173–1185. [Google Scholar] [CrossRef]
- Hatami, M.; Codreanu, M. On the age-optimality of relax-then-truncate approach under partial battery knowledge in energy harvesting IoT networks. In 2023 21st International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks (WiOpt); IEEE: New York, NY, USA, 2023; pp. 589–596. [Google Scholar]
- Chen, G.; Liew, S.C.; Shao, Y. Uncertainty-of-information scheduling: A restless multiarmed bandit framework. IEEE Trans. Inf. Theory 2022, 68, 6151–6173. [Google Scholar] [CrossRef] [Scilit]
- Kleinberg, R.; Niculescu-Mizil, A.; Sharma, Y. Regret bounds for sleeping experts and bandits. Mach. Learn. 2010, 80, 245–272. [Google Scholar] [CrossRef] [Scilit]
- Saha, A.; Gaillard, P.; Valko, M. Improved sleeping bandits with stochastic action sets and adversarial rewards. In International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2020; pp. 8357–8366. [Google Scholar]
- Nguyen, Q.M.; Mehta, N. Near-optimal per-action regret bounds for sleeping bandits. In International Conference on Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2024; pp. 2827–2835. [Google Scholar]
- Verloop, I. Asymptotically optimal priority policies for indexable and nonindexable restless bandits. Ann. Appl. Probab. 2016, 26, 1947–1995. [Google Scholar] [CrossRef] [Scilit]
- Xiong, G.; Wang, S.; Li, J. Learning infinite-horizon average-reward restless multi-action bandits via index awareness. Adv. Neural Inf. Process. Syst. 2022, 35, 17911–17925. [Google Scholar]
- Blum, A. Empirical support for winnow and weighted-majority algorithms: Results on a calendar scheduling domain. Mach. Learn. 1997, 26, 5–23. [Google Scholar] [CrossRef] [Scilit]
- Freund, Y.; Schapire, R.E.; Singer, Y.; Warmuth, M.K. Using and combining predictors that specialize. In Proceedings of the Twenty-Ninth Annual ACM Symposium on Theory of Computing, El Paso, TX, USA, 4–6 May 1997; pp. 334–343. [Google Scholar]
- Blum, A.; Mansour, Y. From external to internal regret. J. Mach. Learn. Res. 2007, 8, 1307–1324. [Google Scholar]
- Kanade, V.; Steinke, T. Learning hurdles for sleeping experts. ACM Trans. Comput. Theory (TOCT) 2014, 6, 1–16. [Google Scholar] [CrossRef] [Scilit]
- Kanade, V.; McMahan, H.B.; Bryan, B. Sleeping experts and bandits with stochastic action availability and adversarial rewards. In Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2009; pp. 272–279. [Google Scholar]
- Gaillard, P.; Saha, A.; Dan, S. One arrow, two kills: A unified framework for achieving optimal regret guarantees in sleeping bandits. In International Conference on Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2023; pp. 7755–7773. [Google Scholar]
- Neu, G.; Valko, M. Online combinatorial optimization with stochastic decision sets and adversarial losses. Adv. Neural Inf. Process. Syst. 2014, 27, 2780–2788. [Google Scholar]
- Kale, S.; Lee, C.; Pál, D. Hardness of online sleeping combinatorial optimization problems. Adv. Neural Inf. Process. Syst. 2016, 29, 2189–2197. [Google Scholar]
- Li, F.; Liu, J.; Ji, B. Combinatorial sleeping bandits with fairness constraints. IEEE Trans. Netw. Sci. Eng. 2019, 7, 1799–1813. [Google Scholar] [CrossRef] [Scilit]
- Papadimitriou, C.H.; Tsitsiklis, J.N. The complexity of optimal queuing network control. Math. Oper. Res. 1999, 24, 293–305. [Google Scholar] [CrossRef] [Scilit]
- Weber, R.R.; Weiss, G. On an index policy for restless bandits. J. Appl. Probab. 1990, 27, 637–648. [Google Scholar] [CrossRef] [Scilit]
- Akbarzadeh, N.; Mahajan, A. Restless bandits with controlled restarts: Indexability and computation of Whittle index. In 2019 IEEE 58th Conference on Decision and Control (CDC); IEEE: New York, NY, USA, 2019; pp. 7294–7300. [Google Scholar]
- Xiong, G.; Li, J.; Singh, R. Reinforcement learning augmented asymptotically optimal index policy for finite-horizon restless bandits. In Proceedings of the AAAI Conference on Artificial Intelligence, Online, 22 February–1 March 2022; Volume 36, pp. 8726–8734. [Google Scholar]
- Jiang, B.; Jiang, B.; Li, J.; Lin, T.; Wang, X.; Zhou, C. Online restless bandits with unobserved states. In International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2023; pp. 15041–15066. [Google Scholar]
- Hatami, M.; Leinonen, M.; Chen, Z.; Pappas, N.; Codreanu, M. On-demand AoI minimization in resource-constrained cache-enabled IoT networks with energy harvesting sensors. IEEE Trans. Commun. 2022, 70, 7446–7463. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Xiong, G.; Li, J. Online restless multi-armed bandits with long-term fairness constraints. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–28 February 2024; Volume 38, pp. 15616–15624. [Google Scholar]
- Puterman, M.L. Markov Decision Processes: Discrete Stochastic Dynamic Programming; John Wiley & Sons: Hoboken, NJ, USA, 2014. [Google Scholar]
- Ethier, S.N.; Kurtz, T.G. Markov Processes: Characterization and Convergence; John Wiley & Sons: Hoboken, NJ, USA, 2009. [Google Scholar]
- Gast, N.; Bruno, G. A mean field model of work stealing in large-scale systems. ACM SIGMETRICS Perform. Eval. Rev. 2010, 38, 13–24. [Google Scholar] [CrossRef] [Scilit]
- Besbes, O.; Gur, Y.; Zeevi, A. Stochastic multi-armed-bandit problem with non-stationary rewards. Adv. Neural Inf. Process. Syst. 2014, 27, 199–207. [Google Scholar]
- Cheung, W.C.; Simchi-Levi, D.; Zhu, H. Learning to optimize under non-stationarity. In The 22nd International Conference on Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2019; pp. 1079–1087. [Google Scholar]
- Ortner, R.; Gajane, P.; Auer, P. Variational regret bounds for reinforcement learning. In Uncertainty in Artificial Intelligence; PMLR: Cambridge, MA, USA, 2020; pp. 81–90. [Google Scholar]
- Hoeffding, W. Probability inequalities for sums of bounded random variables. In The Collected Works of Wassily Hoeffding; Springer: New York, NY, USA, 1994; pp. 409–426. [Google Scholar]
- Herlihy, C.; Prins, A.; Srinivasan, A.; Dickerson, J.P. Planning to fairly allocate: Probabilistic fairness in the restless bandit setting. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Long Beach, CA, USA, 6–10 August 2023; pp. 732–740. [Google Scholar]
- Lattimore, T.; Szepesvári, C. Bandit Algorithms; Cambridge University Press: Cambridge, UK, 2020. [Google Scholar]
- Gurobi Optimization, LLC. Gurobi Optimizer Reference Manual; Gurobi Optimization, LLC: Beaverton, OR, USA, 2024. [Google Scholar]
- Sun, Y.; Uysal-Biyikoglu, E.; Yates, R.D.; Koksal, C.E.; Shroff, N.B. Update or wait: How to keep your data fresh. IEEE Trans. Inf. Theory 2017, 63, 7492–7508. [Google Scholar] [CrossRef] [Scilit]



Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Sun, S.; Fu, C.; Xu, Y.; Wu, W. Adaptive Task Scheduling for Edge-Intelligent Systems: An Online Sleeping Restless Bandits Framework. Symmetry 2026, 18, 951. https://doi.org/10.3390/sym18060951
Sun S, Fu C, Xu Y, Wu W. Adaptive Task Scheduling for Edge-Intelligent Systems: An Online Sleeping Restless Bandits Framework. Symmetry. 2026; 18(6):951. https://doi.org/10.3390/sym18060951
Chicago/Turabian StyleSun, Sujunjie, Chenchen Fu, Yuhang Xu, and Weiwei Wu. 2026. "Adaptive Task Scheduling for Edge-Intelligent Systems: An Online Sleeping Restless Bandits Framework" Symmetry 18, no. 6: 951. https://doi.org/10.3390/sym18060951
APA StyleSun, S., Fu, C., Xu, Y., & Wu, W. (2026). Adaptive Task Scheduling for Edge-Intelligent Systems: An Online Sleeping Restless Bandits Framework. Symmetry, 18(6), 951. https://doi.org/10.3390/sym18060951

