RL-PMO: A Reinforcement Learning-Based Optimization Algorithm for Parallel SFC Migration
Abstract
1. Introduction
- Modeling the migration of all VNFs on the failed node as a single-step multi-objective optimization problem enables the system to maintain a high migration success rate while keeping the overall migration delay and resource overhead within acceptable limits.
- This paper designs an improved CSSA–PSO hybrid heuristic method to generate diverse and feasible migration trajectories under various failure patterns and load levels.
- To capture the dependencies among VNFs, network resources, and migration constraints, we adopt Decision Mamba as the policy network. In addition, we incorporate a twin-critic architecture together with a Conservative Q-Learning (CQL) regularization term to mitigate the inherent overestimation and distribution shift issues in offline reinforcement learning.
2. Related Work
3. System Model and Problem Description
3.1. Network Model
3.2. SFC and VNF
3.3. SFC Migration
3.4. SFC Migration Constraints Under Failure Scenarios
3.4.1. Complete Migration Constraint for VNFs on the Faulty Node
3.4.2. Resource Constraints
- (a)
- CPU Resource Constraint:
- (b)
- Memory Resource Constraint:
- (c)
- Link Bandwidth Constraint
3.4.3. End-to-End Delay Constraint
3.5. Evaluation Metrics for SFC Migration
3.5.1. Migration Success Rate
3.5.2. Migration Time
3.5.3. Average Migration Cost
Migration Cost
Rerouting Cost
Total Migration Cost
4. Algorithm Design
4.1. Offline Dataset Generation
4.2. Markov Decision Process (MDP) Modeling
- State Space
- Action Space
- State Transition Probability
- Reward Function
4.3. Offline Reinforcement Learning Policy
| Algorithm 1: Reinforcement Learning-Driven Parallel Migration Optimization Algorithm | |
| Inputs: Offline data offline dataset ; number of epochs ; number of training steps ; soft-update coefficient ; CQL factor ; batch size ; update frequency ; learning rate | |
| Output: SFC migration policy | |
| Initialize Critic1, Critic2 and Actor (DM) parameters as Initialize target Critic parameters , | |
| 1 | for epoch = 1 to E do |
| 2 | for i = 1 to T do |
| 3 | Sample a mini-batch from |
| 4 | with no_grad: |
| 5 | Generates policy actions . |
| 6 | The main Critic networks predict Q-values: Q1, Q2. |
| 7 | Calculate Critic loss using Equations (23)–(25) |
| 8 | Update main Critic parameters by gradient descent. |
| 9 | if i % d == 0 then |
| 10 | Generate the policy action in the current state: |
| 11 | Calculate actor loss by Equation (26) |
| 12 | Update actor parameters by Equation (24) |
| 13 | end if |
| 14 | Soft-update target critics by Equation (28): |
| 15 | end for |
| 16 | end for |
| 17 | return |
5. Performance Evaluation
5.1. Simulation Setup
Network Architecture and Hyperparameter Configuration
5.2. Simulation Result and Analysis
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Bhamare, D.; Jain, R.; Samaka, M.; Erbad, A. A Survey on Service Function Chaining. J. Netw. Comput. Appl. 2016, 75, 138–155. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Ye, Q.; Zhou, P.; Li, B.; Wang, X.; Zhang, L. Research on fault prediction of computer network nodes driven by log information. Telecommun. Sci. 2024, 40, 11–22. [Google Scholar] [CrossRef]
- Zhai, D.; Meng, X.; Yu, Z.; Hu, H.; Liang, Y. A Migration Method for Service Function Chain Based on Failure Prediction. Comput. Netw. 2023, 222, 109554. [Google Scholar] [CrossRef] [Scilit]
- Kabdjou, J.; Shinomiya, N. A Method for Service Function Chain Migration Based on Server Failure Prediction in Mobile Edge Computing Environment. IEEE Access 2025, 13, 78664–78678. [Google Scholar] [CrossRef] [Scilit]
- Tang, L.; He, X.; Zhao, P.; Zhao, G.; Zhou, Y.; Chen, Q. Virtual Network Function Migration Based on Dynamic Resource Requirements Prediction. IEEE Access 2019, 7, 112348–112362. [Google Scholar] [CrossRef] [Scilit]
- Cho, D.; Taheri, J.; Zomaya, A.Y.; Bouvry, P. Real-Time Virtual Network Function (VNF) Migration toward Low Network Latency in Cloud Environments. In Proceedings of the 2017 IEEE 10th International Conference on Cloud Computing (CLOUD), Honolulu, HI, USA, 25–30 June 2017. [Google Scholar]
- Xia, J.; Cai, Z.; Xu, M. Optimized Virtual Network Functions Migration for NFV. In Proceedings of the 2016 IEEE 22nd International Conference on Parallel and Distributed Systems (ICPADS), Wuhan, China, 13–16 December 2016; pp. 340–346. [Google Scholar]
- Qu, K.; Zhuang, W.; Shen, X.; Li, X.; Rao, J. Dynamic Resource Scaling for VNF Over Nonstationary Traffic: A Learning Approach. IEEE Trans. Cogn. Commun. Netw. 2021, 7, 648–662. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Chen, J.; Chen, J.; Cheng, X.; Guo, K.; Qin, Y. A Deep Q-Learning Based VNF Migration Strategy for Elastic Control in SDN/NFV Network. In Proceedings of the 2021 International Conference on Wireless Communications and Smart Grid (ICWCSG), Hangzhou, China, 13–15 August 2021; pp. 217–223. [Google Scholar]
- Zhang, Q.; Liu, F.; Zeng, C. Online Adaptive Interference-Aware VNF Deployment and Migration for 5G Network Slice. IEEE/ACM Trans. Netw. 2021, 29, 2115–2128. [Google Scholar] [CrossRef] [Scilit]
- Chen, R.; Lu, H.; Lu, Y.; Liu, J. MSDF: A Deep Reinforcement Learning Framework for Service Function Chain Migration. In Proceedings of the 2020 IEEE Wireless Communications and Networking Conference (WCNC), Seoul, Republic of Korea, 25–28 May 2020; pp. 1–6. [Google Scholar]
- Afrasiabi, S.N.; Ebrahimzadeh, A.; Promwongsa, N.; Mouradian, C.; Li, W.; Recse, Á.; Szabó, R.; Glitho, R.H. Cost-Efficient Cluster Migration of VNFs for Service Function Chain Embedding. IEEE Trans. Netw. Serv. Manag. 2024, 21, 979–993. [Google Scholar] [CrossRef] [Scilit]
- Chebotar, Y.; Hausman, K.; Lu, Y.; Xiao, T.; Kalashnikov, D.; Varley, J.; Irpan, A.; Eysenbach, B.; Julian, R.; Finn, C.; et al. Actionable Models: Unsupervised Offline Reinforcement Learning of Robotic Skills. arXiv 2021, arXiv:2104.07749. [Google Scholar] [CrossRef] [Scilit]
- Diehl, C.; Sievernich, T.S.; Krüger, M.; Hoffmann, F.; Bertram, T. Uncertainty-Aware Model-Based Offline Reinforcement Learning for Automated Driving. IEEE Robot. Autom. Lett. 2023, 8, 1167–1174. [Google Scholar] [CrossRef] [Scilit]
- Lu, B.; Wang, K.; Xu, J.; Xie, R.; Song, L.; Zhang, W. Pioneer: Offline Reinforcement Learning based Bandwidth Estimation for Real-Time Communication. In Proceedings of the 15th ACM Multimedia Systems Conference, Bari, Italy, 15–18 April 2024; pp. 306–312. [Google Scholar]
- Gouareb, R.; Friderikos, V.; Aghvami, A.H. Virtual network functions routing and placement for edge cloud latency minimization. IEEE J. Sel. Areas Commun. 2018, 36, 2346–2357. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Lu, Z.; Wen, X.; Knopp, R.; Gupta, R. Joint optimization of service function chaining and resource allocation in network function virtualization. IEEE Access 2016, 4, 8084–8094. [Google Scholar] [CrossRef] [Scilit]
- Eramo, V.; Ammar, M.; Lavacca, F.G. Migration energy aware reconfigurations of virtual network function instances in NFV architectures. IEEE Access 2017, 5, 4927–4938. [Google Scholar] [CrossRef] [Scilit]
- Ota, T. Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces. arXiv 2024, arXiv:2403.19925. [Google Scholar] [CrossRef] [Scilit]
- Qu, H.; Ning, L.; An, R.; Fan, W.; Derr, T.; Liu, H.; Xu, X.; Li, Q. A Survey of Mamba. arXiv 2024, arXiv:2408.01129. [Google Scholar] [PubMed]
- Kumar, A.; Zhou, A.; Tucker, G.; Levine, S. Conservative Q-Learning for Offline Reinforcement Learning. In Advances in Neural Information Processing Systems 33; Neural Information Processing Systems Foundation, Inc.: San Diego, CA, USA, 2020. [Google Scholar]
- Xu, L.; Hu, H.; Liu, Y. SFCSim: A Network Function Virtualization Resource Allocation Simulation Platform. Clust. Comput. 2022, 26, 423–436. [Google Scholar] [CrossRef] [Scilit]
- Kostrikov, I.; Nair, A.; Levine, S. Offline reinforcement learning with implicit q-learning. arXiv 2021, arXiv:2110.06169. [Google Scholar] [CrossRef] [Scilit]








| Algorithm | Migration Success Rate | Migration Time | Migration Cost |
|---|---|---|---|
| RL-PMO | 0.95257 | 3.09393 | 24.41 |
| DM | 0.85597 | 3.0985 | 23.7936 |
| DT | 0.7969 | 3.0436 | 22.4975 |
| BC | 0.74727 | 3.1017 | 32.2223 |
| IQL | 0.8409 | 3.0188 | 21.94397 |
| IQL+FT | 0.86083 | 3.05413 | 23.19163 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hu, H.; Liu, Z.; Wu, F. RL-PMO: A Reinforcement Learning-Based Optimization Algorithm for Parallel SFC Migration. Sensors 2026, 26, 242. https://doi.org/10.3390/s26010242
Hu H, Liu Z, Wu F. RL-PMO: A Reinforcement Learning-Based Optimization Algorithm for Parallel SFC Migration. Sensors. 2026; 26(1):242. https://doi.org/10.3390/s26010242
Chicago/Turabian StyleHu, Hefei, Zining Liu, and Fan Wu. 2026. "RL-PMO: A Reinforcement Learning-Based Optimization Algorithm for Parallel SFC Migration" Sensors 26, no. 1: 242. https://doi.org/10.3390/s26010242
APA StyleHu, H., Liu, Z., & Wu, F. (2026). RL-PMO: A Reinforcement Learning-Based Optimization Algorithm for Parallel SFC Migration. Sensors, 26(1), 242. https://doi.org/10.3390/s26010242

