Deadline-Aware Scheduler-Weight Adaptation for 5G NR V2X Networks Using Probabilistic Prediction and Reinforcement Learning
Abstract
1. Introduction
1.1. Motivation
1.2. Problem Statement
1.3. Proposed Approach and Scope
1.4. Contributions
- A closed-loop ns-3/ns3-ai framework for AI-assisted scheduler-weight control in 5G NR V2X networks, designed to be modular with respect to the underlying NR MAC scheduler.
- A formulation of deadline-violation prediction as a binary next-window supervised classification task using rolling-window telemetry, together with the training of three probabilistic classifiers (GMM, HMM, BLR) for this task.
- A reinforcement-learning controller based on PPO with a deadline-aware reward function, evaluated both standalone and with classifier-derived violation probabilities included as state features.
- An evaluation of ten scheduling strategies: the PF baseline, the non-learning SB-DAS heuristic, three classifier-only controllers, three classifier-assisted PPO variants, PPO-only, and PPO-only with safety shield, using DC-PRR and supporting metrics across multiple vehicle densities and random seeds.
- An empirical assessment of whether classifier-derived violation probabilities improve PPO-based scheduler control. Our results suggest that, under the evaluated conditions, the hybrid combination does not provide a consistent benefit beyond standalone classifier or standalone PPO control.
- A non-learning deadline-aware baseline, the Slack-Based Deadline-Aware Scheduler (SB-DAS), which maps minimum packet slack and a recent violation ratio onto the same scheduler-weight set as the proposed controllers. Including SB-DAS enables a fair PF → heuristic → learned comparison and isolates the contribution of deadline awareness from that of learning.
- Full reproducibility material (software version, hyperparameters, algorithms, reward coefficients, normalization, seeds, and a consolidated notation/parameter table), together with a reward-contribution and coefficient-sensitivity analysis, a scheduler-weight-set rationale and sensitivity analysis, and a per-controller inference-latency and computational-complexity analysis that grounds the deployment-simplicity discussion.
1.5. Paper Organization
2. Related Work
2.1. Reinforcement Learning Based Resource Allocation for V2X
2.2. QoS Prediction and Violation Forecasting
2.3. Deadline-Aware Scheduling in 5G NR
2.4. Safe Reinforcement Learning and Shielding
2.5. Summary and Positioning
3. System Model and Simulation Setup
3.1. Network Topology
3.2. 5G NR Radio Configuration
3.3. V2X Traffic Model
3.4. Mobility Model
3.5. MAC Scheduler and Weight Adaptation
4. Proposed Deadline-Aware Scheduler-Weight Adaptation Framework
4.1. Closed-Loop Control Architecture
| Algorithm 1 Closed-loop deadline-aware scheduler control |
| 1: Initialize ns-3 V2X simulation, NR scheduler, and ns3-ai shared-memory interface |
| 2: Load feature normalizer, trained classifier, and/or trained PPO policy |
| 3: for each control interval t do |
| 4: ns-3 collects telemetry over the current 100 ms window |
| 5: Export network state through ns3-ai shared memory |
| 6: Python controller extracts and normalizes feature vector |
| 7: if classifier-only mode then |
| 8: Estimate violation probability |
| 9: Map to scheduler weight using Algorithm 2 |
| 10: else if SB-DAS mode then |
| 11: Compute minimum slack and recent violation ratio from telemetry |
| 12: Map to scheduler weight using Algorithm 3 |
| 13: else if PPO mode then |
| 14: Construct PPO state from telemetry and optional classifier probability |
| 15: Select scheduler-weight action using the trained PPO policy |
| 16: if safety shield is active then |
| 17: Override using the deterministic safety action |
| 18: end if |
| 19: end if |
| 20: Return scheduler weights to ns-3 through shared memory |
| 21: Apply updated weights to the NR TDMA QoS scheduler |
| 22: end for |
4.2. Deadline-Violation Prediction Task
4.3. Feature Extraction
4.4. Probabilistic Classifiers
4.4.1. Gaussian Mixture Model
4.4.2. Hidden Markov Model
4.4.3. Bayesian Logistic Regression
4.5. Classifier-Only Scheduler Control
| Algorithm 2 Classifier-only probability-to-weight mapping with hysteresis |
| Require: Violation probability , previous weight , thresholds , hysteresis margin |
| Ensure: Scheduler weight |
| 1: if critical slack is below emergency threshold then |
| 2: |
| 3: else if then |
| 4: |
| 5: else if then |
| 6: |
| 7: else if then |
| 8: |
| 9: else |
| 10: |
| 11: end if |
| 12: if then |
| 13: if and then |
| 14: |
| 15: else if and then |
| 16: |
| 17: else if and then |
| 18: |
| 19: end if |
| 20: end if |
| 21: return |
4.6. Slack-Based Deadline-Aware Baseline (SB-DAS)
| Algorithm 3 SB-DAS slack-to-level mapping with hysteresis (per controlled flow, every 100 ms) | |
| Require: minimum slack (ms), recent violation ratio , previous level , edges , , hysteresis margin ms | |
| Ensure: scheduler weight | |
| |
| |
| |
| ▹ normal |
| |
| ▹ recent violations |
| |
| ▹ emergency |
| |
| ▹ strong |
| |
| ▹ mild |
| |
| ▹ safe |
| |
| |
| |
| |
4.7. PPO-Based Scheduler Control
4.7.1. Reward Function
- Network health.
- Slack reward.
- Personal term.
- Altruism term and total.
4.7.2. Safety Shield
5. Experimental Methodology
5.1. Dataset Generation and Train/Evaluation Split
Generalization Split
- Load. PPO policies were trained under stressed conditions: 60-vehicle density with offered packet rates increased by 40% relative to Table 2. Evaluation used the unmodified offered rates and includes 30- and 40-vehicle densities that the agent never saw during training. The 30- and 40-vehicle evaluations therefore constituted a zero-shot generalization test for the learned policies.
- Mobility. Inference used different SUMO mobility traces (different vehicle trajectories on the same urban grid) and different random seeds from those used for training-data generation. The classifiers and PPO policies were not exposed to the evaluation traces.
5.2. Classifier Training
5.3. PPO Hyperparameter Search
5.4. PPO Training Procedure and Compute Budget
5.5. Compared Scheduling Strategies
- 1.
- PF baseline: PF scheduling with equal weights and no AI intervention.
- 2.
- SB-DAS: The non-learning slack-based deadline-aware heuristic of Section 4.6, on the same weight set .
- 3.
- GMM-only: GMM violation probability mapped directly to scheduler weights.
- 4.
- HMM-only: HMM violation probability mapped directly to scheduler weights.
- 5.
- BLR-only: BLR violation probability mapped directly to scheduler weights.
- 6.
- GMM+PPO: PPO control with GMM violation probability included in the state.
- 7.
- HMM+PPO: PPO control with HMM violation probability included in the state.
- 8.
- BLR+PPO: PPO control with BLR violation probability included in the state.
- 9.
- PPO-only without shield: PPO control without explicit classifier-derived violation probability.
- 10.
- PPO+shield: PPO-only control with emergency shielding at inference.
5.6. Evaluation Metrics
5.7. Statistical Validation
6. Results
6.1. Offline Classifier Performance
6.2. Adaptive Scheduling vs. PF Baseline
6.3. Deadline-Aware Heuristic Baseline and Source of the Gain
6.4. Per-Flow Analysis
6.5. Secondary Metrics: Latency, Throughput, and Jitter
6.6. Scaling with Vehicle Density
6.7. Comparing AI-Assisted Methods
- The differences are small relative to the seed variability. Mean differences of one to three percentage points between AI methods are comparable to or smaller than the across-seed standard deviations (Table 6). With three seeds per density, paired analyses cannot reliably separate methods that differ by less than two percentage points. For instance, the paired bootstrap confidence score for PPO-only exceeding GMM-only is only 67.5%, and the paired CI includes zero. PPO-only and GMM-only should therefore be regarded as statistically comparable at this sample size.
- The hybrid PPO+classifier configurations did not deliver the expected benefit. Adding classifier-derived violation probabilities to the PPO state did not consistently improve performance: GMM+PPO underperformed both GMM-only and PPO-only; HMM+PPO showed the largest density sensitivity and the largest secondary-metric degradation; BLR+PPO sat between the standalone methods. This is a substantive negative finding for the privileged-feature design hypothesis. A plausible explanation is that PPO can extract equivalent information about deadline risk from the raw telemetry features in its state, so adding a classifier-derived probability is partially redundant. Further experimentation is needed to distinguish confidently between AI-assisted methods with statistically valid differences.
- The robust comparative result is that all adaptive variants clearly outperform the PF baseline. Within the simulated load range, four learned/probabilistic controllers (GMM-only, HMM-only, BLR-only, PPO-only) form a tight performance band around 97.9–99.1% mean DC-PRR, and the non-learning SB-DAS heuristic falls inside the same band (Section 6.3). Three classifier-assisted PPO variants form a slightly broader and more variable band (95.5–98.1%). The main conclusion is that every adaptive variant, including the deterministic heuristic, consistently and clearly outperforms the PF baseline. Among the controllers, the practically interesting result is that GMM-only matches PPO-only within statistical noise, and SB-DAS in turn matches the learned controllers within statistical noise while being the simplest to train (it requires no training at all), deploy, and reason about.
6.8. Inference Latency and Computational Complexity
6.9. Reward Contribution and Sensitivity
7. Discussion
7.1. Deployment
7.2. Limitations
- Single-gNB topology due to tool constraints. The evaluation uses a single-cell deployment because the 5G-LENA NR module in ns-3 does not currently support handover. Multi-gNB scenarios with handover, inter-cell interference coordination, and multi-cell load balancing are therefore outside the scope of this study.
- Compute-bounded PPO training. Each PPO training run requires more than two days of continuous wall-clock time because each control step calls a full ns-3 simulation step. PPO did not converge in a strict asymptotic sense within the available budget, and the PPO results should be interpreted as those of an online controller under fixed bounded training rather than of a fully converged policy. Longer training may improve policy quality, particularly for the classifier-assisted PPO variants, although inference results indicate a learned policy and improvement over the PF baseline.
- Limited statistical power and inability of current experimental setup for ranking AI-assisted methods. With three seeds per density and method, the paired evidence is sufficient to support the primary PF-vs-adaptive comparison, because every adaptive method improved over PF in every matched density/seed pair with large effect sizes. However, distinguishing among the adaptive methods may require a more concentrated experimental design. Paired confidence scores and bootstrap CIs partially address this limitation, but absolute rankings among the top methods remain uncertain.
- Discrete scheduler-weight action set. The action set keeps the control problem tractable but may limit fine-grained adaptation. A continuous or finer-grained action space could improve performance but would require redesigning the classifier-only mapping rule, adapting the PPO action head, and, most importantly, using a much larger compute budget.
- Open-loop sensitivity analyses. The reward-coefficient sensitivity and contribution decomposition (Section 6.9) and the weight-set sensitivity (Appendix B) were computed offline over recorded trajectories and decisions, with the policy and trajectory held fixed. They quantified how the reward/control signal responded to parameter changes but did not re-learn the policy. A fully retrained reward ablation and a closed-loop re-evaluation under alternative weight sets are left to future work due to the compute budget constraint.
8. Conclusions and Future Work
Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Consolidated Notation and Parameter Table
| Symbol/Constant | Meaning | Value/Domain |
|---|---|---|
| i | UE (or flow) index | — |
| t | Control-cycle index (period 100 ms) | — |
| Packet deadline for UE i | ms | |
| Observed latency | ms | |
| Minimum remaining slack | ms | |
| Recent deadline-violation ratio (5-cycle window) | ||
| Discrete priority level | ||
| Applied scheduler weight | ||
| Mean critical-UE health | ||
| Buffer occupancy normalized to capacity | ||
| Control period | All controllers | 100 ms |
| Rolling window | Feature/violation window | 5 cycles (500 ms) |
| Deadlines ToD UL/DL | Slack, DC-PRR | 100/20 ms |
| Deadlines AIC UL/DL | Slack, DC-PRR | 100/100 ms |
| Deadlines RTSA UL/DL | DC-PRR | 500/500 ms |
| Deadlines HDM UL/DL | DC-PRR | 1000/1000 ms |
| Weight map | Scheduler | , , , |
| SB-DAS slack edges | Level rule | 50/20/5 ms |
| SB-DAS violation-ratio threshold | Level rule | 0.2 |
| SB-DAS hysteresis margin | Downgrade guard | 5 ms |
| Classifier thresholds | level quantization | 0.30/0.60/0.85 |
| Classifier L4 hysteresis hold | Level hold | 3 cycles (300 ms) |
| GMM components K | Classifier | 5 (diag. cov.) |
| HMM states S | Classifier | 8 (diag. cov.) + LR head |
| BLR prior precision/MC samples M | Classifier | 1.0/100 |
| ToD/AIC buffer capacity | Buffer normalization | 256/64 KiB |
Appendix B. Scheduler-Weight-Set Sensitivity
| Weight Set | Mean Applied Weight | p95 Applied Weight | Boosted Weight-Mass % |
|---|---|---|---|
| (default) | 1.654 | 5 | 46.1 |
| (octave) | 1.503 | 4 | 40.7 |
| (linear) | 1.716 | 6 | 48.1 |
Appendix C. SB-DAS Threshold Sensitivity
| Threshold | Value | Boosted % () | Mean Weight |
|---|---|---|---|
| SLACK_SAFE_MS | 30 | 10.57 | 1.641 |
| 50 * | 10.86 | 1.644 | |
| 70 | 11.78 | 1.653 | |
| 100 | 100.0 | 2.536 | |
| SLACK_MILD_MS | 10 | 10.86 | 1.483 |
| 20 * | 10.86 | 1.644 | |
| 40 | 10.86 | 1.650 | |
| SLACK_HIGH_MS | 2 | 10.86 | 1.642 |
| 5 * | 10.86 | 1.644 | |
| 15 | 10.86 | 1.944 | |
| VIOL_RATIO_THR | 0.1 | 11.37 | 1.697 |
| 0.2 * | 10.86 | 1.644 | |
| 0.8 | 10.73 | 1.630 |
Appendix D. Reproducibility and Implementation Parameters
| Component | Version/Reference |
|---|---|
| ns-3 | 3.46.1 (dev), git ns-3.46 |
| nr (5G-LENA) | v4.1.1 (contrib/nr) |
| ns3-ai | git b8c9858 (contrib/ai) |
| Build profile | optimized (-d optimized, release) |
| Compiler/CMake | gcc 13.3.0/CMake 3.28.3 |
| OS/kernel | Ubuntu 24.04.4 LTS/6.17.0-35-generic |
| Python | 3.11.14 (conda env ns3ai_env) |
| numpy/pandas/scipy | 2.3.5/2.3.3/1.17.0 |
| torch | 2.9.1+cu128 |
| scikit-learn/hmmlearn | 1.8.0/0.3.3 |
| optuna/matplotlib/joblib | 4.8.0/3.10.7/1.5.3 |
| Constant | Value |
|---|---|
| SLACK_SAFE_MS | 50 |
| SLACK_MILD_MS | 20 |
| SLACK_HIGH_MS | 5 |
| VIOLATION_RATIO_THRESHOLD | 0.2 |
| HYSTERESIS_SLACK_MS | 5 |
| CONTROL_CYCLE_S | 0.1 |
| WINDOW_SAMPLES | 5 |
- 1.
- Background flows. HDM/RTSA flows are clamped to L1.
- 2.
- Classifier-only common quantization. Probability thresholds L2/L3/L4 (else L1); L4 hysteresis hold three cycles; 100 ms control cycle; five-sample rolling feature window. Per-classifier: GMM two class-conditional models, n_components=5, diag. cov., reg_covar=1e-4; HMM n_components=8, diag. cov., logistic head , online forward filtering; BLR prior_precision=1.0, n_mc_samples=100, Laplace approximation.
- 3.
- PPO hyperparameters (shared across all PPO variants). Learning rate: ; discount: ; GAE: ; clip ; entropy coef.; value-loss coef. ; max grad norm: ; rollout length: 32; PPO epochs: four; mini-batch: 32; network shared: (ReLU), actor: , critic: ; action → level map: ; shield emergency-slack threshold: 5 ms. Training used a 500 s trace at 60-vehicle density with offered load; evaluation seeds: one, two, three; densities: 30, 40, 60.
References
- TS 22.186; Service Requirements for Enhanced V2X Scenarios (Release 16). Technical Specification 22.186. 3rd Generation Partnership Project (3GPP): Valbonne, France, 2019.
- Bagheri, H.; Noor-A-Rahim, M.; Liu, Z.; Lee, H.; Pesch, D.; Moessner, K.; Xiao, P. 5G NR-V2X: Toward Connected and Cooperative Autonomous Driving. IEEE Commun. Stand. Mag. 2021, 5, 48–54. [Google Scholar] [CrossRef]
- Capozzi, F.; Piro, G.; Grieco, L.A.; Boggia, G.; Camarda, P. Downlink Packet Scheduling in LTE Cellular Networks: Key Design Issues and a Survey. IEEE Commun. Surv. Tutor. 2013, 15, 678–700. [Google Scholar] [CrossRef]
- Patriciello, N.; Lagén, S.; Bojović, B.; Giupponi, L. An E2E Simulator for 5G NR Networks. Simul. Model. Pract. Theory 2019, 96, 101933. [Google Scholar] [CrossRef]
- Liang, L.; Ye, H.; Li, G.Y. Spectrum Sharing in Vehicular Networks Based on Multi-Agent Reinforcement Learning. IEEE J. Sel. Areas Commun. 2019, 37, 2282–2292. [Google Scholar] [CrossRef]
- Ye, H.; Li, G.Y.; Juang, B.H.F. Deep Reinforcement Learning Based Resource Allocation for V2V Communications. IEEE Trans. Veh. Technol. 2019, 68, 3163–3173. [Google Scholar] [CrossRef]
- Shao, Z.; Wu, Q.; Fan, P.; Cheng, N.; Fan, Q.; Wang, J. Semantic-Aware Resource Allocation Based on Deep Reinforcement Learning for 5G-V2X HetNets. arXiv 2024, arXiv:2406.07996. [Google Scholar]
- Jabeen, S. A Deep Reinforcement Learning Based Scheduler for IoT Devices in Co-existence with 5G-NR. arXiv 2025, arXiv:2501.11574. [Google Scholar]
- AL-Tam, F.; Correia, N.; Rodriguez, J. Learn to Schedule (LEASCH): A Deep Reinforcement Learning Approach for Radio Resource Scheduling in the 5G MAC Layer. IEEE Access 2020, 8, 108088–108101. [Google Scholar] [CrossRef]
- Gu, Z.; She, C.; Hardjawana, W.; Lumb, S.; McKechnie, D.; Essery, T.; Vucetic, B. Knowledge-Assisted Deep Reinforcement Learning in 5G Scheduler Design: From Theoretical Framework to Implementation. IEEE J. Sel. Areas Commun. 2021, 39, 2014–2028. [Google Scholar] [CrossRef]
- Sun, S.; Li, X. Deep-Reinforcement-Learning-Based Scheduling with Contiguous Resource Allocation for Next-Generation Wireless Systems. In Proceedings of the Intelligent Computing; Lecture Notes in Networks and Systems; Springer: Cham, Switzerland, 2021; Volume 285, pp. 637–657. [Google Scholar] [CrossRef]
- Barmpounakis, S.; Maroulis, N.; Koursioumpas, N.; Kousaridas, A.; Kalamari, A.; Kontopoulos, P.; Alonistioti, N. AI-Driven QoS Prediction for V2X Communications in Beyond 5G Systems. Comput. Netw. 2022, 217, 109341. [Google Scholar] [CrossRef]
- Palaios, A.; Vielhaus, C.L.; Külzer, D.F.; Watermann, C.; Hernangomez, R.; Partani, S.; Geuer, P.; Krause, A.; Sattiraju, R.; Kasparick, M.; et al. Machine Learning for QoS Prediction in Vehicular Communication: Challenges and Solution Approaches. IEEE Access 2023, 11, 92459–92477. [Google Scholar] [CrossRef]
- Koursioumpas, N.; Magoula, L.; Stavrakakis, I.; Alonistioti, N.; Gutierrez-Estevez, M.A.; Khalili, R. DISTINQT: A Distributed Privacy Aware Learning Framework for QoS Prediction for Future Mobile and Wireless Networks. arXiv 2024, arXiv:2401.10158. [Google Scholar]
- Allahdadi, A.; Morla, R. Anomaly Detection and Modeling in 802.11 Wireless Networks. J. Netw. Syst. Manag. 2019, 27, 3–38. [Google Scholar] [CrossRef]
- TS 23.501; System Architecture for the 5G System (5GS) (Release 17). Technical Specification 23.501. 3rd Generation Partnership Project (3GPP): Valbonne, France, 2022.
- Liu, C.L.; Layland, J.W. Scheduling Algorithms for Multiprogramming in a Hard-Real-Time Environment. J. ACM 1973, 20, 46–61. [Google Scholar] [CrossRef]
- Andrews, M.; Kumaran, K.; Ramanan, K.; Stolyar, A.; Vijayakumar, R.; Whiting, P. Providing Quality of Service over a Shared Wireless Link. IEEE Commun. Mag. 2001, 39, 150–154. [Google Scholar] [CrossRef]
- Hendaoui, S.; Hendaoui, F.; Zangar, N. Dynamic Proactive–Reactive Scheduling for URLLC in 5G: Leveraging XGBoost and Network Virtualization. Phys. Commun. 2025, 68, 102553. [Google Scholar] [CrossRef]
- Cohen, K.M.; Park, S.; Simeone, O.; Popovski, P.; Shamai, S. Guaranteed Dynamic Scheduling of Ultra-Reliable Low-Latency Traffic via Conformal Prediction. IEEE Signal Process. Lett. 2023, 30, 473–477. [Google Scholar] [CrossRef]
- Robaglia, B.M.; Destounis, A.; Coupechoux, M.; Tsilimantos, D. Deep Reinforcement Learning for Scheduling Uplink IoT Traffic with Strict Deadlines. In Proceedings of the IEEE Global Communications Conference (GLOBECOM), Madrid, Spain, 7–11 December 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 1–6. [Google Scholar]
- Kanavos, A.; Barmpounakis, S.; Kaloxylos, A. An Adaptive Scheduling Mechanism Optimized for V2N Communications over Future Cellular Networks. Telecom 2023, 4, 378–392. [Google Scholar] [CrossRef]
- Alshiekh, M.; Bloem, R.; Ehlers, R.; Könighofer, B.; Niekum, S.; Topcu, U. Safe Reinforcement Learning via Shielding. Proc. AAAI Conf. Artif. Intell. 2018, 32, 2669–2678. [Google Scholar] [CrossRef]
- Carr, S.; Jansen, N.; Junges, S.; Topcu, U. Safe Reinforcement Learning via Shielding under Partial Observability. Proc. AAAI Conf. Artif. Intell. 2023, 37, 14748–14756. [Google Scholar] [CrossRef]














| Parameter | Value |
|---|---|
| Simulator | ns-3 with 5G-LENA NR module |
| Deployment | Single-gNB urban V2X cell |
| Carrier frequency | 4 GHz |
| System bandwidth | 25 MHz |
| Numerology | 2 |
| Subcarrier spacing | 60 kHz |
| MAC scheduler | NR TDMA QoS scheduler |
| RLC mode | AM |
| Mobility source | SUMO Manhattan-grid traces |
| AI interface | ns3-ai shared memory |
| Control period | 100 ms |
| Flow | Direction | Packet Size | Period | Offered Rate |
|---|---|---|---|---|
| HDM | DL | 1200 B | 3 ms | 3.20 Mbps |
| HDM | UL | 600 B | 10 ms | 0.48 Mbps |
| AIC | DL | 1000 B | 10 ms | 0.80 Mbps |
| AIC | UL | 1200 B | 10 ms | 0.96 Mbps |
| ToD | DL | 600 B | 10 ms | 0.48 Mbps |
| ToD | UL | 1200 B | 3 ms | 3.20 Mbps |
| RTSA | DL | 1000 B | 10 ms | 0.80 Mbps |
| RTSA | UL | 1000 B | 10 ms | 0.80 Mbps |
| Symbol | Value | Role/Where It Enters |
|---|---|---|
| Violation penalty multiplier | 3.0 | Negative branch of |
| Positive-slack saturation | 1.0 | tanh upper bound () |
| Buffer-pressure weight (UL, DL) | 1.0 each | , critical UEs |
| Background rescale gain | 4.0 | |
| Background rescale offset | ||
| Altruism gain | 2.0 | |
| Altruism offset | ||
| ToD buffer capacity | 256 KiB | (ToD) |
| AIC buffer capacity | 64 KiB | (AIC) |
| ToD deadline (UL/DL) | 100/20 ms | Slack normalizer D |
| AIC deadline (UL/DL) | 100/100 ms | Slack normalizer D |
| Trial | Score | Elapsed (s) | LR | Rollout | |||
|---|---|---|---|---|---|---|---|
| 9 | 0.9999 | 8057.5 | 0.000814 | 0.9019 | 0.9089 | 0.0578 | 32 |
| 1 | 0.9978 | 8043.2 | 0.000025 | 0.9907 | 0.9706 | 0.1402 | 128 |
| 7 | 0.9974 | 8059.1 | 0.000103 | 0.9435 | 0.9872 | 0.1655 | 32 |
| 5 | 0.9966 | 8040.4 | 0.000150 | 0.9054 | 0.9874 | 0.1946 | 64 |
| 4 | 0.9939 | 8001.4 | 0.000040 | 0.9501 | 0.9365 | 0.1197 | 128 |
| 2 | 0.9915 | 7975.2 | 0.000457 | 0.9258 | 0.9601 | 0.0683 | 32 |
| 8 | 0.9893 | 8010.1 | 0.000033 | 0.9219 | 0.9199 | 0.1674 | 32 |
| 6 | 0.9845 | 8020.1 | 0.000737 | 0.9081 | 0.9034 | 0.2027 | 256 |
| 0 | 0.9716 | 8024.4 | 0.000024 | 0.9448 | 0.9834 | 0.1029 | 64 |
| 3 | 0.9343 | 7764.7 | 0.000433 | 0.9664 | 0.9385 | 0.2240 | 64 |
| Classifier | Accuracy | F1 | Precision | Recall | ROC-AUC | Avg. Prec. | ECE |
|---|---|---|---|---|---|---|---|
| GMM | 0.9366 | 0.7709 | 0.7089 | 0.8467 | 0.9544 | 0.8157 | – |
| HMM | 0.9408 | 0.7156 | 0.9017 | 0.5933 | 0.9256 | 0.8199 | – |
| BLR | 0.9405 | 0.7140 | 0.9010 | 0.5915 | 0.9284 | 0.8220 | 0.0225 |
| Method | 30 Veh. | 40 Veh. | 60 Veh. | Overall | Std. | 95% CI |
|---|---|---|---|---|---|---|
| PF baseline | 75.20 | 65.33 | 44.11 | 61.55 | 14.54 | [52.37, 70.21] |
| SB-DAS | 99.14 | 99.39 | 97.22 | 98.58 | 1.51 | [97.56, 99.35] |
| GMM-only | 99.62 | 99.77 | 97.53 | 98.97 | 1.11 | [98.25, 99.64] |
| HMM-only | 98.33 | 98.70 | 98.11 | 98.38 | 0.28 | [98.21, 98.53] |
| BLR-only | 97.78 | 98.95 | 96.93 | 97.89 | 1.61 | [96.78, 98.73] |
| GMM+PPO | 98.85 | 97.82 | 97.67 | 98.11 | 1.40 | [97.09, 98.86] |
| HMM+PPO | 99.40 | 95.83 | 91.18 | 95.47 | 6.89 | [90.73, 98.48] |
| BLR+PPO | 98.57 | 97.27 | 93.94 | 96.59 | 4.52 | [93.48, 98.67] |
| PPO-only | 99.77 | 98.69 | 98.82 | 99.09 | 0.73 | [98.63, 99.54] |
| PPO-only+shield | 98.57 | 97.01 | 96.04 | 97.21 | 2.34 | [95.70, 98.52] |
| Method | Mean Gain | 95% CI Low | 95% CI High | Conf. | Wilcoxon |
|---|---|---|---|---|---|
| (pp) | (pp) | (pp) | Score | ||
| SB-DAS | 37.04 | 28.85 | 45.61 | 100% | 0.00195 |
| GMM-only | 37.43 | 29.17 | 46.02 | 100% | 0.00195 |
| HMM-only | 36.83 | 28.13 | 45.78 | 100% | 0.00195 |
| BLR-only | 36.34 | 27.79 | 45.38 | 100% | 0.00195 |
| GMM+PPO | 36.57 | 28.48 | 45.29 | 100% | 0.00195 |
| HMM+PPO | 33.92 | 26.57 | 42.03 | 100% | 0.00195 |
| BLR+PPO | 35.05 | 27.15 | 43.68 | 100% | 0.00195 |
| PPO-only | 37.55 | 28.90 | 46.65 | 100% | 0.00195 |
| PPO-only+shield | 35.66 | 27.28 | 44.34 | 100% | 0.00195 |
| Rung | 30 Veh. | 40 Veh. | 60 Veh. |
|---|---|---|---|
| PF baseline | 75.20 | 65.33 | 44.11 |
| SB-DAS | 99.14 (+23.94) | 99.39 (+34.05) | 97.22 (+53.12) |
| GMM-only | 99.62 (+0.48) | 99.77 (+0.39) | 97.53 (+0.31) |
| GMM+PPO | 98.85 (−0.76) | 97.82 (−1.96) | 97.67 (+0.14) |
| PPO-only+shield | 98.57 (−0.28) | 97.01 (−0.81) | 96.04 (−1.63) |
| Flow | PF | SB-DAS | GMM-Only | HMM-Only | BLR-Only | GMM+PPO | HMM+PPO | BLR+PPO | PPO-Only | PPO+Shield |
|---|---|---|---|---|---|---|---|---|---|---|
| ToD-UL | 99.99 | 99.83 | 99.87 | 99.70 | 99.61 | 99.71 | 99.99 | 99.90 | 99.90 | 99.90 |
| ToD-DL | 95.21 | 99.65 | 99.62 | 99.06 | 97.86 | 99.82 | 100.00 | 99.93 | 100.00 | 99.93 |
| AIC-UL | 42.62 | 96.82 | 99.34 | 95.98 | 86.02 | 92.65 | 96.02 | 81.42 | 98.53 | 81.42 |
| AIC-DL | 98.39 | 99.73 | 99.65 | 98.98 | 99.04 | 99.75 | 100.00 | 99.96 | 99.98 | 99.96 |
| RTSA-UL | 46.54 | 99.82 | 99.79 | 99.92 | 99.53 | 99.89 | 99.92 | 99.43 | 99.99 | 99.43 |
| RTSA-DL | 98.66 | 99.76 | 99.54 | 97.69 | 98.88 | 99.64 | 100.00 | 99.96 | 99.99 | 99.96 |
| HDM-UL | 63.11 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 99.98 | 100.00 |
| HDM-DL | 98.93 | 99.72 | 99.34 | 96.65 | 97.83 | 99.20 | 99.99 | 99.85 | 99.93 | 99.85 |
| Flow | PF | SB-DAS | GMM-Only | HMM-Only | BLR-Only | GMM+PPO | HMM+PPO | BLR+PPO | PPO-Only | PPO+Shield |
|---|---|---|---|---|---|---|---|---|---|---|
| ToD-UL | 99.93 | 99.95 | 99.98 | 99.71 | 99.81 | 99.99 | 99.24 | 99.96 | 99.98 | 99.99 |
| ToD-DL | 96.23 | 99.58 | 99.79 | 98.26 | 99.17 | 97.76 | 95.82 | 97.40 | 98.99 | 96.68 |
| AIC-UL | 24.91 | 99.72 | 99.84 | 99.45 | 98.15 | 99.38 | 99.58 | 96.64 | 99.72 | 97.89 |
| AIC-DL | 87.46 | 99.32 | 99.76 | 98.30 | 99.26 | 97.33 | 95.57 | 96.99 | 98.87 | 96.59 |
| RTSA-UL | 39.07 | 99.86 | 99.95 | 99.95 | 98.97 | 99.93 | 99.68 | 99.92 | 99.34 | 99.88 |
| RTSA-DL | 89.52 | 98.99 | 99.76 | 97.53 | 99.02 | 96.10 | 91.80 | 96.09 | 98.26 | 94.61 |
| HDM-UL | 25.91 | 100.00 | 100.00 | 100.00 | 100.00 | 99.99 | 99.99 | 100.00 | 99.97 | 100.00 |
| HDM-DL | 95.39 | 97.84 | 98.94 | 96.45 | 98.33 | 92.65 | 87.28 | 92.75 | 94.77 | 92.42 |
| Flow | PF | SB-DAS | GMM-Only | HMM-Only | BLR-Only | GMM+PPO | HMM+PPO | BLR+PPO | PPO-Only | PPO+Shield |
|---|---|---|---|---|---|---|---|---|---|---|
| ToD-UL | 44.44 | 99.59 | 99.66 | 99.74 | 93.55 | 99.96 | 99.97 | 99.53 | 99.93 | 99.83 |
| ToD-DL | 78.13 | 97.26 | 97.45 | 96.68 | 97.24 | 98.20 | 91.64 | 94.38 | 98.86 | 95.18 |
| AIC-UL | 23.14 | 99.78 | 98.64 | 99.70 | 96.85 | 99.77 | 98.56 | 99.71 | 99.76 | 99.78 |
| AIC-DL | 54.66 | 96.75 | 98.05 | 96.97 | 97.84 | 97.17 | 89.68 | 93.61 | 98.57 | 95.12 |
| RTSA-UL | 35.35 | 99.92 | 98.59 | 99.90 | 98.29 | 99.95 | 99.91 | 99.97 | 99.96 | 99.92 |
| RTSA-DL | 54.44 | 95.05 | 96.35 | 96.54 | 96.74 | 96.10 | 80.18 | 87.01 | 97.96 | 91.98 |
| HDM-UL | 33.91 | 99.99 | 100.00 | 100.00 | 100.00 | 99.99 | 99.95 | 99.97 | 99.99 | 99.97 |
| HDM-DL | 54.92 | 89.68 | 92.18 | 94.90 | 94.03 | 89.50 | 72.67 | 82.24 | 95.41 | 87.40 |
| Method | 30 Veh. | 40 Veh. | 60 Veh. |
|---|---|---|---|
| PF baseline | 71.97 | 98.50 | 148.29 |
| SB-DAS | 14.22 | 12.24 | 26.50 |
| GMM-only | 11.19 | 11.44 | 19.96 |
| HMM-only | 14.91 | 11.71 | 12.64 |
| BLR-only | 19.00 | 12.08 | 17.17 |
| GMM+PPO | 15.88 | 22.77 | 26.69 |
| HMM+PPO | 13.58 | 34.16 | 75.11 |
| BLR+PPO | 18.13 | 28.87 | 70.09 |
| PPO-only | 10.89 | 19.30 | 13.89 |
| PPO-only+shield | 18.13 | 31.30 | 32.21 |
| Method | 30 Veh. | 40 Veh. | 60 Veh. |
|---|---|---|---|
| PF baseline | 103.14 | 132.20 | 320.50 |
| SB-DAS | 22.17 | 42.07 | 139.78 |
| GMM-only | 19.39 | 19.68 | 123.88 |
| HMM-only | 69.40 | 71.32 | 85.17 |
| BLR-only | 38.67 | 31.84 | 118.21 |
| GMM+PPO | 28.51 | 114.55 | 112.76 |
| HMM+PPO | 15.39 | 198.77 | 175.00 |
| BLR+PPO | 27.56 | 99.26 | 151.59 |
| PPO-only | 19.67 | 51.66 | 67.17 |
| PPO-only+shield | 27.56 | 139.89 | 227.35 |
| Method | 30 Veh. | 40 Veh. | 60 Veh. |
|---|---|---|---|
| PF baseline | 0.916 | 0.875 | 0.733 |
| SB-DAS | 1.023 | 1.027 | 1.016 |
| GMM-only | 1.027 | 1.027 | 1.021 |
| HMM-only | 1.022 | 1.022 | 1.017 |
| BLR-only | 1.019 | 1.027 | 1.022 |
| GMM+PPO | 1.026 | 1.020 | 1.020 |
| HMM+PPO | 1.027 | 1.011 | 1.002 |
| BLR+PPO | 1.022 | 1.020 | 1.004 |
| PPO-only | 1.027 | 1.024 | 1.024 |
| PPO-only+shield | 1.022 | 1.016 | 1.010 |
| Method | 30 Veh. | 40 Veh. | 60 Veh. |
|---|---|---|---|
| PF baseline | 7.63 | 10.84 | 16.72 |
| SB-DAS | 0.94 | 1.18 | 2.04 |
| GMM-only | 0.83 | 1.02 | 1.87 |
| HMM-only | 1.22 | 1.16 | 1.79 |
| BLR-only | 1.63 | 1.14 | 2.01 |
| GMM+PPO | 1.24 | 1.72 | 1.93 |
| HMM+PPO | 1.04 | 1.80 | 3.00 |
| BLR+PPO | 1.59 | 2.24 | 2.48 |
| PPO-only | 0.80 | 1.49 | 1.65 |
| PPO-only+shield | 1.59 | 1.91 | 2.10 |
| Method | Complexity | Params | 30 Vehicles | 40 Vehicles | 60 Vehicles |
|---|---|---|---|---|---|
| PF baseline | — (no control loop) | none | n/a | n/a | n/a |
| SB-DAS | 0 | 0.0060 ± 0.0004/0.0061 | 0.0056 ± 0.0001/0.0057 | 0.0058 ± 0.0001/0.0059 | |
| GMM-only | 6096 | 0.6703 ± 0.0202/0.6823 | 0.6893 ± 0.0108/0.7114 | 0.7167 ± 0.0110/0.7444 | |
| HMM-only | 1564 | 11.7808 ± 0.1358/11.9972 | 16.2086 ± 0.0905/16.3384 | 23.4847 ± 0.1921/23.6609 | |
| BLR-only | 420 | 0.1268 ± 0.0099/0.1339 | 0.1314 ± 0.0090/0.1374 | 0.1374 ± 0.0043/0.1440 | |
| PPO-only | 135,557 | 0.1228 ± 0.0018/0.1246 | 0.1478 ± 0.0073/0.1527 | 0.1923 ± 0.0031/0.1971 | |
| GMM+PPO | 141,653 | 0.8522 ± 0.0302/0.8861 | 0.9161 ± 0.0438/0.9719 | 0.9752 ± 0.0126/0.9960 | |
| HMM+PPO | 137,121 | 12.0662 ± 0.1272/12.2414 | 16.3170 ± 0.1535/16.5852 | 24.1710 ± 0.2201/24.4745 | |
| BLR+PPO | 135,977 | 0.2849 ± 0.0122/0.2972 | 0.3185 ± 0.0073/0.3316 | 0.3734 ± 0.0071/0.3851 |
| Reward Term | Mean Contribution | Share of Signal |
|---|---|---|
| Background | 35.8% | |
| Slack | 33.9% | |
| Altruism | 28.3% | |
| Buffer penalty | 2.0% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Papanikolaou-Ntais, G.; Sotiropoulos, D.N.; Kanavos, A.; Kaloxylos, A. Deadline-Aware Scheduler-Weight Adaptation for 5G NR V2X Networks Using Probabilistic Prediction and Reinforcement Learning. Telecom 2026, 7, 80. https://doi.org/10.3390/telecom7040080
Papanikolaou-Ntais G, Sotiropoulos DN, Kanavos A, Kaloxylos A. Deadline-Aware Scheduler-Weight Adaptation for 5G NR V2X Networks Using Probabilistic Prediction and Reinforcement Learning. Telecom. 2026; 7(4):80. https://doi.org/10.3390/telecom7040080
Chicago/Turabian StylePapanikolaou-Ntais, Gerasimos, Dionysios N. Sotiropoulos, Athanasios Kanavos, and Alexandros Kaloxylos. 2026. "Deadline-Aware Scheduler-Weight Adaptation for 5G NR V2X Networks Using Probabilistic Prediction and Reinforcement Learning" Telecom 7, no. 4: 80. https://doi.org/10.3390/telecom7040080
APA StylePapanikolaou-Ntais, G., Sotiropoulos, D. N., Kanavos, A., & Kaloxylos, A. (2026). Deadline-Aware Scheduler-Weight Adaptation for 5G NR V2X Networks Using Probabilistic Prediction and Reinforcement Learning. Telecom, 7(4), 80. https://doi.org/10.3390/telecom7040080

