Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (28)

Search Parameters:
Keywords = Markov chains with rewards

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
31 pages, 1828 KB  
Article
Service-Level Agreement-Aware Scheduling Algorithm Based on Heterogeneous Computing Collaboration in Smart Video Surveillance Scenarios
by Jiayang Song, Jing Wang, Jun Yan, Ping Ma and Shuihan Yi
Appl. Sci. 2026, 16(17), 8783; https://doi.org/10.3390/app16178783 - 3 Sep 2026
Viewed by 209
Abstract
To address the challenge of satisfying strict Service-Level Agreement (SLA) requirements for concurrent smart video surveillance tasks in heterogeneous edge computing environments, an SLA-aware adaptive scheduling algorithm for heterogeneous computing collaboration is proposed. First, a mixed-task flow model is constructed, and a finite-state [...] Read more.
To address the challenge of satisfying strict Service-Level Agreement (SLA) requirements for concurrent smart video surveillance tasks in heterogeneous edge computing environments, an SLA-aware adaptive scheduling algorithm for heterogeneous computing collaboration is proposed. First, a mixed-task flow model is constructed, and a finite-state Markov chain is utilized to dynamically model the time-varying wireless channel. Second, a Dueling Double Deep Q-Network (Dueling DDQN) scheduling algorithm based on SLA awareness and channel adaptation is proposed, with a designed SLA action-masking mechanism. This mechanism advances hard delay constraints to the decision-generation stage, dynamically prunes the action space based on real-time channel conditions and node loads, and filters out actions predicted to violate the SLA before execution. Experimental results show that the proposed algorithm coordinates heterogeneous computing resources between the cloud center and the edge and exhibits earlier empirical reward stabilization and lower task-violation rates than the compared learning-based baselines under the tested workload conditions. Full article
(This article belongs to the Special Issue Applications of Wireless and Mobile Communications, 2nd Edition)
Show Figures

Figure 1

32 pages, 6635 KB  
Article
Design of a Risk Assessment Model for Grassroots Agricultural Product Quality and Safety Based on Bayesian Networks and Evidential Reasoning
by Yijia Qiu and Yuheng Li
Symmetry 2026, 18(8), 1382; https://doi.org/10.3390/sym18081382 - 17 Aug 2026
Viewed by 247
Abstract
The quality and safety supervision of agricultural products at the grassroots level has long faced the triple superposition dilemma of small-sample sampling, multi-source evidence conflict, and risk chain evolution. Although existing data-driven models have considerable accuracy, they are difficult to leverage for intervention [...] Read more.
The quality and safety supervision of agricultural products at the grassroots level has long faced the triple superposition dilemma of small-sample sampling, multi-source evidence conflict, and risk chain evolution. Although existing data-driven models have considerable accuracy, they are difficult to leverage for intervention decisions, and the simple serial connection of traditional Bayesian networks and evidence theory cannot respond to dynamic scenarios. Aiming at this research gap, this paper constructs a dynamic risk assessment model, CIBE-DR, that deeply couples Bayesian networks with evidential reasoning. It contains three core innovations. First, the structure learning method of the causally identifiable Bayesian network embeds a graded do-calculus identifiability score covering both back-door and front-door criteria into the BDeu scoring function and combines this reward with an expert-prior divergence penalty that breaks Markov equivalence so as to realize the transition from relevance modeling to intervention decision modeling. Second, the conflict-aware adaptive evidence synthesis rule orthogonally decomposes multi-source conflict into an epistemic component and an ontological component, which are modeled respectively by Tsallis belief entropy and abductive inference over a discrete twenty-seven-point heterogeneity hypothesis space and are then fused under a reparameterized Dempster–Yager interpolation in which the two endpoints recover the two named rules under a single consistent interpretation. Third, the bidirectional closed-loop coupling mechanism between BN and ER realizes the mutual calibration between the conditional probability table and the evidence credibility prior under a Lyapunov monotone descent argument with the explicit Lipschitz bound Lθ ≤ 0.028 < 1, endowing the model with time-varying self-correction ability. Based on experiments on 156,847 sampling samples from counties and townships in East China, Central China, and Southwest China from 2021 to 2024, the proposed method achieved the best value in six of the seven evaluation indicators, with a minority recall of 0.864 ± 0.014, an intervention effect estimation error of 0.063 ± 0.005, and a dynamic response delay of 2.8 ± 0.3 days, significantly ahead of eleven mainstream baselines under the McNemar test on classification (p < 0.001) and the Wilcoxon signed-rank test on intervention-effect estimation (p < 0.001). The only indicator on which CIBE-DR does not lead is overall accuracy, which is 0.002 lower than that of Transformer; this difference does not reach statistical significance under the McNemar test (p = 0.32) and does not weaken the value of grassroots supervision in the strong-imbalance scenario where the positive rate is only 1.04%. The robustness advantage of the model is particularly prominent in the scenarios of sparse data, adversarial perturbation, and prior-graph incompleteness, and the intervention-effect estimates were additionally validated against two post-2022 policy interventions with absolute deviations of 1.4 and 1.2 percentage points respectively. These results verify the product gain and grassroots deployability of the three mechanisms. Full article
Show Figures

Figure 1

24 pages, 2424 KB  
Article
Adaptive Capacity Optimization Algorithm Leveraging Joint PHY-MAC Layer Modeling for Dual-Mode Communication Systems
by Yuerong Zhao, Bo Jiang and Zhixiong Chen
Electronics 2026, 15(15), 3461; https://doi.org/10.3390/electronics15153461 - 5 Aug 2026
Cited by 1 | Viewed by 292
Abstract
The extensive deployment of the Power Internet of Things (PIoT) relies on dual-mode communication (HPLC + HRF) for robust data acquisition. However, under massive bursty traffic, conventional static MAC superframe scheduling struggles to reconcile high throughput with stringent reliability constraints. To mitigate this, [...] Read more.
The extensive deployment of the Power Internet of Things (PIoT) relies on dual-mode communication (HPLC + HRF) for robust data acquisition. However, under massive bursty traffic, conventional static MAC superframe scheduling struggles to reconcile high throughput with stringent reliability constraints. To mitigate this, we propose a dynamic adaptive scheduling scheme. Initially, a joint PHY-MAC layer dual-mode system architecture is proposed. At the MAC layer, a dual-link parallel multiplexing contention access mechanism is applied; at the physical layer, a capacity bottleneck determination model is established, incorporating log-normal–Bernoulli–Gaussian mixed noise and multipath fading. Subsequently, an extended two-dimensional Markov chain analytically derives key performance indicators, including equivalent collision probability, joint outage probability, access delay, and network throughput. Building upon this, a Q-learning-based algorithm is proposed. By constructing an asymmetric penalty–reward function, the central coordinator (CCO) autonomously optimizes the Contention Access Period (CAP) to Contention-Free Period (CFP) ratio under dynamic node scales. Simulations demonstrate this methodology effectively averts channel congestion during extreme concurrent traffic surges. Ultimately, it strictly preserves service reliability while substantially augmenting the concurrent carrying capacity and resource utilization of the dual-mode network. Full article
(This article belongs to the Special Issue Advances in Networked Systems and Communication Protocols)
Show Figures

Figure 1

19 pages, 1098 KB  
Article
A Consistent Markov Chain-Based Framework for Life-Cycle Optimization and Cost–Benefit Evaluation of Infrastructure Maintenance Policies
by Artur Zbiciak, Dariusz Walasek, Aleksander Nicał, Mariola Książek-Nowak and Paweł Nowak
Sustainability 2026, 18(15), 7611; https://doi.org/10.3390/su18157611 - 27 Jul 2026
Viewed by 367
Abstract
A consistent computational framework is presented that integrates Markov chain modeling, decision optimization, and cost–benefit analysis for the life-cycle management of engineering assets. The approach combines deterioration modeling with a complete economic evaluation and optimization of maintenance decisions. Each condition state of the [...] Read more.
A consistent computational framework is presented that integrates Markov chain modeling, decision optimization, and cost–benefit analysis for the life-cycle management of engineering assets. The approach combines deterioration modeling with a complete economic evaluation and optimization of maintenance decisions. Each condition state of the system is associated with possible actions such as do-nothing, preventive maintenance, major repair, and replacement, each defined by its own transition matrix or generator describing state changes. The expected one-step reward is formulated as the difference between benefits and all relevant costs including operating, action, and failure costs. The optimization problem is expressed as a discounted Markov decision process and solved by linear programming. The resulting stationary policy specifies the optimal decision rule for every state. Both discrete-time and continuous-time variants are implemented. Transition matrices and generator matrices are linked using a matrix exponential mapping for the selected step length. Under state-dependent policies, the discrete step model and the continuous-time feedback model may lead to different long-run state shares. This is caused by different decision timing. The continuous-time variant also provides reliability indicators such as survival and hazard. It can also provide mean time to absorption under an absorbing failure interpretation. Simulation under the optimal policy yields state trajectories, present values of benefits and costs, and key economic indicators such as net present value, benefit–cost ratio, equivalent annual cost, and equivalent annual net benefit. The framework forms a unified and practical tool that connects reliability analysis, Markov optimization, and life-cycle cost–benefit evaluation for long-term infrastructure management. Full article
(This article belongs to the Section Sustainable Engineering and Science)
Show Figures

Figure 1

22 pages, 663 KB  
Article
State-Dependent Asymmetry in Soft-Pity Gacha Waiting-Time Models: Exact Recurrences, Tail Risk, and Featured-Target Extensions
by Saisai Hou, Yunzhi Zhu and Sen Zhang
Symmetry 2026, 18(6), 1051; https://doi.org/10.3390/sym18061051 - 18 Jun 2026
Viewed by 468
Abstract
Randomized reward mechanisms are often described as repeated trials with a fixed success probability. This constant-hazard reference case is symmetric in the limited finite-state sense that, conditional on non-absorption, the next-draw success probability is invariant with respect to the current draw count. Pity [...] Read more.
Randomized reward mechanisms are often described as repeated trials with a fixed success probability. This constant-hazard reference case is symmetric in the limited finite-state sense that, conditional on non-absorption, the next-draw success probability is invariant with respect to the current draw count. Pity and guarantee rules break this draw-count homogeneity by making the hazard depend on the current state. This paper studies that state-dependent asymmetry for a finite soft-pity waiting-time model. The waiting time for one rare item is represented as an absorption time of a Markov chain whose transient state is the pity counter. We write the corresponding absorbing transition matrix explicitly and then derive the equivalent first-step recurrences for the expectation, variance, and full probability mass function. A simple stochastic-ordering proposition shows how increasing the statewise success probabilities decreases the waiting-time distribution in the usual tail order. Repeated convolution then yields the distribution for multiple independent stages. The numerical section reports quantiles, tail probabilities, VaR/CVaR-type summaries, expected excess values, sensitivity analyses, normal-approximation diagnostics, and distributional asymmetry indicators. A featured-target variant with a binary guarantee state is also included. Throughout, the reported quantities are consequences of the stated transition rule; Monte Carlo simulation is used only as a numerical check. Full article
(This article belongs to the Section B: Mathematics)
Show Figures

Figure 1

23 pages, 3616 KB  
Article
Motion Planning-Augmented Hierarchical Reinforcement Learning for Long-Horizon Mobile Manipulation
by Hyungtai Kim and Mun-Taek Choi
Sensors 2026, 26(12), 3845; https://doi.org/10.3390/s26123845 - 17 Jun 2026
Viewed by 581
Abstract
Long-horizon mobile manipulation requires a robot to execute a sequence of heterogeneous subtasks such as navigation, picking, and articulated-object manipulation in indoor environments. Standard reinforcement learning suffers from reward sparsity and inefficient exploration in this setting, and hierarchical methods often fail at the [...] Read more.
Long-horizon mobile manipulation requires a robot to execute a sequence of heterogeneous subtasks such as navigation, picking, and articulated-object manipulation in indoor environments. Standard reinforcement learning suffers from reward sparsity and inefficient exploration in this setting, and hierarchical methods often fail at the hand-off between consecutive subtasks when the terminal state of one subtask is kinematically infeasible for the next. We propose a motion planning-augmented hierarchical reinforcement learning architecture to resolve the fundamental trade-offs between sample efficiency and hand-off reliability in long-horizon mobile manipulation. The mission is decomposed into subtasks via a Semi-Markov Decision Process; within each subtask, a collision-free reference trajectory generated by RRT* in the full joint configuration space is embedded into the reward as a per-step shaping signal; and a region-goal mechanism, defined analytically from inverse kinematics feasibility, replaces rigid coordinate hand-offs with a continuous feasible region. The architecture is evaluated in the ManiSkill-HAB simulation under teleport-free sequential execution and challenging initialization. The proposed method improves subtask success rate and sample efficiency over the baseline across all six evaluated subtasks, and the advantage compounds along the long-horizon task chain. Full article
(This article belongs to the Topic Robot Manipulation Learning and Interaction Control)
Show Figures

Figure 1

22 pages, 1755 KB  
Article
Dynamic Optimization of Incoming Quality Control Policies for Cost, Carbon, and Energy Reduction Using Bayesian Reinforcement Learning
by David Massetti, Mehdi Raoofi, Tiziano Miroglio, Marco Mosca and Flavio Tonelli
Sustainability 2026, 18(12), 6094; https://doi.org/10.3390/su18126094 - 13 Jun 2026
Viewed by 555
Abstract
The transition towards sustainable manufacturing necessitates complex optimization that integrates economic goals with environmental factors, such as energy consumption and greenhouse gas emissions. This research addresses the critical challenge of optimizing the Incoming Quality Control (IQC) policy for raw material batches. The primary [...] Read more.
The transition towards sustainable manufacturing necessitates complex optimization that integrates economic goals with environmental factors, such as energy consumption and greenhouse gas emissions. This research addresses the critical challenge of optimizing the Incoming Quality Control (IQC) policy for raw material batches. The primary objective is formulated as a multi-criteria control problem that jointly minimizes the weekly final product cost, carbon footprint, and energy consumption. To handle sequential decision making under uncertainty, we adopt a scalarized reinforcement learning (RL) reward that combines these objectives into a single value function and explores different trade-offs through alternative weight configurations. To effectively handle the uncertainty in incoming quality and the sequential decision making required for dynamic control, the optimization problem is modeled as a Bayesian Adaptive Markov Decision Process (BAMDP). To maintain computational tractability despite the continuous belief space inherent in the BAMDP formulation, we employ a Deep Q-Network (DQN) architecture acting as an approximate dynamic programming solver. The Bayesian framework represents model uncertainty explicitly, updates beliefs as new inspection evidence becomes available, and allows prior domain knowledge on supplier quality to be incorporated into the learning process. The BAMDP formulation is used to learn a set of adaptive inspection policies that adjust the IQC strategy over time to achieve conflicting goals: reducing inspection costs while maintaining standard quality, minimizing energy consumption, and lowering CO2-equivalent emissions. The goal is to find robust policies that balance these trade-offs under different quality and demand conditions. This methodology aligns with the principles of Industry 5.0 by leveraging advanced artificial intelligence (AI) methods, such as reinforcement learning (RL), coupled with a stochastic simulation of the production system, based on a geometric/physical model of the component’s tolerance chains, to support decision-makers in designing and assessing sustainable IQC strategies. Comparative simulations on the case study, including a benchmark against ISO 2859-1 sampling plans, confirm that this dynamic and risk-aware optimization paradigm can reduce overall cost, energy use, and environmental impact across various quality conditions, while preserving outgoing quality. Full article
Show Figures

Figure 1

17 pages, 6836 KB  
Article
CReSCENT for Long-Term Value-Driven Scheduling in Multi-Layer Industrial Networks with Milestone-Triggered Rewards
by Wei Xu, Yi Wan and Tianyu Zuo
Algorithms 2026, 19(6), 443; https://doi.org/10.3390/a19060443 - 1 Jun 2026
Cited by 1 | Viewed by 540
Abstract
This paper studies long-term value-driven scheduling in multi-layer industrial networks where useful external reward is released mainly at milestone transitions. The delayed feedback causes supervision degeneracy before milestones, because many trajectory prefixes receive the same external return even when their downstream potential is [...] Read more.
This paper studies long-term value-driven scheduling in multi-layer industrial networks where useful external reward is released mainly at milestone transitions. The delayed feedback causes supervision degeneracy before milestones, because many trajectory prefixes receive the same external return even when their downstream potential is different. We formulate the setting as a time-delayed multi-industrial-chain Markov decision process and present CReSCENT as a milestone-aware structural exploration framework rather than a new reinforcement learning principle. CReSCENT combines macro structural feature construction, contrastive representation learning, milestone-weighted online clustering, and cross-layer credit allocation. It moves intrinsic learning from raw observations to milestone-relevant structural states, then distributes the resulting signal to layers according to their contribution to cross-layer progress. The revised evaluation gives the simulator transition rules, a formal utility metric, an external-reward-only baseline, confidence intervals, and statistical tests. Experiments show that CReSCENT improves utility across layer, task-load, worker-count, and episode-length perturbations against ETA-PSI, EMU, MIMEx, and the external-only baseline. Sensitivity studies further show that the clustering radius and credit-allocation weights have stable operating ranges. Full article
Show Figures

Figure 1

21 pages, 436 KB  
Article
Mean Extinction Times in Multi-Metastable Systems: A Discrete Coarse-Grained Approach
by Santosh Kumar Kudtarkar
Physics 2026, 8(1), 30; https://doi.org/10.3390/physics8010030 - 2 Mar 2026
Viewed by 634
Abstract
The paper develops a coarse-grained framework for computing mean extinction times in multi-metastable systems modeled as one-step continuous-time Markov chains with an absorbing state. At the microscopic level, backward equations on finite corridors are solved to obtain closed-form series for committors, mean first-passage [...] Read more.
The paper develops a coarse-grained framework for computing mean extinction times in multi-metastable systems modeled as one-step continuous-time Markov chains with an absorbing state. At the microscopic level, backward equations on finite corridors are solved to obtain closed-form series for committors, mean first-passage times, and intrawell (basin) waiting times. A renewal–reward construction then yields effective interwell transition rates written as a success probability divided by a mean cycle duration, providing an interpretable effective rate constant. These rates define a reduced Markov chain on the wells together with extinction; mean extinction times follow from a linear system, and the associated fundamental matrix quantifies pre-extinction residence times in each coarse state. This framework makes explicit how multiple escape pathways and intrawell dwell times contribute to extinction statistics in finite systems. The method is illustrated on a double-well landscape with an extinction state, using a reversible potential-to-rates mapping for the numerical example. Comparisons of alternative intrawell models and validation against exact one-step computations demonstrate accuracy at finite system sizes, including regimes where diffusion approximations are unreliable. The resulting formulas require only local rate data, remain numerically stable under strong bias, and extend directly to multiple wells and flexible boundary conditions. Full article
(This article belongs to the Section Statistical Physics and Nonlinear Phenomena)
Show Figures

Figure 1

20 pages, 4053 KB  
Article
Higher-Order Markov Model-Based Analysis of Reinforcement Learning in 6G Mobile Retrial Queueing Systems
by Djamila Talbi and Zoltan Gal
Sensors 2025, 25(23), 7245; https://doi.org/10.3390/s25237245 - 27 Nov 2025
Cited by 1 | Viewed by 1303
Abstract
The dynamic behavior of the retrial queueing system following the incorporation of Deep Q-Network Reinforcement Learning in 6G mobile communication services is examined in this study. The proposed method lies in analyzing the DQN-RL agent’s learning convergence by using the first- and second-order [...] Read more.
The dynamic behavior of the retrial queueing system following the incorporation of Deep Q-Network Reinforcement Learning in 6G mobile communication services is examined in this study. The proposed method lies in analyzing the DQN-RL agent’s learning convergence by using the first- and second-order Markov chain method. By simulating the temporal evolution of reward sequences as Markov and second-order Markov chains, we can quantify convergence characteristics through mixing time analysis. To capture a wide operational landscape, a thorough simulation framework with 120 independent parameter combinations is created. The obtained results indicate that Markov chain analysis confirms 10 training episodes are more than sufficient for policy convergence, and in some cases, as few as 5 episodes allow the agent to enhance the mobile network performance while maintaining low energy consumption. To assess learning stability and system responsiveness, the mixing time of DQN RL rewards is calculated for every episode and configuration. A deeper understanding of the temporal dependencies in the reward process can be gained by incorporating higher-order Markov models. This paper concentrates on studying the learning convergence using an analysis of the Markov model’s spectral gap properties as an indicator. The results provide a rigorous foundation for optimizing 6G queueing strategies under uncertainty by highlighting the sensitivity of DQN convergence to system parameters and retrial dynamics. Full article
(This article belongs to the Special Issue Feature Papers in Communications Section 2025–2026)
Show Figures

Figure 1

19 pages, 509 KB  
Article
Zero-Inflated Distributions of Lifetime Reproductive Output
by Hal Caswell
Populations 2025, 1(3), 19; https://doi.org/10.3390/populations1030019 - 23 Aug 2025
Viewed by 1924
Abstract
Lifetime reproductive output (LRO), also called lifetime reproductive success (LRS) is often described by its mean (total fertility rate or net reproductive rate), but it is in fact highly variable among individuals and often positively skewed. Several approaches exist to calculating the variance [...] Read more.
Lifetime reproductive output (LRO), also called lifetime reproductive success (LRS) is often described by its mean (total fertility rate or net reproductive rate), but it is in fact highly variable among individuals and often positively skewed. Several approaches exist to calculating the variance and skewness of LRO. These studies have noted that a major factor contributing to skewness is the fraction of the population that dies before reaching a reproductive age or stage. The existence of that fraction means that LRO has a zero-inflated distribution. This paper shows how to calculate that fraction and to fit a zero-inflated Poisson or zero-inflated negative binomial distribution to the LRO. We present a series of applications to populations before and after demographic transitions, to populations with particularly high probabilities of death before reproduction, and a couple of large mammal populations for good measure. The zero-inflated distribution also provides extinction probabilities from a Galton-Watson branching process. We compare the zero-inflated analysis with a recently developed analysis using convolution methods that provides exact distributions of LRO. The agreement is strikingly good. Full article
Show Figures

Figure 1

49 pages, 2072 KB  
Article
Game Theory for Predicting Stocks’ Closing Prices
by João Costa Freitas, Alberto Adrego Pinto and Óscar Felgueiras
Mathematics 2024, 12(17), 2676; https://doi.org/10.3390/math12172676 - 28 Aug 2024
Cited by 1 | Viewed by 5506
Abstract
We model the financial markets as a game and make predictions using Markov chain estimators. We extract the possible patterns displayed by the financial markets, define a game where one of the players is the speculator, whose strategies depend on his/her risk-to-reward preferences, [...] Read more.
We model the financial markets as a game and make predictions using Markov chain estimators. We extract the possible patterns displayed by the financial markets, define a game where one of the players is the speculator, whose strategies depend on his/her risk-to-reward preferences, and the market is the other player, whose strategies are the previously observed patterns. Then, we estimate the market’s mixed probabilities by defining Markov chains and utilizing its transition matrices. Afterwards, we use these probabilities to determine which is the optimal strategy for the speculator. Finally, we apply these models to real-time market data to determine its feasibility. From this, we obtained a model for the financial markets that has a good performance in terms of accuracy and profitability. Full article
(This article belongs to the Special Issue Econophysics, Financial Markets, and Artificial Intelligence)
Show Figures

Figure 1

16 pages, 392 KB  
Article
Some Probabilistic Interpretations Related to the Next-Generation Matrix Theory: A Review with Examples
by Florin Avram, Rim Adenane and Lasko Basnarkov
Mathematics 2024, 12(15), 2425; https://doi.org/10.3390/math12152425 - 4 Aug 2024
Cited by 1 | Viewed by 3698
Abstract
The fact that the famous basic reproduction number R0, i.e., the largest eigenvalue of the next generation matrix FV1, sometimes has a probabilistic interpretation is not as well known as it deserves to be. It is well [...] Read more.
The fact that the famous basic reproduction number R0, i.e., the largest eigenvalue of the next generation matrix FV1, sometimes has a probabilistic interpretation is not as well known as it deserves to be. It is well understood that half of this formula, V, is a Markovian generating matrix of a continuous-time Markov chain (CTMC) modeling the evolution of one individual on the compartments. It has also been noted that the not well-enough-known rank-one formula for R0 of Arino et al. (2007) may be interpreted as an expected final reward of a CTMC, whose initial distribution is specified by the rank-one factorization of F. Here, we show that for a large class of ODE epidemic models introduced in Avram et al. (2023), besides the rank-one formula, we may also provide an integral renewal representation of R0 with respect to explicit “age kernels” a(t), which have a matrix exponential form.This latter formula may be also interpreted as an expected reward of a probabilistic continuous Markov chain (CTMC) model. Besides the rather extensively studied rank one case, we also provide an extension to a case with several susceptible classes. Full article
Show Figures

Figure 1

18 pages, 2319 KB  
Article
Handling Efficient VNF Placement with Graph-Based Reinforcement Learning for SFC Fault Tolerance
by Seyha Ros, Prohim Tam, Inseok Song, Seungwoo Kang and Seokhoon Kim
Electronics 2024, 13(13), 2552; https://doi.org/10.3390/electronics13132552 - 28 Jun 2024
Cited by 16 | Viewed by 4142
Abstract
Network functions virtualization (NFV) has become the platform for decomposing the sequence of virtual network functions (VNFs), which can be grouped as a forwarding graph of service function chaining (SFC) to serve multi-service slice requirements. NFV-enabled SFC consists of several challenges in reaching [...] Read more.
Network functions virtualization (NFV) has become the platform for decomposing the sequence of virtual network functions (VNFs), which can be grouped as a forwarding graph of service function chaining (SFC) to serve multi-service slice requirements. NFV-enabled SFC consists of several challenges in reaching the reliability and efficiency of key performance indicators (KPIs) in management and orchestration (MANO) decision-making control. The problem of SFC fault tolerance is one of the most critical challenges for provisioning service requests, and it needs resource availability. In this article, we proposed graph neural network (GNN)-based deep reinforcement learning (DRL) to enhance SFC fault tolerance (GRL-SFT), which targets the chain graph representation, long-term approximation, and self-organizing service orchestration for future massive Internet of Everything applications. We formulate the problem as the Markov decision process (MDP). DRL seeks to maximize the cumulative rewards by maximizing the service request acceptance ratios and minimizing the average completion delays. The proposed model solves the VNF management problem in a short time and configures the node allocation reliably for real-time restoration. Our simulation result demonstrates the effectiveness of the proposed scheme and indicates better performance in terms of total rewards, delays, acceptances, failures, and restoration ratios in different network topologies compared to reference schemes. Full article
(This article belongs to the Special Issue Recent Advances of Cloud, Edge, and Parallel Computing)
Show Figures

Figure 1

20 pages, 10725 KB  
Article
AARF: Autonomous Attack Response Framework for Honeypots to Enhance Interaction Based on Multi-Agent Dynamic Game
by Le Wang, Jianyu Deng, Haonan Tan, Yinghui Xu, Junyi Zhu, Zhiqiang Zhang, Zhaohua Li, Rufeng Zhan and Zhaoquan Gu
Mathematics 2024, 12(10), 1508; https://doi.org/10.3390/math12101508 - 11 May 2024
Cited by 5 | Viewed by 3507
Abstract
Highly interactive honeypots can form reliable connections by responding to attackers to delay and capture intranet attacks. However, current research focuses on modeling the attacker as part of the environment and defining single-step attack actions by simulation to study the interaction of honeypots. [...] Read more.
Highly interactive honeypots can form reliable connections by responding to attackers to delay and capture intranet attacks. However, current research focuses on modeling the attacker as part of the environment and defining single-step attack actions by simulation to study the interaction of honeypots. It ignores the iterative nature of the attack and defense game, which is inconsistent with the correlative and sequential nature of actions in real attacks. These limitations lead to insufficient interaction of the honeypot response strategies generated by the study, making it difficult to support effective and continuous games with attack behaviors. In this paper, we propose an autonomous attack response framework (named AARF) to enhance interaction based on multi-agent dynamic games. AARF consists of three parts: a virtual honeynet environment, attack agents, and defense agents. Attack agents are modeled to generate multi-step attack chains based on a Hidden Markov Model (HMM) combined with the generic threat framework ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge). The defense agents iteratively interact with the attack behavior chain based on reinforcement learning (RL) to learn to generate honeypot optimal response strategies. Aiming at the sample utilization inefficiency problem of random uniform sampling widely used in RL, we propose the dynamic value label sampling (DVLS) method in the dynamic environment. DVLS can effectively improve the sample utilization during the experience replay phase and thus improve the learning efficiency of honeypot agents under the RL framework. We further couple it with a classic DQN to replace the traditional random uniform sampling method. Based on AARF, we instantiate different functional honeypot models for deception in intranet scenarios. In the simulation environment, honeypots collaboratively respond to multi-step intranet attack chains to defend against these attacks, which demonstrates the effectiveness of AARF. The average cumulative reward of the DQN with DVLS is beyond eight percent, and the convergence speed is improved by five percent compared to a classic DQN. Full article
(This article belongs to the Special Issue Advanced Research on Information System Security and Privacy)
Show Figures

Figure 1

Back to TopTop