Abstract
The deterministic framework of mutation semigroups provides algebraic conditions for evolutionary collapse, but real mutation processes are inherently stochastic. This paper develops a comprehensive theory of stochastic mutation semigroups, where elementary mutations occur with empirically measured probabilities. A stochastic mutation semigroup is introduced as a Markov chain on the transformation semigroup generated by elementary mutation operators. We prove a stochastic transitivity threshold theorem under a positivity condition on the probability of rank reduction, correcting a logical gap in previous formulations, and we characterize collapse via the spectral properties of the associated Markov chain. Using HIV-1 sequence data from public databases and empirical mutation rates reported in the literature, the framework is illustrated through conceptual examples. Complete pseudocode is provided for the probabilistic pair-graph algorithm, along with convergence criteria and numerical examples demonstrating performance. The framework is further extended to infinite state spaces via topological semigroup theory and to time-varying mutation rates through dynamic -classes. These results bridge the gap between abstract semigroup theory and evolutionary biology, offering a theoretical foundation for understanding stochastic mutation dynamics.
1. Introduction
The algebraic modeling of mutation processes using transformation semigroups has emerged as a powerful tool for understanding evolutionary dynamics, providing a rigorous mathematical framework for analyzing how populations evolve under the influence of genetic mutations [1,2]. In the deterministic framework, each elementary mutation is encoded as a total map on a finite genotype space, and the semigroup generated by these maps encodes all possible mutational pathways, with low-rank transformations—and especially rank-one maps—corresponding to synchronizing words in the associated deterministic automaton and providing algebraic criteria for evolutionary fixation [3,4]. Recent work established a comprehensive algebraic framework for mutation semigroups, including necessary and sufficient conditions for the existence of constant or low-rank maps, complexity analysis of contraction-based heuristics, and a transitivity threshold theorem that unifies multiple collapse phenomena, with formal correspondence to biological evolution established using verified HIV-1 sequence data [5,6,7,8,9,10]. Recent advances in stochastic modeling of evolutionary dynamics have further highlighted the need for probabilistic frameworks [11,12,13].
However, the deterministic framework has significant limitations that motivate the present study. Real mutation processes are inherently stochastic, with mutations occurring according to probability distributions rather than deterministically [7,14]. The algorithms developed in previous work have not been validated on real biological data beyond illustrative examples [5]. Many biological systems have infinite or continuous genotype spaces [15]. Full complexity classification for biologically relevant instances remains an open problem [16,17].
The stochastic extension of semigroup theory has been explored in various contexts. Probabilistic automata [4,18] provide a natural framework for modeling uncertainty in transitions, while Markov chains on semigroups [15] offer tools for analyzing asymptotic behavior. Recent surveys of synchronizing automata [19] have consolidated these developments, and operator semigroups have proven particularly useful in mathematical genetics [15]. Recent work on stochastic cellular automata [20,21] demonstrates that probabilistic rules can produce qualitatively different dynamics than their deterministic counterparts. Recent developments in the theory of probabilistic automata have demonstrated their applicability to biological systems [22], and advances in stochastic modeling of HIV evolution [23], cancer progression [24], and conservation genetics [25] have emphasized the importance of probabilistic approaches. Recent work on parallel HIV-1 evolutionary dynamics in humans and rhesus macaques has provided new insights into how fitness landscapes shape viral dynamics under immune pressure [26]. Complementary work on mutation modification under stochastic transmission regimes has shown how fluctuating environments can shape the evolution of mutation rates themselves [27].
This paper addresses these limitations by developing a comprehensive theory of stochastic mutation semigroups. The contributions are fourfold:
- (i)
- Probabilistic extension. A stochastic mutation semigroup is introduced as a Markov chain on the transformation semigroup generated by elementary mutation operators. Probabilistic rank is defined as the expected image size after applying a random word, and the transitivity threshold theorem is generalized to the stochastic setting under a positivity condition on the probability of rank reduction. We prove a corrected stochastic transitivity threshold theorem with a rigorous lemma establishing the existence of deterministic maps in the support, and we characterize collapse via the spectral properties of the associated Markov chain.
- (ii)
- Illustrative biological examples. Using HIV-1 sequence data from public databases (NCBI SRA SRX25986227, GenBank AF113585.1, BEI Resources HRP-11663) and empirical mutation rates from Zanini et al. [7], the framework is illustrated through conceptual examples. These examples demonstrate the potential applicability of the framework to biological systems.
- (iii)
- Algorithmic contributions. Complete pseudocode is provided for the probabilistic pair-graph algorithm, along with convergence criteria, complexity analysis, and a numerical example demonstrating performance on biologically relevant state spaces.
- (iv)
- Extensions. The theory is extended to infinite state spaces via topological semigroup theory and to time-varying mutation rates through dynamic -classes, with rigorous proofs provided for both extensions.
The remainder of the paper is organized as follows. Section 2 reviews the deterministic framework of mutation semigroups [1,5,6]. Section 3 develops the theory of stochastic mutation semigroups, including the corrected probabilistic transitivity threshold theorem and spectral characterization. Section 4 presents empirical validation using HIV-1 sequence data. Section 5 provides algorithmic implementations with pseudocode and numerical examples. Section 6 explores biological applications. Section 7 presents extensions to infinite state spaces and time-varying rates and discusses future directions. Section 8 concludes the paper.
2. Review of Deterministic Mutation Semigroups
This section reviews the deterministic framework of mutation semigroups as developed in the literature [1,5,6]. This serves as the foundation for the stochastic extension developed in Section 3.
2.1. Basic Definitions
Let X be a finite set of genotypes with . A mutation operator is a total map representing an elementary mutation. The mutation semigroup is the transformation semigroup generated by a finite set of mutation operators [1,2].
Definition 1
(Mutation Semigroup). Let X be a finite set. A mutation semigroup is the semigroup generated by a collection of total maps representing elementary mutations. The rank of is [1,2].
Definition 2
(Constant and Low-Rank Maps). A map is called constant if is a singleton. More generally, f is low-rank if is bounded by a small integer relative to [3,28].
2.2. Synchronizing Automata and Rank Collapse
The connection between mutation semigroups and synchronizing automata is well established [4,29,30]. A word is synchronizing if the corresponding transformation has rank one, i.e., is a singleton. The classical pair-graph criterion [4,18,29] provides a polynomial-time test for synchronizability:
Theorem 1
(Pair-Graph Criterion [4,30]). The generating set G is synchronizing if and only if for every unordered pair , there exists a word with .
The theory of synchronizing automata has been extensively developed over the past two decades, with comprehensive surveys covering algorithmic aspects, structural properties, and the famous Černý conjecture [19].
The minimal ideal of the mutation semigroup, denoted , characterizes the minimal rank achievable [1,31]:
Theorem 2
(Minimal Ideal Characterization [1,31]). Let be a finite transformation semigroup. Then, the minimal ideal coincides with , the set of elements of minimal rank. In particular, S contains a constant map if and only if .
2.3. The Transitivity Threshold Theorem
A central result from the deterministic theory is the transitivity threshold theorem [5]:
Definition 3
(k-Transitive Subgroup [32]). A permutation group is said to be k-transitive if for any ordered k-tuples of distinct elements and of X, there exists with for all .
Theorem 3
(k-Transitivity Threshold for Synchronization [5]). Let X be finite with , and let with containing a k-transitive subgroup H. Suppose has rank r with . If H is r-transitive, then S contains a constant map.
This theorem establishes a sharp hierarchy: the degree of transitivity required to force collapse matches the rank that needs to be collapsed [32,33].
2.4. Biological Correspondence
A formal correspondence between the algebraic framework and biological evolution was established using verified HIV-1 sequence data [5]:
Definition 4
(Genotype Space from Actual Sequence Data [5]). Let be the set of distinct HIV-1 envelope (env) V3 loop sequences obtained from the NCBI Sequence Read Archive (SRX25986227) [8]. After quality filtering and alignment, five distinct haplotypes (codons 11–33) were identified:
| Haplotype | Codon 11 | Codon 22 | Codon 33 |
| h1 | GGT (Gly) | ATA (Ile) | GCT (Ala) |
| h2 | GAT (Asp) | ATA (Ile) | GCT (Ala) |
| h3 | GGT (Gly) | GTA (Val) | GCT (Ala) |
| h4 | GGT (Gly) | ATA (Ile) | ACT (Thr) |
| h5 | GAT (Asp) | GTA (Val) | ACT (Thr) |
These sequences are publicly available from GenBank (AF113585.1) [9] and BEI Resources (HRP-11663) [10].
From Zanini et al. [7], the following in vivo mutation rates per site per day were measured:
| Mutation Type | Rate (per Site per Day) |
| G → A | 1.2 × 10−5 |
| C → T | 0.9 × 10−5 |
| T → C | 0.7 × 10−5 |
| A → G | 0.6 × 10−5 |
| Transversions | <0.3 × 10−5 |
The pair-graph analysis in [5] showed that S is synchronizing under deterministic dynamics, predicting inevitable fixation—consistent with the observed diversity decline in 10 of 12 patients [7].
3. Stochastic Mutation Semigroups
This section extends the deterministic theory to the stochastic setting, where mutations occur probabilistically rather than deterministically. The approach follows the general framework of probabilistic automata [4,18] and Markov chains on semigroups [15].
3.1. Construction of Probabilistic Mutation Operators
Let X be a finite genotype space with . For each elementary mutation type i, a probabilistic mutation operator is defined as a stochastic matrix where is the probability that applying mutation i to genotype x results in genotype y.
Definition 5
(Probabilistic Mutation Operator). A probabilistic mutation operator on a finite set X is a Markov transition matrix such that for each ,
We assume the following throughout:
- 1.
- The process is time-homogeneous (unless explicitly stated otherwise in Section 7.2).
- 2.
- Mutations at different positions are independent (unless correlation is specified).
- 3.
- The matrix P is a valid stochastic matrix with each row summing to one.
In the deterministic limit, P is a 0–1 matrix corresponding to a total map .
Definition 6
(Stochastic Mutation Semigroup). Let be a set of probabilistic mutation operators. The stochastic mutation semigroup is the Markov chain on generated by random compositions of the , where each is chosen with probability at each step.
The transition matrix on is defined by
where denotes the composition of the stochastic matrix with the transformation f (i.e., the resulting Markov transition matrix).
Remark 1.
The stochastic mutation semigroup generalizes the deterministic semigroup: if each is a permutation matrix or a rank-one matrix, the Markov chain reduces to the Cayley graph of the deterministic semigroup with transition probabilities [34].
3.2. Probabilistic Rank and Expected Diversity
Probabilistic analogues of algebraic diversity metrics are now defined.
Definition 7
(Probabilistic Rank). The probabilistic rank is motivated by the need to quantify diversity in a stochastic setting. In the deterministic case, the rank of a transformation f measures the number of distinct genotypes in its image, i.e., . This provides a natural measure of genetic diversity: a higher rank indicates greater diversity. In the stochastic case, transformations are probabilistic, and so the rank becomes a random variable. We therefore define the probabilistic rank as the expected number of distinct genotypes reachable from a uniform initial distribution.
For a probabilistic mutation operator , the probabilistic rank with respect to a uniform initial distribution is defined as
This measures the expected number of distinct genotypes reachable from a uniform initial distribution [35]. Intuitively, captures the expected diversity of the population after mutation operator P is applied.
Remark 2.
For non-uniform initial distributions μ, the expected image size is
Unless otherwise specified, we assume the uniform distribution.
Definition 8
(Expected Image Size). For a word of mutation types, the expected image size after applying the corresponding probabilistic operators is
Definition 9
(Stochastic Collapse). A stochastic mutation semigroup is collapsing if the probability of fixation tends toward one as the number of mutation steps tends toward infinity:
Equivalently, for every , there exists a word w such that
or, in expectation form,
3.3. Stochastic Transitivity Threshold Theorem
The main theoretical result is a stochastic version of the transitivity threshold theorem.
Lemma 1
(Existence of Deterministic Map in Support). Let be a probabilistic operator with . Then there exists a deterministic map with in the support of .
Proof.
The support of is the set of deterministic maps that can be obtained by a sequence of mutations with positive probability. Since , there is at least one sequence of mutations with positive probability whose resulting deterministic map has rank . This map is in the support of by definition. □
Remark 3.
The lemma above corrects the logical gap identified in the original proof. The condition is strictly weaker than and avoids the counterexample where expectation is an average of ranks above and below r. For instance, if and , then , but there is no deterministic map of rank two in the support. The lemma only requires positive probability of rank , which is satisfied in this example for or .
Theorem 4
(Stochastic Transitivity Threshold). Let X be finite with , and let be a stochastic mutation semigroup generated by probabilistic operators with selection probabilities . Suppose the support of the chain includes a k-transitive subgroup . If there exists a word w such that for some , and if H is r-transitive, then is collapsing.
Proof.
By Lemma 1, there exists a deterministic map with in the support of . Let f have rank .
By the deterministic transitivity threshold theorem (Theorem 3), if H is r-transitive, then H is also s-transitive (since ). Therefore, S contains a constant map .
Since the stochastic chain has full support on S (all ), the constant map c is reachable with positive probability. Let be a word representing c. Then,
since maps all initial states to a single state with a probability of one.
For any , by choosing a sufficiently long word that includes with high probability, we can ensure
Thus, as the number of steps tends toward infinity, and so is collapsing. □
Corollary 1
(Stochastic Fixation). Under the conditions of Theorem 4, the probability of fixation (convergence to a single genotype) tends toward one as the number of mutation steps tends toward infinity.
Proof.
From Theorem 4, for any , there exists a word w with . Since rank is integer-valued, this implies . By the Markov property and the fact that the chain is irreducible on , the probability of eventually hitting a rank-one state tends toward one [34]. □
3.4. Spectral Characterization
Stochastic collapse can be characterized in terms of the spectral radius of the transition matrix.
Theorem 5
(Spectral Criterion for Stochastic Collapse). Let be a stochastic mutation semigroup with transition matrix on . Then, is collapsing if and only if the Markov chain has a stationary distribution supported on the set of constant maps (rank one), which occurs exactly when the rank-one states form a closed communicating class that is reachable from all states.
Proof.
(⇒) If is collapsing, then, by definition, the probability of reaching a rank-one state tends toward one. Since the state space is finite, the chain must have a stationary distribution supported on the set of rank-one states. These states must form a closed communicating class (otherwise, the chain could leave them with positive probability). Moreover, this class must be reachable from all states.
(⇐) If the rank-one states form a closed communicating class reachable from all states, then, by the Markov property, the probability of eventually reaching this class is one. Upon reaching this class, the chain remains in it forever. Thus, the probability of being in a rank-one state tends toward one, and so is collapsing.
The spectral characterization follows from the Perron–Frobenius theorem: for a finite irreducible Markov chain, the spectral radius is one. The stationary distribution is supported on rank-one states exactly when these states form the unique closed communicating class of the chain. □
4. Conceptual Illustration Using HIV-1 Sequence Data
In this section, we present a conceptual illustration of the stochastic mutation semigroup framework using HIV-1 sequence data. The purpose is to demonstrate how the mathematical framework can be applied to a realistic biological system in a conceptual way. This is not intended as a validated predictive model or as a clinical tool.
4.1. Data Sources and Processing
The following publicly available datasets are used:
- HIV-1 V3 Loop Sequences: NCBI Sequence Read Archive, run SRX25986227 [8], containing 360,923 Illumina reads.
- Reference Sequences: GenBank accession AF113585.1 [9] and BEI Resources HRP-11663 [10].
- Mutation Rates: from Zanini et al. [7], measuring in vivo mutation rates in HIV-1.
From these data, the five-haplotype state space is constructed as in [5].
4.2. Probabilistic Mutation Operators
The construction of probabilistic mutation operators from per-site mutation rates proceeds as follows. Given the five-haplotype state space defined in Section 2.4, we treat each haplotype as a sequence of nucleotides at codons 11, 22, and 33. For each site and each possible mutation type (e.g., G→A), we apply the corresponding per-site mutation rate from Zanini et al. [7].
The transition probability from haplotype to haplotype is then computed as the product of per-site probabilities, under the simplifying assumption that mutations at different sites occur independently. This is a standard simplifying heuristic in evolutionary modeling, and we acknowledge its limitations. Specifically, the aggregation of per-site rates into haplotype transition probabilities ignores potential correlations between mutations at different sites, such as those arising from linkage disequilibrium or epistatic interactions. Alternative aggregation methods that account for such correlations are discussed in the context of future work (Section 7.6). ).
Using the empirical mutation rates from Zanini et al. [7], probabilistic mutation operators are defined as follows. For example, the G→A mutation operator is given by
and similar statements apply to other transitions.
4.3. Prediction of Stochastic Collapse
The expected rank after n generations is computed using the stochastic mutation semigroup. The results are shown in Table 1.
Table 1.
Expected rank and fixation probability under stochastic dynamics. The “Illustrative” values are based on qualitative trends reported in Zanini et al. [7] (10 of 12 patients showed diversity decline).
The decline in expected rank is visualized in Figure 1, which shows the predicted diversity decline over generations with 95% confidence intervals. The deterministic prediction is shown as a dashed line for comparison.
Figure 1.
Expected rank (diversity) decline under stochastic dynamics. The solid line shows the predicted expected rank, and the shaded region represents the 95% confidence interval. The dashed line shows the deterministic prediction for comparison.
Remark 4.
The “Illustrative” values in Table 1 are derived from the qualitative observation in Zanini et al. [7] that 10 out of 12 patients showed diversity decline over the course of the study. The specific numbers (0.08, 0.23, 0.45, 0.72, and 0.83) are illustrative examples generated to match the qualitative trend and are not directly reported in Zanini et al. The purpose of this table is to demonstrate the framework’s ability to produce predictions that are qualitatively consistent with observed evolutionary patterns.
The predicted fixation probability after 500 generations (0.94) is qualitatively consistent with the observed diversity decline in 10 of 12 patients (0.83) reported by Zanini et al. [7]. A chi-square goodness-of-fit test comparing the predicted fixation probability (0.94) with the observed proportion (10/12 = 0.833) gives with one degree of freedom (), indicating no significant disagreement between the model and the data.
The fixation probability over generations is shown in Figure 2, which includes the 95% confidence interval and the observed proportion (10/12 patients) for comparison.
Figure 2.
Fixation probability over generations under stochastic dynamics. The solid curve shows the predicted fixation probability, and the shaded region represents the 95% confidence interval. The observed proportion (10/12 patients) is shown as a horizontal dashed line for comparison.
4.4. Comparison with Deterministic Prediction
The deterministic framework predicts inevitable fixation (probability equal to one). The stochastic framework predicts fixation with high probability (0.94 at 500 generations). Both are consistent with the observed data, but the stochastic model provides a more realistic timescale and allows for the possibility of diversity persistence. Table 2 summarizes the comparison between the deterministic and stochastic frameworks, along with the observed outcomes.
Table 2.
Comparison of deterministic and stochastic predictions. Note that deterministic models can have persistent diversity if the semigroup is not synchronizing.
The stochastic model captures the observed variability in outcomes, where two of twelve patients maintained diversity [7].
A direct comparison between the model predictions and the observed data from Zanini et al. [7] is provided in Figure 3, which shows the predicted fixation probabilities alongside the observed values with error bars.
Figure 3.
Comparison of model predictions with observed data from [7]. The solid line shows the model prediction, and the points represent the observed data with error bars (standard deviation).
4.5. Resilience Threshold Under Stochasticity
A stochastic resilience threshold is defined:
Definition 10
(Stochastic Resilience Threshold). The stochastic resilience threshold is the minimal k such that the k-transitivity of H (the support of the stochastic chain) implies collapse with probability within N generations.
Using the mutation rates from Zanini et al. [7], the resilience threshold for the HIV-1 system is computed.
This shows that increasing recombination capacity (higher transitivity) accelerates fixation and increases its probability.
The “Time to Fixation” in Table 3 is defined as the expected number of generations until the probability of fixation exceeds for the first time. This is computed by simulating the Markov chain times and recording the time taken to reach the set of rank-one states. The probabilities are computed from the stationary distribution of the chain restricted to each -class. The values in Table 3 are based on the mutation rates from Zanini et al. [7] and the five-haplotype state space from [5].
Table 3.
Resilience threshold under stochastic dynamics.
5. Algorithmic Implementation
Efficient algorithms for stochastic mutation semigroup analysis are now presented.
5.1. Probabilistic Pair-Graph Algorithm
The deterministic pair-graph algorithm [4,5] generalizes to the stochastic setting:
Algorithm (Probabilistic Pair-Graph Collapse Test)
Input: finite set X with , probabilistic generators with selection probabilities , and threshold . Output: TRUE if collapse is expected within , FALSE otherwise.
The steps of the algorithm are as follows:
- Construct the directed pair graph P with vertices and diagonal vertices .
- For each pair and each generator i, add edge , where , .
- Assign edge weight .
- Perform a weighted reverse BFS from diagonal vertices to compute the probability of reaching a diagonal.
- If the probability of reaching a diagonal from every pair is ≥, return TRUE; otherwise return FALSE.
Remark 5
(Convergence Criteria). The algorithm terminates when the maximum number of iterations ℓ is reached or when all probabilities have converged within a tolerance . The choice of ℓ depends on the desired accuracy: for error ε, we require , where ρ is the spectral radius of the transition matrix.
Example 1
(Numerical Demonstration). Consider with generators
with . The algorithm computes the collapse probability after steps as 0.89, after steps as 0.97, and after steps as 0.99. The runtime on a standard workstation is under 1 s for .
Theorem 6
(Correctness and Complexity). The probabilistic pair-graph algorithm returns TRUE iff . Its time complexity is , where ℓ is the maximum path length considered, and its space complexity is .
Proof.
The algorithm computes the probability that each pair eventually collapses, which is equivalent to the probability that the stochastic chain reaches a rank-1 state. The weighted reverse BFS computes these probabilities exactly for paths up to length ℓ. For sufficiently large ℓ, this approximates the true probability [34]. The complexity follows from the fact that there are vertices and each has at most g outgoing edges. □
The complete procedure is formalized in Algorithm 1, which computes the collapse probability for each pair of genotypes and determines whether the stochastic mutation semigroup is collapsing within the specified tolerance.
| Algorithm 1 Probabilistic Pair-Graph Collapse Test |
| Require: X: finite set, ; : generators; : selection probabilities; : tolerance Ensure: TRUE if collapse probability , FALSE otherwise 1: Construct pair-graph G with vertices 2: Initialize for diagonal vertices, otherwise 3: Initialize for diagonal vertices, otherwise 4: Initialize queue Q with all diagonal vertices 5: while Q is not empty do 6: pop from Q 7: for each predecessor u of v do 8: Compute 9: 10: if increased and then 11: push u to Q 12: end if 13: end for 14: end while 15: return |
5.2. Probabilistic Invariant Partition Detection
Detecting invariant partitions in the stochastic setting requires considering probabilities rather than deterministic containment.
Algorithm (Probabilistic Invariant Partition Detection)
Input: finite set X, probabilistic generators , threshold . Output: a partition of X such that for each block B and each generator i,
for some block , or NO if none exists.
The algorithm uses a modified version of the deterministic backtracking search [5], with invariance checked probabilistically.
5.3. Complexity Analysis
Theorem 7
(Complexity of Probabilistic Detection). The probabilistic invariant partition detection problem is in NP and is NP-hard in the worst case.
Proof.
Membership in NP follows from the fact that a partition is a polynomial-size certificate, and checking the probabilistic invariance condition can be performed in polynomial time by computing the relevant probabilities. NP-hardness follows from the deterministic case, which is a special case of the probabilistic problem with [5,17]. □
6. Conceptual Biological Illustrations
6.1. Predicting the Emergence of Drug Resistance
The stochastic framework can be used to illustrate, conceptually, how the emergence of drug resistance in HIV might be modeled.
Example 2
(Conceptual Illustration: HIV Drug Resistance). Using the stochastic mutation semigroup model, the probability of resistance emergence under different drug regimens is computed. The results are shown in Table 4.
Table 4.
Predicted probability of drug-resistance emergence.
This provides a conceptual illustration of how the framework could, in principle, be used to explore drug resistance dynamics. It does not constitute a validated clinical tool.
The probability of resistance emergence for different drug regimens is shown in Table 5, where the number of ■ symbols represents the relative probability magnitude.
Table 5.
Probability of drug-resistance emergence over time for different drug regimens.
Example 3
(Nevirapine Resistance in HIV-1). Consider the mutation G→A at codon 103 (K103N), which confers high-level resistance to nevirapine. Using the stochastic framework, the probability of resistance emergence is approximated by
where is the probability of the G→A mutation at time t. For a typical patient with a mutation rate of per site per day, the probability of resistance after days is approximately 0.12, which is consistent with clinical observations.
This calculation assumes the following:
- 1.
- Independence of mutations at each time step;
- 2.
- A constant mutation rate of per day;
- 3.
- A single mutation event sufficient for resistance;
- 4.
- No back-mutation to the wild type.
This is a simplified illustration of the framework, not a full application of the stochastic mutation semigroup.
Using the full stochastic mutation semigroup framework with the five-haplotype state space and all mutation operators, the probability of resistance emergence is computed as
which accounts for all possible mutational pathways and their probabilities.
The dynamics of drug resistance emergence over time are illustrated in Figure 4, which shows the probability of resistance as a function of generations for each drug regimen.
Figure 4.
Probability of drug resistance emergence over time for three drug regimens: a single drug with a low barrier, a single drug with a high barrier, and combination therapy (3 drugs). The curves show the increasing probability of resistance as a function of the number of generations.
6.2. Cancer Clonal Evolution (Conceptual Illustration)
The following example illustrates how the framework could, in principle, be applied to model clonal evolution in cancer. This is a conceptual demonstration only and does not constitute a validated model of cancer progression.
Example 4
(Tumor Clonal Evolution). Let represent four tumor subclones. Probabilistic mutation operators correspond to proliferation and migration with rates inferred from sequencing data. The model predicts the probability of clonal sweep (fixation) and polyclonal persistence [36]. We emphasize that this is a simplified illustration; real cancer evolution involves complex selective pressures and spatial heterogeneity that are not captured here.
6.3. Conservation Genetics (Conceptual Illustration)
The following example illustrates how the framework could, in principle, inform conservation decisions by predicting genetic homogenization. This is a conceptual demonstration and does not constitute a validated conservation model.
Example 5
(Conservation Decision Support). For an endangered species with four isolated populations, the stochastic model predicts the probability of genetic homogenization under different migration scenarios. This informs decisions about protected zone management [37]. As with the cancer example, this is a simplified illustration; real conservation decisions require detailed demographic, environmental, and genetic data that are not included in this model.
8. Conclusions
The deterministic theory of mutation semigroups [5,6] has been extended to the stochastic setting, developing a comprehensive theoretical framework for understanding evolutionary dynamics under uncertainty. The stochastic mutation semigroup framework provides a probabilistic generalization of the transitivity threshold theorem (Theorem 4), with the corrected theorem now rigorously establishing collapse under a positivity condition on the probability of rank reduction. A lemma ensuring the existence of deterministic maps in the support of probabilistic operators addresses the logical gap identified in earlier formulations, while the spectral characterization (Theorem 5) correctly identifies collapse with the existence of a closed communicating class of rank-one states reachable from all states.
The theoretical framework is complemented by efficient algorithmic implementations, including complete pseudocode for the probabilistic pair-graph algorithm with convergence criteria and complexity analysis. The framework is illustrated using HIV-1 sequence data and mutation rates from Zanini et al. [7], demonstrating its potential applicability to biological systems. The data sources used for these illustrative examples are publicly available from the NCBI Sequence Read Archive (SRX25986227) [8], GenBank (AF113585.1) [9], and BEI Resources (HRP-11663) [10]. Conceptual illustrations in the contexts of drug resistance prediction, cancer clonal evolution [36], and conservation genetics [37] indicate possible future applications of the framework.
The extensions developed in this paper significantly broaden the theoretical applicability of the framework. The topological approach to infinite state spaces (Section 7.1) provides a rigorous foundation for modeling continuous genotype spaces such as HIV quasispecies [7,38], where the state space is astronomically large. The time-varying framework (Section 7.2) captures the dynamic nature of real mutation processes, accommodating changes in mutation rates due to environmental factors, immune pressure, and drug treatment [7,14].
The framework bridges the gap between abstract semigroup theory and evolutionary biology, providing a mathematically rigorous foundation for understanding stochastic mutation dynamics. Future work will focus on the rigorous development of correlated mutation models incorporating recombination and epistasis, the extension to multi-species systems with coupled dynamics [39,40], the large-scale implementation of the algorithms on genomic datasets, and the development of statistical inference methods for estimating mutation operators from empirical data. These future directions, if pursued, could potentially transform the theoretical framework presented here into a practical tool for empirical applications. At present, however, the framework remains primarily a mathematical contribution with conceptual biological illustrations.
Author Contributions
Conceptualization, M.I.S. and R.G.; Methodology: C.F.I., R.G. and J.S.G.; Software: M.I.S.; Validation, C.F.I. and R.G.; Formal analysis: R.G. and J.S.G.; Investigation: C.F.I. and J.S.G.; Resources: J.S.G.; Writing—original draft: M.I.S.; Supervision: R.G.; Funding acquisition: R.G. All authors have read and agreed to the published version of the manuscript.
Funding
This research was supported by Prince Sattam bin Abdulaziz University, Saudi Arabia, through project number PSAU/2025/01/38808.
Data Availability Statement
The data presented in this study are available in publicly accessible repositories. HIV-1 V3 loop sequence data are available from the NCBI Sequence Read Archive under run accession SRX25986227 (https://www.ncbi.nlm.nih.gov/sra/SRX25986227, accessed on 5 August 2026). Reference sequence AF113585.1 is available from GenBank (https://www.ncbi.nlm.nih.gov/nuccore/AF113585.1, accessed on 5 August 2026). HIV-1 subtype B Env clones are available from BEI Resources under catalog number HRP-11663 (https://www.beiresources.org/Catalog/Clones/HRP-11663.aspx, accessed on 5 August 2026). Empirical mutation rates used in this study are from Zanini et al. [7]. No new primary data were generated in this study. The code implementing the algorithms described in the paper is available from the corresponding author upon reasonable request.
Acknowledgments
The authors gratefully acknowledge the use of public HIV-1 sequence data from the NCBI Sequence Read Archive (SRX25986227), GenBank (AF113585.1), and BEI Resources (HRP-11663).
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Howie, J.M. Fundamentals of Semigroup Theory; Oxford University Press: Oxford, UK, 1995. [Google Scholar] [CrossRef] [Scilit]
- Sampson, M.I.; Lipcsey, Z.; Offiong, A.E.; Efiong, F.A.; Essien, M.A. Algorithm for Semigroup Bases I. Int. J. Math. Anal. Model. 2024, 7, 144–152. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Volkov, M.V. Slowly Synchronizing Automata with Idempotent Letters of Low Rank. J. Autom. Lang. Comb. 2019, 24, 375–386. [Google Scholar]
- Pin, J.-E. Varieties of Formal Languages; North Oxford Academic: London, UK, 1983; Available online: https://zbmath.org/0655.68095 (accessed on 1 May 2026).
- Sampson, M.I.; George, R.; Abubakar, R.B.; George, J.S. From Local Mutations to Global Fixation: A Semigroup Approach to Evolutionary Collapse. Math. Comput. Appl. 2026, 31, 138. [Google Scholar] [CrossRef] [Scilit]
- Sampson, M.I.; George, R.; Abubakar, R.B.; George, J.S. Low-Rank Behavior in Mutation Semigroups: Structural and Computational Methods. Contemp. Math. 2026, 7, 2652–2673. [Google Scholar] [CrossRef] [Scilit]
- Zanini, F.; Puller, V.; Brodin, J.; Albert, J.; Neher, R.A. In vivo mutation rates and the landscape of fitness costs of HIV-1. Virus Evol. 2017, 3, vex003. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- NCBI Sequence Read Archive. NGS of HIV-1 V3 Loop, Run SRX25986227. Available online: https://www.ncbi.nlm.nih.gov/sra/SRX25986227 (accessed on 1 May 2026).
- GenBank. HIV-1 Isolate CN37 Subtype B’ from China Envelope Glycoprotein (env) Gene, Partial cds. Available online: https://www.ncbi.nlm.nih.gov/nuccore/AF113585 (accessed on 1 May 2026).
- BEI Resources. Panel of SGA Human Immunodeficiency Virus Type 1 (HIV-1) Subtype B Env Clones. Catalog No. HRP-11663. Available online: https://www.beiresources.org (accessed on 1 May 2026).
- Baake, M.; Birkner, M. Stochastic processes in evolutionary biology: A review. Stoch. Process. Appl. 2020, 130, 6100–6120. [Google Scholar] [CrossRef] [Scilit]
- Landau, C.; Rosenfeld, J.; Shapira, B. Stochastic semigroups in population genetics: Recent advances. J. Math. Biol. 2024, 89, 45–78. [Google Scholar]
- Hill, S.; Ahn, J.; Park, S. Recent advances in stochastic modeling of mutation dynamics. Math. Biosci. 2025, 378, 109278. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ewens, W.J. Mathematical Population Genetics I: Theoretical Introduction, 2nd ed.; Interdisciplinary Applied Mathematics; Springer: New York, NY, USA, 2004; Volume 27. [Google Scholar]
- Bobrowski, A.; Kimmel, M. An Operator Semigroup in Mathematical Genetics; Springer: Berlin/Heidelberg, Germany, 2015. [Google Scholar] [CrossRef] [Scilit]
- Shabana, H.; Volkov, M.V. Using SAT solvers for synchronization issues in nondeterministic automata. arXiv 2018, arXiv:1801.05391. [Google Scholar] [CrossRef] [Scilit]
- Ruszil, J. Careful Synchronization of One-Cluster Automata. arXiv 2023, arXiv:2311.15020. [Google Scholar] [CrossRef] [Scilit]
- Birget, J.-C. Intersection and union of regular languages and state complexity. Inf. Process. Lett. 1992, 43, 185–190. [Google Scholar] [CrossRef] [Scilit]
- Volkov, M.V. Synchronization of Finite Automata. Russ. Math. Surv. 2022, 77, 819–891. [Google Scholar] [CrossRef] [Scilit]
- Plénet, T.; Bagnoli, F.; El Yacoubi, S.; Raïevsky, C.; Lefèvre, L. Synchronization of Elementary Cellular Automata. Nat. Comput. 2024, 23, 31–40. [Google Scholar] [CrossRef] [Scilit]
- Bagnoli, F.; El Yacoubi, S.; Rechtman, R. Synchronization and Control of Cellular Automata. In ACRI 2010; LNCS 6350; Springer: Berlin/Heidelberg, Germany, 2010; pp. 188–197. [Google Scholar] [CrossRef] [Scilit]
- Cheng, Y.; Liu, X.; Zhang, W. Probabilistic automata and their applications in biological systems. Theor. Comput. Sci. 2023, 945, 113678. [Google Scholar] [CrossRef] [Scilit]
- Neher, R.A.; Leitner, T. HIV evolution and the dynamics of drug resistance. Annu. Rev. Biophys. 2023, 52, 157–179. [Google Scholar]
- Yin, A.; Moes, D.J.A.R.; van Hasselt, J.G.C.; Swen, J.J.; Guchelaar, H.-J. A Review of Mathematical Models for Tumor Dynamics and Treatment Resistance Evolution of Solid Tumors. CPT Pharmacomet. Syst. Pharmacol. 2019, 8, 720–737. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Williams, M.; Thompson, J.; Anderson, R. Conservation genetics in the 21st century: Mathematical approaches to biodiversity management. Conserv. Biol. 2021, 35, 145–160. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sung, K.S.; Barton, J.P. Parallel HIV-1 evolutionary dynamics in humans and rhesus macaques who develop broadly neutralizing antibodies. bioRxiv 2024. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Heinrich-Mora, E.; Feldman, M.W. Evolution Under Stochastic Transmission: Mutation-Rate Modifiers. 2025. Available online: https://github.com/elisaheinrichmora/Stochastic_Transmission_Mutation (accessed on 1 May 2026).
- Ananichev, D.S.; Volkov, M.V. Synchronizing monotonic automata. Theor. Comput. Sci. 2004, 327, 225–239. [Google Scholar] [CrossRef] [Scilit]
- Černý, J. Poznámka k homogénnym eksperimentom s konečnými automatami. Mat.-Fyz. Cas. Slov. Akad. Vied 1964, 14, 208–216. [Google Scholar]
- Volkov, M.V. Synchronizing automata and the Černý conjecture. In Languages and Automata: Theory and Applications (LATA 2008); Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2008; Volume 5196, pp. 11–27. [Google Scholar] [CrossRef] [Scilit]
- Clifford, A.H.; Preston, G.B. The Algebraic Theory of Semigroups; Mathematical Surveys and Monographs; American Mathematical Society: Providence, RI, USA, 1961; Volume I. [Google Scholar] [CrossRef] [Scilit]
- Cameron, P.J. Permutation Groups; Cambridge University Press: Cambridge, UK, 1999. [Google Scholar] [CrossRef] [Scilit]
- Neumann, P.M. Synchronizing groups and the Černý conjecture. Bull. Lond. Math. Soc. 2008, 40, 569–577. [Google Scholar]
- Norris, J.R. Markov Chains; Cambridge University Press: Cambridge, UK, 1997. [Google Scholar] [CrossRef] [Scilit]
- Karlin, S.; Taylor, H.M. A First Course in Stochastic Processes, 2nd ed.; Academic Press: New York, NY, USA, 1975. [Google Scholar]
- Cosenza, M.R.; Rodriguez-Martin, B.; Korbel, J.O. Structural Variation in Cancer: Role, Prevalence, and Mechanisms. Annu. Rev. Genom. Hum. Genet. 2022, 23, 123–152. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kimura, M. The Neutral Theory of Molecular Evolution; Cambridge University Press: Cambridge, UK, 1983. [Google Scholar] [CrossRef] [Scilit]
- Eigen, M. Selforganization of matter and the evolution of biological macromolecules. Naturwissenschaften 1971, 58, 465–523. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Koonin, E.V. Viruses and mobile elements as drivers of evolutionary transitions. Philos. Trans. R. Soc. B 2016, 371, 20150442. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Novozhilov, A.S.; Karev, G.P.; Koonin, E.V. Biological applications of the theory of birth-and-death processes. Brief. Bioinform. 2006, 7, 70–85. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.



