Next Article in Journal
Downstream Purification Strategies for Virus-like Particles: A Systematic Review of Structure Preservation, Impurity Control, and Viral Safety
Next Article in Special Issue
Information-Entropy-Based Single Amino Acid Polymorphism Analysis Reveals Functional Variance of Enterovirus 2A Proteases
Previous Article in Journal
Rigidifying Flexible Regions of a Bacterial Laccase Enables High-Temperature Aflatoxin B1 Degradation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Complex Recombination Landscape and Lineage Turnover in Classical Human Astroviruses

by
Yulia Aleshina
1,2,*,
Vladimir Frantsuzov
3 and
Alexander Lukashev
1
1
Martsinovsky Institute of Medical Parasitology, Tropical and Vector Borne Diseases, Sechenov First Moscow State Medical University, 119435 Moscow, Russia
2
Faculty of Bioengineering and Bioinformatics, Lomonosov Moscow State University, 119234 Moscow, Russia
3
Department of Normal Physiology, Sechenov First Moscow State Medical University, 119435 Moscow, Russia
*
Author to whom correspondence should be addressed.
Microorganisms 2026, 14(4), 857; https://doi.org/10.3390/microorganisms14040857
Submission received: 13 March 2026 / Revised: 3 April 2026 / Accepted: 8 April 2026 / Published: 10 April 2026
(This article belongs to the Special Issue Molecular Epidemiology and Surveillance of Major Enteric Viruses)

Abstract

Human astroviruses are small, non-enveloped RNA viruses belonging to the family Astroviridae. Among the four species known to infect humans, the species Mamastrovirus hominis (the classical human astroviruses, formerly MAstV1) is associated with gastrointestinal illness worldwide, while three more recently identified species have been linked to lethal central nervous system infections. High substitution rates and recombination drive their rapid evolution, yet recombination patterns in classical human astroviruses remain poorly characterized. This study systematically analyzes patterns and temporal dynamics of natural recombination in classical human astroviruses. Publicly available genomes of classical human astroviruses were analyzed to identify recombination hotspots. Recombinant forms were defined as stable phylogenetic lineages unaffected by recombination, and their half-lives were estimated based on time-scaled phylogenies (BEAST2v2.7.7). Recombination in classical human astroviruses occurred most frequently at the ORF1b/ORF2 junction, but also within ORF1a, at the ORF1a/ORF1b junction, and within ORF2. Only the 3′-part of ORF1a and a fragment of ORF1b exhibited robust temporal signal, yielding substitution rates of 2.35 × 10−3 and 2.14 × 10−3 s/s/y, respectively. The half-lives of recombinant forms varied considerably by genomic region: longest for exchanges between the parts of ORF1a (21 years), intermediate for ORF1a/ORF1b recombinants (7–9 years), and shortest for ORF1ab/ORF2 recombinants (2.5–3.6 years). The estimated half-lives for recombinants align with those reported for human enteroviruses and noroviruses. These findings highlight the dynamics of the generation of astrovirus diversity and may inform advanced surveillance of emerging strains.

1. Introduction

Human astroviruses are small non-enveloped RNA viruses that belong to the family Astroviridae. Four species of human astroviruses are recognized to date. Viruses belonging to the species Mamastrovirus hominis (MAstV1), also referred to as classic astroviruses, were discovered in 1975 and are associated with gastrointestinal illness worldwide [1,2]. Based on antigenic properties, classic human astroviruses are classified into eight serotypes, HAstV1-HAstV8 [1]. The rest of the human-infecting species—Mamastrovirus melbournense (MAstV6), Mamastrovirus homustovis (MAstV8), Mamastrovirus virginiaense (MAstV9)—were identified in the late 2000s, along with human astroviruses VA3-VA6 and MLB2-MLB3, which remain unclassified to species level. While their pathogenicity has not yet been explicitly established, these species are known to be associated with lethal central neural system infections [3].
Astroviruses have a single stranded positive-sense RNA genome ranging from 6.7 to 7.0 kb in length. The genome is covalently linked at the 5′ end to viral genome-linked protein (VPg), which is essential for virus infectivity [4], and polyadenylated at the 3′ untranslated region. Astrovirus genome contains three main open reading frames (ORFs): ORF1a encodes a nonstructural polyprotein (nsP1a) that includes VPg and serine protease, while ORF1b and ORF2 encode RNA-dependent RNA polymerase (RdRp) and capsid proteins, respectively. ORF1b is expressed through programmed ribosomal frameshifting, which occurs at a conserved AAAAAAC sequence within the ORF1a/ORF1b overlap region [5]. ORF2 partially overlaps with ORF1b and is translated from a subgenomic RNA. Most human astroviruses have an additional ORF named ORFX that overlaps with ORF2 in the +1 reading frame and encodes viroporin. Notable exceptions include M. virginiaense (VA1), M. homustovis (VA2), and unclassified VA isolates (VA3–VA6) [6].
Similar to other positive-sense RNA viruses, classical human astroviruses evolve rapidly due to mutations and recombination. The published evolutionary rate estimates vary from 4.5 × 10−4 substitutions/site/year in the capsid-encoding gene [7] to 3.7 × 10−3 in the other genome fragments [8]. Homologous recombination occurs when divergent viruses co-infect a single cell, resulting in large-scale genomic changes. In mammalian astroviruses, recombination is frequent [9,10,11,12,13,14,15,16,17,18,19,20,21,22], but occurs strictly within a virus species. Therefore, the capacity for recombination has been proposed as an additional criterion for species demarcation, alongside with pairwise genetic distances in the ORF1b and ORF2 regions and host species [23]. Recombination plays a key role in the generation of novel astrovirus strains, which are reported regularly and may arise from exchanges between viruses of the same or different genotypes/serotypes (Table 1). Although recombination occurs most commonly at the ORF1b/ORF2 junction, events at the ORF1a/ORF1b junction and within the open reading frames have also been documented, including triple recombinants [15,24,25]. It is noteworthy that recombination within ORF2 was reported close to the ends of regions encoding individual proteins and spared their cores.
Recombination has been reported between human astroviruses and astroviruses from other hosts—including feline [20,26], porcine [27], and California sea lion [28]. However, it may be hard to distinguish recombination events that occurred in ancestral viruses that later became established in distinct host species from a true recombination between viruses infecting distinct species. In any case, there is no evidence that such recombination is routine and ubiquitous as observed among classical human astroviruses.
This study aims to systematically investigate natural recombination in classical human astroviruses and determine the frequency at which novel recombinant variants emerge and become established in circulation.
Table 1. Recombinant classical human astroviruses reported previously. Astrovirus serotypes and genotypes are indicated as in the original publication.
Table 1. Recombinant classical human astroviruses reported previously. Astrovirus serotypes and genotypes are indicated as in the original publication.
Location of Recombination BreakpointRecombinant VirusPotential Parents (the Genome Regions Derived from Parents Are Indicated in Brackets)Reference
Within ORF1aHAstV1HAstV7 (5′-part of ORF1a),
HAstV3 (3′-part of ORF1a)
[14]
* HAstV4HAstV3 (ORF1a), HAstV1 (ORF1b)[15]
ORF1a/ORF1b junction* HAstV1HAstV5 (ORF1a), HAstV8 (ORF1b)[24]
* HAstV2cHAstV6 (ORF1a), HAstV3 (ORF1b)[15]
* HAstV2bHAstV5 (ORF1a), HAstV4 (ORF1b)
ORF1b/ORF2 junctionHAstV1HAstV8 (ORF1b), HAstV1 (ORF2)[12,29]
HAstV1HAstV8 (ORF1a), HAstV1 (ORF2)[30]
* HAstV1HAstV8 (ORF1b), HAstV1 (ORF2)[24]
HAstV2aAmbiguous (ORF1ab), HAstV2 (ORF2)[15]
HAstV2dHAstV1b (ORF1b), HAstV2d (ORF2)
HAstV2HAstV3 (ORF1b), HAstV2 (ORF2)[14,15,29]
HAstV2HAstV6 (ORF1b), HAstV2 (ORF2)[12]
* HAstV2bHAstV5 (ORF1ab), HAstV4 (ORF2)[15]
* HAstV2cHAstV3a (ORF1b),
HAstV2c (ORF2)
[9]
HAstV2HAstV8 (ORF1a), HAstV2 (ORF2)[30]
HAstV3cHAstV1 (ORF1b), HAstV3 (ORF2)[10]
HAstV3HAstV2 (ORF1a), HAstV3 (ORF2)[30]
HAstV3HAstV8 (ORF1b), HAstV3 (ORF2)[29]
* HAstV4HAstV1 (ORF1b), HAstV4 (ORF2)[15]
HAstV5HAstV3 (ORF1b), HAstV5 (ORF2)[13]
HAstV8HAstV2 (ORF1b), HAstV8 (ORF2)[12]
HAstV8HAstV2 (ORF1ab), HAstV8 (ORF2)[11]
Within ORF1bHAstV4HAstV1 (ORF1a, 5′-half of ORF1b),
HAstV4 (3′-half of ORF1b, ORF2)
[15]
* HAstV8HAstV8, HAstV1, HAstV2[25]
Within ORF2HAstV2c-2bHAstV2c (5′-part), HAstV2b (3′-part)[9]
HAstV1a-1dHAstV1a (5′part), HAstV1d (3′-part)[31]
HAstV1bHeterologous non-HAstV RNA (5′-end of VP34), HAstV1b (the rest of ORF2)
HAstV3aHAstV3 (5′-part), HAstV3a (3′-part)
HAstV5aHAstV5a/GO-12 (5′-part),
HAstV5a/Oxford-S5 (3′-part)
* triple recombinants.

2. Materials and Methods

2.1. Data Collection

Complete genome sequences of the species Mamastrovirus hominis (formerly Mamastrovirus 1) (N = 125) were collected from GenBank database as of March 2024 using the query: “Mamastrovirus 1”[porgn] AND (“6000”[SLEN]: “8000”[SLEN])”. Viral metadata, including isolate, strain, isolation source, country of isolation, and collection date, were extracted from the corresponding qualifiers in GenBank entries using a custom Python script and manually curated. Missing sampling dates could be retrieved from associated publications for 7 out of 17 incomplete records.
Coding sequences (CDS)—specifically ORF1a, ORF1b, and ORF2—were extracted from complete genomes according to the coordinates indicated in GenBank records and manually verified. The CDS were translated to amino acid sequences, which were then aligned separately using the ClustalW algorithm in MEGA v11 [32]. The resulting alignments were back-translated to produce codon-based alignments. Records lacking CDS coordinates were discarded. Where possible, missing coordinates of ORFs were inferred using similar sequences from the same study; this curation resulted in the exclusion of only 7 sequences. Sequences isolated from sewage exhibited many insertions showing no homology to other sequences in the datasets and were excluded from the analysis. The resulting alignment of complete genomes comprised 104 sequences.
Genomes containing > 5 ambiguous nucleotides (or >3 consecutive ambiguities), or >50 gaps in a row, were discarded, and ambiguous sites were in a few cases resolved by selecting the most probable base from the best-matching BLAST v2.16.0 hit [33] within a 100-nt window (E-value threshold ≤ 1 × 10−10), using a custom script (https://github.com/v-julia/resolve_ambiguous; accessed on 9 April 2026). After quality control, the dataset comprised 77 sequences (Table S1). For recombination analyses, alignments of ORF1a, ORF1b, and ORF2 were concatenated.
To verify the serotype assignment of individual sequences, a preliminary Maximum Likelihood (ML) phylogenetic tree was constructed for the capsid-encoding region (ORF2) using the IQ-TREE web server [34]. For sequences with serotype not indicated in Genbank entry, the serotype was inferred based on phylogenetic grouping.

2.2. Recombination Analysis

A full exploratory recombination scan was performed to identify recombination breakpoints using RDP4 v101 software [35] applying nine standard recombination detection methods: RDP [36], GENECONV [37], MaxChi [38], Chimaera [39], BootScan [40], SiScan [41], 3seq [42], PhylPro [43], and LARD [42]. Only events supported by at least 5 methods were considered. Putative events were further validated by manually examining the phylogenetic placement of recombinant sequences and their parents.
The overall recombination patterns were visualized using breakpoint distribution plot (sliding window 200 nt, step size 100 nt) and recombination region count matrices. Recombination region count matrices indicate the number of times the detected recombination events separated that specific pair of genome positions.
Breakpoint distribution was inspected to identify regions with minimal evidence of recombination. Subsequently, these recombination-free regions were excised from the alignments for downstream phylogenetic analysis.

2.3. Phylogenetic Analysis

ML phylogenetic trees for recombination-free regions were built on the IQ-TREE web server (auto-detection of substitution model, ultrafast bootstrap analysis with 1000 replicates) [34,44,45].
The preliminary assessment of temporal signal for each recombination-free region was performed using TempEst [46]. The best-fitting tree root was selected using “heuristic residual mean squared” function. Root-to-tip divergence and sampling dates were exported from TempEst; linear regression was then performed in Python v3.10 (scikit-learn, SciPy v1.15.3 packages). A recombination-free region was considered to have a temporal signal if (i) the correlation between root-to-tip divergence and sampling date was positive and (ii) the p-value of the regression slope was <0.05. The correspondence between root-to-tip divergence and sampling dates was visualized with the Matplotlib v3.10.3 Python package, using the exported TempEst data.
For genome regions that exhibited a temporal signal in TempEst, time-calibrated phylogenies were inferred using BEAST v2.7.7 [47]. First, the best-fitting nucleotide substitution models and partitioning schemes were selected using PartitionFinder 2.1.1 software [48] under the AIC criterion and with substitution models restricted to those implemented in BEAST v.2.7.7. Marginal likelihoods were then estimated for every combination of molecular clock models (strict, relaxed lognormal, relaxed exponential) and tree prior models (coalescent constant population, exponential population, and bayesian skyline) using Nested Sampling plugin in BEAST2 [49]. Along with marginal likelihood, Nested sampling provides estimates of standard deviation of marginal likelihood. If the difference between marginal likelihoods of two models was less than 2×sqrt(SD1×SD1 + SD2×SD2), the analysis was repeated with more particles to ensure that the difference was not due to randomization. The combination with the highest marginal likelihood was selected as the best-fitting model; Bayes factors (BF) were calculated for pairwise comparisons (Table S2).
The final assessment of the temporal signal was performed using Bayesian Estimation of Temporal Signal (BETS) [50,51]. Under the best-fitting model identified above, marginal likelihoods were estimated for two models—one incorporating the actual sampling times and another suggesting all samples to be contemporaneous (sampling dates set to zero). If log Bayes factor was >5 in favor of the model with actual sampling dates, the dataset of genome regions was interpreted as having a robust temporal signal. Only regions satisfying this criterion were used for full MCMC analysis.
For alignments of genome regions with temporal signal confirmed by BETS, five independent Markov chain Monte Carlo (MCMC) runs were performed to assess the convergence. Chain lengths and sampling frequencies were adjusted per dataset (see Table S2). MCMC convergence was verified in Tracer v1.7.2 [52]; all parameters achieved effective sample size > 200. The resulting posterior tree distributions from independent runs were combined using LogCombiner. Maximum Clade Credibility (MCC) trees were then summarized from the combined files using TreeAnnotator v.2.7.7. after discarding 10% of each chain as burn-in.
The phylogenetic trees were visualized in R environment using ggtree package [53].

2.4. Temporal Dynamics of Recombination

To characterize the turnover of recombinant lineages, we estimated the half-life of recombinant forms—the median period during which a lineage arising from a distinct recombination event persists in the host population before it is no longer detectable. A recombinant form is defined here as a viral lineage descending from a single ancestral recombination event between two distinguishable parental genomes, with no subsequent detectable recombination.
We employed a phylogenetic tree comparison approach previously applied to enteroviruses, noroviruses [54,55], and to foot-and-mouth disease virus [56]. This method identifies recombinant forms by comparing time-calibrated phylogenies inferred from genomic regions that are frequently exchanged through recombination. The lifetime of each recombinant form is measured as the interval between the estimated time of its most recent common ancestor (tMRCA), marking its emergence, and the isolation date of its most recently sampled descendant. The recombinant half-life for a group of viruses is then taken as the median of lifetimes of individual recombinant forms.
In earlier studies [54,55], tree comparisons were performed manually. To overcome this limitation, we previously developed RF-HL, a Python tool that automates the identification of coinciding clades and the calculation of recombinant half-lives [56]. RF-HL requires two input trees: a reference time-scaled phylogeny (Nexus format, e.g., from BEAST) that is used for dating clades and a phylogeny inferred by any other method (Newick or nexus format). It implements two complementary detection strategies—common subtree search and common bipartition search—and outputs the matching subtrees together with the estimated lifetimes of all detected recombinant forms. A detailed description of the algorithm and its validation is provided in [56]. The tool is freely available at https://github.com/v-julia/RF_HL (Accessed on 9 April 2026).
Half-life time calculations were performed under two scenarios. When both genomic regions exhibited robust temporal signal (3′-part of ORF1a and ORF1b), two Bayesian time-scaled phylogenies were compared. When only one region possessed a temporal signal, its Bayesian phylogeny was compared to a maximum likelihood phylogeny of the partner region.

3. Results

3.1. Identification of Recombination Patterns in Classical Human Astroviruses

The refined dataset of complete genomes of classical human astroviruses comprised 77 sequences and included all serotypes, except for serotype 7 (Table 2). Serotype 1 was the most prevalent (N = 32), followed by serotype 4 (N = 19), while the other serotypes were represented by less than 10 sequences.
Recombinant human astroviruses have been detected in field studies (Table 1). Most recombinants exhibited breakpoints at the ORF1b/ORF2 junction, although recombinants with breakpoints within ORFs have also been reported. For many of the detected recombinants, complete genome sequences were not available. To estimate recombination prevalence across the genome and identify genome regions that were more frequently exchanged and contributed to the generation of recombinant forms, we conducted an exploratory analysis using the RDP4 program. A permutation test was performed to compare the observed incidence of breakpoints with a random distribution, corrected for sequence similarity.
A total of 22 unique recombination events, supported by more than five methods and confirmed via manual inspection, were detected in the dataset (Figure 1). The permutation test identified a single recombination hotspot located within approximately 500 nucleotides around the ORF1b/ORF2 junction (8 events) (Figure 1a). Although breakpoint densities in other regions did not reach statistical significance as defined by the permutation test (i.e., they fell below the 99% local confidence interval threshold), the distribution of recombination breakpoints was not uniform. Clusters of breakpoints with densities approaching the upper bound of the confidence interval were observed in three regions: within ORF1a, at the beginning of the region encoding the serine protease p27 (positions 1301–1327 of the reference sequence OR371570/HAstV4); near the ORF1a/ORF1b junction (within the p20 protein); and within ORF2, where breakpoints clustered in the middle of the gene (approximately between the VP34 and VP25/VP27 encoding regions) and in the 3′ end of ORF2 (corresponding to the start of the acidic C-terminal domain).
Recombination region count matrices, which indicate the number of times unique recombination events separate genome regions, confirmed the division of ORF1ab into three blocks with no recombination within them: the 5′-half of ORF1a, the 3′-half of ORF1a, and ORF1b (Figure 1b). Within ORF2, two blocks with no evidence of internal recombination were identified, corresponding to the part of VP34 and the part of VP25/27. The matrices further revealed that recombination breakpoints within ORF1a and at the ORF1a/ORF1b junction contribute to the exchange of ORF1a relative to the ORF2 region. As a result, ORF1a was found to be recombinant relative to ORF2 more frequently than ORF1ab as a whole (Figure 1b).

3.2. Evaluating Temporal Signal Across the Astrovirus Genome

To further identify the genomic regions suitable for molecular dating, we first assessed temporal signal in all recombination-free regions by regressing root-to-tip distances from maximum likelihood (ML) phylogenetic trees against collection dates (Figure 2). Only two regions in our dataset of complete human astrovirus genomes exhibited clock-like evolution: the 3′ end of ORF1a (regression slope p = 4.61 × 10−5, R2 = 0.20) and ORF1b (regression slope p = 1.43 × 10−5, R2 = 0.48). Bayesian estimation of temporal signal subsequently supported a heterochronous model (which incorporates actual sampling dates), confirming the presence of temporal signal in these datasets (Table 3).
Substitution rates inferred by BEAST v2.7.7 were 2.35 × 10−3 s/s/y for 3′-part of ORF1a and 2.14 × 10−3 s/s/y for ORF1b, values that are consistent with those typical of positive-sense RNA viruses. The time to the most recent common ancestor (tMRCA) was estimated at 251 years before present (ybp) for 3′-part of ORF1a and 208 ybp for ORF1b, with overlapping 95% highest posterior density (HPD) intervals for both rates and tree heights (Table 4).

3.3. Estimating the Persistence of Recombinant Lineages

The inferred time-scaled phylogenies for 3′-part of ORF1a and ORF1b enabled dating the recombinant forms. Following the phylogenetic approach established in our previous work on foot-and-mouth disease virus [56], we define recombinant forms as virus lineages that originate from a recombination event and subsequently persist without undergoing additional detectable recombination. Such lineages can be identified through comparison of tree topologies inferred from different genomic regions, appearing as coinciding clades—well-supported groups of sequences that share identical topology across trees. To systematically identify these clades, we employed the RF-HL program, which adopts a bipartition-based approach for tree comparison. This method treats each tree as a set of bipartitions (branches) that split the leaf set into two groups, then searches for coinciding bipartitions that meet predefined support thresholds. Nested bipartitions corresponding to internal branches of the resulting coinciding clades are omitted to avoid redundancy. The lifetime of a recombinant form is calculated as the time between the time of its most recent common ancestor (tMRCA) and the collection date of its most recent isolate; the recombinant form half-life is then defined as the median lifetime across all identified recombinant forms. We searched specifically for bipartitions that were reliably supported, requiring at least 95% ultrafast bootstrap support in ML trees and posterior probability > 0.9 in Bayesian time-scaled phylogenies (Figure 3 and Figures S1–S3).
To date the recombinant form lifetimes, we compared two time-scaled phylogenies—for 3′-part of ORF1a and for ORF1b—with ML phylogenies inferred from other genome regions exhibiting minimal or no recombination: 5′-part of ORF1a and two parts of ORF2. The phylogenetic tree for the 5′-part of ORF2 showed limited resolution, with only 23 out of 75 branches achieving ultrafast bootstrap support above 95%; this region was therefore excluded from further analysis. As time-scaled phylogenies were available for both 3′-part of ORF1a and ORF1b, we dated each set of coinciding clades twice—first using 3′-part of ORF1a tree as the temporal reference, then using the ORF1b tree—allowing to assess the consistency of lifetime estimates across the reference trees.
Applying this approach, we found that lifetimes of coinciding clades ranged from 0.24 to 66 years, depending on the regions compared (Figure 4). The resulting half-lives of recombinant forms are summarized in Figure 5. Comparison between the 5′ and 3′-parts of ORF1a revealed the longest persistence of recombinant forms in circulation, with lifetimes ranging from 9 to 27 years (Figure 4a-i) and a half-life of 21.4 years (Figure 5). This extended persistence aligns with the rare recombination incidence observed between these regions (Figure 1). When comparing the 5′-part of ORF1a with ORF1b, the median lifetime of coinciding clades was nearly threefold shorter (7.8 years). Recombinant forms observed when comparing 3′-part of ORF1a and ORF1b exhibited half-lives of 7.2 years when dated using 3′-part of ORF1a as the reference and 9.2 years when using ORF1b, reflecting modest discrepancies in clade dating between the two reference trees (Figure 5).
The highest degree of phylogenetic incongruence was observed between nonstructural regions (3′-part of ORF1a and ORF1b) and the structural ORF2 (Figure 6 and Figure S3). Sequences from different HAstV serotypes were extensively intermixed in trees built using nonstructural regions, indicative of frequent recombination between viruses from different serotypes (Figure 6). Subsequently, the highest number of recombinant forms could be identified between non-structural and structural genome regions (Figure 4). Comparisons of these genome regions also yielded the shortest half-lives of recombinant forms: 2.5 years for 3′-ORF1a versus 3′-part of ORF2 and 3.6 years for ORF1b versus 3′-part of ORF2 (Figure 5). These values are consistent with the recombination hotspot at the ORF1b/ORF2 junction and underscore the rapid turnover of lineages generated by recombination in this genomic region.

4. Discussion

Recombination is a well-established driver of genetic diversity in astroviruses, yet its patterns remain less explored than in other positive-sense RNA virus families. While multiple natural recombinants of classical human astroviruses (Mamastrovirus hominis) have been reported in field studies (Table 1), expanding availability of complete genome sequences has enabled a systematic evaluation of recombination patterns. Our analysis confirmed the single significant recombination hotspot located at the ORF1b/ORF2 junction, which was also observed at the genus Mamastrovirus level [23]. This finding aligns with a general pattern among positive-sense RNA viruses: frequent recombination between genomic regions encoding structural and non-structural proteins [55,57,58,59,60,61,62,63].
In addition to this canonical hotspot, several additional clusters of breakpoints were observed (i) within ORF1a, specifically in the 3′-end of region encoding the serine protease p27; (ii) near the ORF1a/ORF1b junction, within the p20 protein; and (iii) within ORF2 itself, with breakpoints clustering in the middle of the ORF2 (between the VP34 and VP25/VP27 encoding regions) and at the 3′ end corresponding to the acidic C-terminal domain. These observations are consistent with previously reported events (Table 1) and mirror patterns observed across the Mamastrovirus genus, where recombination was infrequent in ORF1b and in the ORF2 regions encoding VP34 [23].
Based on the recombination patterns identified in classical human astroviruses, five distinct genomic regions that were spared by recombination could be identified: the 5′-part of ORF1a, the 3′-part of ORF1a, the majority of ORF1b, and two segments of ORF2 corresponding approximately to the VP34 and VP25/VP27 coding regions (termed here the 5′-part of ORF2 and the 3-part of ORF2). The generation of recombinant variants in human astroviruses thus appears to occur through frequent recombination at the major ORF1b/ORF2 region hotspot, and lower-frequency recombination events at other genome locations, particularly at domain boundaries within ORF1ab and ORF2.
Consistent with previous reports (Table 1), we found that recombination breakpoints in ORF2 occur at domain boundaries, preserving the integrity of the functional cores of individual capsid proteins. This pattern resembles the modular evolution observed in the coronavirus spike gene, where recombination occurs between protein domains but not within them, enabling the exchange of functional modules while maintaining the overall protein architecture [59,64].
To investigate the temporal dynamics of these recombination events, we first established a reliable timescale for astrovirus evolution. This required identifying genomic regions that were spared by recombination and retained a sufficient temporal signal for molecular clock analysis. Only two genome regions—3′-part of ORF1a (~1500 nt in length) and a fragment of ORF1b (~1100 nt in length)—exhibited robust temporal signal and yielded comparable substitution rates of 2.35 × 10−3 s/s/y and 2.14 × 10−3 s/s/y, respectively. These values are lower than 3.7 × 10−3 s/s/y for 600 nt fragment of ORF1a and 4.1 × 10−3 s/s/y for 430 nt fragment of ORF2 reported by Babkin et al. [8]. However, the analysis by Babkin et al. was based only on 16 complete genomes and did not include serotype 2 and most serotype 4 sequences, many of which have since become available and were incorporated into our study. More recently, Zhou et al. estimated a substitution rate of 4.51 × 10−4 s/s/y for the ORF2 [7]. In our dataset, which comprised complete genomes of classical human astroviruses, the complete ORF2 was not usable for phylogenetic analysis due to multiple recombination events within, and recombination-free fragments of ORF2 did not exhibit a sufficient temporal signal to permit reliable rate estimation. Therefore, it is not clear if there is a true difference in substitution rates between ORF1 and ORF2.
To determine the lifetimes of recombinant forms in human astroviruses, we employed a phylogenetic approach. A key advantage of this method is that it does not rely on pre-defined genotype or lineage classifications. While such classifications are well established for some virus groups—such as the dual nomenclature for polymerase and capsid proteins in noroviruses [65]—no comparable nomenclature exists for human astroviruses. Moreover, given the complex evolutionary dynamics of astroviruses, any attempt to establish a similar classification would likely be arbitrary.
The phylogenetic approach bypasses some limitations but is a subject to others. It requires reliable phylogenies with robust branch support. In practice, inferred trees often lack high bootstrap support across all branches, especially when analyzing datasets with closely related sequences. For clades with highly similar sequences, low bootstrap support precludes precise topology comparisons and thus detection of recombination. To address this, we employed a softer definition of a recombinant form that does not require exact topological matches within clades but instead identifies highly supported branches that are congruent between trees. However, even with this relaxed definition, comparisons become uninformative when most branches in the tree lack high bootstrap support.
The problem of low bootstrap support was aggravated by the need to partition the genome. Dividing the genome into regions with no internal recombination breakpoints is necessary because recombination can disrupt tree topology. However, this partitioning may yield short genomic fragments with limited phylogenetic resolution. Among the five genome blocks involved in recombination that showed no internal recombination, the phylogenetic tree for the 5′-part of ORF2 exhibited especially low branch support (only 30.6% of branches with >95% bootstrap), making phylogenetic comparisons uninformative.
The estimation of recombinant form lifetimes may vary depending on the sequence dataset and the reference time-scaled phylogeny used. The median lifetimes of clades that were congruent between the 3′-part of ORF1a and ORF1b phylogenies differed depending on which tree served as the dating reference. Specifically, half-life estimates for recombinant forms varied by approximately two years–7.2 years when calibrated against the ORF1a phylogeny versus 9.2 years when using the ORF1b phylogeny. Recently, we demonstrated for FMDV that, across different sequence subsamples, the standard deviation of recombinant form half-lives ranged from 0.79 to 1.58 years [56]. Thus, the observed two-year difference between estimates falls within the range of expected variability given the inherent uncertainties of molecular clock dating.
Even though molecular dating of recombination events was rather consistent between the 3′-part of ORF1a and ORF1b, a significant limitation comes from the sampling bias. Serotype coverage was very uneven among the published sequences (Table 2). Moreover, most genomes were sampled over the last five years, and mainly in just a few countries. One can speculate how a better sampling can either increase (by being able to trace circulation of lineages for a longer time) or decrease (by detecting more recombination events) the observed recombinant form half-lives. However, it is unlikely that better sampling would yield radically different recombination rate estimates. In any case, the timing reported here should be taken with caution and interpreted more as an order of magnitude than as precise values.
The patterns and dynamics of recombination in classical human astroviruses correspond very well to those in other enteric RNA viruses. For enterovirus echovirus 30, turnover of recombinant forms made up of structural and non-structural genome regions occurs every 3–5 years, for echovirus 9 approximately every year, and for echovirus 11 every 10 years [66,67]. In enterovirus A71, the half-life of recombinant forms varies by genotype (ranging from 6 to 10 years), with some prevalent recombinant forms persisting without detectable recombination for several decades [68]. Among noroviruses of the GI and GII genogroups—the major causes of adult diarrhea—recombinant forms’ half-lives were longer, estimated at 8 and 10 years respectively, yet remain within the same order of magnitude as those of enteroviruses [55].
In human astroviruses, the half-lives of recombinant forms involving structural and nonstructural regions fell within this same range: 2.5 years for recombinants between the 3′-part of ORF1a and the 3′-part of ORF2, and 3.6 years for those between ORF1b and the 3′-part of ORF2. However, the more complex recombination landscape of astroviruses also includes exchanges confined to the nonstructural region, which exhibited markedly different temporal dynamics. In contrast to the short-lived recombinant forms arising from exchanges between nonstructural and structural regions (2.5–3.6 years), those defined by recombination events within the nonstructural part of the genome exhibited markedly longer persistence. Recombinants involving exchanges between the 5′- and 3′-parts of ORF1a had a half-life of 21.4 years, while those between ORF1a and ORF1b showed intermediate half-lives of 7.2–9.2 years. Therefore, recombination in enteric RNA viruses follows not just similar patterns [69], but also comparable dynamics on a global scale.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/microorganisms14040857/s1, Table S1: Human astrovirus GenBank entries used in this study and corresponding metadata. Table S2: Model selection for BEAST analysis. Figure S1: Inference of half-life time of recombinant forms that were made up by recombination between 3′-part of ORF1a and 5′-part of ORF1a genome regions of classical human astroviruses. (A) Bayesian time-scaled tree for 3′-part of ORF1a genome region. (B) ML phylogeny inferred using the 5′-part of ORF1a region. Black circles indicate nodes with high support (posterior probability > 0.9 in (A), ultrafast bootstrap > 95% in (B)). Clades (bipartitions) that coincide between trees are colored red. Taxa labels are colored according to the serotype. Figure S2: Inference of half-life time of recombinant forms that were made up by recombination between ORF1b and 5′-part of ORF1a genome regions of classical human astroviruses. (A) Bayesian time-scaled tree for ORF1b genome region. (B) ML phylogeny inferred using the 5′-part of ORF1a region. Black circles indicate nodes with high support (posterior probability > 0.9 in (A), ultrafast bootstrap > 95% in (B)). Clades (bipartitions) that coincide between trees are colored red. Taxa labels are colored according to the serotype. Figure S3: Inference of half-life time of recombinant forms that were made up by recombination between 3′-part of ORF1a and 3′-part of ORF2 genome regions of classical human astroviruses. (A) Bayesian time-scaled tree for 3′-part of ORF1a genome region. (B) ML phylogeny inferred using the 3′-part of ORF2 region. Black circles indicate nodes with high support (posterior probability > 0.9 in (A), ultrafast bootstrap > 95% in (B)). Clades (bipartitions) that coincide between trees are colored red. Taxa labels are colored according to the serotype.

Author Contributions

Conceptualization, Y.A.; methodology, Y.A.; software, Y.A. and V.F.; validation, Y.A. and A.L.; formal analysis, V.F.; investigation, V.F. and Y.A.; resources, Y.A.; data curation, V.F. and Y.A.; writing—original draft preparation, Y.A. and V.F.; writing—review and editing, Y.A. and A.L.; visualization, V.F. and Y.A.; supervision, Y.A. and A.L.; project administration, Y.A.; funding acquisition, Y.A. All authors have read and agreed to the published version of the manuscript.

Funding

The study was supported by Russian Science Foundation, grant number 24-74-00097.

Data Availability Statement

The data presented in this study (sequence names and metadata, alignments, phylogenetic trees) are openly available at https://github.com/v-julia/HAstV_recombination_dynamics (accessed on 15 March 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
CDSCoding sequence
HAstVHuman astrovirus
HPDHighest posterior density
MLMaximum likelihood
ORFOpen reading frame
tMRCATime to the most recent common ancestor

References

  1. Bosch, A.; Pintó, R.M.; Guix, S. Human Astroviruses. Clin. Microbiol. Rev. 2014, 27, 1048–1074. [Google Scholar] [CrossRef] [Scilit]
  2. Farahmand, M.; Khales, P.; Salavatiha, Z.; Sabaei, M.; Hamidzade, M.; Aminpanah, D.; Tavakoli, A. Worldwide Prevalence and Genotype Distribution of Human Astrovirus in Gastroenteritis Patients: A Systematic Review and Meta-Analysis. Microb. Pathog. 2023, 181, 106209. [Google Scholar] [CrossRef] [Scilit]
  3. Vu, D.L.; Cordey, S.; Brito, F.; Kaiser, L. Novel Human Astroviruses: Novel Human Diseases? J. Clin. Virol. 2016, 82, 56–63. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Fuentes, C.; Bosch, A.; Pintó, R.M.; Guix, S. Identification of Human Astrovirus Genome-Linked Protein (VPg) Essential for Virus Infectivity. J. Virol. 2012, 86, 10070–10078. [Google Scholar] [CrossRef] [Scilit]
  5. Marczinke, B.; Bloys, A.J.; Brown, T.D.; Willcocks, M.M.; Carter, M.J.; Brierley, I. The Human Astrovirus RNA-Dependent RNA Polymerase Coding Region Is Expressed by Ribosomal Frameshifting. J. Virol. 1994, 68, 5588–5595. [Google Scholar] [CrossRef] [Scilit]
  6. Lulla, V.; Firth, A.E. A Hidden Gene in Astroviruses Encodes a Viroporin. Nat. Commun. 2020, 11, 4070. [Google Scholar] [CrossRef] [Scilit]
  7. Zhou, N.; Zhou, L.; Wang, B. Molecular Evolution of Classic Human Astrovirus, as Revealed by the Analysis of the Capsid Protein Gene. Viruses 2019, 11, 707. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Babkin, I.V.; Tikunov, A.Y.; Zhirakovskaia, E.V.; Netesov, S.V.; Tikunova, N.V. High Evolutionary Rate of Human Astrovirus. Infect. Genet. Evol. 2012, 12, 435–442. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. De Grazia, S.; Medici, M.C.; Pinto, P.; Moschidou, P.; Tummolo, F.; Calderaro, A.; Bonura, F.; Banyai, K.; Giammanco, G.M.; Martella, V. Genetic Heterogeneity and Recombination in Human Type 2 Astroviruses. J. Clin. Microbiol. 2012, 50, 3760–3764. [Google Scholar] [CrossRef] [Scilit]
  10. Medici, M.C.; Tummolo, F.; Martella, V.; Banyai, K.; Bonerba, E.; Chezzi, C.; Arcangeletti, M.C.; De Conto, F.; Calderaro, A. Genetic Heterogeneity and Recombination in Type-3 Human Astroviruses. Infect. Genet. Evol. 2015, 32, 156–160. [Google Scholar] [CrossRef] [Scilit]
  11. Ha, H.J.; Lee, S.G.; Cho, H.G.; Jin, J.Y.; Lee, J.W.; Paik, S.Y. Complete Genome Sequencing of a Recombinant Strain between Human Astrovirus Antigen Types 2 and 8 Isolated from South Korea. Infect. Genet. Evol. 2016, 39, 127–131. [Google Scholar] [CrossRef] [Scilit]
  12. Zaraket, H.; Abou-El-Hassan, H.; Kreidieh, K.; Soudani, N.; Ali, Z.; Hammadi, M.; Reslan, L.; Ghanem, S.; Hajar, F.; Inati, A.; et al. Characterization of Astrovirus-Associated Gastroenteritis in Hospitalized Children under Five Years of Age. Infect. Genet. Evol. 2017, 53, 94–99. [Google Scholar] [CrossRef] [Scilit]
  13. Walter, J.E.; Briggs, J.; Guerrero, M.L.; Matson, D.O.; Pickering, L.K.; Ruiz-Palacios, G.; Berke, T.; Mitchell, D.K. Molecular Characterization of a Novel Recombinant Strain of Human Astrovirus Associated with Gastroenteritis in Children. Arch. Virol. 2001, 146, 2357–2367. [Google Scholar] [CrossRef] [Scilit]
  14. Wolfaardt, M.; Kiulia, N.M.; Mwenda, J.M.; Taylor, M.B. Evidence of a Recombinant Wild-Type Human Astrovirus Strain from a Kenyan Child with Gastroenteritis. J. Clin. Microbiol. 2011, 49, 728–731. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Babkin, I.V.; Tikunov, A.Y.; Sedelnikova, D.A.; Zhirakovskaia, E.V.; Tikunova, N.V. Recombination Analysis Based on the HAstV-2 and HAstV-4 Complete Genomes. Infect. Genet. Evol. 2014, 22, 94–102. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Ito, M.; Kuroda, M.; Masuda, T.; Akagami, M.; Haga, K.; Tsuchiaka, S.; Kishimoto, M.; Naoi, Y.; Sano, K.; Omatsu, T.; et al. Whole Genome Analysis of Porcine Astroviruses Detected in Japanese Pigs Reveals Genetic Diversity and Possible Intra-Genotypic Recombination. Infect. Genet. Evol. 2017, 50, 38–48. [Google Scholar] [CrossRef] [Scilit]
  17. Zhao, C.; Chen, C.; Li, Y.; Dong, S.; Tan, K.; Tian, Y.; Zhang, L.L.; Huang, J.; Zhang, L.L. Genomic Characterization of a Novel Recombinant Porcine Astrovirus Isolated in Northeastern China. Arch. Virol. 2019, 164, 1469–1473. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Zhang, W.; Wang, R.; Liang, J.; Zhao, N.; Li, G.; Gao, Q.; Su, S. Epidemiology, Genetic Diversity and Evolution of Canine Astrovirus. Transbound. Emerg. Dis. 2020, 67, 2901–2910. [Google Scholar] [CrossRef] [Scilit]
  19. Li, M.; Yan, N.; Ji, C.; Wang, M.; Zhang, B.; Yue, H.; Tang, C. Prevalence and Genome Characteristics of Canine Astrovirus in Southwest China. J. Gen. Virol. 2018, 99, 880–889. [Google Scholar] [CrossRef] [Scilit]
  20. Hata, A.; Kitajima, M.; Haramoto, E.; Lee, S.; Ihara, M.; Gerba, C.P.; Tanaka, H. Next-Generation Amplicon Sequencing Identifies Genetically Diverse Human Astroviruses, Including Recombinant Strains, in Environmental Waters. Sci. Rep. 2018, 8, 11837. [Google Scholar] [CrossRef] [Scilit]
  21. Ahmed, S.F.; Sebeny, P.J.; Klena, J.D.; Pimentel, G.; Mansour, A.; Naguib, A.M.; Bruton, J.; Young, S.Y.N.; Holtz, L.R.; Wang, D. Novel Astroviruses in Children, Egypt. Emerg. Infect. Dis. 2011, 17, 2391–2393. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Amimo, J.O.; Machuka, E.M.; Abworo, E.O.; Vlasova, A.N.; Pelle, R. Whole Genome Sequence Analysis of Porcine Astroviruses Reveals Novel Genetically Diverse Strains Circulating in East African Smallholder Pig Farms. Viruses 2020, 12, 1262. [Google Scholar] [CrossRef] [Scilit]
  23. Aleshina, Y.; Lukashev, A. Mamastrovirus Species Are Shaped by Recombination and Can Be Reliably Distinguished in ORF1b Genome Region. Virus Evol. 2025, 11, veaf006. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Wei, H.; Kumthip, K.; Khamrin, P.; Yodmeeklin, A.; Jampanil, N.; Phengma, P.; Xie, Z.; Ukarapol, N.; Ushijima, H.; Maneekarn, N. Triple Intergenotype Recombination of Human Astrovirus 5, Human Astrovirus 8, and Human Astrovirus 1 in the Open Reading Frame 1a, Open Reading Frame 1b, and Open Reading Frame 2 Regions of the Human Astrovirus Genome. Microbiol. Spectr. 2023, 11, e0488822. [Google Scholar] [CrossRef] [Scilit]
  25. Pativada, M.; Bhattacharya, R.; Krishnan, T. Novel Human Astrovirus Strains Showing Multiple Recombinations within Highly Conserved ORF1b Detected from Hospitalized Acute Watery Diarrhea Cases in Kolkata, India. Infect. Genet. Evol. 2013, 20, 284–291. [Google Scholar] [CrossRef] [Scilit]
  26. Atkins, A.; Wellehan, J.F.X.; Childress, A.L.; Archer, L.L.; Fraser, W.A.; Citino, S.B. Characterization of an Outbreak of Astroviral Diarrhea in a Group of Cheetahs (Acinonyx Jubatus). Vet. Microbiol. 2009, 136, 160–165. [Google Scholar] [CrossRef] [Scilit]
  27. Ulloa, J.C.; Gutiérrez, M.F. Genomic Analysis of Two ORF2 Segments of New Porcine Astrovirus Isolates and Their Close Relationship with Human Astroviruses. Can. J. Microbiol. 2010, 56, 569–577. [Google Scholar] [CrossRef] [Scilit]
  28. Rivera, R.; Nollens, H.H.; Venn-Watson, S.; Gulland, F.M.D.D.; Wellehan, J.F.X.X. Characterization of Phylogenetically Diverse Astroviruses of Marine Mammals. J. Gen. Virol. 2010, 91, 166–173. [Google Scholar] [CrossRef] [Scilit]
  29. Wei, H.; Khamrin, P.; Kumthip, K.; Yodmeeklin, A.; Maneekarn, N. Emergence of Multiple Novel Inter-Genotype Recombinant Strains of Human Astroviruses Detected in Pediatric Patients with Acute Gastroenteritis in Thailand. Front. Microbiol. 2021, 12, 789636. [Google Scholar] [CrossRef] [Scilit]
  30. Pativada, M.S.; Chatterjee, D.; Mariyappa, N.S.; Rajendran, K.; Bhattacharya, M.K.; Ghosh, M.; Kobayashi, N.; Krishnan, T. Emergence of Unique Variants and Inter-Genotype Recombinants of Human Astroviruses Infecting Infants, Children and Adults in Kolkata, India. Int. J. Mol. Epidemiol. Genet. 2011, 2, 228–235. [Google Scholar] [PubMed]
  31. Martella, V.; Pinto, P.; Tummolo, F.; De Grazia, S.; Giammanco, G.M.; Medici, M.C.; Ganesh, B.; L’Homme, Y.; Farkas, T.; Jakab, F.; et al. Analysis of the ORF2 of Human Astroviruses Reveals Lineage Diversification, Recombination and Rearrangement and Provides the Basis for a Novel Sub-Classification System. Arch. Virol. 2014, 159, 3185–3196. [Google Scholar] [CrossRef] [Scilit]
  32. Tamura, K.; Stecher, G.; Kumar, S. MEGA11: Molecular Evolutionary Genetics Analysis Version 11. Mol. Biol. Evol. 2021, 38, 3022–3027. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Altschul, S.F.; Gish, W.; Miller, W.; Myers, E.W.; Lipman, D.J. Basic Local Alignment Search Tool. J. Mol. Biol. 1990, 215, 403–410. [Google Scholar] [CrossRef] [PubMed]
  34. Trifinopoulos, J.; Nguyen, L.-T.; von Haeseler, A.; Minh, B.Q. W-IQ-TREE: A Fast Online Phylogenetic Tool for Maximum Likelihood Analysis. Nucleic Acids Res. 2016, 44, W232–W235. [Google Scholar] [CrossRef] [Scilit]
  35. Martin, D.P.; Murrell, B.; Golden, M.; Khoosal, A.; Muhire, B. RDP4: Detection and Analysis of Recombination Patterns in Virus Genomes. Virus Evol. 2015, 1, vev003. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Martin, D.; Rybicki, E. RDP: Detection of Recombination amongst Aligned Sequences. Bioinformatics 2000, 16, 562–563. [Google Scholar] [CrossRef] [Scilit]
  37. Sawyer, S. Statistical Tests for Detecting Gene Conversion. Mol. Biol. Evol. 1989, 6, 526–538. [Google Scholar] [CrossRef] [Scilit]
  38. Smith, J.M. Analyzing the Mosaic Structure of Genes. J. Mol. Evol. 1992, 34, 126–129. [Google Scholar] [CrossRef] [Scilit]
  39. Posada, D.; Crandall, K.A. Evaluation of Methods for Detecting Recombination from DNA Sequences: Computer Simulations. Proc. Natl. Acad. Sci. USA 2001, 98, 13757–13762. [Google Scholar] [CrossRef] [Scilit]
  40. Martin, D.P.; Posada, D.; Crandall, K.A.; Williamson, C. A Modified Bootscan Algorithm for Automated Identification of Recombinant Sequences and Recombination Breakpoints. AIDS Res. Hum. Retroviruses 2005, 21, 98–102. [Google Scholar] [CrossRef] [Scilit]
  41. Gibbs, M.J.; Armstrong, J.S.; Gibbs, A.J. Sister-Scanning: A Monte Carlo Procedure for Assessing Signals in Rebombinant Sequences. Bioinformatics 2000, 16, 573–582. [Google Scholar] [CrossRef] [Scilit]
  42. Boni, M.F.; Posada, D.; Feldman, M.W. An Exact Nonparametric Method for Inferring Mosaic Structure in Sequence Triplets. Genetics 2007, 176, 1035–1047. [Google Scholar] [CrossRef] [Scilit]
  43. Weiller, G.F. Phylogenetic Profiles: A Graphical Method for Detecting Genetic Recombinations in Homologous Sequences. Mol. Biol. Evol. 1998, 15, 326–335. [Google Scholar] [CrossRef] [Scilit]
  44. Kalyaanamoorthy, S.; Minh, B.Q.; Wong, T.K.F.; von Haeseler, A.; Jermiin, L.S. ModelFinder: Fast Model Selection for Accurate Phylogenetic Estimates. Nat. Methods 2017, 14, 587–589. [Google Scholar] [CrossRef] [Scilit]
  45. Minh, B.Q.; Nguyen, M.A.T.; von Haeseler, A. Ultrafast Approximation for Phylogenetic Bootstrap. Mol. Biol. Evol. 2013, 30, 1188–1195. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Rambaut, A.; Lam, T.; Carvalho, L.; Pybus, O. Exploring the Temporal Structure of Heterochronous Sequences Using TempEst (Formerly Path-O-Gen). Virus Evol. 2016, 2, vew007. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Bouckaert, R.; Vaughan, T.G.; Barido-Sottani, J.; Duchêne, S.; Fourment, M.; Gavryushkina, A.; Heled, J.; Jones, G.; Kühnert, D.; De Maio, N.; et al. BEAST 2.5: An Advanced Software Platform for Bayesian Evolutionary Analysis. PLoS Comput. Biol. 2019, 15, e1006650. [Google Scholar] [CrossRef] [Scilit]
  48. Lanfear, R.; Frandsen, P.B.; Wright, A.M.; Senfeld, T.; Calcott, B. PartitionFinder 2: New Methods for Selecting Partitioned Models of Evolution for Molecular and Morphological Phylogenetic Analyses. Mol. Biol. Evol. 2017, 34, 772–773. [Google Scholar] [CrossRef] [Scilit]
  49. Russel, P.M.; Brewer, B.J.; Klaere, S.; Bouckaert, R.R. Model Selection and Parameter Inference in Phylogenetics Using Nested Sampling. Syst. Biol. 2019, 68, 219–233. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. Duchene, S.; Featherstone, L.; Haritopoulou-Sinanidou, M.; Rambaut, A.; Lemey, P.; Baele, G. Temporal Signal and the Phylodynamic Threshold of SARS-CoV-2. Virus Evol. 2020, 6, veaa061. [Google Scholar] [CrossRef] [Scilit]
  51. Duchene, S.; Lemey, P.; Stadler, T.; Ho, S.Y.W.; Duchene, D.A.; Dhanasekaran, V.; Baele, G. Bayesian Evaluation of Temporal Signal in Measurably Evolving Populations. Mol. Biol. Evol. 2020, 37, 3363–3379. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Rambaut, A.; Drummond, A.; Xie, D.; Baele, G.; Suchard, M. Posterior Summarisation in Bayesian Phylogenetics Using Tracer 1.7. Syst. Biol. 2018, 67, 901–904. [Google Scholar] [CrossRef] [Scilit]
  53. Yu, G. Using Ggtree to Visualize Data on Tree-like Structures. Curr. Protoc. Bioinforma. 2020, 69, e96. [Google Scholar] [CrossRef] [Scilit]
  54. Lukashev, A.N.; Shumilina, E.Y.; Belalov, I.S.; Ivanova, O.E.; Eremeeva, T.P.; Reznik, V.I.; Trotsenko, O.E.; Drexler, J.F.; Drosten, C. Recombination Strategies and Evolutionary Dynamics of the Human Enterovirus A Global Gene Pool. J. Gen. Virol. 2014, 95, 868–873. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Vakulenko, Y.A.; Orlov, A.V.; Lukashev, A.N. Patterns and Temporal Dynamics of Natural Recombination in Noroviruses. Viruses 2023, 15, 372. [Google Scholar] [CrossRef] [Scilit]
  56. Malichava, M.; Lukashev, A.; Aleshina, Y. Temporal Dynamics of Recombination in Field Isolates of Foot-and-Mouth Disease Virus. Viruses 2026, 18, 262. [Google Scholar] [CrossRef] [Scilit]
  57. Lukashev, A. Recombination among Picornaviruses. Rev. Med. Virol. 2010, 20, 327–337. [Google Scholar] [CrossRef] [Scilit]
  58. de Klerk, A.; Swanepoel, P.; Lourens, R.; Zondo, M.; Abodunran, I.; Lytras, S.; MacLean, O.A.; Robertson, D.; Kosakovsky Pond, S.L.; Zehr, J.D.; et al. Conserved Recombination Patterns across Coronavirus Subgenera. Virus Evol. 2022, 8, veac054. [Google Scholar] [CrossRef] [Scilit]
  59. Vakulenko, Y.; Deviatkin, A.; Drexler, J.F.; Lukashev, A. Modular Evolution of Coronavirus Genomes. Viruses 2021, 13, 1270. [Google Scholar] [CrossRef] [Scilit]
  60. Begall, L.F.L.; Mauroy, A.; Thiry, E. Norovirus Recombinants: Recurrent in the Field, Recalcitrant in the Lab—A Scoping Review of Recombination and Recombinant Types of Noroviruses. J. Gen. Virol. 2018, 99, 970–988. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. Tohma, K.; Kulka, M.; Coughlan, S.; Green, K.Y.; Parra, G.I. Genomic Analyses of Human Sapoviruses Detected over a 40-Year Period Reveal Disparate Patterns of Evolution among Genotypes and Genome Regions. Viruses 2020, 12, 516. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. dos Anjos, K.; Lima, L.M.P.; Silva, P.A.; Inoue-Nagata, A.K.; Nagata, T. The Possible Molecular Evolution of Sapoviruses by Inter- and Intra-Genogroup Recombination. Arch. Virol. 2011, 156, 1953–1959. [Google Scholar] [CrossRef] [Scilit]
  63. Mahar, J.E.; Hall, R.N.; Shi, M.; Mourant, R.; Huang, N.; Strive, T.; Holmes, E.C. The Discovery of Three New Hare Lagoviruses Reveals Unexplored Viral Diversity in This Genus. Virus Evol. 2019, 5, vez005. [Google Scholar] [CrossRef] [Scilit]
  64. Graham, R.L.; Baric, R.S. Recombination, Reservoirs, and the Modular Spike: Mechanisms of Coronavirus Cross-Species Transmission. J. Virol. 2010, 84, 3134–3146. [Google Scholar] [CrossRef] [Scilit]
  65. Chhabra, P.; de Graaf, M.; Parra, G.I.; Chan, M.C.W.; Green, K.; Martella, V.; Wang, Q.; White, P.A.; Katayama, K.; Vennema, H.; et al. Updated Classification of Norovirus Genogroups and Genotypes. J. Gen. Virol. 2019, 100, 1393–1406, Corrigendum in J. Gen. Virol. 2020, 101, 893. https://doi.org/10.1099/jgv.0.001475. [Google Scholar] [CrossRef] [Scilit]
  66. McWilliam Leitch, E.C.; Cabrerizo, M.; Cardosa, J.; Harvala, H.; Ivanova, O.E.; Kroes, A.C.M.; Lukashev, A.; Muir, P.; Odoom, J.; Roivainen, M.; et al. Evolutionary Dynamics and Temporal/Geographical Correlates of Recombination in the Human Enterovirus Echovirus Types 9, 11, and 30. J. Virol. 2010, 84, 9292–9300. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  67. McWilliam Leitch, E.C.; Bendig, J.; Cabrerizo, M.; Cardosa, J.; Hyypia, T.; Ivanova, O.E.; Kelly, A.; Kroes, A.C.M.; Lukashev, A.; MacAdam, A.; et al. Transmission Networks and Population Turnover of Echovirus 30. J. Virol. 2009, 83, 2109–2118. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  68. McWilliam Leitch, E.C.; Cabrerizo, M.; Cardosa, J.; Harvala, H.; Ivanova, O.E.; Koike, S.; Kroes, A.C.M.; Lukashev, A.; Perera, D.; Roivainen, M.; et al. The Association of Recombination Events in the Founding and Emergence of Subgenogroup Evolutionary Lineages of Human Enterovirus 71. J. Virol. 2012, 86, 2676–2685. [Google Scholar] [CrossRef] [Scilit]
  69. Simmonds, P. Recombination and Selection in the Evolution of Picornaviruses and Other Mammalian Positive-Stranded RNA Viruses. J. Virol. 2006, 80, 11124–11140. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Recombination analysis of classical human astroviruses. The analysis was performed in RDP4 v101 software [35]. (a) Distribution of recombination breakpoints across astrovirus ORFs. A schematic of the astrovirus open reading frames (ORFs) and mature proteins are shown at the top (TM—transmembrane protein). Individual breakpoint positions are indicated by tick marks below this schematic. The lower panel displays the density of detected breakpoints (solid black line) calculated using a 200-nt sliding window with a 10-nt step. The light and dark gray areas represent the local 95% and 99% confidence intervals, respectively, generated from 1000 permutations. Regions where the observed recombination event count exceeds these confidence intervals represent significant recombination hotspots. Genome regions selected for subsequent temporal analysis are highlighted: those with a temporal signal are shaded green, while those without are shaded red. Nucleotide positions are shown in relation to OR371570 (HAstV4) sequence. (b) Recombination region count matrix. Each cell of the matrix represents a pair of nucleotide positions along the genome; the color scale (from blue to red) indicates the number of unique recombination events (mapped in panel (a)) partitioning that specific site pair. Warmer colors indicate genomic regions more frequently exchanged through recombination.
Figure 1. Recombination analysis of classical human astroviruses. The analysis was performed in RDP4 v101 software [35]. (a) Distribution of recombination breakpoints across astrovirus ORFs. A schematic of the astrovirus open reading frames (ORFs) and mature proteins are shown at the top (TM—transmembrane protein). Individual breakpoint positions are indicated by tick marks below this schematic. The lower panel displays the density of detected breakpoints (solid black line) calculated using a 200-nt sliding window with a 10-nt step. The light and dark gray areas represent the local 95% and 99% confidence intervals, respectively, generated from 1000 permutations. Regions where the observed recombination event count exceeds these confidence intervals represent significant recombination hotspots. Genome regions selected for subsequent temporal analysis are highlighted: those with a temporal signal are shaded green, while those without are shaded red. Nucleotide positions are shown in relation to OR371570 (HAstV4) sequence. (b) Recombination region count matrix. Each cell of the matrix represents a pair of nucleotide positions along the genome; the color scale (from blue to red) indicates the number of unique recombination events (mapped in panel (a)) partitioning that specific site pair. Warmer colors indicate genomic regions more frequently exchanged through recombination.
Microorganisms 14 00857 g001
Figure 2. Evaluation of temporal signal across five recombination-free regions of human astroviruses. Root-to-tip regression analysis was performed in TempEst v.1.5.3 on ML phylogenetic trees for: 5′ half of ORF1a (a), 3′ half of ORF1a (b), ORF1b (c), 5′ end of ORF2 (d), and 3′ end of ORF2 (e). Each panel shows the linear regression fit (red line), regression equation, p-value for regression slope. The regression slope indicates substitution rate (substitutions/site/year).
Figure 2. Evaluation of temporal signal across five recombination-free regions of human astroviruses. Root-to-tip regression analysis was performed in TempEst v.1.5.3 on ML phylogenetic trees for: 5′ half of ORF1a (a), 3′ half of ORF1a (b), ORF1b (c), 5′ end of ORF2 (d), and 3′ end of ORF2 (e). Each panel shows the linear regression fit (red line), regression equation, p-value for regression slope. The regression slope indicates substitution rate (substitutions/site/year).
Microorganisms 14 00857 g002
Figure 3. Inference of half-life of recombinant forms that occurred due to recombination between 3′-part of ORF1a and ORF1b genome regions of classical human astroviruses. Bayesian time-scaled trees were built using 3′-part of ORF1a (a) and ORF1b (b) genome regions. In both trees, black circles indicate nodes with high support (posterior probability > 0.9). Branches of clades that coincide between the trees are colored red. Taxa labels are colored according to the serotype.
Figure 3. Inference of half-life of recombinant forms that occurred due to recombination between 3′-part of ORF1a and ORF1b genome regions of classical human astroviruses. Bayesian time-scaled trees were built using 3′-part of ORF1a (a) and ORF1b (b) genome regions. In both trees, black circles indicate nodes with high support (posterior probability > 0.9). Branches of clades that coincide between the trees are colored red. Taxa labels are colored according to the serotype.
Microorganisms 14 00857 g003
Figure 4. Distribution of lifetimes (time between tMRCA and the most recent isolate) of coinciding clades (recombinant forms) inferred from phylogenetic comparisons between Bayesian time-scaled phylogenies of reference regions with temporal signal, the 3′-part of ORF1a (a) and ORF1b (b) and trees from other genome regions: (i) 5′-part of ORF1a (ML tree), (ii) ORF1b (Bayesian phylogeny), (iii) 3′-part of ORF2 (ML tree). Dashed red line indicates recombinant form half-life (median of clade lifetimes).
Figure 4. Distribution of lifetimes (time between tMRCA and the most recent isolate) of coinciding clades (recombinant forms) inferred from phylogenetic comparisons between Bayesian time-scaled phylogenies of reference regions with temporal signal, the 3′-part of ORF1a (a) and ORF1b (b) and trees from other genome regions: (i) 5′-part of ORF1a (ML tree), (ii) ORF1b (Bayesian phylogeny), (iii) 3′-part of ORF2 (ML tree). Dashed red line indicates recombinant form half-life (median of clade lifetimes).
Microorganisms 14 00857 g004
Figure 5. Half-lives of recombinant forms (in years) inferred from phylogenetic comparisons between genome regions involved in recombination. Top: Schematic representation of the human astrovirus genome, with regions used for phylogenetic comparison highlighted in brighter colors and marked with a diagonal pattern. Bottom: Median lifetimes of clades coinciding between Bayesian time-scaled phylogenies of the 3′-part of ORF1a and ORF1b (rows) and maximum likelihood trees of other genome regions (columns). For comparison of the 3′-part of ORF1a and ORF1b, the Bayesian phylogenies were utilized. Dashes indicate cells corresponding to the same tree in a row and a column.
Figure 5. Half-lives of recombinant forms (in years) inferred from phylogenetic comparisons between genome regions involved in recombination. Top: Schematic representation of the human astrovirus genome, with regions used for phylogenetic comparison highlighted in brighter colors and marked with a diagonal pattern. Bottom: Median lifetimes of clades coinciding between Bayesian time-scaled phylogenies of the 3′-part of ORF1a and ORF1b (rows) and maximum likelihood trees of other genome regions (columns). For comparison of the 3′-part of ORF1a and ORF1b, the Bayesian phylogenies were utilized. Dashes indicate cells corresponding to the same tree in a row and a column.
Microorganisms 14 00857 g005
Figure 6. Inference of half-life time of recombinant forms that were made up by recombination between ORF1b and 3′-part of ORF2 genome regions of classical human astroviruses. (a) Bayesian time-scaled tree for ORF1b genome region. (b) ML phylogeny inferred using the 3′-part of ORF2 region. Black circles indicate nodes with high support (posterior probability > 0.9 in (a), ultrafast bootstrap >95% in (b)). Clades (bipartitions) that coincide between trees are colored red. Taxa labels are colored according to the serotype.
Figure 6. Inference of half-life time of recombinant forms that were made up by recombination between ORF1b and 3′-part of ORF2 genome regions of classical human astroviruses. (a) Bayesian time-scaled tree for ORF1b genome region. (b) ML phylogeny inferred using the 3′-part of ORF2 region. Black circles indicate nodes with high support (posterior probability > 0.9 in (a), ultrafast bootstrap >95% in (b)). Clades (bipartitions) that coincide between trees are colored red. Taxa labels are colored according to the serotype.
Microorganisms 14 00857 g006
Table 2. The number of sequences of each serotype in the refined dataset.
Table 2. The number of sequences of each serotype in the refined dataset.
HAstV SerotypeNumber of Complete Genomes
Serotype 132
Serotype 25
Serotype 39
Serotype 419
Serotype 59
Serotype 62
Serotype 70
Serotype 81
Table 3. Results of BETS analysis for two genome regions with a detectable temporal signal. BETS analysis compares the fit of a heterochronous model (utilizing true sampling dates) to an isochronous null model (all dates set equal). Marginal likelihood estimates were calculated using the nested sampling algorithm implemented in BEAST2 v2.7.7. A log Bayes factor greater than 3 indicates strong support for the presence of temporal signal.
Table 3. Results of BETS analysis for two genome regions with a detectable temporal signal. BETS analysis compares the fit of a heterochronous model (utilizing true sampling dates) to an isochronous null model (all dates set equal). Marginal likelihood estimates were calculated using the nested sampling algorithm implemented in BEAST2 v2.7.7. A log Bayes factor greater than 3 indicates strong support for the presence of temporal signal.
3′-Part of ORF1aORF1b
Isochronous model−11,916.34−8013.75
Heterochronous−11,779.94−7969.57
Log Bayes Factor136.4044.10
Table 4. The evolutionary parameters of human astroviruses estimated for recombination free regions using BEAST2 v2.7.7.
Table 4. The evolutionary parameters of human astroviruses estimated for recombination free regions using BEAST2 v2.7.7.
Genome RegionSubstitution Rate [95% HPD Confidence Interval] × 10−3, s/s/ytMRCA [95% HPD Confidence Interval], Years
3′-part of ORF1a2.35 [1.85–2.86]251.38 [195.83–313.11]
ORF1b2.13 [1.52–2.78]208.04 [145.07–278.21]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Aleshina, Y.; Frantsuzov, V.; Lukashev, A. Complex Recombination Landscape and Lineage Turnover in Classical Human Astroviruses. Microorganisms 2026, 14, 857. https://doi.org/10.3390/microorganisms14040857

AMA Style

Aleshina Y, Frantsuzov V, Lukashev A. Complex Recombination Landscape and Lineage Turnover in Classical Human Astroviruses. Microorganisms. 2026; 14(4):857. https://doi.org/10.3390/microorganisms14040857

Chicago/Turabian Style

Aleshina, Yulia, Vladimir Frantsuzov, and Alexander Lukashev. 2026. "Complex Recombination Landscape and Lineage Turnover in Classical Human Astroviruses" Microorganisms 14, no. 4: 857. https://doi.org/10.3390/microorganisms14040857

APA Style

Aleshina, Y., Frantsuzov, V., & Lukashev, A. (2026). Complex Recombination Landscape and Lineage Turnover in Classical Human Astroviruses. Microorganisms, 14(4), 857. https://doi.org/10.3390/microorganisms14040857

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop