Next Article in Journal
Transcriptome Dynamics Reveal the Potential Roles of Long Non-Coding RNAs in Regulating Flower Color of Safflowers (Carthamus tinctorius)
Previous Article in Journal
Novel Disease-Specific Panel of Salivary microRNAs for the Detection of Oral Squamous Cell Carcinoma from Early Invasion to Stage IV Disease
Previous Article in Special Issue
Potyvirus HcPro Suppressor of RNA Silencing Induces PVY Superinfection Exclusion in a Strain-Specific Manner
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Identifying Conserved Regions in HIV-1 Proteins by Entropy Analysis of Sequence Variability

by
Alexandr N. Shchemelev
*,
Elena N. Serikova
,
Yulia V. Ostankova
,
Vladimir S. Davydenko
,
Edward S. Ramsay
and
Areg A. Totolian
Saint Petersburg Pasteur Institute, 197101 St. Petersburg, Russia
*
Author to whom correspondence should be addressed.
Int. J. Mol. Sci. 2026, 27(11), 5139; https://doi.org/10.3390/ijms27115139
Submission received: 27 March 2026 / Revised: 2 June 2026 / Accepted: 3 June 2026 / Published: 5 June 2026
(This article belongs to the Special Issue Viral Infections and Viral Pathogenesis)

Abstract

The extraordinary genetic diversity of human immunodeficiency virus type 1 (HIV-1), driven by high mutation and recombination rates, poses significant challenges for diagnostics, therapy, and vaccine development. While variable regions enable immune escape, hyperconserved regions are critical for viral function and represent promising targets for novel therapeutic interventions. This study aimed to develop and validate a bioinformatic algorithm for quantitative assessment of sequence conservation and automated identification of functionally significant conserved regions across all major HIV-1 proteins. A total of 1119 full-length HIV-1 genome sequences representing major subtypes (A1, A2, A6, B, C, D, F1, F2, G, H, J, K) were analyzed. Normalized Shannon entropy (S-index) was calculated for each alignment column. Statistical thresholds for conserved regions were established using 95% confidence intervals derived from bootstrap resampling. Two complementary algorithms, clustering and local maxima detection, were applied to identify conserved regions, which were subsequently mapped to known functional domains based on literature data. Protein conservation varied markedly, with Sm values ranging from 0.784 (Vpu) to 0.920 (Pol). Gag, Pol, and Vpr demonstrated the highest overall conservation, while Env, Rev, Tat, and Vpu exhibited pronounced variability interspersed with conserved domains. In total, 25 conserved regions in Gag, 49 in Pol, 28 in Env, and 6–4 regions in accessory proteins (Vif, Vpr, Rev, Tat, Nef, Vpu) were identified. These regions corresponded to critical functional elements including enzyme catalytic centers, zinc fingers, receptor-binding sites, protein interaction interfaces, and membrane-anchoring domains. The developed computational framework enables statistically grounded identification of evolutionarily constrained regions across analyzed HIV-1 subtypes. The identified conserved regions represent candidate sites for further investigation and may inform downstream studies focused on antiviral target prioritization, immunogen design, and diagnostic assay development. However, their translational applicability requires additional analytical, structural, and experimental validation.

1. Introduction

1.1. Global HIV Variability

The human immunodeficiency virus (HIV) remains one of the most serious global public health challenges. According to data from the World Health Organization (WHO), as of the end of 2024, approximately 40.8 million people worldwide were living with HIV, and 1.3 million new infections were reported [1]. One of the key characteristics that underlies the complexity of controlling this infection is the extraordinary genetic diversity of the virus, which exhibits distinct geographic patterns.
The virus is classified into two main types: HIV-1, which accounts for the overwhelming majority of infections globally, and HIV-2, which is less virulent and is predominantly distributed in West Africa. HIV-1 has the greatest epidemiological significance and is further divided into several groups. Group M (“major”) is responsible for more than 90% of all infections and comprises multiple subtypes (A, B, C, D, F, G, H, J, K) as well as circulating recombinant forms (CRFs), which differ from one another by approximately 25–35% at the nucleotide sequence level. Subtype B has historically predominated in Europe, North America, and Australia. Subtype C is widespread in Southern Africa and India. Subtype A and recombinant forms derived from it, such as CRF02_AG, circulate in West and East Africa [2], as well as in countries of Eastern Europe and Central Asia [3].
The Russian Federation provides an illustrative example of changes in the landscape of circulating HIV-1 variants. Whereas in 2008–2010, subtype A accounted for 91.3% of cases and subtype B for 8.7%, during 2011–2014 the spectrum of genetic diversity expanded due to the emergence of recombinant forms (AB, AG, CRF06_cpx) and subtype C [4,5,6]. This genetic diversity affects diagnostic accuracy, therapeutic efficacy, and vaccine development strategies, rendering global surveillance of HIV strains critically important.

1.2. Biological Nature of HIV Variability

The fundamental basis of the high variability of HIV lies in the characteristics of its replicative cycle as a retrovirus. A key role is played by the enzyme reverse transcriptase, which synthesizes DNA using viral RNA as a template. This enzyme lacks proofreading and error-correction mechanisms, resulting in the frequent occurrence of mutations at a rate of (4.1 ± 1.7) × 10−3 per nucleotide, which is the highest value reported for any biological entity [7]. More than one billion new viral particles are produced daily in an infected organism; this intensity, combined with the high mutation rate, generates an enormous number of genetic variants [8].
This process is further exacerbated by the virus’s capacity for recombination. When a single cell is infected by two different HIV strains, their genetic material can “mix” during the assembly of new virions, leading to the emergence of CRFs that may possess novel properties [9]. Natural selection exerted by the human immune system and, critically, by antiretroviral therapy (ART) favors mutant variants capable of evading immune control or exhibiting resistance to antiretroviral drugs.
Of particular interest is the interaction between the virus and host chemokine receptors. The majority of primary infections are caused by viral variants (R5) that utilize the CCR5 receptor. Human genetic polymorphisms, such as the well-characterized Δ32 deletion in the CCR5 gene, result in the production of a defective receptor and may confer resistance to HIV infection [10]. Other mutations, for example in the CCR2 or CXCL12 genes, are associated with delayed disease progression.

1.3. Methods of Genetic Variability Research

Contemporary analyses of the genetic variability of HIV-1 are based on integrated bioinformatic approaches that enable investigation of viral evolution at multiple structural and functional levels. Phylogenetic analysis remains a fundamental tool for reconstructing the evolutionary history of viral isolates and tracing transmission pathways. With the advent of next-generation sequencing (NGS) technologies, it has become possible to study not individual strains, but the entire spectrum of HIV quasispecies within a single patient, which is critically important for understanding the mechanisms underlying the development of drug resistance [11]. However, a mere description of genetic diversity has proven insufficient for addressing applied problems in antiviral therapy.
To identify functionally significant regions of the viral genome, analyses of evolutionary conservation are widely employed based on the calculation of Shannon entropy or similar metrics of nucleotide sequence variability [12,13]. Hyperconserved regions, which exhibit minimal variability even among different HIV-1 subtypes, are considered promising targets for the development of novel therapeutic agents as their stability indicates critical functional importance for the viral life cycle. In particular, the structural proteins capsid p24 and nucleocapsid p7 demonstrate a high degree of conservation, which explains their attractiveness as targets for new classes of drugs, such as capsid inhibitors [14].
Studies have shown that even highly conserved regions, such as the 5′ untranslated region (5′ UTR), display substantial interstrain variability, up to 17% in the U5–PBS region and up to 20% in the gag leader sequence (GLS), among different subtypes and [15]. This variability may create strain-specific transcription factor binding sites (for example, an E-box in CRF22_01A1 or Stat6 in subtypes A and G), potentially affecting the efficiency of viral replication. Hyperconserved regions exhibiting minimal variability even between different HIV-1 subtypes are therefore regarded as promising targets for the development of new therapeutic strategies, as their stability reflects their critical functional role in the viral life cycle.
Of particular importance is the analysis of covariation networks of amino acid residues, which enables the identification of compensatory mutations arising in response to the selective pressure exerted by antiretroviral therapy. For example, mutations conferring resistance to reverse transcriptase inhibitors are often accompanied by specific changes in other regions of the protein that restore the functional activity of the enzyme. Modern algorithms, such as methods for detecting coevolving positions (e.g., Direct Coupling Analysis), allow prediction of the emergence of such compensatory mutations and may be used to optimize antiretroviral treatment regimens.
Molecular modeling and prediction of the three-dimensional structures of viral proteins harboring resistance-associated mutations make it possible to interpret observed genetic changes at the structural level. Machine learning-based approaches are increasingly applied for the classification of viral subtypes, prediction of antiretroviral drug resistance, and identification of novel potential therapeutic targets. Nevertheless, despite the abundance of bioinformatic methods, there remains a lack of standardized solutions for comprehensive screening of large sequence datasets with simultaneous assessment of conservation, variability, and functional relevance of identified regions. The development of integrated algorithms combining phylogenetic analysis, evaluation of evolutionary pressure, and prediction of the structural and functional consequences of mutations represents a pressing challenge in HIV bioinformatics.
Thus, the aim of the present study was to develop an algorithm for the screening and analysis of nucleotide and amino acid sequence datasets that enables quantitative assessment of their conservation and variability, as well as automated identification of hyperconserved regions with potentially critical functional significance.

2. Results

2.1. Determination of Conservation Parameters for Sequence Alignment and Establishment of Cutoff Thresholds

Calculation results for amino acid sequence alignments are presented in Figure 1 and Table 1. Among the sufficiently well-represented HIV-1 subtypes (i.e., those for which more than 20 sequences were present in the alignment), subtypes D, F1, and G exhibited higher levels of diversity, whereas subtypes A6, B, and H were the most conserved. For several genotypes (A2, F2, J, K), insufficient numbers of sequences meeting the inclusion criteria were available to allow an unambiguous assessment of their degree of conservation. Nevertheless, cutoff thresholds for the identification of highly conserved and highly variable regions were also calculated for these genotypes for subsequent analyses. The confidence intervals presented in Table 1 were adopted as the cutoff thresholds.

2.2. Determination of Viral Protein Amino Acid Homogeneity and Identification of Region Type (Conserved, Variable) Within Sequence Alignments

2.2.1. Group-Specific Antigen

For amino acid sequence alignments of the group-specific antigen (Gag) polyprotein, an analysis was performed to examine the dynamics of conservation, and averaged S-index panoramas were obtained (Figure 2), along with generalized statistics describing the conservation of this region (Table 2).
Gag was shown to be a relatively conserved region with distinct variable segments. Clearly defined conserved regions can be observed in amino acid sequences, and these regions largely coincide with one another. Both conserved and variable regions were identified within the alignments, and the summarized results of this analysis for the protein alignments are presented in Figure 3 and Table 3.
Across alignments of different HIV-1 subtypes, 23–30 regions with Sm values ranging from 0.9856 to 1.0000 were identified. Alignment of all sequences made it possible to detect 26 conserved regions with an Sm value of 0.9856. In addition, the alignments were analyzed using an algorithm based on the calculation of local maxima. The results of this supplementary analysis, together with filtering of the identified conserved regions whose consensus sequences consisted predominantly of gaps, allowed the identification of 25 principal conserved regions of Gag, which are presented in Table 4.

2.2.2. Polymerase

The polymerase (Pol) polyprotein as a whole demonstrates a uniformly high level of conservation throughout its entire length. The averaged panorama for the aggregate of all subtypes does not reveal clearly pronounced peaks of conservation or variability, indicating relatively homogeneous evolutionary preservation of this protein (Figure 4, Table 5). However, alignments of individual subtypes do exhibit distinctly defined regions with increased conservation, while subtypes K and G display anomalously low levels of amino acid homogeneity. For subtype K, this may be explained by the limited number of sequences included in the alignment. However, the result obtained for subtype G remains anomalous and is likely attributable to an algorithmic artifact.
Although the Pol region proved to be relatively uniformly conserved along its entire length, cluster analysis enabled the identification of 44 to 54 conserved regions depending on subtype (excluding subtype K), with Sm values ranging from 0.92775 to 1.0000 (Figure 5, Table 6).
Supplementary Analysis using the local maxima algorithm, followed by final filtering, enabled the identification of 49 major conserved regions within Pol, which are presented in Table 7.

2.2.3. Envelope

In contrast to the more conserved Gag and Pol polyproteins, envelope (Env) exhibits the highest level of variability. The averaged S-index panoramas obtained from amino acid sequence alignments (Figure 6) and the low conservation index values (Table 8) indicate pronounced heterogeneity in this region.
For systematic identification of conserved elements, a clustering algorithm was first applied, which detected between 31 and 41 such regions depending on subtype. These identified regions exhibited high internal conservation (Sm ranging from 0.919 to 0.988; Figure 7, Table 9).
In the second stage, using a local maxima algorithm followed by filtering, the obtained list was refined and reduced to 28 major conserved regions common to all analyzed Env sequences. The final data, including the coordinates of these regions, are presented in Table 10.

2.2.4. Viral Infectivity Factor

Initial analysis of alignment containing sequences of HIV viral infectivity factor (Vif) from all subtypes revealed relatively low S-index values (Figure 8, Table 11), indicating the heterogeneity of this protein across the viral population as a whole. However, analysis of subtype-stratified alignments revealed a different pattern: within each subtype, Vif exhibits a high degree of conservation. This suggests the presence of significant inter-subtype differences alongside tight stabilization of the protein within distinct viral evolutionary lineages.
In the next stage, using a clustering algorithm, between 5 and 10 conserved regions were identified in the individual subtype alignments, characterized by high S-index values (0.978–1.000; Figure 9, Table 12).
The final stage of analysis, employing a local maxima algorithm, enabled the consolidation and refinement of these data, identifying six major conserved regions common to all investigated HIV-1 subtypes. Detailed information on these regions, including their coordinates and Sm-index values, is provided in Table 13.

2.2.5. Viral Protein R

Averaged S-index panoramas (Figure 10) for HIV viral protein R (Vpr) alignments demonstrated uniformly high values of the homogeneity index across most of the protein length. The summarized statistical data presented in Table 14 confirm this pattern: Sm-index values for individual subtypes range from 0.902 to 0.944, indicating a relatively high degree of overall Vpr conservation.
To identify localized regions with the highest degree of conservation, a clustering algorithm was applied. Analysis of subtype-stratified alignments enabled the identification of 3 to 6 highly conserved regions, depending on subtype. These regions were characterized by high Sm values ranging from 0.984 to 1.000 (Figure 11, Table 15), indicating minimal variability at the positions they encompass.
In the final stage, a local maxima algorithm was used to refine the boundaries of the conserved regions and eliminate redundant fragments. This approach enabled the consolidation and filtering of data obtained for different subtypes, leading to the identification of four major conserved regions common to all investigated HIV-1 subtypes. Detailed information on each region, including coordinates relative to the HXB2 reference sequence, as well as the alignment, length, and Sm values, is provided in Table 16.

2.2.6. Regulator of Virion Expression

In the first stage of analysis, the averaged S-index panoramas for regulator of virion expression (Rev) sequences alignments revealed a non-uniform conservation profile: against a background of overall variability, several distinct regions with elevated homogeneity index values were clearly distinguishable (Figure 12, Table 17). The Sm value for the alignment of all subtypes was 0.825, while for individual subtypes this value ranged from 0.836 (subtype D) to 0.930 (subtype H), indicating substantial inter-subtype differences in the degree of Rev conservation.
To identify localized regions with the highest degree of conservation, a clustering algorithm was applied at the second stage (Figure 13). Analysis of subtype-stratified alignments enabled the identification of 1 to 5 highly conserved regions, depending on subtype. These regions were characterized by high Sm values ranging from 0.9600 to 1.000 (Table 18), indicating minimal variability at the positions they encompass. The smallest number of conserved regions (1–2) was detected for subtypes A2 and J, which is likely associated with the most uniform sequence conservation, whereas for most other subtypes, their number ranged from 4 to 5.
In the final stage, a local maxima algorithm was used to refine the boundaries of the conserved regions and eliminate fragmented segments. This approach enabled the consolidation and filtering of data obtained for different subtypes, leading to the identification of four major conserved regions common to all investigated HIV-1 subtypes. Detailed information on each region, including coordinates relative to the HXB2 reference sequence, as well as the alignment, length, and Sm values, is provided in Table 19.

2.2.7. Trans-Activator of Transcription

The averaged S-index panorama for alignments of trans-activator of transcription (Tat) sequences obtained at the first stage of analysis revealed a pronounced non-uniformity in the conservation profile (Figure 14). The N-terminal region of the protein is characterized by relatively high homogeneity index values. The remainder of the sequence exhibits significantly higher variability. The summarized statistical data presented in Table 20 confirm this overall pattern: the Sm value for the alignment of all subtypes was 0.8205. When analyzing individual subtypes, this value ranged from 0.818 (subtype D) to 0.905 (subtypes A6, H), indicating substantial differences in Tat between sequences of different subtypes.
To identify localized regions with the highest degree of conservation, a clustering algorithm was applied at the second stage. The analysis enabled the identification of 2 to 5 highly conserved regions, depending on subtype (Figure 15). These regions were characterized by high Sm values ranging from 0.963 to 1.000 (Table 21), indicating minimal variability at the positions they encompass. The smallest number of conserved regions (2) was detected for subtypes A6, C, F2, G, and K. For subtype F1, it reached 5.
In the final stage of analysis, two major conserved regions common to all investigated HIV-1 subtypes were identified. Detailed information on both regions, including coordinates relative to the HXB2 reference sequence, as well as the alignment, length, and Sm values, is provided in Table 22.

2.2.8. Negative Regulatory Factor

Sequence analysis revealed that negative regulatory factor (Nef) is characterized by a pronounced non-uniform distribution of conserved and variable regions (Figure 16). The averaged S-index panoramas display clearly defined peaks of conservation in the N-terminal region and several internal domains, whereas a substantial portion of the sequence exhibits increased variability. According to the summarized statistical data (Table 23), the Sm-index value for the alignment of all subtypes was 0.866. When examining individual subtypes, this value varies over a wide range: from 0.847 (subtype K) to 0.926 (subtype A6).
This protein proved to be relatively highly variable; however, its N-terminus demonstrated high conservation within subtypes. Depending on the subtype, between 7 and 11 highly conserved regions were identified (Figure 17, Table 24). The identified regions exhibit high Sm values ranging from 0.966 to 0.994, indicating extremely low variability at the amino acid positions they encompass. The highest number of conserved regions (10–11) was characteristic of subtypes A6, B, F1, and J. For subtypes A2, C, F2, G, H, and K, their number ranged from 7 to 8.
In the final stage, using a local maxima algorithm followed by the filtering of data obtained for different subtypes, six major conserved regions common to all investigated HIV-1 subtypes were identified. General information on these regions, including coordinates relative to the HXB2 reference sequence, as well as the alignment, length, and Sm values, is provided in Table 25.

2.2.9. Viral Protein U

According to data obtained in the first stage of analysis of viral protein U (Vpu) sequences and the averaged S-index diagram (Figure 18, Table 26), Vpu exhibits the lowest conservation values among all HIV-1 proteins investigated. The Sm value for the alignment of all subtypes combined was 0.784, indicating high variability of this protein across the viral population. Nevertheless, it can be noted that the central regions of Vpu contain a relatively conserved area, which is particularly pronounced within individual subtypes. When analyzing specific subtypes, this value varies over a wide range: from 0.773 (subtype K) to 0.905 (subtype B). The highest homogeneity index values are observed for subtypes B (0.905) and A6 (0.886). Subtypes D, G, J, and K are characterized by the lowest values (0.773–0.803), reflecting substantial inter-subtype differences in the degree of Vpu conservation.
The applied clustering algorithm enabled the identification of 1 to 4 highly conserved regions, depending on subtype (Figure 19, Table 27). The identified regions were characterized by Sm values ranging from 0.927 to 1.000. The highest number of conserved regions was detected for subtypes A1 and K. For subtypes A2, F2, and for all subtypes combined (all_sequences), only one conserved region was found. The boundaries of this region were confirmed at the final stage of analysis. It is located within alignment coordinates 57–66 and spans 10 amino acid positions. The Sm value for this region is 0.954, indicating a high degree of conservation in this area, despite the overall variability of the Vpu protein as a whole.

3. Discussion

3.1. General Characteristics of the Obtained Alignments

The performed analysis of HIV-1 protein sequences provided a comprehensive overview of the distribution of conserved and variable regions across the major viral proteins. The applied computational approach, based on normalized Shannon entropy combined with clustering and local maxima algorithms, enabled quantitative assessment of sequence homogeneity and systematic identification of highly conserved regions potentially associated with critical viral functions.
The analyzed datasets differed substantially in size depending on subtype and genomic region. The largest number of sequences was available for globally prevalent subtypes, whereas several subtypes were represented by comparatively limited numbers of full-length genomes. This factor should be considered when interpreting the results, particularly for low-representation subtypes, in which the observed conservation may partially reflect limited sampling. Nevertheless, inclusion of these subtypes allowed a broader overview of HIV-1 diversity within Group M.
The use of a unified statistical framework based on entropy analysis enabled direct comparison of conservation profiles across proteins with distinct evolutionary characteristics. Although this approach does not fully account for differences in selective pressures between proteins, the identified conserved regions demonstrated strong correspondence with functionally important domains previously described in structural and biochemical studies, supporting the biological relevance of the obtained results.
Although the present study is purely computational and does not include experimental validation, the identified conserved regions may be relevant for future research aimed at HIV control and potential eradication strategies. Long-term viral persistence remains one of the principal barriers to HIV cure efforts, and highly conserved viral regions represent attractive candidates for approaches intended to minimize immune escape and resistance-associated variability. In particular, conserved elements identified within Gag, Pol, and structurally constrained regions of Env may support downstream development of broad-reactive immunogens, antiviral targets, or molecular diagnostic strategies intended to maintain effectiveness across genetically diverse HIV-1 variants.
In the following sections, an attempt is made to relate the identified conserved regions to functional and structural elements of HIV-1 proteins previously described in the literature, including catalytic centers, intermolecular interaction interfaces, membrane-associated domains, and other regions subject to strong evolutionary constraints.

3.2. Characterization of the Conserved Amino Acid Regions Identified

3.2.1. Group-Specific Antigen

Gag (pr55) is synthesized in the cytoplasm as a multidomain protein with a molecular mass of 55 kDa. Gag consists of four major structural domains (from N- to C-terminus): MA (matrix), CA, NC (nucleocapsid), and p6. There are also two small peptides, spacer peptide 1 (SP1) and spacer peptide 2 (SP2), which link the CA-NC and NC-p6 domains, respectively. Historically, the understanding of the structure, maturation, and function of Gag has undergone changes. Consequently, the original annotations of the HXB2 reference sequence are partially outdated and do not fully align with the current understanding of this polyprotein’s organization. In the following text, we will endeavor to rely on contemporary insights into the arrangement of regions within this and subsequent proteins, guided by the coordinates of the obtained alignment and accompanying them with HXB2 coordinates in parentheses.
Synthesis of the analysis performed for different subtypes allowed the identification of 24 highly conserved regions within the Gag protein sequence (Table 4). These conserved regions are represented in all final protein products resulting from Gag processing and, as expected, largely correspond to known functional areas of these products.
  • Matrix Protein
Within Gag, the matrix protein is represented by the MA domain, which in our alignment corresponds to coordinates 1–143 (HXB2: 1–131). Upon final maturation, it becomes the p17 protein. This structural protein typically consists of 132 amino acids and is encoded by the Gag gene, playing an important role in most stages of the HIV-1 life cycle. The protein is partially globular, comprising four helices that form a dense central domain, capped by a β-sheet containing basic amino acids.
Region CR_1 corresponds to the beginning of the peptide and is located at the very start of the N-terminal domain, which, due to subsequent active myristoylation, performs the function of attaching the capsid to the lipoprotein envelope of the future virion. This same connection requires the highly basic region (HBR), within which the identified conserved region CR_2 is located. This region is rich in arginine and lysine residues and electrostatically interacts with negatively charged phospholipids located in the inner leaflet of the plasma membrane [16].
Region CR_3 is located entirely within α-helix A and largely corresponds to its boundaries. A 1997 study described the role of this helix in the interaction of p17 with the cell membrane and noted that virions with mutations in this region lacked infectious activity [17]. However, the full functional significance of this highly conserved region has not been completely elucidated. Regions CR_4 and CR_5 (partially) are located within α-helix D, whose role in p17 function has not been revealed. Nevertheless, its relatively high conservation, as in the case of helix A, may be explained by the significant role of helices in maintaining the correct structure of p17. It is known that due to the presence of specific structural motifs defined as “coiled-coil” sequences, p17 exhibits a high propensity for misfolding and aggregation [18]. This HIV-1 matrix protein is known to possess self-interaction properties, as it forms trimers and hexamers [19,20,21]. The main identified sites for aggregation are the C-terminal region (capable of interacting with the same region of other p17 molecules [22]) and the central part of the matrix (amino acids 42–47) [23], which partially coincides with the identified conserved regions CR_3–4. It is likely that the high conservation of these regions is necessary to prevent excessive aggregation of p17 molecules and subsequent disruptions in capsid assembly.
  • Capsid Protein
The HIV capsid protein is represented within the Pr55 polyprotein by the CA domain, which in our alignment corresponds to coordinates 144–374 (HXB2: 132–362) and, upon final maturation, forms the p24 protein. The CA domain partially includes region CR_5 and entirely encompasses 12 identified conserved regions, CR_6–17. Among these, five are located in the N-terminal capsid domain (CA-NTD, HXB2: 142–249), one is partially within the linker region (HXB2: 250–274) and partially within the C-terminal capsid domain (CA-CTD, HXB2: 275–348), and five are entirely within the CA-CTD. Interactions between CA domains within Gag play a central role in forming the immature lattice observed in the budding particle. The assembly of CA proteins into a fullerene structure, consisting of approximately 250 CA hexamers and 12 CA pentamers, forms the viral capsid. It serves as a protective shell for the viral genomic RNA inside the mature particle and during entry into the cell cytoplasm [24,25].
The CA-NTD consists of seven α-helices and a cyclophilin A (CypA) binding loop, with an N-terminus that is unstructured in Gag, but folds into a β-hairpin in mature CA [26,27]. The CA-NTD forms a stable inner core of the hexamer, with the β-hairpin essentially lining the interior of its center, the R18 pore. This pore is a size-selective, positively charged channel that facilitates the diffusion of nucleotides into the capsid core while simultaneously preventing access to nucleases, host cell restriction factors, and sensors [28].
The CA-CTD consists of a helix, the major homology region (MHR) loop, which is highly conserved and essential for viral replication, and four α-helices. The C-terminus of the CA-CTD undergoes an important structural rearrangement during viral maturation [29,30]. The C-terminal domains form a flexible outer ring of hexamers, playing a crucial role in forming interfaces for hexamer assembly and for contacts with other hexamers [31].
In addition to the R18 pore, there is a critical interprotomer pocket that forms between hexameric subunits and is generated by the interaction of the NTD of one subunit with the CTD of another. This NTD-CTD interprotomer interface is present in the mature HIV-1 capsid and is crucial for proper capsid assembly, stability, and recognition of host cell factors [32]. Evidence suggests that this pocket binds host cell factors involved in nuclear import, indicating that intact CA oligomers are imported into the nucleus during nuclear entry of the preintegration complex (PIC) [33]. Amino acid residues comprising CR_17, which captures the end of the CA domain within Gag (specifically CTD helix H10), play an important role in interface formation. Additionally, CTD helix H9, which coincides with CR_16, plays a significant role in forming CTD-CTD interfaces within the hexamer [34].
Additionally, CR_15–17 may participate in interactions with human lysyl-tRNA synthetase for the selective packaging of tRNA-Lys, which plays a key role in initiating HIV reverse transcription. For the remaining identified conserved regions, no precise correspondence to described functional domains could be established. However, based on the observed relationships, it can be inferred that they are also critical for the functioning of the R18 pore (CR_5–6) and for the formation of NTD-CTD interfaces (CR_7–15) [35].
  • Nucleocapsid Protein
The Gag nucleocapsid domain (align. coordinates 390–445, HXB2: 376–430) and its corresponding mature protein form (NCp7) are basic polypeptides of 55 amino acids. They are characterized by the presence of two conserved structures, zinc fingers (ZF), containing the invariant CCHC motif. The ZFs are separated by a short linker and flanked by small domains rich in basic residues.
Among the identified conserved regions, four are entirely within NC (CR_19–22), and one is partially included (CR_23). Specifically, CR_19 corresponds to the first zinc finger, including the amino acid preceding it. CR_20 encompasses the end of the zinc finger and part of the linker sequence. CR_21–22 are entirely within the second zinc finger, including the amino acid preceding it. Finally, CR_23 comprises the last six amino acids of the NC domain [36]. The zinc fingers and the hydrophobic plateau formed upon ZF folding bind to the backbone and nitrogenous bases of nucleic acids. Moreover, the high flexibility of the protein enables it to bind to a wide variety of nucleic acid sequences [37,38,39,40]. The importance of preserving NC structure explains the high sequence conservation across HIV-1 subtypes and isolates from treated patients, as well as the low probability of detecting mutations [41,42].
  • Gag p6
Gag p6 is a multifunctional domain that plays a key role in the late phase of the viral cycle. It typically comprises approximately 52 amino acids located at the C-terminus of the Gag polyprotein, corresponding to alignment coordinates 464–543 (HXB2: 447–498). This domain contains the identified conserved region CR_25, situated at the C-terminal end of the domain. This conserved region participates in Env binding and regulates the packaging of cleaved Pol proteins [43,44,45]. Functional domains are also known to exist at the N-terminus of p6. However, their function is primarily associated with amino acid motifs rather than strictly conserved sequences [16,46,47,48]. The role of the central region of p6 remains unclear and is likely non-essential for p6 function insofar as its polymorphism does not impair either viral infectivity or replication kinetics [49].
  • Spacer Regions SP1 and SP
The Gag spacer regions SP1 (align. coordinates 375–389, HXB2: 363–375) and SP2 (align. coordinates 446–463, HXB2: 431–446) are regulators of Gag assembly and maturation. Within SP1 (a conserved region), CR_18 was identified, which is presumably associated with the functional role of this domain. Mutations in the first seven residues of SP1 proximal to the CA domain are known to disrupt viral assembly and lead to the formation of tubular structures containing unprocessed Gag at the plasma membrane [50,51]. A single conserved region, CR_24, was also identified within the SP2 region. Although the role of this region has not been fully elucidated, the identified conserved region is most likely associated with Gag processing and proper targeting of the HIV protease.

3.2.2. Polymerase

Whereas the structural proteins of HIV-1 are initially translated as part of the Gag polyprotein, the enzymes that enable the virus to replicate and spread are synthesized as part of the Gag-Pol precursor polyproteins. Both polyproteins are cleaved by the viral protease (PR), which is itself synthesized as part of Gag-Pol during maturation [52]. Pol is processed to yield the viral enzymes PR, reverse transcriptase (RT), and integrase (IN). Gag-Pol is synthesized via translational readthrough; different retroviruses employ various readthrough mechanisms. HIV-1 utilizes a ribosomal frameshifting mechanism. The efficiency of this frameshift results in a Gag:Gag-Pol ratio of approximately 20:1 [53,54,55,56], making Gag-Pol a relatively minor component of the virion (~100 copies per virion) [57]. Because the enzymes comprising Pol are key participants in the viral life cycle, they are also targets for antiretroviral therapy. As the Pol polyprotein is the precursor to these viral enzymes, it represents an extremely conserved region, and the conducted analysis enabled the identification of 49 highly conserved regions associated with the functional domains of the mature enzymes (Table 7). The coordinates used correspond to alignments of the full polyprotein rather than its individual cleavage products; the same applies to the HXB2 reference sequence coordinates. However, the subsequent description of conserved regions will be structured according to the Pol cleavage products.
  • Retroviral Aspartyl Protease
HIV-1 protease is a member of the aspartic protease family due to the presence of the conserved catalytic Asp-Thr/Ser-Gly triad [58]. The mature protease is catalytically active as a dimer composed of two 99-residue subunits, with each subunit contributing one copy of the catalytic triplet. PR recognizes specific amino acid sequences at various cleavage sites within the Gag and Gag-Pol polyproteins and hydrolyzes the peptide bonds to release the individual structural proteins and enzymes. These cleavage sites must be hydrolyzed in the correct sequential order to generate an infectious virus [59,60,61].
The identified Pol_CR_1 contains the beginning of the Pol domain sequence corresponding to the PR, including the 4 amino acids preceding it and the first 5 residues of the enzyme. The subsequent conserved regions Pol_CR_2–5 are located within PR. Pol_CR_1 and Pol_CR_5 are part of the interface between monomers of the protease dimer, which upon folding is formed by the N- and C-termini of the monomer (alignment coordinates: 86–90 and 181–184). Pol_CR_2 is associated with hinge points in the fulcrum element of the enzyme and also encompasses the conserved catalytic DTG triplet. Pol_CR_3 corresponds to the active site flaps, which modulate substrate entry into the active site cavity and consist of two β-hairpin loops (alignment coordinates: 128–141). Pol_CR_4 comprises the sequence lining the interior of the enzyme active site, situated between the cantilever and an α-helix [62].
According to genotype and phenotype data available in HIVdb [63,64], the major protease inhibitor resistance mutations are located within the Pol_CR_3 region. This aligns with the mechanism of action of these inhibitors: they typically enter the active site and block it, whereas altered active site flaps prevent this access, leaving the enzyme active site free. However, these same changes negatively impact enzyme activity. Consequently, such mutations are maintained by selection only under the pressure of antiretroviral drugs.
  • Reverse Transcriptase
The RT enzyme is represented in Pol by several domains: the retroviral reverse transcriptase (RT_Rtv), the RT thumb domain (RVT_thumb), the reverse transcriptase connection domain (RVT_connect), and the retroviral ribonuclease H (RNase_H). The RT_Rtv, RVT_thumb, and RVT_connect domains together constitute the polymerase proper and contain the fingers, palm, thumb, and connection motifs of reverse transcriptase [65,66,67]. The RNase H domain facilitates degradation of RNA from the DNA-RNA duplex during reverse transcription. Other functions of RNase H include the removal of tRNA3Lys and the removal of the polypurine tract (PPT), which serve as primers for negative-strand DNA synthesis and positive-strand DNA synthesis, respectively [68,69,70].
Conserved regions Pol_CR_6 through Pol_CR_31 are located within reverse transcriptase, with Pol_CR_6–15 corresponding to the RT_Rtv domain; Pol_CR_16–18 to the RVT_thumb domain; Pol_CR_20–24 to the RVT_connect domain; and Pol_CR_25–31 to the RNase_H domain. Three major studies [71,72,73] provide detailed reviews of conserved patterns in RT. The conserved regions identified in those studies generally coincide with those identified in our current analysis. However, some discrepancies exist, mainly consisting of certain large regions in our study being subdivided into several smaller ones due to separation by less conserved areas. Since the nature of conservation in these regions has been thoroughly described in those works, we will not dwell in detail on the coinciding regions and will instead focus only on those that differ.
Regions Pol_CR_7–9 are located within the RT_Rtv domain and were not annotated as conserved in a study [71]. Nevertheless, the amino acids comprising these regions were noted as highly conserved but, for some reason, were not delineated by the authors as distinct conserved areas. In a different study [74], however, these regions are identified as conserved regions of RT. All three conserved regions are part of DNA-binding elements and, evidently, play a key role in this process, which accounts for their high conservation.
  • Integrase
Integrase (IN) is a crucial viral enzyme consisting of 288 amino acids and encoded by the 3′-end of the HIV polymerase gene. Integrase catalyzes the integration of newly synthesized double-stranded DNA into the host genomic DNA. It also plays a role in stabilizing the preintegration complex (PIC), which consists of the 3′-processed viral genome and one or more cellular cofactors involved in PIC nuclear import [75]. HIV IN comprises three main domains: the integrase zinc binding domain (IN_Zn); the integrase core domain (retroviral-like integrase, rve); and the integrase DNA binding domain (IN_DBD_C). The rve domain also contains the transposase InsO and inactivated derivatives (Tra5) region, which exhibits transposase activity.
A total of 18 conserved regions were identified within IN: Pol_CR_32–49. Regions Pol_CR_33 and Pol_CR_34 belong to the IN_Zn domain and are part of the zinc fingers, indicating conservation not only of positions corresponding to the IN_Zn motif, but also of the overall zinc finger structure. Regions Pol_CR_35–45 correspond to the integrase core domain, with Pol_CR_39–45 residing within the region possessing transposase activity. A study [76] described several of the identified conserved regions within the core domain and DBD_C domain, specifically: Pol_CR_36, 37, 40, 41, 47, and 48. However, the remaining conserved regions also play important roles in IN function. For instance, region Pol_CR_36 performs a critical structural role, linking the core domain and the C-terminal DBD domain within helix α6. Regions Pol_CR_42–45 are part of β-sheets 1–5, which are essential for proper folding of the core domain and for correct communication between the catalytic core and the DNA-binding domain.

3.2.3. Envelope

HIV-1 Env (gp160) consists of two subunits, gp120 and gp41, along with a spacer peptide that is cleaved off during proteolysis. The gp120/gp41 complex is presented on the virion surface as trimers. Env directs the fusion of viral and cellular membranes to initiate infection of a susceptible cell [77]. Conformational changes accompany the binding of the native Env trimer to the receptor (CD4) and coreceptor (e.g., CCR5 or CXCR4), leading to a cascade of refolding events in gp41. Subsequent folding of the C-terminal region of gp41 into a hairpin conformation creates a post-fusion six-helix bundle [78,79], which brings the viral and cellular membranes together, resulting in fusion and viral entry.
  • gp120
Of the twenty-eight highly conserved regions identified within the Env polyprotein (Table 10), sixteen (Env_CR_1–16) correspond to the gp120 subunit. This protein is well known for the high variability of its epitopes; nevertheless, regions exhibiting conservation are also quite distinctly delineated [80]. Among these, the first three are located in the N-terminal region of the protein and correspond primarily to hydrophobic areas of gp120 with a not yet fully understood functional role. Env_CR_4 coincides with the β1 strand, preceding the α1 helix which contains the conserved region Env_CR_5. The conserved region Env_CR_6 is situated in the area corresponding to the V1/V2 loop and represents a region of high conservation within this relatively variable loop. The next conserved region (Env_CR_7) belongs to the β3 strand, located immediately after the V1/V2 loop. Conserved regions Env_CR_8–9 are found in the β4 and β5 strands, respectively. Following these strands lies the variable loop A region. Conserved region Env_CR_10 corresponds to the β7 and β8 strands. Region Env_CR_11 encompasses the B loop and the β9 strand. Further on, between regions Env_CR_11 and Env_CR_12, there is an extended variable region, and the twelfth conserved region includes the β16 and β17 strands. Subsequently, region Env_CR_13 corresponds to the β19 and β20 strands, region Env_CR_14 to the β21 strand, and Env_CR_15 pertains to the final helix, α5. Env_CR_15 corresponds to a hydrophilic region at the end of gp120, which constitutes the cleavage site between gp120 and gp41.
A similar structure of conserved and variable regions has been described [81]. Clearly, the conserved areas of gp120 are, first and foremost, important structural elements required for the formation of correct monomers and subsequently for gp120 trimers. Partially conserved regions are also found in the interfaces involved in host receptor binding. However, it is also known that key areas within the structure of such interfaces are considerably variable.
  • gp41
The remaining conserved regions (Env_CR_17–28) correspond to gp41 sequence. Within these, Env_CR_17–22 belong to the N-terminal part of the protein, encompassing regions responsible for membrane fusion. The subsequent regions correspond to the C-terminus, which includes the transmembrane and intracellular domains of gp41. The conserved region Env_CR_17 is located within the N-terminal fusion peptide, which is responsible for merging the viral membrane with the host membrane. The following regions, Env_CR_18–22, are part of the extended heptad repeat region 1, which precedes the cysteine loop region containing the conserved regions Env_CR_23 and Env_CR_24. Subsequently, after a long variable region, Env_CR_25 begins, corresponding to the membrane-proximal external region. The conserved region Env_CR_26 is situated within the transmembrane region, and regions Env_CR_27–28 are located within the cytoplasmic domain. A similar conservation pattern for gp41 has been reported [82]. The authors emphasize the high conservation of the gp41 ectodomains and associate this with the high functional relevance of this region.

3.2.4. Viral Infectivity Factor

Vif is a small, intrinsically disordered protein with a molecular mass of approximately 23 kDa, consisting of roughly 192–216 amino acids depending on HIV-1 subtype. Vif acts as a substrate receptor for A3 proteins within the Cullin–RING E3 ubiquitin ligase complex. E3 Cullin–RING complexes represent the largest subfamily of RING domain-containing ligases. These complexes target cellular proteins for ubiquitin-mediated degradation via the 26S proteasome [83]. Two functionally distinct domains of Vif mediate its role as a structural hub for the complex, the N-terminal A3 substrate-binding domain (referred to as the α/β domain) and the C-terminal adapter domain (α domain), which interacts with the E3 ligase machinery. The Vif α/β domain forms a central five-stranded antiparallel β-sheet (strands β2–β6), flanked on its convex side by three helices (α1, α2, α5). The α domain contains two α-helices and includes a BC-box motif, as well as the conserved HCCH motif. The HCCH motif contains His and Cys residues that tetrahedrally coordinate a single Zn2+ ion with high affinity [84,85,86].
Analysis of the Vif sequence alignments enabled the identification of six highly conserved regions, which are presented in Table 13. The first conserved region, located at HXB2 coordinates 1–16, resides in the N-terminal domain of the protein and corresponds to one of the A3F protein binding sites. Vif_CR_6 also participates in forming this site upon Vif folding. The A3G protein binding site is primarily formed by Vif_CR_2 and Vif_CR_5; however, the region corresponding to HXB2 coordinates 40–45 is also known to participate in this interaction. Nevertheless, this region was not classified as conserved in our analysis, as its conservation was observed only within individual subtypes and diminishes significantly when comparing sequences from different subtypes. Additionally, Vif_CR_3 contributes to the formation of the A3 protein-binding domain. The fourth conserved region constitutes the Elongin C binding site, a component of the E3 ubiquitin ligase complex.
Beyond the aforementioned HXB2 40–45 region, the Vif protein contains numerous other areas that exhibit high conservation within subtypes but display considerable inter-subtype variability. This may indicate active adaptation of Vif during viral evolution and could contribute to differential infectivity among HIV-1 subtypes.

3.2.5. Viral Protein R

The product of the accessory gene vpr encodes a 14 kDa protein consisting of 96 amino acid residues; it is expressed during the late stages of viral replication from an open reading frame located in the central region of the viral genome. This protein is highly conserved among primate lentiviruses, including HIV-1, HIV-2, and simian immunodeficiency virus, supporting the hypothesis that it plays an important role in the viral life cycle [87]. Although our understanding of how Vpr contributes to HIV-1 replication and pathogenesis remains incomplete, one well-characterized mechanism involves the recruitment of the CRL4 E3 ubiquitin ligase and its substrate receptor DCAF1 (CRL4 DCAF1) to deplete cellular proteins that directly or indirectly target viral components [88]. Nevertheless, other potential functions of Vpr have been proposed, including facilitation of nuclear pore transport, participation in preintegration complex formation, and induction of cell cycle arrest at the G2 phase, among others.
The functional domains of the protein remain poorly characterized. Therefore, it is not possible to definitively associate the identified conserved regions with specific Vpr functions. However, it can be noted that two regions, Vpr_CR_2 and Vpr_CR_3, are located within the hydrophobic core, which consists of several sterically converging hydrophobic residues from helices α1 and α3. Both regions are situated at the boundary between organized helices and disordered regions and are likely important for stabilizing the complex structure. Vpr_CR_1 and Vpr_CR_4 are located in disordered regions at the N- and C-termini of Vpr, respectively, yet their functional roles remain unclear.

3.2.6. Regulator of Virion Expression

Rev is a protein consisting of 116 amino acids (~13 kDa) on average. It is encoded by two exons that overlap with other genes: tat and env [89,90]. Rev plays a crucial role in the viral life cycle by coordinating the nuclear export of unspliced and incompletely spliced viral mRNAs [91]. Rev, expressed from fully spliced viral RNA, translocates to the nucleus, binds to the Rev response element (RRE) through oligomerization, and induces the export of incompletely spliced mRNAs to the cytoplasm [89,91,92]. Rev is also involved in RNA splicing, stability, and translation. However, its impact on these processes remains poorly understood [91].
The four most conserved regions we identified (Table 26) are located within the N-terminal portion of Rev and are situated in the N-terminal domain itself (Rev_CR_1), as well as within the oligomerization domain, the turn region (Rev_CR_2–3), and the arginine-rich motif (ARM) domain (Rev_CR_4). The ARM domain represents the most highly conserved region of the Rev protein [93]. This finding is consistent with the results of competitive deep mutational scanning, in which the majority of residues within the ARM domain exhibited strong selection [90]. Therefore, it can be inferred that the well-conserved structure of the ARM domain is essential for Rev function. The N-terminal region as a whole also demonstrated a high degree of conservation, which may be explained by its overlap with the functionally significant arginine-rich motif of the Tat protein [90,94].

3.2.7. Trans-Activator of Transcription

Tat is a short protein (averaging 101 residues) expressed during the early stages of infection, initially described as a transactivator of HIV-1 genes [95]. However, another intriguing property of Tat is its extracellular role and high level of secretion from HIV-1-infected cells, suggesting that it plays an important part in HIV-1′s ability to evade the immune response in infected patients [96]. Extracellular Tat penetrates the membrane of numerous uninfected bystander cells, such as cytotoxic T lymphocytes, to induce apoptosis [97].
Six major regions are distinguished within the Tat structure [98]. The conserved regions we identified (Table 27) correspond to region I, located at the N-terminus of the protein (Tat_CR_1), as well as to regions III and, predominantly, IV (Tat_CR_2). Other studies have highlighted the high functional significance of region IV. It is required for crossing the cell membrane and binding to the trans-activating response element, a stem-loop structure at the 5′-end of viral mRNA. The role of region I has not been fully elucidated. However, based on its involvement in folding the central β-sheet of the protein, this region primarily serves a structural function for Tat.

3.2.8. Negative Regulatory Factor

Nef is a myristoylated peripheral membrane protein of 23–35 kDa (~206 amino acid residues in most HIV-1 strains) expressed by primate lentiviruses. The full-length structure of Nef consists of six α-helices (α1–α6) and a β-sheet composed of five antiparallel β-strands (β1–β5). Nef structure is divided into four units: a flexible myristoylated membrane-anchoring region of variable length (residues 1–56); followed by the PxxP loop (residues 57–80); the core domain (residues 81–206, with Δ148–180); and a flexible C-terminal loop (residues 148–180). While the anchor and the flexible C-terminal loop are not conformationally constrained relative to the core domain, the PxxP loop is weakly associated with the core domain through hydrophobic interactions [99].
The identified region Nef_CR_1 is located at the very beginning of the N-domain responsible for anchoring Nef to the membrane. This highly conserved region and its myristoylation are known to be critically important for membrane localization of the protein. Within the proline-rich loop of the N-terminus, two conserved regions, Nef_CR_2–3, can be observed at the beginning and end of the loop. The proline-rich motif itself mediates interactions between Nef and signaling molecules, such as Hck and Vav, and is central to Nef’s ability to induce cellular activation, a function that may be necessary for maintaining viral replication in resting cells [100,101]. However, the identified conserved regions may also be required for protease cleavage at the loop termini during final Nef maturation [102]. The next highly conserved regions identified, Nef_CR_4 and Nef_CR_5, are part of the Nef core domain and are primarily associated with the SH3 binding site. The latter is necessary for interactions with Fyn and MHC class I molecules, and this interaction is apparently important for Nef function [103] Nef_CR_6 belongs to the C-terminal loop of the protein, which likely also participates in interactions with AP-1 and AP-2 molecules to influence MHC class I. However, the function of this loop is insufficiently characterized to definitively assess the role of the identified conserved region in its activity.

3.2.9. Viral Protein U

The HIV-1-specific Vpu protein is an integral class I membrane phosphoprotein consisting of 81 amino acids. Its amino acid sequence exhibits extensive diversity, which often increases as infection progresses. Key conserved functions of Vpu include downregulation of CD4, antagonism of tetherin (BST-2/CD317), and modulation of other cellular receptors [104,105,106]. Vpu is amphipathic in nature and consists of a hydrophobic N-terminal membrane anchor located adjacent to a polar C-terminal cytoplasmic domain.
Despite the presence of conserved intra-subtype regions within the transmembrane domain of Vpu, the areas exhibiting the highest conservation across all subtypes are located within the cytoplasmic domain of the protein. The highly conserved region identified during analysis is situated entirely within the second α-helix of the cytoplasmic domain. A distinctive feature of this region is a dileucine-like sorting motif. This motif functions as an intracellular trafficking signal. It influences protein localization and is required for efficient removal of tetherin from the cell surface, as well as for controlling the volume of virus-containing compartments in macrophages [104,105,106,107,108]. Mutations at positions G59 and E62 also impair the ability of Vpu to suppress tetherin-induced NF-κB signaling [104].

3.3. Study Limitations

The authors recognize that the present study has several limitations that should be considered when interpreting the results and planning future investigations. In this work, the analysis was focused on sequences belonging to HIV-1 Group M, which predominates in the global epidemiology of HIV infection. Groups O, N, and P were not included due to their extremely limited representation in public databases in the form of full-length genomes with sufficiently reliable annotation quality.
In addition, within Group M, the number of sequences available for different subtypes varied substantially. For subtypes A2, F2, J, and K, fewer than 15 full-length genomes met the inclusion criteria, which may have affected the statistical robustness of the analyses. Therefore, results obtained for low-representation subtypes should be interpreted with particular caution, and conclusions regarding pan-subtype conservation require validation using expanded datasets.
The present study attempted to identify conserved regions using a unified statistical threshold applied across all proteins regardless of differences in the biological and evolutionary pressures acting on them. Although many of the identified conserved regions corresponded to functionally important domains previously described in the literature, this approach may underestimate conservation in certain protein regions. Consequently, the results should be interpreted in conjunction with complementary studies addressing the structural, biochemical, and physicochemical properties of HIV-1 proteins. Similarly, the minimum conserved-region length threshold (5 amino acids) was selected empirically to prioritize extended conserved regions; therefore, shorter biologically significant motifs may have been overlooked.
Finally, this study is purely bioinformatic in nature and should be regarded as a framework for prioritization of hypothetical targets rather than direct translational validation. The identified conserved regions do not by themselves guarantee suitability for the development of diagnostic primers, vaccine immunogens, or therapeutic antibodies. Additional experimental studies are required to evaluate their accessibility, immunogenicity, and functional relevance in appropriate biological systems.

4. Materials and Methods

4.1. Materials

We downloaded random full-length HIV-1 genomic sequences from the NCBI Nucleotide database, selected according to several criteria: sequences were deposited up to the year 2025; HIV strains were obtained from different patients; and the annotations explicitly indicated viral subtype and the boundaries of genomic regions. In this study, we focused on sequences belonging to subtypes of group M, as this group is the most prevalent in the population; therefore, belonging to other groups was considered an exclusion criterion. Additionally, the number of sequences per viral subtype was limited to a maximum of 201. A total of 1119 sequences were included in the study (Table 28). An overview of the next steps in the analysis is provided in Figure 20.

4.2. Normalized Shannon Entropy

Prior to analysis, multiple sequence alignment of the investigated sequences was performed using the MUSCLE v5 algorithm [109]. Subsequently, amino acid frequencies were calculated for each column of the multiple sequence alignment, based on which the Shannon entropy (1) was computed. For amino acid sequences, all symbols represented in GenBank translations, as well as the gap symbol, were taken into account.
H = i = 1 q p i l o g 2   p i
  • H—Shannon index;
  • q—maximum possible number of distinct symbols;
  • p i —frequency of a given symbol in an alignment column.
Additionally, normalization was performed against the maximum possible entropy. The normalized entropy measure was calculated as follows (2).
H ( n o r m ) = 1 q p i l o g q   p i
  • H(norm)—normalized Shannon index;
  • q—maximum possible number of distinct symbols;
  • p i —frequency of a given symbol in an alignment column.
For the analysis, the amino acid residue homogeneity measure S was used (3).
S = 1 H ( n o r m )
  • S—amino acid residue homogeneity;
  • H(norm)—normalized Shannon index.
The next stage involved analysis of the resulting amino acid residue heterogeneity profile. For this purpose, the profile of S-values obtained for each position (the full S-profile) was used.

4.3. Ranking and Threshold Methods

The complete S-profile can be represented as a histogram (Figure 1), wherein each alignment column is assigned an S-index value. It is evident that the more heterogeneous the investigated alignment, the smaller the area of the figure bounded by the upper edge of the histogram bars. Since the area of each individual bar equals the S-value at that position, the total area corresponds to the sum of S-values across all alignment columns. Normalizing this value by the maximum possible area, which would equal one, yields the Sm value for the entire alignment. Based on this reasoning, we selected it as a metric reflecting the overall sequence similarity within the alignment.
Furthermore, it was hypothesized that Sm represents a meaningful measure of conservation for an alignment fragment or the entire alignment. To more distinctly identify highly conserved and highly variable positions, a confidence interval at a 95% significance level was determined for Sm, with its upper and lower bounds serving as the primary cutoff thresholds for detecting regions exhibiting differential conservation. For amino acid sequences, the threshold corresponding to the upper bound of the confidence interval determined collectively for all protein products was applied.

4.4. Conserved Region Detection

Two independent algorithms based on analysis of the S-index profile were used to identify conserved regions within the sequences. The Clustering Algorithm was applied as the primary method to identify extended conserved domains by grouping adjacent positions. All positions with an S-index value exceeding the upper bound of the 95% confidence interval were considered candidate conserved sites. Their indices in the ordered sequence were used to determine breaks as follows. Positions were separated into distinct regions where the gap between indices exceeded 1, allowing the delineation of natural conservation clusters while ignoring isolated non-conserved insertions. To eliminate random short clusters, only regions with a minimum length of 5 amino acid positions were retained for further analysis. This approach is effective for detecting extended conserved blocks with well-defined boundaries.
The Peaks Algorithm was selected as a complementary method, aimed at identifying local conservation maxima. Local maxima were identified in the S-index profile according to the following criteria: the peak height had to exceed the upper bound of the 95% confidence interval; and its relative prominence had to be at least 10% of the difference between the maximum S-index value and the threshold, ensuring statistical significance. The minimum distance between peaks was set to 10 positions. From each detected peak point, expansion was performed in both directions along the sequence as long as S-index values remained above the threshold and the distance from the peak did not exceed 10 positions, preventing excessive merging of adjacent peaks. As with the first method, the final regions were filtered by a minimum length of 5 positions. This approach enables the identification of compact, yet highly conserved, motifs that might be overlooked when analyzing average values over extended regions.
Conservative regions confirmed by both algorithms that included regions consisting predominantly of gaps (i.e., insertions occurring in a small number of sequences) were filtered out from the overall list.

4.5. Statistics

To assess the distribution of Sm within a selected alignment region, the bootstrap method was employed with a pseudosample size equal to half of the original dataset. To ensure accuracy in estimating the distribution, the number of bootstrap iterations was set to 1000. The 95% confidence interval of the distribution was selected as the confidence interval for Sm. This approach provided a statistically justified basis for determining the threshold values subsequently used to identify conserved and variable regions within the alignments.

4.6. Visualizations

For visualization of the obtained results, the averaged S-profile was used as it is more convenient for visual assessment of the data and is associated with the value over a sequence interval. Two methods were employed to display the averaged profile: a moving average heatmap and a line plot generated using cubic spline interpolation. This method allows approximation of the original data with a high degree of smoothness, particularly in regions where the S-index varies unevenly. Cubic spline interpolation is especially useful for noise reduction and identifying long-term trends in the data.
The use of splines is justified in bioinformatics research, where data are dense and preserving both global and local trends is crucial for interpretation. This method is recommended as a powerful tool for the analysis and visualization of genetic sequences due to its flexibility and mathematical precision. Images were generated using the Python v3.12 and libraries Plotly v3.0.1 and Matplotlib v 3.10.

5. Conclusions

In the present study, an algorithm for screening and analysis of amino acid sequence conservation was developed and validated. This approach is based on calculation of the normalized Shannon index, followed by the application of two complementary methods: clustering and local maxima detection. The employed approach enabled not only the quantitative assessment of sequence homogeneity, but also the identification of localized regions exhibiting extreme conservation values, including compact motifs that might have been overlooked when using traditional sliding window analytical methods. The application of the bootstrap method for confidence interval calculation provided statistical justification for threshold values. The use of cubic spline interpolation for visualization effectively reduced noise impact and revealed long-term trends in the distribution of conserved regions.
Application of the algorithm to HIV-1 protein alignments revealed a wide range of homogeneity index values: from 0.784 for the highly variable Vpu protein to 0.920 for the highly conserved Pol polyprotein. This reflects differences in functional constraints, evolutionary age of genes, and the intensity of selective pressure from the immune system and antiretroviral therapy. The highest conservation was observed for enzymatic and structural proteins (Pol, Gag, Vpr). Regulatory proteins (Rev, Tat) and proteins interacting with the immune system (Env, Vpu) demonstrated significantly higher variability, with conserved regions clearly localized within functionally significant domains.
As a result of the analysis, the major conserved regions were cataloged for all investigated proteins. The identified regions in most cases correspond to known functional domains such as enzyme catalytic centers, zinc fingers, sites of interaction with cellular factors, membrane-binding regions, and signaling motifs. This correspondence confirms the validity of the employed approach. Regions conserved across all subtypes are of particular value insofar as they represent promising targets for the development of therapeutic strategies and diagnostic systems effective in the context of global HIV-1 diversity. The obtained results can be used to optimize existing approaches and develop new strategies for HIV therapy and diagnostics.

Author Contributions

Conceptualization, A.N.S. and E.N.S.; methodology, A.N.S.; software, A.N.S.; validation, Y.V.O. and A.A.T.; formal analysis, A.N.S., E.N.S. and V.S.D.; investigation, V.S.D.; resources, A.A.T.; data curation, A.N.S.; writing—original draft preparation, A.N.S. and E.S.R.; writing—review and editing, Y.V.O. and E.S.R.; visualization, A.N.S.; supervision, A.A.T.; project administration, A.A.T.; funding acquisition, A.A.T. All authors have read and agreed to the published version of the manuscript.

Funding

Funding was provided by the research project «Molecular genetic characteristics of human immunodeficiency virus (HIV) in isolated or co-infection settings, considering human genetic factors, during and outside of antiretroviral therapy».

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
HIVhuman immunodeficiency virus
WHOWorld Health Organization
CRFcirculating recombinant form
ARTantiretroviral therapy
NGSnext-generation sequencing
UTRuntranslated region
GLSGag leader sequence
CIconfidence interval
Gag, pr55group-specific antigen
CRconserved region
Polpolymerase
Env, gp160envelope
Vifviral infectivity factor
Vprviral protein R
Revregulator of expression
Tattrans-activator of transcription
Nefnegative regulatory factor
Vpuviral protein U
CAcapsid domain
MAmatrix domain
NCnucleocapsid domain
SPspacer peptide
HBRhighly basic region
NTDN-terminal domain
CTDC-terminal domain
CypAcyclophilin A
MHRmajor homology region
PICpreintegration complex
ZFzinc finger
PRprotease
RTreverse transcriptase
INintegrase
RT_Rtvretroviral reverse transcriptase
RNase_Hretroviral ribonuclease H
RVT_thumbRT thumb domain
RVT_connectreverse transcriptase connection domain
PPTpolypurine tract
rveretroviral-like integrase
IN_DBD_Cintegrase DNA binding domain
IN_Znintegrase zinc binding domain
RRErev response element
ARMarginine-rich motif

References

  1. World Health Organization. HIV and AIDS. Available online: https://www.who.int/en/news-room/fact-sheets/detail/hiv-aids (accessed on 13 December 2025).
  2. Shchemelev, A.N.; Boumbaly, S.; Ostankova, Y.V.; Zueva, E.B.; Semenov, A.V.; Totolian, A.A. Prevalence of drug resistant HIV-1 forms in patients without any history of antiretroviral therapy in the Republic of Guinea. J. Med. Virol. 2022, 95, e28184. [Google Scholar] [CrossRef] [Scilit]
  3. Williams, A.; Menon, S.; Crowe, M.; Agarwal, N.; Biccler, J.; Bbosa, N.; Ssemwanga, D.; Adungo, F.; Moecklinghoff, C.; Macartney, M.; et al. Geographic and Population Distributions of Human Immunodeficiency Virus (HIV)–1 and HIV-2 Circulating Subtypes: A Systematic Literature Review and Meta-analysis (2010–2021). J. Infect. Dis. 2023, 228, 1583–1591. [Google Scholar] [CrossRef] [Scilit]
  4. Lebedev, A.; Kireev, D.; Kirichenko, A.; Mezhenskaya, E.; Antonova, A.; Bobkov, V.; Lapovok, I.; Shlykova, A.; Lopatukhin, A.; Shemshura, A.; et al. The Molecular Epidemiology of HIV-1 in Russia, 1987-2023: Subtypes, Transmission Networks and Phylogenetic Story. Pathogens 2025, 14, 738. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  5. Shchemelev, A.N.; Semenov, A.V.; Ostankova, Y.V.; Zueva, E.B.; Valutite, D.E.; Semenova, D.A.; Davydenko, V.S.; Totolian, A.A. Genetic diversity and drug resistance mutations of HIV-1 in Leningrad Region. J. Microbiol. Epidemiol. Immunobiol. 2022, 99, 28–37. [Google Scholar] [CrossRef] [Scilit]
  6. Shchemelev, A.N.; Semenov, A.V.; Ostankova, Y.V.; Naidenova, E.V.; Zueva, E.B.; Valutite, D.E.; Churina, M.A.; Virolainen, P.A.; Totolian, A.A. Genetic diversity of the human immunodeficiency virus (HIV-1) in the Kaliningrad region. Vopr. Virusol. 2022, 67, 310–321. [Google Scholar]
  7. Cuevas, J.M.; Geller, R.; Garijo, R.; López-Aldeguer, J.; Sanjuán, R. Extremely High Mutation Rate of HIV-1 In Vivo. PLoS Biol. 2015, 13, e1002251. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  8. Coffin, J.M. HIV population dynamics in vivo: Implications for genetic variation, pathogenesis, and therapy. Science 1995, 267, 483–489. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Steain, M.C.; Wang, B.; Dwyer, D.E.; Saksena, N.K. HIV-1 co-infection, superinfection and recombination. Sex Health 2004, 1, 239–250. [Google Scholar] [CrossRef] [Scilit]
  10. Dean, M.; Carrington, M.; Winkler, C.; Huttley, G.A.; Smith, M.W.; Allikmets, R.; Goedert, J.J.; Buchbinder, S.P.; Vittinghoff, E.; Gomperts, E.; et al. Genetic restriction of HIV-1 infection and progression to AIDS by a deletion allele of the CKR5 structural gene. Hemophilia Growth and Development Study, Multicenter AIDS Cohort Study, Multicenter Hemophilia Cohort Study, San Francisco City Cohort, ALIVE Study. Science 1996, 273, 1856–1862. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Zanini, F.; Brodin, J.; Thebo, L.; Lanz, C.; Bratt, G.; Albert, J.; Neher, R.A. Population genomics of intrapatient HIV-1 evolution. eLife 2015, 4, e11282. [Google Scholar] [CrossRef] [Scilit]
  12. Skittrall, J.P.; Ingemarsdotter, C.K.; Gog, J.R.; Lever, A.M.L. A scale-free analysis of the HIV-1 genome demonstrates multiple conserved regions of structural and functional importance. PLoS Comput. Biol. 2019, 15, e1007345. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  13. Yll, M.; Cortese, M.F.; Guerrero-Murillo, M.; Orriols, G.; Gregori, J.; Casillas, R.; González, C.; Sopena, S.; Godoy, C.; Vila, M.; et al. Conservation and variability of hepatitis B core at different chronic hepatitis stages. World J. Gastroenterol. 2020, 26, 2584–2598. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  14. Troyano-Hernáez, P.; Reinosa, R.; Holguín, Á. HIV Capsid Protein Genetic Diversity Across HIV-1 Variants and Impact on New Capsid-Inhibitor Lenacapavir. Front Microbiol. 2022, 13, 854974. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  15. Mbondji-Wonje, C.; Dong, M.; Zhao, J.; Wang, X.; Nanfack, A.; Ragupathy, V.; Sanchez, A.M.; Denny, T.N.; Hewlett, I. Genetic variability of the U5 and downstream sequence of major HIV-1 subtypes and circulating recombinant forms. Sci. Rep. 2020, 10, 13214. [Google Scholar] [CrossRef] [Scilit]
  16. Klingler, J.; Anton, H.; Réal, E.; Zeiger, M.; Moog, C.; Mély, Y.; Boutant, E. How HIV-1 Gag Manipulates Its Host Cell Proteins: A Focus on Interactors of the Nucleocapsid Domain. Viruses 2020, 12, 888. [Google Scholar] [CrossRef] [Scilit]
  17. Cannon, P.M.; Matthews, S.; Clark, N.; Byles, E.D.; Iourin, O.; Hockley, D.J.; Kingsman, S.M.; Kingsman, A.J. Structure-function studies of the human immunodeficiency virus type 1 matrix protein, p17. J. Virol. 1997, 71, 3474–3483. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  18. Doherty, R.S.; De Oliveira, T.; Seebregts, C.; Danaviah, S.; Gordon, M.; Cassol, S. BioAfrica’s HIV-1 proteomics resource: Combining protein data with bioinformatics tools. Retrovirology 2005, 2, 18. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  19. Alfadhli, A.; Still, A.; Barklis, E. Analysis of human immunodeficiency virus type 1 matrix binding to membranes and nucleic acids. J. Virol. 2009, 83, 12196–12203. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  20. Saad, J.S.; Miller, J.; Tai, J.; Kim, A.; Ghanam, R.H.; Summers, M.F. Structural basis for targeting HIV-1 Gag proteins to the plasma membrane for virus assembly. Proc. Natl. Acad. Sci. USA 2006, 103, 11364–11369. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  21. De Matteis, M.A.; Godi, A. PI-loting membrane traffic. Nat. Cell Biol. 2004, 6, 487–492. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Verli, H.; Calazans, A.; Brindeiro, R.; Tanuri, A.; Guimarães, J.A. Molecular dynamics analysis of HIV-1 matrix protein: Clarifying differences between crystallographic and solution structures. J. Mol. Graph. Model. 2007, 26, 62–68. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Hill, C.P.; Worthylake, D.; Bancroft, D.P.; Christensen, A.M.; Sundquist, W.I. Crystal structures of the trimeric human immunodeficiency virus type 1 matrix protein: Implications for membrane association and assembly. Proc. Natl. Acad. Sci. USA 1996, 93, 3099–3104. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  24. Ganser, B.K.; Li, S.; Klishko, V.Y.; Finch, J.T.; Sundquist, W.I. Assembly and analysis of conical models for the HIV-1 core. Science 1999, 283, 80–83. [Google Scholar] [CrossRef] [Scilit]
  25. Pornillos, O.; Ganser-Pornillos, B.K.; Yeager, M. Atomic-level modelling of the HIV capsid. Nature 2011, 469, 424–427. [Google Scholar] [CrossRef] [Scilit]
  26. Gamble, T.R.; Vajdos, F.F.; Yoo, S.; Worthylake, D.K.; Houseweart, M.; Sundquist, W.I.; Hill, C.P. Crystal structure of human cyclophilin A bound to the amino-terminal domain of HIV-1 capsid. Cell 1996, 87, 1285–1294. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Gitti, R.K.; Lee, B.M.; Walker, J.; Summers, M.F.; Yoo, S.; Sundquist, W.I. Structure of the Amino-Terminal Core Domain of the HIV-1 Capsid Protein. Science 1996, 273, 231–235. [Google Scholar] [CrossRef] [Scilit]
  28. Jacques, D.A.; McEwan, W.A.; Hilditch, L.; Price, A.J.; Towers, G.J.; James, L.C. HIV-1 uses dynamic capsid pores to import nucleotides and fuel encapsidated DNA synthesis. Nature 2016, 536, 349–353. [Google Scholar] [CrossRef] [Scilit]
  29. Gamble, T.R. Structure of the Carboxyl-Terminal Dimerization Domain of the HIV-1 Capsid Protein. Science 1997, 278, 849–853. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Von Schwedler, U.K. Proteolytic refolding of the HIV-1 capsid protein amino-terminus facilitates viral core assembly. EMBO J. 1998, 17, 1555–1568. [Google Scholar]
  31. Pornillos, O.; Ganser-Pornillos, B.K.; Kelly, B.N.; Hua, Y.; Whitby, F.G.; Stout, C.D.; Sundquist, W.I.; Hill, C.P.; Yeager, M. X-ray structures of the hexameric building block of the HIV capsid. Cell 2009, 137, 1282–1292. [Google Scholar] [CrossRef] [Scilit]
  32. Rossi, E.; Meuser, M.E.; Cunanan, C.J.; Cocklin, S. Structure, Function, and Interactions of the HIV-1 Capsid Protein. Life 2021, 11, 100. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  33. Bhattacharya, A.; Alam, S.L.; Fricke, T.; Zadrozny, K.; Sedzicki, J.; Taylor, A.B.; Demeler, B.; Pornillos, O.; Ganser-Pornillos, B.K.; Diaz-Griffero, F.; et al. Structural basis of HIV-1 capsid recognition by PF74 and CPSF6. Proc. Natl. Acad. Sci. USA 2014, 111, 18625–18630. [Google Scholar] [CrossRef] [Scilit]
  34. Zhao, G.; Perilla, J.; Yufenyuy, E.; Meng, X.; Chen, B.; Ning, J.; Ahn, J.; Gronenborn, A.M.; Schulten, K.; Aiken, C.; et al. Mature HIV-1 capsid structure by cryo-electron microscopy and all-atom molecular dynamics. Nature 2013, 497, 643–646. [Google Scholar] [CrossRef] [Scilit]
  35. Mascarenhas, A.P.; Musier-Forsyth, K. The capsid protein of human immunodeficiency virus: Interactions of HIV-1 capsid with host protein factors. FEBS J. 2009, 276, 6118–6127. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Post, K.; Olson, E.D.; Naufer, M.N.; Gorelick, R.J.; Rouzina, I.; Williams, M.C.; Musier-Forsyth, K.; Levin, J.G. Mechanistic differences between HIV-1 and SIV nucleocapsid proteins and cross-species HIV-1 genomic RNA recognition. Retrovirology 2016, 13, 89. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  37. Amarasinghe, G.K.; De Guzman, R.N.; Turner, R.B.; Chancellor, K.J.; Wu, Z.R.; Summers, M.F. NMR structure of the HIV-1 nucleocapsid protein bound to stem-loop SL2 of the Ψ-RNA packaging signal. Implications for genome recognition. J. Mol. Biol. 2000, 301, 491–511. [Google Scholar] [CrossRef] [Scilit]
  38. Bourbigot, S.; Ramalanjaona, N.; Boudier, C.; Salgado, G.F.J.; Roques, B.P.; Mély, Y.; Bouaziz, S.; Morellet, N. How the HIV-1 Nucleocapsid Protein Binds and Destabilises the (−)Primer Binding Site During Reverse Transcription. J. Mol. Biol. 2008, 383, 1112–1128. [Google Scholar] [CrossRef] [Scilit]
  39. De Guzman, R.N. Structure of the HIV-1 Nucleocapsid Protein Bound to the SL3 -RNA Recognition Element. Science 1998, 279, 384–388. [Google Scholar] [CrossRef] [Scilit]
  40. Godet, J.; Kenfack, C.; Przybilla, F.; Richert, L.; Duportail, G.; Mély, Y. Site-selective probing of cTAR destabilization highlights the necessary plasticity of the HIV-1 nucleocapsid protein to chaperone the first strand transfer. Nucleic Acids Res. 2013, 41, 5036–5048. [Google Scholar] [CrossRef] [Scilit]
  41. Darlix, J.-L.; Godet, J.; Ivanyi-Nagy, R.; Fossé, P.; Mauffret, O.; Mély, Y. Flexible nature and specific functions of the HIV-1 nucleocapsid protein. J. Mol. Biol. 2011, 410, 565–581. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Godet, J.; Boudier, C.; Humbert, N.; Ivanyi-Nagy, R.; Darlix, J.-L.; Mély, Y. Comparative nucleic acid chaperone properties of the nucleocapsid protein NCp7 and Tat protein of HIV-1. Virus Res. 2012, 169, 349–360. [Google Scholar] [CrossRef] [Scilit]
  43. Ott, D.E.; Chertova, E.N.; Busch, L.K.; Coren, L.V.; Gagliardi, T.D.; Johnson, D.G. Mutational Analysis of the Hydrophobic Tail of the Human Immunodeficiency Virus Type 1 p6Gag Protein Produces a Mutant That Fails To Package Its Envelope Protein. J. Virol. 1999, 73, 19–28. [Google Scholar] [CrossRef] [Scilit]
  44. Dettenhofer, M.; Yu, X.-F. Proline Residues in Human Immunodeficiency Virus Type 1 p6Gag Exert a Cell Type-Dependent Effect on Viral Replication and Virion Incorporation of Pol Proteins. J. Virol. 1999, 73, 4696–4704. [Google Scholar] [CrossRef] [Scilit]
  45. Yu, X.-F.; Dawson, L.; Tian, C.-J.; Flexner, C.; Dettenhofer, M. Mutations of the Human Immunodeficiency Virus Type 1 p6Gag Domain Result in Reduced Retention of Pol Proteins during Virus Assembly. J. Virol. 1998, 72, 3412–3417. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Göttlinger, H.G.; Dorfman, T.; Sodroski, J.G.; Haseltine, W.A. Effect of mutations affecting the p6 gag protein on human immunodeficiency virus particle release. Proc. Natl. Acad. Sci. USA 1991, 88, 3195–3199. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Huang, M.; Orenstein, J.M.; Martin, M.A.; Freed, E.O. p6Gag is required for particle production from full-length human immunodeficiency virus type 1 molecular clones expressing protease. J. Virol. 1995, 69, 6810–6818. [Google Scholar]
  48. Accola, M.A.; Bukovsky, A.A.; Jones, M.S.; Göttlinger, H.G. A conserved dileucine-containing motif in p6gag governs the particle association of Vpx and Vpr of simian immunodeficiency viruses SIVmac and SIVagm. J. Virol. 1999, 73, 9992–9999. [Google Scholar] [CrossRef] [Scilit]
  49. Bleiber, G.; Peters, S.; Martinez, R.; Cmarko, D.; Meylan, P.; Telenti, A. The central region of human immunodeficiency virus type 1 p6 protein (Gag residues S14–I31) is dispensable for the virus in vitro. J. Gen. Virol. 2004, 85, 921–927. [Google Scholar] [CrossRef] [Scilit]
  50. Datta, S.A.K.; Clark, P.K.; Fan, L.; Ma, B.; Harvin, D.P.; Sowder, R.C.; Nussinov, R.; Wang, Y.-X.; Rein, A. Dimerization of the SP1 Region of HIV-1 Gag Induces a Helical Conformation and Association into Helical Bundles: Implications for Particle Assembly. J. Virol. 2016, 90, 1773–1787. [Google Scholar] [CrossRef] [Scilit]
  51. Datta, S.A.K.; Temeselew, L.G.; Crist, R.M.; Soheilian, F.; Kamata, A.; Mirro, J.; Harvin, D.; Nagashima, K.; Cachau, R.E.; Rein, A. On the Role of the SP1 Domain in HIV-1 Particle Assembly: A Molecular Switch? J. Virol. 2011, 85, 4111–4121. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Ganser-Pornillos, B.K.; Yeager, M.; Sundquist, W.I. The structural biology of HIV assembly. Curr. Opin. Struct. Biol. 2008, 18, 203–217. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  53. Jacks, T.; Power, M.D.; Masiarz, F.R.; Luciw, P.A.; Barr, P.J.; Varmus, H.E. Characterization of ribosomal frameshifting in HIV-1 gag-pol expression. Nature 1988, 331, 280–283. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Freed, E.O. HIV-1 assembly, release and maturation. Nat. Rev. Microbiol. 2015, 13, 484–496. [Google Scholar] [CrossRef] [Scilit]
  55. Adamson, C.S. Protease-mediated maturation of HIV: Inhibitors of protease and the maturation process. Mol. Biol. Int. 2012, 2012, 604261. [Google Scholar] [CrossRef] [Scilit]
  56. Adamson, C.S.; Freed, E.O. Human immunodeficiency virus type 1 assembly, release, and maturation. Adv. Pharmacol. 2007, 55, 347–387. [Google Scholar] [PubMed]
  57. Harrison, J.J.E.K.; Passos, D.O.; Bruhn, J.F.; Bauman, J.D.; Tuberty, L.; DeStefano, J.J.; Ruiz, F.X.; Lyumkis, D.; Arnold, E. Cryo-EM structure of the HIV-1 Pol polyprotein provides insights into virion maturation. Sci. Adv. 2022, 8, eabn9874. [Google Scholar] [CrossRef] [Scilit]
  58. Toh, H.; Ono, M.; Saigo, K.; Miyata, T. Retroviral Protease-like Sequence in the Yeast Transposon Ty 1. Nature 1985, 315, 691. [Google Scholar] [CrossRef] [Scilit]
  59. Wiegers, K.; Rutter, G.; Kottler, H.; Tessmer, U.; Hohenberg, H.; Krausslich, H.G. Sequential Steps in Human Immunodeficiency Virus Particle Maturation Revealed by Alterations of Individual Gag Polyprotein Cleavage Sites. J. Virol. 1998, 72, 2846–2854. [Google Scholar] [CrossRef] [Scilit]
  60. Pettit, S.C.; Lindquist, J.N.; Kaplan, A.H.; Swanstrom, R. Processing Sites in the Human Immunodeficiency Virus Type 1 (HIV-1) Gag-Pro-Pol Precursor are Cleaved by the Viral Protease at Different Rates. Retrovirology 2005, 2, 66. [Google Scholar] [PubMed]
  61. Deshmukh, L.; Tugarinov, V.; Louis, J.M.; Clore, G.M. Binding Kinetics and Substrate Selectivity in HIV-1 Protease-Gag Interactions Probed at Atomic Resolution by Chemical Exchange NMR. Proc. Natl. Acad. Sci. USA 2017, 114, E9855–E9862. [Google Scholar] [CrossRef] [Scilit]
  62. Sherry, D.; Sayed, Y. Unveiling a Hidden Pocket in HIV-1 Protease: New Insights Into Retroviral Protease Cantilever-Tip Region Characteristics. Proteins 2024, 92, 1398–1412. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  63. Rhee, S.Y.; Gonzales, M.J.; Kantor, R.; Betts, B.J.; Ravela, J.; Shafer, R.W. Human Immunodeficiency Virus Reverse Transcriptase and Protease Sequence Database. Nucleic Acids Res. 2003, 31, 298–303. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  64. Shafer, R.W. Rationale and Uses of a Public HIV Drug-resistance Database. J. Infect. Dis. 2006, 194, S51–S58. [Google Scholar] [CrossRef] [Scilit]
  65. Jacobo-Molina, A.; Ding, J.; Nanni, R.G.; Clark, A.D., Jr.; Lu, X.; Tantillo, C.; Williams, R.L.; Kamer, G.; Ferris, A.L.; Clark, P.; et al. Crystal structure of human immunodeficiency virus type 1 reverse transcriptase complexed with double-stranded DNA at 3.0 A resolution shows bent DNA. Proc. Natl. Acad. Sci. USA 1993, 90, 6320–6324. [Google Scholar] [CrossRef] [Scilit]
  66. Kohlstaedt, L.A.; Wang, J.; Friedman, J.M.; Rice, P.A.; Steitz, T.A. Crystal structure at 3.5 A resolution of HIV-1 reverse transcriptase complexed with an inhibitor. Science 1992, 256, 1783–1790. [Google Scholar] [CrossRef] [Scilit]
  67. Rodgers, D.W.; Gamblin, S.J.; Harris, B.A.; Ray, S.; Culp, J.S.; Hellmig, B.; Woolf, D.J.; Debouck, C.; Harrison, S.C. The structure of unliganded reverse transcriptase from the human immunodeficiency virus type 1. Proc. Natl. Acad. Sci. USA 1995, 92, 1222–1226. [Google Scholar] [CrossRef] [Scilit]
  68. Coffin, J.M.; Hughes, S.H.; Varmus, H.E. Retroviruses; Cold Spring Harbor Laboratory Press: New York, NY, USA, 1997. [Google Scholar]
  69. Smith, J.S.; Roth, M.J. Specificity of human immunodeficiency virus-1 reverse transcriptase-associated ribonuclease H in removal of the minus-strand primer, tRNA(Lys3). J. Biol. Chem. 1992, 267, 15071–15079. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  70. Wohrl, B.M.; Moelling, K. Interaction of HIV-1 ribonuclease H with polypurine tract containing RNA-DNA hybrids. Biochemistry 1990, 29, 10141–10147. [Google Scholar] [CrossRef] [Scilit]
  71. Ceccherini-Silberstein, F.; Gago, F.; Santoro, M.; Gori, C.; Svicher, V.; Rodríguez-Barrios, F.; d’Arrigo, R.; Ciccozzi, M.; Bertoli, A.; d’Arminio Monforte, A.; et al. High sequence conservation of human immunodeficiency virus type 1 reverse transcriptase under drug pressure despite the continuous appearance of mutations. J. Virol. 2005, 79, 10718–10729. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  72. Alcaro, S.; Artese, A.; Ceccherini-Silberstein, F.; Chiarella, V.; Dimonte, S.; Ortuso, F.; Perno, C.F. Computational analysis of Human Immunodeficiency Virus (HIV) Type-1 reverse transcriptase crystallographic models based on significant conserved residues found in Highly Active Antiretroviral Therapy (HAART)-treated patients. Curr. Med. Chem. 2010, 17, 290–308. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  73. Santos, A.F.; Lengruber, R.B.; Soares, E.A.; Jere, A.; Sprinz, E.; Martinez, A.M.; Silveira, J.; Sion, F.S.; Pathak, V.K.; Soares, M.A. Conservation patterns of HIV-1 RT connection and RNase H domains: Identification of new mutations in NRTI-treated patients. PLoS ONE 2008, 3, e1781. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  74. Nagata, S.; Imai, J.; Makino, G.; Tomita, M.; Kanai, A. Evolutionary Analysis of HIV-1 Pol Proteins Reveals Representative Residues for Viral Subtype Differentiation. Front. Microbiol. 2017, 8, 2151. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  75. Mckee, C.J.; Kessl, J.J.; Shkriabia, N.; Dar, M.J.; Engelman, A.; Kvaratskhelia, M. Dynamic modulation of HIV-1 integrase structure and function by cellular Lens Epithelium-derived Growth Factor (LEDGF) Protein. J. Biol. Chem. 2008, 283, 31802–31812. [Google Scholar] [CrossRef] [Scilit]
  76. Ceccherini-Silberstein, F.; Malet, I.; D’Arrigo, R.; Antinori, A.; Marcelin, A.G.; Perno, C.F. Characterization and structural analysis of HIV-1 integrase conservation. AIDS Rev. 2009, 11, 17–29. [Google Scholar] [PubMed]
  77. Harrison, S.C. Mechanism of membrane fusion by viral envelope proteins. Adv. Virus Res. 2005, 64, 231–261. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  78. Weissenhorn, W.; Dessen, A.; Harrison, S.C.; Skehel, J.J.; Wiley, D.C. Atomic structure of the ectodomain from HIV-1 gp41. Nature 1997, 387, 426–430. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  79. Chan, D.C.; Fass, D.; Berger, J.M.; Kim, P.S. Core structure of gp41 from the HIV envelope glycoprotein. Cell 1997, 89, 263–273. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  80. Modrow, S.; Hahn, B.H.; Shaw, G.M.; Gallo, R.C.; Wong-Staal, F.; Wolf, H. Computer-assisted analysis of envelope protein sequences of seven human immunodeficiency virus isolates: Prediction of antigenic epitopes in conserved and variable regions. J. Virol. 1987, 61, 570–578. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  81. Klasse, P.J.; Sanders, R.W.; Ward, A.B.; Wilson, I.A.; Moore, J.P. The HIV-1 envelope glycoprotein: Structure, function and interactions with neutralizing antibodies. Nat. Rev. Microbiol. 2025, 23, 734–752. [Google Scholar] [CrossRef] [Scilit]
  82. Valadés-Alcaraz, A.; Reinosa, R.; Holguín, Á. HIV Transmembrane Glycoprotein Conserved Domains and Genetic Markers Across HIV-1 and HIV-2 Variants. Front. Microbiol. 2022, 13, 855232. [Google Scholar] [CrossRef] [Scilit]
  83. Azimi, F.C.; Lee, J.E. Structural perspectives on HIV-1 Vif and APOBEC3 restriction factor interactions. Protein Sci. 2020, 29, 391–406. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  84. Xiao, Z.; Ehrlich, E.; Yu, Y.; Luo, K.; Wang, T.; Tian, C.; Yu, X.-F. Assembly of HIV-1 Vif-Cul5 E3 ubiquitin ligase through a novel zinc-binding domain-stabilized hydrophobic interface in Vif. Virology 2006, 349, 290–299. [Google Scholar] [CrossRef] [Scilit]
  85. Mehle, A.; Thomas, E.R.; Rajendran, K.S.; Gabuzda, D. A zinc-binding region in Vif binds Cul5 and determines cullin selection. J. Biol. Chem. 2006, 281, 17259–17265. [Google Scholar] [CrossRef] [Scilit]
  86. Paul, I.; Cui, J.; Maynard, E.L. Zinc binding to the HCCH motif of HIV-1 virion infectivity factor induces a conformational change that mediates protein-protein interactions. Proc. Natl. Acad. Sci. USA 2006, 103, 18475–18480. [Google Scholar] [CrossRef] [Scilit]
  87. Morellet, N.; Bouaziz, S.; Petitjean, P.; Roques, B.P. NMR structure of the HIV-1 regulatory protein VPR. J. Mol. Biol. 2003, 327, 215–227. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  88. Byeon, I.J.L.; Calero, G.; Wu, Y.; Byeon, C.-H.; Jung, J.; DeLucia, M.; Zhou, X.; Weiss, S.; Ahn, J.; Hao, C.; et al. Structure of HIV-1 Vpr in complex with the human nucleotide excision repair protein hHR23A. Nat. Commun. 2021, 12, 6864. [Google Scholar] [CrossRef] [Scilit]
  89. Rausch, J.W.; Le Grice, S.F. HIV Rev Assembly on the Rev Response Element (RRE): A Structural Perspective. Viruses 2015, 7, 3053–3075. [Google Scholar] [CrossRef] [Scilit]
  90. Jayaraman, B.; Fernandes, J.D.; Yang, S.; Smith, C.; Frankel, A.D. Highly Mutable Linker Regions Regulate HIV-1 Rev Function and Stability. Sci. Rep. 2019, 9, 5139. [Google Scholar] [CrossRef] [Scilit]
  91. Truman, C.T.; Järvelin, A.; Davis, I.; Castello, A. HIV Rev-isited. Open Biol. 2020, 10, 200320. [Google Scholar] [CrossRef] [Scilit]
  92. Stoltzfus, C.M. Chapter 1 Regulation of HIV-1 alternative RNA splicing and its role in virus replication. Adv. Virus Res. 2009, 74, 1–40. [Google Scholar] [CrossRef] [Scilit]
  93. Rolland, M.; Nickle, D.C.; Mullins, J.I. HIV-1 group M conserved elements vaccine. PLoS Pathog. 2007, 3, e157. [Google Scholar] [CrossRef] [Scilit]
  94. Kuznetsova, A.; Kim, K.; Tumanov, A.; Munchak, I.; Antonova, A.; Lebedev, A.; Ozhmegova, E.; Orlova-Morozova, E.; Drobyshevskaya, E.; Pronin, A.; et al. Features of Tat Protein in HIV-1 Sub-Subtype A6 Variants Circulating in the Moscow Region, Russia. Viruses 2023, 15, 2212. [Google Scholar] [CrossRef] [Scilit]
  95. Wong-Staal, F.; Gallo, R.C. Human T-lymphotropic retroviruses. Nature 1985, 317, 395–403. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  96. Debaisieux, S.; Rayne, F.; Yezid, H.; Beaumelle, B. The ins and outs of HIV-1 Tat. Traffic 2012, 13, 355–363. [Google Scholar] [CrossRef] [Scilit]
  97. Westendorp, M.O.; Frank, R.; Ochsenbauer, C.; Stricker, K.; Dhein, J.; Walczak, H.; Debatin, K.M.; Krammer, P.H. Sensitization of T cells to CD95-mediated apoptosis by HIV-1 Tat and gp120. Nature 1995, 375, 497–500. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  98. Mediouni, S.; Darque, A.; Ravaux, I.; Baillat, G.; Devaux, C.; Loret, E.P. Identification of a highly conserved surface on Tat variants. J. Biol. Chem. 2013, 288, 19072–19080. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  99. Geyer, M.; Peterlin, B.M. Domain assembly, surface accessibility and sequence conservation in full length HIV-1 Nef. FEBS Lett. 2001, 496, 91–95. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  100. Saksela, K.; Cheng, G.; Baltimore, D. Proline-rich (PxxP) motifs in HIV-1 Nef bind to SH3 domains of a subset of Src kinases and are required for the enhanced growth of Nef+ viruses but not for down-regulation of CD4. EMBO J. 1995, 14, 484–491. [Google Scholar] [CrossRef] [Scilit]
  101. Fackler, O.T.; Luo, W.; Geyer, M.; Alberts, A.S.; Peterlin, B.M. Activation of Vav by Nef induces cytoskeletal rearrangements and downstream effector functions. Mol. Cell 1999, 3, 729–739. [Google Scholar] [CrossRef] [Scilit]
  102. Geyer, M.; Fackler, O.T.; Peterlin, B.M. Structure–function relationships in HIV-1 Nef. EMBO Rep. 2001, 2, 580–585. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
  103. Greenberg, M.E.; Iafrate, A.; Skowronski, J. The SH3 domain-binding surface and an acidic motif in HIV-1 Nef regulate trafficking of class I MHC complexes. EMBO J. 1998, 17, 2777–2789. [Google Scholar] [CrossRef] [Scilit]
  104. Pickering, S.; Hué, S.; Kim, E.-Y.; Reddy, S.; Wolinsky, S.M.; Neil, S.J.D. Preservation of Tetherin and CD4 Counter-Activities in Circulating Vpu Alleles despite Extensive Sequence Variation within HIV-1 Infected Individuals. PLoS Pathog. 2014, 10, e1003895, Correction in PLoS Pathog. 2014, 10, E1004118. [Google Scholar] [CrossRef] [Scilit]
  105. Umviligihozo, G.; Cobarrubias, K.D.; Chandrarathna, S.; Jin, S.W.; Reddy, N.; Byakwaga, H.; Muzoora, C.; Bwana, M.B.; Lee, G.Q.; Hunt, P.W.; et al. Differential Vpu-Mediated CD4 and Tetherin Downregulation Functions among Major HIV-1 Group M Subtypes. J. Virol. 2020, 94, e00293-20. [Google Scholar] [CrossRef] [Scilit]
  106. Pawlak, E.N.; Dirk, B.S.; Jacob, R.A.; Johnson, A.L.; Dikeakos, J.D. The HIV-1 accessory proteins Nef and Vpu downregulate total and cell surface CD28 in CD4+ T cells. Retrovirology 2018, 15, 6. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  107. Federau, T.; Schubert, U.; Flossdorf, J.; Henklein, P.; Schomburg, D.; Wray, V. Solution structure of the cytoplasmic domain of the human immunodeficiency virus type 1 encoded virus protein U (Vpu). Int. J. Pept. Protein Res. 1996, 47, 297–310. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  108. Leymarie, O.; Lepont, L.; Versapuech, M.; Judith, D.; Abelanet, S.; Janvier, K.; Berlioz-Torrent, C. Contribution of the Cytoplasmic Determinants of Vpu to the Expansion of Virus-Containing Compartments in HIV-1-Infected Macrophages. J. Virol. 2019, 93, 11. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  109. Edgar, R.C. Muscle5: High-accuracy alignment ensembles enable unbiased assessments of sequence homology and phylogeny. Nat. Commun. 2022, 13, 6968. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Aggregate values for amino acid sequence alignments of different HIV-1 subtypes. Red asterisks indicate subtypes represented by a limited number of sequences in the alignment.
Figure 1. Aggregate values for amino acid sequence alignments of different HIV-1 subtypes. Red asterisks indicate subtypes represented by a limited number of sequences in the alignment.
Ijms 27 05139 g001
Figure 2. Averaged S-index panorama for the alignment of Gag amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Figure 2. Averaged S-index panorama for the alignment of Gag amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Ijms 27 05139 g002
Figure 3. Generalized representation of conserved regions identified in Gag amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Figure 3. Generalized representation of conserved regions identified in Gag amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Ijms 27 05139 g003
Figure 4. Averaged S-index panorama for the alignment of Pol amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Figure 4. Averaged S-index panorama for the alignment of Pol amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Ijms 27 05139 g004
Figure 5. Generalized representation of conserved regions identified in Pol amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Figure 5. Generalized representation of conserved regions identified in Pol amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Ijms 27 05139 g005
Figure 6. Averaged S-index panorama for the alignment of Env amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Figure 6. Averaged S-index panorama for the alignment of Env amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Ijms 27 05139 g006
Figure 7. Generalized representation of conserved regions identified in Env amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Figure 7. Generalized representation of conserved regions identified in Env amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Ijms 27 05139 g007
Figure 8. Averaged S-index panorama for the alignment of Vif amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Figure 8. Averaged S-index panorama for the alignment of Vif amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Ijms 27 05139 g008
Figure 9. Generalized representation of conserved regions identified in Vif amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Figure 9. Generalized representation of conserved regions identified in Vif amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Ijms 27 05139 g009
Figure 10. Averaged S-index panorama for the alignment of Vpr amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Figure 10. Averaged S-index panorama for the alignment of Vpr amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Ijms 27 05139 g010
Figure 11. Generalized representation of conserved regions identified in Vpr amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Figure 11. Generalized representation of conserved regions identified in Vpr amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Ijms 27 05139 g011
Figure 12. Averaged S-index panorama for the alignment of Rev amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Figure 12. Averaged S-index panorama for the alignment of Rev amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Ijms 27 05139 g012
Figure 13. Generalized representation of conserved regions identified in Rev amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Figure 13. Generalized representation of conserved regions identified in Rev amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Ijms 27 05139 g013
Figure 14. Averaged S-index panorama for the alignment of Tat amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Figure 14. Averaged S-index panorama for the alignment of Tat amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Ijms 27 05139 g014
Figure 15. Generalized representation of conserved regions identified in Tat amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Figure 15. Generalized representation of conserved regions identified in Tat amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Ijms 27 05139 g015
Figure 16. Averaged S-index panorama for the alignment of Nef amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Figure 16. Averaged S-index panorama for the alignment of Nef amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Ijms 27 05139 g016
Figure 17. Generalized representation of conserved regions identified in Nef amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Figure 17. Generalized representation of conserved regions identified in Nef amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Ijms 27 05139 g017
Figure 18. Averaged S-index panorama for the alignment of Vpu amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Figure 18. Averaged S-index panorama for the alignment of Vpu amino acid sequences obtained using a sliding window of 50 amino acids. The colors correspond to the mean S-index value according to the color scale to the right of the figure. The marks above the figure indicate MSA regions corresponding to gaps in the HXB2 sequence. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates, with colors chosen randomly for contrast.
Ijms 27 05139 g018
Figure 19. Generalized representation of conserved regions identified in Vpu amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Figure 19. Generalized representation of conserved regions identified in Vpu amino acid sequence alignments. The colors correspond to the mean S-index value according to the color scale to the right of the figure. Below the figure, annotations of the HXB2 sequence are shown in alignment coordinates.
Ijms 27 05139 g019
Figure 20. Workflow of the computational pipeline used for identification of conserved regions in HIV-1 proteins.
Figure 20. Workflow of the computational pipeline used for identification of conserved regions in HIV-1 proteins.
Ijms 27 05139 g020
Table 1. Statistical analysis of S-index values for establishing cutoff thresholds for subsequent analysis.
Table 1. Statistical analysis of S-index values for establishing cutoff thresholds for subsequent analysis.
SubtypeSmCI−CI+
All sequences0.87460.86710.8822
A10.91580.90980.9227
A20.90300.89610.9095
A60.92700.92080.9327
B0.93120.92570.9370
C0.90860.90170.9153
D0.88120.87330.8890
F10.90630.89890.9133
F20.90080.89330.9075
G0.86850.86020.8757
H0.92300.91630.9289
J0.87170.86420.8792
K0.83560.82870.8427
Table 2. S-index statistics for Gag protein aligns.
Table 2. S-index statistics for Gag protein aligns.
SubtypeSmCI−CI+
All sequences0.90580.88960.9209
A10.93930.92690.9516
A20.92480.91050.9383
A60.94590.93480.9560
B0.95170.94230.9605
C0.92420.90870.9374
D0.90870.89370.9223
F10.92160.90740.9347
F20.92160.90640.9368
G0.91430.89910.9300
H0.94320.93090.9547
J0.90870.89510.9228
K0.92700.91280.9404
Table 3. Statistics of conserved regions detected in Gag.
Table 3. Statistics of conserved regions detected in Gag.
SubtypeConserved RegionsMean Conserved Region LengthSm in Conserved Regions
All sequences268.350.9856
A12410.670.9953
A2277.671.0000
A6249.080.9957
B308.130.9949
C268.120.9952
D257.920.9863
F1259.240.9912
F2238.521.0000
G288.710.9896
H2410.040.9982
J258.040.9909
K237.961.0000
Table 4. Conserved region summary for Gag polyprotein aligns.
Table 4. Conserved region summary for Gag polyprotein aligns.
Region NameHXB2 CoordinatesAlignment CoordinatesLengthSm
Gag_CR_11–61–660.9979
Gag_CR_221–2521–2550.9916
Gag_CR_336–4536–45100.9871
Gag_CR_496–10196–10160.9812
Gag_CR_5128–136140–14890.9708
Gag_CR_6147–157159–169110.9900
Gag_CR_7176–180188–19250.9777
Gag_CR_8190–201202–213120.9878
Gag_CR_9203–213215–225110.9830
Gag_CR_10230–240242–252110.9959
Gag_CR_11260–278272–290190.9841
Gag_CR_12280–284292–29650.9858
Gag_CR_13286–299298–311140.9972
Gag_CR_14303–308315–32060.9880
Gag_CR_15319–324331–33660.9843
Gag_CR_16326–330338–34250.9924
Gag_CR_17342–355354–367140.9938
Gag_CR_18362–367374–37960.9928
Gag_CR_19389–398403–412100.9876
Gag_CR_20402–407416–42160.9800
Gag_CR_21410–415424–42960.9967
Gag_CR_22417–424431–43880.9885
Gag_CR_23425–433440–44890.9741
Gag_CR_24440–446457–46370.9836
Gag_CR_25487–495532–54090.9821
Table 5. S-index statistics for Pol protein aligns.
Table 5. S-index statistics for Pol protein aligns.
SubtypeSmCI−CI+
All sequences0.92000.91020.9286
A10.95920.95230.9661
A20.94630.93760.9547
A60.96170.95440.9683
B0.96540.95930.9711
C0.95440.94660.9613
D0.93600.92710.9444
F10.94620.93770.9546
F20.94440.93570.9532
G0.88120.87320.8887
H0.95650.94910.9635
J0.93570.92650.9452
K0.75880.75170.7661
Table 6. Statistics of conserved regions detected in Pol.
Table 6. Statistics of conserved regions detected in Pol.
SubtypeConserved RegionsMean Conserved Region LengthSm in Conserved Regions
All sequences509.960.9709
A15011.740.9971
A25410.151.0000
A65310.210.9963
B4911.140.9965
C4611.170.9965
D479.890.9890
F1519.980.9945
F2489.851.0000
G4810.460.9278
H5010.520.9961
J448.181.0000
K308.600.8091
Table 7. Conserved region summary for Pol polyprotein aligns.
Table 7. Conserved region summary for Pol polyprotein aligns.
Region NameHXB2 CoordinatesAlignment CoordinatesLengthSm
Pol_CR_1 82–9090.9703
Pol_CR_2 102–115140.9696
Pol_CR_311–21127–137110.9713
Pol_CR_448–53164–16960.9723
Pol_CR_559–67175–18390.9730
Pol_CR_678–83194–19960.9728
Pol_CR_785–95201–211110.9656
Pol_CR_8114–123230–239100.9710
Pol_CR_9127–131243–24750.9645
Pol_CR_10134–165250–281320.9684
Pol_CR_11168–184284–300170.9684
Pol_CR_12188–198304–314110.9771
Pol_CR_13207–221323–337150.9770
Pol_CR_14249–259367–382110.9666
Pol_CR_15279–301402–424230.9744
Pol_CR_16303–308426–43160.9695
Pol_CR_17316–335440–459200.9739
Pol_CR_18362–367486–49160.9729
Pol_CR_19369–374493–49860.9690
Pol_CR_20400–408524–53290.9720
Pol_CR_21411–419535–54390.9700
Pol_CR_22425–429549–55350.9640
Pol_CR_23469–476593–60080.9762
Pol_CR_24478–494602–618170.9683
Pol_CR_25505–509629–63350.9769
Pol_CR_26517–521641–64550.9679
Pol_CR_27536–540660–66450.9726
Pol_CR_28548–553672–67760.9664
Pol_CR_29557–575681–699190.9665
Pol_CR_30599–611723–735130.9776
Pol_CR_31613–617737–74150.9768
Pol_CR_32628–634752–75870.9634
Pol_CR_33650–654774–77850.9576
Pol_CR_34657–662781–78660.9697
Pol_CR_35675–683799–80790.9711
Pol_CR_36688–695812–81980.9727
Pol_CR_37699–707823–83190.9755
Pol_CR_38709–720833–844120.9675
Pol_CR_39738–742862–86650.9774
Pol_CR_40751–757875–88170.9755
Pol_CR_41761–779885–903190.9719
Pol_CR_42781–786905–91060.9681
Pol_CR_43792–796916–92050.9756
Pol_CR_44801–811925–935110.9680
Pol_CR_45818–824942–94870.9674
Pol_CR_46847–857971–981110.9708
Pol_CR_47859–874983–998160.9762
Pol_CR_48881–8881005–101280.9766
Pol_CR_49894–9011018–102580.9707
Table 8. S-index statistics for Env protein aligns.
Table 8. S-index statistics for Env protein aligns.
SubtypeSmCI−CI+
All sequences0.83200.81420.8487
A10.87540.86110.8900
A20.87160.85640.8858
A60.88860.87420.9013
B0.90000.88720.9120
C0.87310.85910.8867
D0.83400.81770.8506
F10.87330.85820.8878
F20.86790.85290.8826
G0.84660.83020.8619
H0.88680.87230.9002
J0.79890.78290.8126
K0.86520.84970.8799
Table 9. Statistics of conserved regions detected in Env.
Table 9. Statistics of conserved regions detected in Env.
SubtypeConserved RegionsMean Conserved Region LengthSm in Conserved Regions
All sequences338.700.9666
A1348.940.9885
A2408.750.9746
A6409.200.9815
B419.150.9840
C418.680.9794
D348.180.9642
F1328.880.9846
F2398.080.9806
G337.970.9749
H378.540.9881
J367.420.9192
K318.480.9831
Table 10. Conserved region summary for Env polyprotein aligns.
Table 10. Conserved region summary for Env polyprotein aligns.
Region NameHXB2 CoordinatesAlignment CoordinatesLengthSm
Env_CR_134–4436–46110.9920
Env_CR_249–5951–61110.9635
Env_CR_365–7867–80140.9791
Env_CR_492–9794–9960.9478
Env_CR_5106–119108–121140.9603
Env_CR_6121–128123–13080.9752
Env_CR_7202–206272–27650.9872
Env_CR_8211–217281–28770.9846
Env_CR_9223–227293–29750.9624
Env_CR_10244–250314–32070.9882
Env_CR_11252–266322–336150.9718
Env_CR_12377–384456–46380.9603
Env_CR_13417–425519–52790.9303
Env_CR_14430–436532–53870.9540
Env_CR_15474–486584–596130.9790
Env_CR_16505–509615–61950.9542
Env_CR_17518–531630–643140.9834
Env_CR_18533–539645–65170.9601
Env_CR_19541–549653–66190.9828
Env_CR_20555–561667–67370.9774
Env_CR_21565–579677–691150.9796
Env_CR_22586–591698–70360.9559
Env_CR_23593–598705–71060.9829
Env_CR_24610–614722–72760.9378
Env_CR_25675–679788–79250.9937
Env_CR_26685–694798–807100.9584
Env_CR_27703–713816–826110.9630
Env_CR_28843–847964–96850.9748
Table 11. S-index statistics for Vif protein aligns.
Table 11. S-index statistics for Vif protein aligns.
SubtypeSmCI−CI+
All sequences0.88090.85380.9075
A10.91830.89500.9372
A20.91320.88690.9366
A60.93220.91370.9497
B0.92210.90130.9434
C0.90810.88330.9299
D0.88380.85770.9101
F10.89990.87600.9254
F20.90840.88080.9317
G0.87950.85290.9070
H0.93140.91190.9503
J0.90100.87350.9251
K0.89820.87000.9239
Table 12. Statistics of conserved regions detected in Vif.
Table 12. Statistics of conserved regions detected in Vif.
SubtypeConserved RegionsMean Conserved Region LengthSm in Conserved Regions
All sequences77.000.9832
A1107.400.9941
A297.221.0000
A687.130.9900
B87.880.9930
C106.400.9913
D87.250.9794
F187.380.9955
F287.001.0000
G86.380.9854
H75.860.9958
J107.600.9785
K57.001.0000
Table 13. Conserved region summary for Vif polyprotein aligns.
Table 13. Conserved region summary for Vif polyprotein aligns.
Region NameHXB2 CoordinatesAlignment CoordinatesLengthSm
VIf_CR_11–161–16160.9761
VIf_CR_252–5953–6080.9738
VIf_CR_368–7270–7450.9923
VIf_CR_4141–150146–155100.9777
VIf_CR_5161–166166–17160.9877
VIf_CR_6171–175176–18050.9906
Table 14. S-index statistics for Vpr protein aligns.
Table 14. S-index statistics for Vpr protein aligns.
SubtypeSmCI−CI+
All sequences0.90910.86800.9441
A10.93780.90680.9649
A20.90260.86020.9398
A60.94340.91340.9697
B0.94000.90790.9663
C0.92900.89760.9554
D0.90930.86550.9447
F10.93360.90580.9589
F20.90700.86970.9429
G0.90630.86490.9403
H0.94420.91660.9666
J0.91000.87420.9442
K0.90240.86410.9392
Table 15. Statistics of conserved regions detected in Vpr.
Table 15. Statistics of conserved regions detected in Vpr.
SubtypeConserved RegionsMean Conserved Region LengthSm in Conserved Regions
All sequences46.500.9843
A156.400.9946
A245.251.0000
A665.660.9929
B55.400.9937
C56.200.9935
D36.000.9854
F146.250.9932
F246.501.0000
G56.000.9892
H66.500.9940
J47.001.0000
K36.331.0000
Table 16. Conserved region summary for Vpr polyprotein aligns.
Table 16. Conserved region summary for Vpr polyprotein aligns.
Region NameHXB2 CoordinatesAlignment CoordinatesLengthSm
Vpr_CR_17–127–1260.9749
Vpr_CR_229–3629–3680.9891
Vpr_CR_349–5449–5460.9864
Vpr_CR_4 78–8360.9867
Table 17. S-index statistics for Rev protein aligns.
Table 17. S-index statistics for Rev protein aligns.
SubtypeSmCI−CI+
All sequences0.82460.78660.8613
A10.87130.83580.9057
A20.86220.82600.8968
A60.90640.88260.9306
B0.90470.87500.9320
C0.85620.81870.8868
D0.83650.80180.8696
F10.87510.83840.9079
F20.87720.84210.9111
G0.84080.80260.8778
H0.92950.90920.9493
J0.85240.81610.8871
K0.86100.82870.8932
Table 18. Statistics of conserved regions detected in Rev.
Table 18. Statistics of conserved regions detected in Rev.
SubtypeConserved RegionsMean Conserved Region LengthSm in Conserved Regions
All sequences46.500.9694
A146.000.9919
A215.001.0000
A625.500.9928
B47.250.9864
C46.500.9877
D55.800.9702
F159.200.9872
F246.001.0000
G56.200.9715
H58.401.0000
J29.000.9905
K56.600.9603
Table 19. Conserved region summary for Rev polyprotein aligns.
Table 19. Conserved region summary for Rev polyprotein aligns.
Region NameHXB2 CoordinatesAlignment CoordinatesLengthSm
Rev_CR_11–61–660.9877
Rev_CR_222–2722–2760.9711
Rev_CR_333–3933–3970.9466
Rev_CR_441–4741–4770.9721
Table 20. S-index statistics for Tat protein aligns.
Table 20. S-index statistics for Tat protein aligns.
SubtypeSmCI−CI+
All sequences0.82050.77310.8662
A10.87520.83890.9083
A20.85320.81280.8922
A60.90450.87280.9330
B0.89120.85600.9208
C0.87260.83580.9102
D0.81790.77300.8641
F10.86750.82730.9058
F20.84140.79960.8810
G0.82850.78760.8695
H0.90450.87240.9327
J0.85390.80680.8912
K0.84130.80020.8797
Table 21. Statistics of conserved regions detected in Tat.
Table 21. Statistics of conserved regions detected in Tat.
SubtypeConserved RegionsMean Conserved Region LengthSm in Conserved Regions
All sequences37.660.9659
A145.750.9922
A235.661.0000
A6211.500.9923
B38.660.9821
C26.000.9910
D35.330.9729
F155.600.9770
F228.500.9849
G28.500.9631
H35.330.9917
J36.660.9855
K27.500.9907
Table 22. Conserved region summary for Tat polyprotein aligns.
Table 22. Conserved region summary for Tat polyprotein aligns.
Region NameHXB2 CoordinatesAlignment CoordinatesLengthSm
Tat_CR_113–1813–1860.9807
Tat_CR_241–5241–52120.9676
Table 23. S-index statistics for Nef protein aligns.
Table 23. S-index statistics for Nef protein aligns.
SubtypeSmCI−CI+
All sequences0.86620.83810.8936
A10.90160.87730.9254
A20.86010.83650.8837
A60.92630.90580.9442
B0.90470.88250.9274
C0.89200.86940.9156
D0.85620.82920.8826
F10.89110.86880.9127
F20.86890.84360.8926
G0.85480.82670.8806
H0.91170.88900.9317
J0.85110.82280.8766
K0.84740.82190.8716
Table 24. Statistics of conserved regions detected in Nef.
Table 24. Statistics of conserved regions detected in Nef.
SubtypeConserved RegionsMean Conserved Region LengthSm in Conserved Regions
All sequences87.250.9755
A198.220.9889
A279.280.9736
A6107.500.9895
B107.100.9836
C78.000.9823
D86.880.9657
F1117.000.9754
F279.710.9705
G78.420.9700
H77.000.9940
J107.000.9747
K77.570.9797
Table 25. Conserved region summary for Nef polyprotein aligns.
Table 25. Conserved region summary for Nef polyprotein aligns.
Region NameHXB2 CoordinatesAlignment CoordinatesLengthSm
Nef_CR_11–71–780.9654
Nef_CR_266–7094–9850.9885
Nef_CR_372–80100–10890.9893
Nef_CR_490–97118–12580.9887
Nef_CR_5109–115137–14370.9784
Nef_CR_6 154–16070.9745
Table 26. S-index statistics for Vpu protein aligns.
Table 26. S-index statistics for Vpu protein aligns.
GenotypeSmCI−CI+
All sequences0.78390.73070.8321
A10.86020.81550.9004
A20.84460.80250.8844
A60.88610.84480.9221
B0.90460.87220.9346
C0.83270.78890.8733
D0.79750.75040.8457
F10.84120.79820.8820
F20.81670.76540.8665
G0.80290.75930.8437
H0.84790.80690.8862
J0.80250.76040.8412
K0.77250.72320.8194
Table 27. Statistics of conserved regions detected in Vpu.
Table 27. Statistics of conserved regions detected in Vpu.
GenotypeConserved RegionsMean Conserved Region LengthSm in Conserved Regions
All sequences110.000.9541
A145.000.9873
A215.001.0000
A639.330.9799
B37.330.9778
C26.000.9687
D28.000.9687
F128.000.9642
F219.000.9766
G28.000.9553
H27.500.9783
J37.330.9315
K46.250.9270
Table 28. Number of analyzed sequences by region, protein, and genotype.
Table 28. Number of analyzed sequences by region, protein, and genotype.
SubtypeEnvGagNefPolRevTatVpu
A1111105106106105105106
A28989998
A6193196196194194195195
B199200199196197200200
C190197196189191201200
D93989799959395
F161626156556061
F21113101291310
G89919396969692
H100989799109896
J16221016161616
K12121216111111
All subtypes108311031085108898810971090
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Shchemelev, A.N.; Serikova, E.N.; Ostankova, Y.V.; Davydenko, V.S.; Ramsay, E.S.; Totolian, A.A. Identifying Conserved Regions in HIV-1 Proteins by Entropy Analysis of Sequence Variability. Int. J. Mol. Sci. 2026, 27, 5139. https://doi.org/10.3390/ijms27115139

AMA Style

Shchemelev AN, Serikova EN, Ostankova YV, Davydenko VS, Ramsay ES, Totolian AA. Identifying Conserved Regions in HIV-1 Proteins by Entropy Analysis of Sequence Variability. International Journal of Molecular Sciences. 2026; 27(11):5139. https://doi.org/10.3390/ijms27115139

Chicago/Turabian Style

Shchemelev, Alexandr N., Elena N. Serikova, Yulia V. Ostankova, Vladimir S. Davydenko, Edward S. Ramsay, and Areg A. Totolian. 2026. "Identifying Conserved Regions in HIV-1 Proteins by Entropy Analysis of Sequence Variability" International Journal of Molecular Sciences 27, no. 11: 5139. https://doi.org/10.3390/ijms27115139

APA Style

Shchemelev, A. N., Serikova, E. N., Ostankova, Y. V., Davydenko, V. S., Ramsay, E. S., & Totolian, A. A. (2026). Identifying Conserved Regions in HIV-1 Proteins by Entropy Analysis of Sequence Variability. International Journal of Molecular Sciences, 27(11), 5139. https://doi.org/10.3390/ijms27115139

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop