1. Introduction
Nuclear receptors comprise one of the largest and most extensively studied families of ligand-activated transcription factors in animals. These receptors regulate diverse physiological processes, including development, reproduction, metabolism, immune responses, cellular differentiation, and maintenance of homeostasis [
1]. The transcriptional activity of nuclear receptors is controlled through interactions with a complex network of coregulatory proteins that either enhance or repress gene expression [
2,
3]. Among these regulatory molecules, Steroid Receptor Coactivators (SRCs), also known as the p160 coactivator family, represent one of the most important groups of transcriptional coactivators involved in nuclear receptor signaling pathways [
3,
4].
The SRC family was first identified with the discovery of SRC-1, a protein capable of enhancing steroid hormone receptor-mediated transcriptional activation [
4]. Subsequent studies led to the identification of two additional paralogs, SRC-2 and SRC-3, collectively forming the p160 steroid receptor coactivator family [
5]. These proteins are encoded by the nuclear receptor coactivator genes
NCOA1,
NCOA2, and
NCOA3, respectively, and function as transcriptional integrators that bridge ligand-activated nuclear receptors with the basal transcription machinery and chromatin-remodeling complexes [
3,
6]. Through these interactions, SRC proteins facilitate transcriptional activation of target genes and contribute to the regulation of numerous physiological processes.
SRC proteins are characterized by a conserved modular architecture that underlies their biological functions. The N-terminal region contains a basic helix–loop–helix (bHLH) domain and Per–Arnt–Sim (PAS) domains that mediate protein–protein interactions, cellular localization, and transcriptional regulation. The central receptor interaction domain contains multiple LXXLL motifs, commonly referred to as nuclear receptor boxes, which directly interact with ligand-bound nuclear receptors [
7]. The C-terminal region contains activation domains that recruit additional coactivators, histone acetyltransferases, methyltransferases, and chromatin-remodeling complexes required for efficient transcriptional activation [
3,
5]. This modular organization enables SRC proteins to function as molecular scaffolds that coordinate multiple signaling pathways and transcriptional networks.
Although members of the SRC family share a common structural framework, extensive evidence indicates that they possess both overlapping and distinct biological functions. SRC-1 plays critical roles in reproductive physiology, neuroendocrine signaling, and hormone-responsive gene regulation [
8,
9]. Studies using knockout mouse models have demonstrated that SRC-1 contributes to fertility, mammary gland development, and behavioral responses mediated by steroid hormones [
5,
8]. SRC-2, also known as transcriptional intermediary factor 2 (TIF2) or glucocorticoid receptor-interacting protein 1 (GRIP1), is involved in metabolic regulation, energy homeostasis, glucose metabolism, and reproductive function [
5,
10]. In contrast, SRC-3, also known as amplified in breast cancer 1 (AIB1) or activator of thyroid and retinoid receptors (ACTR), has been strongly associated with cellular proliferation, developmental regulation, and oncogenic signaling pathways [
11,
12].
The physiological significance of SRC proteins extends far beyond their classical role as nuclear receptor coactivators. These proteins interact with numerous transcription factors and signaling molecules, including Nuclear Factor Kappa B (NF-κB), Activator Protein 1 (AP-1), Signal Transducer and Activator of Transcription (STAT proteins), Early Region 2 Binding Factor (E2F) transcription factors, and other regulators of cellular growth and differentiation [
6,
13,
14]. Consequently, SRC proteins influence a broad spectrum of biological processes, including embryonic development, tissue differentiation, immune regulation, inflammatory responses, and metabolic adaptation [
13,
14,
15]. Their ability to integrate signals from multiple pathways has established SRC proteins as central regulators of cellular physiology [
3].
Given their widespread regulatory functions, dysregulation of SRC proteins has been implicated in numerous human diseases. Elevated expression or aberrant activation of SRC family members has been reported in breast cancer, prostate cancer, ovarian cancer, endometrial cancer, and several other malignancies [
13,
14,
16]. SRC-3 has emerged as a prominent oncogenic coactivator whose amplification and overexpression contribute to tumor progression, metastasis, and therapeutic resistance [
11,
16]. Similarly, alterations in SRC-1 and SRC-2 have been associated with endocrine disorders, reproductive abnormalities, obesity, metabolic syndrome, and inflammatory diseases [
9,
13,
17]. These observations have stimulated considerable interest in SRC proteins as potential therapeutic targets and biomarkers for disease diagnosis and prognosis.
Despite considerable advances in understanding the evolution of vertebrate steroid receptor families [
18] and the molecular functions of SRC proteins [
5,
14], the evolutionary origin, structural diversification, and conservation of the SRC family across vertebrates remain incompletely understood. Most previous studies have focused on the molecular mechanisms, physiological functions, and pathological significance of SRC family members [
5,
6,
19], whereas comparatively little attention has been devoted to their evolutionary history, structural diversification, and genome-wide conservation across vertebrate lineages. Previous evolutionary investigations have largely examined steroid receptor signaling pathways and the functional specialization of nuclear receptor coactivators following gene duplication events [
5,
8,
18], rather than the evolutionary diversification of the SRC family itself.
Comparative genomics and phylogenetic analyses provide powerful approaches for reconstructing the evolutionary history of protein families by examining sequence conservation, domain architecture, conserved motifs, and phylogenetic relationships across diverse taxa [
19]. Such analyses have successfully elucidated the origin and diversification of several vertebrate regulatory protein families, including steroid receptors [
18]. Applying these approaches to the SRC family can provide valuable insights into the evolutionary mechanisms underlying the emergence of complex transcriptional regulatory systems and identify conserved structural domains and functional motifs associated with SRC-mediated transcriptional regulation.
However, comprehensive genome-wide studies integrating SRC protein identification, comparative domain architecture, conserved receptor-interaction motifs, amino acid conservation, structural analyses, and phylogenetic reconstruction across diverse vertebrate lineages remain limited. Furthermore, the evolutionary origin of the modular SRC protein architecture and the potential contribution of ancient protein domains to the emergence of modern SRC proteins have not been systematically investigated.
To address these knowledge gaps, the present study integrates genome-wide SRC identification, comparative domain architecture, LXXLL motif characterization, amino acid conservation, structural analyses, and phylogenetic reconstruction to investigate the origin, structural diversification, and evolutionary relationships of vertebrate SRC family proteins, thereby providing a comparative genomic framework for investigating the evolutionary origin, structural diversification, and functional evolution of vertebrate SRC proteins.
2. Results and Discussion
2.1. Evaluation of Domain-Screening Tools for SRC Protein Identification
The structural characteristics, conserved domains, and functional features of SRCs are well defined in the literature [
9,
17,
20]. These proteins are characterized by distinct domain architectures that enable their identification across diverse species. However, the large-scale identification of SRC proteins requires a domain-screening approach capable of distinguishing among the three closely related SRC paralogs. Therefore, before conducting genome-wide analyses, we compared three commonly used domain annotation tools to determine which most effectively differentiated the three well-characterized human SRC reference proteins used in this study. This comparison was intended to guide tool selection for the present analysis rather than to establish the overall performance of these programs for all protein families.
To address this, human SRC-1, SRC-2, and SRC-3 proteins (UniProt IDs: SRC-1: Q15788; SRC-2: Q15596; SRC-3: Q9Y6Q9) were used as reference sequences to evaluate the performance of different domain-based screening tools. Three widely used computational programs were assessed: the MOTIF Search tool (
https://www.genome.jp/tools/motif/, accessed on 23 March 2026), the Hidden Markov Model Scan (HMMSCAN) [
21,
22] implemented via the HMMER website [
23], and the National Center for Biotechnology Information (NCBI) Batch Web CD-Search Tool [
24] (
Figure 1).
Comparative analysis of the three domain annotation tools showed that, among the reference SRC-1, SRC-2, and SRC-3 proteins examined in this study, the NCBI Batch CD-Search Tool provided the clearest differentiation of the characteristic domain architectures associated with each paralog (
Figure 1). Therefore, this tool was selected for subsequent genome-wide SRC identification and classification. This comparison was performed solely to select the most suitable tool for the present study and should not be interpreted as a comprehensive benchmark of domain-screening software. According to the domain annotations generated by the NCBI Batch CD-Search Tool [
24], all three SRC proteins share a common core of four conserved domains, reflecting their evolutionary relatedness. In addition, SRC-2 and SRC-3 share three additional domains, further highlighting their closer relationship with each other compared with SRC-1.
SRC-1 exhibited the most complex domain architecture among the proteins analyzed, containing five unique domains and an additional pfam07469 domain, which may contribute to its broader functional versatility (
Figure 1). In contrast, SRC-2 and SRC-3 each possess one unique domain, consistent with their more specialized and lineage-restricted roles. Notably, the reliability and applicability of the NCBI Batch CD-Search Tool [
24] for large-scale transcription factor family analysis have been demonstrated in recent studies, including its successful use in the identification and classification of NF-κB transcription factor family members [
25].
2.2. Distribution and Taxonomic Representation of SRC Family Proteins Across Vertebrates
Genome-wide analysis of canonical SRC family proteins (SRC-1, SRC-2, and SRC-3) retrieved from the UniProt database revealed distinct distribution patterns among currently available vertebrate protein records (
Figure 2 and
Tables S1–S3).
A total of 298 SRC proteins were identified, comprising 160 SRC-1 proteins, 43 SRC-2 proteins, and 95 SRC-3 proteins, indicating that SRC-1 represented the largest number of retrieved protein records, followed by SRC-3, whereas SRC-2 represents the least abundant paralog (
Figure 2 and
Tables S1–S3). The distribution analysis further suggested that all three SRC paralogs are associated with vertebrate lineages, particularly Mammalia and Aves. However, SRC-3 exhibited the broadest taxonomic spread, extending into Actinopteri, Amphibia, Chondrichthyes, and Sarcopterygii (
Figure 2 and
Tables S1–S3). The broader taxonomic representation of SRC-3 compared with SRC-1 and SRC-2 highlights differences in the distribution patterns of SRC family members across vertebrate lineages, indicating considerable variation in their taxonomic representation.
The observed distribution patterns are consistent with the known biological functions of SRC family members. SRC-1 was the most abundant paralog identified in this study and was particularly well represented in mammalian species. Previous studies have demonstrated that SRC-1 participates in steroid hormone signaling, reproductive physiology, and neuroendocrine regulation, functions that are especially prominent in vertebrates with complex endocrine systems [
4,
8]. The widespread occurrence of SRC-1 across mammalian taxa therefore reflects its important role in regulating hormone-responsive transcriptional networks.
SRC-2 exhibited the lowest abundance among the three paralogs and was detected in a more restricted range of vertebrate taxa. Functional studies have shown that SRC-2 is involved in metabolic regulation, energy homeostasis, and reproductive physiology [
27,
28]. The relatively limited distribution of SRC-2 observed in this study may reflect greater functional specialization than that of other members of the SRC family.
In contrast, SRC-3 was identified across a broader range of vertebrate groups, including fishes, amphibians, birds, reptiles, and mammals. SRC-3 has been implicated in developmental regulation, cellular growth, and transcriptional activation of numerous signaling pathways [
16,
29]. Its extensive distribution suggests that SRC-3 participates in fundamental biological processes that are widely conserved across vertebrates.
2.3. Comparative Domain Architecture of SRC Family Proteins
Domain architecture analysis of SRC-1, SRC-2, and SRC-3 proteins suggested a high degree of structural conservation across the SRC family while also revealing notable paralog-specific differences in domain organization (
Figure 3). All three SRC proteins retained the characteristic modular arrangement associated with the p160 nuclear receptor coactivator family, including conserved bHLH and PAS domains, central receptor interaction regions, and C-terminal transcriptional activation domains (
Figure 3). The conservation of these domains across SRC paralogs suggests that the fundamental molecular functions of nuclear receptor recognition and transcriptional coactivation have been evolutionarily maintained. However, clear differences in the number, arrangement, and extension of specific domains were observed among the paralogs, suggesting functional divergence following ancestral gene duplication events (
Figure 3). Differences in domain architecture compared with human SRC family members include the presence of Smart00091 (PAS) domains in one of the SRC-1 members, which are common to both SRC-2 and SRC-3 (
Figure 3A). Four SRC-2 members have the CL26621 domain, which is characteristic of SRC-1 members (
Figure 3A). Only 3 SRC-3 members have the domain architecture as human SRC-3 proteins, and the rest of the SRC-3 members (92) have the CL26621 domain, which is characteristic of SRC-1 members (
Figure 3A). Overall, SRC-1 displayed the most structurally elaborate architecture, with multiple conserved domains distributed throughout the protein, whereas SRC-2 exhibited a comparatively reduced, compact domain organization. SRC-3 retained a moderately complex architecture that shared structural features with both SRC-1 and SRC-2, indicating a combination of conserved and lineage-specific domain characteristics within this paralog (
Figure 3A). Overall, six domains were shared among all SRC family members, whereas two additional domains were common only to SRC-2 and SRC-3 (
Figure 3B). SRC-1 contained four unique domains together with the additional pfam07469 domain (
Figure 3B). These differences in domain composition highlight the structural diversification that has occurred within the SRC family while preserving a conserved core architecture required for coactivator function.
2.4. Evolutionary Conservation and Functional Variation of LXXLL Nuclear Receptor Interaction Motifs in SRC Family Proteins
Analysis of SRC proteins revealed that SRC-1, SRC-2, and SRC-3 possess multiple LXXLL motifs distributed throughout their sequences (
Figure 4A and
Tables S4–S6). The number of LXXLL motifs in SRC-1 proteins ranged from 3–8 (22 SRC-1 had 8 motifs, 84 SRC-1 had 7 motifs, 53 SRC-1 had 6 motifs and one SRC-1 had only 3 motifs); in SRC-2 proteins the motif number ranged from 3–6 (one SRC-3 had 6 motifs, 33 SRC-3 had 4 motifs and 9 SRC-3 had 3 motifs); and in SRC-3 proteins the motif number ranged from 5–8 (3 SRC-3 had 8 motifs; one SRC-3 had 7 motifs; 79 SRC-3 had 6 motifs and 12 SRC-3 had 5 motifs) (
Figure 4A and
Tables S4–S6). Analysis of motif amino acids revealed the presence of a specific pattern of amino acids in these motifs across all the SRCs, where in some motifs a single pattern was present, and in some motifs more than one pattern of amino acid sequences was observed (
Figure 4A).
Seven motif sequence patterns at relevant positions are shared among all three SRC families (
Figure 4B), suggesting the presence of highly conserved ancestral receptor-interaction sites that have been maintained throughout vertebrate evolution. Five motif sequence patterns were shared between SRC-1 and SRC-3 (
Figure 4B). In contrast, additional motifs were found exclusively within individual SRC paralogs. SRC-1 contained 3 unique motif patterns, whereas SRC-2 and SRC-3 exhibited 3 and 7 distinct motif patterns (
Figure 4B). The presence of both shared and lineage-specific LXXLL motif patterns suggests that, following duplication of an ancestral SRC gene, the paralogs retained a conserved core set of receptor-binding motifs while simultaneously acquiring novel interaction motifs through sequence diversification. Such diversification may have expanded the spectrum of nuclear receptors and transcriptional complexes that each SRC protein could engage, thereby promoting functional specialization. Similar evolutionary patterns have been reported in other nuclear receptor coactivator families, in which duplication and modification of receptor-interaction motifs have contributed to increased regulatory complexity and tissue-specific functions [
3,
30].
The distribution of motifs further suggests differential evolutionary pressures among the SRC paralogs. SRC-3 appears to possess the greatest motif complexity (19 LXXLL motif patterns), including both conserved and unique LXXLL motifs, which may explain its broad functional involvement in development, growth signaling, and oncogenic pathways. SRC-1 exhibits a combination of conserved motifs and several unique motif patterns (3), consistent with its prominent role in steroid hormone signaling and reproductive regulation. Conversely, SRC-2 appears to retain a more restricted repertoire of motifs, supporting previous observations that it is evolutionarily more conserved and functionally specialized for metabolic regulation and energy homeostasis. Collectively, these findings suggest that the evolution of SRC proteins involved both the preservation of ancestral nuclear receptor interaction motifs and the acquisition of paralog-specific LXXLL motifs, suggesting one possible molecular mechanism for the functional diversification of the vertebrate SRC/p160 coactivator family.
The observed diversity of LXXLL motif patterns suggests that SRC family members possess distinct capacities to interact with nuclear receptors and associated transcriptional complexes. The coexistence of highly conserved motifs and paralog-specific motif variants suggests that the fundamental mechanism of receptor recognition has been maintained throughout evolution while allowing flexibility in receptor specificity and regulatory interactions. Such diversification may contribute to differences in transcriptional activity, tissue specificity, and signaling pathway utilization among SRC paralogs. These findings further support the importance of LXXLL motifs as key structural determinants of SRC-mediated transcriptional regulation.
2.5. Structural Conservation and Sequence Variation of SRC Family Proteins
The domain architectures and LXXLL motif compositions described above provide insight into the functional organization of SRC proteins. To determine whether these conserved structural features are also reflected at the sequence level, amino acid conservation patterns among SRC family members were subsequently investigated using Profile Multiple Alignment with predicted Local Structures and 3D constraints (PROMALS3D) [
31], which integrates sequence and structural information to provide a robust assessment of conservation patterns. The analysis revealed distinct differences in amino acid conservation among SRCs (
Table 1). Analysis of all SRC proteins revealed 143 completely conserved amino acids (score 9) across three different family members. Comparative analysis of individual members revealed the highest amino acid conservation in SRC-2 (500 amino acids), followed by SRC-1 (306 amino acids) and SRC-3 (235 amino acids). These observations suggest that the conservation patterns reflect intrinsic evolutionary properties of each SRC family rather than differences in sequence representation. This high amino acid conservation suggests that SRC-2 has been subject to greater evolutionary constraints, indicating the preservation of essential structural and functional features throughout vertebrate evolution. In contrast, SRC-1 and SRC-3 contained fewer highly conserved amino acids than SRC-2, indicating comparatively greater sequence variability that may contribute to differences in structural flexibility and biological function. Overall, the findings suggest that SRC-2 is the most structurally conserved member of the SRC family, retaining ancestral sequence features important for its biological function.
To visualize the spatial distribution of conserved amino acid residues and secondary structural elements, predicted three-dimensional models of human SRC-1, SRC-2, and SRC-3 were obtained from the AlphaFold3 database [
32] (
Figure 5). Because SRC proteins are known to contain extensive intrinsically disordered regions, these models were used primarily to visualize local secondary-structure features rather than to interpret detailed full-length tertiary-structure information.
Structural analysis revealed that SRCs 1–3 proteins are predominantly composed of intrinsically disordered regions, characterized by abundant loops, with only limited regions forming defined secondary structural elements, such as α-helices and β-sheets (
Figure 5). This structural organization is consistent with the known properties of transcriptional coactivators, which often exhibit structural flexibility to facilitate interactions with multiple binding partners. However, the predicted three-dimensional models exhibited relatively low overall confidence, with predicted template modeling (pTM) and interface predicted template modeling (ipTM) scores ranging from 0.23 to 0.28. These low confidence values are consistent with the intrinsically disordered nature of SRC proteins and therefore limit the reliability of the predicted full-length tertiary structures. Consequently, the structural interpretation presented in this study is limited to local secondary-structure elements and conserved domains. These limitations likely stem from the lack of experimentally resolved crystal structures of SRCs and are consistent with the intrinsically disordered nature of SRC proteins, which substantially reduces the confidence of full-length structural predictions despite recent advances in AlphaFold3 modelling [
32]. Despite the relatively low confidence of the full-length structural models, AlphaFold predictions were considered sufficiently reliable for the limited purpose of visualizing local secondary structural features, including α-helices and β-sheets, within conserved regions of the proteins. Consequently, the structural interpretation presented in this study is restricted to these higher-confidence local structural features rather than the overall three-dimensional architecture.
To further investigate the distribution of these secondary structural elements across SRCs 1–3 members, a multiple sequence alignment (MSA) of these proteins was performed, and the secondary structural elements were mapped alongside conserved amino acids and domains (
Figure S1). The MSA showed that several predicted α-helices and β-sheets were conserved among SRC family members and frequently coincided with conserved domains and highly conserved amino acid residues (
Table 1 and
Figure S1). Notably, all 306 highly conserved amino acids identified in SRC-1 were located within conserved domains, several of which corresponded to predicted α-helical and β-sheet regions (
Table 1 and
Figure S1). In contrast, only 50% of the conserved amino acids were found to be part of different domains in SRC-2 and SRC-3 (
Table 1 and
Figure S1).
Although AlphaFold3 [
32] provides valuable structural predictions for proteins lacking experimentally determined structures, the extensive intrinsically disordered regions present in SRC proteins limit the confidence of full-length structural models. Therefore, the present analysis should be interpreted primarily as a visualization of conserved secondary-structure features and domain organization, rather than as definitive evidence of the complete tertiary structure. Future experimental structural studies will be required to validate these predicted conformations.
2.6. Phylogenetic Relationships and Evolutionary History of SRC Family Proteins
Phylogenetic analysis of SRCs 1–3 revealed clear clustering of the three paralogous protein families, suggesting distinct evolutionary trajectories across vertebrate lineages (
Figure 6). The phylogenetic tree showed that SRC-1, SRC-2, and SRC-3 form separate monophyletic clades, supporting the hypothesis that the modern SRC family arose through ancient gene duplication events early during vertebrate evolution. Among the three paralogs, SRC-3 displayed the broadest phylogenetic distribution, containing representatives from Chondrichthyes, Actinopteri, Amphibia, Aves, Sarcopterygii, and Mammalia, whereas SRC-1 and SRC-2 were comparatively enriched in mammalian taxa as described in
Section 2.2.
The extensive distribution of SRC-3 across basal and derived vertebrate groups suggests that this coactivator has been broadly retained throughout vertebrate evolution and across diverse vertebrate lineages. SRC-3 sequences from cartilaginous fishes, ray-finned fishes, coelacanths, amphibians, birds, and mammals formed a large and deeply branching clade, suggesting ancient functional importance. This observation agrees with previous studies showing that members of the p160/SRC family originated before vertebrate diversification and subsequently evolved specialized physiological functions while maintaining conserved nuclear receptor interaction domains [
3,
5]. The retention of SRC-3 across highly divergent taxa suggests that this paralog performs fundamental regulatory functions that were required throughout vertebrate evolution. Furthermore, the broader phylogenetic distribution and deeper branching topology of SRC-3 suggests that this paralog has retained evolutionary features that have persisted across a wide range of vertebrate lineages.
SRC-1 exhibited strong diversification within Mammalia, particularly among primates, carnivores, cetaceans, bats, rodents, and marsupials. Although avian SRC-1 sequences were also present, mammalian sequences dominated this clade, suggesting extensive lineage-specific expansion and adaptation after the mammalian radiation. SRC-1 is known to function prominently in steroid hormone signaling, reproductive biology, and neuroendocrine regulation [
4,
8]. The high abundance and diversification of mammalian SRC-1 sequences observed in the tree may therefore reflect adaptive evolution associated with increasingly complex endocrine and developmental systems in mammals. In addition, the clustering of closely related mammalian taxa within SRC-1 suggests relatively recent diversification events superimposed on an ancient conserved framework.
SRC-2 formed a comparatively smaller and more compact clade dominated by mammals, with fewer representatives from birds, amphibians, and reptiles. This restricted phylogenetic spread suggests stronger evolutionary constraints on SRC-2 than on SRC-1 and SRC-3. Previous studies have demonstrated that SRC-2 plays major roles in energy metabolism, reproductive physiology, and systemic metabolic homeostasis [
27,
28]. Because metabolic regulation is highly conserved across vertebrates, the relatively compact topology of SRC-2 may reflect functional conservation and slower evolutionary divergence. The reduced taxonomic breadth of SRC-2 relative to SRC-3 is consistent with greater functional specialization following duplication of an ancestral SRC-like gene.
Collectively, the phylogenetic topology supports an evolutionary scenario in which an ancestral SRC-like coactivator existed before vertebrate diversification and subsequently gave rise to the modern SRC family through ancient gene duplication events. Following duplication, SRC-1 underwent substantial expansion and diversification, particularly within mammals, whereas SRC-2 and SRC-3 followed distinct evolutionary trajectories characterized by differences in taxonomic representation and lineage-specific divergence. Together, these findings suggest that gene duplication, divergence, and functional specialization were major evolutionary forces that shaped the modern SRC family and may have contributed to the increasing complexity of vertebrate nuclear receptor signaling networks.
Although phylogenetic analysis provides comparative evidence for the duplication and divergence of the modern SRC paralogs, it does not reveal how the ancestral SRC architecture originally arose. To address this question, a broader evolutionary investigation was conducted by examining the occurrence and combinations of SRC-associated domains across the domains of life. This approach enabled reconstruction of the potential evolutionary steps that preceded the emergence of the canonical vertebrate SRC proteins.
2.7. Evolutionary Origin and Structural Diversification of the SRC Family
The combined evidence from domain architecture, motif composition, amino acid conservation, and phylogenetic analyses supports a common evolutionary origin for the SRC family and provides a framework for reconstructing its emergence across the domains of life. To investigate the evolutionary origin and emergence of SRCs across the domains of life, a comprehensive domain architecture analysis was conducted using more than 20,000 initially identified protein hits. Among these, over 4000 proteins contained one or more domains characteristic of SRC proteins, in addition to the canonical SRC-1, SRC-2, and SRC-3 proteins identified in this study (
Figure 7 and
Table S7). The analysis revealed a complex evolutionary history possibly involving progressive domain acquisition, recombination, and conservation across diverse lineages.
Among the three life domains, the earliest SRC-associated signature identified was the Smart00091 domain, detected in archaeal proteins and conserved across all three vertebrate SRC paralogs (
Figure 7). In Bacteria, proteins containing Smart00091 were identified alongside additional domains, including cd00130 and cl37882, both characteristic of modern SRC-1 proteins. Similarly, fungal proteins exhibited combinations of Smart00091 and cd00130, suggesting that these domains were retained and propagated during the transition from prokaryotic to lower-eukaryotic lineages. The occurrence of these domains across Archaea, Bacteria, and Fungi suggests that several structural components of modern SRC proteins predate the emergence of multicellular animals and likely served as evolutionary building blocks for the subsequent development of the SRC family.
A substantial increase in the complexity of SRC-associated domains was observed in invertebrates. Six characteristic SRC domains, including Smart00091, cd00130, cl37882, cd18949, cl26621, and Pfam14598, were identified in various combinations (
Figure 7). Proteins containing only a single SRC-associated domain, as observed in Archaea, Bacteria, and Fungi, were also detected in invertebrates, suggesting the inheritance of ancestral domain architectures followed by gradual structural elaboration. In addition, numerous proteins containing combinations of two or three SRC-associated domains were identified across multiple invertebrate phyla (
Figure 7). The widespread occurrence of these domain arrangements across taxonomically diverse invertebrates may represent ancestral modules that were successfully retained during speciation and subsequently served as substrates for further evolutionary innovation.
The greatest diversification of SRC domain architectures was observed within chordates. Analysis revealed numerous proteins with distinct combinations of SRC-associated domains that differed from those of the canonical vertebrate SRC-1, SRC-2, and SRC-3 proteins. Notably, the remaining eight characteristic SRC domains were detected exclusively in chordate lineages, indicating that the assembly of complete SRC architectures occurred relatively late during metazoan evolution. The presence of multiple intermediate domain combinations in chordates suggests that these proteins may represent transitional evolutionary forms that preceded the emergence of the canonical vertebrate SRC paralogs. Furthermore, the appearance of canonical SRC structural arrangements coincided with the emergence of vertebrate lineages, supporting the hypothesis that extensive domain shuffling and recombination preceded the formation of the contemporary SRC family.
Interestingly, several vertebrate groups contained proteins harbouring characteristic SRC domains despite the apparent absence of specific SRC paralogs (
Figure 7 and
Table S7). For example, although SRC-2 proteins were not identified in Actinopteri, proteins containing the SRC-2-associated domain cd18950 were detected. Similarly, Amphibia lacked identifiable SRC-1 proteins but contained proteins possessing characteristic SRC-1 domains. Likewise, Chondrichthyes lacked detectable SRC-1 and SRC-2 proteins, yet proteins carrying domains diagnostic of both paralogs were present. These observations suggest that key SRC structural modules may have evolved before the emergence of the complete SRC proteins. Two possible explanations could account for this observation: either these vertebrate lineages diverged before the complete assembly of specific SRC paralogs, or the corresponding SRC proteins were subsequently lost or remain undiscovered. However, given the extensive dataset analyzed in this study, comprising more than 20,000 protein sequences, the former explanation appears more consistent with the extensive taxonomic sampling employed here.
Overall, the results suggest that the evolution of SRC proteins was not a simple linear process but rather involved the gradual accumulation, rearrangement, and recombination of ancestral domains across the domains of life. The presence of SRC-associated domains in Archaea, Bacteria, Fungi, Invertebrates, and Vertebrates (Chordates) is consistent with a stepwise evolutionary model in which ancient protein modules were progressively combined to generate increasingly complex architectures. This process might have played a role in the emergence of the vertebrate SRC-1, SRC-2, and SRC-3 proteins, which subsequently diversified to fulfil specialized roles in nuclear receptor signaling and transcriptional regulation.
When considered together, the domain architecture, LXXLL motif composition, amino acid conservation, and phylogenetic analyses presented in this study support a model of progressive SRC evolution. Conserved domains and receptor-interaction motifs may represent ancient ancestral modules that were gradually assembled through domain acquisition and recombination events. Subsequent gene duplication and lineage-specific evolution gave rise to the modern SRC-1, SRC-2, and SRC-3 paralogs, each retaining a conserved functional core while acquiring distinct structural, regulatory, and functional characteristics. This integrated evolutionary framework provides new insight into the emergence and functional specialization of the vertebrate SRC family.
4. Conclusions
This study presents a comparative genomic framework of the steroid receptor coactivator (SRC) family across the domains of life. Using a combination of genome-wide screening, domain architecture analysis, phylogenetic reconstruction, motif characterization, and amino acid conservation analyses, the evolutionary history of SRC proteins was systematically investigated. The results demonstrated that the NCBI Batch CD-Search Tool was suitable for identifying and classifying the reference SRC proteins examined in this study and was therefore used for subsequent genome-wide analyses. Although more than 20,000 proteins were initially examined during genome-wide data mining, only proteins that satisfied the predefined conserved-domain criteria were classified as canonical SRC family members, thereby minimizing the inclusion of unrelated proteins despite potential differences in database representation.
Domain-based evolutionary reconstruction suggested that the ancestral SRC architecture likely emerged through the progressive accumulation, recombination, and structural refinement of pre-existing protein domains traceable to lineages predating the emergence of vertebrates. The presence of SRC-associated domains across Archaea, Bacteria, Fungi, Invertebrates, and Chordates supports a stepwise evolutionary model in which ancient protein modules were gradually assembled into increasingly complex architectures. Following the establishment of this ancestral SRC framework, phylogenetic analyses suggested that ancient vertebrate gene duplication events may have given rise to the modern SRC paralogs. Subsequent lineage-specific evolution resulted in distinct taxonomic distributions, structural characteristics, and functional properties among SRC-1, SRC-2, and SRC-3. Although full-length protein sequences provide an overview of the evolutionary relationships among SRC paralogs, these proteins contain extensive intrinsically disordered regions that may reduce phylogenetic resolution. Future studies employing conserved-domain-only phylogenies together with orthology-based analyses may provide additional insights into the evolutionary history of the SRC family.
Comparative analyses of LXXLL nuclear receptor interaction motifs revealed extensive conservation of core receptor-binding motifs, alongside the emergence of paralog-specific motif patterns that may have contributed to functional specialization. Structural analyses further suggested that, despite extensive diversification, members of the SRC family retain highly conserved domain organization, secondary structural elements, and amino acid residues possibly associated with essential biological functions. Notably, SRC-2 exhibited the highest amino acid conservation, suggesting stronger evolutionary constraints than SRC-1 and SRC-3.
Collectively, these findings suggest an evolutionary scenario in which ancestral protein modules progressively assembled into an early SRC architecture that subsequently evolved through gene duplication, structural innovation, and functional specialization. Integrating domain architecture, motif composition, sequence conservation, and phylogenetic evidence provides a unified framework for understanding the emergence and evolutionary history of the SRC family across the domains of life. Although the widespread occurrence of SRC-associated domains across bacteria, archaea, fungi, invertebrates, and vertebrates is consistent with a stepwise domain acquisition model, these observations alone do not establish a direct evolutionary pathway leading to modern SRC proteins. Alternative evolutionary scenarios, including convergent evolution, independent domain shuffling, repeated domain reuse, incomplete lineage sampling, and annotation bias, may also explain the observed distribution of individual domains. Therefore, the evolutionary model proposed in this study should be regarded as a hypothesis, supported by comparative genomic observations, that warrants further validation through reciprocal homology analyses, domain-specific phylogenies, and orthology-based evolutionary reconstruction.
Furthermore, because this study is based on publicly available protein records, the observed distribution patterns should be interpreted in the context of current database coverage. Differences in genome sequencing, annotation quality, and taxonomic sampling may influence the number of proteins retrieved for individual vertebrate groups. Nevertheless, all proteins included in this study satisfied the same domain-based identification criteria, providing a consistent framework for comparative evolutionary analyses.