Next Article in Journal
Extracellular Molecular Repertoire of Xerotolerant Actinobacteria Colonizing Serpentinite Rocks
Previous Article in Journal
Unfolding Immune Dysregulation in COPD: Identification of a Three-Gene Signature and Functional Validation of TCF7 in Human Lung Tissue and T Lymphocytes
Previous Article in Special Issue
A Comparative Review of Artificial Intelligence Applications in Small Molecule Versus Peptide Drug Discovery
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Mapping of Phenotype Specific Host–Microbiome Protein–Protein Interaction Networks in Colorectal Cancer Using Deep Learning

by
Despoina P. Kiouri
1,
Georgios C. Batsis
2,
Ippokratis Messaritakis
3,4,
John Souglakos
3,5 and
Christos T. Chasapis
1,*
1
Laboratory of Organic Chemistry, Department of Chemistry, National and Kapodistrian University of Athens, 15772 Athens, Greece
2
AI Lab, Department of Digital Systems, University of Piraeus, Gr. Lampraki 126, 18534 Piraeus, Greece
3
Laboratory of Translational Oncology, Medical School, University of Crete, 70013 Heraklion, Greece
4
Laboratory of Clinical Microbiology, German Medical Institute, Yiannoukas Labs Ltd., Bioiatriki Group, Limassol 4105, Cyprus
5
Laboratory of Pathology, University General Hospital of Heraklion, 70013 Heraklion, Greece
*
Author to whom correspondence should be addressed.
Int. J. Mol. Sci. 2026, 27(10), 4232; https://doi.org/10.3390/ijms27104232
Submission received: 14 March 2026 / Revised: 26 April 2026 / Accepted: 6 May 2026 / Published: 9 May 2026
(This article belongs to the Special Issue New Horizons in Structure and AI-Based Drug Design)

Abstract

Colorectal cancer (CRC) pathogenesis is driven by complex protein–protein interactions (PPIs) between the host and the gut microbiome, yet these molecular dialogs remain largely unmapped. This study utilizes a Deep Learning framework, enhanced by protein structure embeddings, to predict approximately 8.9 billion interspecies PPIs from clinical metagenomic data. The model achieved high accuracy with an AUROC of 0.9960, identifying a high-confidence interactome representing roughly 16% of evaluated protein pairs. Phenotype-specific analysis revealed that while microbial hubs shift—transitioning from metabolic enzymes in healthy states to transport and regulatory proteins in CRC—the primary human targets remain remarkably consistent across both cohorts. These core human interactors are predominantly metalloproteins and regulators of ubiquitination, apoptosis, and zinc transport, suggesting these pathways are primary focal points for microbial manipulation regardless of disease state. Furthermore, co-occurring bacterial genera exhibit over 99% overlap in host target profiles, indicating significant functional redundancy in microbial engagement with the host. These findings suggest that CRC probably arises from network-level perturbations of stable host signaling hubs, offering a blueprint for identifying novel therapeutic targets and biomarkers.

1. Introduction

Hippocrates, the father of medicine, famously stated that “All disease begins in the gut,” a prescient observation that reflects the centrality of intestinal health to systemic well-being [1]. Over the past two decades, advances in sequencing technologies have underscored the importance of the gut microbiome (GM), a densely populated microbial ecosystem comprising bacteria, viruses, fungi, archaea, and protozoa [2]. Among these, bacteria have been most extensively characterized, with human microbiome projects identifying over 2000 species spanning multiple phyla, including Firmicutes, Bacteroidetes, Actinobacteria, Proteobacteria, and Verrucomicrobia [3]. These organisms play critical roles in nutrient metabolism, immune system development, and epithelial barrier function, earning the microbiome the designation of a “super-organ” [4,5,6,7].
The balance of microbial populations is crucial for maintaining homeostasis. Factors such as pH, bile acids, oxygen tension, and host immune responses shape microbial communities along the gastrointestinal tract [8]. When this balance is disrupted, a state of dysbiosis arises, which has been associated with numerous conditions, ranging from inflammatory bowel disease and irritable bowel syndrome [9] to metabolic disorders [10], cardiovascular diseases [11], and autoimmune conditions [12]. Dysbiosis also affects the gut–brain axis, a bidirectional communication system linking the enteric and central nervous systems, and has been implicated in neurodevelopmental [13] and psychiatric conditions [14] such as autism spectrum disorder, major depression, and Parkinson’s disease [15,16]. These diverse links highlight the microbiome’s wide-reaching influence across human physiology and pathology.
Within this broader landscape, colorectal cancer (CRC) has emerged as one of the most compelling examples of a disease with direct microbial involvement. CRC is the second most common cancer and the third leading cause of cancer-related death worldwide [17]. The rising incidence of early-onset CRC, diagnosed before the age of 50, further emphasizes the urgency of understanding non-genetic drivers of disease. Importantly, more than 80% of CRC cases are sporadic and arise from complex interactions between host genetics, environmental factors, and the gut microbiome rather than inherited mutations [18]. The microbiome is therefore positioned not only as a marker of CRC risk but also as a potential mediator of disease onset and progression.
Specific bacterial taxa have been consistently associated with CRC. Pathogens such as Helicobacter pylori [19], Salmonella [18], and Campylobacter jejuni [20] contribute to carcinogenesis via chronic inflammation, toxin production, and genotoxic effects. Other taxa, including Bacteroides fragilis, Escherichia coli, and Streptococcus gallolyticus [21], participate in biofilm formation and secrete metabolites that modulate epithelial signaling pathways [22]. Fusobacterium nucleatum [23] and Porphyromonas gingivalis, in particular, have been linked to chemoresistance in CRC, underscoring the clinical impact of microbial activity on treatment outcomes [24]. Collectively, these findings suggest that CRC is not merely influenced by microbial imbalance but may in fact be driven by direct microbial contributions.
The molecular mechanisms that underpin these associations are increasingly recognized as pivotal for advancing both diagnostics and therapy. In recent years, research has shifted from taxonomic descriptions to mechanistic investigations at the molecular level. Protein–protein interactions (PPIs) between microbial proteins and host proteins represent one such mechanism. These interactions can alter host signaling networks, immune responses, and cellular processes in ways that favor tumor initiation or progression. While experimental approaches such as yeast two-hybrid assays, cross-linking mass spectrometry, and photo-reactive probes have shed light on microbe–host interactions [25,26,27,28,29,30,31], the scale of experimental mapping remains limited, leaving large portions of the interactome unexplored. A critical limitation in our current understanding is the paucity of experimental evidence detailing protein interactions between gut bacteria and the human host. While general human–bacterial interaction data exist in public repositories, the experimentally determined pan-human–bacterial protein interactome consists of fewer than 20,000 interactions. This limited scope, especially when contrasted with the 300–500 bacterial species estimated to inhabit the human gut, creates a research deficit that likely obstructs a deeper comprehension of how GM–host imbalances drive disease, highlighting that these molecular dialogs are largely unmapped.
Consequently, computational methods—including domain–domain interaction modeling and machine learning algorithms—have become indispensable for predicting host–microbe PPIs [32,33,34,35,36].
In this context, this work introduces a computational approach to map these crucial molecular dialogs. Next-Generation Sequencing (NGS) was applied to meticulously profile bacterial strain abundances within the gut microbiomes of both healthy individuals and colorectal cancer patients. Building on this detailed community characterization, we then applied a novel Deep Learning (DL) algorithm specifically designed to PPIs between proteins from the identified bacterial strains and proteins of the human host’s gut. The objective of this interdisciplinary study is to uncover potential protein interaction networks that could explain how specific bacterial strains contribute to, or protect against, colorectal cancer development through direct molecular engagement with the host.

2. Results

2.1. Model Re-Training Results on Test PPIs

The training procedure followed the same optimization strategy and early stopping protocol as our previous framework [37], with convergence achieved at the thirteenth epoch. The updated model, incorporating ProstT5-based embeddings, was evaluated on a held-out test set of 3,338,020 protein pairs (204,564 positive; 3,133,456 negative). Despite the pronounced class imbalance, the model demonstrated strong generalization and classification performance.
Evaluation metrics (Table 1) confirm high overall accuracy, with a macro-averaged F1-score of 0.9828 and strong performance on the interacting class (F1 = 0.9676). Imbalance-aware metrics, including a Matthews Correlation Coefficient of 0.9657 and AUROC of 0.9960 (Figure 1), further attest to the model’s robustness. The confusion matrix (Table 2) shows minimal false positive and false negative rates, while the recall at a minimum precision threshold of 0.5 reached 0.9902, validating the model’s utility for high-confidence interaction prediction.

2.2. Host–Microbiome PPI Network

The re-trained DL model was applied to systematically predict PPIs between human and bacterial proteins across all identified gut bacterial species from the clinical data. Using the microbiome profiles derived from the study samples, 904,554 bacterial proteins were paired exhaustively with a reference set of 9864 human proteins. This yielded a total of approximately 8.9 billion potential interspecies protein pairs. Interaction probabilities were estimated for each pair using the calibrated prediction model. To ensure high confidence in reported interactions, we applied a posterior probability threshold of 1.0. After filtering, a total of 1,481,068,857 interactions were retained as high-confidence predictions, representing approximately 16% of the evaluated protein pairs. These interactions constitute the inferred host–microbiome interactome used in downstream analyses. Detailed examples of PPIs involving the highest-degree human and bacterial proteins are documented in Supplementary Files S7 and S8, respectively.

2.3. Microbial Phenotype Association with Disease State

To investigate the relationship between microbial genera and disease status, bacterial genera were systematically assigned to CRC-associated, health-associated, or unassigned (“Both”) categories based on their differential abundance between healthy and CRC cohorts. Out of the 204 evaluated genera, 48 were classified as CRC-associated, 40 as health-associated, and 116 showed no clear association (Table 3). These phenotype labels were subsequently used to stratify predicted PPIs by microbial origin.
CRC-associated data exhibited significant enrichment in the CRC cohort with genera such as Peptoniphilus, Finegoldia, Porphyromonas, and Fusobacterium demonstrating highly significant q-values (e.g., Peptoniphilus: q = 2.38 × 10−37) (Table 4).
Conversely, health-associated genera included well-known beneficial taxa such as Faecalibacterium (q = 2.91 × 10−13), Subdoligranulum (q = 1.94 × 10−9), and Bifidobacterium (q = 4.97 × 10−5). These genera were significantly enriched in healthy individuals (Table 5).

2.4. Phenotype-Specific Networks

To investigate the distinct interspecies PPI landscapes associated with health and disease, the predicted PPI network was stratified based on bacterial phenotype assignment. Specifically, two separate subnetworks were constructed: one for Health, and another for CRC-associated bacteria, by filtering interactions where each bacterial protein was annotated with a corresponding phenotype.
The Health-associated network comprised 1,258,444,460 predicted PPIs, while the CRC-associated network included 1,087,290,486 interactions. Subsequently, the proteins were also ranked by degree centrality (i.e., the number of unique interaction partners) to identify the most central host and microbial proteins in each phenotype-specific context. Results are presented in Table 6 and Table 7 for health and CRC-associated proteins, accordingly. The complete table with all the host and bacterial proteins participating in the predicted networks of both disease states can be found in Supplementary Materials File S1.

2.5. Analytical Protocol for the Evaluation of Molecular Network Architectures

To investigate the alignment between microbial co-occurrence and functional interactions with the human host, a comparative analysis integrating co-occurrence patterns with DL-based interspecies PPI predictions was performed. Bacterial genera participating in the same co-occurrence sub-network were cross-mapped to their corresponding host protein predicted interactors.
The overlap in host interaction profiles among bacterial pairs was assessed by computing the proportion of shared human protein interactors. A total of 126 CRC-associated bacterial pairs were analyzed. The distribution of human host interaction overlap revealed a bimodal pattern: most pairs shared >99% of human targets, while a minority showed no overlap. The mean percentage of shared targets was 84.9% (±34.8%), with a median of 99.08%.
In the health-associated cohort, 196 bacterial pairs were examined. The overlap distribution mirrored the CRC pattern, with an average similarity of 84.9% (±34.7%) and a median of 99.05% shared targets. The co-occurrence files of the Healthy-associated and Bacterial-associated cohorts can be found in Supplementary Materials Files S2 and S3.

3. Discussion

Beyond colorectal cancer, dysbiosis of the gut microbiome has been implicated in a wide spectrum of human diseases. Altered microbial populations have been associated with neurodevelopmental and psychiatric conditions such as autism spectrum disorder, attention deficit hyperactivity disorder, depression, Alzheimer’s, and Parkinson’s disease [15,16]. Dysbiotic microbiota have also been linked to metabolic disorders, including obesity, type 2 diabetes, and non-alcoholic fatty liver disease [38,39], as well as cardiovascular diseases like atherosclerosis and hypertension [40]. Additionally, autoimmune diseases such as rheumatoid arthritis, multiple sclerosis, and type 1 diabetes have shown connections with gut microbial imbalances [12,41].
Among the various diseases influenced by the gut microbiome, colorectal cancer (CRC) stands out due to the particularly strong and consistent evidence linking microbial dysbiosis to its pathogenesis. A key characteristic of the CRC gut microbiome is its altered composition of bacterial strains relative to healthy individuals. Until now, the bacterium that has been mostly associated with the development of CRC is Helicobacter pylori. In line with its role as a potent pathogen, individuals infected with Helicobacter pylori harbor a nearly twofold increased risk to develop CRC [19]. Several large-scale epidemiological studies in the Netherlands strongly indicate that Salmonella infection elevates the risk of CRC. Key findings include a standardized incidence ratio (SIR 1.54) of early-onset CRC in the proximal colon following Salmonella exposure, a sustained higher risk associated with non-Enteritis or Typhimurium serovars and increased serological markers of Salmonella exposure (FliC antibodies) in CRC patients [18]. Furthermore, inflammation resulting from Salmonella colonization may be a contributing factor to this increased CRC risk [18]. Another bacterial strain that has been linked to CRC is Campylobacter jejuni, which produces a DNA-altering cytolethal distending toxin (CDT) [20]. Although it is known that CDT contributes to the development of inflammation in the GI tract, it was recently demonstrated that not only the production of cdtB advances CRC and promotes metastasis, but also that even the presence of this type of bacteria can affect the components and transcriptional activity of the gut microbial population [20,42]. A study conducted in Taiwan that analyzed pyogenic liver abscesses (PLA), an early sign of CRC, caused by Klebsiella pneumoniae, revealed that those strain-specific abscesses resulted in a far greater rate of subsequent CRC than abscesses derived from other bacteria [43]. Clostridium difficile is another bacterium that is implicated in gastrointestinal infections and antibiotic-associated colitis. C. difficile produces three different toxins, two of which (i.e., toxin A (TcdA), toxin B (TcdB)) cause detrimental effects to the GI tract’s epithelial barrier, damage the cells’ genetic material as well as activate STAT3 and NF-κΒ chronic inflammation-related pathways, potentially triggering CRC pathogenesis [44]. Even though Bacteroides fragilis and Escherichia coli are a normal part of the enteric microbiome, some pathogenic strains of them have been linked to CRC [45]. Toxin-expressing strains of B. fragilis (Bacteroides fragilis toxin (bft)) and E. coli (colibactin (clbB)) have not only been spatially associated in biofilms in the gut, but also their synergistic pro-carcinogenic involvement in CRC has emerged [22]. A recent study by Ding et. al. revealed that B. fragilis promotes chemoresistance in CRC, and at the same time, phage elimination experiments conducted in mice uncovered restored chemosensitivity of those CRC cells [46]. A great number of studies have connected CRC pathology with B. fragilis, but its exact role remains unclear [47,48,49]. Streptococcus gallolyticus is an opportunistic pathogen that has been associated with numerous studies with stimulation of cell reproduction and augmentation of tumor burden [21]. Additionally, Fusobacterium nucleatum also promotes chemoresistance of CRC cells through modulation of autophagy [23]. Although Enterococcus faecalis has been described as both a stimulant and a protector against CRC [50], recent studies have elucidated the role of biliverdin (i.e., one of its metabolites) as a tumor-stimulating compound that affects the host’s PI3K/AKT/mTOR pathway and thus promotes cell proliferation and angiogenesis [51,52]. Moreover, an experiment conducted by Chang et al. demonstrated that Parvimonas micra activated the Ras/ERK/c-Fos signaling pathway via micro-RNA upregulation and enhanced cellular proliferation in CRC [53]. Other studies have also highlighted P. micra’s implication in the immune response of CRC patients, potentially serving as a predictive biomarker for poor patient survival in CRC [54]. Peptostreptococcus anaerobius interacts directly with colonic cells via one of its surface proteins and also activates the proliferation-related integrin α2/β1-PI3K-Akt-NF-κB pathway [55]. Furthermore, P. anaerobius has been shown to intensify chemoresistance to oxaliplatin [56]. Porphyromonas gingivalis contributes to the proliferation of colorectal cancer cells through a mechanism involving cellular invasion and the subsequent activation of the MAPK/ERK signaling pathway [24]. Besides this role, a recent experiment showed that P. gingivalis upregulates chitinase 3-like-1 protein (CHI3L1) in invariant natural killer T (iNKT) cells, leading to detrimental effects on their cytotoxic function that result in the immune evasion of the tumors [57].
In parallel, recent studies highlight the importance of protein-level interactions between gut microbiota and the host. However, the experimental interactome remains limited, motivating the use of computational methods, including machine learning and domain-domain interaction prediction, to expand our understanding [32,33,34,35,36]. By integrating computational predictions with clinical microbiome data, this work contributes to mapping the molecular dialogs that underlie CRC development and potentially other microbiome-associated diseases.
The two phenotype-specific subnetworks, which were created after the categorization of the bacteria genera by phenotype association (i.e., health and CRC associated) after excluding the non-associated genera, revealed that the human proteins that are present in both subnetworks are the same. Interestingly, those human proteins also, when ranked according to their centrality degree, remain in the same order in both cases. This finding highlights not only the robustness of the human protein interactome and its key interactors but also emphasizes the complexity of the human proteins that can interact with different proteins at different health statuses. These most connected human proteins in these networks are mainly metalloproteins (E3 ubiquitin protein ligases, zinc transporters, etc.), transmembrane proteins, proteins related to ADP ribosylation and proteins related to mechanisms of cell proliferation and survival (i.e., apoptosis, programmed cell death, growth factors). It has been demonstrated that various species of pathogenic bacteria encode E3 ligases that have the ability to hijack the host’s ubiquitination mechanisms via a series of different strategies, including mimicking host-derived E3 ligases and encoding novel E3 ligases, as well as encoding deubiquitinases. Those proteins are transported into the host cell via type III or type IV secretion systems (T3SS and T4SS, respectively) and can then manipulate the host’s machinery for their proliferation [58,59]. Furthermore, there are studies that demonstrate the influence of the host’s zinc transporters on the homeostasis of the gut. A recent study showed that deletion of zinc (Zn) transporter ZIP14 creates a reduction in Zn in the entire intestinal tract and ultimately leads to lower microbial diversity [60,61]. Additionally, ATP-binding cassette (ABC) transporters have been associated with the microbial populations of the gut. More specifically, gut bacteria have been shown to both utilize them as binding receptors but also up- and downregulate their expression [62]. Research has also demonstrated that the bacterial ADP-ribosylation system, and mainly ADP-ribosyl transferases (ARTs), irreversibly modify host proteins with key function to the cellular cycle [63]. For example, Bxa of Bacteroides modifies non-muscle myosin II and triggers cellular remodeling that ultimately leads to inosine secretion that is then used by the microorganism as a carbon source [64]. Finally, since the gut bacteria have long been associated with cancer, mainly types of cancer that affect the GI tract. Therefore, it is no surprise that the gut microbiome influences growth factors, like the insulin-like growth factor 1 (IGF-1), which is essential for bone growth [65]. With respect to apoptosis, gut bacteria exert dual effects by either promoting or inhibiting cell death. Certain pathogens, such as non-typhoidal Salmonella, induce macrophage apoptosis via SPI-1 expression, thereby limiting inflammatory cytokine production [66]. Conversely, many bacterial species actively suppress host cell apoptosis to evade efferocytosis and enhance survival within the host [67].
On one hand, in the health-specific network, the most important proteins are those belonging to Lachnospiraceae, which are among the most abundant taxa in the GI tract. Evidence from multiple studies suggests that members of the Lachnospiraceae family contribute to maintaining the host’s physiological functions [68]. More specifically, they produce short-chain fatty acids that are converted into secondary bile acids that hinder the colonization of pathogenic strains [69].
On the other hand, in the bacteria-specific network, the most important proteins include several outer membrane, transport, and regulatory elements with distinct roles in microbial adaptation and survival. Among them, SusE and SusF are two outer membrane proteins from Bacteroides composed of tandem starch-specific carbohydrate-binding modules (CBMs) [70]. Although they lack enzymatic activity, these proteins are thought to play an essential role in starch metabolism by sequestering polysaccharides at the bacterial surface, thereby limiting access to competitor host cells [70]. Additional key proteins include alanine racemase, an enzyme required for the conversion of L-alanine to D-alanine and thus the synthesis of peptidoglycan, a critical component of bacterial cell walls [71], and ATP-binding cassette (ABC) transporters, which form a large superfamily of membrane complexes responsible for nutrient uptake, protein secretion, and resistance to environmental stressors, but also drug transfer [72,73]. Regulatory and DNA-binding proteins also contribute substantially to this network. These include lactose-binding proteins, such as the LacI repressor, which controls metabolic gene expression [74]. Enzymes such as thymidylate synthase, essential for deoxythymidine monophosphate (dTMP) synthesis [75], and adenine-specific DNA methyltransferases, which protect bacterial genomes from restriction enzymes [76], further emphasize the centrality of DNA-modifying activities. Finally, metalloproteins, particularly those containing iron–sulfur clusters, act as critical cofactors in electron transfer, catalysis, and gene regulation [77]. Collectively, these proteins reflect the molecular strategies employed by bacteria to secure nutrients, maintain genomic stability, and compete effectively within the intestinal ecosystem.
Apart from the most central proteins, proteins with centrality values near the mean also merit particular attention. From a graph-theoretical perspective, these nodes provide critical redundancy within the network: if only the most central proteins were perturbed, network function would collapse unless compensated by the surrounding intermediate nodes. Pathway enrichment analysis revealed that these moderately central proteins are implicated in the same biological pathways as the top-ranking hubs, thereby reinforcing their functional relevance. Their involvement suggests that they may act as auxiliary regulators or stabilizers, ensuring continuity of pathway activity and buffering the network against perturbations that target the primary hubs. Collectively, this highlights that both highly central and near-mean centrality proteins contribute to the robustness of bacteria–host interactions, and their combined roles are essential for maintaining network integrity under physiological and pathological conditions. The table with all the host and bacterial proteins with centrality values near the mean for both disease states can be found in Supplementary Materials File S4.
Interestingly, analysis of cross-distribution patterns revealed that CRC-associated taxa were also detectable within healthy samples, while health-associated taxa were observed in CRC samples. Quantitative analysis demonstrated that 48 bacterial taxa classified as CRC-associated were also detected in healthy samples, with an average relative abundance of approximately 0.27. Conversely, 40 taxa typically considered health-associated were found in CRC samples, with a higher mean abundance of about 1.03. These findings suggest that bacterial classification as CRC- or health-associated is not absolute but rather context-dependent, reflecting shifts in abundance and ecological balance rather than strict presence or absence. The observation that these taxa appear across both sample groups underscores the importance of relative abundance and community composition in shaping host–microbiome interactions and highlights that disease associations are likely driven by dysbiosis and altered network dynamics rather than by individual taxa alone. The tables with the Health-associated strains present in CRC patients and the CRC-associated strains in Healthy individuals, along with their relative abundances, can be found in Supplementary Materials Files S5 and S6.
Concerning the DL-based methodology, the integration of ProstT5-based embeddings and clinical metagenomics offers a high-resolution view of the CRC interactome, but it is essential to acknowledge the inherent constraints of a deep-learning-driven approach. The scientific rigor of this work is supported by the deliberate selection of primary, peer-reviewed repositories such as HPIDB, IntAct, and PHISTO for training, ensuring the model is grounded in the most comprehensive and experimentally validated interspecies data available. However, the dependence on high-fidelity input from predictive frameworks like AlphaFold and the “black-box” nature of deep neural networks present interpretability challenges regarding the precise biochemical drivers of each interaction. To mitigate the risks of dataset bias and overfitting common in high-dimensional biological spaces, we implemented technical safeguards, including focal loss functions and strict early stopping protocols. By prioritizing transparency in data filtering and calibrating thresholds to ensure a precision-driven interactome, this study provides a robust framework for identifying novel therapeutic targets and biomarkers while recognizing that these findings reflect computational predictions that require subsequent mechanistic validation. Finally, a limitation of this study is the lack of external validation in independent and ethnically diverse cohorts. While such validation is essential for assessing the generalizability of our findings, the present work was designed as a proof-of-concept study focusing on the development and internal evaluation of a deep learning framework for PPI mapping in a well-characterized colorectal cancer cohort from Crete/Greece. This setting provided a relatively homogeneous population, which facilitated controlled model development and reduced confounding variability during the initial phase of analysis. Future studies should aim to validate these findings across multicenter cohorts and diverse populations, as well as integrate additional omics datasets to further evaluate robustness and translational potential.

4. Materials and Methods

4.1. Training Dataset

This study began by compiling protein–protein interaction (PPI) data from multiple sources. Initially, we retrieved the pan-human–bacterial dataset comprising 19,686 experimentally validated interactions between 5714 bacterial and 4287 human proteins from four public databases: HPIDB [78,79], IntAct [80,81,82], PHISTO [83], and MorCVD [84]. To broaden our scope, we also incorporated a more inclusive PPI dataset from six widely used interaction databases (i.e., IntAct [80], MINT [85], DIP [86], HPRD [87], BioGRID [88], and SIFTS [89,90]). The selection of these specific repositories was based on their status as primary, peer-reviewed sources for experimentally validated interspecies interactions, ensuring the model was grounded in the most comprehensive and high-confidence data currently available in the field. This curated approach was designed to minimize noise and maximize the biological relevance of the training features. The dataset contained 1,081,401 PPIs, including 330,530 human inter-species PPIs and 750,871 inter- and intra-species interactions across diverse organisms. Notably, within this larger set, only 13 PPIs between host and gut bacterial proteins were identified, none involving proteoforms of the same gene.
To standardize structural information, all proteins from both datasets were mapped to their corresponding structures using the AlphaFold database API [91,92], mitigating potential biases from varying structural quality.
For constructing the positive dataset, both the original and larger PPI collections were filtered, retaining only interactions where both participating proteins had available structures. Conversely, the negative dataset, representing non-interacting pairs, was constructed using human proteins known to reside in different organs (data from Human Protein Atlas [93]) and whose domains (i.e., Pfam domains [94]) are not known to interact; the complete human proteome was sourced from UniProt Proteomes.
Furthermore, a gold-standard dataset of 17,278 experimentally supported domain-domain interactions (DDIs) from PDB complexes was retrieved from the 3did database [95]. These DDIs were then filtered to include only those not already present in our positive dataset or the known human interactome, ensuring unique, high-confidence examples.
The final curated dataset comprised a total of 16,690,098 PPI samples. To ensure robust model training and evaluation, these data were partitioned into training (60%—10,681,662 samples), validation (20%—2,670,416 samples), and test (20%—3,338,020 samples) subsets. A consistent class distribution was maintained across all subsets, reflecting the inherent imbalance of the interactome, with approximately 6% positive (interacting) and 94% negative (non-interacting) pairs. This distribution strategy was implemented to prevent overfitting and ensure that the performance metrics accurately reflect the model’s generalization capabilities on large-scale clinical data.

4.2. Deep Learning Model Architecture

The current study builds upon our previously published structure-based DL framework for human–gut bacterial PPI prediction [37]. While the architectural foundation remains unchanged, the protein feature extraction pipeline was modified to accommodate the unique demands of the current dataset and application domain. In the original model, protein structures were converted into graph representations and encoded using a pre-trained variational autoencoder (VAE). In this study, we replaced the VAE-based structural embedder with the Protein structure-sequence T5 (ProstT5) framework, a protein language model capable of capturing both sequential and structural information [96].
ProstT5 is a bilingual protein language model from the ProtT5 transformer architecture, fine-tuned to perform bi-directional translation between amino acid sequences and 3D structural representations. For the purposes of this study, we leveraged the encoder component of ProstT5 solely for protein feature extraction. Each protein sequence was converted into a fixed-size embedding vector using average pooling over the token-level embeddings obtained from ProstT5’s encoder. These embeddings were then used as inputs to the subsequent components of our established DL pipeline.
Unlike the structure-based MAPE-PPI encoder [97] used in prior work [37], which generates embeddings from a single processed representation of protein structure, ProstT5 leverages a bidirectional translation framework that learns jointly from amino acid sequences and structure-derived 3Di tokens. During pre-training, ProstT5 is trained not only to predict masked segments within either modality but also to translate from sequence to structure and from structure back to sequence, capturing the underlying correspondence between linear sequence patterns and their three-dimensional context. This bidirectional translation enables the model to internalize the mutual constraints between folding, residue interactions, and sequence evolution, producing embeddings that are structurally aware yet sequence-complete. At inference time, ProstT5 can generate informative embeddings directly from sequence data while implicitly encoding structural features, giving it a strong generalization advantage over models that depend on explicit sequence or structural inputs.
The downstream architecture, comprising a bi-directional Cross-Attention module and a fully connected interaction classifier, was preserved without modification. Briefly, the embeddings of a protein pair, i.e., one host and one and bacterial, were independently projected into a common subspace and fused via a multi-head cross-attention mechanism. This fusion step allows the model to capture inter-protein contextual dependencies and structural complementarity. The resulting interaction embedding was passed through a sequence of fully connected layers to produce a scalar interaction score. Class imbalance in the dataset was addressed using the focal loss function, and training was performed using the Adam optimizer with early stopping to prevent overfitting. Through this update, the model integrates ProstT5’s capacity to generalize across diverse protein sequences and structures while maintaining our previously validated attention-based interaction modeling and training strategy.

4.3. Application of DL Model on Clinical Data

4.3.1. Clinical Data

The patient cohort (N = 152) comprised individuals diagnosed with colorectal cancer, collected at the time of initial diagnosis and before the application of any type of treatment. The patients were recruited from 3 major centers of colorectal surgery in Crete, Greece, thereby representing a geographically distinct population with inherent commonalities in dietary patterns, climatic exposure, and potentially other lifestyle factors. The control group (N = 91) consisted of people without any known pathology that could influence the homeostasis of the gut microbiome. The demographic information of the clinical cohort is presented in Figure 2. The 204 bacterial genera identified in both control and CRC samples were meticulously mapped to their respective strains. The primary resource for this mapping was the Human Gut Microbiome Atlas [98]. In cases where strain information was absent from the Atlas, data from UniProt [99,100] were employed for manual strain assignment. This detailed strain-level resolution was a necessary precursor to the subsequent computational investigations.

4.3.2. Host–Microbiome PPI Network Prediction

Using the trained DL model, large-scale predictions of PPIs between human gut-expressed proteins and proteins derived from bacterial strains identified in each subject’s microbiome profile were performed. All pairwise combinations of host and microbial proteins were evaluated, and interaction probabilities were computed. The calibrated decision threshold, ensuring a minimum precision of 0.5 based on validation set performance, was applied to define interaction presence. The resulting subset of protein pairs was thus classified as interacting, forming an individual-level predicted interspecies PPI network. To further focus on the most confident predictions, a second filtering step was applied to the predicted network by retaining only protein pairs with posterior interaction probability close to 1.0.

4.3.3. Post-Prediction Analysis of Interaction Probabilities

1.
Mapping to Clinical Phenotypes: Each bacterial genus was annotated based on the disease status of the individual (colorectal cancer vs. healthy control), facilitating a group-specific comparison of PPI patterns. Non-parametric statistical testing at the taxonomic level was performed to detect significant differences in relative abundance between the two phenotypic groups. For each taxon, a Mann–Whitney U test (two-sided) was conducted to compare abundance distributions between healthy and cancer samples, selected for its robustness to non-normal distributions and unequal variances [101]. The median abundance in each phenotypic group, the Mann–Whitney U statistic, and the associated p-value were computed. For multiple hypothesis testing correction, the Benjamini–Hochberg procedure was applied to control the false discovery rate (FDR), reporting q-values alongside significance indicators (q-value < 0.05) [102]. Each genus was classified based on its differential abundance profile. Taxa with statistically significant differences (FDR-corrected q-value < 0.05) were labeled as either:
  • CRC-associated: higher median abundance in the CRC cohort.
  • Health-associated: higher median abundance in the healthy cohort.
Samples without significant differences, or with identical medians across groups, were labeled as Both (No clear association).
2.
Aggregation at the Genus Level: To generalize interaction patterns across microbial taxonomy, predicted interactions were aggregated per bacterial genus.
3.
Comparative Analysis by Disease State: Identification of genera or host targets with distinct interaction profiles in cancer versus health state.

4.3.4. Topological Analysis

The resulting high-confidence network was subjected to graph theory analysis. Degree centrality was computed to identify central nodes within the host and bacterial sub-networks, highlighting putative key mediators of host–microbiome molecular interaction. All bacterial proteins were remapped to their corresponding genus and disease state to facilitate downstream comparative analyses.

4.3.5. Co-Occurrence Network Construction

In parallel, co-occurrence networks of bacterial genera were constructed separately for the healthy and CRC cohorts. Pairwise Pearson correlation coefficients were calculated between genus-level abundances across samples. Only statistically significant and strong associations (|r| > 0.6, p < 0.05) were retained, resulting in genus–genus networks representing co-existence patterns. These networks were analyzed using degree centrality, and community detection was also applied to identify microbiome clusters.

4.3.6. Comparative Analysis of Co-Occurrence and Predicted Molecular Networks

To explore convergence between co-occurrence and molecular interaction patterns, an examination was conducted to determine whether bacterial genera co-occurring in the ecological network also exhibited shared or functionally similar PPI profiles with the human host. Bacterial proteomes from the co-occurrence networks were cross-referenced with those used in the DL-based predictions. Further filtering was applied to ensure taxonomic consistency, excluding unmatched strains. Similarity of host target profiles was assessed among co-occurring genera to evaluate potential functional redundancy or cooperative behavior.

4.3.7. Functional Characterization of Host Targets

Finally, human proteins identified as central nodes in the predicted PPI network were subjected to functional enrichment analysis to determine their involvement in cancer-related processes such as inflammation, immune modulation, apoptosis, and metabolism. This analysis was further contextualized by mapping bacterial network clusters (derived from co-occurrence patterns) to their corresponding host interaction signatures.

5. Conclusions

This study integrates clinical microbiome data with a deep learning framework for host–microbe protein–protein interaction (PPI) prediction to provide new insights into the molecular basis of colorectal cancer (CRC). By mapping phenotype-specific PPI networks, we identified highly central bacterial and human proteins that may serve as key mediators of host–microbiome crosstalk. Interestingly, human proteins maintained consistent centrality across both healthy- and CRC-associated networks, underscoring their robustness and pivotal roles in cellular processes such as ubiquitination, apoptosis, zinc transport, and growth regulation. On the microbial side, central proteins highlighted distinct strategies of adaptation and survival, ranging from starch sequestration and ABC transporter activity to DNA modification and electron transfer.
Beyond the most central proteins, nodes with near-mean centrality were shown to contribute substantially to network resilience, participating in the same pathways as the hubs and acting as auxiliary stabilizers of host–microbiome interactions. Therefore, CRC pathogenicity cannot be addressed solely by targeting hubs but requires intervention across the entire pathway that secondary nodes stabilize. Furthermore, cross-distribution analyses revealed that CRC-associated taxa were present in healthy samples, and health-associated taxa were detectable in CRC, indicating that disease associations are more strongly linked to dysbiosis and ecological imbalance than to strict taxonomic exclusivity.
Collectively, these findings reinforce the concept that colorectal cancer arises not from the influence of single bacterial species alone but from complex, network-level perturbations in host–microbiome interactions. By demonstrating the utility of computational PPI prediction in conjunction with clinical microbiome profiling, this work provides a framework for uncovering mechanistic pathways of CRC development and opens avenues for the identification of novel therapeutic targets and biomarkers.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/ijms27104232/s1.

Author Contributions

D.P.K. and G.C.B.: Methodology, Investigation, Data Curation, and Writing—Original Draft; I.M. and J.S.: Writing—Review and Editing; C.T.C.: Conceptualization, Writing—Review and Editing, Methodology, Investigation, and Supervision. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

This study has been approved by the Ethics Committee/Institutional Review Board of the University Hospital of Heraklion (Number 27/31 January 2020 and 17 February 2020). All procedures performed were in accordance with the ethical standards of the institutional and/or national research committee and the 1964 Helsinki Declaration and its later amendments or comparable ethical standards.

Informed Consent Statement

All patients signed a written informed consent form for their participation in the study.

Data Availability Statement

The original contributions presented in this study are included in the article/Supplementary Materials. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

Author Ippokratis Messaritakis was employed by the company Yiannoukas Labs Ltd., Bioiatriki Group. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Abavisani, M.; Faraji, N.; Ebadpour, N.; Kesharwani, P.; Sahebkar, A. Beyond digestion: Exploring how the gut microbiota modulates human social behaviors. Neuroscience 2025, 565, 52–62. [Google Scholar] [CrossRef] [PubMed]
  2. Sorboni, S.G.; Moghaddam, H.S.; Jafarzadeh-Esfehani, R.; Soleimanpour, S. A Comprehensive Review on the Role of the Gut Microbiome in Human Neurological Disorders. Clin. Microbiol. Rev. 2022, 35, e0033820. [Google Scholar] [CrossRef] [PubMed]
  3. Chandrasekaran, P.; Weiskirchen, S.; Weiskirchen, R. Effects of Probiotics on Gut Microbiota: An Overview. Int. J. Mol. Sci. 2024, 25, 6022. [Google Scholar] [CrossRef]
  4. Zysset-Burri, D.C.; Morandi, S.; Herzog, E.L.; Berger, L.E.; Zinkernagel, M.S. The role of the gut microbiome in eye diseases. Prog. Retin. Eye Res. 2023, 92, 101117. [Google Scholar] [CrossRef] [PubMed]
  5. Binda, C.; Lopetuso, L.R.; Rizzatti, G.; Gibiino, G.; Cennamo, V.; Gasbarrini, A. Actinobacteria: A relevant minority for the maintenance of gut homeostasis. Dig. Liver Dis. 2018, 50, 421–428. [Google Scholar] [CrossRef]
  6. Shin, N.R.; Whon, T.W.; Bae, J.W. Proteobacteria: Microbial signature of dysbiosis in gut microbiota. Trends Biotechnol. 2015, 33, 496–503. [Google Scholar] [CrossRef]
  7. Solar, C.; Escalona, A.; Garrido, D. Chapter 24—The Gut Microbiome After Bariatric Surgery. In Microbiome and Metabolome in Diagnosis, Therapy, and Other Strategic Applications; Faintuch, J., Faintuch, S., Eds.; Academic Press: Cambridge, MA, USA, 2019; pp. 235–242. [Google Scholar]
  8. Arun, K.B.; Madhavan, A.; Sindhu, R.; Emmanual, S.; Binod, P.; Pugazhendhi, A.; Sirohi, R.; Reshmy, R.; Awasthi, M.K.; Gnansounou, E.; et al. Probiotics and gut microbiome—Prospects and challenges in remediating heavy metal toxicity. J. Hazard. Mater. 2021, 420, 126676. [Google Scholar] [CrossRef]
  9. Singh, R.; Zogg, H.; Wei, L.; Bartlett, A.; Ghoshal, U.C.; Rajender, S.; Ro, S. Gut Microbial Dysbiosis in the Pathogenesis of Gastrointestinal Dysmotility and Metabolic Disorders. J. Neurogastroenterol. Motil. 2021, 27, 19–34. [Google Scholar] [CrossRef]
  10. Xu, W.-T.; Nie, Y.-Z.; Yang, Z.; Lu, N.-H. The Crosstalk Between Gut Microbiota and Obesity and Related Metabolic Disorders. Future Microbiol. 2016, 11, 825–836. [Google Scholar] [CrossRef] [PubMed]
  11. Witkowski, M.; Weeks, T.L.; Hazen, S.L. Gut Microbiota and Cardiovascular Disease. Circ. Res. 2020, 127, 553–570. [Google Scholar] [CrossRef]
  12. Xu, H.; Liu, M.; Cao, J.; Li, X.; Fan, D.; Xia, Y.; Lu, X.; Li, J.; Ju, D.; Zhao, H. The Dynamic Interplay between the Gut Microbiota and Autoimmune Diseases. J. Immunol. Res. 2019, 2019, 7546047. [Google Scholar] [CrossRef] [PubMed]
  13. Vuong, H.E.; Hsiao, E.Y. Emerging Roles for the Gut Microbiome in Autism Spectrum Disorder. Biol. Psychiatry 2017, 81, 411–423. [Google Scholar] [CrossRef]
  14. Wang, J.; Chen, W.-D.; Wang, Y.-D. The Relationship Between Gut Microbiota and Inflammatory Diseases: The Role of Macrophages. Front. Microbiol. 2020, 11, 1065. [Google Scholar] [CrossRef]
  15. Wang, Q.; Yang, Q.; Liu, X. The microbiota-gut-brain axis and neurodevelopmental disorders. Protein Cell 2023, 14, 762–775. [Google Scholar] [CrossRef]
  16. Hashimoto, K. Emerging role of the host microbiome in neuropsychiatric disorders: Overview and future directions. Mol. Psychiatry 2023, 28, 3625–3637. [Google Scholar] [CrossRef]
  17. Kim, J.; Lee, H.K. Potential Role of the Gut Microbiome in Colorectal Cancer Progression. Front. Immunol. 2022, 12, 807648. [Google Scholar] [CrossRef] [PubMed]
  18. Dougherty, M.W.; Jobin, C. Intestinal bacteria and colorectal cancer: Etiology and treatment. Gut Microbes 2023, 15, 2185028. [Google Scholar] [CrossRef] [PubMed]
  19. Ralser, A.; Dietl, A.; Jarosch, S.; Engelsberger, V.; Wanisch, A.; Janssen, K.P.; Middelhoff, M.; Vieth, M.; Quante, M.; Haller, D.; et al. Helicobacter pylori promotes colorectal carcinogenesis by deregulating intestinal immunity and inducing a mucus-degrading microbiota signature. Gut 2023, 72, 1258–1270. [Google Scholar] [CrossRef]
  20. He, Z.; Gharaibeh, R.Z.; Newsome, R.C.; Pope, J.L.; Dougherty, M.W.; Tomkovich, S.; Pons, B.; Mirey, G.; Vignard, J.; Hendrixson, D.R.; et al. Campylobacter jejuni promotes colorectal tumorigenesis through the action of cytolethal distending toxin. Gut 2019, 68, 289–300. [Google Scholar] [CrossRef]
  21. Taylor, J.C.; Kumar, R.; Xu, J.; Xu, Y. A pathogenicity locus of Streptococcus gallolyticus subspecies gallolyticus. Sci. Rep. 2023, 13, 6291. [Google Scholar] [CrossRef]
  22. Dejea, C.M.; Fathi, P.; Craig, J.M.; Boleij, A.; Taddese, R.; Geis, A.L.; Wu, X.; DeStefano Shields, C.E.; Hechenbleikner, E.M.; Huso, D.L.; et al. Patients with familial adenomatous polyposis harbor colonic biofilms containing tumorigenic bacteria. Science 2018, 359, 592–597. [Google Scholar] [CrossRef] [PubMed]
  23. Yu, T.; Guo, F.; Yu, Y.; Sun, T.; Ma, D.; Han, J.; Qian, Y.; Kryczek, I.; Sun, D.; Nagarsheth, N.; et al. Fusobacterium nucleatum Promotes Chemoresistance to Colorectal Cancer by Modulating Autophagy. Cell 2017, 170, 548–563.e16. [Google Scholar] [CrossRef]
  24. Mu, W.; Jia, Y.; Chen, X.; Li, H.; Wang, Z.; Cheng, B. Intracellular Porphyromonas gingivalis Promotes the Proliferation of Colorectal Cancer Cells via the MAPK/ERK Signaling Pathway. Front. Cell. Infect. Microbiol. 2020, 10, 584798. [Google Scholar] [CrossRef] [PubMed]
  25. Dyer, M.D.; Neff, C.; Dufford, M.; Rivera, C.G.; Shattuck, D.; Bassaganya-Riera, J.; Murali, T.M.; Sobral, B.W. The Human-Bacterial Pathogen Protein Interaction Networks of Bacillus anthracis, Francisella tularensis, and Yersinia pestis. PLoS ONE 2010, 5, e12089. [Google Scholar] [CrossRef]
  26. Acharya, D.; Dutta, T.K. Elucidating the network features and evolutionary attributes of intra- and interspecific protein–protein interactions between human and pathogenic bacteria. Sci. Rep. 2021, 11, 190. [Google Scholar] [CrossRef] [PubMed]
  27. Yang, F.; Lei, Y.; Zhou, M.; Yao, Q.; Han, Y.; Wu, X.; Zhong, W.; Zhu, C.; Xu, W.; Tao, R.; et al. Development and application of a recombination-based library versus library high- throughput yeast two-hybrid (RLL-Y2H) screening system. Nucleic Acids Res. 2018, 46, e17. [Google Scholar] [CrossRef]
  28. Li, X.-M.; Huang, S.; Li, X.D. Photo-ANA enables profiling of host–bacteria protein interactions during infection. Nat. Chem. Biol. 2023, 19, 614–623. [Google Scholar] [CrossRef]
  29. Walch, P.; Selkrig, J.; Knodler, L.A.; Rettel, M.; Stein, F.; Fernandez, K.; Viéitez, C.; Potel, C.M.; Scholzen, K.; Geyer, M.; et al. Global mapping of Salmonella enterica-host protein-protein interactions during infection. Cell Host Microbe 2021, 29, 1316–1332.e12. [Google Scholar] [CrossRef]
  30. Post, S.E.; Brito, I.L. Structural insight into protein–protein interactions between intestinal microbiome and host. Curr. Opin. Struct. Biol. 2022, 74, 102354. [Google Scholar] [CrossRef]
  31. Schweppe, D.K.; Harding, C.; Chavez, J.D.; Wu, X.; Ramage, E.; Singh, P.K.; Manoil, C.; Bruce, J.E. Host-Microbe Protein Interactions during Bacterial Infection. Chem. Biol. 2015, 22, 1521–1530. [Google Scholar] [CrossRef]
  32. Deng, M.; Mehta, S.; Sun, F.; Chen, T. Inferring domain-domain interactions from protein-protein interactions. Genome Res. 2002, 12, 1540–1548. [Google Scholar] [CrossRef]
  33. Guimarães, K.S.; Jothi, R.; Zotenko, E.; Przytycka, T.M. Predicting domain-domain interactions using a parsimony approach. Genome Biol. 2006, 7, R104. [Google Scholar] [CrossRef] [PubMed]
  34. Singhal, M.; Resat, H. A domain-based approach to predict protein-protein interactions. BMC Bioinform. 2007, 8, 199. [Google Scholar] [CrossRef]
  35. Chen, X.-W.; Liu, M. Prediction of protein–protein interactions using random decision forest framework. Bioinformatics 2005, 21, 4394–4400. [Google Scholar] [CrossRef]
  36. Alborzi, S.Z.; Ahmed Nacer, A.; Najjar, H.; Ritchie, D.W.; Devignes, M.-D. PPIDomainMiner: Inferring domain-domain interactions from multiple sources of protein-protein interactions. PLoS Comput. Biol. 2021, 17, e1008844. [Google Scholar] [CrossRef] [PubMed]
  37. Kiouri, D.P.; Batsis, G.C.; Chasapis, C.T. Structure-Based Deep Learning Framework for Modeling Human–Gut Bacterial Protein Interactions. Proteomes 2025, 13, 10. [Google Scholar] [CrossRef]
  38. Ispas, S.; Tuta, L.A.; Botnarciuc, M.; Ispas, V.; Staicovici, S.; Ali, S.; Nelson-Twakor, A.; Cojocaru, C.; Herlo, A.; Petcu, A. Metabolic Disorders, the Microbiome as an Endocrine Organ, and Their Relations with Obesity: A Literature Review. J. Pers. Med. 2023, 13, 1602. [Google Scholar] [CrossRef]
  39. Zhang, X.; Cai, X.; Zheng, X. Gut microbiome-oriented therapy for metabolic diseases: Challenges and opportunities towards clinical translation. Trends Pharmacol. Sci. 2021, 42, 984–987. [Google Scholar] [CrossRef]
  40. Rahman, M.M.; Islam, F.; Or-Rashid, M.H.; Mamun, A.A.; Rahaman, M.S.; Islam, M.M.; Meem, A.F.K.; Sutradhar, P.R.; Mitra, S.; Mimi, A.A.; et al. The Gut Microbiota (Microbiome) in Cardiovascular Disease and Its Therapeutic Regulation. Front. Cell. Infect. Microbiol. 2022, 12, 903570. [Google Scholar] [CrossRef] [PubMed]
  41. Christovich, A.; Luo, X.M. Gut Microbiota, Leaky Gut, and Autoimmune Diseases. Front. Immunol. 2022, 13, 946248. [Google Scholar] [CrossRef] [PubMed]
  42. He, Z.; Yu, J.; Gong, J.; Wu, J.; Zong, X.; Luo, Z.; He, X.; Cheng, W.M.; Liu, Y.; Liu, C.; et al. Campylobacter jejuni-derived cytolethal distending toxin promotes colorectal cancer metastasis. Cell Host Microbe 2024, 32, 2080–2091.e6. [Google Scholar] [CrossRef]
  43. Huang, W.K.; Chang, J.W.; See, L.C.; Tu, H.T.; Chen, J.S.; Liaw, C.C.; Lin, Y.C.; Yang, T.S. Higher rate of colorectal cancer among patients with pyogenic liver abscess with Klebsiella pneumoniae than those without: An 11-year follow-up study. Color. Dis. 2012, 14, e794–e801. [Google Scholar] [CrossRef]
  44. Nezhadi, J.; Lahouty, M.; Rezaee, M.A.; Fadaee, M. Clostridium difficile as a potent trigger of colorectal carcinogenesis. Discov. Oncol. 2025, 16, 910. [Google Scholar] [CrossRef] [PubMed]
  45. Lichtenstern, C.R.; Lamichhane-Khadka, R. A tale of two bacteria—Bacteroides fragilis, Escherichia coli, and colorectal cancer. Front. Bacteriol. 2023, 2, 1229077. [Google Scholar] [CrossRef]
  46. Ding, X.; Ting, N.L.-N.; Wong, C.C.; Huang, P.; Jiang, L.; Liu, C.; Lin, Y.; Li, S.; Liu, Y.; Xie, M.; et al. Bacteroides fragilis promotes chemoresistance in colorectal cancer, and its elimination by phage VA7 restores chemosensitivity. Cell Host Microbe 2025, 33, 941–956.e10, Erratum in Cell Host Microbe 2025, 33, 1796. [Google Scholar] [CrossRef] [PubMed]
  47. Wu, Z.; Yu, M.; Zeng, Y.; Huang, Y.; Zheng, W. LRP11-AS1 mediates enterotoxigenic Bacteroides fragilis-related carcinogenesis in colorectal Cancer via the miR-149-3p/CDK4 pathway. Cancer Gene Ther. 2025, 32, 184–197. [Google Scholar] [CrossRef] [PubMed]
  48. Matsumiya, Y.; Suenaga, M.; Ishikawa, T.; Hanaoka, M.; Iwata, N.; Masuda, T.; Yamauchi, S.; Tokunaga, M.; Kinugasa, Y. Clinical significance of Bacteroides fragilis as potential prognostic factor in colorectal cancer patients. J. Clin. Oncol. 2022, 40, 137. [Google Scholar] [CrossRef]
  49. Nazarinejad, N.; Hajikhani, B.; Vaezi, A.A.; Firoozeh, F.; Sameni, F.; Yaslianifard, S.; Goudarzi, M.; Dadashi, M. Association between colorectal cancer, the frequency of Bacteroides fragilis, and the level of mismatch repair genes expression in the biopsy samples of Iranian patients. BMC Gastroenterol. 2024, 24, 82. [Google Scholar] [CrossRef]
  50. de Almeida, C.V.; Taddei, A.; Amedei, A. The controversial role of Enterococcus faecalis in colorectal cancer. Ther. Adv. Gastroenterol. 2018, 11, 1756284818783606. [Google Scholar] [CrossRef]
  51. Zhang, L.; Liu, J.; Deng, M.; Chen, X.; Jiang, L.; Zhang, J.; Tao, L.; Yu, W.; Qiu, Y. Enterococcus faecalis promotes the progression of colorectal cancer via its metabolite: Biliverdin. J. Transl. Med. 2023, 21, 72. [Google Scholar] [CrossRef]
  52. Zhang, L.; Deng, M.; Liu, J.; Zhang, J.; Wang, F.; Yu, W. The pathogenicity of vancomycin-resistant Enterococcus faecalis to colon cancer cells. BMC Infect. Dis. 2024, 24, 230. [Google Scholar] [CrossRef]
  53. Chang, Y.; Huang, Z.; Hou, F.; Liu, Y.; Wang, L.; Wang, Z.; Sun, Y.; Pan, Z.; Tan, Y.; Ding, L.; et al. Parvimonas micra activates the Ras/ERK/c-Fos pathway by upregulating miR-218-5p to promote colorectal cancer progression. J. Exp. Clin. Cancer Res. 2023, 42, 13. [Google Scholar] [CrossRef]
  54. Zhao, L.; Zhang, X.; Zhou, Y.; Fu, K.; Lau, H.C.-H.; Chun, T.W.-Y.; Cheung, A.H.-K.; Coker, O.O.; Wei, H.; Wu, W.K.-K.; et al. Parvimonas micra promotes colorectal tumorigenesis and is associated with prognosis of colorectal cancer patients. Oncogene 2022, 41, 4200–4210. [Google Scholar] [CrossRef]
  55. Long, X.; Wong, C.C.; Tong, L.; Chu, E.S.H.; Ho Szeto, C.; Go, M.Y.Y.; Coker, O.O.; Chan, A.W.H.; Chan, F.K.L.; Sung, J.J.Y.; et al. Peptostreptococcus anaerobius promotes colorectal carcinogenesis and modulates tumour immunity. Nat. Microbiol. 2019, 4, 2319–2330. [Google Scholar] [CrossRef]
  56. Gu, J.; Lv, X.; Li, W.; Li, G.; He, X.; Zhang, Y.; Shi, L.; Zhang, X. Deciphering the mechanism of Peptostreptococcus anaerobius-induced chemoresistance in colorectal cancer: The important roles of MDSC recruitment and EMT activation. Front. Immunol. 2023, 14, 1230681. [Google Scholar] [CrossRef]
  57. Díaz-Basabe, A.; Lattanzi, G.; Perillo, F.; Amoroso, C.; Baeri, A.; Farini, A.; Torrente, Y.; Penna, G.; Rescigno, M.; Ghidini, M.; et al. Porphyromonas gingivalis fuels colorectal cancer through CHI3L1-mediated iNKT cell-driven immune evasion. Gut Microbes 2024, 16, 2388801. [Google Scholar] [CrossRef]
  58. Berglund, J.; Gjondrekaj, R.; Verney, E.; Maupin-Furlow, J.A.; Edelmann, M.J. Modification of the host ubiquitome by bacterial enzymes. Microbiol. Res. 2020, 235, 126429. [Google Scholar] [CrossRef]
  59. Pisano, A.; Albano, F.; Vecchio, E.; Renna, M.; Scala, G.; Quinto, I.; Fiume, G. Revisiting Bacterial Ubiquitin Ligase Effectors: Weapons for Host Exploitation. Int. J. Mol. Sci. 2018, 19, 3576. [Google Scholar] [CrossRef]
  60. Thorn, T.L.; Mitchell, S.B.; Kim, Y.; Lee, M.-T.; Comrie, J.M.C.; Johnson, E.L.; Aydemir, T.B. Metal transporter SLC39A14/ZIP14 modulates regulation between the gut microbiome and host metabolism. bioRxiv 2021, 2021.12.22.473859. [Google Scholar] [CrossRef]
  61. Mitchell, S.B.; Thorn, T.L.; Lee, M.T.; Kim, Y.; Comrie, J.M.C.; Bai, Z.S.; Johnson, E.L.; Aydemir, T.B. Metal transporter SLC39A14/ZIP14 modulates regulation between the gut microbiome and host metabolism. Am. J. Physiol. Gastrointest. Liver Physiol. 2023, 325, G593–G607. [Google Scholar] [CrossRef]
  62. Mercado-Lubo, R.; McCormick, B.A. The interaction of gut microbes with host ABC transporters. Gut Microbes 2010, 1, 301–306. [Google Scholar] [CrossRef]
  63. Mikolčević, P.; Hloušek-Kasun, A.; Ahel, I.; Mikoč, A. ADP-ribosylation systems in bacteria and viruses. Comput. Struct. Biotechnol. J. 2021, 19, 2366–2383. [Google Scholar] [CrossRef]
  64. Brown, E.M.; Arellano-Santoyo, H.; Temple, E.R.; Costliow, Z.A.; Pichaud, M.; Hall, A.B.; Liu, K.; Durney, M.A.; Gu, X.; Plichta, D.R.; et al. Gut microbiome ADP-ribosyltransferases are widespread phage-encoded fitness factors. Cell Host Microbe 2021, 29, 1351–1365.e11. [Google Scholar] [CrossRef]
  65. Zheng, X.; Qian, Y.; Wang, L. Causal relationship between gut microbiota and insulin-like growth factor 1: A bidirectional two-sample Mendelian randomization study. Front. Cell. Infect. Microbiol. 2024, 14, 1406132. [Google Scholar] [CrossRef]
  66. Lin, H.H.; Chen, H.L.; Weng, C.C.; Janapatla, R.P.; Chen, C.L.; Chiu, C.H. Activation of apoptosis by Salmonella pathogenicity island-1 effectors through both intrinsic and extrinsic pathways in Salmonella-infected macrophages. J. Microbiol. Immunol. Infect. 2021, 54, 616–626. [Google Scholar] [CrossRef]
  67. Behar, S.M.; Briken, V. Apoptosis inhibition by intracellular bacteria and its consequence on host immunity. Curr. Opin. Immunol. 2019, 60, 103–110. [Google Scholar] [CrossRef]
  68. Vacca, M.; Celano, G.; Calabrese, F.M.; Portincasa, P.; Gobbetti, M.; De Angelis, M. The Controversial Role of Human Gut Lachnospiraceae. Microorganisms 2020, 8, 573. [Google Scholar] [CrossRef]
  69. Sorbara, M.T.; Littmann, E.R.; Fontana, E.; Moody, T.U.; Kohout, C.E.; Gjonbalaj, M.; Eaton, V.; Seok, R.; Leiner, I.M.; Pamer, E.G. Functional and Genomic Variation between Human-Derived Isolates of Lachnospiraceae Reveals Inter- and Intra-Species Diversity. Cell Host Microbe 2020, 28, 134–146.e4. [Google Scholar] [CrossRef]
  70. Cameron, E.A.; Maynard, M.A.; Smith, C.J.; Smith, T.J.; Koropatkin, N.M.; Martens, E.C. Multidomain Carbohydrate-binding Proteins Involved in Bacteroides thetaiotaomicron Starch Metabolism. J. Biol. Chem. 2012, 287, 34614–34625. [Google Scholar] [CrossRef]
  71. Wei, Y.; Qiu, W.; Zhou, X.-D.; Zheng, X.; Zhang, K.-K.; Wang, S.-D.; Li, Y.-Q.; Cheng, L.; Li, J.-Y.; Xu, X.; et al. Alanine racemase is essential for the growth and interspecies competitiveness of Streptococcus mutans. Int. J. Oral Sci. 2016, 8, 231–238. [Google Scholar] [CrossRef]
  72. Mahendran, A.; Orlando, B.J. Genome wide structural prediction of ABC transporter systems in Bacillus subtilis. Front. Microbiol. 2024, 15, 1469915. [Google Scholar] [CrossRef]
  73. Nigam, S.K. What do drug transporters really do? Nat. Rev. Drug Discov. 2015, 14, 29–44. [Google Scholar] [CrossRef]
  74. Hoover, T.R. Bacterial Transcription Factors. In Encyclopedia of Genetics; Brenner, S., Miller, J.H., Eds.; Academic Press: New York, NY, USA, 2001; pp. 163–165. [Google Scholar]
  75. Pozzi, C.; Lopresti, L.; Tassone, G.; Mangani, S. Targeting Methyltransferases in Human Pathogenic Bacteria: Insights into Thymidylate Synthase (TS) and Flavin-Dependent TS (FDTS). Molecules 2019, 24, 1638. [Google Scholar] [CrossRef]
  76. Wion, D.; Casadesús, J. N6-methyl-adenine: An epigenetic signal for DNA-protein interactions. Nat. Rev. Microbiol. 2006, 4, 183–192. [Google Scholar] [CrossRef]
  77. Esquilin-Lebron, K.; Dubrac, S.; Barras, F.; Boyd, J.M. Bacterial Approaches for Assembling Iron-Sulfur Proteins. mBio 2021, 12, e0242521. [Google Scholar] [CrossRef]
  78. Ammari, M.G.; Gresham, C.R.; McCarthy, F.M.; Nanduri, B. HPIDB 2.0: A curated database for host–pathogen interactions. Database 2016, 2016, baw103. [Google Scholar] [CrossRef]
  79. Kumar, R.; Nanduri, B. HPIDB—A unified resource for host-pathogen interactions. BMC Bioinform. 2010, 11, S16. [Google Scholar] [CrossRef]
  80. del Toro, N.; Shrivastava, A.; Ragueneau, E.; Meldal, B.; Combe, C.; Barrera, E.; Perfetto, L.; How, K.; Ratan, P.; Shirodkar, G.; et al. The IntAct database: Efficient access to fine-grained molecular interaction data. Nucleic Acids Res. 2021, 50, D648–D653. [Google Scholar] [CrossRef]
  81. Hermjakob, H.; Montecchi-Palazzi, L.; Lewington, C.; Mudali, S.; Kerrien, S.; Orchard, S.; Vingron, M.; Roechert, B.; Roepstorff, P.; Valencia, A.; et al. IntAct: An open source molecular interaction database. Nucleic Acids Res. 2004, 32, D452–D455. [Google Scholar] [CrossRef]
  82. Kerrien, S.; Aranda, B.; Breuza, L.; Bridge, A.; Broackes-Carter, F.; Chen, C.; Duesbury, M.; Dumousseau, M.; Feuermann, M.; Hinz, U.; et al. The IntAct molecular interaction database in 2012. Nucleic Acids Res. 2012, 40, D841–D846. [Google Scholar] [CrossRef]
  83. Durmuş Tekir, S.; Çakır, T.; Ardıç, E.; Sayılırbaş, A.S.; Konuk, G.; Konuk, M.; Sarıyer, H.; Uğurlu, A.; Karadeniz, İ.; Özgür, A.; et al. PHISTO: Pathogen–host interaction search tool. Bioinformatics 2013, 29, 1357–1358. [Google Scholar] [CrossRef]
  84. Singh, N.; Bhatia, V.; Singh, S.; Bhatnagar, S. MorCVD: A Unified Database for Host-Pathogen Protein-Protein Interactions of Cardiovascular Diseases Related to Microbes. Sci. Rep. 2019, 9, 4039. [Google Scholar] [CrossRef]
  85. Licata, L.; Briganti, L.; Peluso, D.; Perfetto, L.; Iannuccelli, M.; Galeota, E.; Sacco, F.; Palma, A.; Nardozza, A.P.; Santonico, E.; et al. MINT, the molecular interaction database: 2012 update. Nucleic Acids Res. 2012, 40, D857–D861. [Google Scholar] [CrossRef]
  86. Salwinski, L.; Miller, C.S.; Smith, A.J.; Pettit, F.K.; Bowie, J.U.; Eisenberg, D. The Database of Interacting Proteins: 2004 update. Nucleic Acids Res. 2004, 32, D449–D451. [Google Scholar] [CrossRef]
  87. Keshava Prasad, T.S.; Goel, R.; Kandasamy, K.; Keerthikumar, S.; Kumar, S.; Mathivanan, S.; Telikicherla, D.; Raju, R.; Shafreen, B.; Venugopal, A.; et al. Human Protein Reference Database—2009 update. Nucleic Acids Res. 2009, 37, D767–D772. [Google Scholar] [CrossRef]
  88. Oughtred, R.; Rust, J.; Chang, C.; Breitkreutz, B.J.; Stark, C.; Willems, A.; Boucher, L.; Leung, G.; Kolas, N.; Zhang, F.; et al. The BioGRID database: A comprehensive biomedical resource of curated protein, genetic, and chemical interactions. Protein Sci. 2021, 30, 187–200. [Google Scholar] [CrossRef]
  89. Velankar, S.; Dana, J.M.; Jacobsen, J.; van Ginkel, G.; Gane, P.J.; Luo, J.; Oldfield, T.J.; O’Donovan, C.; Martin, M.-J.; Kleywegt, G.J. SIFTS: Structure Integration with Function, Taxonomy and Sequences resource. Nucleic Acids Res. 2012, 41, D483–D489. [Google Scholar] [CrossRef]
  90. Dana, J.M.; Gutmanas, A.; Tyagi, N.; Qi, G.; O’Donovan, C.; Martin, M.; Velankar, S. SIFTS: Updated Structure Integration with Function, Taxonomy and Sequences resource allows 40-fold increase in coverage of structure-based annotations for proteins. Nucleic Acids Res. 2018, 47, D482–D489. [Google Scholar] [CrossRef]
  91. Jumper, J.; Evans, R.; Pritzel, A.; Green, T.; Figurnov, M.; Ronneberger, O.; Tunyasuvunakool, K.; Bates, R.; Žídek, A.; Potapenko, A.; et al. Highly accurate protein structure prediction with AlphaFold. Nature 2021, 596, 583–589. [Google Scholar] [CrossRef]
  92. Varadi, M.; Bertoni, D.; Magana, P.; Paramval, U.; Pidruchna, I.; Radhakrishnan, M.; Tsenkov, M.; Nair, S.; Mirdita, M.; Yeo, J.; et al. AlphaFold Protein Structure Database in 2024: Providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2023, 52, D368–D375. [Google Scholar] [CrossRef]
  93. Uhlén, M.; Fagerberg, L.; Hallström, B.M.; Lindskog, C.; Oksvold, P.; Mardinoglu, A.; Sivertsson, Å.; Kampf, C.; Sjöstedt, E.; Asplund, A.; et al. Tissue-based map of the human proteome. Science 2015, 347, 1260419. [Google Scholar] [CrossRef]
  94. Chasapis, C.T. Building Bridges Between Structural and Network-Based Systems Biology. Mol. Biotechnol. 2019, 61, 221–229. [Google Scholar] [CrossRef]
  95. Mosca, R.; Céol, A.; Stein, A.; Olivella, R.; Aloy, P. 3did: A catalog of domain-based interactions of known three-dimensional structure. Nucleic Acids Res. 2014, 42, D374–D379. [Google Scholar] [CrossRef] [PubMed]
  96. Heinzinger, M.; Weissenow, K.; Sanchez, J.G.; Henkel, A.; Mirdita, M.; Steinegger, M.; Rost, B. Bilingual language model for protein sequence and structure. NAR Genom. Bioinform. 2024. [Google Scholar] [CrossRef]
  97. Wu, L.; Tian, Y.; Huang, Y.; Li, S.; Lin, H.; Chawla, N.V.; Li, S.Z. Mape-ppi: Towards effective and efficient protein-protein interaction prediction via microenvironment-aware protein embedding. arXiv 2024, arXiv:2402.14391. [Google Scholar]
  98. Human Gut Microbiome Atlas. Available online: https://www.microbiomeatlas.org (accessed on 2 July 2024).
  99. UniProt (Universal Protein Resource). Available online: https://www.uniprot.org/ (accessed on 16 January 2023).
  100. The UniProt Consortium. UniProt: The Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 2022, 51, D523–D531. [Google Scholar] [CrossRef]
  101. McKnight, P.E.; Najab, J. Mann-Whitney U Test. In The Corsini Encyclopedia of Psychology; John Wiley & Sons, Inc.: Hoboken, NJ, USA, 2010; p. 1. [Google Scholar]
  102. Thissen, D.; Steinberg, L.; Kuang, D. Quick and easy implementation of the Benjamini-Hochberg procedure for controlling the false positive rate in multiple comparisons. J. Educ. Behav. Stat. 2002, 27, 77–83. [Google Scholar] [CrossRef]
Figure 1. ROC curve for the test portion of the PPI dataset.
Figure 1. ROC curve for the test portion of the PPI dataset.
Ijms 27 04232 g001
Figure 2. The demographic information of the clinical cohort.
Figure 2. The demographic information of the clinical cohort.
Ijms 27 04232 g002
Table 1. Evaluation metrics for prediction on the test portion of the PPI dataset.
Table 1. Evaluation metrics for prediction on the test portion of the PPI dataset.
Precision (macro)0.9911
Recall (macro)0.9747
F1 (macro)0.9827
Precision (positive)0.9855
Recall (positive)0.9503
F1 (Positive)0.9676
MCC0.9657
Balanced Accuracy0.9747
AP0.9867
AUROC0.9959
Recall@precision_500.9901
Table 2. Confusion Matrix for the prediction on the test portion of the PPI dataset.
Table 2. Confusion Matrix for the prediction on the test portion of the PPI dataset.
Predicted NegativePredicted Positive
Actual Negative3,129,9492853
Actual Positive10,133194,112
Table 3. Phenotype association information.
Table 3. Phenotype association information.
Phenotype AssociationNumber of Samples
Both (No clear association)116
CRC-associated48
Health-associated40
Table 4. CRC-associated samples.
Table 4. CRC-associated samples.
TaxonomyMedianp-Valueq-Value
g__unknown_Finegoldia0.94801.472 × 10−391.502 × 10−37
g__unknown_Anaerococcus0.57841.238 × 10−391.502 × 10−37
g__unknown_Peptoniphilus1.21813.501 × 10−392.381 × 10−37
g__Corynebacterium0.07591.440 × 10−346.923 × 10−33
g__Porphyromonas0.48211.697 × 10−346.923 × 10−33
g__unknown_Ezakiella0.10241.342 × 10−314.562 × 10−30
g__Campylobacter0.14192.343 × 10−316.828 × 10−30
g__Lawsonella0.01395.945 × 10−311.516 × 10−29
g__S5-A14a0.07198.633 × 10−301.957 × 10−28
g__Fusobacterium0.06263.299 × 10−286.729 × 10−27
g__unknown_Fenollaria0.12801.485 × 10−272.753 × 10−26
g__Mobiluncus0.01074.018 × 10−266.306 × 10−25
g__unknown_Murdochiella0.04974.575 × 10−256.222 × 10−24
g__Peptostreptococcus0.03764.423 × 10−256.222 × 10−24
g__Varibaculum0.00538.724 × 10−241.112 × 10−22
g__Peptococcus0.01844.041 × 10−234.849 × 10−22
g__unknown_Fastidiosipila0.00735.704 × 10−236.464 × 10−22
g__Pyramidobacter0.00342.621 × 10−202.814 × 10−19
g__unknown_Parvimonas0.01563.295 × 10−203.361 × 10−19
g__Negativicoccus0.00241.407 × 10−191.367 × 10−18
g__Escherichia-Shigella0.62941.608 × 10−181.491 × 10−17
g__Lactobacillus0.01206.432 × 10−175.652 × 10−16
g__unknown_Gallicola0.00156.650 × 10−175.652 × 10−16
g__Facklamia0.00071.215 × 10−159.911 × 10−15
g__Enterococcus0.02981.829 × 10−151.435 × 10−14
g__Gemella0.00473.007 × 10−152.272 × 10−14
g__Streptococcus0.72632.605 × 10−131.715 × 10−12
g__Negativibacillus0.10242.994 × 10−121.745 × 10−11
g__Prevotella1.93242.928 × 10−111.615 × 10−10
g__Staphylococcus0.00628.023 × 10−114.307 × 10−10
g__Actinomyces0.01033.922 × 10−102.000 × 10−9
g__Solobacterium0.00166.767 × 10−93.068 × 10−8
g__Ruminococcus torques group1.26612.016 × 10−87.758 × 10−8
g__Methanobrevibacter0.00032.611 × 10−89.863 × 10−8
g__Klebsiella0.00511.352 × 10−74.924 × 10−7
g__unknown_Atopobiaceae0.00131.624 × 10−75.812 × 10−7
g__Blautia4.50435.155 × 10−71.724 × 10−6
g__Sellimonas0.01956.426 × 10−72.081 × 10−6
g__Clostridium innocuum group0.00641.110 × 10−63.538 × 10−6
g__unknown_Peptostreptococcaceae0.24282.851 × 10−68.681 × 10−6
g__Dorea0.20933.859 × 10−61.158 × 10−5
g__Collinsella0.53575.222 × 10−61.522 × 10−5
g__Clostridium sensu stricto 10.12321.408 × 10−53.729 × 10−5
g__Desulfovibrio0.00364.232 × 10−51.066 × 10−4
g__Slackia0.00141.788 × 10−33.286 × 10−3
g__Anaerotruncus0.00401.480 × 10−22.305 × 10−2
g__unknown_Murib.302aculaceae0.00441.561 × 10−22.395 × 10−2
g__Paraprevotella0.00082.506 × 10−23.704 × 10−2
Table 5. Healthy-associated samples.
Table 5. Healthy-associated samples.
TaxonomyMedianp-Valueq-Value
g__Eubacterium eligens group1.99333.020 × 10−265.134 × 10−25
g__Monoglobus0.26256.650 × 10−154.845 × 10−14
g__Faecalibacterium9.58034.137 × 10−142.910 × 10−13
g__Lachnospiraceae NK4A136 group2.47901.731 × 10−131.177 × 10−12
g__UCG-0030.34458.407 × 10−135.197 × 10−12
g__GCA-9000665750.04461.344 × 10−128.064 × 10−12
g__unknown_Ruminococcaceae0.19301.042 × 10−115.902 × 10−11
g__Subdoligranulum2.93053.706 × 10−101.938 × 10−9
g__Alistipes2.01033.612 × 10−91.714 × 10−8
g__Coprobacter0.01748.986 × 10−93.819 × 10−8
g__Colidextribacter0.16491.162 × 10−84.838 × 10−8
g__Eubacterium ventriosum group0.13611.443 × 10−85.888 × 10−8
g__DTU0890.00821.874 × 10−87.444 × 10−8
g__Christensenellaceae R-7 group0.78012.074 × 10−77.296 × 10−7
g__Adlercreutzia0.05403.301 × 10−71.141 × 10−6
g__unknown_Clostridia UCG-0140.71404.265 × 10−71.450 × 10−6
g__Bacteroides21.56575.576 × 10−71.835 × 10−6
g__Oscillospira0.02272.372 × 10−67.331 × 10−6
g__Fusicatenibacter1.86784.795 × 10−61.418 × 10−5
g__Haemophilus0.02096.277 × 10−61.778 × 10−5
g__Oscillibacter0.16946.731 × 10−61.873 × 10−5
g__Anaerostipes0.37676.794 × 10−61.873 × 10−5
g__Erysipelotrichaceae UCG-0030.24231.202 × 10−53.226 × 10−5
g__Bifidobacterium0.86481.901 × 10−54.971 × 10−5
g__Butyricicoccus0.38582.780 × 10−57.088 × 10−5
g__UCG-0021.93355.882 × 10−51.463 × 10−4
g__unknown_Lachnospiraceae6.92466.741 × 10−51.637 × 10−4
g__UBA18190.01991.925 × 10−44.411 × 10−4
g__Veillonella0.02053.616 × 10−47.683 × 10−4
g__Barnesiella0.65695.741 × 10−41.195 × 10−3
g__NK4A214 group0.32001.046 × 10−32.072 × 10−3
g__unknown_Clostridia vadinBB60 group0.04481.087 × 10−32.132 × 10−3
g__UCG-0050.47691.141 × 10−32.197 × 10−3
g__Roseburia0.01641.355 × 10−32.536 × 10−3
g__Sutterella0.97262.152 × 10−33.785 × 10−3
g__UCG-0090.00586.594 × 10−31.112 × 10−2
g__Coprococcus0.94421.291 × 10−22.058 × 10−2
g__Eubacterium siraeum group0.13971.386 × 10−22.174 × 10−2
g__Parasutterella0.06001.505 × 10−22.325 × 10−2
g__Lachnospiraceae ND3007 group0.22391.706 × 10−22.577 × 10−2
Table 6. Top-20 Degree: Health-associated interactions.
Table 6. Top-20 Degree: Health-associated interactions.
Protein TypeUniprot IDDegreeNameSpecies
Bacterial ProteinsA0A1T5GC3829,343Cellulase (Glycosyl hydrolase family 5)Lachnospiraceae bacterium
A0A1T5D2V229,307Sugar phosphate isomerase/epimeraseLachnospiraceae bacterium
A0A1T5GAW129,307Thymidylate synthase (EC 2.1.1.45)Lachnospiraceae bacterium
A0A1T5BRU829,307ATPase/GTPase, AAA15 familyLachnospiraceae bacterium
A0A1T5CV3629,280GTP cyclohydrolase 1 type 2 homologLachnospiraceae bacterium
A0A1T5CWM929,259HPr Serine kinase C-terminal domain-containing proteinLachnospiraceae bacterium
A0A1T5EZK829,259ATPase/GTPase, AAA15 familyLachnospiraceae bacterium
A0A1T5FSD029,226Extracellular solute-binding proteinLachnospiraceae bacterium
A0A1T5BTE929,226Sugar phosphate isomerase/epimeraseLachnospiraceae bacterium
A0A1T5D3P229,097ATPase/GTPase, AAA15 familyLachnospiraceae bacterium
A0A1T5FKU629,094M6 family metalloprotease domain-containing proteinLachnospiraceae bacterium
A0A1T5FSP829,085Predicted dehydrogenaseLachnospiraceae bacterium
A0A1T5FSV429,079Sugar phosphate isomerase/epimeraseLachnospiraceae bacterium
A0A1T5C7W128,980L-ribulose-5-phosphate 3-epimeraseLachnospiraceae bacterium
A0A1T5EFX028,974Condensation domain-containing proteinLachnospiraceae bacterium
A0A1T5E9H328,854Ribose transport system permease proteinLachnospiraceae bacterium
A0A1T5CBP828,836Autoinducer 2 import system permease protein LsrDLachnospiraceae bacterium
A0A1T5G1V428,830Putative selenium metabolism hydrolaseLachnospiraceae bacterium
A0A1T5FWX828,830Putative peptidoglycan binding domain-containing proteinLachnospiraceae bacterium
A0A1T5FKR228,785D-alanyl-D-alanine carboxypeptidase (Penicillin-binding protein 5/6)Lachnospiraceae bacterium
Human ProteinsO00165693,752HCLS1-associated protein X-1 (HS1-associating protein X-1) (HAX-1) (HS1-binding protein 1) (HSP1BP-1)Homo sapiens (Human)
O15105664,028Mothers against decapentaplegic homolog 7 (MAD homolog 7) (Mothers against DPP homolog 7) (Mothers against decapentaplegic homolog 8) (MAD homolog 8) (Mothers against DPP homolog 8) (SMAD family member 7) (SMAD 7) (Smad7) (hSMAD7)Homo sapiens (Human)
Q92504634,221Zinc transporter SLC39A7 (Histidine-rich membrane protein Ke4) (Really interesting new gene 5 protein) (Solute carrier family 39 member 7) (Zrt-, Irt-like protein 7) (ZIP7)Homo sapiens (Human)
Q6E0U4626921Dermokine (Epidermis-specific secreted protein SK30/SK89)Homo sapiens (Human)
P54826614,560Growth arrest-specific protein 1 (GAS-1)Homo sapiens (Human)
Q9HC07599,200Putative divalent cation/proton antiporter TMEM165 (Transmembrane protein 165) (Transmembrane protein PT27) (Transmembrane protein TPARL)Homo sapiens (Human)
Q9NVW2586,421E3 ubiquitin-protein ligase RLIM (EC 2.3.2.27) (LIM domain-interacting RING finger protein) (RING finger LIM domain-binding protein) (R-LIM) (RING finger protein 12) (RING-type E3 ubiquitin transferase RLIM) (Renal carcinoma antigen NY-REN-43)Homo sapiens (Human)
Q15773586,010Myeloid leukemia factor 2 (Myelodysplasia-myeloid leukemia factor 2)Homo sapiens (Human)
Q9Y252581,832E3 ubiquitin-protein ligase RNF6 (EC 2.3.2.27)Homo sapiens (Human)
Q9BZR8576,099Apoptosis facilitator Bcl-2-like protein 14 (Bcl2-L-14) (Apoptosis regulator Bcl-G)Homo sapiens (Human)
Q9UHA4552,657Ragulator complex protein LAMTOR3 (Late endosomal/lysosomal adaptor and MAPK and MTOR activator 3) (MEK-binding partner 1) (Mp1) (Mitogen-activated protein kinase 1-interacting protein 1) (Mitogen-activated protein kinase scaffold protein 1)Homo sapiens (Human)
Q9BRP1535,201Programmed cell death protein 2-likeHomo sapiens (Human)
P62993535,073Growth factor receptor-bound protein 2 (Adapter protein GRB2) (Protein Ash) (SH2/SH3 adapter GRB2)Homo sapiens (Human)
Q969M3531,043Protein YIPF5 (Five-pass transmembrane protein localizing in the Golgi apparatus and the endoplasmic reticulum 5) (Smooth muscle cell-associated protein 5) (SMAP-5) (YIP1 family member 5) (YPT-interacting protein 1 A)Homo sapiens (Human)
Q9Y4L5528,189E3 ubiquitin-protein ligase RNF115 (EC 2.3.2.27) (RING finger protein 115) (RING-type E3 ubiquitin transferase RNF115) (Rab7-interacting RING finger protein) (Rabring 7) (Zinc finger protein 364)Homo sapiens (Human)
Q8N6T3517,063ADP-ribosylation factor GTPase-activating protein 1 (ARF GAP 1) (ADP-ribosylation factor 1 GTPase-activating protein) (ARF1 GAP) (ARF1-directed GTPase-activating protein)Homo sapiens (Human)
Q9BW91512,704ADP-ribose pyrophosphatase, mitochondrial (EC 3.6.1.13) (ADP-ribose diphosphatase) (ADP-ribose phosphohydrolase) (Adenosine diphosphoribose pyrophosphatase) (ADPR-PPase) (Nucleoside diphosphate-linked moiety X motif 9) (Nudix motif 9)Homo sapiens (Human)
Q02535511,921DNA-binding protein inhibitor ID-3 (Class B basic helix-loop-helix protein 25) (bHLHb25) (Helix-loop-helix protein HEIR-1) (ID-like protein inhibitor HLH 1R21) (Inhibitor of DNA binding 3) (Inhibitor of differentiation 3)Homo sapiens (Human)
Q9H0V1511,224Transmembrane protein 168Homo sapiens (Human)
Q9NZ45510,746CDGSH iron-sulfur domain-containing protein 1 (Cysteine transaminase CISD1) (EC 2.6.1.3) (MitoNEET)Homo sapiens (Human)
Table 7. Top-20 Degree: CRC-associated interactions.
Table 7. Top-20 Degree: CRC-associated interactions.
Protein TypeUniprot IDDegreeNameSpecies
Bacterial ProteinsR6WAI119,562ABC-type transport system substrat-binding componentRuminococcus sp. CAG:382
R6XB8219,562Outer membrane protein SusFPrevotella sp. CAG:732
S0IX7619,556Solute-binding protein family 5 domain-containing proteinEubacterium sp. 14-2
R6XFW819,556Putative cellulasePrevotella sp. CAG:732
Q8RG2819,556Alanine racemase (EC 5.1.1.1)Fusobacterium nucleatum subsp. nucleatum (strain ATCC 25586/DSM 15643/BCRC 10681/CIP 101130/JCM 8532/KCTC 2640/LMG 13131/VPI 4355)
A0A173W5P619,556Lactose-binding protein[Ruminococcus] torques
R7KUC719,556ABC transporter solute-binding proteinRuminococcus sp. CAG:353
Q8REC719,554Type I-B CRISPR-associated protein Cas7/Cst2/DevRFusobacterium nucleatum subsp. nucleatum (strain ATCC 25586/DSM 15643/BCRC 10681/CIP 101130/JCM 8532/KCTC 2640/LMG 13131/VPI 4355)
R6Q3S419,554Outer membrane protein SusF/SusE-like C-terminal domain-containing proteinPrevotella sp. CAG:386
R6XJ0119,548Concanavalin A-like lectin/glucanases family proteinPrevotella sp. CAG:732
R5LV0519,548Outer membrane protein SusEPrevotella sp. CAG:1185
R6XA8919,548Putative iron-sulfur cluster-binding proteinRuminococcus sp. CAG:382
R6W4K019,546Carbohydrate-binding domain-containing proteinRuminococcus sp. CAG:382
R6F9J019,542Right-handed beta helix domain-containing proteinPrevotella sp. CAG:520
A0A239RL2519,542Thymidylate synthasePrevotellaceae bacterium KH2P17
R5PWH519,542site-specific DNA-methyltransferase (adenine-specific) (EC 2.1.1.72)Prevotella sp. CAG:1092
R7KSV119,542Glycoside hydrolase family 16Ruminococcus sp. CAG:353
R6W02119,542Oligopeptide ABC superfamily ATP binding cassette transporter binding proteinRuminococcus sp. CAG:382
R5PW2719,540LipoproteinPrevotella sp. CAG:1092
R6EGC319,540Concanavalin A-like lectin/glucanases family proteinPrevotella sp. CAG:1320
Human ProteinsO00165618,870HCLS1-associated protein X-1 (HS1-associating protein X-1) (HAX-1) (HS1-binding protein 1) (HSP1BP-1)Homo sapiens (Human)
O15105591,247Mothers against decapentaplegic homolog 7 (MAD homolog 7) (Mothers against DPP homolog 7) (Mothers against decapentaplegic homolog 8) (MAD homolog 8) (Mothers against DPP homolog 8) (SMAD family member 7) (SMAD 7) (Smad7) (hSMAD7)Homo sapiens (Human)
Q92504563,786Zinc transporter SLC39A7 (Histidine-rich membrane protein Ke4) (Really interesting new gene 5 protein) (Solute carrier family 39 member 7) (Zrt-, Irt-like protein 7) (ZIP7)Homo sapiens (Human)
Q6E0U4557,079Dermokine (Epidermis-specific secreted protein SK30/SK89)Homo sapiens (Human)
P54826545,772Growth arrest-specific protein 1 (GAS-1)Homo sapiens (Human)
Q9HC07531,560Putative divalent cation/proton antiporter TMEM165 (Transmembrane protein 165) (Transmembrane protein PT27) (Transmembrane protein TPARL)Homo sapiens (Human)
Q9NVW2519,970E3 ubiquitin-protein ligase RLIM (EC 2.3.2.27) (LIM domain-interacting RING finger protein) (RING finger LIM domain-binding protein) (R-LIM) (RING finger protein 12) (RING-type E3 ubiquitin transferase RLIM) (Renal carcinoma antigen NY-REN-43)Homo sapiens (Human)
Q15773519,551Myeloid leukemia factor 2 (Myelodysplasia-myeloid leukemia factor 2)Homo sapiens (Human)
Q9Y252515,797E3 ubiquitin-protein ligase RNF6 (EC 2.3.2.27)Homo sapiens (Human)
Q9BZR8510,681Apoptosis facilitator Bcl-2-like protein 14 (Bcl2-L-14) (Apoptosis regulator Bcl-G)Homo sapiens (Human)
Q9UHA4489,358Ragulator complex protein LAMTOR3 (Late endosomal/lysosomal adaptor and MAPK and MTOR activator 3) (MEK-binding partner 1) (Mp1) (Mitogen-activated protein kinase 1-interacting protein 1) (Mitogen-activated protein kinase scaffold protein 1)Homo sapiens (Human)
Q9BRP1473,519Programmed cell death protein 2-likeHomo sapiens (Human)
P62993473,461Growth factor receptor-bound protein 2 (Adapter protein GRB2) (Protein Ash) (SH2/SH3 adapter GRB2)Homo sapiens (Human)
Q969M3469,851Protein YIPF5 (Five-pass transmembrane protein localizing in the Golgi apparatus and the endoplasmic reticulum 5) (Smooth muscle cell-associated protein 5) (SMAP-5) (YIP1 family member 5) (YPT-interacting protein 1 A)Homo sapiens (Human)
Q9Y4L5467,166E3 ubiquitin-protein ligase RNF115 (EC 2.3.2.27) (RING finger protein 115) (RING-type E3 ubiquitin transferase RNF115) (Rab7-interacting RING finger protein) (Rabring 7) (Zinc finger protein 364)Homo sapiens (Human)
Q8N6T3457,183ADP-ribosylation factor GTPase-activating protein 1 (ARF GAP 1) (ADP-ribosylation factor 1 GTPase-activating protein) (ARF1 GAP) (ARF1-directed GTPase-activating protein)Homo sapiens (Human)
Q9BW91453,329ADP-ribose pyrophosphatase, mitochondrial (EC 3.6.1.13) (ADP-ribose diphosphatase) (ADP-ribose phosphohydrolase) (Adenosine diphosphoribose pyrophosphatase) (ADPR-PPase) (Nucleoside diphosphate-linked moiety X motif 9) (Nudix motif 9)Homo sapiens (Human)
Q02535452,521DNA-binding protein inhibitor ID-3 (Class B basic helix-loop-helix protein 25) (bHLHb25) (Helix-loop-helix protein HEIR-1) (ID-like protein inhibitor HLH 1R21) (Inhibitor of DNA binding 3) (Inhibitor of differentiation 3)Homo sapiens (Human)
Q9H0V1452,025Transmembrane protein 168Homo sapiens (Human)
Q9NZ45451,580CDGSH iron-sulfur domain-containing protein 1 (Cysteine transaminase CISD1) (EC 2.6.1.3) (MitoNEET)Homo sapiens (Human)
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Kiouri, D.P.; Batsis, G.C.; Messaritakis, I.; Souglakos, J.; Chasapis, C.T. Mapping of Phenotype Specific Host–Microbiome Protein–Protein Interaction Networks in Colorectal Cancer Using Deep Learning. Int. J. Mol. Sci. 2026, 27, 4232. https://doi.org/10.3390/ijms27104232

AMA Style

Kiouri DP, Batsis GC, Messaritakis I, Souglakos J, Chasapis CT. Mapping of Phenotype Specific Host–Microbiome Protein–Protein Interaction Networks in Colorectal Cancer Using Deep Learning. International Journal of Molecular Sciences. 2026; 27(10):4232. https://doi.org/10.3390/ijms27104232

Chicago/Turabian Style

Kiouri, Despoina P., Georgios C. Batsis, Ippokratis Messaritakis, John Souglakos, and Christos T. Chasapis. 2026. "Mapping of Phenotype Specific Host–Microbiome Protein–Protein Interaction Networks in Colorectal Cancer Using Deep Learning" International Journal of Molecular Sciences 27, no. 10: 4232. https://doi.org/10.3390/ijms27104232

APA Style

Kiouri, D. P., Batsis, G. C., Messaritakis, I., Souglakos, J., & Chasapis, C. T. (2026). Mapping of Phenotype Specific Host–Microbiome Protein–Protein Interaction Networks in Colorectal Cancer Using Deep Learning. International Journal of Molecular Sciences, 27(10), 4232. https://doi.org/10.3390/ijms27104232

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop