1. Introduction
Hippocrates, the father of medicine, famously stated that “All disease begins in the gut,” a prescient observation that reflects the centrality of intestinal health to systemic well-being [
1]. Over the past two decades, advances in sequencing technologies have underscored the importance of the gut microbiome (GM), a densely populated microbial ecosystem comprising bacteria, viruses, fungi, archaea, and protozoa [
2]. Among these, bacteria have been most extensively characterized, with human microbiome projects identifying over 2000 species spanning multiple phyla, including
Firmicutes,
Bacteroidetes,
Actinobacteria,
Proteobacteria, and
Verrucomicrobia [
3]. These organisms play critical roles in nutrient metabolism, immune system development, and epithelial barrier function, earning the microbiome the designation of a “super-organ” [
4,
5,
6,
7].
The balance of microbial populations is crucial for maintaining homeostasis. Factors such as pH, bile acids, oxygen tension, and host immune responses shape microbial communities along the gastrointestinal tract [
8]. When this balance is disrupted, a state of dysbiosis arises, which has been associated with numerous conditions, ranging from inflammatory bowel disease and irritable bowel syndrome [
9] to metabolic disorders [
10], cardiovascular diseases [
11], and autoimmune conditions [
12]. Dysbiosis also affects the gut–brain axis, a bidirectional communication system linking the enteric and central nervous systems, and has been implicated in neurodevelopmental [
13] and psychiatric conditions [
14] such as autism spectrum disorder, major depression, and Parkinson’s disease [
15,
16]. These diverse links highlight the microbiome’s wide-reaching influence across human physiology and pathology.
Within this broader landscape, colorectal cancer (CRC) has emerged as one of the most compelling examples of a disease with direct microbial involvement. CRC is the second most common cancer and the third leading cause of cancer-related death worldwide [
17]. The rising incidence of early-onset CRC, diagnosed before the age of 50, further emphasizes the urgency of understanding non-genetic drivers of disease. Importantly, more than 80% of CRC cases are sporadic and arise from complex interactions between host genetics, environmental factors, and the gut microbiome rather than inherited mutations [
18]. The microbiome is therefore positioned not only as a marker of CRC risk but also as a potential mediator of disease onset and progression.
Specific bacterial taxa have been consistently associated with CRC. Pathogens such as
Helicobacter pylori [
19],
Salmonella [
18], and
Campylobacter jejuni [
20] contribute to carcinogenesis via chronic inflammation, toxin production, and genotoxic effects. Other taxa, including
Bacteroides fragilis,
Escherichia coli, and
Streptococcus gallolyticus [
21], participate in biofilm formation and secrete metabolites that modulate epithelial signaling pathways [
22].
Fusobacterium nucleatum [
23] and
Porphyromonas gingivalis, in particular, have been linked to chemoresistance in CRC, underscoring the clinical impact of microbial activity on treatment outcomes [
24]. Collectively, these findings suggest that CRC is not merely influenced by microbial imbalance but may in fact be driven by direct microbial contributions.
The molecular mechanisms that underpin these associations are increasingly recognized as pivotal for advancing both diagnostics and therapy. In recent years, research has shifted from taxonomic descriptions to mechanistic investigations at the molecular level. Protein–protein interactions (PPIs) between microbial proteins and host proteins represent one such mechanism. These interactions can alter host signaling networks, immune responses, and cellular processes in ways that favor tumor initiation or progression. While experimental approaches such as yeast two-hybrid assays, cross-linking mass spectrometry, and photo-reactive probes have shed light on microbe–host interactions [
25,
26,
27,
28,
29,
30,
31], the scale of experimental mapping remains limited, leaving large portions of the interactome unexplored. A critical limitation in our current understanding is the paucity of experimental evidence detailing protein interactions between gut bacteria and the human host. While general human–bacterial interaction data exist in public repositories, the experimentally determined pan-human–bacterial protein interactome consists of fewer than 20,000 interactions. This limited scope, especially when contrasted with the 300–500 bacterial species estimated to inhabit the human gut, creates a research deficit that likely obstructs a deeper comprehension of how GM–host imbalances drive disease, highlighting that these molecular dialogs are largely unmapped.
Consequently, computational methods—including domain–domain interaction modeling and machine learning algorithms—have become indispensable for predicting host–microbe PPIs [
32,
33,
34,
35,
36].
In this context, this work introduces a computational approach to map these crucial molecular dialogs. Next-Generation Sequencing (NGS) was applied to meticulously profile bacterial strain abundances within the gut microbiomes of both healthy individuals and colorectal cancer patients. Building on this detailed community characterization, we then applied a novel Deep Learning (DL) algorithm specifically designed to PPIs between proteins from the identified bacterial strains and proteins of the human host’s gut. The objective of this interdisciplinary study is to uncover potential protein interaction networks that could explain how specific bacterial strains contribute to, or protect against, colorectal cancer development through direct molecular engagement with the host.
3. Discussion
Beyond colorectal cancer, dysbiosis of the gut microbiome has been implicated in a wide spectrum of human diseases. Altered microbial populations have been associated with neurodevelopmental and psychiatric conditions such as autism spectrum disorder, attention deficit hyperactivity disorder, depression, Alzheimer’s, and Parkinson’s disease [
15,
16]. Dysbiotic microbiota have also been linked to metabolic disorders, including obesity, type 2 diabetes, and non-alcoholic fatty liver disease [
38,
39], as well as cardiovascular diseases like atherosclerosis and hypertension [
40]. Additionally, autoimmune diseases such as rheumatoid arthritis, multiple sclerosis, and type 1 diabetes have shown connections with gut microbial imbalances [
12,
41].
Among the various diseases influenced by the gut microbiome, colorectal cancer (CRC) stands out due to the particularly strong and consistent evidence linking microbial dysbiosis to its pathogenesis. A key characteristic of the CRC gut microbiome is its altered composition of bacterial strains relative to healthy individuals. Until now, the bacterium that has been mostly associated with the development of CRC is
Helicobacter pylori. In line with its role as a potent pathogen, individuals infected with
Helicobacter pylori harbor a nearly twofold increased risk to develop CRC [
19]. Several large-scale epidemiological studies in the Netherlands strongly indicate that
Salmonella infection elevates the risk of CRC. Key findings include a standardized incidence ratio (SIR 1.54) of early-onset CRC in the proximal colon following
Salmonella exposure, a sustained higher risk associated with
non-Enteritis or
Typhimurium serovars and increased serological markers of
Salmonella exposure (FliC antibodies) in CRC patients [
18]. Furthermore, inflammation resulting from
Salmonella colonization may be a contributing factor to this increased CRC risk [
18]. Another bacterial strain that has been linked to CRC is
Campylobacter jejuni, which produces a DNA-altering cytolethal distending toxin (CDT) [
20]. Although it is known that CDT contributes to the development of inflammation in the GI tract, it was recently demonstrated that not only the production of cdtB advances CRC and promotes metastasis, but also that even the presence of this type of bacteria can affect the components and transcriptional activity of the gut microbial population [
20,
42]. A study conducted in Taiwan that analyzed pyogenic liver abscesses (PLA), an early sign of CRC, caused by
Klebsiella pneumoniae, revealed that those strain-specific abscesses resulted in a far greater rate of subsequent CRC than abscesses derived from other bacteria [
43].
Clostridium difficile is another bacterium that is implicated in gastrointestinal infections and antibiotic-associated colitis.
C. difficile produces three different toxins, two of which (i.e., toxin A (TcdA), toxin B (TcdB)) cause detrimental effects to the GI tract’s epithelial barrier, damage the cells’ genetic material as well as activate STAT3 and NF-κΒ chronic inflammation-related pathways, potentially triggering CRC pathogenesis [
44]. Even though
Bacteroides fragilis and
Escherichia coli are a normal part of the enteric microbiome, some pathogenic strains of them have been linked to CRC [
45]. Toxin-expressing strains of
B. fragilis (
Bacteroides fragilis toxin (bft)) and
E. coli (colibactin (clbB)) have not only been spatially associated in biofilms in the gut, but also their synergistic pro-carcinogenic involvement in CRC has emerged [
22]. A recent study by Ding et. al. revealed that
B. fragilis promotes chemoresistance in CRC, and at the same time, phage elimination experiments conducted in mice uncovered restored chemosensitivity of those CRC cells [
46]. A great number of studies have connected CRC pathology with
B. fragilis, but its exact role remains unclear [
47,
48,
49].
Streptococcus gallolyticus is an opportunistic pathogen that has been associated with numerous studies with stimulation of cell reproduction and augmentation of tumor burden [
21]. Additionally,
Fusobacterium nucleatum also promotes chemoresistance of CRC cells through modulation of autophagy [
23]. Although
Enterococcus faecalis has been described as both a stimulant and a protector against CRC [
50], recent studies have elucidated the role of biliverdin (i.e., one of its metabolites) as a tumor-stimulating compound that affects the host’s PI3K/AKT/mTOR pathway and thus promotes cell proliferation and angiogenesis [
51,
52]. Moreover, an experiment conducted by Chang et al. demonstrated that
Parvimonas micra activated the Ras/ERK/c-Fos signaling pathway via micro-RNA upregulation and enhanced cellular proliferation in CRC [
53]. Other studies have also highlighted
P. micra’s implication in the immune response of CRC patients, potentially serving as a predictive biomarker for poor patient survival in CRC [
54].
Peptostreptococcus anaerobius interacts directly with colonic cells via one of its surface proteins and also activates the proliferation-related integrin α2/β1-PI3K-Akt-NF-κB pathway [
55]. Furthermore,
P. anaerobius has been shown to intensify chemoresistance to oxaliplatin [
56].
Porphyromonas gingivalis contributes to the proliferation of colorectal cancer cells through a mechanism involving cellular invasion and the subsequent activation of the MAPK/ERK signaling pathway [
24]. Besides this role, a recent experiment showed that
P. gingivalis upregulates chitinase 3-like-1 protein (CHI3L1) in invariant natural killer T (iNKT) cells, leading to detrimental effects on their cytotoxic function that result in the immune evasion of the tumors [
57].
In parallel, recent studies highlight the importance of protein-level interactions between gut microbiota and the host. However, the experimental interactome remains limited, motivating the use of computational methods, including machine learning and domain-domain interaction prediction, to expand our understanding [
32,
33,
34,
35,
36]. By integrating computational predictions with clinical microbiome data, this work contributes to mapping the molecular dialogs that underlie CRC development and potentially other microbiome-associated diseases.
The two phenotype-specific subnetworks, which were created after the categorization of the bacteria genera by phenotype association (i.e., health and CRC associated) after excluding the non-associated genera, revealed that the human proteins that are present in both subnetworks are the same. Interestingly, those human proteins also, when ranked according to their centrality degree, remain in the same order in both cases. This finding highlights not only the robustness of the human protein interactome and its key interactors but also emphasizes the complexity of the human proteins that can interact with different proteins at different health statuses. These most connected human proteins in these networks are mainly metalloproteins (E3 ubiquitin protein ligases, zinc transporters, etc.), transmembrane proteins, proteins related to ADP ribosylation and proteins related to mechanisms of cell proliferation and survival (i.e., apoptosis, programmed cell death, growth factors). It has been demonstrated that various species of pathogenic bacteria encode E3 ligases that have the ability to hijack the host’s ubiquitination mechanisms via a series of different strategies, including mimicking host-derived E3 ligases and encoding novel E3 ligases, as well as encoding deubiquitinases. Those proteins are transported into the host cell via type III or type IV secretion systems (T3SS and T4SS, respectively) and can then manipulate the host’s machinery for their proliferation [
58,
59]. Furthermore, there are studies that demonstrate the influence of the host’s zinc transporters on the homeostasis of the gut. A recent study showed that deletion of zinc (Zn) transporter ZIP14 creates a reduction in Zn in the entire intestinal tract and ultimately leads to lower microbial diversity [
60,
61]. Additionally, ATP-binding cassette (ABC) transporters have been associated with the microbial populations of the gut. More specifically, gut bacteria have been shown to both utilize them as binding receptors but also up- and downregulate their expression [
62]. Research has also demonstrated that the bacterial ADP-ribosylation system, and mainly ADP-ribosyl transferases (ARTs), irreversibly modify host proteins with key function to the cellular cycle [
63]. For example, Bxa of
Bacteroides modifies non-muscle myosin II and triggers cellular remodeling that ultimately leads to inosine secretion that is then used by the microorganism as a carbon source [
64]. Finally, since the gut bacteria have long been associated with cancer, mainly types of cancer that affect the GI tract. Therefore, it is no surprise that the gut microbiome influences growth factors, like the insulin-like growth factor 1 (IGF-1), which is essential for bone growth [
65]. With respect to apoptosis, gut bacteria exert dual effects by either promoting or inhibiting cell death. Certain pathogens, such as non-typhoidal
Salmonella, induce macrophage apoptosis via SPI-1 expression, thereby limiting inflammatory cytokine production [
66]. Conversely, many bacterial species actively suppress host cell apoptosis to evade efferocytosis and enhance survival within the host [
67].
On one hand, in the health-specific network, the most important proteins are those belonging to
Lachnospiraceae, which are among the most abundant taxa in the GI tract. Evidence from multiple studies suggests that members of the
Lachnospiraceae family contribute to maintaining the host’s physiological functions [
68]. More specifically, they produce short-chain fatty acids that are converted into secondary bile acids that hinder the colonization of pathogenic strains [
69].
On the other hand, in the bacteria-specific network, the most important proteins include several outer membrane, transport, and regulatory elements with distinct roles in microbial adaptation and survival. Among them, SusE and SusF are two outer membrane proteins from
Bacteroides composed of tandem starch-specific carbohydrate-binding modules (CBMs) [
70]. Although they lack enzymatic activity, these proteins are thought to play an essential role in starch metabolism by sequestering polysaccharides at the bacterial surface, thereby limiting access to competitor host cells [
70]. Additional key proteins include alanine racemase, an enzyme required for the conversion of L-alanine to D-alanine and thus the synthesis of peptidoglycan, a critical component of bacterial cell walls [
71], and ATP-binding cassette (ABC) transporters, which form a large superfamily of membrane complexes responsible for nutrient uptake, protein secretion, and resistance to environmental stressors, but also drug transfer [
72,
73]. Regulatory and DNA-binding proteins also contribute substantially to this network. These include lactose-binding proteins, such as the LacI repressor, which controls metabolic gene expression [
74]. Enzymes such as thymidylate synthase, essential for deoxythymidine monophosphate (dTMP) synthesis [
75], and adenine-specific DNA methyltransferases, which protect bacterial genomes from restriction enzymes [
76], further emphasize the centrality of DNA-modifying activities. Finally, metalloproteins, particularly those containing iron–sulfur clusters, act as critical cofactors in electron transfer, catalysis, and gene regulation [
77]. Collectively, these proteins reflect the molecular strategies employed by bacteria to secure nutrients, maintain genomic stability, and compete effectively within the intestinal ecosystem.
Apart from the most central proteins, proteins with centrality values near the mean also merit particular attention. From a graph-theoretical perspective, these nodes provide critical redundancy within the network: if only the most central proteins were perturbed, network function would collapse unless compensated by the surrounding intermediate nodes. Pathway enrichment analysis revealed that these moderately central proteins are implicated in the same biological pathways as the top-ranking hubs, thereby reinforcing their functional relevance. Their involvement suggests that they may act as auxiliary regulators or stabilizers, ensuring continuity of pathway activity and buffering the network against perturbations that target the primary hubs. Collectively, this highlights that both highly central and near-mean centrality proteins contribute to the robustness of bacteria–host interactions, and their combined roles are essential for maintaining network integrity under physiological and pathological conditions. The table with all the host and bacterial proteins with centrality values near the mean for both disease states can be found in
Supplementary Materials File S4.
Interestingly, analysis of cross-distribution patterns revealed that CRC-associated taxa were also detectable within healthy samples, while health-associated taxa were observed in CRC samples. Quantitative analysis demonstrated that 48 bacterial taxa classified as CRC-associated were also detected in healthy samples, with an average relative abundance of approximately 0.27. Conversely, 40 taxa typically considered health-associated were found in CRC samples, with a higher mean abundance of about 1.03. These findings suggest that bacterial classification as CRC- or health-associated is not absolute but rather context-dependent, reflecting shifts in abundance and ecological balance rather than strict presence or absence. The observation that these taxa appear across both sample groups underscores the importance of relative abundance and community composition in shaping host–microbiome interactions and highlights that disease associations are likely driven by dysbiosis and altered network dynamics rather than by individual taxa alone. The tables with the Health-associated strains present in CRC patients and the CRC-associated strains in Healthy individuals, along with their relative abundances, can be found in
Supplementary Materials Files S5 and S6.
Concerning the DL-based methodology, the integration of ProstT5-based embeddings and clinical metagenomics offers a high-resolution view of the CRC interactome, but it is essential to acknowledge the inherent constraints of a deep-learning-driven approach. The scientific rigor of this work is supported by the deliberate selection of primary, peer-reviewed repositories such as HPIDB, IntAct, and PHISTO for training, ensuring the model is grounded in the most comprehensive and experimentally validated interspecies data available. However, the dependence on high-fidelity input from predictive frameworks like AlphaFold and the “black-box” nature of deep neural networks present interpretability challenges regarding the precise biochemical drivers of each interaction. To mitigate the risks of dataset bias and overfitting common in high-dimensional biological spaces, we implemented technical safeguards, including focal loss functions and strict early stopping protocols. By prioritizing transparency in data filtering and calibrating thresholds to ensure a precision-driven interactome, this study provides a robust framework for identifying novel therapeutic targets and biomarkers while recognizing that these findings reflect computational predictions that require subsequent mechanistic validation. Finally, a limitation of this study is the lack of external validation in independent and ethnically diverse cohorts. While such validation is essential for assessing the generalizability of our findings, the present work was designed as a proof-of-concept study focusing on the development and internal evaluation of a deep learning framework for PPI mapping in a well-characterized colorectal cancer cohort from Crete/Greece. This setting provided a relatively homogeneous population, which facilitated controlled model development and reduced confounding variability during the initial phase of analysis. Future studies should aim to validate these findings across multicenter cohorts and diverse populations, as well as integrate additional omics datasets to further evaluate robustness and translational potential.