Next Article in Journal
Mechanisms of Diapause in Major Crustacean Taxa: Environmental Cues, Gene Regulation, and Adaptive Strategies
Previous Article in Journal
Benthic Community Composition of Coral Reefs Is Associated with the Time Since Last Blast Fishing Event
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Data Mining of Groundwater Genomes for Metagenome-Assembled Genomes (MAGs) Containing Monooxygenase Operons Associated with Contaminant Biodegradation

by
Alison M. Cupples
*,
James Richards
and
Maria Basaldua Del Cid
Department of Civil and Environmental Engineering, Michigan State University, East Lansing, MI 48824, USA
*
Author to whom correspondence should be addressed.
Biology 2026, 15(18), 1586; https://doi.org/10.3390/biology15181586
Submission received: 7 August 2026 / Revised: 31 August 2026 / Accepted: 3 September 2026 / Published: 9 September 2026
(This article belongs to the Section Bioinformatics)

Simple Summary

Environmental contaminants are often degraded by bacteria already present in soil and water. The occurrence of particularly active bacteria can be investigated by targeting specific nucleic acid (DNA) sequences. The objective of the current study was to investigate freely available nucleic acid sequences obtained from numerous previous studies across multiple countries to determine which species of active bacteria are common and which active genes they contain. For this, the Department of Energy Systems Biology Knowledgebase (KBase) was used to analyze the raw data and create assembled genomes for each microorganism of interest. The sequences of numerous active genes were compared between these assembled genomes. These sequences were then made available to the public through KBase websites. The most commonly detected active genes included those encoding for the enzymes methane monooxygenase, propane monooxygenase, phenol monooxygenase and toluene monooxygenase. The information will be particularly useful to environmental engineers involved in the design of molecular assays for the quantification of active bacteria in contaminated groundwater and soils.

Abstract

This study examined freely available whole genome sequencing (WGS) data for genes associated with contaminant biodegradation. Thirteen WGS datasets (>600 individual samples) involving more than 12,000 Gbases from multiple countries were examined. The Department of Energy Systems Biology Knowledgebase (KBase) was used to create metagenome-assembled genomes (MAGs) containing the operons of interest. The arrangement and length of subunits for each operon were compared. Phylogenetic trees were created for common biomarkers (tmoA, pmoA, prmA, mmoX, dmpN). MAGs were uploaded into publicly available KBase narratives. Thirty-four MAGs, within the phyla Actinomycetota, Pseudomonadota and Chloroflexota, were identified with the full propane monooxygenase operon (prmABCD). Twenty-six MAGs, within the classes Gammaproteobacteria and Alphaproteobacteria, contained the full operon for soluble methane monooxygenase (mmoXYBZDC). More than 100 MAGs contained the full operon for ammonia/particulate methane monooxygenase (pmoCAB) and were classified within the Gammaproteobacteria and Alphaproteobacteria groups, as well as other phyla. Fifty-five MAGs, within Burkholderiales (Gammaproteobacteria) and Alphaproteobacteria, contained the full operon for toluene-4-monooxygenase (tmoABCDEF). From the MAGs containing the full operon for toluene monooxygenase, thirty-three also contained the full operon for phenol monooxygenase (dmpKLMNOP). The MAGs generated and their associated functional gene sequences have the potential to improve molecular detection methods for site bioremediation.

1. Introduction

A promising approach for site bioremediation is the use of aerobic metabolism and co-metabolism because these processes can degrade multiple contaminants, and there is the potential to degrade these contaminants to low levels [1,2,3,4,5,6,7,8,9]. A number of oxygenases and microorganisms have been associated with such detoxification processes. In particular, a large amount of evidence exists for the co-metabolic biodegradation of contaminants comprising methane, propane, toluene or phenol monooxygenases. As bioremediation often involves multiple lines of evidence for proof of effective biodegradation, engineers frequently monitor the occurrence of the genes encoding for these enzymes at contaminated sites. To optimize these monitoring approaches, it is important to understand the diversity associated with these biomarkers across groundwater metagenomes.
Early research on contaminant co-metabolism focused on methanotrophs, microorganisms that use methane as a sole carbon and energy source. These microorganisms contain methane monooxygenase (MMO), an enzyme associated with the oxidation of methane to methanol. This enzyme is occurs in a soluble cytoplasmic form (sMMO) or a particulate membrane-associated form (pMMO) [10]. Particulate methane monooxygenase has been detected in all methanotrophs, except for in the genera Methylocella [11] and Methyloferula [12], whereas sMMO has been found in fewer strains [13]. The substrate range for sMMO includes monoaromatics, diaromatics, alkanes, alicyclic hydrocarbons, alkenes, and substituted methane derivatives [10]. Particulate methane monooxygenase is similar to ammonia monooxygenase from ammonia-oxidizing bacteria [14,15] and has been linked to the degradation of TCe and short linear hydrocarbons [10]. Propane monooxygenase is another enzyme important for the removal of environmental pollutants. For example, indigenous microorganisms degraded 1,4-dioxane, TCE and 1,2-dichloroethane when propane and oxygen were added to groundwater [16]. Other co-metabolic contaminant degraders include microorganisms associated with toluene oxidation [17,18,19,20]. For example, toluene-4-monooxygenase (encoded by tmoABCDEF) [21,22,23] has been linked with the degradation of N-nitrosodimethylamine; polycyclic aromatic substrates; cDCE; TCE; 1,1-DCE; 1,4-dioxane; alkenes; chloroform and substituted benzenes [22,23,24,25,26,27,28,29,30,31]. Catechol and substituted catechols are intermediate chemicals in the biodegradation of many pollutants [32,33,34]. For example, phenol monooxygenase from Pseudomonas sp. strain CF600 is important for the biodegradation of phenol to catechol [35].
Knowledge regarding the occurrence of microorganisms with such degradative capabilities should enhance our understanding of the removal of many contaminants. Typically, quantitative PCR (qPCR) is used to determine the presence and abundance of these functional genes in environmental samples. However, these assays are frequently based on pure cultures sequences and often do not take into account the gene sequences actually present in groundwater or subsurface sediments. For this evaluation, the current study examined freely available whole genome sequencing (WGS) data to determine gene sequences actually found in environmental groundwater. In the past decade, WGS has exploded as an investigative tool for mixed microbial communities across a large number of disciplines. However, the use of this method in environmental engineering to examine bioremediation potential at contaminated sites is still in its infancy, primarily due to (1) sequencing costs and (2) challenges relating to data analysis. The development of an on-line sequencing analysis platform, the Department of Energy Systems Biology Knowledgebase, or KBase [36]., has the potential to remove one of these limitations. KBase is free, easy to use, sharable across users and offers fully developed tools (e.g., quality control, annotation, creation of metagenome-assembled genomes) to analyze WGS datasets. The increase in publicly available WGS data from environmental samples from the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA) database has the potential to remove the other limitation. There are several hundreds of gigabytes from WGS runs from environmental samples, and these datasets have yet to be tapped as a resource for understanding contaminant biodegradation potential.
The current project involved several objectives. One objective was to collect and analyze freely available groundwater sequencing data (WGS datasets) for the functional genes (listed above) previously associated with contaminant biodegradation from multiple studies across multiple countries. A second objective was to summarize the phylotypes associated with these functional genes. A third objective was to compare the gene sequences of the subunits commonly used as biomarkers for each functional gene (tmoA, pmoA, prmA, mmoX, dmpN) and, when applicable, to identify patterns for each di-iron center. A final objective was to make the resulting data freely available so others can update genetic detection methods for these key genes. For this, metagenome-assembled genomes (MAGs), with full operons for the targeted genes, were collected and made available in a group of KBase narratives.

2. Methods

2.1. Selection of Publicly Available Sequencing Data

Thirteen WGS datasets, with more than 600 individual samples, involving more than 12,000 Gbases, were downloaded from previous studies focusing on the molecular analysis of groundwater (Table 1). The datasets originated from nine different countries (USA, New Zealand, Japan, Canada, Sweden, China, Bangladesh, Germany and The Netherlands). The datasets were selected based on the size of data for each sample (i.e., larger files), as preliminary data analysis with smaller files did not generate MAGs with the desired genes. To obtain the datasets, NCBI’s SRA database was searched using keywords such as “groundwater” and “shotgun sequencing” or “WGS”. The results were analyzed further to exclude those that were not relevant (e.g., amplicon sequencing data, non-typical groundwater). The datasets were downloaded to the High Performance Computing Cluster (HPCC) at MSU using the SRA toolkit, as most files were >5 GB. The International Nucleotide Sequence Database Collaboration policy allows unrestricted access to all data records to enable publication, with appropriate credit being given by citing the original submission (accession numbers). The final datasets were obtained primarily from groundwater studies (pristine as well as contaminated sites), with a range of samples (from 6 to >300) and a total number of bases. However, we recognize that microorganisms in pristine environments may not survive at contaminated sites. Additional information for each dataset and study can be found in the references listed (Table 1).

2.2. Analysis of Whole Genome Sequencing Data

The WGS files were uploaded from HPCC to the United States Department of Energy Systems Biology Knowledgebase (KBase) [36] using Globus v3.3.0 (freely available at MSU). Due to the large amount of data, more than eighty narratives were generated. For example, the data collected from BioProject Accession number PRJNA858913 was analyzed in forty narratives. The steps involved for the analysis of each narrative were the same. The sequencing files were first subjected to quality control and filtering using FastQC (version 0.12.1) [53] as well as Trimmomatic (version 0.36) (sliding window size of 4; sliding window minimum quality of 15) [54]. Megahit (version 1.2.9) [55] was then used to assemble the read libraries using 2000 bp as a minimum read length. RASTtk (version 1.073) [56] and Prokka (version 1.14.16) [57] were then used to annotate the read libraries. The occurrence of each target gene in each assembly was determined using the search function in KBase, e.g., searching for “methane” or “toluene”.
For metagenomes containing the target genes, the contigs were binned using three binning approaches: MaxBin2 (version 2.2.4) [58], CONCOCT (version 1.1) [59], and MetaBAT2 (version 1.7) [60]. The bins from the three approaches were then optimized using Das Tool (version 1.1.2) [61]. Then, CheckM (version 1.0.18) [62] was used to filter the bins for quality. The KBase app called Extract Bins as Assemblies from BinnedContigs (version 1.0.2) [36] was used to extract the bins that passed the quality filtering step. The KBase app Annotate Multiple Microbial Assemblies with RASTtk (version 1.073) [56] was used to annotate the extracted bins. Finally, the KBase app called Classify Microbes with GTDB-Tk (version 1.7.0) [63] was used to determine the taxonomy of each extracted bin. Only MAGs with completeness ≥ 80% and contamination ≤ 20% were included.

2.3. Selection of Metagenome-Assembled Genomes

The annotated bins, also called MAGs, were examined for the presence of the target genes using the KBase search function. When present, the subunit sequence representing key biomarkers for each gene (e.g., tmoA, pmoA, prmA, mmoX) were collected and stored in text files. Only MAGs containing all subunits (of the correct length and orientation) of each target gene on the same contig were selected for downloading. This involved four subunits for propane monooxygenase (prmABCD), including propane monooxygenase large subunit (prmA), reductase subunit (prmB), small subunit (prmC) and coupling protein (prmD) [64]. Soluble methane monooxygenase (sMMO) involves a six-gene operon (mmoXYBZDC), encoding for the α, β, and γ subunits of the hydroxylase (mmoXYZ), reductase (mmoC) and a regulatory or coupling protein (mmoB) [10]. MMOD is a regulatory entity to repress expression of sMMO [65]. Particulate MMO (pMMO) is encoded by the pmoCAB operon [65], which shares similarities with ammonia monooxygenase (AMO, encoded by amoCAB) resulting from ammonia-oxidizing bacteria [14,15].
Toluene-4-monooxygenase is encoded by the gene cluster tmoABCDEF [21,22,23]. The genes tmoA, tmoB, and tmoE encode the α, β, and γ subunits, respectively, that form the (αβγ)2 quaternary structure of the hydroxylase component [66]. The genes tmoC, tmoD, and tmoF encode the [2Fe-2S] ferredoxin, the effector protein, and the NADH oxidoreductase, respectively [66,67,68,69]. All MAGs containing the target genes were downloaded from individual narratives, then uploaded to publicly available KBase narratives under the following names: “Data Mining Propane Monooxygenase”, “Data Mining Soluble Methane Monooxygenase”, “Data Mining Particulate Methane/Ammonia Monooxygenase Burkholderiales Pseudomonadales and Nevskiales”, “Data Mining Particulate Methane/Ammonia Monooxygenase Methylococcales”, “Data Mining Particulate Methane/Ammonia Monooxygenase Other Phyla” and “Data Mining Toluene Monooxygenase”. The MAGs containing particulate methane/ammonia monooxygenase operons were uploaded to three separate narratives due to the large number of MAGs obtained. Links to all KBase narratives containing all MAGs are available at the following narrative “Data Mining of Groundwater to Identify MAGs with Methane, Propane and Toluene Monooxygenases” (https://narrative.kbase.us/narrative/242870, accessed on 1 July 2026).
The toluene monooxygenase containing MAGs were also examined for the presence of phenol monooxygenase, as recent studies documented the co-occurrence of toluene monooxygenase and phenol monooxygenase in Rhodoferax and Azonexus MAGs [70,71]. The gene cluster dmpKLMNOP encodes for phenol monooxygenase [35]. DmpL-DmpN-DmpO form the oxygenating component of the hydroxylase. DmpN contains the dinuclear iron center and the active site [72,73,74]. The enzyme also has a reductase component, containing one iron–sulfur cluster and FAD (DmpP), and a regulatory component (DmpM) [74,75,76,77]. DmpK is important for the iron-dependent assembly of an active oxygenase [76]. Transcription is positively regulated by the product of dmpR, with expression occurring only when the substrates are present [78,79,80].

2.4. Analysis of Metagenome-Assembled Genomes and Genes

Phylogenetic information and the source of each MAG containing the genes of interest were summarized across all samples. As stated above, when present, the subunit sequence representing key biomarkers for each gene (e.g., tmoA, pmoA, prmA, mmoX, dmpN) were collected, stored in text files, and used to generate phylogenetic trees using MEGA 12 [81]. Due to the large number of sequences collected for pmoA, only one representative sequence for each phylotype was used for each tree. The subunit size and orientations were illustrated in gene arrow plots created with R (version 4.2.1) [82] in RStudio (version 2022.12.0) [83] and the R packages readxl (version 1.4.2) [84], ggplot2 (version 3.4.4) [85], and gggenes (version 0.5.1.) [86]. Due to the large number of sequences obtained, the gene arrow diagrams were divided into groups according to their phylogenetic classification.

3. Results

3.1. Phylogenetic Classification of MAGs with the Target Operons

Tables were generated to summarize the classification of the MAGs containing the targeted operons (Supplementary Tables S1–S10). Each table contains bin and SRR numbers that correspond to the MAGs in the publicly available KBase narratives (see below). In all, thirty-four MAGs, obtained from five different datasets (different BioProject Accession Numbers), were identified as containing the full propane monooxygenase operon (prmABCD). From these, eighteen were classified with the phylum Actinomycetota (Supplementary Table S1) and sixteen with the phyla Pseudomonadota and Chloroflexota (Supplementary Table S2). The majority of those within the phylum Actinomycetota were in the order Mycobacteriales and family Mycobacteriaceae (Rhodococcus, Mycobacterium, Williamsia and Gordonia). The remaining Actinomycetota MAGs were in the orders Rubrobacterales and Solirubacterales. MAGs classifying within the Pseudomonadota belonged to the orders Rhizobiales, Acetobacterales, Rhodobacterales, Burkholderiales and Ktedonobacterales.
Twenty-six MAGs (from six datasets) contained the full operon for soluble methane monooxygenase (mmoXYBZDC) (Supplementary Table S3). The majority (twenty-four) classified within the class Gammaproteobacteria, and the remaining two classified within the class Alphaproteobacteria. Notably, the MAGs classified within only three families (Methylomonadaceae, Methylophilaceae and Beijerinckiaceae). At the genus level, sMMO containing MAGs classifying as Methylobacter, Methylomonas and Methylocella were identified multiple times.
A large number of MAGs (>100), from nine datasets, contained the full operon for ammonia/particulate methane monooxygenase (Supplementary Tables S4–S6). Forty-seven classified within the order Methylococcales (Gammaproteobacteria), with the majority of these belonging to the family Methylomonadaceae, and commonly detected genera including Methylobacter, Methylicorpusculum, Methylovulum, Methylomonas, and Methyloglobulus (Supplementary Table S4). Thirty-two classified within the orders Burkholderiales and Pseudomonadales and Nevskiales (Gammaproteobacteria) (Supplementary Table S5). The remaining MAGs containing this operon classified with class Alphaproteobacteria (e.g., Methylosinus, Methylocystis, Methylocella) and phyla Methylomirabilota, Bacteroidota, CSP1-3, Nitrospirota, Actinomycetota and Desulfobacterota_B (Supplementary Table S6).
Fifty-five MAGs (from ten datasets) contained the full operon for toluene monooxygenase (Supplementary Tables S7 and S8). From these, thirty classified within the Burkholderiaceae (Burkholderiales, Gammaproteobacteria) (Supplementary Table S7), twenty-four classified within the Rhodocyclaceae (Burkholderiales, Gammaproteobacteria) (Supplementary Table S8), and two classified within the Alphaproteobacteria (Supplementary Table S8). The majority of MAGs classified in a small number of genera (Azonexus, Hydrogenophaga, Rhodoferax and Trinickia). Interestingly, from the MAGs containing the full operon for toluene monooxygenase, thirty-three (Azonexus, Hydrogenophaga, Rhodoferax, Trinickia, Sphaerotilus, Polaromonas, Ramlibacter, Limnobacter, Zoogloea) also contained the full operon for phenol monooxygenase (Supplementary Tables S9 and S10).

3.2. Operon Characterization and Phylogenetic Trees

The size and alignment of each subunit was determined for each operon of interest obtained from the sequencing data. In general, the subunit lengths and orientation were similar to those previously reported (NCBI) and determined in our previous studies [9,71,87,88,89]. MAGs containing the four subunits of the propane monooxygenase operon (prmABCD) classified within the phyla Actinomycetota, Pseudomonadota and Chloroflexota (Figure 1 and Figure 2). For operons within MAGs classifying as phylum Actinomycetota, the average ± standard deviation (range) for the number of bases for prmA, prmB, prmC and prmD was 1643 ± 14 (1628–1664), 1053 ± 17 (1043–1097), 1120 ± 21 (1100–1169), and 350 ± 10 (341–374), respectively (Figure 1). Similarly, for operons within MAGs classifying within the phyla Pseudomonadota and Chloroflexota, the average ± standard deviation (range) for the number of bases for prmA, prmB, prmC and prmD was 1668 ± 9 (1655–1694), 1057 ± 18 (1034–1088), 1083 ± 11 (1058–1097), and 354 ± 16 (320–374), respectively (Figure 2).
MAGs containing the soluble methane monooxygenase operon (mmoXYBZDC) classified within the Gammaproteobacteria and Alphaproteobacteria (Supplementary Figures S1 and S2). For operons within MAGs classifying within the Gammaproteobacteria and Alphaproteobacteria, the average ± standard deviation (range) for the number of bases for mmoX, mmoY, mmoB, mmoZ, mmoD and mmoC was 1584 ± 6 (1577–1604), 1178 ± 1 (1175–1181), 424 ± 2 (419–425), 499 ± 4 (494–506), 264 ± 8 (242–269), and 1036 ± 2 (1031–1037), respectively (Supplementary Figure S1). Notably, eleven MAGs containing the soluble methane monooxygenase operon did not annotate the subunit D; however, the hypothetical gene was determined to be in the same position at the same length (Supplementary Figure S2). For these, the characteristics of the subunits were similar to those listed above, and the average ± standard deviation (range) for the number of bases for mmoX, mmoY, mmoB, mmoZ, hypothetical and mmoC was 1582 ± 2 (1580–1583), 1176 ± 2 (1175–1178), 417 ± 3 (413–425), 502 ± 6 (497–512), 274 ± 25 (245–311), and 1032 ± 23 (1037–1085), respectively.
Due to the large number of particulate ammonia/methane monooxygenase operons obtained, only one representative operon for each phylotype (in bold for Supplementary Tables S5–S7) was selected for determining subunit lengths and for generating the gene arrow figures (Supplementary Figures S3–S5). For operons within MAGs classifying within the order Methylococcales (Gammaproteobacteria) (Supplementary Table S4), the average ± standard deviation (range) for the number of bases for pmoC, pmoA and pmoB was 753 ± 5 (752–773), 744 ± 2 (743–749), and 1243 ± 3 (1232–1244), respectively (Supplementary Figure S3). For operons within MAGs classifying within the orders Burkholderiales and Pseudomonadales and Nevskiales (Gammaproteobacteria) (Supplementary Table S5) the average ± standard deviation (range) for the number of bases for pmoC, pmoA and pmoB was 787 ± 32 (710–851), 811 ± 22 (770–833), and 1270 ± 14 (1250–1292), respectively (Supplementary Figure S4). For operons within MAGs classifying within the class Alphaproteobacteria and phyla Methylomirabilota, Bacteroidota, CSP1-3, Nitrospirota, Actinomycetota and Desulfobacterota_B (Supplementary Table S6), the average ± standard deviation (range) for the number of bases for pmoC, pmoA and pmoB was 795 ± 37 (701–860), 779 ± 37 (725–830), and 1262 ± 24 (1214–1307), respectively (Supplementary Figure S5).
All MAGs containing the toluene monooxygenase operon classified within the order Burkholderiales (Gammaproteobacteria) (Supplementary Figures S6 and S7). For operons within MAGs classifying within the Burkholderiaceae (Burkholderiales, Gammaproteobacteria) (Supplementary Table S7), the average ± standard deviation (range) for the number of bases for tmoA, tmoB, tmoC, tmoD, tmoE and tmoF was 1499 ± 32 (1331–1505), 263 ± 3 (260–272), 355 ± 0, 315 ± 3 (308–326), 987 ± 3 (986–998), and 1021 ± 11 (1004–1037), respectively (Supplementary Figure S6). For MAGs classifying within the Rhodocyclaceae (Burkholderiales, Gammaproteobacteria) and Alphaproteobacteria (Supplementary Table S8), the average ± standard deviation (range) for the number of bases for tmoA, tmoB, tmoC, tmoD, tmoE and tmoF was 1505 ± 1 (1499–1508), 266 ± 5 (257–284), 335 ± 2 (335–344), 391 ± 62 (308–440), 986 ± 1 (986–989), and 1016 ± 10 (989–1046), respectively (Supplementary Figure S7).
For MAGS containing both operons for toluene monooxygenase and phenol monooxygenase, classifying within the Azonexus, the average ± standard deviation (range) for the number of bases for 2,3-catechol 2,3-dioxygenase, dmpK, dmpL, dpmM, dpmN, dpmO and dpmL was 926 ± 0, 275 ± 0, 989 ± 0, 269 ± 0, 1553 ± 0, 389 ± 14 (359–395), and 1061 ± 0, respectively (Supplementary Figure S8). For all other MAGs (Hydrogenophaga, Ramlibacter, Rhodoferax, Trinickia and Zoogloea) with both operons, the average ± standard deviation (range) for the number of bases for 2,3-catechol 2,3-dioxygenase, dmpK, dmpL, dpmM, dpmN, dpmO and dpmL was 937 ± 7 (926–944), 250 ± 22 (224–287), 994 ± 2 (992–998), 269 ± 0, 1544 ± 19 (1520–1562), 356 ± 1 (356–359), and 1062 ± 5 (1055–1073), respectively (Supplementary Figure S9). There were also nine MAGs that contained both operons in close proximity (Supplementary Figure S10).
Phylogenetic trees were constructed for the most common biomarker subunit genes for each operon (Figure 3, Figure 4, Figure 5, Figure 6 and Figure 7 and S11–S13). For propane monooxygenase, prmA sequences from each MAG were used to construct the tree (Figure 3) using sequences from the phyla Actinomycetota, Pseudomonadota and Chloroflexota (Supplementary Tables S1 and S2). As expected, prmA sequences from the same phyla and families clustered together. For soluble methane monooxygenase, the tree was constructed from mmoX sequences from MAGs classifying within Methylomonadaceae, Methylophilaceae, and Beijerinckiaceae (Figure 4). Again, mmoX sequences from the same families clustered together. Due to the large number of particulate ammonia/methane monooxygenase operons retrieved, three trees were constructed for pmoA (Supplementary Figures S11–S13), and each tree corresponds to data in Supplementary Tables S4, S5 and S6, respectively. For all three trees, pmoA sequences from similar phylotypes (order, families, genera) cluster together.
The MAG-derived tmoA sequenced were separated into two trees. One tree corresponds to data from MAGs classifying as Burkholderiaceae (Burkholderiales, Gammaproteobacteria) (Supplementary Table S7) (Figure 5). The other tree contains data corresponding to tmoA sequences from MAGs classifying within Rhodocyclaceae, Nevskiaceae (Gammaproteobacteria) and Alphaproteobacteria (Supplementary Table S8) (Figure 6). For both trees, tmoA sequences from similar genera cluster together, e.g., those from Hydrogenophaga, Rhodoferax, Trinickia and Azonexus. All dmpN sequences obtained from the toluene-containing MAGs are presented in one tree (Figure 7). Similar to the trends illustrated for tmoA, dmpN sequences from the same genera clustered together.

3.3. Di-Iron Center Amino Acid Sequences

The amino acid sequences of both di-iron centers were determined from the prmA, mmoX, tmoA and dmpN nucleic acid sequences for all MAGs (Table 2). All MAGs classifying in the phyla Actinomycetota and Chloroflexota illustrated similar patterns for both di-iron centers for prmA (DEVRH, DESRH). In contrast, MAGs classifying within the Pseudomonadota contained a different amino acid in the first prmA di-iron center (DEFRH). For mmoX, all MAGs contained the same two di-iron center amino acids (DEIRH, DELRH). More variability was observed for the first di-iron center of tmoA, with four different trends being noted. All second di-iron centers for tmoA were the same (DESRH). For dmpN, two patterns were present for both di-iron centers (DEIRH, DELRH for the first and DEARH and DESRH for the second).

3.4. Publicly Available KBase Narratives

As stated in the Methods Section, the MAGs containing the target genes were downloaded from individual narratives, then uploaded to six publicly available KBase narratives. Links to the six KBase narratives can be easily accessed at the following DOI site https://doi.org/10.25982/242870.6/3013398. These narratives are particularly useful, as they provide a user-friendly approach for obtaining nucleotide and amino acid sequence data for the targeted operons and surrounding genes, as well as the size of the genome. Each MAG, with identifying information in Tables S1–S10, is opened in the main window. To retrieve information from any MAG, simply select the browse features tab, then search for the appropriate gene name (“propane”, “methane”, “toluene”, “phenol”, depending on the narrative). This will produce information concerning the start position, strand (+ or −), length and the appropriate contig for each subunit. Following this, select a feature ID (e.g., gene_4358) to get access to the nucleotide and amino acid sequence of each subunit for each operon. Also, in the feature context box generated, users can obtain annotations and sequence data for neighboring genes.

4. Discussion

This research involved the analysis of a large amount of whole genome sequencing data (>600 individual samples, >12,000 Gbases, >80 KBase narratives) from groundwater samples to generate groups of MAGs containing the target operons. Following this, the MAGs were uploaded into publicly available KBase narratives to allow others to obtain nucleic and amino acid sequences for all subunits for each operon, as well as the surrounding genes. The number of MAGs identified differed considerably for each target operon: 34 contained the full propane monooxygenase operon; 26 contained the full operon for soluble methane monooxygenase; 55 contained the full operon for toluene monooxygenase; 33 contained both toluene and phenol monooxygenase, and more than 100 contained the full operon for particulate ammonia/particulate methane monooxygenase. The low number of MAGs obtained with the operons of interest is perhaps surprising, given the large amount of data examined. MAGs with toluene monooxygenase and particulate ammonia/methane monooxygenase were detected across 10 and 9 datasets, respectively. In contrast, MAGs with propane monooxygenase and soluble methane monooxygenase were only detected across five and six datasets, respectively. This pattern perhaps indicates the more common occurrence of the former target operons compared to the latter in groundwater. However, additional research is needed to confirm this.
The identification of propane monooxygenase containing MAGs in the phylum Actinomycetota is consistent with the results of previous research. Specifically, reported propanotrophs in this phylum include those in the genera Mycobacterium, Rhodococcus, Gordonia, Nocardioides and Pseudonocardia [2,3,4,90,91,92,93,94,95,96,97]. The current study also identified propane monooxygenase-containing MAGs in the phylum Pseudomonadota. Although, based on sequence data [98,99], phylotypes in this phylum have been reported to contain the propane monooxygenase, little is known about their ability to degrade propane. Exceptions to this include Methylocella sp. PC1 and PC4, Methylocella silvestris BL2, and Methylocella silvestris TVC (Rhizobiales, Beijerinckiaceae), as these can use propane to support growth and contain the propane monooxygenase operon [100,101,102,103]. Also, Azoarcus DD4 (Rhodocyclales, Zoogloeaceae) contains a putative propane monooxygenase and can use propane as a carbon source [104]. Consistent with previous pure and mixed culture work, in the current study for the phylum Pseudomonadota, propane monooxygenase-containing MAGs were identified in groundwater from the genera Methylocella, Azoarcus and Methylibium. MAGs in both phyla Actinomycetota (Mycobacterium, Rhodococcus opacus, Rhodococcus wratislaviensis, Pseudonocardia) and Pseudomonadota (Methylibium) were also recently identified with the full propane monooxygenase operon in propanotrophic enrichment cultures [9,105].
Methanotrophs classify within two phyla, Proteobacteria (classes Alphaproteobacteria and Gammaproteobacteria) and Verrucomicrobia [106]. The genes for sMMO have been well studied in Methylosinus trichosporium OB3b [107,108] and Methylococcus capsulatus Bath [109,110,111]. In the current study, all MAGs containing sMMO classified within the Proteobacteria. Specifically, the majority of sMMO operons were detected in MAGs in the family Methylomonadaceae (Methylococcales, Gammaproteobacteria), including, for example, those in the genera Methylobacter, Methylovulum and Methylomonas. Consistent with this, sMMO genes have previously been reported for species in the same genera: Methylobacter methanoversatilis, Methylobacter tundripaludum, Methylobacter psychrophilus [112], Methylovulum miyakonense HT12 [113], Methylomonas sp. strain 11b, Methylomonas sp. strain LW13 [114], Methylomonas methanica strain 68-1 [115], and Methylomonas methanica MC09 [116]. Two Methylocella (Beijerinckiaceae, Rhizobiales, Alphaproteobacteria) MAGs containing sMMO operons were also identified in the current study. This is also consistent with previous reports of this genus containing sMMO, for example, for three strains of Methylocella palustris [117], Methylocella silvestris BL2T [118], and three strains of Methylocella tundrae [119]. The identification of sMMO-containing MAGs in the family Methylophilaceae is more puzzling, as this family typically utilizes methanol or methylamine as a sole source of carbon and energy [120,121]. As the MAG was only classified to the family level, it is possible that a mis-identification occurred. It is surprising that no MAGs containing sMMO were detected in the genera Methylosinus, Methylococcus Methylocystis, as these phylotypes have previously been detected in groundwater aquifers [122,123].
The majority of MAGs containing the full operon for ammonia/particulate methane monooxygenase classified within the Gammaproteobacteria (Methylococcales, Burkholderiales, Pseudomonadales and Nevskiales), with only six within the Alphaproteobacteria (primarily Rhizobiales). The remaining MAGs with this operon classified with the phyla Methylomirabilota, Bacteroidota, CSP1-3, Nitrospirota, Actinomycetota and Desulfobacterota_B. Others have also documented the presence of Alphaproteobacteria and Gammaproteobacteria containing this operon in groundwater. For example, Methylobacter spp. were active in groundwater microcosms capable of the biodegradation of chlorinated volatile organic compounds [124]. In another study focused on the stimulation of methanotrophs in multiple acidic aquifers, various genera (Methylomonas, Methylobacter, Methylosinus, and Methylococcus) were detected [122]. Molecular analysis of uncontaminated groundwater wells and methanotrophic enrichment cultures derived from the water also revealed sequences aligning with the genera Methylobacter, Methylosinus, Methylomonas, Methylocaldum, and Methylocystis [123]. Little is known about the occurrence of methanotrophs with this operon in the additional phyla listed above.
MAGs containing the full operon for toluene monooxygenase classified within the Burkholderiaceae (Burkholderiales, Gammaproteobacteria), Rhodocyclaceae (Burkholderiales, Gammaproteobacteria) and the Alphaproteobacteria families. The majority classified in a small number of genera, including Azonexus, Hydrogenophaga, Rhodoferax and Trinickia. The detection of the genes encoding for toluene monooxygenase has significant implications, as this enzymes has been linked to the biodegradation of numerous pollutants, such as alkenes, 1,4-dioxane, N-nitrosodimethylamine, chloroform, TCE, various polycyclic aromatic substrates, and substituted benzenes [24,25,26,27,28,29,30,31]. For example, tmoABCDEF has been associated with the co-metabolism of 1,1-DCE, cis-DCE, and 1,4-dioxane in Azoarcus sp. DD4 [22,23]. The same gene cluster was found within a Rhodoferax MAG in enrichment cultures capable of xylene degradation, established from groundwater from a xylene-contaminated site [70]. Additional research will be needed to more conclusively link contaminant biodegradation to each toluene monooxygenase-containing MAG.

5. Conclusions

The current study provides novel gene sequences and MAGs from groundwater metagenomes for a group of important biomarkers for the aerobic biodegradation of pollutants. Of course, it is important to state that the presence of genetic material does not necessarily indicate biodegradation activity. Such results have important implications because these functional genes have been associated with the removal of numerous groundwater contaminants. The work highlights the advantages of KBase for the analysis of large quantities of whole genome sequencing data. Further, KBase narratives are publicly available, providing an excellent platform for sharing data. For example, qPCR assays could be updated based on the collected sequences. From the biomarkers investigated, MAGs containing particulate methane/ammonia monooxygenase were most commonly detected (>100). Multiple (>20 for each) MAGs containing the full operons for propane monooxygenase, soluble methane monooxygenase, toluene monooxygenase and phenol monooxygenase were also detected. The MAGs generated and their associated functional gene sequences have the potential to improve molecular detection methods for site bioremediation.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/biology15181586/s1, Supplementary Figure S1. Subunit length and arrangement of soluble monooxygenase operons in groundwater metagenomes associated with the identified metagenome assembled genomes within the Gammaproteobacteria and Alphaproteobacteria, Supplementary Figure S2. Subunit length and arrangement of soluble monooxygenase operons in groundwater metagenomes associated with the identified metagenome assembled genomes within the Gammaproteobacteria and Alphaproteobacteria. The mmoD subunit was not annotated in these cases. Supplementary Figure S3. Subunit length and arrangement of ammonia/methane monooxygenase operons in groundwater metagenomes associated with the identified metagenome assembled genomes within the Methylococcales (Gammaproteobacteria). Supplementary Figure S4. Subunit length and arrangement of ammonia/methane monooxygenase operons in groundwater metagenomes associated with the identified metagenome assembled genomes within the Burkholderiales and Pseudomonadales and Nevskiales (Gammaproteobacteria). Supplementary Figure S5. Subunit length and arrangement of ammonia/methane monooxygenase operons in groundwater metagenomes associated with the identified metagenome assembled genomes within the the class Alphaproteobacteria, and phyla Methylomirabilota, Bacteroidota, CSP1-3, Nitrospirota, Actinomycetota and Desulfobacterota_B. Supplementary Figure S6. Subunit length and arrangement of toluene monooxygenase operons in groundwater metagenomes associated with the identified metagenome assembled genomes within the Burkholderiaceae (Burkholderiales, Gammaproteobacteria). Supplementary Figure S7. Subunit length and arrangement of toluene monooxygenase operons in groundwater metagenomes associated with the identified metagenome assembled genomes within the Rhodocyclaceae, Nevskiaceae (Gammaproteobacteria) and Alphaproteobacteria. Supplementary Figure S8. Subunit length and arrangement of phenol monooxygenase operons in groundwater metagenomes associated with Azonexus metagenome assembled genomes also containing toluene monooxygenase operons. Supplementary Figure S9. Subunit length and arrangement of phenol monooxygenase operons in groundwater metagenomes associated with Hydrogenophaga, Ramlibacter, Rhodoferax, Trinickia and Zoogloea metagenome assembled genomes also containing toluene monooxygenase operons. Supplementary Figure S10. Subunit length and arrangement of phenol monooxygenase operons in close proximity to a toluene monooxygenase operon. Supplementary Figure S11. Alignment of pmoA from MAGs containing the particulate ammonia/methane monooxygenase operon within the Methylococcales (Gammaproteobacteria). Supplementary Figure S12. Alignment of pmoA from MAGs containing the particulate ammonia/methane monooxygenase operon within the Burkholderiales and Pseudomonadales and Nevskiales (Gammaproteobacteria). Supplementary Figure S13. Alignment of pmoA from MAGs containing the particulate ammonia/methane monooxygenase operon within the class Alphaproteobacteria, and phyla Methylomirabilota, Bacteroidota, CSP1-3, Nitrospirota, Actinomycetota and Desulfobacterota_B. Supplementary Table S1. MAGs containing four subunits of propane monooxygenase classifying within the phylum Actinomycetota. Supplementary Table S2. MAGs containing four subunits of propane monooxygenase classifying within the phyla Pseudomonadota and Chloroflexota. Supplementary Table S3. MAGs containing soluble methane monooxygenase operons classifying within the Gammaproteobacteria and Alphaproteobacteria. Supplementary Table S4. MAGs containing particulate ammonia/particulate methane monooxygenase operons classifying within the Methylococcales (Gammaproteobacteria). Supplementary Table S5. MAGs containing particulate ammonia/particulate methane monooxygenase operons classifying within the Burkholderiales and Pseudomonadales and Nevskiales (Gammaproteobacteria). Supplementary Table S6. MAGs containing particulate ammonia/particulate methane monooxygenase operons classifying within the class Alphaproteobacteria, and phyla Methylomirabilota, Bacteroidota, CSP1-3, Nitrospirota, Actinomycetota and Desulfobacterota_B. Supplementary Table S7. MAGs containing toluene monooxygenase operons classifying within the family Burkholderiaceae (Burkholderiales, Gammaproteobacteria). Supplementary Table S8. MAGs containing toluene monooxygenase operons classifying within the Rhodocyclaceae, Nevskiaceae (Gammaproteobacteria) and Alphaproteobacteria. Supplementary Table S9. MAGs containing toluene monooxygenase operons classifying within the family Burkholderiaceae (Burkholderiales, Gammaproteobacteria). Supplementary Table S10. MAGs containing both toluene monooxygenase and phenol monooxygenase operons within the Rhodocyclaceae.

Author Contributions

Conceptualization, A.M.C.; methodology, A.M.C., J.R. and M.B.D.C.; software, A.M.C., J.R. and M.B.D.C.; validation, A.M.C., J.R. and M.B.D.C.; formal analysis, A.M.C., J.R. and M.B.D.C.; investigation, A.M.C., J.R. and M.B.D.C.; resources, A.M.C.; data curation, A.M.C., J.R. and M.B.D.C.; writing—original draft preparation, A.M.C.; writing—review and editing, A.M.C., J.R. and M.B.D.C.; visualization, A.M.C.; supervision, A.M.C.; project administration, A.M.C.; funding acquisition, A.M.C. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by NSF (2413523) and the College of Engineering (MSU) through the ENSURE program.

Institutional Review Board Statement

The study did not involve animal or human subjects.

Data Availability Statement

All sequencing data was obtained from NCBI. Links to all KBase narratives containing all MAGs are available using the following narrative: “Data Mining of Groundwater to Identify MAGs with Methane, Propane and Toluene Monooxygenases” (https://narrative.kbase.us/narrative/242870, accessed on 1 July 2025).

Acknowledgments

The authors thank the researchers that prepared the sequencing datasets analyzed in this study.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Steffan, R.J.; McClay, K.; Vainberg, S.; Condee, C.W.; Zhang, D. Biodegradation of the gasoline oxygenates methyl tert-butyl ether, ethyl tert-butyl ether, and tert-amyl methyl ether by propane-oxidizing bacteria. Appl. Environ. Microbiol. 1997, 63, 4216–4222. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Sharp, J.O.; Sales, C.M.; Alvarez-Cohen, L. Functional characterization of propane-enhanced N-nitrosodimethylamine degradation by two actinomycetales. Biotechnol. Bioeng. 2010, 107, 924–932. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Cappelletti, M.; Pinelli, D.; Fedi, S.; Zannoni, D.; Frascari, D. Aerobic co-metabolism of 1,1,2,2-tetrachloroethane by Rhodococcus aetherivorans TPA grown on propane: Kinetic study and bioreactor configuration analysis. J. Chem. Technol. Biotechnol. 2018, 93, 155–165. [Google Scholar] [CrossRef] [Scilit]
  4. Wang, B.; Chu, K.H. Cometabolic biodegradation of 1,2,3-trichloropropane by propane-oxidizing bacteria. Chemosphere 2017, 168, 1494–1497. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Rolston, H.; Hyman, M.; Semprini, L. Single-well push-pull tests evaluating isobutane as a primary substrate for promoting in situ cometabolic biotransformation reactions. Biodegradation 2022, 33, 349–371. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Danko, A.; Lippincott, D.; Hatzinger, P.; Lavorgna, G.; Semprini, L.; Hyman, M.R. Evaluation of a Novel Multiple Primary Substrate (MPS) Cometabolic Approach for In Situ Bioremediation of 1,4-Dioxane and Chlorinated Solvents in Groundwater; ESTCP Project ER-201733; ESTCP: Alexandria, VA, USA, 2023. [Google Scholar]
  7. Hatzinger, P.B.; Lippincott, D.R. Field demonstration of N-Nitrosodimethylamine (NDMA) treatment in groundwater using propane biosparging. Water Res. 2019, 164, 114923. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Stohr, H.; Vaidya, R.; Wilson, C.; Pruden, A.; Salazar-Benites, G.; Bott, C. Cometabolic treatment of 1,4-dioxane in biologically active carbon filtration with tetrahydrofuran and propane at relevant concentrations for potable reuse. ACS EST Water 2023, 9, 2948–2954. [Google Scholar] [CrossRef] [Scilit]
  9. Eshghdoostkhatami, Z.; Li, Z.; Faghihinezhad, M.; Cupples, A.M. Characterization of propanotrophic enrichments from agricultural soils capable of 1,4-dioxane biodegradation to sub-µg/L levels. Sci. Total Environ. 2025, 1005, 180824. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Jiang, H.; Chen, Y.; Murrell, J.C.; Jiang, P.; Zhang, C.; Xing, X.-H.; Smith, T.J. Methanotrophs: Multifunctional Bacteria with promising applicationsin environmental bioengineering. Biochem. Eng. J. 2010, 49, 277–288. [Google Scholar] [CrossRef] [Scilit]
  11. Theisen, A.R.; Ali, M.H.; Radajewski, S.; Dumont, M.G.; Dunfield, P.F.; McDonald, I.R.; Dedysh, S.N.; Miguez, C.B.; Murrell, J.C. Regulation of methane oxidation in the facultative methanotroph Methylocella silvestris BL2. Mol. Microbiol. 2005, 58, 682–692. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Vorobev, A.V.; Baani, M.; Doronina, N.V.; Brady, A.L.; Liesack, W.; Dunfield, P.F.; Dedysh, S.N. Methyloferula stellata gen. nov., sp. nov., an acidophilic, obligately methan otrophic bacterium that possesses only a soluble methane monooxygenase. Int. J. Syst. Evol. Microbiol. 2011, 61, 2456–2463. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Murrell, J.C.; McDonald, I.R.; Gilbert, B. Regulation of expression of methane monooxygenases by copper ions. Trends Microbiol. 2000, 8, 221–225. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Holme, A.J.; Costello, A.; Lidstrom, M.E.; Murrell, J.C. Evidence that particulate methane monooxygenase and ammonia monooxygenase may be evolutionarily related. FEMS Microbiol. Lett. 1995, 132, 203–208. [Google Scholar] [CrossRef]
  15. Wendeborn, S. The chemistry, biology, and modulation of ammonium nitrification in soil. Angew. Chem. Int. Ed. Engl. 2020, 59, 2182–2202. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Chu, M.Y.J.; Bennett, P.J.; Dolan, M.E.; Hyman, M.R.; Peacock, A.D.; Bodour, A.; Anderson, R.H.; Mackay, D.M.; Goltz, M.N. Concurrent treatment of 1,4-dioxane and chlorinated aliphatics in a groundwater recirculation system via aerobic cometabolism. Groundw. Monit. Remediat. 2018, 38, 53–64. [Google Scholar] [CrossRef] [Scilit]
  17. Guo, G.L.; Tseng, D.H.; Huang, S.L. Co-metabolic degradation of trichloroethylene by Pseudomonas putida in a fibrous bed bioreactor. Biotechnol. Lett. 2001, 23, 1653–1657. [Google Scholar] [CrossRef] [Scilit]
  18. Sun, A.K.; Wood, T.K. Trichloroethylene mineralization in a fixed-film bioreactor using a pure culture expressing constitutively toluene ortho-monooxygenase. Biotechnol. Bioeng. 1997, 55, 674–685. [Google Scholar] [CrossRef] [Scilit]
  19. Fries, M.R.; Forney, L.J.; Tiedje, J.M. Phenol-and toluene degrading microbial populations from an aquifer in which successful trichloroethene cometabolism occurred. Appl. Environ. Microbiol. 1997, 63, 1523–1530. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Chang, H.L.; Alvarez-Cohen, L. Transformation capacities of chlorinated organics by mixed cultures enriched on methane, propane, toluene, or phenol. Biotechnol. Bioeng. 1995, 45, 440–449. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Yen, K.M.; Karl, M.R.; Blatt, L.M.; Simon, M.J.; Winter, R.B.; Fausset, P.R.; Lu, H.S.; Harcourt, A.A.; Chen, K.K. Cloning and characterization of a Pseudomonas mendocina KR1 gene cluster encoding toluene-4-monooxygenase. J. Bacteriol. 1991, 173, 5315–5327. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Deng, D.; Pham, D.N.; Li, F.; Li, M. Discovery of an inducible toluene monooxygenase that co-oxidizes 1, 4-dioxane and 1, 1-dichloroethylene in propanotrophic Azoarcus sp. DD4. Appl. Environ. Microbiol. 2020, 86, e01163-20. [Google Scholar] [PubMed]
  23. Li, F.; Deng, D.; Zeng, L.; Abrams, S.; Li, M. Sequential anaerobic and aerobic bioaugmentation for commingled groundwater contamination of trichloroethene and 1,4-dioxane. Sci. Total Environ. 2021, 774, 145118. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Mahendra, S.; Alvarez-Cohen, L. Kinetics of 1,4-dioxane biodegradation by monooxygenase-expressing bacteria. Environ. Sci. Technol. 2006, 40, 5435–5442. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Parales, R.E.; Parales, J.V.; Pelletier, D.A.; Ditty, J.L. Diversity of microbial toluene degradation pathways. Adv. Appl. Microbiol. 2008, 64, 1–73. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Sharp, J.O.; Wood, T.K.; Alvarez-Cohen, L. Aerobic biodegradation of N-nitrosodimethylamine (NDMA) by axenic bacterial strains. Biotechnol. Bioeng. 2005, 89, 608–618. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. McClay, K.; Fox, B.G.; Steffan, R.J. Chloroform mineralization by toluene-oxidizing bacteria. Appl. Environ. Microbiol. 1996, 62, 2716–2722. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. McClay, K.; Fox, B.G.; Steffan, R.J. Toluene monooxygenase-catalyzed epoxidation of alkenes. Appl. Environ. Microbiol. 2000, 66, 1877–1882. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Oppenheim, S.F.; Studts, J.M.; Fox, B.G.; Dordick, J.S. Aromatic hydroxylation catalyzed by toluene 4-monooxygenase in organic solvent/aqueous buffer mixtures. Appl. Biochem. Biotechnol. 2001, 90, 187–197. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Pikus, J.D.; Studts, J.M.; McClay, K.; Steffan, R.J.; Fox, B.G. Changes in the regiospecificity of aromatic hydroxylation produced by active site engineering in the diiron enzyme toluene 4-monooxygenase. Biochemistry 1997, 36, 9283–9289. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Winter, R.B.; Yen, K.-M.; Ensley, B.D. Efficient degradation of trichloroethylene by a recombinant Escherichia coli. Nat. Biotechnol. 1989, 7, 282–285. [Google Scholar] [CrossRef] [Scilit]
  32. Reineke, W.; Knackmuss, H.J. Microbial degradation of haloaromatics. Annu. Rev. Microbiol. 1988, 42, 263–287. [Google Scholar] [CrossRef] [PubMed]
  33. Dagley, S. Biochemistry of aromatic hydrocarbon degradation in pseudomonads. In The Biology of Pseudomonas; Sokatch, J.R., Ed.; Academic Press, Inc.: London, UK, 1986; Volume 10, pp. 527–556. [Google Scholar]
  34. Shingler, V.; Powlowski, J.; Marklund, U. Nucleotide sequence and functional analysis of the complete phenol/3,4-dimethylphenol catabolic pathway of Pseudomonas sp. strain CF600. J. Bacteriol. 1992, 174, 711–724. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Powlowski, J.; Shingler, V. Genetics and biochemistry of phenol degradation by Pseudomonas sp. CF 600. Biodegradation 1994, 5, 219–236. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Arkin, A.P.; Cottingham, R.W.; Henry, C.S.; Harris, N.L.; Stevens, R.L.; Maslov, S.; Dehal, P.; Ware, D.; Perez, F.; Canon, S.; et al. KBase: The United States Department of Energy Systems Biology Knowledgebase. Nat. Biotechnol. 2018, 36, 566–569. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Anantharaman, K.; Brown, C.T.; Burstein, D.; Castelle, C.J.; Probst, A.J.; Thomas, B.C.; Williams, K.H.; Banfield, J.F. Analysis of five complete genome sequences for members of the class Peribacteria in the recently recognized Peregrinibacteria bacterial phylum. PeerJ 2016, 4, e1607. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Anantharaman, K.; Brown, C.T.; Hug, L.A.; Sharon, I.; Castelle, C.J.; Probst, A.J.; Thomas, B.C.; Singh, A.; Wilkins, M.J.; Karaoz, U.; et al. Thousands of microbial genomes shed light on interconnected biogeochemical processes in an aquifer system. Nat. Commun. 2016, 7, 13219. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Zaremba-Niedzwiedzka, K.; Caceres, E.F.; Saw, J.H.; Backstrom, D.; Juzokaite, L.; Vancaester, E.; Seitz, K.W.; Anantharaman, K.; Starnawski, P.; Kjeldsen, K.U.; et al. Asgard archaea illuminate the origin of eukaryotic cellular complexity. Nature 2017, 541, 353–358. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  40. Danczak, R.E.; Johnston, M.D.; Kenah, C.; Slattery, M.; Wilkins, M.J. Capability for arsenic mobilization in groundwater is distributed across broad phylogenetic lineages. PLoS ONE 2019, 14, e0221694. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Mosley, O.E.; Gios, E.; Weaver, L.; Close, M.; Daughney, C.; van der Raaij, R.; Martindale, H.; Handley, K.M. Metabolic Diversity and Aero-Tolerance in Anammox Bacteria from Geochemically Distinct Aquifers. mSystems 2022, 7, e0125521. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Tian, R.; Ning, D.; He, Z.; Zhang, P.; Spencer, S.J.; Gao, S.; Shi, W.; Wu, L.; Zhang, Y.; Yang, Y.; et al. Small and mighty: Adaptation of superphylum Patescibacteria to groundwater environment drives their genome simplicity. Microbiome 2020, 8, 51. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Lui, L.M.; Nielsen, T.N.; Smith, H.J.; Chandonia, J.M.; Kuehl, J.V.; Song, F.; Sczesnak, A.; Hendrickson, A.; Hazen, T.C.; Fields, M.W.; et al. Sediment and groundwater metagenomes from subsurface microbial communities from the Oak Ridge National Laboratory Oak Ridge Reservation, Oak Ridge, Tennessee, USA. Microbiol. Resour. Announc. 2025, 14, e0001425. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Goff, J.L.; Lui, L.M.; Nielsen, T.N.; Poole, F.L.; Smith, H.J.; Walker, K.F.; Hazen, T.C.; Fields, M.W.; Arkin, A.P.; Adams, M.W.W. Mixed waste contamination selects for a mobile genetic element population enriched in multiple heavy metal resistance genes. ISME Commun. 2024, 4, ycae064. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Hernsdorf, A.W.; Amano, Y.; Miyakawa, K.; Ise, K.; Suzuki, Y.; Anantharaman, K.; Probst, A.; Burstein, D.; Thomas, B.C.; Banfield, J.F. Potential for microbial H2 and metal transformations associated with novel bacteria and archaea in deep terrestrial subsurface sediments. ISME J. 2017, 11, 1915–1929. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Ruff, S.E.; Humez, P.; de Angelis, I.H.; Diao, M.; Nightingale, M.; Cho, S.; Connors, L.; Kuloyo, O.O.; Seltzer, A.; Bowman, S.; et al. Hydrogen and dark oxygen drive microbial productivity in diverse groundwater ecosystems. Nat. Commun. 2023, 14, 3194. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Dang, H.; Cupples, A.M. Diversity and abundance of the functional genes and bacteria associated with RDX degradation at a contaminated site pre- and post-biostimulation. Appl. Microbiol. Biotechnol. 2021, 105, 6463–6475. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Wu, X.; Holmfeldt, K.; Hubalek, V.; Lundin, D.; Astrom, M.; Bertilsson, S.; Dopson, M. Microbial metagenomes from three aquifers in the Fennoscandian shield terrestrial deep biosphere reveal metabolic partitioning among populations. ISME J. 2016, 10, 1192–1203. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  49. Wu, Z.; Liu, T.; Chen, Q.; Chen, T.; Hu, J.; Sun, L.; Wang, B.; Li, W.; Ni, J. Unveiling the unknown viral world in groundwater. Nat. Commun. 2024, 15, 6788. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. Diba, F.; Hoque, M.N.; Rahman, M.S.; Haque, F.; Rahman, K.M.J.; Moniruzzaman, M.; Khan, M.; Hossain, M.A.; Sultana, M. Metagenomic and culture-dependent approaches unveil active microbial community and novel functional genes involved in arsenic mobilization and detoxification in groundwater. BMC Microbiol. 2023, 23, 241. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  51. Chaudhari, N.M.; Overholt, W.A.; Figueroa-Gonzalez, P.A.; Taubert, M.; Bornemann, T.L.V.; Probst, A.J.; Holzer, M.; Marz, M.; Kusel, K. The economical lifestyle of CPR bacteria in groundwater allows little preference for environmental drivers. Environ. Microbiome 2021, 16, 24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Hauptfeld, E.; Pappas, N.; van Iwaarden, S.; Snoek, B.L.; Aldas-Vargas, A.; Dutilh, B.E.; von Meijenfeldt, F.A.B. Integrating taxonomic signals from MAGs and contigs improves read annotation and taxonomic profiling of metagenomes. Nat. Commun. 2024, 15, 3373. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Andrews, S. FastQC: A Quality Control Tool for High Throughput Sequence Data. 2010. Available online: https://www.bioinformatics.babraham.ac.uk/projects/fastqc/ (accessed on 1 July 2025).
  54. Bolger, A.M.; Lohse, M.; Usadel, B. Trimmomatic: A flexible trimmer for Illumina sequence data. Bioinformatics 2014, 30, 2114–2120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Li, D.; Liu, C.M.; Luo, R.; Sadakane, K.; Lam, T.W. MEGAHIT: An ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics 2015, 31, 1674–1676. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  56. Aziz, R.K.; Bartels, D.; Best, A.A.; DeJongh, M.; Disz, T.; Edwards, R.A.; Formsma, K.; Gerdes, S.; Glass, E.M.; Kubal, M.; et al. The RAST Server: Rapid Annotations using Subsystems Technology. BMC Genom. 2008, 9, 75. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Seemann, T. Prokka: Rapid prokaryotic genome annotation. Bioinformatics 2014, 30, 2068–2069. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  58. Wu, Y.W.; Simmons, B.A.; Singer, S.W. MaxBin 2.0: An automated binning algorithm to recover genomes from multiple metagenomic datasets. Bioinformatics 2016, 32, 605–607. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  59. Alneberg, J.; Bjarnason, B.S.; de Bruijn, I.; Schirmer, M.; Quick, J.; Ijaz, U.Z.; Lahti, L.; Loman, N.J.; Andersson, A.F.; Quince, C. Binning metagenomic contigs by coverage and composition. Nat. Methods 2014, 11, 1144–1146. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Kang, D.D.; Froula, J.; Egan, R.; Wang, Z. MetaBAT, an efficient tool for accurately reconstructing single genomes from complex microbial communities. PeerJ 2015, 3, e1165. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. Sieber, C.M.K.; Probst, A.J.; Sharrar, A.; Thomas, B.C.; Hess, M.; Tringe, S.G.; Banfield, J.F. Recovery of genomes from metagenomes via a dereplication, aggregation and scoring strategy. Nat. Microbiol. 2018, 3, 836–843. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. Parks, D.H.; Imelfort, M.; Skennerton, C.T.; Hugenholtz, P.; Tyson, G.W. CheckM: Assessing the quality of microbial genomes recovered from isolates, single cells, and metagenomes. Genome Res. 2015, 25, 1043–1055. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  63. Chaumeil, P.A.; Mussig, A.J.; Hugenholtz, P.; Parks, D.H. GTDB-Tk: A toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics 2019, 36, 1925–1927. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  64. Sharp, J.O.; Sales, C.M.; LeBlanc, J.C.; Liu, J.; Wood, T.K.; Eltis, L.D.; Mohn, W.W.; Alvarez-Cohen, L. An inducible propane monooxygenase is responsible for N-nitrosodimethylamine degradation by Rhodococcus sp. strain RHA1. Appl. Environ. Microbiol. 2007, 73, 6930–6938. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  65. Koo, C.W.; Rosenzweig, A.C. Biochemistry of aerobic biological methane oxidation. Chem. Soc. Rev. 2021, 50, 3424–3436. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  66. Brouk, M.; Nov, Y.; Fishman, A. Improving biocatalyst performance by integrating statistical methods into protein engineering. Appl. Environ. Microbiol. 2010, 76, 6397–6403. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  67. Bailey, L.J.; McCoy, J.G.; Phillips, G.N., Jr.; Fox, B.G. Structural consequences of effector protein complex formation in a diiron hydroxylase. Proc. Natl. Acad. Sci. USA 2008, 105, 19194–19198. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  68. Leahy, J.G.; Batchelor, P.J.; Morcomb, S.M. Evolution of the soluble diiron monooxygenases. FEMS Microbiol. Rev. 2003, 27, 449–479. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. Mitchell, K.H.; Studts, J.M.; Fox, B.G. Combined participation of hydroxylase active site residues and effector protein binding in a para to ortho modulation of toluene 4-monooxygenase regiospecificity. Biochemistry 2002, 41, 3176–3188. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  70. Tancsics, A.; Banerjee, S.; Soares, A.; Bedics, A.; Kriszt, B. Combined omics approach reveals key differences between aerobic and microaerobic xylene-degrading enrichment bacterial communities: Rhodoferax—A hitherto unknown player emerges from the microbial dark matter. Environ. Sci. Technol. 2023, 57, 2846–2855. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Cupples, A.M.; Dang, H.; Foss, K.; Bernstein, A.; Thelusmond, J.R. An investigation of soil and groundwater metagenomes for genes encoding soluble and particulate methane monooxygenase, toluene-4-monoxygenase, propane monooxygenase and phenol hydroxylase. Arch. Microbiol. 2024, 206, 363. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  72. Qian, H.; Edlund, U.; Powlowski, J.; Shingler, V.; Sethson, I. Solution structure of phenol hydroxylase protein component P2 determined by NMR spectroscopy. Biochemistry 1997, 36, 495–504. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  73. Enroth, C.; Neujahr, H.; Schneider, G.; Lindqvist, Y. The crystal structure of phenol hydroxylase in complex with FAD and phenol provides evidence for a concerted conformational change in the enzyme and its cofactor during catalysis. Structure 1998, 6, 605–617. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. El-Sayed, W.S.; Ibrahim, M.K.; Ouf, S.A. Molecular characterization of the alpha subunit of multicomponent phenol hydroxylase from 4-chlorophenol-degrading Pseudomonas sp. strain PT3. J. Microbiol. 2014, 52, 13–19. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  75. Nordlund, I.; Powlowski, J.; Shingler, V. Complete nucleotide sequence and polypeptide analysis of multicomponent phenol hydroxylase from Pseudomonas sp. strain CF600. J. Bacteriol. 1990, 172, 6826–6833. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  76. Powlowski, J.; Shingler, V. In vitro analysis of polypeptide requirements of multicomponent phenol hydroxylase from Pseudomonas sp. strain CF600. J. Bacteriol. 1990, 172, 6834–6840. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  77. Powlowski, J.; Sealy, J.; Shingler, V.; Cadieux, E. On the role of DmpK, an auxiliary protein associated with multicomponent phenol hydroxylase from Pseudomonas sp. strain CF600. J. Biol. Chem. 1997, 272, 945–951. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  78. Shingler, V.; Moore, T. Sensing of aromatic compounds by the DmpR transcriptional activator of phenol-catabolizing Pseudomonas sp. strain CF600. J. Bacteriol. 1994, 176, 1555–1560. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  79. Shingler, V.; Franklin, F.C.; Tsuda, M.; Holroyd, D.; Bagdasarian, M. Molecular analysis of a plasmid-encoded phenol hydroxylase from Pseudomonas CF600. Microbiology 1989, 135, 1083–1092. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  80. Shingler, V.; Bartilson, M.; Moore, T. Cloning and nucleotide sequence of the gene encoding the positive regulator (DmpR) of the phenol catabolic pathway encoded by pVI150 and identification of DmpR as a member of the NtrC family of transcriptional activators. J. Bacteriol. 1993, 175, 1596–1604. [Google Scholar] [CrossRef] [Scilit] [PubMed][Green Version]
  81. Kumar, S.; Stecher, G.; Suleski, M.; Sanderford, M.; Sharma, S.; Tamura, K. MEGA12: Molecular Evolutionary Genetic Analysis Version 12 for Adaptive and Green Computing. Mol. Biol. Evol. 2024, 41, msae263. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  82. R Core Team. R: A Language and Environment for Statistical Computing; R Foundation for Statistical Computing: Vienna, Austria, 2018. [Google Scholar]
  83. RStudio_Team. RStudio: Integrated Development for R. Rstudio; PBC: Boston, MA, USA, 2020. [Google Scholar]
  84. Wickham, H.; Bryan, J. Readxl: Read Excel Files_. R Package Version 1.4.2. 2023. Available online: https://CRAN.R-project.org/package=readxl (accessed on 1 July 2025).
  85. Wickham, H. Ggplot2: Elegant Graphics for Data Analysis; Springer: New York, NY, USA, 2016; Available online: https://ggplot2.tidyverse.org (accessed on 1 July 2025).
  86. Wilkins, D. Gggenes: Draw Gene Arrow Maps in ‘ggplot2’_. R Package Version 0.5.1. 2023. Available online: https://CRAN.R-project.org/package=gggenes (accessed on 1 July 2025).
  87. Cupples, A.M.; Thelusmond, J.R. Predicting the occurrence of monooxygenases and their associated phylotypes in soil microcosms. J. Microbiol. Methods 2022, 193, 106401. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  88. Faghihinezhad, M.; Eshghdoostkhatami, Z.; Bernstein, A.; Cupples, A.M. Identification of the dominant methanotrophs in trichloroethene degrading enrichment cultures from multiple sources. J. Hazard. Mater. 2025, 499, 140268. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  89. Cupples, A.M.; Li, Z.; Wilson, F.P.; Ramalingam, V.; Kelly, A. In silico analysis of soil, sediment and groundwater microbial communities to predict biodegradation potential. J. Microbiol. Methods 2022, 202, 106595. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  90. Hand, S.; Wang, B.; Chu, K.-H. Biodegradation of 1,4-dioxane: Effects of enzyme inducers and trichloroethylene. Sci. Total Environ. 2015, 520, 154–159. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  91. Lippincott, D.; Streger, S.H.; Schaefer, C.E.; Hinkle, J.; Stormo, J.; Steffan, R.J. Bioaugmentation and propane biosparging for in situ biodegradation of 1,4-dioxane. Groundw. Monit. Remediat. 2015, 35, 81–92. [Google Scholar] [CrossRef] [Scilit]
  92. Bell, C.H.; Wong, J.; Parsons, K.; Semel, W.; McDonough, J.; Gerbe, K. First full-scale in situ propane biosparging for co-metabolic bioremediation of 1,4-dioxane. Groundw. Monit. Remediat. 2022, 42, 54–66. [Google Scholar] [CrossRef] [Scilit]
  93. Kotani, T.; Yurimoto, H.; Kato, N.; Sakai, Y. Novel acetone metabolism in a propane-utilizing bacterium, Gordonia sp. strain TY-5. J. Bacteriol. 2007, 189, 886–893. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  94. Kotani, T.; Kawashima, Y.; Yurimoto, H.; Kato, N.; Sakai, Y. Gene structure and regulation of alkane monooxygenases in propane-utilizing Mycobacterium sp. TY-6 and Pseudonocardia sp. TY-7. J. Biosci. Bioeng. 2006, 102, 184–192. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  95. Vainberg, S.; McClay, K.; Masuda, H.; Root, D.; Condee, C.; Zylstra, G.J.; Steffan, R.J. Biodegradation of ether pollutants by Pseudonocardia sp. strain ENV478. Appl. Environ. Microbiol. 2006, 72, 5218–5224. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  96. Hamamura, N.; Arp, D.J. Isolation and characterization of alkane-utilizing Nocardioides sp. strain CF8. FEMS Microbiol. Lett. 2000, 186, 21–26. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  97. Tupa, P.R.; Masuda, H. Comparative proteomic analysis of propane metabolism in Mycobacterium sp. Strain ENV421 and Rhodococcus sp. Strain ENV425. J. Mol. Microbiol. Biotechnol. 2018, 28, 107–115. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  98. Cupples, A.M. Propane monooxygenases in soil associated metagenomes align most closely to those in the genera Kribbella, Amycolatopsis, Bradyrhizobium, Paraburkholderia and Burkholderia. Curr. Microbiol. 2024, 81, 314. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  99. Hristova, K.R.; Schmidt, R.; Chakicherla, A.Y.; Legler, T.C.; Wu, J.; Chain, P.S.; Scow, K.M.; Kane, S.R. Comparative transcriptome analysis of Methylibium petroleiphilum PM1 exposed to the fuel oxygenates methyl tert-butyl ether and ethanol. Appl. Environ. Microbiol. 2007, 73, 7347–7357. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  100. Farhan Ul Haque, M.; Crombie, A.T.; Murrell, J.C. Novel facultative Methylocella strains are active methane consumers at terrestrial natural gas seeps. Microbiome 2019, 7, 134. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  101. Crombie, A.T.; Murrell, J.C. Trace-gas metabolic versatility of the facultative methanotroph Methylocella silvestris. Nature 2014, 510, 148–151. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  102. Chen, Y.; Crombie, A.; Rahman, M.T.; Dedysh, S.N.; Liesack, W.; Stott, M.B.; Alam, M.; Theisen, A.R.; Murrell, J.C.; Dunfield, P.F. Complete genome sequence of the aerobic facultative methanotroph Methylocella silvestris BL2. J. Bacteriol. 2010, 192, 3840–3841. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  103. Wang, J.; Geng, K.; Farhan Ul Haque, M.; Crombie, A.; Street, L.E.; Wookey, P.A.; Ma, K.; Murrell, J.C.; Pratscher, J. Draft genome sequence of Methylocella silvestris TVC, a facultative methanotroph isolated from permafrost. Genome Announc. 2018, 6, e00040-18. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  104. Deng, D.; Li, F.; Wu, C.; Li, M. Synchronic biotransformation of 1,4-dioxane and 1,1-dichloroethylene by a gram-negative propanotroph Azoarcus sp. DD4. Environ. Sci. Technol. Lett. 2018, 5, 526–532. [Google Scholar] [CrossRef] [Scilit]
  105. Faghihinezhad, M.; Eshghdoostkhatami, Z.; Cupples, A.M. Characterization of multiple trichloroethene, cis-dichloroethene and 1,1-dichloroethene degrading propanotrophic communities. J. Environ. Manag. 2026, 408, 129957. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  106. Op den Camp, H.J.; Islam, T.; Stott, M.B.; Harhangi, H.R.; Hynes, A.; Schouten, S.; Jetten, M.S.; Birkeland, N.K.; Pol, A.; Dunfield, P.F. Environmental, genomic and taxonomic perspectives on methanotrophic Verrucomicrobia. Environ. Microbiol. Rep. 2009, 1, 293–306. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  107. Cardy, D.L.; Laidler, V.; Salmond, G.P.; Murrell, J.C. Molecular analysis of the methane monooxygenase (MMO) gene cluster of Methylosinus trichosporium OB3b. Mol. Microbiol. 1991, 5, 335–342. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  108. Cardy, D.L.; Laidler, V.; Salmond, G.P.; Murrell, J.C. The methane monooxygenase gene cluster of Methylosinus trichosporium: Cloning and sequencing of the mmoC gene. Arch. Microbiol. 1991, 156, 477–483. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  109. Stainthorpe, A.C.; Murrell, J.C.; Salmond, G.P.; Dalton, H.; Lees, V. Molecular analysis of methane monooxygenase from Methylococcus capsulatus (Bath). Arch. Microbiol. 1989, 152, 154–159. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  110. Stainthorpe, A.C.; Lees, V.; Salmond, G.P.; Dalton, H.; Murrell, J.C. The methane monooxygenase gene cluster of Methylococcus capsulatus (Bath). Gene 1990, 91, 27–34. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  111. Csaki, R.; Bodrossy, L.; Klem, J.; Murrell, J.C.; Kovacs, K.L. Genes involved in the copper-dependent regulation of soluble methane monooxygenase of Methylococcus capsulatus (Bath): Cloning, sequencing and mutational analysis. Microbiology 2003, 149, 1785–1795. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  112. Wutkowska, M.; Nweze, J.A.; Tlaskal, V.; Nweze, J.E.; Daebeler, A. Uncovering hidden phylo- and ecogenomic diversity of the widespread methanotrophic genus Methylobacter. FEMS Microbiol. Ecol. 2026, 102, fiaf127. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  113. Iguchi, H.; Yurimoto, H.; Sakai, Y. Soluble and particulate methane monooxygenase gene clusters of the type I methanotroph Methylovulum miyakonense HT12. FEMS Microbiol. Lett. 2010, 312, 71–76. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  114. Kalyuzhnaya, M.G.; Lamb, A.E.; McTaggart, T.L.; Oshkin, I.Y.; Shapiro, N.; Woyke, T.; Chistoserdova, L. Draft genome sequences of gammaproteobacterial methanotrophs isolated from lake washington sediment. Genome Announc. 2015, 3, e00103-15. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  115. Koh, S.C.; Bowman, J.P.; Sayler, G.S. Soluble methane monooxygenase production and trichloroethylene degradation by a Type I methanotroph, Methylomonas methanica 68-1. Appl. Environ. Microbiol. 1993, 59, 960–967. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  116. Boden, R.; Cunliffe, M.; Scanlan, J.; Moussard, H.; Kits, K.D.; Klotz, M.G.; Jetten, M.S.; Vuilleumier, S.; Han, J.; Peters, L.; et al. Complete genome sequence of the aerobic marine methanotroph Methylomonas methanica MC09. J. Bacteriol. 2011, 193, 7001–7002. [Google Scholar] [CrossRef] [Scilit] [PubMed][Green Version]
  117. Dedysh, S.N.; Liesack, W.; Khmelenina, V.N.; Suzina, N.E.; Trotsenko, Y.A.; Semrau, J.D.; Bares, A.M.; Panikov, N.S.; Tiedje, J.M. Methylocella palustris gen. nov., sp. nov., a new methane-oxidizing acidophilic bacterium from peat bogs, representing a novel subtype of serine-pathway methanotrophs. Int. J. Syst. Evol. Microbiol. 2000, 50, 955–969. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  118. Dunfield, P.F.; Khmelenina, V.N.; Suzina, N.E.; Trotsenko, Y.A.; Dedysh, S.N. Methylocella silvestris sp. nov., a novel methanotroph isolated from an acidic forest cambisol. Int. J. Syst. Evol. Microbiol. 2003, 53, 1231–1239. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  119. Dedysh, S.N.; Berestovskaya, Y.Y.; Vasylieva, L.V.; Belova, S.E.; Khmelenina, V.N.; Suzina, N.E.; Trotsenko, Y.A.; Liesack, W.; Zavarzin, G.A. Methylocella tundrae sp. nov., a novel methanotrophic bacterium from acidic tundra peatlands. Int. J. Syst. Evol. Microbiol. 2004, 54, 151–156. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  120. Sheu, C.; Cai, C.Y.; Sheu, S.Y.; Li, Z.H.; Chen, W.M. Pseudomethylobacillus aquaticus gen. nov., sp. nov., a new member of the family Methylophilaceae isolated from an artificial reservoir. Int. J. Syst. Evol. Microbiol. 2019, 69, 3551–3559. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  121. Yu, Z.; Beck, D.A.C.; Chistoserdova, L. Natural selection in synthetic communities highlights the roles of Methylococcaceae and Methylophilaceae and suggests differential roles for alternative methanol dehydrogenases in methane consumption. Front. Microbiol. 2017, 8, 2392. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  122. Shao, Y.; Hatzinger, P.B.; Streger, S.H.; Rezes, R.T.; Chu, K.H. Evaluation of methanotrophic bacterial communities capable of biodegrading trichloroethene (TCE) in acidic aquifers. Biodegradation 2019, 30, 173–190. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  123. Newby, D.T.; Reed, D.W.; Petzke, L.M.; Igoe, A.L.; Delwiche, M.E.; Roberto, F.F.; McKinley, J.P.; Whiticar, M.J.; Colwell, F.S. Diversity of methanotroph communities in a basalt aquifer. FEMS Microbiol. Ecol. 2004, 48, 333–344. [Google Scholar] [CrossRef] [PubMed]
  124. Hwangbo, M.; Rezes, R.; Chu, K.H.; Hatzinger, P.B. Evaluation of microbial community dynamics and chlorinated solvent biodegradation in methane-amended microcosms from an acidic aquifer. Biodegradation 2024, 36, 8. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Subunit length and arrangement of propane monooxygenase operons in groundwater metagenomes associated with the identified metagenome-assembled genomes within the phylum Actinomycetota.
Figure 1. Subunit length and arrangement of propane monooxygenase operons in groundwater metagenomes associated with the identified metagenome-assembled genomes within the phylum Actinomycetota.
Biology 15 01586 g001
Figure 2. Subunit length and arrangement of propane monooxygenase operons in groundwater metagenomes associated with the identified metagenome-assembled genomes within the phyla Pseudomonadota and Chloroflexota.
Figure 2. Subunit length and arrangement of propane monooxygenase operons in groundwater metagenomes associated with the identified metagenome-assembled genomes within the phyla Pseudomonadota and Chloroflexota.
Biology 15 01586 g002
Figure 3. Alignment of prmA from MAGs containing the propane monooxygenase operon within the phyla Actinomycetota (class Actinomycetia, red circles; class Rubrobacteria, orange circles; class Thermoleophilia, pink circles), Pseudomonadota (class Alphaproteobacteria, blue squares; class Gammaproteobacteria, green squares) and Chloroflexota (purple triangles). The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown next to the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 36 nucleotide sequences. There were a total of 1714 positions in the final dataset. Evolutionary analyses were conducted in MEGA12.
Figure 3. Alignment of prmA from MAGs containing the propane monooxygenase operon within the phyla Actinomycetota (class Actinomycetia, red circles; class Rubrobacteria, orange circles; class Thermoleophilia, pink circles), Pseudomonadota (class Alphaproteobacteria, blue squares; class Gammaproteobacteria, green squares) and Chloroflexota (purple triangles). The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown next to the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 36 nucleotide sequences. There were a total of 1714 positions in the final dataset. Evolutionary analyses were conducted in MEGA12.
Biology 15 01586 g003
Figure 4. Alignment of mmoX from MAGs containing the soluble monooxygenase operon within the Methylomonadaceae (green circles), Methylophilaceae (blue triangles) and Beijerinckiaceae (red diamonds) families. The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown next to the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to a matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 27 nucleotide sequences. There were a total of 1605 positions in the final dataset. Evolutionary analyses were conducted in MEGA12.
Figure 4. Alignment of mmoX from MAGs containing the soluble monooxygenase operon within the Methylomonadaceae (green circles), Methylophilaceae (blue triangles) and Beijerinckiaceae (red diamonds) families. The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown next to the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to a matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 27 nucleotide sequences. There were a total of 1605 positions in the final dataset. Evolutionary analyses were conducted in MEGA12.
Biology 15 01586 g004
Figure 5. Alignment of tmoA from MAGs containing the soluble monooxygenase operon within the family Burkholderiaceae (Burkholderiales, Gammaproteobacteria): Hydrogenophaga (red circles), Rhodoferax (blue circles), Malikia (purple circles), Sphaerotilus (orange circle), Polaromonas (brown circle), JAAFJR01 (light blue circle), Ramlibacter (light green circle), Trinickia (green circle), Limnobacter (light brown circle) and LX47W (light purple circle). The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown next to the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to a matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 30 nucleotide sequences. There were a total of 1508 positions in the final dataset. Evolutionary analyses were conducted in MEGA12.
Figure 5. Alignment of tmoA from MAGs containing the soluble monooxygenase operon within the family Burkholderiaceae (Burkholderiales, Gammaproteobacteria): Hydrogenophaga (red circles), Rhodoferax (blue circles), Malikia (purple circles), Sphaerotilus (orange circle), Polaromonas (brown circle), JAAFJR01 (light blue circle), Ramlibacter (light green circle), Trinickia (green circle), Limnobacter (light brown circle) and LX47W (light purple circle). The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown next to the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to a matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 30 nucleotide sequences. There were a total of 1508 positions in the final dataset. Evolutionary analyses were conducted in MEGA12.
Biology 15 01586 g005
Figure 6. Alignment of tmoA from MAGs containing the soluble monooxygenase operon within the Rhodocyclaceae, Nevskiaceae (Gammaproteobacteria) and Alphaproteobacteria: Azonexus (blue triangle), Rugosibacter (green triangle), Zoogloea (red triangle), Nevskia (purple triangle), Pinisolibacter (grey triangle), and CAIXXC01 (dark blue triangle) families. The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown below the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to a matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 26 nucleotide sequences. There were a total of 1510 positions in the final dataset. Evolutionary analyses were conducted in MEGA11.
Figure 6. Alignment of tmoA from MAGs containing the soluble monooxygenase operon within the Rhodocyclaceae, Nevskiaceae (Gammaproteobacteria) and Alphaproteobacteria: Azonexus (blue triangle), Rugosibacter (green triangle), Zoogloea (red triangle), Nevskia (purple triangle), Pinisolibacter (grey triangle), and CAIXXC01 (dark blue triangle) families. The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown below the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to a matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 26 nucleotide sequences. There were a total of 1510 positions in the final dataset. Evolutionary analyses were conducted in MEGA11.
Biology 15 01586 g006
Figure 7. Alignment of dmpN from MAGs containing the phenol monooxygenase operon, including Hydrogenophaga (red triangle), Rhodoferax (blue circles), Sphaerotilus (orange circle), Polaromonas (brown circle), Ramlibacter (light green circle), Trinickia (green circle), Limnobacter (light brown circle), Azonexus (blue triangle), and Zoogloea (red triangle). The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown next to the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to a matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 34 nucleotide sequences. There were a total of 1573 positions in the final dataset. Evolutionary analyses were conducted in MEGA11.
Figure 7. Alignment of dmpN from MAGs containing the phenol monooxygenase operon, including Hydrogenophaga (red triangle), Rhodoferax (blue circles), Sphaerotilus (orange circle), Polaromonas (brown circle), Ramlibacter (light green circle), Trinickia (green circle), Limnobacter (light brown circle), Azonexus (blue triangle), and Zoogloea (red triangle). The evolutionary history was inferred by using the Maximum Likelihood method and the General Time Reversible model. The percentage of trees in which the associated taxa clustered together is shown next to the branches. Initial tree(s) for the heuristic search were obtained automatically by applying Neighbor-Join and BioNJ algorithms to a matrix of pairwise distances estimated using the Maximum Composite Likelihood (MCL) approach and then selecting the topology with superior log likelihood value. The tree is drawn to scale, with branch lengths measured in the number of substitutions per site. This analysis involved 34 nucleotide sequences. There were a total of 1573 positions in the final dataset. Evolutionary analyses were conducted in MEGA11.
Biology 15 01586 g007
Table 1. Examined whole genome sequencing datasets obtained from NCBI’s SRA.
Table 1. Examined whole genome sequencing datasets obtained from NCBI’s SRA.
BioProject Accession Site InformationNumber of Samples and the Total Number of Bases Analyzed in the Current Study Reference
PRJNA288027Groundwater metagenomes, Rifle, CO, USA25 samples, ~500 Gbases[37,38,39]
PRJNA512237Groundwater metagenomes, Ohio9 samples, ~13 Gbases[40]
PRJNA699054Groundwater and sediment metagenomes, New Zealand25 samples, ~55 Gbases[41]
PRJNA513876Groundwater metagenomes, Tennessee, USA12 samples, ~400 Gbases[42]
PRJNA1001011Sediment and groundwater metagenomes from subsurface microbial communities from the Oak Ridge National Laboratory Field Research Center, Oak Ridge, TN, USA114 samples, ~926 Gbases[43,44]
PRJNA321556Groundwater metagenomes, Horonobe Underground Research Laboratory, Hokkaido, Japan9 samples, ~130 Gbases[45]
PRJNA700657Groundwater metagenomes, Alberta, Cananda26 samples, ~395 Gbases[46]
PRJNA1162924US Navy Site, USA22 samples, ~107 Gbases[47]
PRJNA279923Metagenome of deep subsurface groundwater in Äspö, Sweden15 samples, ~120 Gbases[48]
PRJNA858913Groundwater metagenomes from samples from
monitoring wells in seven geo-environment zones across China
>300 samples, >30 Gbases per sample[49]
PRJNA916093Groundwater metagenomes, Bangladesh6 samples, ~22 Gbases[50]
PRJEB36505Groundwater metagenomes, within the Hainich Critical Zone Exploratory, Turingia, Germany32 samples, ~626 Gbases[51]
PRJNA947390Groundwater metagenomes from three groundwater monitoring wells in an agricultural area in the Netherlands18 samples, ~112 Gbases[52]
Table 2. Amino acid composition of both di-iron centers (DE*RH) in the MAGs.
Table 2. Amino acid composition of both di-iron centers (DE*RH) in the MAGs.
Operon and MAGFirst Di-Iron CenterSecond Di-Iron Center
Propane monooxygenase
All MAGs in the phyla Actinomycetota and ChloroflexotaDE V RHDE S RH
All MAGs in the phylum PseudomonadotaDE F RHDE S RH
Soluble methane monooxygenase
All MAGsDE I RHDE L RH
Toluene monooxygenase
Rhodocyclaceae, Nevskiaceae (Gammaproteobacteria), Alphaproteobacteria
Azonexus (all except 1)DE I RHDE S RH
Rugosibacter (2 of 3)DE I RHDE S RH
ZoogloeaDE I RHDE S RH
NevskiaDE I RHDE S RH
Azonexus (1)DE V RHDE S RH
Rugosibacter (1 of 3)DE V RHDE S RH
PinisolibacterDR N RHDE S RH
Burkholderiaceae (Burkholderiales, Gammaproteobacteria)
HydrogenophagaDR N RHDE S RH
SphaerotilusDR N RHDE S RH
RamlibacterDR N RHDE S RH
TrinickiaDR N RHDE S RH
StellaceaeDR N RHDE S RH
Rhodoferax (1)DE I RHDE S RH
BurkholderiaceaeDE I RHDE S RH
LimnbacterDE I RHDE S RH
LX47WDE I RHDE S RH
PolaromonasDE M RHDE S RH
Rhodoferax (3)DE M RHDE S RH
MalikiaDE V RHDE S RH
Rhodoferax (4)DE V RHDE S RH
Phenol monooxygenase
Rhodocyclaceae, Nevskiaceae (Gammaproteobacteria), Alphaproteobacteria
AzonexusDE I RHDE A RH
ZoogloeaDE I RHDE A RH
Burkholderiaceae (Burkholderiales, Gammaproteobacteria)
HydrogenophagaDE L RHDE S RH
SphaerotilusDE I RHDE A RH
RamlibacterDE L RHDE S RH
TrinickiaDE L RHDE S RH
LimnobacterDE L RHDE S RH
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Cupples, A.M.; Richards, J.; Basaldua Del Cid, M. Data Mining of Groundwater Genomes for Metagenome-Assembled Genomes (MAGs) Containing Monooxygenase Operons Associated with Contaminant Biodegradation. Biology 2026, 15, 1586. https://doi.org/10.3390/biology15181586

AMA Style

Cupples AM, Richards J, Basaldua Del Cid M. Data Mining of Groundwater Genomes for Metagenome-Assembled Genomes (MAGs) Containing Monooxygenase Operons Associated with Contaminant Biodegradation. Biology. 2026; 15(18):1586. https://doi.org/10.3390/biology15181586

Chicago/Turabian Style

Cupples, Alison M., James Richards, and Maria Basaldua Del Cid. 2026. "Data Mining of Groundwater Genomes for Metagenome-Assembled Genomes (MAGs) Containing Monooxygenase Operons Associated with Contaminant Biodegradation" Biology 15, no. 18: 1586. https://doi.org/10.3390/biology15181586

APA Style

Cupples, A. M., Richards, J., & Basaldua Del Cid, M. (2026). Data Mining of Groundwater Genomes for Metagenome-Assembled Genomes (MAGs) Containing Monooxygenase Operons Associated with Contaminant Biodegradation. Biology, 15(18), 1586. https://doi.org/10.3390/biology15181586

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop