Next Article in Journal
Kinetic Analysis of the Photocatalytic Degradation of Indigo Carmine Using a Heterogeneous MgAl–LDH Catalyst
Next Article in Special Issue
Synthetic Dyes in Textile Wastewater: Classification, Environmental Risks, and Microbiological and Enzymatic Remediation Strategies
Previous Article in Journal
Correction: Heba et al. Green Dynamic Kinetic Resolution—Stereoselective Acylation of Secondary Alcohols by Enzyme-Assisted Ruthenium Complexes. Catalysts 2022, 12, 1395
Previous Article in Special Issue
Functional Expression of the Aromatic Prenyltransferase NphB in Chlamydomonas reinhardtii Highlights Challenges in Cannabinoid Biocatalysis
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Machine Learning-Guided Enzyme Engineering Approaches for Enhanced Biocatalytic Efficiency: Concepts, Mechanisms, and Future Directions

Department of Pharmaceutical Chemistry and Pharmacognosy, College of Pharmacy, Jazan University, P.O. Box 114, Jazan 45142, Saudi Arabia
Catalysts 2026, 16(7), 598; https://doi.org/10.3390/catal16070598
Submission received: 15 April 2026 / Revised: 3 June 2026 / Accepted: 16 June 2026 / Published: 30 June 2026
(This article belongs to the Special Issue Biocatalysis and Biosynthesis: Opportunities and Challenges)

Abstract

Biocatalysis has emerged as a mainstay in the field of sustainable chemical synthesis owing to its high selectivity, mild reaction conditions, and reduced environmental impact. Traditional enzyme engineering approaches, such as rational design and directed evolution, are often associated with limited throughput and a limited understanding of sequence–structure–function relationships, despite high experimental costs. In recent years, the integration of machine learning (ML) into enzyme engineering has emerged as a transformative approach, enabling data-driven prediction, design, and optimization of biocatalysts, thereby enhancing performance and applications. This review provides a comprehensive overview of ML-guided strategies to improve key enzymatic parameters, including the turnover number (kcat), substrate affinity (Km), and catalytic efficiency (kcat/Km), with a focus on mechanistic insights and performance outcomes. The integration of ML models into design–build–test–learn (DBTL) cycles accelerated directed evolution, reduced screening efforts, and enabled targeted mutagenesis. Beyond applications, this review also discusses the current limitations of ML-guided approaches, including data scarcity, model interpretability, and challenges in predicting complex mutations and allosteric effects. The gap between computational predictions and experimental outcomes is identified, and the role of ML integration with enzyme kinetics, molecular dynamics, and high-throughput experimentation is emphasized. Future directions, such as generative AI, explainable ML, and autonomous laboratories, are discussed for next-generation biocatalytic applications.

Graphical Abstract

1. Introduction

Enzyme engineering is considered a cornerstone of modern biocatalysis, having led to the development of highly selective and sustainable catalytic systems with potential applications in the pharmaceutical, chemical, and biotechnology sectors. The increasing applications of enzymes and their ability to catalyze complex and stereoselective transformations have provided numerous advantages over traditional chemical catalysis. Enzymatic processes have been successfully employed in the production of active pharmaceutical ingredients (APIs), particularly chiral intermediates, for which high regio- and enantioselectivity are required to achieve therapeutic efficacy and regulatory compliance [1,2].
Large-scale adoption of enzymes in industrial settings is often limited by numerous constraints, such as suboptimal catalytic efficiency, narrow substrate scope, and limited stability under processing conditions, including high temperatures, extreme pH, and solvent exposure. Catalytic efficiency is commonly expressed as the ratio of the turnover number (kcat) to the Michaelis constant (Km) and is a crucial parameter governing enzyme performance and process feasibility. Therefore, a central challenge in enzyme engineering remains the improvement of these kinetic parameters while maintaining or enhancing the stability and selectivity of enzymes [3,4,5].
Traditional enzyme engineering approaches, such as directed evolution and rational design, have been successfully used to improve enzyme functions. Directed evolution involves iterative cycles of random mutagenesis and high-throughput screening to identify improved variants, whereas rational design relies primarily on structural and mechanistic insights to introduce targeted mutations in enzymes. Despite their successful applications, these approaches suffer from numerous limitations. Directed evolution is labor-intensive and explores only a small fraction of the vast protein sequence space, whereas rational design is constrained by inadequate understanding of complex sequence–structure–function relationships and often fails to anticipate synergistic mutational effects [6,7]. Therefore, the development of enzymes as efficient and robust biocatalysts remains a challenging, time-consuming, and resource-intensive process.
Recently, machine learning (ML) has emerged as a powerful strategy to address these limitations. ML algorithms have enabled data-driven exploration of protein sequence space and helped predict enzyme properties efficiently [7,8,9,10]. It can learn complex, non-linear relationships among enzyme sequences, structures, and functions from large-scale biological datasets, which could help identify beneficial mutations with substantially reduced experimental effort. ML can be applied to enzyme engineering to predict catalytic activity, thermostability, substrate specificity, and mutational effects, as well as to design novel protein sequences with tailored functionalities [11,12,13].
ML has facilitated a paradigm shift from empirical, trial-and-error enzyme engineering approaches to a predictive, iterative optimization framework. The integration of ML into the design–build–test–learn (DBTL) cycles also enabled continuous refinement of the predictive models through experimental feedback, thereby accelerating enzyme optimization and minimizing screening requirements. Recent advancements in this field have demonstrated that ML-guided approaches have been shown to deliver substantial improvements in enzyme performance, particularly in catalytic efficiency and substrate specificity [14,15,16].
Despite these advancements, critical challenges remain in integrating ML into enzyme engineering. The foremost challenge is the predictive accuracy of ML models, as it is often limited by the availability and quality of experimental data, particularly for kinetic parameters such as kcat and Km. Moreover, many ML models act as black boxes and offer limited mechanistic interpretability, restricting their utility in the rational enzyme design. Discrepancies between in silico predictions and experimental outcomes are frequently observed because protein dynamics, solvent effects, and complex allosteric interactions are not fully captured by currently developed ML models [12,17,18].
Therefore, there is a growing need for more focused and critically evaluated ML-guided enzyme engineering strategies that improve catalytic efficiency and mechanistic understanding. Although several previously published reviews have addressed the applications of ML in enzyme engineering, many focus on algorithm development, protein design workflows, or broad applications. Limited emphasis was given to the quantitative improvements in the catalytic parameters and their mechanistic interpretation. This review primarily adopts a performance- and mechanism-oriented perspective on ML-based enzyme engineering approaches and discusses ML’s contributions to improving key enzymatic parameters, including kcat, Km, and kcat/Km. It aims to provide a comprehensive overview of enzyme efficiency and its determinants, examine ML-based approaches specifically for enzyme optimization, analyze quantitative advancements achieved through ML-assisted strategies, and critically evaluate their current limitations. This review includes the most recent advancements in this field from literature published after 2023, including protein language models (PLMs), inverse-folding approaches, generative diffusion-based frameworks, benchmarking datasets, and autonomous optimization systems. Future directions that could help achieve the next generation of biocatalysts are also provided. This review will help researchers working in biocatalysis, computational biology, and enzyme engineering by providing both conceptual understanding and practical insights for efficiency-driven enzyme design.

2. Fundamentals of Enzyme Efficiency

The biocatalytic performance and industrial applicability of an enzyme are determined by its efficiency, which is governed by the rate, selectivity, and overall productivity of the enzymatic reactions. The enzyme efficiency is quantitatively evaluated based on kinetic parameters derived from the Michaelis–Menten framework, including kcat, Km, and the kcat/Km ratio, which collectively characterize catalytic performance across varying substrate concentrations.

2.1. Kinetic Parameters

There are three key kinetic parameters kcat, Km, and kcat/Km, which are measures of the enzyme efficiency. The kcat represents the maximum number of substrate molecules that are converted to products per active site per unit time under saturating substrate conditions. It is a measure of an enzyme’s intrinsic catalytic capability and is directly associated with the rate-limiting step of the catalytic cycle. In contrast, the Km reflects substrate affinity and represents the substrate concentration at which the reaction rate reaches half of its maximum velocity. A lower Km value indicates stronger substrate binding; however, it does not necessarily correlate with the faster catalysis process [19,20]. Similarly, the ratio kcat/Km is often referred to as the catalytic efficiency or specificity constant, and it integrates both catalytic turnover and substrate binding. It provides a comprehensive measure of enzyme performance, specifically under low substrate concentrations. This parameter is particularly relevant in physiological and industrial contexts where substrate availability is limited. Enzymes with diffusion-controlled limits of 108–109 M−1s−1 are considered catalytically perfect, as their reaction rates are constrained primarily by substrate diffusion rather than by the chemical transformation [21,22].

2.2. Mechanistic Determinants of Enzyme Efficiency

Enzyme efficiency is governed by a complex interplay among structural, dynamic, and physicochemical factors that influence substrate binding, transition-state stabilization, and product release. These mechanistic determinants include active-site structure, protein dynamics, substrate binding, and the stability–activity trade-off, as shown in Figure 1.

2.2.1. Active Site Structure and Transition State Stabilization

The active site of an enzyme provides a highly specialized microenvironment that facilitates the catalysis process through the precise positioning of catalytic residues, substrates, and cofactors. The rate acceleration achieved by enzymes is primarily by stabilizing the transition state, which lowers the activation energy barrier. Even minor changes in the geometry and electrostatics of the active site can influence catalytic rates; therefore, this region is the primary target for enzyme engineering [23].

2.2.2. Protein Dynamics and Conformational Flexibility

In contrast to the classical lock-and-key theory, enzymes are dynamic molecules that can undergo conformational changes during catalysis. These include motions ranging from local side-chain rearrangements to large-domain movements and can affect substrate recognition, catalytic turnover, and product release. These dynamic effects critically influence enzyme efficiency and are generally challenging to predict using static structural models [24].

2.2.3. Substrate Binding and Specificity

An optimal balance between substrate binding and turnover is required for efficient catalysis, as excessively tight binding can reduce catalytic turnover by limiting product release, and weak binding may reduce reaction rates due to poor substrate recognition [25]. Enzyme specificity is therefore determined by a combination of steric, electronic, and hydrophobic interactions inside the binding pocket, which together define the substrate scope and selectivity [26].

2.2.4. Stability–Activity Trade-Off

The trade-off between stability and catalytic activity remains a critical challenge in enzyme engineering, as mutations that enhance thermostability often reduce conformational flexibility and may impair an enzyme’s catalytic function. Mutations that increase activity can destabilize protein structure; therefore, balancing these competing effects is a major focus of modern enzyme design strategies [27,28].

2.3. Limitations of Conventional Optimization Approaches

Although traditional methods for improving enzyme efficiency, such as directed evolution and rational design, have achieved substantial success over the years, they are often associated with inherent limitations. Directed evolution mainly relies on random mutagenesis and screening, which makes it inefficient while targeting the improvements in multiple kinetic parameters simultaneously. In addition, the vast combinatorial sequence space, which typically contains >1013 possible variants even for smaller proteins, makes exhaustive exploration practically impossible [3,29]. Similarly, rational design depends upon the prior knowledge of enzyme structure and mechanism. The structure-guided approaches, although they can identify key mutation residues, often fail to explain long-range interactions, epistasis, and dynamic properties influencing catalytic efficiency at large. Eventually, several engineered variants failed to achieve the desired improvements in kcat or kcat/Km upon experimental validation [30].

2.4. Implications of ML-Guided Enzyme Engineering

The non-linear interactions among sequence, structure, and dynamics complicate enzyme efficiency and pose challenges for conventional enzyme engineering approaches. ML offers a potential alternative to these approaches by enabling the identification of hidden patterns in large datasets and the successful prediction of mutation effects on key kinetic parameters without exhaustive experimentation [31]. In addition, the ML-based models can efficiently capture epistatic interactions and multidimensional relationships that are challenging to model using traditional enzyme engineering approaches. Importantly, the success of ML-guided approaches depends on their ability to accurately report key efficiency parameters such as kcat, Km, and kcat/Km. Thereby, a thorough understanding of the fundamental determinants of enzyme efficiency is warranted for the rational application of the ML-based techniques. Integration of kinetic theory with data-driven modeling enables the design of enzymes with improved catalytic performance using minimal experimental effort [32].

3. ML-Based Approaches in Enzyme Optimization

Predicting and improving catalytic properties directly from sequence and structural data make ML-based approaches a central tool in enzyme engineering. Conventional artificial intelligence (AI) overviews generally describe algorithms and tools, whereas ML in enzyme optimization focuses on mapping sequence–structure–function relationships to predict advantageous mutations and improve catalytic performance. Several ML-guided strategies that directly enhance enzyme efficiency are described below.

3.1. Data Representation and Feature Engineering

One of the most critical factors for the success of ML models in enzyme optimization is the representation of protein sequence and structure data. Previous approaches utilized handcrafted features, including the amino acid composition, physicochemical descriptors, and evolutionary conservation scores obtained from multiple sequence alignments. Although these characteristics yielded valuable insights, they often failed to capture the intricate non-linear relationships that govern enzyme functionality. Recent advancements in representation learning have improved the ability to capture information from protein sequences and structures relevant to enzyme function and catalysis [33,34,35,36]. When combined with sequence-based representations, structural features such as the active-site geometry, solvent accessibility, and residue–residue interactions provide complementary inputs for model training.

3.2. Predictive Modeling of Enzyme Properties

ML models are widely adopted to predict key enzymatic properties relevant to their optimization, such as catalytic activity, thermostability, and substrate specificity. When trained on curated datasets, classical algorithms showed strong performance in the modeling of enzyme function [37,38]. These models can capture non-linear dependencies and potential interactions among amino acid residues, which are crucial for modeling the effects of mutations on enzyme efficiency [39,40].

3.3. Prediction of Mutation Effects and Sequence Optimization

Predicting the effects of mutations on enzyme performance is one of the most important applications of ML-based methods in enzyme engineering. Sequence-to-function ML models trained on mutational datasets can effectively predict the effects of single- and multiple-amino-acid substitutions on enzymatic activity, stability, and specificity. These models may capture patterns consistent with epistatic interactions, where the effects of one mutation depend on the presence of others, although accurately predicting higher-order epistasis remains challenging [37]. ML-based sequence optimization can also be used to propose candidate variants for experimental validation and functional improvement [41].

3.4. ML-Guided Library Design

Another important advantage of ML in enzyme optimization is its ability to guide the design of focused mutant libraries. The classical directed evolution process depends on large, randomly generated libraries, which often require screening thousands to millions of variants. On the contrary, ML-guided approaches enable the selection of a small subset of highly promising candidates, thereby reducing the number of variants to be screened [42]. Prediction of sequence–function relationships allows the ML models to identify hotspot residues and optimal mutation combinations. Such strategies substantially reduce the requirements for experimental screening while maintaining a high probability of identifying the beneficial variants [14,41].

3.5. Integration with Multimodal Data and Hybrid Modeling

Modern ML-guided approaches are designed to integrate multiple data types, such as sequence, structure, and experimental measurements, to improve the prediction accuracy. Hybrid models combine ML with molecular docking and simulations to integrate corresponding physicochemical information [43]. These integrative approaches are useful for predicting not only static properties but also dynamic behaviors that influence catalysis, such as conformational flexibility and substrate-binding pathways [12].

4. Advanced ML Frameworks, Benchmarking, and Model Selection

The effectiveness of ML-guided enzyme optimization processes not only depends on the choice of ML model but also on data representation, benchmarking strategies, and appropriate validation protocols. Several advanced ML frameworks are applied in enzyme engineering and are considered while selecting models and benchmarking.

4.1. Classical ML Approaches

Random forests (RF), support vector machines (SVM), and gradient boosting machines (GBM) are traditional ML algorithms widely adopted in enzyme engineering for their unmatched robustness and interpretability [31,44]. These models rely mainly on the engineered features, including amino acid composition, physicochemical descriptors, structural parameters, and evolutionary conservation scores [45,46,47].

4.2. Deep Learning (DL) Architectures

Classical ML approaches are associated with several limitations, and DL-based approaches have been applied to overcome them by enabling end-to-end learning directly from raw sequence or structural data. Identification of local sequence motifs can be achieved using CNNs [47], while graph neural networks (GNNs) model proteins as residue interaction networks, capturing spatial relationships within three-dimensional structures [48]. PLMs, including the ESM-1v, ESM-2, and ProtTrans, have been trained on millions of protein sequences and can generate embeddings that encode structural and functional information without explicit labeling [49,50,51].

4.3. Zero-Shot and Transfer Learning Approaches

The development of zero-shot models has been a major advancement in ML-guided enzyme engineering, enabling prediction of mutation effects without task-specific training. The Evolutionary Model of Variant Effect (EVE) and ESM-based predictor models can provide evolutionary information derived from multiple sequence alignments or from unsupervised training on large-scale datasets [50,51].

4.4. Generative Models and De Novo Enzyme Design

Generative ML models have enabled the design of novel protein sequences rather than the optimization of existing ones. These include variational autoencoders (VAEs), generative adversarial networks (GANs), and diffusion-based models such as RFdiffusion, which are increasingly used to explore unexplored regions of the protein sequence space [52,53,54,55]. Sequences compatible with the provided protein structure can be designed using inverse folding models such as Protein MPNN and ESM-IF, ultimately enabling structure-guided enzyme optimization [56,57].

4.5. Benchmarking and Validation Strategies

The ML models must be evaluated for reliability, an essential factor for ensuring their applicability in enzyme engineering. To standardize their performance assessment, several benchmarking frameworks have been developed. These include Tasks Assessing Protein Embeddings (TAPE), which evaluates models on multiple protein prediction tasks [58]; Fitness Landscape Inference for Proteins (FLIP), which focuses on the prediction of mutation effects across diverse fitness landscapes [59]; and ProteinGym, which is a large-scale benchmark dataset used to predict variant effects using deep mutational scanning data [60].

4.6. Model Selection Criteria

The selection of an appropriate ML model depends on several factors, including the size and type of the dataset, the target objective, and computational resources. Additionally, interpretability requirements play a major role in selecting an appropriate ML model. While DL models offer higher predictive power, classical and explainable AI approaches offer mechanistic insights valuable for the rational enzyme design. Generally, for small datasets, classical ML models such as RF and SVM are preferred to reduce the risk of overfitting. For medium-sized datasets, hybrid approaches are preferred, combining feature engineering with DL. Large datasets require DL and PLMs as they provide superior performance. For structure-driven tasks, GNNs and inverse folding models are preferred, whereas generative models such as VAEs, GANs, and diffusion models are applied for exploratory design. A practical decision-making workflow that maps dataset size, prediction objectives, and the availability of structural information to the most appropriate ML model families is shown in Figure 2.
In practice, the ML model selection typically depends on dataset size, prediction target, availability of structural information, and computational resources. Zero-shot PLMs are particularly useful when experimentally labeled datasets are limited, whereas supervised deep-learning approaches generally require substantially larger datasets for robust generalization. Structure-guided tasks may further benefit from graph neural networks and inverse-folding frameworks that explicitly incorporate three-dimensional structural information.
Various ML frameworks used in enzyme engineering are analyzed and compared in Table 1.

5. ML-Driven Catalytic Performance Enhancement

Overall, ML-guided approaches have improved enzyme engineering workflows by enabling more targeted, data-driven optimization of catalytic performance. The true potential of ML lies not only in predicting enzyme properties but also in systematically enhancing their catalytic performance through targeted enzyme design strategies. ML contributes to improving kcat, Km, and kcat/Km and emphasizes mechanistic interpretation and experimentally validated outcomes.

5.1. Enhancement of Turnover Number (kcat)

One of the most challenging aspects of enzyme engineering is improving kcat, as it is governed by the rate-limiting step of the catalytic cycle, which involves transition-state stabilization, bond formation or cleavage, and product release [61,62,63]. These findings highlighted the importance of long-range interactions and demonstrated ML’s ability to uncover non-intuitive mutational hotspots. Mechanistically, these enhancements were attributed to improved transition-state stabilization, alignment of catalytic residues, or rapid conformational transitions required for catalysis [64]. Those ML models, which incorporate structural and dynamic descriptors, can predict mutations that influence catalytic turnover without affecting enzyme stability.

5.2. Optimization of Substrate Affinity (Km) and Binding Interactions

The ML-based approaches have also been applied to optimizing substrate binding, as reflected by changes in the Michaelis constant (Km). Traditional structure-based methods primarily focus on active-site residues, whereas ML models can assess the effects of both local and distal residues on substrate recognition and binding affinity [65]. ML algorithms can predict mutations that could enhance binding interactions through improved electrostatic complementarity, hydrophobic packing, and hydrogen bonding networks [66]. More importantly, the ML-based optimization of Km should balance the binding strength with catalytic turnover. Excessively tight binding may reduce turnover, as products are released slowly, whereas moderate improvements in binding affinity can enhance overall catalytic efficiency [67]. Therefore, ML models that simultaneously evaluate kcat and Km are particularly important for achieving optimal enzyme performance.

5.3. Enhancement of Catalytic Efficiency (kcat/Km)

Enzyme performance is measured by its catalytic efficiency, which integrates both turnover number and substrate affinity. The ML-guided approaches have been applied to improve this parameter by identifying mutations that enhance both catalytic turnover and substrate binding simultaneously [68,69]. The enhancement in catalytic efficiency was primarily attributed to coordinated changes in the active-site geometry, substrate positioning, and protein dynamics. ML-based models that are trained on combinatorial datasets can capture patterns consistent with epistatic interactions. However, accurate prediction of higher-order epistasis remains challenging, particularly when training data are limited and combinatorial diversity is insufficient. Therefore, ML can assist in identifying optimal combinations of mutations to maximize catalytic efficiency, but experimental validation is essential.

5.4. Epistasis Modeling in ML-Guided Enzyme Engineering

Epistasis, defined as a non-additive interaction between mutations, is one of the major challenges in enzyme engineering. The functional effects of a mutation largely depend on the presence or absence of other mutations within the protein sequence [70]. These interactions complicate predictions of enzyme activity, stability, and selectivity, particularly in combinatorial mutagenesis and directed evolution strategies, where multiple substitutions are introduced simultaneously. Conventional enzyme engineering approaches often assume that mutational effects are additive; however, it has been demonstrated that even beneficial mutations can become neutral or deleterious when combined [71]. It is therefore imperative to accurately predict the higher-order mutational interactions to efficiently explore the protein fitness landscapes.
Advanced ML-based approaches have shown potential to capture these epistatic relationships by learning complex sequence–function patterns from large mutational datasets [72,73]. Deep learning architectures, such as transformer-based PLMs and GNNs, have been shown to model long-range residue dependencies and co-evolutionary relationships that are difficult to identify with conventional statistical approaches [72]. It has been demonstrated that the ML-assisted direct evolution can improve navigation of complex combinatorial fitness landscapes while reducing the experimental screening burden [74]. In addition, PLMs such as ESM and EVE have been shown to predict mutational fitness effects from large-scale sequence data [50,51,75]. However, accurate prediction of higher-order epistasis remains challenging as most models rely primarily on statistical correlations rather than mechanistic understanding [76]. Therefore, integrating ML with structural biology, molecular dynamics simulations, and mechanistic enzymology is a forward-looking strategy to improve the reliability and interpretability of epistasis-aware enzyme engineering.

5.5. Mechanistic Insights from ML-Guided Mutations

In addition to enhancing catalytic performance, ML-guided enzyme engineering provides valuable mechanistic insights into the factors that influence catalytic efficiency [77]. A critical analysis of ML-predicted mutations revealed a few common mechanisms. The first is distal mutations, which affect active-site dynamics, demonstrating the importance of protein-wide communication networks [78]. Another mechanism is the alteration of conformational flexibility, thereby facilitating substrate binding and product release [79]. The third mechanism is the improvement in transition-state stabilization through optimization of the electrostatic environment [80]. Eventually, the hydrogen bonding networks are reorganized to improve the positioning of catalytic residues [81]. These insights emphasize that enzyme efficiency is not solely governed by active-site interactions but reflects a comprehensive interplay among the protein’s structure and dynamics. ML models that utilize the structural and evolutionary information can predict these complex relationships and can translate them into optimal design strategies [82,83,84,85,86]. The quantitative impact of ML on the kinetic parameters and their key contributing factors is illustrated in Figure 3.
Previous investigations reported integration of ML with enzyme/protein engineering, and the key outcomes are summarized in Table 2. The diversity of enzyme classes, model strategies, training datasets, validation approaches, and reported performance gains is summarized. As the table shows, most studies report activity, conversion, selectivity, or enrichment rather than complete kinetic descriptors such as kcat, Km, and kcat/Km, thereby limiting direct cross-study comparison. The quantitative ranges discussed in Table 2 and Figure 3 were derived from experimentally validated studies that reported activity, conversion, selectivity, specific activity, and improvements in kinetic performance.
Currently used ML models rely on statistical correlations between sequence features and enzyme performance; however, next-generation ML approaches are expected to leverage both mechanistic and physicochemical insights, providing more accurate and interpretable predictions [95]. Moreover, the integration of ML with MD simulations, quantum-mechanical calculations, and enzyme kinetics modeling will offer deeper insights into catalytic mechanisms and improved predictions of mutational effects [96]. Therefore, mechanism-aware ML models are warranted and have the potential to move beyond empirical optimization to enable the rational design of highly efficient biocatalysts.

6. Integration of ML with Directed Evolution

ML integrated with directed evolution marks a radical shift from random, high-throughput experimentation to iterative, data-driven frameworks. The Design–Build–Test–Learn (DBTL) cycle is essential to this transformation, enabling continual fine-tuning of enzyme variants through feedback-driven learning (Figure 4). It serves as a structured framework for integrating ML into enzyme optimization. It is an iterative process that ensures predictive models inform subsequent experimental designs, resulting in faster development and refinement of enzymes with enhanced catalytic properties and efficiencies [97]. Traditional directed evolution relies on random mutagenesis, while ML-guided DBTL workflows improve the efficiency and precision of enzyme optimization and reduce experimental efforts [98].

6.1. The Design Phase (Generation of a Hypothesis)

The design phase in the DBTL cycle uses predictive models to identify promising mutations or sequence variants. When trained on sequence–function datasets, ML algorithms can rank residues and mutation combinations most likely to enhance catalytic efficiency. This focused approach differs greatly from random mutagenesis, in which vast libraries are generated with low hit rates. It is important to note that ML-driven design can simultaneously address multiple objectives, such as increasing catalytic turnover while maintaining stability [99]. ML models can generate hypotheses that capture both local active-site features and global protein properties using sequence embeddings, structural descriptors, and prior experimental data. This enables detection of non-obvious mutations, including distal residues that can alter enzyme dynamics and overall catalytic performance.

6.2. The Build Phase (Construction of the Library)

Once the design phase is completed, the selected mutations are tested to create focused mutant libraries. The ML-guided libraries are typically smaller and more enriched than those constructed using traditional methods, and the number of variants is often reduced from millions to merely hundreds or even fewer [91]. To generate these libraries, techniques including site-directed mutagenesis, combinatorial mutagenesis, and synthetic gene assembly are used. The smaller library size not only reduces experimental costs but also makes it more efficient to characterize each variant, including determining key kinetic parameters important for assessing catalytic efficiency.

6.3. The Test Phase (High-Throughput Screening and Kinetic Evaluation)

The test phase in the cycle involves experimental evaluation of the engineered variants using the high-throughput screening or the targeted kinetic assay techniques. ML-guided workflows often prioritize quality over quantity, emphasizing precise measurement of enzyme performance over large-scale screening. Recent investigations have revealed that ML-integrated platforms, when combined with automated screening systems and cell-free expression technologies, enable rapid evaluation of enzyme variants across multiple conditions [100,101]. It enables the construction of high-quality datasets that capture not only activity but also kinetic parameters, thereby providing valuable insights for subsequent learning cycles [13]. The incorporation of quantitative kinetic data into the ML models enhances their predictive capacity and enables more accurate optimization of catalytic efficiency in the following iterations.

6.4. The Learning Phase (Model Refinement and Feedback Integration)

This is the defining feature of the DBTL cycle, in which the experimental data obtained are fed back into the ML models, thereby improving their predictive accuracy. It is an iterative learning process that enables ML models to continuously refine themselves and learn sequence–function relationships, including epistatic interactions and non-linear effects [102]. ML models can update their predictions by incorporating the newly generated experimental data and guide the next round of design with enhanced precision. This closed-loop system transforms the enzyme engineering into a self-improving system, and each feedback iteration improves both the model and the enzyme variants being developed [103]. Active learning strategies are developed and implemented within the DBTL cycle, in which the refined model selectively identifies the most informative experiments to perform. Therefore, this approach maximizes information gain and minimizes experimental effort, further accelerating the optimization process [7].

6.5. Closed-Loop Optimization and Autonomous Systems

The DBTL framework has recently been extended to closed-loop and semi-autonomous systems, in which ML models are integrated with robotic platforms and automated experimental setups [104]. These systems ensure that the entire cycle, from design to learning, is executed automatically with minimal human intervention. In these systems, ML-driven prediction is combined with high-throughput experimentation and real-time data analysis, enabling rapid optimization of enzyme performance [31]. Closed-loop systems have previously achieved remarkable improvements in catalytic efficiency within a few iterative cycles, highlighting the potential of autonomous enzyme engineering [14].

6.6. Advantages of ML-Integrated Directed Evolution over Traditional Approaches

The integration of ML into the DBTL cycle offers several advantages over conventional directed evolution, including a reduced screening burden, as it enables focused library design (Figure 5). It also results in an improved hit rate due to targeted mutation selection, faster convergence toward optimal variants, the ability to optimize multiple parameters simultaneously, including kcat, Km, stability, etc., and an enhanced understanding of sequence–function relationships through iterative learning [92]. These advantages make the ML-guided DBTL workflows particularly well-suited for industrial applications, where time, cost, and scalability are crucial considerations.

6.7. Challenges in the Integration of ML into the DBTL Workflow

Despite its immense potential, integration of ML-guided DBTL cycles is associated with numerous challenges. The foremost challenge is that the cycle’s effectiveness depends heavily on the quality and diversity of experimental data, which are limited for certain classes of enzymes [12]. In addition, discrepancies between predicted and experimental outcomes can propagate through subsequent iterative cycles if not addressed properly. Another challenge is integrating heterogeneous data types, such as sequence, structural, and kinetic data, into a unified ML model. Therefore, data consistency and standardization must be ensured across iterations to maintain the reliability of these models. Adopting automated and closed-loop systems requires substantial computational infrastructure and expertise, thereby limiting their accessibility in certain research settings.

7. Case Studies Concerning ML-Guided Enzyme Engineering

Representative case studies demonstrate the impact of ML in enzyme engineering and illustrate how ML-guided approaches have achieved quantitative improvements in catalytic efficiency, reduced experimentation, and revealed mechanistic insights. The experimentally validated ML-assisted enzyme engineering workflows are depicted in Figure 6, and the important case studies are summarized in Table 3.
One of the pioneering case studies in this field is the ML-guided optimization of an amide synthetase enzyme, in which an augmented ridge-regression model trained on high-throughput sequence-activity datasets was used to prioritize productive variants for experimental validation [87]. Focused mutant libraries were constructed using ML-based predictions and evaluated using high-throughput screening and kinetic assay methods. Experimentally validated variants demonstrated substantial improvements in catalytic activity (kcat) relative to the parent enzyme and a marked reduction in the number of screened variants compared with conventional random mutagenesis (Table 2). The sequence-activity data generated through HTS experiments were used to train the supervised learning model. Subsequently, an experimental evaluation of computationally prioritized variants was conducted to enable focused exploration of sequence space compared with conventional random-library screening.
Another breakthrough study reported improved catalytic efficiency and operational stability of the transaminase enzyme, which was subsequently applied to chiral amine synthesis [89]. Sequence-derived physicochemical descriptors and structure-informed features were integrated into supervised ML models to prioritize beneficial mutations. Selected variants of the enzyme were expressed and evaluated for kinetic performance, and the parameters kcat and Km were calculated. A significant improvement in catalytic efficiency (kcat/Km), improved performance at neutral pH, and a reduced need for extensive screening were reported. The sequence- and structure-derived descriptors were utilized as model inputs to prioritize mutations for experimental evaluation. The ML-guided strategy reduced the need for experimental screening by identifying the variants with improved catalytic efficiency under near-neutral conditions, which could not be achieved using conventional directed evolution workflows.
In another study, generative ML-based approaches were used to expand the functional sequence space, enabling exploration of novel enzyme variants beyond naturally occurring sequences [31]. In this study, GANs and related ML models were used to design new protein sequences with predicted functional activity, and the models were trained on large datasets to detect underlying patterns in protein structures and functions. The selected generated sequences were then synthesized and experimentally validated to assess enzymatic activity, and the experimentally validated enzyme variants retained catalytic properties, demonstrating the feasibility of examining non-natural regions of sequence space [41].
ML-guided directed evolution was applied to enhance the catalytic activity and regioselectivity of cytochrome P450 enzymes in selective oxidation reactions [105]. The optimization workflow combined ML-assisted fitness prediction with iterative directed-evolution cycles using experimentally generated sequence-activity datasets. The developed model incorporated prior mutation data and experimental results from previous cycles. A closed-loop system was implemented, in which the variants predicted by the model were tested via targeted screening, with the results then fed back into the model for subsequent optimization cycles. The findings showed that iterative ML-assisted optimization improved the catalytic activity and regioselectivity toward desired products, reduced the number of experimental cycles required for variant identification, and achieved a more rapid convergence than traditional approaches [106]. The predictive framework integrated experimental data from prior optimization rounds, thereby iteratively refining mutation prioritization within the DBTL workflow.
ML was subsequently applied to the enzyme optimization using cell-free expression systems to evaluate a large number of enzyme variants simultaneously [94]. In these systems, thousands of enzyme variants were screened in parallel, and the resulting data were used to refine ML predictions. The key findings of this study included the identification of enzyme variants with substantially improved activity, a marked reduction in experimental complexity, and improved predictive accuracy through iterative learning. The integration of ML into cell-free systems enabled efficient exploration of sequence space and detection of variants with improved catalytic efficiency.
Therefore, the integration of ML has had a transformative impact on enzyme engineering, particularly in enhancing catalytic efficiency through targeted, data-driven approaches. ML-guided strategies enabled quantitative improvements, reduced experimental effort, and uncovered key mechanical insights, thereby offering a robust framework for the rational design of high-performance biocatalysts.

8. Applications of ML-Guided Enzyme Engineering

The ML-driven optimization of catalytic efficiency has been applied to pharmaceutical synthesis, antimicrobial development, natural products synthesis, and sustainable biocatalysis.

8.1. Pharmaceutical Biocatalysis and Chiral Synthesis

One of the most important applications of ML-guided enzyme engineering is in the pharmaceutical industry, where enzymes are used to synthesize complex and chiral molecules [87]. Enzymes with high catalytic efficiency and selectivity are essential in this context to afford product purity, reduce unwanted reactions, and meet stringent regulatory requirements. Several enzymes, including transaminases, ketoreductases, and cytochrome P450 enzymes, have been optimized for enantioselective transformations using ML-guided approaches. ML models enable the development of biocatalysts with enhanced performance in the synthesis of active pharmaceutical ingredients (APIs) by predicting beneficial mutations that improve substrate binding and catalytic turnover [107]. Previously, the ML-assisted enzyme engineering resulted in remarkable improvements in kcat/Km with faster reaction rates and reduced enzyme loading [93]. Using ML, multiple parameters such as activity, selectivity, and stability can be optimized simultaneously across different process conditions, including high substrate concentrations and non-aqueous media [47].

8.2. Antimicrobial Enzyme Engineering and Biotherapeutics

Another important application of ML-guided enzyme engineering is in the development of antimicrobial agents and enzyme-based therapeutics. This application is particularly useful given the global rise in antimicrobial resistance. Several enzymes, such as hydrolases, oxidoreductases, and lyases, can be engineered using ML-based models and used to disrupt cell wall components, biofilms, and microbial virulence factors [108,109]. ML models can help identify mutations that can enhance enzyme activity against specific microbial targets, improve substrate specificity, and increase stability under physiological conditions. For instance, ML-guided optimization of substrate–enzyme interactions has been used to enhance the catalytic efficiency of enzymes targeting biofilms, thereby improving antimicrobial activity. These improvements are associated with better substrate recognition and increased catalytic turnover, driven by the efficiency-driven design of enzymes [12]. Subsequently, ML can assist in the design of enzymes with improved pharmacokinetic properties and reduced immunogenicity, and it demonstrates remarkable potential as a biotherapeutic.

8.3. Green Chemistry and Sustainable Chemical Synthesis

Sustainability is one of the major driving forces behind the adoption of biocatalysis in chemical engineering, and ML has contributed to this field by enabling the development of efficient biocatalysts that work under mild conditions, minimize the usage of hazardous chemicals, and reduce waste generation [110]. ML-guided approaches have led to the development of enzymes that have been applied to the synthesis of fine chemicals, agrochemicals, and specialty compounds with improved catalytic efficiency [87]. It has even resulted in higher yields, lower energy consumption, faster reaction rates, improved substrate conversion, and reduced overall production costs. Using ML-based strategies, enzymes capable of catalyzing non-natural or challenging reactions can be identified. This further expands the scope of biocatalysts beyond traditional transformations [1].

8.4. Biocatalysts for the Synthesis and Modification of Natural Products

Natural products and their derivatives are a rich source of bioactive molecules with pharmaceutical and cosmeceutical applications. However, their structures are complex and are often challenging to synthesize chemically. ML-guided enzyme engineering can help tailor biosynthetic pathways and modify natural products with high specificity [63]. The enzymes involved in biosynthetic pathways can be optimized using ML, thereby enhancing catalytic efficiency and substrate flexibility and enabling the production of novel derivatives with improved biological properties [111,112]. This is particularly relevant to plant secondary metabolites such as terpenes, alkaloids, and phenolic compounds, where even a minor structural modification can affect pharmacological properties [13].

8.5. Emerging Applications

The applications of ML-guided enzyme engineering have expanded beyond established domains, with new areas emerging, including environmental biocatalysis, biosensing, and materials science [113]. For instance, ML-designed enzymes engineered for enhanced catalytic efficiency are being used to degrade environmental pollutants, such as plastics and toxic chemicals, and are contributing to waste management and environmental remediation [114,115]. Enzymes with optimized catalytic properties also find applications in biosensing, as they can improve sensitivity and response time, enabling more accurate and highly sensitive detection of analytes in clinical and environmental samples [116].

9. Critical Limitations of ML-Guided Enzyme Engineering

Despite remarkable progress in ML-guided enzyme engineering, several limitations undermine its predictive accuracy, generalizability, and practical applicability. These challenges to its widescale adoption arise from both data-related and methodological constraints, as well as the inherent complexity of enzyme catalysis (Figure 7). Understanding and addressing these limitations is essential for the effective application of ML in the biocatalytic process.

9.1. Limited Availability and Quality of Kinetic Data

The scarcity of high-quality, experimentally validated datasets remains a major bottleneck in ML-guided enzyme engineering, particularly for kinetic parameters such as kcat and Km [31]. Although large-scale sequence databases are widely available, quantitative kinetic data are comparatively scarce and heterogeneous, and are often reported under different experimental conditions [117]. The ability of ML models to learn accurate sequence–function relationships is limited by the lack of standardized datasets. In addition, inconsistencies in experimental protocols and reporting formats introduce bias and noise into the training datasets, further reducing model reliability and reproducibility [12].

9.2. Poor Generalization Across Enzyme Classes

Another important limitation is that ML models are largely trained on datasets derived from a limited number of enzymes, resulting in poor generalization when applied to novel and less-characterized enzyme classes [118]. This becomes particularly relevant when predicting the effects of mutations in enzymes with low sequence homology to the training data. The protein sequence space is highly dimensional, and when combined with limited labeled data, this makes it difficult for ML models to extrapolate beyond known regions. Consequently, ML-based predictions may perform well only for specific enzyme classes but fail when applied to structurally or functionally distinct systems [7].

9.3. Inadequate Representation of Protein Dynamics

The enzyme-catalysis process is considered inherently dynamic, involving conformational changes that occur over multiple timescales [119]. However, several ML models rely solely on static protein structure representations and fail to capture the dynamic nature of enzyme function [120]. The catalytic process involves important steps, including substrate binding, transition-state stabilization, and product release, which are often governed by conformational flexibility and transient structural states [121]. The currently developed ML models are unable to fully account for these dynamic effects, which limits their accuracy in predicting the catalytic efficiency. This is particularly important for mutations that influence protein motion rather than the static structure [24].

9.4. Challenges in Modeling Epistasis and Multi-Mutation Effects

Currently employed ML models can predict single-point mutation effects, but accurately capturing epistatic interactions among multiple mutations remains a critical challenge [122]. In most cases, the combined effects of mutations are non-additive, and individual beneficial mutations may result in neutral or even detrimental effects when added together. While several advanced models can account for pairwise interactions, higher-order epistasis remains poorly understood and challenging to model [123]. Therefore, the ability of ML-guided approaches to predict optimal multi-mutation combinations reliably remains limited [13].

9.5. Lack of Mechanistic Interpretability

A major challenge for ML-based approaches in enzyme engineering is their limited interpretability, as many ML models, particularly the deep learning architectures, act as black boxes that provide predictions without proper mechanistic explanations [124,125]. This results in a lack of transparency, which is challenging for the rational enzyme design, as the mechanistic understanding is often required to validate and trust the predicted mutations. This is particularly crucial in the industrial and pharmaceutical settings, where regulatory considerations are important, and the inability of these models to explain predictions can ultimately hinder the adoption of ML-guided approaches [12].

9.6. Mismatch Between Predictions and Experimental Outcomes

This is a recurring issue in ML-guided enzyme engineering, as mutations predicted by the model to enhance catalytic efficiency may fail experimentally [63]. This may be attributed to numerous factors not captured by computational models, including solvent effects and reaction environment, cofactor interactions, protein folding and expression efficiency, and allosteric regulation and long-range structural effects. These discrepancies emphasize the limitations of purely data-driven models and underscore the need to integrate ML with experimental outcomes and mechanistic insights [126]. Recent benchmarking studies have shown that zero-shot PLMs can achieve moderate-to-strong correlations with experimentally determined mutational fitness landscapes [68]. For instance, ProteinGym benchmarking reported Spearman correlation values ranging from ~0.4 to 0.7, depending on the protein family and assay type [60]. In a similar study, it was demonstrated that the unsupervised sequence models could successfully capture substantial functional constraints across diverse proteins. However, predictive accuracy decreased for higher-order mutational combinations and for poorly represented sequence families.

9.7. Data Bias and Overfitting

As ML models are highly sensitive to the composition of training datasets, overrepresentation of certain enzyme classes, substrates, or experimental conditions may result in biased models that perform well only on familiar data and poorly on new systems [127]. Another concern is overfitting, which occurs when models are trained on small, high-dimensional datasets. In such cases, models may memorize the training data but fail to learn generalizable patterns, leading to reduced predictive performance on unseen data [14].

9.8. Computational and Infrastructure Constraints

While the integration of ML reduces experimental workload, it also introduces computational challenges, particularly for large-scale models such as deep neural networks (DNNs) and PLMs [128]. Training and deployment of these models require substantial computational resources, high-performance computing infrastructure, and specialized expertise. Moreover, the integration of ML models with experimental workflows, such as high-throughput screening and automated platforms, further increases resource demands, ultimately limiting their accessibility in some research labs [129].

9.9. Lack of Standardization and Benchmarking

The lack of standardized datasets, evaluation metrics, and benchmarking protocols constitutes another major challenge in the integration of ML with enzyme engineering [130]. Different studies use distinct datasets, performance metrics, and experimental conditions, making it difficult to compare outcomes across studies. The absence of standardization further hinders the development of robust and generalizable ML models and decelerates their progress toward practical implementation [131]. Therefore, establishing community-wide benchmarking and standardized reporting practices is essential for advancing ML-guided enzyme engineering.

10. Future Perspectives

The integration of ML with enzyme engineering and high-throughput experimentation is rapidly transforming the field of biocatalysis into a predictive, design-driven discipline. While current strategies have demonstrated marked improvements in catalytic efficiency, the next phase of development will focus on deeper integration of computational intelligence, mechanistic understanding, and automation. The future directions in ML-guided enzyme engineering are depicted in Figure 8.

10.1. Mechanism-Aware and Physics-Informed ML

A major limitation of current ML approaches, as mentioned earlier, is their reliance on statistical correlations rather than a mechanistic understanding. This can be addressed by developing mechanism-aware ML models that integrate enzyme kinetics, thermodynamics, and structural dynamics into predictive networks [63,132]. These models can be developed by incorporating data from MD simulations, quantum-mechanical calculations, and transition-state modeling into ML pipelines. This will result in a more accurate prediction of catalytic parameters such as kcat and activation energy barriers. These hybrid systems can transform enzyme engineering from empirical optimization to physics-informed design, in which predictions are supported by fundamental catalytic principles.

10.2. Generative AI for De Novo Enzyme Design

Generative AI models, such as transformer-based PLMs, diffusion models, and variational autoencoders, are the future of de novo enzyme design [133]. They differ from traditional approaches that modify existing enzymes by enabling the creation of entirely new protein sequences with desired catalytic properties. Future research focuses on improving the functional ability of generated sequences and ensuring that the designed enzymes achieve high catalytic efficiency in experiments. By integrating generative models with experimental validation platforms, rapid exploration of previously inaccessible regions of the protein sequence space has been achieved, thereby further expanding the scope of biocatalysis [134].

10.3. Autonomous and Self-Driving Laboratories

The emergence of autonomous laboratories has been the most crucial development in enzyme engineering. In this approach, ML models are integrated with robotics and high-throughput experimentation to create closed-loop optimization systems. In these systems, the entire DBTL cycle is automated and continuous, with iterative optimization of enzyme variants achieved with minimal human intervention [135]. These platforms are highly advantageous as they can rapidly explore large sequence spaces, identify optimal mutations, and refine predictive models in real time. These self-driving laboratories can dramatically accelerate enzyme engineering workflows and enable the development of highly efficient biocatalysts within shorter timeframes than traditional approaches [136].

10.4. Explainable Artificial Intelligence (XAI)

The increasing complexity of ML models often requires interpretability and transparency. The XAI aims to provide insights into how ML models make predictions and help understand the underlying factors influencing the enzyme performance [137]. In enzyme engineering, XAI can identify critical residues, interaction networks, and structural features that govern enzyme catalytic efficiency [138]. This interpretability is particularly valuable for the rational design of enzymes and for building trust in the ML-guided approaches.

10.5. Integration with Systems Biology and Metabolic Engineering

Future enzyme engineering efforts should move beyond single-enzyme optimization toward pathway- and system-level design, and ML models should be used to optimize entire biosynthetic pathways by coordinating the activities of multiple enzymes and improving overall metabolic flux and product yield [139]. The integration of ML with systems biology will enable the design of efficient microbial cell factories to produce a variety of pharmaceuticals, biofuels, and high-value chemicals. ML-guided pathway optimization is expected to play a crucial role in balancing enzyme activities, minimizing bottlenecks, and improving the overall process efficiency [140,141].

10.6. Data Standardization and Collaborative Platforms

The availability of high-quality, standardized datasets is essential for advancing ML-guided enzyme engineering, and future efforts should focus on developing open-access databases that include comprehensive kinetic, structural, and experimental data [47]. Collaborative platforms can be developed by integrating experimental data with ML tools, thereby facilitating knowledge sharing and accelerating innovation in the field [98]. Standardization of data formats, experimental protocols, and evaluation metrics will be essential to improve the generalizability and reproducibility of future ML models.

10.7. Toward Precision and Personalized Biocatalysis

An emerging frontier in ML-guided enzyme engineering is precision biocatalysis, in which enzymes are tailored to specific substrates, reaction conditions, or applications [47]. Future ML-driven strategies should enable the customization of enzymes to perform highly specialized tasks such as personalized medicine, targeted drug synthesis, and niche industrial processes [142]. ML can be integrated with advanced screening and validation techniques to design enzymes with unprecedented specificity and efficiency, tailored to meet precise functional requirements [112].

11. Conclusions

In summary, ML has fundamentally reshaped the enzyme engineering process by enabling a shift from empirical, trial-and-error approaches to predictive, efficiency-driven strategies. By capturing the complex sequence–structure–function relationships, ML-driven methods showed substantial improvements in key catalytic parameters, including ~10- to 50-fold increases in catalytic turnover, ~1.5- to 5-fold enhancements in substrate affinity, and up to ~100-fold improvements in catalytic efficiency in selected systems involving nuclease, ketoreductase, and engineered catalytic variants. The ML-guided approaches also reduced the experimental screening burden through focused library design and iterative DBTL-driven optimization. Modern frameworks such as PLMs, transformer architectures, and ML-guided directed evolution have helped identify beneficial and non-intuitive mutational combinations that are difficult to detect using conventional approaches alone. Several case studies in this field have demonstrated that ML-driven approaches can deliver measurable impact in pharmaceutical biocatalysis, antimicrobial drug development, and sustainable chemical and natural product synthesis.
Despite these advancements, critical challenges associated with the ML-guided approaches beset their widespread adoption. The most pressing issues include the scarcity of standardized, experimentally validated kinetic datasets, limited predictive accuracy for higher-order epistatic interactions, and a persistent gap between predictions and experimental outcomes, particularly for certain enzyme systems. However, addressing these limitations through improved integration of ML with enzyme kinetics, molecular modeling, and high-throughput experimentation, along with the development of standardized datasets and interpretable models, is warranted. The PLMs, such as ESM, and benchmarking platforms, such as TAPE, FLIP, and ProteinGym, have provided practical starting points for ML-assisted enzyme engineering and fitness prediction. Future directions, including mechanism-aware ML, generative AI, and autonomous laboratory platforms, have the potential to further accelerate enzyme discovery and optimization. Converging these technologies will enable the rational design of highly efficient, application-specific biocatalysts, transforming biocatalysis into a precise, scalable, and intelligent engineering field.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

The authors gratefully acknowledge the research facilities provided by Jazan University.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Sheldon, R.A.; Woodley, J.M. Role of biocatalysis in sustainable chemistry. Chem. Rev. 2018, 118, 801–838. [Google Scholar] [PubMed]
  2. Bornscheuer, U.T.; Hauer, B.; Jaeger, K.-E.; Schwaneberg, U. Directed evolution empowered redesign of natural proteins for the sustainable production of chemicals and pharmaceuticals. Angew. Chem. Int. Ed. 2019, 58, 36–40. [Google Scholar]
  3. Arnold, F.H. Directed evolution: Bringing new chemistry to life. Angew. Chem. Int. Ed. 2018, 57, 4143–4148. [Google Scholar] [CrossRef]
  4. Wu, S.; Snajdrova, R.; Moore, J.C.; Baldenius, K.; Bornscheuer, U.T. Biocatalysis: Enzymatic synthesis for industrial applications. Angew. Chem. Int. Ed. 2021, 60, 88–119. [Google Scholar]
  5. Bayer, T.; Wu, S.; Snajdrova, R.; Baldenius, K.; Bornscheuer, U.T. An update: Enzymatic synthesis for industrial applications. Angew. Chem. Int. Ed. 2025, 64, e202505976. [Google Scholar] [CrossRef]
  6. Romero, P.A.; Arnold, F.H. Exploring protein fitness landscapes by directed evolution. Nat. Rev. Mol. Cell Biol. 2009, 10, 866–876. [Google Scholar] [CrossRef] [PubMed]
  7. Yang, K.K.; Wu, Z.; Arnold, F.H. Machine-learning-guided directed evolution for protein engineering. Nat. Methods 2019, 16, 687–694. [Google Scholar] [PubMed]
  8. Patsch, D.; Buller, R. Improving enzyme fitness with machine learning. Chimia 2023, 77, 116–121. [Google Scholar] [CrossRef] [PubMed]
  9. Li, Z.L.; Pei, S.; Chen, Z.; Huang, T.Y.; Wang, X.D.; Shen, L.; Chen, X.; Wang, Q.Q.; Wang, D.X.; Ao, Y.F. Machine learning-assisted amidase-catalytic enantioselectivity prediction and rational design of variants for improving enantioselectivity. Nat. Commun. 2024, 15, 8778. [Google Scholar] [PubMed]
  10. Cadet, X.F.; Gelly, J.C.; van Noord, A.; Cadet, F.; Acevedo-Rocha, C.G. Learning strategies in protein directed evolution. Methods Mol. Biol. 2022, 2461, 225–275. [Google Scholar] [CrossRef] [PubMed]
  11. Siedhoff, N.E.; Schwaneberg, U.; Davari, M.D. Machine learning-assisted enzyme engineering. Methods Enzymol. 2020, 643, 281–315. [Google Scholar] [CrossRef] [PubMed]
  12. Mazurenko, S.; Prokop, Z.; Damborsky, J. Machine learning in enzyme engineering. ACS Catal. 2020, 10, 1210–1223. [Google Scholar]
  13. Wittmann, B.J.; Johnston, K.E.; Wu, Z.; Arnold, F.H. Advances in machine learning for directed evolution. Curr. Opin. Struct. Biol. 2021, 69, 11–18. [Google Scholar] [CrossRef] [PubMed]
  14. Fox, R. Directed molecular evolution by machine learning and the influence of nonlinear interactions. J. Theor. Biol. 2005, 234, 187–199. [Google Scholar] [CrossRef] [PubMed]
  15. Huang, C.; Zhang, L.; Tang, T.; Wang, H.; Jiang, Y.; Ren, H.; Zhang, Y.; Fang, J.; Zhang, W.; Jia, X.; et al. Application of directed evolution and machine learning to enhance the diastereoselectivity of ketoreductase for dihydrotetrabenazine synthesis. JACS Au 2024, 4, 2547–2556. [Google Scholar] [CrossRef] [PubMed]
  16. Ao, Y.F.; Dörr, M.; Menke, M.J.; Born, S.; Heuson, E.; Bornscheuer, U.T. Data-driven protein engineering for improving catalytic activity and selectivity. ChemBioChem 2024, 25, e202300754. [Google Scholar] [PubMed]
  17. Strokach, A.; Becerra, D.; Corbi-Verge, C.; Perez-Riba, A.; Kim, P.M. Fast and flexible protein design using deep graph neural networks. Cell Syst. 2020, 11, 402–411. [Google Scholar] [CrossRef] [PubMed]
  18. Zhang, Y.; Chen, Y.; Wang, C.; Lo, C.C.; Liu, X.; Wu, W.; Zhang, J. ProDCoNN: Protein design using a convolutional neural network. Proteins 2020, 88, 819–829. [Google Scholar] [PubMed]
  19. Segel, I.H. Enzyme Kinetics: Behavior and Analysis of Rapid Equilibrium and Steady-State Enzyme Systems; Wiley: New York, NY, USA, 1993. [Google Scholar]
  20. Seibert, E.; Tracy, T.S. Fundamentals of enzyme kinetics. Methods Mol. Biol. 2014, 1113, 9–22. [Google Scholar] [CrossRef] [PubMed]
  21. Bar-Even, A.; Noor, E.; Savir, Y.; Liebermeister, W.; Davidi, D.; Tawfik, D.S.; Milo, R. The moderately efficient enzyme: Evolutionary and physicochemical trends shaping enzyme parameters. Biochemistry 2011, 50, 4402–4410. [Google Scholar] [CrossRef] [PubMed]
  22. Labourel, F.; Rajon, E. Resource uptake and the evolution of moderately efficient enzymes. Mol. Biol. Evol. 2021, 38, 3938–3952. [Google Scholar] [CrossRef] [PubMed]
  23. Warshel, A.; Sharma, P.K.; Kato, M.; Parson, W.W. Modeling electrostatic effects in proteins. Biochim. Biophys. Acta 2006, 1764, 1647–1676. [Google Scholar] [CrossRef] [PubMed]
  24. Henzler-Wildman, K.; Kern, D. Dynamic personalities of proteins. Nature 2007, 450, 964–972. [Google Scholar] [CrossRef] [PubMed]
  25. Li, H.; Xie, Y.; Liu, C.; Liu, S. Physicochemical bases for protein folding, dynamics, and protein–ligand binding. Sci. China Life Sci. 2014, 57, 287–302. [Google Scholar] [PubMed]
  26. Fersht, A. Structure and Mechanism in Protein Science: A Guide to Enzyme Catalysis and Protein Folding; World Scientific: Cambridge, UK, 2017. [Google Scholar]
  27. Tokuriki, N.; Tawfik, D.S. Stability effects of mutations and protein evolvability. Curr. Opin. Struct. Biol. 2009, 19, 596–604. [Google Scholar] [CrossRef] [PubMed]
  28. Tokuriki, N.; Stricher, F.; Schymkowitz, J.; Serrano, L.; Tawfik, D.S. The stability effects of protein mutations appear to be universally distributed. J. Mol. Biol. 2007, 369, 1318–1332. [Google Scholar] [CrossRef] [PubMed]
  29. Gargiulo, S.; Soumillion, P. Directed evolution for enzyme development in biocatalysis. Curr. Opin. Chem. Biol. 2021, 61, 107–113. [Google Scholar] [CrossRef] [PubMed]
  30. Goldsmith, M.; Tawfik, D.S. Enzyme engineering: Reaching the maximal catalytic efficiency peak. Curr. Opin. Struct. Biol. 2017, 47, 140–150. [Google Scholar] [CrossRef] [PubMed]
  31. Malli, A.; Vasyutyn, D.; Kim, J.R. Advances in machine learning models for predicting enzyme kinetic parameters. J. Chem. Inf. Model. 2026, 66, 42–60. [Google Scholar] [PubMed]
  32. Wang, J.; Zhao, Y.; Yang, Z.; Yao, G.; Han, P.; Liu, J.; Chen, C.; Zan, P.; Wan, X.; Bo, X.; et al. IECata: Interpretable bilinear attention network and evidential deep learning improve catalytic efficiency prediction of enzymes. Brief. Bioinform. 2025, 26, bbaf283. [Google Scholar] [CrossRef] [PubMed]
  33. Rives, A.; Meier, J.; Sercu, T.; Goyal, S.; Lin, Z.; Liu, J.; Guo, D.; Ott, M.; Zitnick, C.L.; Ma, J.; et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl. Acad. Sci. USA 2021, 118, e2016239118. [Google Scholar] [CrossRef] [PubMed]
  34. Kim, P.T.; Winter, R.; Clevert, D.A. Unsupervised representation learning for proteochemometric modeling. Int. J. Mol. Sci. 2021, 22, 12882. [Google Scholar] [CrossRef] [PubMed]
  35. Lin, Z.; Akin, H.; Rao, R.; Hie, B.; Zhu, Z.; Lu, W.; Smetanin, N.; Verkuil, R.; Kabeli, O.; Shmueli, Y.; et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 2023, 379, 1123–1130. [Google Scholar] [CrossRef] [PubMed]
  36. Elnaggar, A.; Heinzinger, M.; Dallago, C.; Rehawi, G.; Wang, Y.; Jones, L.; Gibbs, T.; Feher, T.; Angerer, C.; Steinegger, M.; et al. ProtTrans: Toward understanding the language of life through self-supervised learning. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 7112–7127. [Google Scholar] [PubMed]
  37. Kouba, P.; Kohout, P.; Haddadi, F.; Bushuiev, A.; Samusevich, R.; Sedlar, J.; Damborsky, J.; Pluskal, T.; Sivic, J.; Mazurenko, S. Machine learning-guided protein engineering. ACS Catal. 2023, 13, 13863–13895. [Google Scholar] [CrossRef] [PubMed]
  38. Ali, R.; Zhang, Y. Machine learning meets enzyme engineering: Examples in the design of polyethylene terephthalate hydrolases. Front. Chem. Sci. Eng. 2024, 18, 149. [Google Scholar] [CrossRef]
  39. Yang, J.; Li, F.Z.; Arnold, F.H. Opportunities and challenges for machine learning-assisted enzyme engineering. ACS Cent. Sci. 2024, 10, 226–241. [Google Scholar] [CrossRef] [PubMed]
  40. Freschlin, C.R.; Fahlberg, S.A.; Romero, P.A. Machine learning to navigate fitness landscapes for protein engineering. Curr. Opin. Biotechnol. 2022, 75, 102713. [Google Scholar] [CrossRef] [PubMed]
  41. Repecka, D.; Jauniskis, V.; Karpus, L.; Rembeza, E.; Rokaitis, I.; Zrimec, J.; Poviloniene, S.; Laurynenas, A.; Viknander, S.; Abuajwa, W.; et al. Expanding functional protein sequence spaces using generative adversarial networks. Nat. Mach. Intell. 2021, 3, 324–333. [Google Scholar] [CrossRef]
  42. Lindley, S.E.; Lu, Y.; Shukla, D. The experimentalist’s guide to machine learning for small molecule design. ACS Appl. Bio Mater. 2024, 7, 657–684. [Google Scholar] [PubMed]
  43. Moreno, M.; Cuesta, S.A.; Mora, J.R.; Márquez Brazon, E.A.; Paz, J.L.; Agüero-Chapin, G.; Pérez-Pérez, N.; García-Jacas, C.R. Hybrid computational framework integrating ensemble learning, molecular docking, and dynamics for predicting antimalarial efficacy. Int. J. Mol. Sci. 2026, 27, 1875. [Google Scholar] [CrossRef] [PubMed]
  44. Kumar, C.; Choudhary, A. A top-down approach to classify enzyme functional classes and sub-classes using random forest. J. Bioinform. Sys. Biol. 2012, 1, 2012. [Google Scholar]
  45. Shi, Z.; Xu, S.; Xue, S.; Chen, K.; Lu, Y.; Wang, F.; Long, S.; Tian, Y.; Zhang, P.; Wang, J.; et al. From machine learning to multimodal models: The AI revolution in enzyme engineering. Biodes. Res. 2025, 8, 100044. [Google Scholar] [CrossRef] [PubMed]
  46. Salas-Nuñez, L.F.; Barrera-Ocampo, A.; Caicedo, P.A.; Cortes, N.; Osorio, E.H.; Villegas-Torres, M.F.; González Barrios, A.F. Machine learning to predict enzyme–substrate interactions in elucidation of synthesis pathways: A review. Metabolites 2024, 14, 154. [Google Scholar] [CrossRef] [PubMed]
  47. Vornholt, T.; Stockinger, P.; Mutný, M.; Jeschek, M.; Nestl, B.; Oberdorfer, G.; Osuna, S.; Pleiss, J.; Welner, D.H.; Krause, A.; et al. Of revolutions and roadblocks: The emerging role of machine learning in biocatalysis. ACS Cent. Sci. 2025, 11, 1828–1838. [Google Scholar] [CrossRef] [PubMed]
  48. Moorhoff, F.; Zhang, Y.; Qiu, W.; Dong, W.; Medina-Ortiz, D.; Davari, M.D. Machine learning-driven enzyme mining: Opportunities, challenges, and future perspectives. ACS Catal. 2026, 16, 12–30. [Google Scholar]
  49. Kroll, A.; Ranjan, S.; Engqvist, M.K.M.; Lercher, M.J. A general model to predict small molecule substrates of enzymes based on machine and deep learning. Nat. Commun. 2023, 14, 2787. [Google Scholar] [CrossRef] [PubMed]
  50. Yang, Q.; Yu, J.; Zheng, J. A survey of downstream applications of evolutionary scale modeling protein language models. Quant. Biol. 2025, 14, e70013. [Google Scholar] [CrossRef] [PubMed]
  51. Leclercq, M.; Droit, A. Protein language models: Applications and perspectives. J. Proteome Res. 2026, 25, 507–524. [Google Scholar] [PubMed]
  52. Rajagopal, N.; Choudhary, U.; Tsang, K.; Martin, K.P.; Karadag, M.; Chen, H.T.; Kwon, N.Y.; Mozdzierz, J.; Horspool, A.M.; Li, L.; et al. Deep learning-based design and experimental validation of a medicine-like human antibody library. Brief. Bioinform. 2024, 26, bbaf023. [Google Scholar]
  53. Jiang, Y.; Ran, X.; Yang, Z.J. Data-driven enzyme engineering to identify function-enhancing enzymes. Protein Eng. Des. Sel. 2023, 36, gzac009. [Google Scholar] [PubMed]
  54. Tang, M.; Ge, F.; Li, A.; Hu, L.; Wang, C.; Tang, J.; Song, X.; Liu, X.; Shi, H.; Tan, Z. Artificial intelligence-driven de novo design of robust enzymes to enhance their performance. ACS Synth. Biol. 2025, 14, 4178–4201. [Google Scholar] [CrossRef] [PubMed]
  55. Ahern, W.; Yim, J.; Tischer, D.; Salike, S.; Woodbury, S.M.; Kim, D.; Kalvet, I.; Kipnis, Y.; Coventry, B.; Altae-Tran, H.R.; et al. Atom-level enzyme active site scaffolding using RFdiffusion2. Nat. Methods 2026, 23, 96–105. [Google Scholar] [PubMed]
  56. Sumida, K.H.; Núñez-Franco, R.; Kalvet, I.; Pellock, S.J.; Wicky, B.I.M.; Milles, L.F.; Dauparas, J.; Wang, J.; Kipnis, Y.; Jameson, N.; et al. Improving protein expression, stability, and function with ProteinMPNN. J. Am. Chem. Soc. 2024, 146, 2054–2061. [Google Scholar] [CrossRef] [PubMed]
  57. Chen, A.; Peng, X.; Shen, T.; Zheng, L.; Wu, D.; Wang, S. Discovery, design, and engineering of enzymes based on molecular retrobiosynthesis. mLife 2025, 4, 107–125. [Google Scholar] [CrossRef] [PubMed]
  58. Rao, R.; Bhattacharya, N.; Thomas, N.; Duan, Y.; Chen, X.; Canny, J.; Abbeel, P.; Song, Y.S. Evaluating protein transfer learning with TAPE. Adv. Neural Inf. Process. Syst. 2019, 32, 9689–9701. [Google Scholar] [PubMed]
  59. Dallago, C.; Mou, J.; Johnston, K.E.; Wittmann, B.J.; Bhattacharya, N.; Goldman, S.; Madani, A.; Yang, K.K. FLIP: Benchmark tasks in fitness landscape inference for proteins. bioRxiv 2021, preprint. [Google Scholar]
  60. Notin, P.; Kollasch, A.W.; Ritter, D.; van Niekerk, L.; Paul, S.; Spinner, H.; Rollins, N.; Shaw, A.; Weitzman, R.; Frazer, J.; et al. ProteinGym: Large-scale benchmarks for protein design and fitness prediction. bioRxiv 2023, 36, 64331–64379. [Google Scholar]
  61. Venanzi, N.A.E.; Basciu, A.; Vargiu, A.V.; Kiparissides, A.; Dalby, P.A.; Dikicioglu, D. Machine learning integrating protein structure, sequence, and dynamics to predict enzyme activity. J. Chem. Inf. Model. 2024, 64, 2681–2694. [Google Scholar] [CrossRef]
  62. Dolinska, M.B.; Sergeev, Y.V. Insights from computational dynamic active site mapping into substrate recognition. Int. J. Mol. Sci. 2026, 27, 1937. [Google Scholar] [CrossRef] [PubMed]
  63. Khan, M.F.; Khan, M.T. AI-driven enzyme engineering: Emerging models and next-generation biotechnological applications. Molecules 2025, 31, 45. [Google Scholar] [PubMed]
  64. Vajanapanich, P.; Nearmnala, P.; Parkbhorn, J.; Nutho, B.; Rungrotmongkol, T.; Hongdilokkul, N. Catalytic residue reprogramming enhances enzyme activity at alkaline pH. ACS Synth. Biol. 2025, 14, 3612–3623. [Google Scholar] [CrossRef] [PubMed]
  65. Leidner, F.; Kurt Yilmaz, N.; Schiffer, C.A. Target-specific prediction of ligand affinity with structure-based interaction fingerprints. J. Chem. Inf. Model. 2019, 59, 3679–3691. [Google Scholar] [PubMed]
  66. Xu, W.; Li, A.; Zhao, Y.; Peng, Y. Decoding the effects of mutation on protein interactions using machine learning. Biophys. Rev. 2025, 6, 011307. [Google Scholar] [CrossRef]
  67. Deshpande, A.; Ouldridge, T.E. Optimizing enzymatic catalysts for rapid turnover of substrates with low enzyme sequestration. Biol. Cybern. 2020, 114, 653–668. [Google Scholar] [CrossRef] [PubMed]
  68. Li, F.Z.; Yang, J.; Johnston, K.E.; Gürsoy, E.; Yue, Y.; Arnold, F.H. Evaluation of machine learning-assisted directed evolution across diverse combinatorial landscapes. Cell Syst. 2025, 16, 101387. [Google Scholar] [CrossRef] [PubMed]
  69. Thomas, N.; Belanger, D.; Xu, C.; Lee, H.; Hirano, K.; Iwai, K.; Polic, V.; Nyberg, K.D.; Hoff, K.G.; Frenz, L.; et al. Engineering highly active nuclease enzymes with machine learning and high-throughput screening. Cell Syst. 2025, 16, 101236. [Google Scholar] [CrossRef] [PubMed]
  70. Poelwijk, F.J.; Krishna, V.; Ranganathan, R. The context-dependence of mutations: A linkage of formalisms. PLoS Comput. Biol. 2016, 12, e1004771. [Google Scholar] [CrossRef] [PubMed]
  71. Starr, T.N.; Thornton, J.W. Epistasis in protein evolution. Protein Sci. 2016, 25, 1204–1218. [Google Scholar] [CrossRef] [PubMed]
  72. Rao, R.; Meier, J.; Sercu, T.; Ovchinnikov, S.; Rives, A. Transformer protein language models are unsupervised structure learners. bioRxiv 2020. preprint. [Google Scholar] [CrossRef]
  73. Russ, W.P.; Figliuzzi, M.; Stocker, C.; Barrat-Charlaix, P.; Socolich, M.; Kast, P.; Hilvert, D.; Monasson, R.; Cocco, S.; Weigt, M.; et al. An evolution-based model for designing chorismate mutase enzymes. Science 2020, 369, 440–445. [Google Scholar] [CrossRef] [PubMed]
  74. Wittmann, B.J.; Yue, Y.; Arnold, F.H. Informed training set design enables efficient machine learning-assisted directed protein evolution. Cell Syst. 2021, 12, 1026–1045.e7. [Google Scholar] [CrossRef] [PubMed]
  75. Hsu, C.; Nisonoff, H.; Fannjiang, C.; Listgarten, J. Learning protein fitness models from evolutionary and assay-labeled data. Nat. Biotechnol. 2022, 40, 1114–1122. [Google Scholar] [CrossRef] [PubMed]
  76. Riesselman, A.J.; Ingraham, J.B.; Marks, D.S. Deep generative models of genetic variation capture the effects of mutations. Nat. Methods 2018, 15, 816–822. [Google Scholar] [CrossRef] [PubMed]
  77. Song, Z.; Trozzi, F.; Tian, H.; Yin, C.; Tao, P. Mechanistic insights into enzyme catalysis from explaining machine-learned quantum mechanical and molecular mechanical minimum energy pathways. ACS Phys. Chem. Au 2022, 2, 316–330. [Google Scholar] [CrossRef] [PubMed]
  78. Leander, M.; Liu, Z.; Cui, Q.; Raman, S. Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots. eLife 2022, 11, e79932. [Google Scholar] [CrossRef] [PubMed]
  79. Petrović, D.; Risso, V.A.; Kamerlin, S.C.L.; Sanchez-Ruiz, J.M. Conformational dynamics and enzyme evolution. J. R. Soc. Interface 2018, 15, 20180330. [Google Scholar] [CrossRef] [PubMed]
  80. Gradisteanu, V.; Chan, E.W.; Hedges, L.; Malagarriga, M.; David, R.; de la Puente, M.; Laage, D.; Tuñón, I.; van der Kamp, M.W.; Zinovjev, K. Simulating enzyme catalysis with electrostatically embedded machine learning potentials. Chem. Sci. 2026, in press. [Google Scholar] [CrossRef] [PubMed]
  81. Li, G.C.; Srivastava, A.K.; Kim, J.; Taylor, S.S.; Veglia, G. Mapping the hydrogen bond networks in the catalytic subunit of protein kinase A. Biochemistry 2015, 54, 4042–4049. [Google Scholar] [CrossRef] [PubMed]
  82. Heckmann, D.; Lloyd, C.J.; Mih, N.; Ha, Y.; Zielinski, D.C.; Haiman, Z.B.; Desouki, A.A.; Lercher, M.J.; Palsson, B.O. Machine learning applied to enzyme turnover numbers reveals protein structural correlates. Nat. Commun. 2018, 9, 5252. [Google Scholar] [CrossRef] [PubMed]
  83. Xia, W.; Bai, Y.; Shi, P. Improving substrate affinity and catalytic efficiency of β-glucosidase by rational design. Biomolecules 2021, 11, 1882. [Google Scholar] [PubMed]
  84. Li, F.; Yuan, L.; Lu, H.; Li, G.; Chen, Y.; Engqvist, M.K.M.; Kerkhoven, E.J.; Nielsen, J. Deep learning-based kcat prediction enables improved enzyme-constrained model reconstruction. Nat. Catal. 2022, 5, 662–672. [Google Scholar]
  85. Tripathi, N.; Herisson, J.; Faulon, J.-L. Machine learning in predictive biocatalysis: A comparative review. Biotechnol. Adv. 2025, 84, 108698. [Google Scholar] [CrossRef] [PubMed]
  86. Li, Y.; Song, K.; Zhang, J.; Lu, S. A computational method to predict effects of residue mutations on catalytic efficiency. Catalysts 2021, 11, 286. [Google Scholar] [CrossRef]
  87. Landwehr, G.M.; Bogart, J.W.; Magalhaes, C.; Hammarlund, E.G.; Karim, A.S.; Jewett, M.C. Accelerated enzyme engineering by machine-learning-guided cell-free expression. Nat. Commun. 2025, 16, 865. [Google Scholar] [PubMed]
  88. Wu, Z.; Kan, S.B.J.; Lewis, R.D.; Wittmann, B.J.; Arnold, F.H. Machine learning-assisted directed protein evolution with combinatorial libraries. Proc. Natl. Acad. Sci. USA 2019, 116, 8852–8858. [Google Scholar] [CrossRef] [PubMed]
  89. Menke, M.J.; Ao, Y.-F.; Bornscheuer, U.T. Practical machine learning-assisted design protocol for protein engineering: Transaminase engineering for the conversion of bulky substrates. ACS Catal. 2024, 14, 6462–6469. [Google Scholar] [CrossRef]
  90. Ding, K.; Chin, M.; Zhao, Y.; Huang, W.; Mai, B.K.; Wang, H.; Liu, P.; Yang, Y.; Luo, Y. Machine learning-guided co-optimization of fitness and diversity facilitates combinatorial library design in enzyme engineering. Nat. Commun. 2024, 15, 6392. [Google Scholar] [PubMed]
  91. Saito, Y.; Oikawa, M.; Sato, T.; Nakazawa, H.; Ito, T.; Kameda, T.; Tsuda, K.; Umetsu, M. Machine-learning-guided library design cycle for directed evolution of enzymes: The effects of training data composition on sequence space exploration. ACS Catal. 2021, 11, 14615–14624. [Google Scholar]
  92. Trivedi, V.D.; Chappell, T.C.; Krishna, N.B.; Shetty, A.; Sigamani, G.G.; Mohan, K.; Ramesh, A.; Pravin, K.R.; Nair, N.U. In-depth sequence-function characterization reveals multiple pathways to enhance enzymatic activity. ACS Catal. 2022, 12, 2381–2396. [Google Scholar] [PubMed]
  93. Marchal, D.G.; Schulz, L.; Schuster, I.; Ivanovska, J.; Paczia, N.; Prinz, S.; Zarzycki, J.; Erb, T.J. Machine learning-supported enzyme engineering toward improved CO2-fixation of glycolyl-CoA carboxylase. ACS Synth. Biol. 2023, 12, 3521–3530. [Google Scholar] [PubMed]
  94. Thornton, E.L.; Boyle, J.T.; Laohakunakorn, N.; Regan, L. Cell-free protein synthesis as a method to rapidly screen machine learning-generated protease variants. ACS Synth. Biol. 2025, 14, 1710–1718. [Google Scholar] [PubMed]
  95. Erkanli, M.E.; Jang, Y.; Malli, A.; El-Halabi, K.; Ryu, C.; Kim, J.R. Machine learning framework for kcat/Km prediction in β-glucosidases. ACS Synth. Biol. 2025, 14, 3927–3939. [Google Scholar] [CrossRef] [PubMed]
  96. Jurich, C.; Shao, Q.; Ran, X.; Yang, Z.J. Physics-based modeling in the new era of enzyme engineering. Nat. Comput. Sci. 2025, 5, 279–291. [Google Scholar] [CrossRef] [PubMed]
  97. Manan, A.; Qayyum, N.; Ramachandran, R.; Qayyum, N.; Ilyas, S. Digital to biological translation: Algorithmic data-driven design in synthetic biology. SynBio 2025, 3, 17. [Google Scholar]
  98. Lu, X.; Cao, M.; Ma, M.; Wu, Y.; Qu, M.; Du, F.; Ji, R.; Duan, M.; Dong, L.; Liu, K.; et al. Accelerating enzyme engineering with artificial intelligence in biocatalysis. Food Bioeng. 2025, 4, 589–611. [Google Scholar] [CrossRef]
  99. Kitano, S.; Lin, C.; Foo, J.L.; Chang, M.W. Synthetic biology: Learning toward high-precision biological design. PLoS Biol. 2023, 21, e3002116. [Google Scholar] [PubMed]
  100. Vanella, R.; Kovacevic, G.; Doffini, V.; Fernández de Santaella, J.; Nash, M.A. High-throughput screening and machine learning in enzyme engineering. Chem. Commun. 2022, 58, 2455–2467. [Google Scholar]
  101. Callaway, E. Will self-driving robot labs replace biologists? Nature 2026, 650, 809–810. [Google Scholar] [CrossRef] [PubMed]
  102. Clark-ElSayed, A.; Harrison, I.M.; Olsen, M.L.; Lazar, J.T.; Jewett, M.C.; Ellington, A.D. LDBT instead of DBTL: Combining machine learning and cell-free testing. Nat. Commun. 2025, 16, 9782. [Google Scholar] [PubMed]
  103. Zhang, Q.; Chen, W.; Qin, M.; Wang, Y.; Pu, Z.; Ding, K.; Liu, Y.; Zhang, Q.; Li, D.; Li, X.; et al. Integrating protein language models and biofoundry for enhanced protein evolution. Nat. Commun. 2025, 16, 1553. [Google Scholar] [CrossRef] [PubMed]
  104. Hägele, L.; Trachtmann, N.; Takors, R. Knowledge-driven DBTL cycle provides mechanistic insights. Microb. Cell Factories 2025, 24, 111. [Google Scholar] [PubMed]
  105. Jones, B.S.; Soler, J.; Sharratt, J.W.; Hogg, B.N.; Tavanti, M.; Schnepel, C.; Kress, N.; Seibt, L.S.; Osuna, S.; Garcia-Borràs, M.; et al. Mechanistic insight-guided engineering of cytochrome P450 regioselectivity. ACS Catal. 2026, 16, 6673–6684. [Google Scholar]
  106. Zhai, J.; Qi, X.; Cai, L.; Liu, Y.; Tang, H.; Xie, L.; Wang, J. NNKcat: Deep neural network to predict catalytic constants. Brief. Bioinform. 2025, 26, bbaf212. [Google Scholar] [CrossRef] [PubMed]
  107. Markus, B.; Christian, C.G.; Andreas, K.; Arkadij, K.; Stefan, L.; Gustav, O.; Elina, S.; Radka, S. Accelerating biocatalysis discovery with machine learning: A paradigm shift in enzyme engineering, discovery, and design. ACS Catal. 2023, 13, 14454–14469. [Google Scholar] [CrossRef] [PubMed]
  108. Al-Madboly, L.A.; Aboulmagd, A.; El-Salam, M.A.; Kushkevych, I.; El-Morsi, R.M. Microbial enzymes as natural anti-biofilm candidates. Microb. Cell Factories 2024, 23, 343. [Google Scholar] [PubMed]
  109. Efremenko, E.; Stepanov, N.; Aslanli, A.; Lyagin, I.; Senko, O.; Maslova, O. Enzyme-based antimicrobial materials: Trends and perspectives. J. Funct. Biomater. 2023, 14, 64. [Google Scholar] [PubMed]
  110. de Regil, R.; Sandoval, G. Biocatalysis for biobased chemicals. Biomolecules 2013, 3, 812–847. [Google Scholar] [CrossRef] [PubMed]
  111. Mao, S.; Jiang, J.; Xiong, K.; Chen, Y.; Yao, Y.; Liu, L.; Liu, H.; Li, X. Enzyme engineering for food industry applications. Foods 2024, 13, 3846. [Google Scholar] [PubMed]
  112. Farhan, M.; Hasani, I.W.; Khafaga, D.S.R.; Ragab, W.M.; Ahmed Kazi, R.N.; Aatif, M.; Muteeb, G.; Fahim, Y.A. Enzymes as catalysts in industrial biocatalysis. Catalysts 2025, 15, 891. [Google Scholar] [CrossRef]
  113. Ndochinwa, O.G.; Wang, Q.Y.; Amadi, O.C.; Nwagu, T.N.; Nnamchi, C.I.; Okeke, E.S.; Moneke, A.N. Current status in enzyme engineering: Industrial perspective. Heliyon 2024, 10, e32673. [Google Scholar] [CrossRef] [PubMed]
  114. Deivayanai, V.C.; Karishma, S.; Thamarai, P.; Kamalesh, R.; Saravanan, A.; Yaashikaa, P.R.; Vickram, A.S. Plastic remediation using catalytic and ML approaches. J. Contam. Hydrol. 2024, 267, 104449. [Google Scholar] [PubMed]
  115. Gupta, G.K.; Dixit, M.; Chot, E.; Shukla, P. Microbial enzymatic biodegradation of plastics. ACS Environ. Au 2025, 5, 520–542. [Google Scholar] [CrossRef] [PubMed]
  116. Sonowal, K.; Borthakur, P.P.; Pathak, K. Advances in enzyme-based biosensors. Eng. Proc. 2025, 106, 5. [Google Scholar] [CrossRef]
  117. Ji, Z.L.; Chen, X.; Zhen, C.J.; Yao, L.X.; Han, L.Y.; Yeo, W.K.; Chung, P.C.; Puy, H.S.; Tay, Y.T.; Muhammad, A.; et al. KDBI: Kinetic data of biomolecular interactions database. Nucleic Acids Res. 2003, 31, 255–257. [Google Scholar] [CrossRef] [PubMed]
  118. Shah, A.; Bi, F.; Yang, J. Machine learning in drug–food interaction prediction. J. Cheminform. 2025, 18, 8. [Google Scholar] [PubMed]
  119. Doshi, U.; McGowan, L.C.; Ladani, S.T.; Hamelberg, D. Role of enzyme conformational dynamics in catalysis. Proc. Natl. Acad. Sci. USA 2012, 109, 5699–5704. [Google Scholar] [PubMed]
  120. Cui, X.; Ge, L.; Chen, X.; Lv, Z.; Wang, S.; Zhou, X.; Zhang, G. Protein dynamics modeling in the post-AlphaFold era. Brief. Bioinform. 2025, 26, bbaf340. [Google Scholar] [PubMed]
  121. Niazi, S.K. Protein catalysis through structural dynamics. Pharmaceuticals 2025, 18, 951. [Google Scholar] [CrossRef] [PubMed]
  122. Dieckhaus, H.; Kuhlman, B. Protein stability models fail to capture epistatic interactions. Protein Sci. 2025, 34, e70003. [Google Scholar] [PubMed]
  123. Taylor, M.B.; Ehrenreich, I.M. Higher-order genetic interactions in complex traits. Trends Genet. 2015, 31, 34–40. [Google Scholar] [PubMed]
  124. Sidak, D.; Schwarzerová, J.; Weckwerth, W.; Waldherr, S. Interpretable machine learning in systems biology. Front. Mol. Biosci. 2022, 9, 926623. [Google Scholar] [PubMed]
  125. Shi, H.; Bai, X.; Tian, F.; Li, Y.; Li, D.; Yao, L.; Xue, C.; Tang, C. AI-driven enzyme engineering from structure prediction to de novo design. J. Agric. Food Chem. 2026, 74, 9975–9990. [Google Scholar] [PubMed]
  126. Yu, H.; Deng, H.; He, J.; Keasling, J.D.; Luo, X. UniKP: Prediction of enzyme kinetic parameters. Nat. Commun. 2023, 14, 8211. [Google Scholar] [PubMed]
  127. Medina-Ortiz, D.; Khalifeh, A.; Anvari-Kazemabad, H.; Davari, M.D. Explainable ML models for protein engineering. Biotechnol. Adv. 2025, 79, 108495. [Google Scholar] [PubMed]
  128. Dritsas, E.; Trigka, M. Machine learning and big data: A survey. Mach. Learn. Knowl. Extr. 2025, 7, 13. [Google Scholar] [CrossRef]
  129. Le Piane, F.; Vozza, M.; Baldoni, M.; Mercuri, F. ML and HPC integration for nanomaterials. Beilstein J. Nanotechnol. 2024, 15, 1498–1521. [Google Scholar] [PubMed]
  130. Davoudi, S.; Henry, C.S.; Miller, C.S.; Banaei-Kashani, F. EC-Bench: Enzyme classification benchmark. Bioinform. Adv. 2026, 6, vbag004. [Google Scholar] [PubMed]
  131. Nair, M.; Svedberg, P.; Larsson, I.; Nygren, J.M. Barriers to AI implementation in healthcare. PLoS ONE 2024, 19, e0305949. [Google Scholar] [PubMed]
  132. Shao, Q.; Hollenbeak, A.C.; Jiang, Y.; Ran, X.; Bachmann, B.O.; Yang, Z.J. SubTuner leverages physics-based modeling to complement AI in enzyme engineering toward nan-native substrates. Chem. Catal. 2025, 5, 101334. [Google Scholar] [PubMed]
  133. Wen, S.; Zheng, W.; Bornscheuer, U.T.; Wu, S. Generative AI for enzyme design. Curr. Opin. Green Sustain. Chem. 2025, 52, 101010. [Google Scholar]
  134. Xie, W.J.; Warshel, A. Generative AI for enzyme catalysis and evolution. Natl. Sci. Rev. 2023, 10, nwad331. [Google Scholar] [CrossRef] [PubMed]
  135. Sommer, L.M.; Groves, T.; Santos, A. Data infrastructure for autonomous laboratories. Curr. Opin. Biotechnol. 2026, 97, 103434. [Google Scholar] [PubMed]
  136. Singh, N.; Lane, S.; Yu, T.; Lu, J.; Ramos, A.; Cui, H.; Zhao, H. AI-powered autonomous enzyme engineering. Nat. Commun. 2025, 16, 5648. [Google Scholar] [PubMed]
  137. Agrawal, R.; Gupta, T.; Gupta, S.; Chauhan, S.; Patel, P.; Hamdare, S. Explainable AI for decision transparency. Diagn. Pathol. 2025, 20, 105. [Google Scholar] [PubMed]
  138. Feehan, R.; Montezano, D.; Slusky, J.S.G. Machine learning for enzyme engineering, selection and design. Protein Eng. Des. Sel. 2021, 34, gzab019. [Google Scholar] [PubMed]
  139. Helmy, M.; Smith, D.; Selvarajoo, K. Systems biology and AI in metabolic engineering. Metab. Eng. Commun. 2020, 11, e00149. [Google Scholar] [PubMed]
  140. Lawson, C.E.; Martí, J.M.; Radivojevic, T.; Jonnalagadda, S.V.R.; Gentz, R.; Hillson, N.J.; Peisert, S.; Kim, J.; Simmons, B.A.; Petzold, C.J.; et al. Machine learning for metabolic engineering. Metab. Eng. 2021, 63, 34–60. [Google Scholar] [CrossRef] [PubMed]
  141. Cheng, Y.; Bi, X.; Xu, Y.; Liu, Y.; Li, J.; Du, G.; Lv, X.; Liu, L. ML for metabolic pathway optimization. Comput. Struct. Biotechnol. J. 2023, 21, 2381–2393. [Google Scholar] [PubMed]
  142. Ho, D.; Quake, S.R.; McCabe, E.R.B.; Chng, W.J.; Chow, E.K.; Ding, X.; Gelb, B.D.; Ginsburg, G.S.; Hassenstab, J.; Ho, C.M.; et al. Enabling technologies for personalized medicine. Trends Biotechnol. 2020, 38, 497–518. [Google Scholar] [PubMed]
Figure 1. The four major factors governing enzyme efficiency collectively influence the catalytic parameters and provide mechanistic targets for enzyme engineering.
Figure 1. The four major factors governing enzyme efficiency collectively influence the catalytic parameters and provide mechanistic targets for enzyme engineering.
Catalysts 16 00598 g001
Figure 2. Flowchart illustrating practical model selection strategies for ML-guided enzyme engineering, based on dataset size, prediction bias, and the availability of structural information.
Figure 2. Flowchart illustrating practical model selection strategies for ML-guided enzyme engineering, based on dataset size, prediction bias, and the availability of structural information.
Catalysts 16 00598 g002
Figure 3. Impact of ML on enzyme kinetic parameters and the mechanism behind improvement. ML-driven protein engineering improves catalytic efficiency by optimizing turnover, substrate affinity, and overall performance as reported in experimentally validated ML-assisted enzyme engineering studies. The fold-change ranges shown are approximate values compiled from studies summarized in Table 2 and should be interpreted as representative ranges rather than universal performance limits.
Figure 3. Impact of ML on enzyme kinetic parameters and the mechanism behind improvement. ML-driven protein engineering improves catalytic efficiency by optimizing turnover, substrate affinity, and overall performance as reported in experimentally validated ML-assisted enzyme engineering studies. The fold-change ranges shown are approximate values compiled from studies summarized in Table 2 and should be interpreted as representative ranges rather than universal performance limits.
Catalysts 16 00598 g003
Figure 4. Steps in the ML-guided DBTL cycle for enzyme engineering, showing data-driven design, iterative experimentation, and continuous learning to accelerate enzyme optimization.
Figure 4. Steps in the ML-guided DBTL cycle for enzyme engineering, showing data-driven design, iterative experimentation, and continuous learning to accelerate enzyme optimization.
Catalysts 16 00598 g004
Figure 5. Advantages of ML-guided DBTL workflows over the traditional cycle.
Figure 5. Advantages of ML-guided DBTL workflows over the traditional cycle.
Catalysts 16 00598 g005
Figure 6. Quantitative overview of ML-assisted enzyme engineering strategies, data performance gains, and validation workflows. The workflow panel shows a typical sequence of data acquisition, model training, variant prioritization, experimental testing, and model refinement. The circular chart summarizes the distribution of targeted optimization parameters across representative studies, while the fold-improvement plot shows reported performance gains compiled from the experimentally validated studies as listed in Table 2. The key advantages of ML-assisted enzyme engineering, including reduced screening burden, improved variant prioritization, and accelerated exploration of sequence space, are shown in the callout panel.
Figure 6. Quantitative overview of ML-assisted enzyme engineering strategies, data performance gains, and validation workflows. The workflow panel shows a typical sequence of data acquisition, model training, variant prioritization, experimental testing, and model refinement. The circular chart summarizes the distribution of targeted optimization parameters across representative studies, while the fold-improvement plot shows reported performance gains compiled from the experimentally validated studies as listed in Table 2. The key advantages of ML-assisted enzyme engineering, including reduced screening burden, improved variant prioritization, and accelerated exploration of sequence space, are shown in the callout panel.
Catalysts 16 00598 g006
Figure 7. Critical challenges and limitations of ML-guided enzyme engineering approaches.
Figure 7. Critical challenges and limitations of ML-guided enzyme engineering approaches.
Catalysts 16 00598 g007
Figure 8. Key emerging directions in ML-guided enzyme engineering strategies.
Figure 8. Key emerging directions in ML-guided enzyme engineering strategies.
Catalysts 16 00598 g008
Table 1. Comparison of ML frameworks used in enzyme engineering.
Table 1. Comparison of ML frameworks used in enzyme engineering.
ML ModelData SourceDataset RequiredStrengthsLimitationsApplications in Enzyme EngineeringReferences
RFSequence features; physicochemical descriptorsSmall-mediumLow computational cost; interpretable; robustUnable to capture non-linear interactionsEnzyme activity prediction; stability screening[31,44]
SVMEngineered featuresSmall-mediumEffective for classification tasksSensitive to feature selection onlySubstrate specificity prediction[45,46]
GBMStructured featuresMediumHigh predictive accuracyRequires feature engineeringPrediction of kinetic parameters[31,47]
CNNSequence/structure gridsMedium-largeCaptures local motifsLimited long-range interactionsMotif detection; enzyme activity prediction[47]
GNNProtein structure graphsMedium-largeModels spatial relationshipsRequires structural dataStructure–function mapping[48]
PLMsProtein sequencesLargeCaptures long-range dependenciesHigh computational costMutation effect prediction; embedding generation[49]
EVE/ESM-1vSequence + Evolutionary dataNo labeled data requiredUseful for low-data availabilityDependent on sequence homologyMutation prioritization[50,51]
VAE/GANSequence DataLargeGenerates novel sequencesExperimental validation is requiredSequence exploration; diversity generation[52,53]
DiffusionStructural dataLargeGenerates high-quality structure-guided designComputationally intensiveDe novo enzyme design[54,55]
ProteinMPNN/ESM-IFProtein structureMedium-largeGenerates structure-based sequence designRequires accurate structureStructure-guided enzyme engineering[56,57]
Table 2. Quantitative benchmarking of experimentally validated ML-assisted enzyme engineering studies.
Table 2. Quantitative benchmarking of experimentally validated ML-assisted enzyme engineering studies.
Enzyme System/StudyEnzyme Class/Protein TypeML Model/StrategyDataset Size/Training DataBaseline ComparatorParameter AffectedReported ImprovementReferences
Amide synthetaseLigaseAugmented ridge regressionHT screening data across multiple substrates (~103 variants)Parent enzymeEnzymatic activity/Substrate preference1.6- to 42-fold improvement in activity[87]
Cytochrome P450 for carbene Si-H insertionOxidoreductase/P450ML-assisted directed evolution using combinatorial librariesCombinatorial library data (~104 variants)Conventional directed evolution/previous variantsEnantioselectivity (ee)Variants with 93% and 79% enantiomeric excess (ee) for stereodivergent catalysis[88]
Transaminase for bulky substratesTransferase/transaminasePractical ML-assisted variant design protocolSequence/structure-guided dataset (~103 variants)Starting transaminase variantEnzymatic activity/Substrate conversionUp to 3-fold improved conversion of bulky substrates and up to > 99% improved ee.[89]
Combinatorial enzyme librariesMultipleML-assisted optimization of fitness and diversityCombinatorial fitness/diversity datasets (~103 variants)Fitness-only or diversity-free selectionFitness/DiversityImproved library enrichment and diversity[90]
Sortase ATranspeptidaseML-guided iterative library designTraining datasets with different compositions (~103–104 variants)Non-ML library designEnzymatic activity2.2- to 2.5-fold improved enzyme activity; Improved sequence space exploration[91]
Phenylalanine ammonia-lyase (PAL)LyaseSequence–function modeling; Deep mutational scanning (DMS)Deep mutational scanning datasets (~105 variants)Parent/lower-activity variantsEnzymatic activity112 mutations at 79 functionally relevant sites were revealed[92]
Glycolyl-CoA carboxylase (GCC)Carboxylase/CO2-fixation enzymeML-supported enzyme engineeringMutational and activity data (~103 variants)Wild-type/starting enzymeCO2-fixation efficiency~2-fold increased CO2-fixation performance[93]
KetoreductasesOxidoreductase/ketoreductaseDirected evolution + MLDirected evolution activity/selectivity data (~105 variants)Parent ketoreductaseDiastereoselectivity/YieldIsolated yield of 40.7% and an enhanced diastereoselectivity of 91.3%[94]
Nuclease enzymesNucleaseML + high-throughput screening (HTS) TeleProtHTS dataset (~104 variants)Starting nuclease/screening baselineSpecific activityHighly active nuclease variants identified with 11-fold improved specific activity[95]
Protease variantsProteasesML-generated variants + cell-free screeningML-generated library (~103 sequences)Parent/baseline proteaseKinetic properties/ActivityRapid identification of functional variants; 4-fold improvement in kinetic properties[96]
Table 3. Representative ML-guided enzyme engineering workflows and DBTL integration strategies.
Table 3. Representative ML-guided enzyme engineering workflows and DBTL integration strategies.
Enzyme/SystemObjectiveExperimental ApproachLibrary/Variant StrategyVariants Screened/TestedValidation MethodKey Workflow ContributionReferences
Amide synthetaseTo improve the catalytic activity and substrate scopeFocused mutant library + HT screeningComputational prioritization of sequence variants~103 experimentally screened variants across multiple substratesHT activity assays followed by kinetic characterizationEfficient navigation of sequence space[87]
TransaminaseTo enhance catalytic efficiency under neutral conditionsTargeted mutagenesis + kinetic evaluationFocused combinatorial mutagenesisFocused library (~102 variants)Enzyme kinetics and substrate conversion assaysCombined sequence and structural descriptors for mutation prioritization[89]
Cytochrome P450To improve activity and regioselectivityIterative mutation + feedback learningIteratively refined focused librariesMultiple iterative focused libraries (~102–103 variants per cycleRegioselectivity and catalytic activity assaysAccelerated convergence toward optimized variants[9,105,106]
CO2 fixation enzymeTo improve pathway-related catalytic performanceRational mutagenesis guided by computational predictionPrioritized mutation combinationsTargeted variant subsets (~102 variants)Enzymatic pathway assaysApplication of ML in metabolic pathway optimization[93]
Nuclease enzymeTo identify high-activity variantsParallelized experimental testing with iterative model refinementComputationally prioritized screening setsLarge HTS datasets (~104 variants)Activity-based screening assaysCoupling with automated HT workflows[69]
PET hydrolaseTo enhance plastic degradation capabilityStructure-guided engineering workflowRationally prioritized mutation setsTargeted variant libraries (~102–103 variants)PET degradation assaysIntegration of ML with sustainability-focused enzyme engineering[38]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ahsan, W. Machine Learning-Guided Enzyme Engineering Approaches for Enhanced Biocatalytic Efficiency: Concepts, Mechanisms, and Future Directions. Catalysts 2026, 16, 598. https://doi.org/10.3390/catal16070598

AMA Style

Ahsan W. Machine Learning-Guided Enzyme Engineering Approaches for Enhanced Biocatalytic Efficiency: Concepts, Mechanisms, and Future Directions. Catalysts. 2026; 16(7):598. https://doi.org/10.3390/catal16070598

Chicago/Turabian Style

Ahsan, Waquar. 2026. "Machine Learning-Guided Enzyme Engineering Approaches for Enhanced Biocatalytic Efficiency: Concepts, Mechanisms, and Future Directions" Catalysts 16, no. 7: 598. https://doi.org/10.3390/catal16070598

APA Style

Ahsan, W. (2026). Machine Learning-Guided Enzyme Engineering Approaches for Enhanced Biocatalytic Efficiency: Concepts, Mechanisms, and Future Directions. Catalysts, 16(7), 598. https://doi.org/10.3390/catal16070598

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop