Next Article in Journal
Bioactive Compounds from Seaweeds as Natural Preservatives for Seafood
Previous Article in Journal
The Influence of the Ozonation Process on the Quality Parameters and Physicochemical Stability of Horse Meat
Previous Article in Special Issue
Galloylation-Driven Anchoring of the Asp325-Asp336 Ridge: The Molecular Logic Behind the Superior Kinetic Stabilization of HMPV Fusion Protein by Green Tea Dimeric Catechins
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

A Review on Applications of Artificial Intelligence (AI) in Monoclonal Antibody (mAb) Manufacturing

by
Fawad Abidi
and
Dimitrios I. Gerogiorgis
*
Institute for Materials and Processes (IMP), School of Engineering, University of Edinburgh, Edinburgh EH9 3FB, UK
*
Author to whom correspondence should be addressed.
Molecules 2026, 31(18), 3313; https://doi.org/10.3390/molecules31183313 (registering DOI)
Submission received: 5 August 2025 / Revised: 17 April 2026 / Accepted: 20 April 2026 / Published: 18 September 2026
(This article belongs to the Special Issue Development of Computational Approaches in Chemical Biology)

Abstract

This review paper offers a comprehensive exploration of the many Artificial Intelligence (AI) applications in monoclonal antibody (mAb) manufacturing, covering a wide spectrum of topics, including upstream and downstream processes, process control, optimisation, product property monitoring and regulatory compliance. We present several case studies showcasing the predictive modelling capabilities of AI and ML algorithms for cell culture optimisation, media formulation, and bioreactor control, highlighting their potential to enhance cell growth, productivity, and product quality. Furthermore, we explore its potential in process monitoring, fault detection, and real-time decision-making, towards improved process robustness and reduced production costs.

1. Introduction

Artificial Intelligence (AI) shows exceptional potential to revolutionise many aspects of monoclonal antibody (mAb) manufacturing, encompassing many application areas toward improved efficiency, cost-effectiveness, quality control and process optimisation [1]. Because AI can rapidly analyse large datasets collected during manufacturing processes and identify patterns or correlations that humans may overlook, it can optimise operating process parameters (temperature, pH, stirring) for higher mAb yield and purity.
AI algorithms can predict the outcome of different process variables and identify critical parameters affecting mAb production. By leveraging historical data, AI can provide insights into the likelihood of process failure or suboptimal performance, enabling manufacturers to take proactive measures. Quality control is another important area that can benefit from AI, by analysing real-time data from key sources (e.g., in-line sensors), to detect deviations or anomalies in the manufacturing process, provide early warnings of potential quality issues, enabling timely interventions for product quality [2].
AI implementations allow for tracking multiple process parameters simultaneously, supporting the vision of real-time feedback control signal provision to plant operators. Automatic adjustment of process parameters can occur on the basis of predefined criteria, ensuring consistent and reproducible mAb production. AI can guide decision-making by analysing multidimensional legacy datasets, considering a multitude of pivotal factors (e.g., product specifications, manufacturing costs, regulatory frames), assisting technical managers to decide on process modifications, raw materials, and production scheduling.
Beyond biopharma process operations, drug formulation optimisation can also be performed via AI/ML algorithms, via analysing in-house data to identify optimal excipient choices for maximum mAb stability, solubility, product performance, and shelf life. Moreover, AI can analyse sensor data and legacy records to predict equipment failures and determine preventive maintenance needs, whose proactive implementation can thus reduce OpEx, minimise plant downtime and ensure continuous mAb production.

1.1. AI Methodologies

AI can be classified into distinct categories based on the machine’s capacity to use previous data/experiences to predict future outcomes/decisions: the most frequently used AI tools include machine learning (ML), artificial neural networks (ANNs), machine vision (MV) and deep learning (DL). All these AI methodologies have been applied to tackle demanding medicine and biopharma industry problems in the past two decades.
ML is one of the AI tools where machines are not explicitly programmed to perform specific tasks; instead, they automatically learn and improve from experience. The algorithms use computational methods to directly learn from data. There are various machine learning algorithms, e.g., unsupervised, supervised, and reinforcement learning.

1.1.1. Supervised Learning (SL)

Supervised learning is useful when data relevant to a desired prediction is available: therein, previous input-output data is used to predict new outputs based on new inputs. The SL model can make predictions based on evidence, and in the presence of uncertainty; SL algorithms receive a known input dataset and known resulting (output) responses, and then train a model to generate reasonable predictions of output responses for new inputs.
Support Vector Machines (SVM) is an exemplary SL algorithm, offering an exciting approach for model development, which can also be used for estimation and prediction.

1.1.2. Unsupervised Learning (UL)

Unsupervised learning targets the discovery of patterns, structures, or relationships in unlabeled, unstructured data without predefined target outputs. Widespread methods comprise clustering (grouping similar data) and dimensionality reduction (simplifying data), often used for e.g., customer segmentation, anomaly detection, and data visualisation.

1.1.3. Reinforcement Learning (RL)

In reinforcement learning, the machine is instructed to learn by engagement, hence providing better, more reliable results as more experience is gained and accumulated. Thus, RL can maximise a cumulative reward by recording actions via a trial-and-error approach in a set environment. A great RL advantage is that it can be applied with little or even no historical data: unlike other ML methods, it does not require prior information.

1.1.4. Artificial Neural Networks (ANNs) and Deep Learning (DL)

ANNs aim to capture nonlinear patterns in data via structured layers of parameters. Every ANN has inputs, a hidden layer with parameters, and one or more output layers. Deep learning denotes an ANN with hidden layers, encompassing various architectures; DL techniques require large datasets and great computing power (many GPUs) for best performance, since they tune many parameters postulated within elaborate frameworks.
The focus of this review paper is a detailed overview of ANN case studies and their key benefits for mAb manufacturing; we note that such implementations require careful validation, regulatory compliance, and integration with existing processes/workflows. Harnessing AI power thus offers higher productivity, great mAb quality and cost savings.

2. Model Development

Bioprocess modelling plays an important role in biotherapeutics manufacturing, its goal being to estimate optimal production conditions, reduce delays and minimise risks. Bioprocessing modelling tools are broadly classified under: (i) Mechanistic modelling, and (ii) Empirical modelling (Figure 1). Mechanistic models are maps with rigorous mathematical representations of complex mAb production biosystems (structure and functions), whose usage requires extensive process knowledge and parameterisation.
The complex nature of mAbs manufacturing bioprocesses induces convoluted reaction kinetics, hence globally hampering kinetic model fidelity and parameterisation. Rolandi analysed advanced biopharma modelling via first-principles approaches, emphasising a need for model life-cycle frameworks and sustainable model usage [4].
Recent AΙ developments (especially ML and ANN) mitigate the said challenge and offer an alternative to first-principles bottlenecks, thus facilitating bioprocess modelling. Empirical (especially data-driven) modelling does not mandate a prerequisite of profound mAb manufacturing process knowledge: it relies on a lighter set of assumptions about underlying relationships among system variables, and models can predict outputs from input variables on the basis of fitting against a training dataset, under set assumptions.

2.1. ANN Models

Dewasme et al. (2017) published a mAb production model, employing six sequential suspended hybridoma batch cell cultures using two hybridoma strains, HB1 and HB2 [5]. The configuration kept biomass in the bioreactor, with products (lactate, ammonia, mAbs) withdrawn. Glucose and glutamine concentrations ranged between 6 and 7 and 0.3–0.4 g/L−1, respectively, and the initial biomass concentration of the first batch was at 0.1 × 106 cells/mL−1. Cell culture duration was approx. 15 days; culture media renewal occurred after 7 days. Experimental measurements were obtained once daily, with data illustrated in Figure 2 and Figure 3. Cell growth started to decline after 4 days of batch, returning to the initial value at day 7. With fresh culture medium added at this specific time point, culture viability rises again.
Maximum Likelihood Principal Component Analysis (MLPCA) is used with this data to reduce dimensionality, yielding a matrix ρ of MLPCA components as a subspace basis:
ρ = 0.0074       0.0317       0.4314 0.0173 0.0045 0.3108       0.1366 0.5955 0.6581       0.0169 0.0404 0.0561 0.1389 0.7778 0.5300 0.9805 0.1940 0.0153 . . .
Glucose and glutamine stoichiometric coefficients come from (1); the reaction scheme is:
Substrate   oxidation :     k 31 G +   k 41 G n   φ X +   k 61 m A b
Substrate   surplus   ( to   by - product ) :     k 32 G + k 42 G n   φ X + k 52 L
Biomass   death : X   φ X d + k 63 m A b
Ammonium formation is ignored to achieve a smaller reaction scheme (fewer reactions).
The dynamic reaction model based on mass balances is obtained as an ODE system:
d X d t = φ 1 + φ 2 φ 3 X
  d X d d t = φ 3 X
d G d t = k 31 φ 1 k 32 φ 2
    d G n d t = k 41 φ 1 k 42 φ 2
  d L d t = k 52 φ 2
    d m A b d t = k 61 φ 1 + k 63 φ 3
Dewasme et al. propose a viable strategy decoupling stoichiometry identification from reaction kinetics, thus achieving a first estimate of the stoichiometric parameters [2]. Biopharmaceutical processes are extremely complex and require many model parameters.
Model parameterisation for the ODE systems of Equations (5)–(10) was performed via MATLAB® R2025a fmincon (constrained minimisation) algorithm, with box constraints used to limit the search space, and typically comprises three successive phases. The first considers the MLPCA stoichiometry estimates and a guess vector for reaction kinetic parameters; the next are initialised with parameter values from previous minimisations.
Figure 4 presents direct model validation results: coloured circles and error bars (at 99% confidence) denote experiments, and solid black lines are MLPCA model predictions, for a total of six observables (i.e., live and dead biomass, glucose, glutamine, lactate, mAb). The model performs well, but with clear mAb discrepancy after medium renewal (day 8).
Dewasme et al. (2023) later proposed another data-driven strategy for fast dynamic bioprocess modelling with minimal complexity and a minimum number of reactions [6]. MLPCA is used to infer a reduced reaction scheme from a 25-state mammalian cell culture database, and a multi-stage nonlinear model predictive control (MSNMPC) algorithm is used to tackle kinetic structural uncertainty (and then also pursue robust process control).
Larger datasets are used therein (Figure 4 and Figure 5, [6]) vs. an early study (Figure 2 and Figure 3, [5]), employing a richer knowledge basis of five (5) experiments and 25 signals for each, with a total of five (3 low, 2 high) amino acid concentration levels in the culture medium.
Figure 6 thus justifies that a 3-dimensional subspace (three principal components) suffices for describing selected data (Expts. 1 + 4), via a log-likelihood criterion (JP) vs. the MLPCA subspace dimensionality; a dashed line denotes a χ2 distribution quantile of 5%.
Equation (11) shows a biologically consistent stoichiometric matrix constructed via a linear transformation ( K ^ = ρ G ) from MLPCA and χ2 test-derived basis ρ of Equation (1). All three columns of K ^ are normalised with respect to biomass concentration; the second features non-positive biomass, glucose and glutamine coefficients, and the third has a zero ammonium coefficient (the 25 row labels denote metabolite, amino acid and mAb, cf. [6]).
K ^ = ρ G = 1 1 1 0.493 0.151 2.017 0.266 0.443 0.153 0.139 0 0.466 0.074 0.106 0 0.008 0.041 0.069 0.0085 0.013 0.001 0.014 0.084 0.151 0.003 0.099 0.222 0.023 0.126 0.221 0.010 0.014 0.003 0.012 0.105 0.207 0.005 0.013 0.012 0.064 0.105 0.033 0.006 0.004 0.011 0.002 0.015 0.042 0.002 0.003 0.001 0.011 0.011 0.009 0.001 0.021 0.047 0.001 0.038 0.093 0.002 0.059 0.147 0.001 0.021 0.054 0.005 0.001 0.015 0.001 0.010 0.025 0.156 0.191 0.073
This MLPCA model formulation includes many previously [5] neglected by-products (alanine, proline, aspartate, asparagine, etc.), and Table 1 summarises model parameters.
Direct validation vs. Experiments 1–2 (two culture media) is presented in Figure 7 and Figure 8. Cross-validation was also completed, where only initial conditions were identified, with model parameters set to direct validation values. The other 3 fed-batch Experiments 3–5 (2 with low, 1 with high amino acid concentrations) were used to compute a fit metric, JML. Table 2 lists fitting cost function (JML) residuals for all cross-validation experiments.
A recent study by Kotidis et al. (2020) employed an ANN-based data-driven model to predict protein glycosylation in mAb manufacturing processes [7], for which rigorous modelling requires accurate estimation of many protein-specific kinetic parameters, as well as prior quantification of enzyme and transport protein levels in the Golgi membrane. Their ANN model describing protein glycosylation was applied for a case study including four recombinant glycoproteins, two mAbs and two CHO-produced fusion proteins [7].
The glycosylation process is initiated in the endoplasmic reticulum and continues in the Golgi complex: it greatly depends on intra-Golgi membrane glycosylation enzymes and nucleotide sugar donors (NSDs) levels, controlled by nucleotide sugar transporters (NSTs) in the Golgi apparatus. The complex biological phenomena (glycosylation enzyme and NST level variations) governing protein glycosylation and NSD synthesis cannot be captured by Monod-type kinetic models. The proposed data-driven ANN model of [7,8] is aimed at their description using highly nonlinear relationships between nucleotides and NSDs (process inputs) and glycoform (recombinant protein) distribution (process output), circumventing NSD transport modelling challenges if informative training datasets exist. The HyGlycoM tool combines CHO cell metabolism, NSD and mAb synthesis (Figure 9).
Figure 10 shows that the ANN model can greatly suppress kinetic model inaccuracies, which affect nucleotide sugar estimation, featuring high predictive potential and solid estimates; for some, accuracy deteriorates close to the end of runs (days 11–12).

2.2. Hybrid Modelling

Hybrid models combine more than one system description to boost fidelity [9,10,11,12]. Fu and Barford (1996) pioneered a hybrid system embedding rigorous cell metabolism model (first-principles mass and energy balances combined with reaction kinetics) [13] into an AI framework which used an RL ANN for iterative parameterisation (Figure 11).
The RL ANN block is trained via input (X) and output (Y) datasets and implemented in Fortran to yield model parameters P, which are nonlinear state functions (Figure 12). The 11-step algorithm comprises ANN training, weight updating, internal reinforcement and a performance evaluation unit encoding biochemistry knowledge to decide if the previous P vector is suitable or not vs. the error between predictions Y and actual data E; hence, key hybridoma cell cultivation variables can be kept or altered, respectively [13,14].
A comparative study of its predictive potential was performed vs. experimental data, for the first-principles model (solely based on hybridoma growth kinetics for mAb production) and the hybrid model (encompassing both reaction kinetics and the ANN as above). The former involves detailed stoichiometry and cell metabolism energetics (45 variables and 145 parameters) and can simulate all key states (specific growth rates, cell densities, macromolecular composition, substrates, lactic and amino acid profiles, and mAb titres). The latter, conversely, relies on the RL ANN structure by Quantrille and Liu (1991), in which first-layer outputs are fed in both the second and third (output) layer [14].
Figure 13 and Figure 14 show prediction quality comparisons for both simulation approaches vs. experimental data, assessing the first-principles vs. hybrid (reaction + ANN) model. The superior fit performance of the hybrid vs. the conventional kinetic model is thus clear, even though the authors did not provide any metrics for comparative error quantification.
An advance to the study [8] was achieved by Antonakoudis et al. (2021) (Figure 15), who developed an ANN considering NSD fluxes computed by the stoichiometric model as inputs, and the glycan distribution of the protein of interest (here, IgG) as output [15]. Original NSD flux and glycan production data were used for ANN training, and another independent dataset was then employed so as to test ANN predictive accuracy (Table 3).
A product-specific genome-scale model (GeM) constructed as per another study [16] implemented flux balance analysis (FBA), a linear programming method computing flux distributions that optimise a chosen objective function, subject to conservation constraints. Figure 16 illustrates the validation of FBS algorithms for biomass growth and mAb yield.
Figure 16A shows that the model was most accurate in predicting biomass growth rates in exponential rise and stationary phases, but not in the decline phase, where it failed to capture trends. Specific mAb productivity was predicted well in all periods (Figure 16B).
Botton et al. [17] proposed a hybrid digital (semi-parametric) model (HDM) which tracks viable culture cells, glucose, glutamine, lactate, ammonia, and mAb concentrations. The focus of their effort was data augmentation, where generating in silico results assists the expansion of datasets available from actual (in vivo) bioprocess experimental runs.
The HDM structure (Figure 17) has a mechanistic section describing material balances of chemical species (first-principles model), and an ML section (ANN) employed to estimate the complex, unknown kinetic expressions from cell culture experimental data. Two data generation strategies were tested for final (mAb) product titer estimation: the first relied on a first-principles model only, the other on the said hybrid (HDM) model.
Figure 18 shows HDM-generated mAb titer profiles: Solid red lines are the 3 training batches from the process, while the dashed grey lines denote the 10 simulated trajectories. With 100 in silico batches generated, 10 time points from each were used for HDM training: all generated profiles were subsampled at the same 10 time points of real measurements. Resulting data are organised in matrix XHDM [100×5×10] of culture variable trajectories, and vector yHDM [100×1], with the harvest (final time) mAb concentrations. Very promising results were achieved: the HDM identified highly productive cell lines (high mAb titer), even when very few (2–3) actual experimental batches were available.

3. Process Analysis, Control and Optimisation

3.1. Process Control

ML methods hold promise for biopharma manufacturing process control [18,19,20]. NNs can analyse real-time bioreactor sensor data (temperature, pH, dissolved O2, nutrient concentrations) to monitor bioprocess progress. Learning from historical plant data, NNs can detect anomalies, deviations or drifts, and provide early warning for corrective action. By incorporation in hybrid (or standalone) process models and via real-time sensor data, NNs can reliably predict future process behaviour and optimise dynamic manipulations, ensuring optimal performance while accounting for dynamics and respecting constraints.
Integration into model predictive control (MPC) frameworks is an attractive idea; while bioprocess NN applications have abounded for decades, online process control studies are limited due to practical (real-time and hardware) implementation difficulties.
Gadkar et al. [20] used a recurrent neural network (RNN) study with output layer feedback and intra-connections, tracking a fed-batch yeast fermentation with adaptive weights computed from online concentration (dissolved O2) measurements (Figure 19). Substrate, ethanol and biomass concentrations are not measured but predicted by the NN.
Figure 19 shows the weights for connections between input and hidden layers (Vij), those for connections between hidden and output layer (Wjk), and the respective ones for weighted intra-connections (dashed lines) from the dissolved O2 (middle) to other output layer nodes (Uk). Bias nodes (ξj and ξk) usage is also essential for this NN formulation [20]. For NN tuning and architecture optimisation, a training dataset was generated at a constant 6 min sampling interval for 15 h, and random initial weights in the ±0.1 range.
Figure 20 (left) presents recall profiles from training, confirming excellent fit therein. To achieve reliable NN performance outside the training regime, however, all NN weights had to be continuously updated using the online adaptation algorithm described in [21].
Online weight adaptation as per [20] requires continuous information flow from actual experiments, via at least one online measurement of sufficient process importance. Dissolved O2 concentration meets this requirement, reflecting the process metabolic state. Given O2 availability and respiratory capacity, growth occurs on both glucose and ethanol.
Figure 20 (right) shows ΝΝ-driven process control (with/without weight adaptation) vs. noisy O2 experimental data, where NN predictions (B) with online weight updates are noisier but track more accurately concentration state profiles (A) than those without (C).
The NN is used to control cell concentration along a linear rise profile prescribed as:
X = c t +   c 1
with c and c1 constants, t the elapsed time during fermentation and X cell concentration. Fermentation runs used a given medium (with 2% glucose, 2% peptone, 1% yeast extract) at Τ = 30 °C, pH = 5 and stirring (300 rpm), with X measured via optical density (600 nm). The NN was used to obtain state variable estimates for evaluating instantaneous feed rate. Figure 21 (right) shows NN-driven control improves a lot with online weight adaptation due to rapid computations (1–2 s), even if training occurs for variant c via Equation (12).

3.2. Process Optimisation

AI can revolutionise the optimisation of mAb manufacturing processes via improved efficiency, cost-effectiveness, and quality control. Big datasets from manufacturing plants can be analysed, to identify patterns or correlations overlooked by human perception.
Optimising process parameters (e.g., temperature, pH, agitation speed) can enhance mAb yield and purity (plant scale); planning inventory levels vs. demand variation, potential bottlenecks or broader disruptions (supply chain scale) yields financial benefits by reduced OpEx costs and improved efficiency of mAb production and distribution.
Traditional mAb manufacturing process optimisation involves testing a small set of specific (often expert-derived) hypotheses, via costly and time-consuming experiments. Data-driven methods are a promising alternative approach, relying on computer models with parameters inferred from process data, economising on experimentation expenses.
Pham et al. (2023) explore Supervised Learning and Data-driven Optimisation (SLDO) studies in mAb batch process monitoring and control, illustrating bibliometric data on SLDO applications for mAbs [3]. Their keyword-based search strategy (‘SLDO methods’, ‘mAb bioprocesses’, ‘biomanufacturing’) across 3 decades (1994–2022) reveals that most papers concern upstream problems, and most (90%) were published between 2010–2022 (Figure 22).
Manapragada et al. [22] studied mAb manufacturing process optimisation in a fed-batch culture via ML (Figure 23), in case of insufficient data for surrogate model training, under two distinct (uniform and non-uniform) feeding patterns, as per its daily variation. The feeding pattern schedule is portrayed by a vector of 10 features, each representing the daily bioreactor feed amount for the entire duration (day 3 to day 13) of the culture period. Their input space X included the feeding schedule and concentration values for three (expert-selected) amino acids. The objective function to be maximised is product yield.
Figure 23 presents their ML model architectures employed: first, their simpler Naive model has just a 5-neuron hidden layer between the input and output layers. Their ITO model is a dual configuration for capturing the distinct lifecycle stages of cell cultures in mAb manufacturing: it comprises two shallow ANNs, the Input–Throughput and the Throughput–Output model, respectively. An intermediary logical layer (with indexed throughput nodes) forms the output of the first (but also input to the second) ANN [22].
Both in silico and in vivo experimental validation ensued: the first used 100 separate seeds to generate 100 objective function candidates for each (Naïve, ITO) type, and a grid search for 5·5·5 = 125 input combinations (the amino acids 2, 4, 17, varied at 5 levels each) produced the best decision vector, maximising final mAb product concentration (Table 4).
For the in vivo experiments, a six-bioreactor fed-batch study tested two optimised feed flow regimes (non-uniform as well as uniform feeding) vs. the respective reference (control) base cases. Table 4 summarises results for all cases, and Figure 24 compares all in silico runs vs. experiments: both ITO and Naive models predict the outputs accurately.
A review by Rathore et al. [10] summarises AI/ML applications in biopharmaceutical manufacturing, with focus on multivariate data analysis, ANN and RL methods (Figure 25). Optimisation and control of biopharma unit operations is challenging, as it often involves complex phenomena (e.g., protein expression in cell culture bioreactors, binding on composite chromatography resins, pH-driven virus inactivation, ultrafiltration via semi-permeable membranes). The authors opine that there is an exceptionally promising scope for AI/ML at plant-wide manufacturing scale, due to vast datasets from diverse process sensors (e.g., pH, temperature, conductivity, redox, UV, weight, level, flow rate, pressure). Many subsequent studies confirmed this booming field of sustained R&D interest [23,24,25,26]. A focal challenge is accelerating bioprocess scale-up via ML-driven frameworks [27,28].
The Dewasme et al. model was used for optimisation, to compute optimal medium renewal time and composition, to maximise mAb production at lower substrate cost [5]. The goals form an objective function, subject to the validated model of Equations (1)–(11).
J o b j = m A b ( t r e n e w a l ) 2 + m A b ( t f ) 2 α ( G r e n e w a l 2 + G n r e n e w a l 2 )
where α is a weight coefficient accounting for substrate savings vs. mAb production, thereby arbitrarily specifying the predominance of one optimisation target over the other. The MATLAB® 2025a fmincon (constrained optimisation) algorithm is used to find the optimal decision vector θ = [trenewal Grenewal Gnrenewal]T, i.e., the medium renewal time, and the glucose and glutamine concentrations in culture medium at that medium renewal point.
Figure 26 depicts optimisation results for α = 0 (disregarding substrate cost savings) and α = 10 (ascribing a strong emphasis on substrate cost savings vs. mAb production). The optimal medium renewal time is trenewal = 4.54 vs. 7 days (i.e., much later), respectively. The final mAb concentration is 60.92 vs. 75 µg.mL−1, respectively (after the 14-day period); thus, even the first (smaller) improvement is a 30% gain vs. experiments (40–45 µg.mL−1). A more economical fermentation may thus last 56% longer (before first renewal), but also offer a remarkable (23%) extra mAb production increase, in addition to the said for α = 0.
MLPCA thus reliably extracts critical information from experimental data, achieving:
  • Rapid stoichiometry estimation that is independent of reaction kinetics, reducing the number of unknown model parameters (half of the latter are actually stoichiometric).
  • A small number of reactions, to avoid useless model complication vs. experimental rig while capturing underlying biological mechanisms (viability, growth, inhibition).
Their subsequent study via the same model [6] presents a comparative evaluation of multi-stage vs. classical nonlinear model predictive control (MSNMPC vs. NMPC), with the “divide and conquer” approach allowing for stoichiometry and reaction kinetics to be estimated iteratively (separately or simultaneously), via estimates from previous steps. All 25 state variable concentration (and volume) trajectories are illustrated in Figure 27.
Figure 28 presents MSNMPC vs. NMPC optimised feed rate profile comparisons: the first achieves better performance via faster but also smaller actuator variations, while the second features smoother input trajectories with abrupt peak variations in the first 2 days. Glutamine (bottom panel) shows early abrupt rises in both cases, then rapidly attenuates.
The authors conclude that NMPC computes continuous reaction rates but cannot also detect possible metabolic switches cancelling/activating some of the rates at critical times. The MSNMPC model is challenged and compared to the classical economic NMPC, using a simulated plant with discontinuous reaction kinetics representing metabolic switches. Results show that classical NMPC performs well, but the MSNMPC algorithm achieves superior accurate tracking and increased robustness vs. structural uncertainties globally.
Le et al. (2018) explored the development of a rapid ML-based analytical method for mAb manufacturing quality control, focusing on the production of four mAbs (Infliximab, Bevacizumab, Rituximab, Ramucirumab) in a range of conditions and concentrations [29]. This can be critical, e.g., for ensuring mAb quality before administration by medical units. The study used a novel collaborative data challenge platform, namely Rapid Analytics and Model Prototyping (RAMP), collating problem solutions from ca. 300 data scientists.
ML analyses occurred in partnership with the Paris-Saclay Centre for Data Science (CDS) using the Scikit-Learn library [30]; Figure 29 shows Raman spectra for all mAbs. Raman and near-infrared (NIR) spectroscopy enable direct mAb measurements through glass/plastic packaging (unlike HPLC/UV or LC/MS/MS chromatography), so a linear and an ML method were comparatively used towards classification and regression (Table 5). Classification via Principal Components vs. Partial Least Squares versions of Discriminant Analysis (PCA-DA and PLS-DA), and via ML, was used to estimate mAb concentrations.
The challenge here is to predict mAb concentrations in solution: Table 5 results clearly indicate that the ML approach achieves much lower error metrics for all four mAbs, and overall. With linear (PCA-DA and PLS-DA) tools, 63% of samples at a low concentration ≤ 1 mg/mL) had error of over 15%; the error is much smaller with the use of ML (5.6%). ML model bias is negligible at concentrations above 0.1 mg.mL−1; relative error drops sharply as mAb concentration of 15% (y < 1 mg.mL−1) falls to 5% (y > 3 mg.mL−1).
The ML models employed are sensitive to the scale of data: the majority of effective prediction strategies use nonlinear kernel-based methods, e.g., kernel PCA for nonlinear dimensionality reduction [31] and support vector machines (SVM) for prediction [32]. Nitika et al. (2024) used convolutional neural networks (CNN) with Raman spectroscopic data in their process analytics tool for mAb charge variant monitoring and prediction [33].
Another fruitful field for mAbs optimisation applications is production operations. Response surface methods (RSM) offer multi-dimensional visualisation of cause (x, y axes) vs. effect (z axis) [34,35], and are a lot more intuitive than PCA/PLS approaches [36,37]. Bashokouh et al. (2019) [36] used an ANN to optimise cell culture cultivation conditions (i.e., incubation time, temperature and foetal bovine serum/FBS concentration) to maximise IgM mAb production by hybridoma M1A2 cells. Two cell culture media were used: Dulbecco’s Modified Eagle Medium (DMEM) and Roswell Park Memorial Institute Medium (RPMI 1640). Experimental data were analysed via the three-layer feed-forward ANN structure to establish a predictive model that could optimise IgM mAb production.
The ANN comprises a three-neuron input layer (cultivation time, temperature, FBS), an eight-neuron hidden layer, and a single-neuron output layer (IgM mAb concentration), trained with back-propagation to find the optimal topology via RMSE and R2 metrics [36]. IgM concentration measurements were obtained via a protocol by Ishida et al. (2018) [38].
Table 6 shows the highest IgM mAb production (1220 µg.mL−1) is obtained in DMEM with 17% FBS, cultivated for 5 days at 33 °C; the mAb production drops to 719.24 µg.mL−1 for a DMEM at low (7%) FBS, at identical cultivation time and temperature (thus a 41% reduction in IgM mAb production occurs for a 59% FBS concentration reduction).
Figure 30, Figure 31 and Figure 32 present IgM concentration curves vs. each of the three operational variables, indicating similar trends (but with substantial differences in mAb productivity).
The mAb results vs. FBS concentration results agree with Mel et al. (2008) [34], who studied FBS effects (5–15% range) on RC1 hybridoma viability (IgG mAb-secreting cells). Temperature can impede mAb growth (Figure 31), as per Chong et al. (2008) [39], who show hybridoma C2E7 cell growth is slower under mild hypothermic (32 °C) conditions; maximum viable cell density is 10% lower, but mAb productivity is 5% higher vs. 37 °C.
The 3D plot (Figure 33) transcends the univariate RSM view of Figure 30, Figure 31 and Figure 32, as it shows mAb productivity (z-axis) vs. two-at-a-time of three independent variables; beyond capturing evident local maxima, the culture media also clearly affect yield (A–C vs. D–F).
Maximised mAb productivity may suffice at bench- but not production-scale [40,41,42], since technical as well as economic performance metrics are pivotal to biopharma ventures. George and Farid (2007–2009) in stochastic combinatorial multiobjective optimisation (MOO) studies employed [41,42] went beyond previously established mathematical approaches (e.g., mixed-integer (non)linear programming, MILP/MINLP) of optimisation frameworks for manufacturing capacity planning and scheduling in mAb plants [43,44].
A combined ML-evolutionary computation strategy was chosen to tackle the significant complexity of the problem via of an estimation-of-distribution algorithm (EDA). The complete stochastic optimisation framework employed is illustrated in Figure 34 [41].
Model development combined C++ (for fast calculations saving CPU time) and MS Excel/VBasic (versatile in data storing, extraction, manipulation and plotting). The reward (profitability) and risk metrics for identifying the best strategies were the mean positive net present value (NPV) and the probability for it to be positive p (NPV > 0), respectively.
Economic performance vs. multiple objectives in an uncertain environment requires case studies; for instance, that in [42] assumes a biopharma company with 10 mAb drug candidates available for development, which must decide on a drug portfolio of limited size. The optimisation model considers commercial product characteristics, technical success probabilities, durations, revenue and costs (capital and operating expenditure, royalties). Stochastic variables addressed are taken to obey triangular probability distributions, (convenient for small datasets). Four cash flow constraints (i.e., −$75 MM, −$200 MM, −$200 MM, unconstrained) portray company cash limitations to R&D funding.
The stochastic optimisation settings include: max. number of iterations (tmax) = 17, candidates in each generation/iteration (|G(t)|) = 1000, Monte Carlo trials per candidate strategy (U) = 250, and superior candidate strategies in G(t) at each iteration (|S(t)|) = 500.
Figure 35 shows optimisation results for a five-drug portfolio under said constraints, with four clear Pareto fronts and strong relation (trade-off) between NPV and p(NPV > 0). For a given risk level p(NPV > 0), each tighter cash flow constraint reduces the mean NPV.
The rigorous formulation of [42] combines ML and evolutionary computation under uncertainty, achieving efficient decision space search and reliable tracing of a Pareto front.
Table 7 summarises key features of impactful publications discussed in this section.

3.3. Process Monitoring and Control

The biopharmaceutical industry follows good manufacturing practices (GMPs), but process and product variability (often due to feedstock variability) pose great challenges. Multivariate statistical methods, as per [5,6], have been established for decades in batch process monitoring (BPM), and can efficiently monitor and control biopharma processes.
Nikita et al. (2022) and Park et al. (2023) published two papers summarising many advances on real-time quality prediction in continuous mAb manufacturing [45] and on digital twins for multistep forecasting of cell culture evolution profiles, respectively [46]. Biochemical complexity necessitates the development of elaborate frameworks in the case of continuous mAb manufacturing, e.g., a study on control of surge tanks employs four layers (data acquisition, process scheduling, deviation handling, real-time execution) [47].
Tulsyan et al. (2019) [48] addressed the real-time statistical process control (SPC) problem for processes with limited production history or modest data availability (Low-N datasets). Low-N scenarios pose several BPM challenges, so a transition from a Low- to a Large-N scenario can involve generating an arbitrarily large number of in silico batch datasets [48]. This happens in two steps: first, stratified resampling occurs on the original data stored in the high-frequency data acquisition system (DAQ); then, a generator model is trained on the said resampling data, via a Gaussian process state-space model (GP-SSM) (belonging to a class of Bayesian non-parametric methods) which provides the flexibility for portraying complex, nonlinear, non-stationary (batch) process dynamics (Figure 36).
The Hotelling T2 and squared prediction error (SPE) provided the control limits: the first detects shifts and deviations from normal operations, and the second quantifies the magnitude of variations in the data that cannot be explained by the postulated model [48]. Figure 37 presents a T2 chart for a new incoming batch based on the Low-N model (left), depicting large T2 values yet without control limits to flag atypical operation of concern. Employing in silico batch dataset generation (via GP-SSM) yields a richer profile (right), in which we can distinguish three separate areas outside the colour-lined control limits.
The conclusion is that the Large-N outperforms the Low-N model in explaining and predicting data variation; the former has higher statistical stability vs. the latter (due to its Law of Large Numbers/LLN-based assumptions in the BMP theory), and can credibly pinpoint three separate alarm events denoted as (1), (2) and (3) in Figure 37 (right), in which the signal surpasses both 95% and 99% T2 control limits, for several consecutive samples (indeed the right panel is far more indicative of T2 plots derived from batch operations).
A study by Jin et al. (2018) [49] presents a PCA-based bioprocess monitoring method for analysing abnormal variations often experienced in commercial mAb manufacturing, adversely affecting plant performance and consistency of final pharma product quality.
Their case study concerns Immunoglobulin G (IgG), which is the most abundant type of antibody in human blood and extracellular fluid (ca. 75% of serum antibodies); a potent remedy against bacterial and viral infections, it neutralises pathogens and binds to them for destruction by other immune cells. For such IgG plants, such abnormal variations are correlated with high lactate in recombinant CHO cells producing IgG, and lactate titres were also predicted therein via a nonlinear support vector machine (SVM) classifier [30].
Figure 38 presents a typical signal variation (the study used a dataset from 166 runs), observing that high- and low-lactate batches could not be intuitively (or unmistakably) separated in the original (principal) space. Hence, a latent space was then searched and determined, so as to separate them under reduced coordinates (PCA-based monitoring). Figure 38 elucidates how high- and low-lactate runs are inseparable in the principal (vs. the T2 limit), but differentiable in the residual space vs. the respective (SPE) control limit.
Jin et al. (2019) in a later study [50] further observed that the lactate concentration fluctuated widely near the end of batch runs, and did so even more for high-lactate ones. Due to a quasi-bivariate final lactate concentration distribution, it is easy to distinguish high- and low-lactate profiles, but also with many intermediates unclassified (Figure 39).
Figure 40 shows that using batch industrial production data of an 8-year period is ineffective and inconclusive in the PCA subspace: high- and low-lactate batches scarcely differ in terms of T2 values (left), whereas SPE is very effective in the residual subspace. High- vs. low-lactate batches were thus easily separated for most (if not all) cumulative proportion of variance (CPV) values, with the best distinction seen for CPV = 90% (right).

4. Downstream mAb Processing

The foregoing studies discuss AI/ML applications for upstream mAb manufacturing process modelling and optimisation only, within or near the bioreactors producing mAbs. Beyond that, AI/ML can improve many pivotal downstream operations (chromatographic separations, filtration, viral inactivation) for better quality control and plant sustainability.
Figure 41 is a process flow diagram for the downstream part of continuous mAb manufacturing: after the bioreactor and cell culture harvest, the resulting mAb solution undergoes protein A chromatography, viral inactivation, filtration [51], anion exchange membrane (AEM) and cation exchange (CEX) chromatography before final formulation.
The study by Nikita et al. [45] details ML application to develop prediction models for capture and polishing chromatography steps, comparing 5 different ML algorithms: (i) deep neural networks (DNN), (ii) support vector regression (SVR), (iii) decision trees (DT), (iv) random forest regression (RFR), and (v) gradient boosting regression (GBR). Intra-batch real-time data were compiled for key variables (conductivity, pH, UV sensor).
A brief overview of fundamental concepts for some of the methods in [45] follows. The DNN implementations rely on several processing layers that learn based on process conditions xi, predicting parameters wi and targets y. The nonlinear output function is:
y =   α ( x 1 w 1 + x 2 w 2 + x n w n + b )
where α is the activation function, xi are input values, wi are the weights, and b is a bias.
Figure 42 and Figure 43 show RFR and GBR principles: the first averages predictions from many diverse structures, and the second relies on successive training/testing cycles [45].
Model fidelity is measured by the root mean square error (RMSE) and the coefficient of determination (R2), close to 1 for reliable regressions (training datasets are excluded):
R 2 = 1 i y i f i 2 i y i y ¯ 2
where yi are experimental data, y ¯ is their average, and fi are the model predictions.
Figure 44 shows that all methods achieve adequately high accuracy for most observables (R2 > 0.8), exception SVR for aggregate content (first group, last bar), but also DNN for mAb elute concentration of Protein A chromatography (last group, first bar). Akaike (AIC) and Bayes (BIC) information criteria rule out GBR use: AIC/BIC values are 5–65% higher than tree-based models, and low criteria values maximise fidelity [45].
Figure 45 shows all model predictions vs. experiments (left) and error metrics (right). DT and SVR emerge as most accurate (minimal RMSE), with GBR and NN least reliable; RFR is superior to DT and other algorithms for small datasets (due to reduced CPU time), achieving 4.76% max. error in mAb elute level prediction from protein A chromatography.

5. Biopharmaceutical Process Scale-Up

AI can play a crucial role in biopharmaceutical process scale-up, which belabours the transition from laboratory-scale processes to larger production scales. AI algorithms can analyse large datasets generated during laboratory-scale processes and identify patterns, correlations, and process parameters influencing product quality and yield. Data-driven approaches can help us develop accurate models and extrapolate to larger sizes, to predict (via historical data and process knowledge) how bioprocesses behave when scaled up. By simulating various scenarios, AI can help identify potential challenges and optimise scaling-up, reducing the need for costly and time-consuming trial-and-error experiments.
Facco et al. (2020) used multivariate statistics, data analytics and pattern recognition for biopharma process scale-up in mAb production (Figure 46) [27], achieving cell line screening and performance prediction for selecting static, shaken or agitated bioreactors.
Data-driven models used in [27] make use of multivariate statistical methodologies and standard pattern recognition techniques. Principal component analysis (PCA) is thus deployed for data processing; it summarises the information available in a dataset X as:
X = a = 1 R t a P a T = a = 1 A t a P a T + a = A + 1 R t a P a T = a = 1 A t a P a T + E
where R is the rank of X; Pa is the [M × 1] loading vector and ta is the [N × 1] score vector for the ath PC; superscript T is the transpose operator; E is the [N × M] residual matrix (to be minimised in least-squares sense), and A is the number of PCs that need be determined.
Two statistical modelling tools, Partial Least Squares (PLS) and Joint-Y PLS (JY-PLS) were used for predictive analysis: these regression methodologies determine the direction of maximum variability in a set of regressors (inputs) X [N × M] that are most predictive for a given set of responses (outputs) Y [N × U], according to the next matrix equations:
X = T P T + E
Y = T Q T + F
T = X W *
where T is the [N × A] score matrix, P [M × A] and Q [U × A] are the loading matrices for inputs X and outputs Y, E and F are the [N × M] and [N × U] residual matrices of X and Y, respectively (to be minimised in least-squares sense), and W* is the [M × A] weight matrix.
The study considered 4 PCs: Figure 47a shows PC1 load cumulative absolute values of all measured variables across all batches, and Figure 47b proves the model correctly captures titres, and a pronounced anti-correlation of glucose (cell nutrient) vs. lactate. The authors of [27] also report similar results from a dynamic analysis of the 2 L bioreactor, implying the variables affecting variability among batches are the same across both scales.
The PLS (local) model uses available data (metabolites, pH, O2 transfer rates), while the JY-PLS (global/shared) model uses data from all other scales except for the predicted.
Table 8 summarises the validation and performance for both modelling approaches. For PLS and JY-PLS and all bar one scale, the determination coefficient R2 is at least 0.94, indicating reliable estimation. The PLS model is more accurate than JY-PLS in all bar one cases; the best and worst overall validation results at two scales are shown in Figure 48, showing all estimation errors well within a standard deviation of measured data (μ ± 1σ).
Figure 49a portrays Integral of Viable Cells (IVC) estimates from shake-flask data; PLS is good, but JY-PLS (inclusion of other-scale data) worsens estimation performance. Figure 49b, however, shows that the peak viable cell count (VCC) trend is very different: JY-PLS is much better than PLS herein, indicating that information from other scales can improve estimation accuracy, a clear benefit to correlation structure at the particular size. The JY-PLS advantage is greater if only a few samples are available at shake-flask scale.
Figure 49c depicts Peak VCC estimation also, but for only 5 samples available now: PLS accuracy declined enormously, while JY-PLS performance deteriorated very slightly. The conclusion is that data scarcity at a given scale can be mitigated by exploiting valuable information embedded in data at other scales, reducing R&D risk in cell culture selection.
A review on Quality by Design (QbD) via ML by Walsh et al. [52] describes the role of statistical modelling in biomanufacturing, showing how certain ML tools can analyse bioprocess datasets, to accelerate systematic QbD implementation in industry (Figure 50). Critical process parameters (CPP) and quality attributes (CQA) are defined, followed by a comparison of multivariate data analysis (MVDA) vs. supervised and unsupervised ML.

6. ML for mAb Property and Quality Control

Beyond AI/ML applications in plantwide mAb manufacturing, ML methods [53,54,55,56,57,58,59,60,61,62] emerge as useful in predicting, calculating and controlling specific physical properties of mAbs: the next sections describe a variety of successful ML implementations to this end.

6.1. Viscosity

Viscosity is one of the pivotal mAb physical properties determining their final market value: mAb solutions exhibit a significant viscosity increase with increasing mAb titres, which in turn complicates manufacturing processes, limiting options for product delivery as the underlying molecular basis of rheological behaviour is still not widely studied [53].
Lai et al. (2021) [54] measured concentration-dependent solution viscosity of 27 FDA-approved mAbs, in the 50–200 mg.mL−1 range (pH = 6.0, 10 mM histidine-HCl buffer). Predicted net charges and special charge map (SCM) scores were separately used in order to classify experimental viscosity data, and then ML algorithms were applied to search for molecular descriptors for those cases that accurately separated high and low viscosity. A high viscosity index (HVI) was used as a metric for rapid (low vs. high viscosity) mAb screening, and a decision tree model based on the mAb net charge and HVI was proposed.
A comparative analysis of viscosity predictions for immunoglobulin G variants [54] has evaluated the models proposed by Li et al. (2014) [56] and Sharma et al. (2014) [57]. Figure 51 summarises viscosity predictions based on these models: solid and dashed lines therein denote the least squares regression and the identity (diagonal) lines, respectively.
The first model [55] predicted the viscosity of 27 proprietary mAbs at 150 mg.mL−1 (pH = 5.8, 20 mM histidine-HCl buffer), underestimating most experimental data (r = 0.64). The second model [56] pursued viscosity prediction for the same mAbs at 180 mg.mL−1 (pH = 5.5, 200 mM arginine-HCl buffer), at worse correlation to experimental data (r = 0.4); on the latter, it is noted that solution condition mismatch may affect model performance.
Another ML study of viscosity prediction [54] partitioned 27 mAbs into training and test sets for model evaluation and feature selection, including key molecular descriptors (e.g., charge, hydrophobic and hydrophilic properties) which affect rheological behaviour. The area under the precision-recall curve (AUPRC) was chosen as the accuracy metric, so the top 3 features (max. AUPRC) found are the hydrophobic and hydrophilic residues in Fv regions (N_hydrophobic, N_hydrophilic) and the charge symmetry parameter (CSP); achieving 86%, 83% and 81% accuracy, respectively, for the AUPRC values of Figure 52.
Authors note high-viscosity mmAbs are richer in hydrophilic (poorer in hydrophobic) residues, concluding (via charge analysis) that mAbs of high or low net charge have lower viscosity due to the resulting higher repulsive interactions between proteins, or strong attractive interactions that lead to larger cluster formation and opalescence. Conversely, the rheology of mAbs with medium net charge is governed by short-range interactions.
Lai et al. (2022) [55] later predicted and measured mAb aggregation rates and viscosity at 45 °C and 150 mg.mL−1 for a set of 20 AstraZeneca (preclinical and clinical-stage) mAbs. The collection included 18 IgG1- and 2 IgG4P-subclass products; accelerated aggregation tests (at pH = 6.0, using a 20 mM histidine-HCl buffer, for 2 weeks) provided experimental data for essential comparisons vs. mAb aggregation predictions, to evaluate ML model fidelity via the leave-one-out cross-validation (LOOCV) method for various regression models [55]. Molecular Dynamics (MD) simulations provided full-length antibody features (specific aggregation-inducing sequences) for ML model development.
Figure 53 shows the comparison of predicted vs. experimental aggregation rates and fit statistics (RMSE and linear correlation coefficients, r) computed via the best two-feature combination from three regression models: linear, support vector regression (SVR), and k-nearest neighbours (kNN); the latter achieved both max(r) = 0.89 and min(RMSE) = 1.07. Later advances explored concentration-dependent viscosity behaviour [58,59], viscosity reduction [60], and targeted high-concentration mAb product stability improvement [61].

6.2. Thermal Stability

Thermal stability is a critical feature in determining R&D prospects of therapeutics. A recent thermostability study by Harmalkar et al. (2023) [62] addressed multispecific biologics (msAbs), which can successfully engage multiple targets (unlike mAbs that only bind to a single target): msAbs feature a single-chain variable fragment (scFv), a fusion protein of the variable regions of heavy (VH) and light (VL) chains of immunoglobulins, connected with a short peptide linker (ca. 10–25 amino acids). A key challenge hampering the market potential of msAbs is their poor physical properties vs. conventional mAbs, because of scFvs, and the relatively poor product stability in case of varying temperatures.
The study [62] employed DL methods to extract thermal stability characteristics from experimentally generated scFv sequences, then screened them vs. temperature sensitivity. The unsupervised CNN predictive potential for thermostability was tested via pre-trained language models (PTLMs), and SL was subsequently applied to train CNN architectures.
Figure 54 shows key biological challenges to predict antibody thermostability from sequences, e.g., identifying particular sequences for thermally stable structure formation. Both regression and classification were used to predict thermal stability via temperature measurements (TS50) for 2700 scFv sequences. For the first, absolute TS50 data were used; for the latter, TS50 data [63,64] were divided into four bins, to be used with a pre-trained language model (PTLM), and with a supervised CNN with Rosetta energetic features [65].
The ML models received enriched information (including thermodynamic features), unlike conventional ML approaches relying on sequence or structural coordinates only. The conclusion is that fine-tuning pre-trained sequence models with thermostability data can greatly improve their classification performance vs. zero-shot predictions (Figure 55).
Figure 56 analyses SL model performance plotting receiver-operating-characteristic (ROC) curves derived from the prediction of the 70-up bin with the energetics-only model. The smaller sample sizes of test scFv and isolated scFv curves explain their few (3) kinks.
Data from thermal aggregation experimental studies by Koenig et al. (2017) [63] and Warszawski et al. (2019) [64] served to evaluate whether CNNs trained on TS50 data can provide insight into temperature-specific features (e.g., thermal aggregation), and if they can differentiate between thermally enhancing vs. thermally hampering mutations.
Figure 57 compares both models’ predictions vs. experimental thermostable mutants for heavy and light chains, revealing the performance of both supervised CNN and PTLM.
Spheres indicate experimentally validated mutants improving Tm; pink indicates ML predictions matching experimental residue positions (but different amino acid mutation), purple denotes ML predictions matching both experimental residue positions and amino acid mutations, and grey indicates mutations not observed in computational predictions.

6.3. Characterisation

Various stresses during upstream/downstream mAb (and other therapeutic protein) manufacturing generate subvisible particle populations [66,67]: Greenblott et al. (2022) [66] used CNN to identify and classify particle generation root causes by subjecting three different mAbs (from formulations) to common manufacturing stresses, as per Table 9. The particle population generated by stressing was probed, indicating that CNN analyses are sensitive not only to the applied stress but also to the mAb type and buffer conditions.
ML models trained on particle images created with one mAb-buffer system cannot be accurate for root cause analysis if applied to particles generated by other such systems. A lever-rule analysis of CNN-derived fingerprints was employed in that study [66] to quantitatively characterise the composition of mixtures comprising various particle types. The FIM particle analysis method of [66] advances earlier work by Calderon et al. (2018) [67] relying on RGB preprocessing to portray temporal evolution of particle fingerprints.
Flow imaging microscopy (FIM) analyses digital images of particles [66,67] from mAb formulations, e.g., with a FlowCAM® VS instrument (Fluid Imaging Technologies). Figure 58 is an FIM image collection from mAb1 samples subjected to six stress conditions.
A fingerprint CNN was trained on particle images captured after subjecting mAb 1 to Stress F, or Stress GL, or depicting ETFE particles. The trained CNN model was then validated vs. FIM test datasets, and Figure 59 shows embedding outputs for both stresses.
Figure 60a shows a test embedding fingerprint from a 50–50% particle mixture which is generated by pumping mAb 1 via large-D Gore® tubing (Stress GL) and freeze-thawing (Stress F) mAb 1 (green) with the training fingerprint by Stress GL (orange contours). Cross-hairs indicate FEMTest Mixture (white), FEMTrain F and FEMTrain GL (magenta). Euclidian distances between key FEM coordinate points are indicated with curly braces, and contours refer to 90%, 75%, 50%, 25%, and 10% of the embedding point cloud mass. Figure 60b is a scatterplot of DTest Mixture, Train GL vs. the particle fraction generated by Stress F (the shaded area represents a 95% confidence interval on a linear fit of the relevant data).
Beyond mAb properties (e.g., viscosity, thermal stability), a concern for industrial mAb therapeutics manufacturing is the detrimental agglomeration of many biomolecules; Shrivastava et al. (2023) [68] authored an ML-driven dynamic light scattering (DLS) study. DLS is a well-established method for evaluating sample stability and estimating average protein aggregate sizes by measuring size distributions in a wide (nano-to micro-) range, to compute relative percentages of multimers (monomers, dimers, trimers, tetramers) for two products (mAb1, mAb2), in the 10–100 nm range (e.g., size distributions in Figure 61).
Their experimental protocol entailed the manipulation of key operating conditions (e.g., pH, temperature, light intensity) for generating a wide variety of mAb aggregates. Light intensities of scattered (633 nm He–Ne) light from resulting mAb solutions were measured using a Zetasizer Nano ZS 90 instrument (Malvern Instruments, Malvern, UK). These DLS scattered intensities were recorded at a fixed scattering angle, and all samples were allowed to attain equilibrium before performing 11 scans for each measurement. Particle size distributions (PSD) were then calculated by post-processing the said DLS data, via proprietary Zetasizer software (Malvern Instruments, Malvern, UK) (Figure 61).
Thereafter, two ML algorithms were used for quantifying various mAb multimers. The first is an SVR/SVM method (a powerful SL tool), the other is an NN of multiple layers and neurons with adjustable weights; both algorithms were used to describe DLS data.
A total of 150 and 132 data points were used to study mAb1 and mAb2, respectively: 130 (mAb1) and 112 (mAb2) of them were used to train the models, while the remaining 20 (for both mAb cases) data points were used for the equally essential NN model testing. The goal is to quantify the composition (all multimer components) using these algorithms.
NN and SVR predictions for test and validation datasets are directly comparable for mAb1 and mAb2 products, respectively. For mAb1 test data, both algorithms predicted oligomer composition with determination coefficients (R2) of 0.97 (NN) and 0.96 (SVR). For mAb1 validation data, respective R2 values were 0.95 (NN) and 0.94 (SVR) (Figure 62). For mAb2 test data, oligomer prediction R2 values are 0.95 (NN) and 0.93 (SVR), but for mAb2 validation data, both R2 values are the lowest: 0.93 (NN) and 0.91 (SVR) (Figure 63).
The conclusion emerging from the said study is that DLS-ML (NN and SVR) results were in good agreement with mAb experimental datasets: the NN prediction algorithm achieved higher accuracy vs. the SVR model, for both the validation and the test datasets. The proposed ML approach was deemed to have great potential for application at various process R&D and analysis stages, for the given (mAb) and other therapeutic modalities. High-throughput mAb aggregation prediction at high concentration is also achieved [69].
A study by Gentiluomo et al. (2019) [70] proposes interpretable ANNs for therapeutic mAb biophysical property prediction (melting temperature Tm, aggregation onset temperature Tagg, and interaction parameter kD) as a function of pH, salt concentration and amino acid composition (Figure 64). The authors used early-stage screening datasets for ANN training, with accuracy: Figure 65 shows ANN predictions vs. experimental data.
Model Tm, Tagg and kD predictions from mAb amino acid composition and the said formulation conditions (pH, salt concentration) were cross-validated (data for two mAbs were selected and held back from ANN training, for completing an unbiased validation). The higher RMSE for Tagg vs. Tm (2.01 vs. 0.87 °C) can be attributed to high-throughput screening, which stretched the required high data density for determination of the onset. The kD sign predictions are very good, without false negatives or false positives (Table 10).

7. Conclusions

Artificial Intelligence & Machine Learning (AI/ML) methods and software hold great promise in bioprocess R&D [71], with booming focus on biopharma mAb therapeutics. This review highlights pivotal ML applications in plantwide mAb modelling, bioreactor, downstream and portfolio optimisation, process control and property prediction [72,73,74,75], belabouring several AI/ML-driven protein structure-function mapping studies [76,77].

Author Contributions

Conceptualisation, D.I.G.; methodology, F.A. and D.I.G.; investigation, F.A. and D.I.G.; resources, D.I.G.; writing—original draft preparation, F.A. and D.I.G.; writing—review and editing, F.A. and D.I.G.; visualisation, D.I.G.; supervision and project administration, D.I.G.; funding acquisition, F.A. and D.I.G. All authors have read and agreed to the published version of the manuscript.

Funding

The authors gratefully acknowledge the Higher Education Commission (HEC) of Pakistan for a Ph.D. Fellowship awarded to F.A. A now completed Royal Society Short Industrial Fellowship (2020–22) and a recent Royal Society International Exchanges Programme grant IES\R2\232014 (2023–25) have been awarded to D.I.G. We also acknowledge financial support from the Engineering and Physical Sciences Research Council (EPSRC UK), under the auspices of a recent research grant (RAPID: ReAltime Process ModellIng and Diagnostics—Powering Digital Factories, EP/V028618/1).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article, and Literature References provide more details on all topics discussed.

Conflicts of Interest

The authors declare no conflicts of interest in regard to previous or current capacities. The foregoing funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Lawson, C.E.; Martí, J.M.; Radivojevic, T.; Jonnalagadda, S.V.R.; Gentz, R.; Hillson, N.J.; Peisert, S.; Kim, J.; Simmons, B.A.; Petzold, C.J.; et al. Machine learning for metabolic engineering: A review. Metab. Eng. 2021, 63, 34–60. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Abidi, F.; Gerogiorgis, D.I. Process Analytical Technologies (PAT) for accurate quantification of monoclonal antibodies (mAbs). J. Chem. Eng. Jpn. 2026, 59, 2570711. [Google Scholar] [CrossRef] [Scilit]
  3. Pham, T.D.; Manapragada, C.; Sun, Y.; Basset, R.; Aickelin, U. A scoping review of supervised learning modelling and data-driven optimisation in monoclonal antibody process development. Digit. Chem. Eng. 2023, 7, 100080. [Google Scholar] [CrossRef] [Scilit]
  4. Rolandi, P.A. The unreasonable effectiveness of equations: Advanced modeling for biopharmaceutical process development. Comput. Aided Chem. Eng. 2019, 47, 137–150. [Google Scholar] [CrossRef] [Scilit]
  5. Dewasme, L.; Cote, F.; Filee, P.; Hansson, A.; Wouwer, A.V. Macroscopic dynamic modeling of sequential batch cultures of hybridoma cells: An experimental validation. Bioengineering 2017, 4, 17. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Dewasme, L.; Makinen, M.; Chotteau, V. Practical data-driven modeling and robust predictive control of mammalian cell fed-batch process. Comput. Chem. Eng. 2023, 171, 108164. [Google Scholar] [CrossRef] [Scilit]
  7. Kotidis, P.; Kontoravdi, C. Harnessing the potential of artificial neural networks for predicting protein glycosylation. Metab. Eng. Commun. 2020, 10, e00131. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Kotidis, P.; Jedrzejewski, P.; Sou, S.N.; Sellick, C.; Polizzi, K.; Del Val, I.J.; Kontoravdi, C. Model-based optimization of antibody galactosylation in CHO cell culture. Biotechnol. Bioeng. 2019, 116, 1612–1626. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Zhang, D.; Del Rio-Chanona, E.A.; Petsagkourakis, P.; Wagner, J. Hybrid physics-based and data-driven modeling for bioprocess online simulation & optimization. Biotechnol. Bioeng. 2019, 116, 2919–2930. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Rathore, A.S.; Nikita, S.; Thakur, G. Artificial intelligence and machine learning applications in biopharmaceutical manufacturing. Trends Biotechnol. 2023, 41, 497–510. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Pinto, J.; Ramos, J.R.C.; Costa, R.S.; Rossell, S.; Dumas, P.; Oliveira, R. Hybrid deep modelling of a CHO-K1 fed-batch process: Combining first-principles with deep neural networks. Front. Bioeng. Biotechnol. 2023, 11, 1237963. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Das, P.K.; Sahoo, A.; Veeranki, V.D. Modeling and optimization of recombinant Tocilizumab production from Pichia pastoris using response surface methodology and artificial neural network. Biotechnol. Bioeng. 2024, 122, 2093–2110. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Fu, P.C.; Barford, J.P. A hybrid neural network—First principles approach for modelling of cell metabolism. Comput. Chem. Eng. 1996, 20, 951–958. [Google Scholar] [CrossRef] [Scilit]
  14. Quantrille, T.Ε.; Liu, Y.A. Artificial Intelligence in Chemical Engineering; Academic Press: San Diego, CA, USA, 1991; Volume 374. [Google Scholar]
  15. Antonakoudis, A.; Strain, B.; Barbosa, R.; Jimenez del Val, I.; Kontoravdi, C. Synergising stoichiometric modelling with artificial neural networks to predict antibody glycosylation patterns in Chinese hamster ovary cells. Comput. Chem. Eng. 2021, 154, 107471. [Google Scholar] [CrossRef] [Scilit]
  16. Gutierrez, J.M.; Feizi, A.; Li, S.; Kallehauge, T.B.; Hefzi, H.; Grav, L.M.; Ley, D.; Baycin Hizal, D.; Betenbaugh, M.J.; Voldborg, B.; et al. Genome-scale reconstructions of the mammalian secretory pathway predict metabolic costs and limitations of protein secretion. Nat. Commun. 2020, 11, 68. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Botton, A.; Barberi, G.; Facco, P. Data augmentation to support biopharmaceutical process development through digital models—A proof of concept. Processes 2022, 10, 1796. [Google Scholar] [CrossRef] [Scilit]
  18. Buşoniu, L.; de Bruin, T.; Tolić, D.; Kober, J.; Palunko, I. Reinforcement learning for control: Performance, stability, and deep approximators. Annu. Rev. Control 2018, 46, 8–28. [Google Scholar] [CrossRef] [Scilit]
  19. Mowbray, M.; Smith, R.; Del Rio-Chanona, E.A.; Zhang, D. Using process data to generate an optimal control policy via apprenticeship and reinforcement learning. AIChE J. 2021, 67, e17306. [Google Scholar] [CrossRef] [Scilit]
  20. Gadkar, K.; Mehra, S.; Gomes, J. On-line adaptation of neural networks for bioprocess control. Comput. Chem. Eng. 2005, 29, 1047–1057. [Google Scholar] [CrossRef] [Scilit]
  21. Lucia, S.; Finkler, T.; Engell, S. Multi-stage nonlinear model predictive control applied to a semi-batch polymerization reactor under uncertainty. J. Process Control 2013, 23, 1306–1319. [Google Scholar] [CrossRef] [Scilit]
  22. Manapragada, C.; Pham, T.D.; Rajan, N.; Aickelin, U. Pharmaceutical process optimisation: Decision support under high uncertainty. Comput. Chem. Eng. 2023, 170, 108100. [Google Scholar] [CrossRef] [Scilit]
  23. Hashizume, T.; Ozawa, Y.; Ying, B.W. Employing active learning in the optimization of culture medium for mammalian cells. npj Syst. Biol. Appl. 2023, 9, 20. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Baako, T.-M.D.; Kulkarni, S.K.; McClendon, J.L.; Harcum, S.W.; Gilmore, J. Machine learning and deep learning strategies for Chinese hamster ovary cell bioprocess optimization. Fermentation 2024, 10, 234. [Google Scholar] [CrossRef] [Scilit]
  25. Alam, M.N.; Anurag, A.; Gangwar, N.; Ramteke, M.; Kodamana, H.; Rathore, A.S. Physics-informed neural networks guided modelling and multi-objective optimization of a mAb production process. Can. J. Chem. Eng. 2025, 103, 1319–1334. [Google Scholar] [CrossRef] [Scilit]
  26. Ranbhor, R. Advancing monoclonal antibody manufacturing: Process optimization, cost reduction strategies, and emerging technologies. Biol. Targets Ther. 2025, 19, 177–187. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Facco, P.; Zomer, S.; Rowland-Jones, R.C.; Marsh, D.; Diaz-Fernandez, P.; Finka, G.; Bezzo, F.; Barolo, M. Using data analytics to accelerate biopharmaceutical process scale-up. Biochem. Eng. J. 2020, 164, 107791. [Google Scholar] [CrossRef] [Scilit]
  28. Alavijeh, M.K.; Lee, Y.Y.; Gras, S.L. A perspective-driven and technical evaluation of machine learning in bioreactor scale-up: A case-study for potential model developments. Eng. Life Sci. 2024, 24, e2400023. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Le, L.M.M.; Kegl, B.; Gramfort, A.; Marini, C. Optimization of classification and regression analysis of four monoclonal antibodies from Raman spectra using collaborative machine learning approach. Talanta 2018, 184, 260–265. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Pedregosa, G.; Varoquaux, A.; Gramfort, V.; Michel, B.; Thirion, O.; Grisel, M.; Blondel, P.; Prettenhofer, R.; Weiss, V.; Dubourg, J.; et al. Duchesnay, Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
  31. Schölkopf, B.; Smola, A.; Müller, K. Nonlinear component analysis as a kernel eigenvalue problem. Neural Comput. 1998, 10, 1299–1319. [Google Scholar] [CrossRef] [Scilit]
  32. Cristianini, N.; Shawe-Taylor, J. An Introduction to Support Vector Machines and Other Kernel-based Learning Methods; Cambridge University Press: Cambridge, UK, 2000. [Google Scholar]
  33. Nitika, N.; Keerthiveena, B.; Thakur, G.; Rathore, A.S. Convolutional neural networks guided Raman spectroscopy as a process analytical technology (PAT) tool for monitoring and simultaneous prediction of monoclonal antibody charge variants. Pharm. Res. 2024, 41, 463–479. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Mel, M.; Saad, M.M.; Hashim, Y.Z.H.; Salleh, M.R.M. Monoclonal antibody production: Media optimization for enhancement the cell viability of hybridoma cell. Asian J. Sci. Res. 2008, 1, 525–531. [Google Scholar] [CrossRef] [Scilit][Green Version]
  35. Ibrahim, S.; Abdul Wahab, N. Improved artificial neural network training based on response surface methodology for membrane flux prediction. Membranes 2022, 12, 726. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Bashokouh, F.; Abbasilaisi, S.; Tan, J.S. Optimization of cultivation conditions for monoclonal IgM antibody production by M1A2 hybridoma using artificial neural network. Cytotechnology 2019, 71, 849–860. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Sovány, T.; Tislér, Z.; Kristó, K.; Kelemen, A.; Géza Regdon, G. Estimation of design space for an extrusion–spheronization process using response surface methodology and artificial neural network modelling. Eur. J. Pharm. Biopharm. 2016, 106, 79–87. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Ishida, T.; Ichihara, M.; Wang, X.; Yamamoto, K.; Kimura, J.; Majima, E.; Kiwada, H. Injection of PEGylated liposomes in rats elicits PEG specific IgM, which is responsible for rapid elimination of a second dose of PEGylated liposomes. J. Control Release 2006, 112, 15–25. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Chong, S.L.; Mou, D.G.; Ali, A.M.; Lim, S.H.; Tey, B.T. Cell growth, cell-cycle progress, and antibody production in hybridoma cells cultivated under mild hypothermic conditions. Hybridoma 2008, 27, 107–111. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  40. Alam, M.N.; Anupa, A.; Kodamana, H.; Rathore, A.S. A deep learning-aided multi-objective optimization of a downstream process for production of monoclonal antibody products. Biochem. Eng. J. 2024, 208, 109357. [Google Scholar] [CrossRef] [Scilit]
  41. George, E.; Farid, S. Strategic biopharmaceutical portfolio development: An analysis of constraint-induced implications. Biotechnol. Prog. 2008, 24, 698–713. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. George, E.; Farid, S. Combinatorial optimisation algorithms for strategic biopharmaceutical portfolio & capacity management. Comput.-Aided Chem. Eng. 2009, 26, 1063–1068. [Google Scholar] [CrossRef] [Scilit]
  43. Liu, S.; Simaria, A.S.; Farid, S.S.; Papageorgiou, L. Mixed integer optimisation of antibody purification processes. Comput.-Aided Chem. Eng. 2013, 32, 157–162. [Google Scholar] [CrossRef] [Scilit]
  44. Jones, W.; Gerogiorgis, D.I. Plantwide optimisation of monoclonal antibody (mab) manufacturing platforms via Mixed Integer Nonlinear Programming (MINLP). R. Soc. Open Sci. 2026, in press. [Google Scholar]
  45. Nikita, S.; Thakur, G.; Jesubalan, N.G.; Kulkarni, A.; Yezhuvath, V.B.; Rathore, A.S. AI-ML applications in bioprocessing: ML as an enabler of real time quality prediction in continuous manufacturing of mAbs. Comput. Chem. Eng. 2022, 164, 107896. [Google Scholar] [CrossRef] [Scilit]
  46. Park, S.Y.; Kim, S.J.; Park, C.H.; Kim, J.; Lee, D.Y. Data-driven prediction models for forecasting multistep ahead profiles of mammalian cell culture toward bioprocess digital twins. Biotechnol. Bioeng. 2023, 119, 3596–3611. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Thakur, G.; Nikita, S.; Tiwari, A.; Rathore, A.S. Control of surge tanks for continuous manufacturing of monoclonal antibodies. Biotechnol. Bioeng. 2021, 118, 1913–1931. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Tulsyan, A.; Garvin, C.; Undey, C. Industrial batch process monitoring with limited data. J. Process Control 2019, 77, 114–133. [Google Scholar] [CrossRef] [Scilit]
  49. Jin, Y.; Qin, S.J.; Saucedo, V.; Meier, A.; Kunda, S. Process variability source analysis for a multi-step bioprocess. Comput.-Aided Chem. Eng. 2018, 44, 2497–2502. [Google Scholar] [CrossRef] [Scilit]
  50. Jin, Y.; Qin, S.J.; Huang, Q.; Saucedo, V. Classification and diagnosis of bioprocess cell growth productions using early-stage data. Ind. Eng. Chem. Res. 2019, 58, 13469–13480. [Google Scholar] [CrossRef] [Scilit]
  51. Jesubalan, N.G.; Thakur, G.; Rathore, A.S. Deep neural network for prediction and control of permeability decline in single pass tangential flow ultrafiltration in continuous processing of monoclonal antibodies. Front. Chem. Eng. 2023, 5, 1182817. [Google Scholar] [CrossRef] [Scilit]
  52. Walsh, I.; Myint, M.; Nguyen-Khuong, T.; Ho, S.K.N.; Lakshmanan, M. Harnessing the potential of machine learning for advancing “Quality by Design” in biomanufacturing. mAbs 2022, 14, 2013593. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Wu, I.E.; Kalejaye, L.; Lai, P.K. Machine learning models for predicting monoclonal antibody biophysical properties from molecular dynamics simulations and deep learning-based surface descriptors. Mol. Pharm. 2025, 22, 142–153. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Lai, P.K.; Fernando, A.; Cloutier, T.K.; Gokarn, Y.; Zhang, J.F.; Schwenger, W.; Chari, R.; Calero-Rubio, C.; Trout, B.L. Machine learning applied to determine the molecular descriptors responsible for the viscosity behavior of concentrated therapeutic antibodies. Mol. Pharm. 2021, 18, 1167–1175. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Lai, P.; Gallegos, A.; Mody, N.; Sathish, H.A.; Trout, B.L. Machine learning prediction of antibody aggregation and viscosity for high concentration formulation development of protein therapeutics. mAbs 2022, 14, e2026208. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  56. Li, L.; Kumar, S.; Buck, P.M.; Burns, C.; Lavoie, J.; Singh, S.K.; Warne, N.W.; Nichols, P.; Luksha, N.; Boardman, D. Concentration dependent viscosity of monoclonal antibody solutions: Explaining experimental behavior in terms of molecular properties. Pharm. Res. 2014, 31, 3161−3178. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Sharma, V.K.; Patapoff, T.W.; Kabakoff, B.; Pai, S.; Hilario, E.; Zhang, B.; Li, C.; Borisov, O.; Kelley, R.F.; Chorny, I.; et al. In silico selection of therapeutic antibodies for development: Viscosity, clearance, and chemical stability. Proc. Natl. Acad. Sci. USA 2014, 111, 18601−18606. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  58. Tomar, D.S.; Li, L.; Broulidakis, M.P.; Luksha, N.G.; Burns, C.T.; Singh, S.K.; Kumar, S. In-silico prediction of concentration- dependent viscosity curves for monoclonal antibody solutions. mAbs 2017, 9, 476−489. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  59. Schmitt, J.; Razvi, A.; Grapentin, C. Predictive modeling of concentration-dependent viscosity behavior of monoclonal antibody solutions using artificial neural networks. mAbs 2023, 15, 2169440. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Makowski, E.K.; Chen, H.T.; Wang, T.X.; Wu, L.N.; Huang, J.; Mock, M.; Underhill, P.; Pelegri-O’Day, E.; Maglalang, E.; Winters, D. Reduction of monoclonal antibody viscosity using interpretable machine learning. mAbs 2024, 16, 2303781. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. Kalejaye, L.A.; Chu, J.M.; Wu, I.E.; Amofah, B.; Lee, A.M.; Hutchinson, M.; Chakiath, C.; Dippel, A.; Kaplan, G.; Damschroder, M. Accelerating high-concentration monoclonal antibody development with large-scale viscosity data and ensemble deep learning. mAbs 2025, 17, 2483944. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  62. Harmalkar, A.; Rao, R.; Xie, Y.; Honer, J.; Gray, J.; Wei, K. Toward generalizable prediction of antibody thermostability using machine learning on sequence and structure features. mAbs 2023, 15, e2163584. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  63. Koenig, P.; Lee, C.V.; Walters, B.T.; Janakiraman, V.; Stinson, J.; Patapoff, T.W.; Fuh, G. Mutational landscape of antibody variable domains reveals a switch modulating the interdomain conformational dynamics and antigen binding. Proc. Natl. Acad. Sci. USA 2017, 114, E486–E495. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  64. Warszawski, S.; Katz, A.B.; Lipsh, R.; Khmelnitsky, L.; Nissan, G.B.; Javitt, G.; Dym, O.; Unger, T.; Knop, O.; Albeck, S. Optimizing antibody affinity and stability by the automated design of the variable light-heavy chain interfaces. PLoS Comput. Biol. 2019, 15, 1–24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  65. Alford, R.F.; Leaver-Fay, A.; Jeliazkov, J.R.; O’Meara, M.J.; DiMaio, F.P.; Park, H.; Shapovalov, M.V.; Renfrew, P.D.; Mulligan, V.K.; Kappel, K. The Rosetta all-atom energy function for macromolecular modeling and design. J. Chem. Theory Comput. 2017, 13, 3031–3048. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  66. Greenblott, D.N.; Zhang, J.D.; Calderon, C.P.; Randolph, T.W. Machine learning approaches to root cause analysis, characterization, and monitoring of subvisible particles in monoclonal antibody formulations. Biotechnol. Bioeng. 2022, 119, 3596–3611. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  67. Calderon, C.P.; Daniels, A.L.; Randolph, T.W. Deep convolutional neural network analysis of flow imaging microscopy data to classify subvisible particles in protein formulations. J. Pharm. Sci. 2018, 107, 999–1008. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  68. Shrivastava, A.; Mandal, S.; Pattanayak, S.; Rathore, A. Rapid estimation of size-based heterogeneity in monoclonal antibodies by machine learning-enhanced dynamic light scattering. Anal. Chem. 2023, 95, 8299−8309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. Zidar, M.; Šušterič, A.; Ravnik, M.; Kuzman, D. High throughput prediction approach for monoclonal antibody aggregation at high concentration. Pharm. Res. 2017, 34, 1831–1839. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  70. Gentiluomo, L.; Roessner, D.; Augustijn, D.; Svilenova, H.; Kulakova, A. Application of interpretable artificial neural networks to early monoclonal antibodies development. Eur. J. Pharm. Biopharm. 2019, 141, 81–89. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. del Rio-Chanona, E.A.; Cong, X.Y.; Bradford, E.; Zhang, D.D.; Jing, K.J. Review of advanced physical and data-driven models for dynamic bioprocess simulation: Case study of algae-bacteria consortium wastewater treatment. Biotechnol. Bioeng. 2019, 116, 342–353. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  72. Budholiya, N.; Roy, S.; Rathore, A.S. Neural network-based fingerprinting of monoclonal antibody aggregation using biolayer interferometry. Anal. Bioanal. Chem. 2019, 412, 2177–2186. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  73. Gerogiorgis, D.I.; Barton, P.I. Steady-state optimization of a continuous pharmaceutical process. Comput.-Aided Chem. Eng. 2009, 7, 927–932. [Google Scholar] [CrossRef] [Scilit]
  74. Delmar, J.A.; Buehler, E.; Chetty, A.K.; Das, A.; Quesada, G.M.; Wang, J.; Chen, X. Machine learning prediction of methionine and tryptophan photooxidation susceptibility. Mol. Ther. Methods Clin. Dev. 2021, 21, 466–477. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  75. Knez, B.; Erzin, L.; Kos, Ž.; Kuzman, D.; Ravnik, M. Prediction of aggregation in monoclonal antibodies from molecular surface curvature. Sci. Rep. 2025, 15, 28266. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  76. Fuh, G.; Wu, P.; Liang, W.C.; Ultsch, M.; Lee, C.V.; Moffat, B.; Wiesmann, C. Structure-function studies of two synthetic anti-vascular endothelial growth factor Fabs and comparison with the Avastin Fab. J. Biol. Chem. 2006, 281, 6625–6631. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  77. Shapovalov, M.V.; Dunbrack, R.L.J. A smoothed backbone-dependent rotamer library for proteins derived from adaptive kernel density estimates and regressions. Structure 2011, 19, 844–858. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. An overview of approaches for biopharma modelling and optimisation, drawn from [3].
Figure 1. An overview of approaches for biopharma modelling and optimisation, drawn from [3].
Molecules 31 03313 g001
Figure 2. Experimental datasets for sequential batch culture using HB1 hybridoma cells [5].
Figure 2. Experimental datasets for sequential batch culture using HB1 hybridoma cells [5].
Molecules 31 03313 g002
Figure 3. Experimental datasets for sequential batch culture using HB2 hybridoma cells [5].
Figure 3. Experimental datasets for sequential batch culture using HB2 hybridoma cells [5].
Molecules 31 03313 g003
Figure 4. Direct validation of the derived MLPCA model vs. exptl. data for the HB1 strain [5].
Figure 4. Direct validation of the derived MLPCA model vs. exptl. data for the HB1 strain [5].
Molecules 31 03313 g004
Figure 5. Experiments with low (1–3) and high (4–5) amino acid levels in culture medium [6].
Figure 5. Experiments with low (1–3) and high (4–5) amino acid levels in culture medium [6].
Molecules 31 03313 g005
Figure 6. Log-likelihood (JP) of an m-dimensional subspace (relative noise variance: 10%) [6].
Figure 6. Log-likelihood (JP) of an m-dimensional subspace (relative noise variance: 10%) [6].
Molecules 31 03313 g006
Figure 7. Direct validation of the MLPCA model for three distinct η levels vs. Expt. 1 [6].
Figure 7. Direct validation of the MLPCA model for three distinct η levels vs. Expt. 1 [6].
Molecules 31 03313 g007
Figure 8. Direct validation of the MLPCA model for three distinct η levels vs. Expt. 2 [6].
Figure 8. Direct validation of the MLPCA model for three distinct η levels vs. Expt. 2 [6].
Molecules 31 03313 g008
Figure 9. HyGlycoM combines kinetic models with an ANN for glycosylation; adapted from [7].
Figure 9. HyGlycoM combines kinetic models with an ANN for glycosylation; adapted from [7].
Molecules 31 03313 g009
Figure 10. HyGlycoM model NSD estimates vs. exptl. data for 6 nucleotide sugars, data from [7].
Figure 10. HyGlycoM model NSD estimates vs. exptl. data for 6 nucleotide sugars, data from [7].
Molecules 31 03313 g010
Figure 11. Hybrid model architecture combining first principles, ANN and experimental data [13].
Figure 11. Hybrid model architecture combining first principles, ANN and experimental data [13].
Molecules 31 03313 g011
Figure 12. The detailed RL ANN architecture implemented in [13] as first proposed in [14].
Figure 12. The detailed RL ANN architecture implemented in [13] as first proposed in [14].
Molecules 31 03313 g012
Figure 13. Predictions of (a) first-principles and (b) hybrid model: cell density and mAb [13].
Figure 13. Predictions of (a) first-principles and (b) hybrid model: cell density and mAb [13].
Molecules 31 03313 g013
Figure 14. Predictions of (a) first-principles and (b) hybrid model: key cell metabolites [13].
Figure 14. Predictions of (a) first-principles and (b) hybrid model: key cell metabolites [13].
Molecules 31 03313 g014
Figure 15. Hybrid modelling via metabolic kinetics, fluxes and NN for glycan prediction, from [15].
Figure 15. Hybrid modelling via metabolic kinetics, fluxes and NN for glycan prediction, from [15].
Molecules 31 03313 g015
Figure 16. Biomass growth rate (A) and specific mAb productivity (B) FBA predictions, from [15].
Figure 16. Biomass growth rate (A) and specific mAb productivity (B) FBA predictions, from [15].
Molecules 31 03313 g016
Figure 17. Hybrid digital model (HDM) block structure used to generate in silico batches [17].
Figure 17. Hybrid digital model (HDM) block structure used to generate in silico batches [17].
Molecules 31 03313 g017
Figure 18. HDM training (solid red) and simulated (dashed grey) mAb titer profiles [17].
Figure 18. HDM training (solid red) and simulated (dashed grey) mAb titer profiles [17].
Molecules 31 03313 g018
Figure 19. Schematic diagram of the proposed NN architecture with feedback action [20].
Figure 19. Schematic diagram of the proposed NN architecture with feedback action [20].
Molecules 31 03313 g019
Figure 20. NN training (left) and profiles with/without online weight adaptation (right) [20].
Figure 20. NN training (left) and profiles with/without online weight adaptation (right) [20].
Molecules 31 03313 g020
Figure 21. Training dataset (left); NN-driven online control signals vs. exptl. data (right) [20].
Figure 21. Training dataset (left); NN-driven online control signals vs. exptl. data (right) [20].
Molecules 31 03313 g021
Figure 22. Bibliometric analysis of SLDO studies vs. biopharma publications, 1994–2022 [3].
Figure 22. Bibliometric analysis of SLDO studies vs. biopharma publications, 1994–2022 [3].
Molecules 31 03313 g022
Figure 23. The Naive (left) and ITO (right) ANN model architectures, both adapted from [22].
Figure 23. The Naive (left) and ITO (right) ANN model architectures, both adapted from [22].
Molecules 31 03313 g023
Figure 24. Predicted vs. actual mAb yield for uniform and non-uniform feeding cases, data by [22].
Figure 24. Predicted vs. actual mAb yield for uniform and non-uniform feeding cases, data by [22].
Molecules 31 03313 g024
Figure 25. A comprehensive taxonomy of AI/ML methods vs. learning types, adapted from [10].
Figure 25. A comprehensive taxonomy of AI/ML methods vs. learning types, adapted from [10].
Molecules 31 03313 g025
Figure 26. mAb production maximisation vs. weight variation: α = 0 (left), α = 10 (right), as per [5].
Figure 26. mAb production maximisation vs. weight variation: α = 0 (left), α = 10 (right), as per [5].
Molecules 31 03313 g026
Figure 27. State trajectory comparisons: MSNMPC vs. NMPC optimisation algorithms, as per [6].
Figure 27. State trajectory comparisons: MSNMPC vs. NMPC optimisation algorithms, as per [6].
Molecules 31 03313 g027
Figure 28. Feed rate trajectory comparison: MSNMPC vs. NMPC optimisation algorithms [6].
Figure 28. Feed rate trajectory comparison: MSNMPC vs. NMPC optimisation algorithms [6].
Molecules 31 03313 g028
Figure 29. Mean raw Raman spectra for the four mAbs studied (400–4000 cm−1); adapted from [29].
Figure 29. Mean raw Raman spectra for the four mAbs studied (400–4000 cm−1); adapted from [29].
Molecules 31 03313 g029
Figure 30. Effect of FBS concentration on mAb production in DMEM (A) vs. RPMI 1640 (B) [36].
Figure 30. Effect of FBS concentration on mAb production in DMEM (A) vs. RPMI 1640 (B) [36].
Molecules 31 03313 g030
Figure 31. Effect of temperature on mAb production in DMEM (A) and RPMI 1640 (B) [36].
Figure 31. Effect of temperature on mAb production in DMEM (A) and RPMI 1640 (B) [36].
Molecules 31 03313 g031
Figure 32. Effect of cultivation time on mAb production in DMEM (A) vs. RPMI 1640 (B) [36].
Figure 32. Effect of cultivation time on mAb production in DMEM (A) vs. RPMI 1640 (B) [36].
Molecules 31 03313 g032
Figure 33. Bivariate mAb production: (A,D) FBS concentration + cultivation time; (B,E) FBS concentration + T (°C); (C,F) cultivation time + T (°C) (media: DMEM/top; RPMI 1640/bottom) [36].
Figure 33. Bivariate mAb production: (A,D) FBS concentration + cultivation time; (B,E) FBS concentration + T (°C); (C,F) cultivation time + T (°C) (media: DMEM/top; RPMI 1640/bottom) [36].
Molecules 31 03313 g033
Figure 34. The complete biopharma simulation and optimisation framework, adapted from [41].
Figure 34. The complete biopharma simulation and optimisation framework, adapted from [41].
Molecules 31 03313 g034
Figure 35. Mean NPV vs. risk for a 5-drug portfolio and 4 cash flow constraints, adapted from [41].
Figure 35. Mean NPV vs. risk for a 5-drug portfolio and 4 cash flow constraints, adapted from [41].
Molecules 31 03313 g035
Figure 36. Sparse block-wise learning, adapted from the GP-SSM generator model shown in [48].
Figure 36. Sparse block-wise learning, adapted from the GP-SSM generator model shown in [48].
Molecules 31 03313 g036
Figure 37. Hotelling T2 plots for: new batch, with 4 components and 3 historical batches (left), and new batch, again with 4 components but combining 50 in silico and 3 historical batches (right) [48].
Figure 37. Hotelling T2 plots for: new batch, with 4 components and 3 historical batches (left), and new batch, again with 4 components but combining 50 in silico and 3 historical batches (right) [48].
Molecules 31 03313 g037
Figure 38. PCA monitoring results in principal space (left) and in residual space (right) as per [49].
Figure 38. PCA monitoring results in principal space (left) and in residual space (right) as per [49].
Molecules 31 03313 g038
Figure 39. Classified high/low (left) and unclassified intermediate (right) lactate profiles from [50].
Figure 39. Classified high/low (left) and unclassified intermediate (right) lactate profiles from [50].
Molecules 31 03313 g039
Figure 40. T2 (left) and SPE (right) plots for (top) 70%, 80%, 85% and 90% CPV, adapted from [50].
Figure 40. T2 (left) and SPE (right) plots for (top) 70%, 80%, 85% and 90% CPV, adapted from [50].
Molecules 31 03313 g040
Figure 41. A process diagram for downstream continuous mAb manufacturing, adapted from [45].
Figure 41. A process diagram for downstream continuous mAb manufacturing, adapted from [45].
Molecules 31 03313 g041
Figure 42. Random Forest Regression (RFR) for ML based response prediction, adapted from [45].
Figure 42. Random Forest Regression (RFR) for ML based response prediction, adapted from [45].
Molecules 31 03313 g042
Figure 43. Gradient Boosting Regression (GBR) for ML-based trend prediction, adapted from [45].
Figure 43. Gradient Boosting Regression (GBR) for ML-based trend prediction, adapted from [45].
Molecules 31 03313 g043
Figure 44. Comparison of R2 values for key observables and different ML approaches; data by [45].
Figure 44. Comparison of R2 values for key observables and different ML approaches; data by [45].
Molecules 31 03313 g044
Figure 45. AI/ML performance for eluted mAb prediction in protein A chromatography, from [45].
Figure 45. AI/ML performance for eluted mAb prediction in protein A chromatography, from [45].
Molecules 31 03313 g045
Figure 46. Bioprocess scale-up for mAb-producing cell line generation: units and data, from [27].
Figure 46. Bioprocess scale-up for mAb-producing cell line generation: units and data, from [27].
Molecules 31 03313 g046
Figure 47. MPCA (a) time-cumulated absolute loadings, and (b) key state variable profiles, by [27].
Figure 47. MPCA (a) time-cumulated absolute loadings, and (b) key state variable profiles, by [27].
Molecules 31 03313 g047
Figure 48. Viable cell estimation error: (a) 6-well scale (best), (b) late T25 scale (worst), from [27].
Figure 48. Viable cell estimation error: (a) 6-well scale (best), (b) late T25 scale (worst), from [27].
Molecules 31 03313 g048
Figure 49. Estimation error for: (a) IVC and (b,c) Peak VCC estimation at shake-flask scale, for which (b) all the calibration clones and (c) only 5 calibration clones are used for computations - from [27].
Figure 49. Estimation error for: (a) IVC and (b,c) Peak VCC estimation at shake-flask scale, for which (b) all the calibration clones and (c) only 5 calibration clones are used for computations - from [27].
Molecules 31 03313 g049
Figure 50. A superstructure describing ML modelling and multiple QbD applications, from [52].
Figure 50. A superstructure describing ML modelling and multiple QbD applications, from [52].
Molecules 31 03313 g050
Figure 51. Comparison of mAb viscosity (η) predictions of two models (A,B) vs. exptl. data, from [54].
Figure 51. Comparison of mAb viscosity (η) predictions of two models (A,B) vs. exptl. data, from [54].
Molecules 31 03313 g051
Figure 52. Relationship between viscosity and top three (max. AUPRC) mAb features, data by [54].
Figure 52. Relationship between viscosity and top three (max. AUPRC) mAb features, data by [54].
Molecules 31 03313 g052
Figure 53. Correlation coefficients for the best two-feature linear, SVR and kNN models [55].
Figure 53. Correlation coefficients for the best two-feature linear, SVR and kNN models [55].
Molecules 31 03313 g053
Figure 54. Thermostability (a) prediction, (b) data generation, (c) DL training for classification [62].
Figure 54. Thermostability (a) prediction, (b) data generation, (c) DL training for classification [62].
Molecules 31 03313 g054
Figure 55. Fine-tuning improves the accuracy of ML models (right) vs. zero-shot ones (left) [62].
Figure 55. Fine-tuning improves the accuracy of ML models (right) vs. zero-shot ones (left) [62].
Molecules 31 03313 g055
Figure 56. ROC curve demonstrating classification of test sequences for the AUC > 70 bin, by [62].
Figure 56. ROC curve demonstrating classification of test sequences for the AUC > 70 bin, by [62].
Molecules 31 03313 g056
Figure 57. Computational deep mutational scan of mAb fragment agrees with exptl. data [62].
Figure 57. Computational deep mutational scan of mAb fragment agrees with exptl. data [62].
Molecules 31 03313 g057
Figure 58. A collection of FIM particle images subjected to six stresses (Table 9), adapted from [66].
Figure 58. A collection of FIM particle images subjected to six stresses (Table 9), adapted from [66].
Molecules 31 03313 g058
Figure 59. (ad) Training and test clouds from a CNN trained on 16,000 images and 2 stresses, from [66].
Figure 59. (ad) Training and test clouds from a CNN trained on 16,000 images and 2 stresses, from [66].
Molecules 31 03313 g059
Figure 60. Test embedding fingerprint from an approximate 50–50% particle mixture, from [66].
Figure 60. Test embedding fingerprint from an approximate 50–50% particle mixture, from [66].
Molecules 31 03313 g060
Figure 61. (A) SEC chromatogram, (B) DLS-PSD oligomer distribution analysis curves, from [68].
Figure 61. (A) SEC chromatogram, (B) DLS-PSD oligomer distribution analysis curves, from [68].
Molecules 31 03313 g061
Figure 62. Predicted vs. exptl. data for mAb1 test and validation ((A,C): NN; (B,D): SVR) [68].
Figure 62. Predicted vs. exptl. data for mAb1 test and validation ((A,C): NN; (B,D): SVR) [68].
Molecules 31 03313 g062
Figure 63. Predicted vs. exptl. data for mAb2 test and validation ((A,C): NN; (B,D): SVR) [68].
Figure 63. Predicted vs. exptl. data for mAb2 test and validation ((A,C): NN; (B,D): SVR) [68].
Molecules 31 03313 g063
Figure 64. The ML implementation used to derive interpretable ANN-based predictions, from [70].
Figure 64. The ML implementation used to derive interpretable ANN-based predictions, from [70].
Molecules 31 03313 g064
Figure 65. PPI-13&3 Tm1 & Tagg prediction performance (black: training, red: validation), from [70].
Figure 65. PPI-13&3 Tm1 & Tagg prediction performance (black: training, red: validation), from [70].
Molecules 31 03313 g065
Table 1. Final MLPCA model kinetic parameter values for rate expressions φ (i) and amino acids [6].
Table 1. Final MLPCA model kinetic parameter values for rate expressions φ (i) and amino acids [6].
ParameterValueUnitParameterValueUnit
φmax (1)17.166d−1KN,s48.147mM
φmax (2)199.847d−1KAsn,s0.135mM
φmax (3)0.0243d−1KArg,s67.467mM
KCys,s0.099mMKAla,inh0.113mM−1
KVal,s0.288mMKMet,inh6.179mM−1
KIle,s0.294mMKPro,inh19.991mM−1
KLeu,s0.298mMKVal,inh0.221mM−1
KLys,s2.805mM(kinetic rate expressions φ (i) detailed in [3])
Table 2. Experimental data-driven model fitting cost function JML residuals in cross-validations [6].
Table 2. Experimental data-driven model fitting cost function JML residuals in cross-validations [6].
ExperimentJML (n = 103)JML (n = 104)
30.4440.553
40.8791.076
50.8400.974
Table 3. ANN performance analysis: error metrics vs. experimental data for different periods [15].
Table 3. ANN performance analysis: error metrics vs. experimental data for different periods [15].
Days0–88–1010–1212+
Training Error (%)0.140.40.270.07
Prediction Error (%)0.110.220.130.03
Table 4. Final yield for each bioreactor and the two feeding regimes, in the validation study [22].
Table 4. Final yield for each bioreactor and the two feeding regimes, in the validation study [22].
ConditionReplicateProduct Concentration (mg.L−1)
Control with non-uniform feeding13635.036
Control with non-uniform feeding22665.541
Optimised feed with non-uniform feeding12825.806
Optimised feed with non-uniform feeding22927.120
Optimised feed with uniform feeding12828.525
Optimised feed with uniform feeding22752.028
Table 5. Misclassification (Rclf), mean abs. relative (Rreg) and combined error (Rcomb) scores [29].
Table 5. Misclassification (Rclf), mean abs. relative (Rreg) and combined error (Rcomb) scores [29].
Linear (Chemometrics)Machine Learning
  MoleculeRclf (%)Rreg (%)Rcomb (%)Rclf (%)Rreg (%)Rcomb (%)
  Infliximab13.712.313.20.98.43.4
  Bevacizumab19.814.017.91.04.32.1
  Ramucirumab9.07.38.40.03.51.2
  Rituximab16.026.719.61.06.93.0
  Overall14.514.714.60.75.82.4
Table 6. RSM-ANN training and testing for IgM mAb production by hybridoma M1A2 cells [36].
Table 6. RSM-ANN training and testing for IgM mAb production by hybridoma M1A2 cells [36].
Dataset Type and
Medium Type
FBS
(%)
Cultivation Time
(day)
Temperature (°C)Experimental mAb (μg/mL)Predicted mAb (μg/mL)
Training data5333673.97673.88
(DMEM)7533719.24719.34
7233638.02638.08
20333821.39821.37
17233509.83509.88
175331220.01219.8
5337529.14529.14
10437603.78607.98
10437610.00605.77
104331066.71033.8
104331006.01039.5
17537600.00600.04
10637600.00599.98
10133411.87411.78
20337529.14529.15
106331097.21097.1
Testing data10437610.00608.00
(DMEM)104331050.01034.0
7237504.31545.62
7537569.24585.93
10133400.00411.72
Training data5333354.06354.06
(RPMI 1640)7533600.00600.00
7233446.90446.90
20333125.29125.28
17233194.82194.81
17537568.18568.17
5337289.52289.52
10437547.45547.45
10433520.00520.00
104371100.01100.0
17537120.00120.00
10637100.00100.00
10133332.01332.00
Testing data10437520.00533.73
(RPMI 1640)104331002.01100.0
7237451.42504.14
7537110.00213.54
10133153.76243.23
Table 7. Summary of ML algorithms and results of optimisation studies, and authoritative reviews.
Table 7. Summary of ML algorithms and results of optimisation studies, and authoritative reviews.
AuthorsYear ML MethodEffectivenessRemarks
Dewasme et al. [5]2017Principal Component
Analysis (PCA), and
MATLAB 2025a (fmincon)
30% increased
production
Final mAb titre:
60.92 µg.mL−1 vs.
40 to 45 µg.mL−1
Le et al. [29]2018Kernel PCA and
Support Vector
Machines (SVM)
97.6%
prediction accuracy
Combined error of 2.4% vs. 14.6%
by a linear approach
Dewasme et al. [6]2023Max. Likelihood Principal Component Analysis (MLPCA) and Multi-Stage Nonlinear Model Predictive Control (MSNMPC)28% increased
production
Final mAb titre:
17.1 mM vs. 12.5 mM by NMPC
Manapragada et al. [22]2023Input–Throughput–
Output (ITO)
dual ANN model
8.9% increased
production
Product titre via
feed optimisation:
2927 vs. 2665 mg.L−1
Pham et al. [3]2023Supervised Learning (SL)
& data-driven optimisation (SLDDO) use in studies
-(Review paper)
Rathore et al. [10]2023A review of several AI/ML
applications in the modern
biopharmaceutical industry
-(Review paper)
Table 8. Performance of both PLS and JY-PLS models in estimating the number of viable cells [27].
Table 8. Performance of both PLS and JY-PLS models in estimating the number of viable cells [27].
Scale (Figure 46)Number of Available ClonesR2PLSR2JY-PLSeJY-PLS < ePLS (%)
24 w2770.990.9835
24 w L1950.960.8828
6 w820.990.9845
6 w L420.940.9540
T25480.960.9552
T25 L170.860.8652
Table 9. Manufacturing stress condition abbreviations relevant to FIM images of Figure 58 [66].
Table 9. Manufacturing stress condition abbreviations relevant to FIM images of Figure 58 [66].
StressAbbreviation
Freeze–thawStress F
AgitationStress A
Pumping in Gore® large diameter tubingStress GL
Pumping in Gore® small diameter tubingStress GS
Pumping in C-flex® ULTRA tubingStress C
pH swingStress P
Table 10. Validation set results for the predicted vs. true sign of interaction parameter kD [70].
Table 10. Validation set results for the predicted vs. true sign of interaction parameter kD [70].
Predicted Sign of kD
Sign of kD Count:NegativePositive
Negative48 + 240
Positive024 + 12
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Abidi, F.; Gerogiorgis, D.I. A Review on Applications of Artificial Intelligence (AI) in Monoclonal Antibody (mAb) Manufacturing. Molecules 2026, 31, 3313. https://doi.org/10.3390/molecules31183313

AMA Style

Abidi F, Gerogiorgis DI. A Review on Applications of Artificial Intelligence (AI) in Monoclonal Antibody (mAb) Manufacturing. Molecules. 2026; 31(18):3313. https://doi.org/10.3390/molecules31183313

Chicago/Turabian Style

Abidi, Fawad, and Dimitrios I. Gerogiorgis. 2026. "A Review on Applications of Artificial Intelligence (AI) in Monoclonal Antibody (mAb) Manufacturing" Molecules 31, no. 18: 3313. https://doi.org/10.3390/molecules31183313

APA Style

Abidi, F., & Gerogiorgis, D. I. (2026). A Review on Applications of Artificial Intelligence (AI) in Monoclonal Antibody (mAb) Manufacturing. Molecules, 31(18), 3313. https://doi.org/10.3390/molecules31183313

Article Metrics

Back to TopTop