Skip to Content
EnergiesEnergies
  • Review
  • Open Access

8 September 2026

Artificial Intelligence and Machine Learning for Bioenergy Generation from Anaerobic Digestion: Data Processing Pipelines, Predictive Models, and Intelligent Control

,
,
,
and
Faculty of Electronics, Telecommunications and Informatics, and Advanced Materials Centre, Gdańsk University of Technology, 80-233 Gdańsk, Poland
*
Author to whom correspondence should be addressed.

Abstract

Anaerobic digestion (AD) plays a central role in renewable energy generation and sustainable waste management. However, the operation of AD systems remains challenging due to the complex interactions among microbial communities, substrate variability, and non-linear process dynamics. Traditional monitoring and control approaches often fail to anticipate disturbances or maintain optimal conditions. Recent advances in artificial intelligence (AI) and machine learning (ML) provide new opportunities to model, predict, and control AD processes by leveraging high-resolution sensor data and data-driven algorithms. This review synthesizes current progress in ML- and AI-based approaches for prediction, optimization, and intelligent control of AD, with particular emphasis on data processing pipelines, neural network architectures, soft sensors, digital twins, and explainable AI.

1. Introduction

European energy policy has changed considerably in recent years under the combined pressure of climate change mitigation and the need to improve energy security. A major milestone was the introduction of the REPowerEU strategy in 2022, which identified domestic renewable energy production as a priority for reducing dependence on imported fossil fuels, particularly natural gas [1,2,3]. Expanding biomethane production from agricultural residues, industrial by-products, municipal wastewater, and other organic waste streams has consequently become an important component of the European energy transition [4,5,6,7,8].
Anaerobic digestion (AD) converts organic matter into biogas through a sequence of oxygen-free microbial processes. After upgrading, the resulting methane-rich gas can be injected into existing energy infrastructure, whereas digestate provides a nutrient-rich material suitable for agricultural use [9,10]. Beyond renewable energy production, AD supports waste valorisation and circular economy objectives [11]. Achieving stable reactor operation, however, remains challenging because process performance depends on tightly coupled biological and physicochemical interactions that continuously evolve over time [9,12,13]. Microbial communities respond to changes in substrate composition, volatile solids content, Carbon-to-Nitrogen (C/N) ratio, alkalinity, volatile fatty acid (VFA) concentrations, temperature, organic loading rate (OLR), and inhibitory compounds. Since these responses are nonlinear and often delayed, relatively small disturbances may eventually reduce methane production or destabilize reactor operation [14,15].
Mechanistic modelling has provided the theoretical foundation for analysing AD processes. Models such as Anaerobic Digestion Model No. 1 (ADM1) describe biochemical conversion pathways in considerable detail, but their practical application requires extensive parameter calibration and detailed substrate characterization [16]. Such information is rarely available during routine operation of industrial digesters, which limits the direct applicability of purely mechanistic approaches [17,18]. Similar limitations affect conventional control strategies that are commonly based on simplified process assumptions and therefore respond to disturbances only after measurable changes have already occurred [19,20].
The rapid development of sensor technologies and automated monitoring systems has substantially increased the amount of operational data collected in modern AD plants [21,22]. High-frequency measurements describing gas production, substrate characteristics, and reactor operating conditions are now routinely available [22,23,24]. Their interpretation, however, is complicated by nonlinear process dynamics, measurement noise, and strong interactions among process variables. Data-driven methods based on artificial intelligence (AI) and machine learning (ML) offer an effective way to extract useful information from these datasets without requiring explicit mathematical descriptions of every biochemical relationship [22,25,26,27,28].
Recent studies have shown that neural networks and ensemble learning algorithms frequently outperform conventional statistical approaches when predicting methane production and other process indicators characterized by nonlinear behaviour [29,30]. Artificial neural networks (ANNs), recurrent architectures, tree-based ensemble methods, and hybrid modelling frameworks have been successfully applied to biogas production forecasting, soft sensing of difficult-to-measure variables, anomaly detection, and process diagnostics [31,32]. More recently, these techniques have also become components of digital twins (DTs) and advanced control systems capable of supporting increasingly autonomous operation of biogas plants [33,34,35].
Despite this rapid progress, research on AI applications in AD remains dispersed across several disciplines. Existing reviews typically focus on sensor technologies, individual ML algorithms, or selected control strategies, whereas considerably less attention has been devoted to the complete workflow linking data acquisition, preprocessing, feature engineering, predictive modelling, and operational decision support. As AI-based solutions mature, understanding these connections becomes increasingly important for successful industrial implementation.
Although several review articles have summarized the application of AI and ML in AD, including the systematic review by Rutland et al. [30], the recent overview by Xu et al. [27], and our previous review [22], there is still a need for reviews that integrate these complementary aspects into a unified AI framework for AD. Existing reviews generally emphasize selected aspects of AI implementation, such as ML algorithms, sensor integration, or process optimization. In particular, our previous review [22] examined sensing technologies, monitoring strategies, and the integration of sensor data with AI-based methods for AD process monitoring. In contrast, the present review adopts a data-centric and AI-centred perspective, focusing on how operational data are transformed into actionable knowledge through preprocessing, feature engineering, predictive modelling, model validation, deployment, explainable AI, intelligent process control, and decision-support strategies. Rather than concentrating on sensing technologies themselves, this review emphasizes the complete AI workflow and its role in enabling intelligent and increasingly autonomous management of AD systems. Consequently, the two reviews are complementary rather than overlapping, addressing different stages of the AI implementation pipeline in AD. By integrating these complementary aspects into a single framework, this review aims to provide researchers and practitioners with a comprehensive roadmap for the development of next-generation intelligent AD systems.
To address these knowledge gaps, this review examines AI and ML applications in AD from a data-centred perspective. Rather than reviewing individual algorithms in isolation, it follows the complete pathway from operational data acquisition and preprocessing to predictive modelling, soft sensing, intelligent control, and decision support. Attention is given to neural network architectures, feature engineering, explainable AI, and advanced control strategies. Sensor technologies are discussed only briefly because they have been comprehensively reviewed elsewhere, while the main emphasis is placed on methods that transform process data into actionable information for monitoring, optimization, and control of AD systems.
In the context of AD, bioenergy generation is closely linked to the stable and efficient production of biogas and methane, which determines the energy potential of the process. The development of intelligent modelling and control strategies can support this objective by improving process stability and raw biogas production, thereby providing an important basis for subsequent biogas upgrading to biomethane. Accordingly, this review focuses on how AI and ML can support data-driven prediction, monitoring, optimization, and control of AD processes. Detailed techno-economic assessment of energy recovery, downstream biomethane upgrading technologies, and energy conversion systems are beyond the scope of this review.
To ensure a comprehensive overview of the topic, the literature selection for this review targeted peer-reviewed articles, conference proceedings, and technical reports retrieved from major scientific databases, primarily Scopus and Web of Science. The search strategy focused on the intersection of AI, ML, and AD. Rather than applying a rigid chronological restriction, emphasis was placed on recent advancements, predominantly from the last decade, to reflect the rapid evolution of data-driven techniques in bioprocess engineering. Studies were included based on their thematic relevance to data processing pipelines, predictive modelling, soft sensor applications, optimization, and intelligent control strategies, ensuring a focused synthesis of the current state of the art.

2. Data Sources and Their Characteristics in AD

2.1. Types of Measured Variables

AI models applied to AD rely on datasets that differ in their origin, structure, sampling frequency, and analytical complexity. Unlike many industrial processes monitored using only a limited number of operational variables, AD combines continuous sensor measurements with laboratory analyses, gas composition records, spectroscopic measurements, and, increasingly, molecular data describing microbial communities [36,37]. Each type of data reflects a different aspect of reactor performance and contributes distinct information for prediction, monitoring, and process optimization. Their characteristics also influence subsequent preprocessing, feature engineering, and model selection.

2.1.1. Physicochemical Process Data

Most datasets used for AI applications in AD originate from routine monitoring of physicochemical reactor conditions. The biological stages of AD remain closely interconnected, and disturbances affecting one stage rapidly influence the others. Although these stages are often presented sequentially for conceptual simplicity, in practice they occur concurrently, with different microbial groups and substrate fractions being active simultaneously [36,38]. Continuous measurements of operational variables therefore provide the primary source of information about reactor performance. Figure 1 illustrates the main stages of AD together with the principal intermediates formed during each biochemical step.
Figure 1. Simplified scheme of the AD process showing the four biological stages, major intermediates, and final products. For clarity, the stages are shown in linear, sequential form; in practice, they occur concurrently, with different substrate fractions and microbial guilds active simultaneously.
Temperature is one of the most frequently monitored variables because it directly influences microbial metabolism and community composition. Industrial digesters operate under psychrophilic, mesophilic, thermophilic, or extremely thermophilic conditions selected according to substrate characteristics and process requirements [39]. Temperature is usually recorded continuously using probes installed inside the reactor, producing high-resolution time series suitable for predictive modelling [40,41,42]. Even relatively small temperature fluctuations may alter microbial activity, particularly during acidogenesis, where reaction efficiency has been reported to decrease by almost 50% within the range of 15 to 45 °C [43].
Continuous pH measurements constitute another core component of process monitoring. Electrochemical sensors provide reliable measurements at relatively low operating cost [40,44]. Different microbial groups exhibit distinct optimal pH ranges, and buffering agents are commonly added to maintain conditions favourable for methanogenesis [36,45]. Although pH is readily available, it has limited value for early fault detection because the buffering capacity of anaerobic digesters delays measurable changes despite the accumulation of acidic intermediates. Observable pH shifts therefore often appear only after process imbalance has already developed [46].
Among laboratory measurements, VFAs provide some of the most informative indicators of reactor stability. Individual acids, including acetate, propionate, butyrate, valerate, and their branched isomers, respond differently to changes in organic and hydraulic loading [47]. Acetate usually reflects short-term disturbances, whereas elevated propionate concentrations are more closely associated with prolonged deterioration of reactor performance. Routine VFA monitoring remains limited by analytical requirements. Conventional determination by gas chromatography or high-performance liquid chromatography involves centrifugation, filtration, acidification, and refrigerated sample storage before analysis, resulting in datasets with relatively low sampling frequency [24,46]. Automated online gas chromatography systems have been proposed, but their complexity has restricted wider industrial implementation [24]. Continuous estimation of VFAs has therefore become one of the main applications of ML-based soft sensors discussed later in this review.

2.1.2. Gas Production Data

Gas production measurements represent another major source of information used to assess AD performance. Total biogas production is routinely monitored because it reflects reactor productivity and process efficiency. Gas yield may be estimated from substrate composition or measured experimentally using volumetric, manometric, or gas density methods, whereas methane and carbon dioxide concentrations are commonly determined by gas chromatography [48].
From the perspective of ML, gas production data are characterised by continuous acquisition and relatively high measurement reliability. Their usefulness depends on the selected variable. Total biogas flow often changes gradually and may not indicate the onset of process instability. During the early stages of inhibition, methane production decreases while carbon dioxide production increases, allowing the total gas volume to remain relatively stable despite deteriorating reactor performance [47]. Continuous methane measurements therefore provide greater diagnostic value and are more frequently used in predictive models and early warning systems [49].

2.1.3. Spectroscopic Data

Spectroscopic techniques have become an important source of information because they enable simultaneous estimation of multiple process variables from a single measurement. Instead of determining individual compounds directly, these methods analyse characteristic molecular vibrations to estimate variables such as chemical oxygen demand (COD), alkalinity, and VFA concentrations [50,51]. Depending on analytical requirements, ultraviolet–visible, near-infrared (NIR), and mid-infrared (MIR) spectroscopy may be implemented either in laboratory analyses or online monitoring systems [50,52].
Compared with conventional sensors, spectroscopic techniques generate high-dimensional datasets that require dedicated preprocessing before model development. Wavelength selection, feature extraction, and dimensionality reduction are therefore common steps during data analysis. Eccleston et al. demonstrated that ML models combined with attenuated total reflectance Fourier transform infrared spectroscopy accurately predicted total alkalinity and ammonium nitrogen. Prediction of VFAs was reliable only at higher concentrations because longer infrared fibres reduced the signal-to-noise ratio [53]. More recent studies have shown that NIR spectroscopy combined with wavelength selection algorithms can estimate acetate and propionate directly without extensive sample preparation, illustrating the potential of optical sensing for real-time AD monitoring [54].

2.1.4. Microbial and Omics Data

Physicochemical measurements describe reactor performance, whereas molecular datasets provide information about the microbial communities responsible for AD. Because microbial composition and activity determine the efficiency of substrate degradation, biological data are becoming an increasingly important source of information for predictive modelling [55,56].
Most microbial datasets are generated through DNA or RNA extraction followed by sequencing of the 16S rRNA gene, producing detailed profiles of community composition [56,57,58]. Compared with conventional sensor measurements, these datasets are collected less frequently, contain thousands of biological features, and require extensive bioinformatic processing before analysis [57]. ML models trained on sequencing data have already demonstrated the ability to predict methane production and identify microbial signatures associated with reactor performance [59]. Additional information can be obtained from metagenomics, transcriptomics, proteomics, and metabolomics, which describe microbial function and metabolic activity from complementary perspectives [60,61].

2.2. Common Data Limitations

The increasing availability of process measurements has created new opportunities for applying AI in AD. The performance of predictive models, however, remains closely linked to the quality of the available data. Measurements collected from full-scale digesters are affected by technical and operational factors that reduce data completeness, consistency, and transferability. Careful preprocessing and quality control therefore remain essential before model development.
One source of difficulty is the nonlinear behaviour of the process itself. Mechanistic frameworks, particularly, provide a detailed mathematical description of biochemical and physicochemical transformations, but they rely on deterministic assumptions that cannot fully reproduce the dynamic behaviour of industrial digesters [16,62]. Simplified ADM1 variants reduce computational complexity by limiting the number of state variables, while Benchmark Simulation Models extend these concepts to plant-wide wastewater treatment systems [62,63]. Despite these developments, deterministic models remain sensitive to nonlinear interactions between microbial populations and continuously changing operating conditions [62,64]. Experimental datasets therefore often deviate from theoretical predictions and present additional challenges for data-driven modelling.
The availability of informative measurements represents another limitation. Most online monitoring systems continuously record variables such as temperature, pH, gas production, and methane concentration. These parameters are readily available but often respond only after substantial changes have already occurred inside the reactor. Variables that provide earlier indications of instability, particularly individual VFAs, usually require laboratory analyses because robust online measurement techniques are still limited [21,47]. As a result, industrial datasets combine continuously recorded sensor signals with laboratory measurements collected at much lower frequencies. Such asynchronous sampling complicates data integration and reduces the effectiveness of predictive models.
Measurement uncertainty further affects data quality. Sensors operating under anaerobic conditions remain exposed to chemically aggressive media, suspended solids, and microbial biofilms throughout long operating periods. These conditions promote biofouling, accelerate sensor drift, and gradually reduce measurement accuracy [65,66]. Because signal degradation may develop slowly, systematic errors can accumulate in historical datasets before they are detected. Reactor contents also vary in composition, which alters sensor responses and increases measurement variability [66]. Raw industrial datasets therefore frequently contain noise, missing observations, and outliers that require preprocessing before model training.
Spatial and operational heterogeneity introduces additional variability. Measurements are usually collected at a limited number of sampling points, although biological and physicochemical conditions may differ considerably within large industrial digesters [66]. Differences between biogas plants further increase dataset variability. Reactor configuration, substrate composition, operating strategy, sensor infrastructure, and sampling frequency often differ from one installation to another. Monitoring systems are also influenced by economic constraints, with many facilities adopting customised sensor configurations instead of standardised commercial platforms [50]. While these solutions reduce implementation costs, they also increase variability in measurement protocols and data quality.
The lack of standardisation remains one of the main obstacles to developing transferable AI models. Available datasets differ in temporal resolution, measured variables, data formats, and quality control procedures, making it difficult to combine information collected at different facilities. To date, no large-scale public dataset represents the diversity of industrial AD systems while providing consistent measurement protocols suitable for benchmarking ML models [50].
The limitations discussed above originate primarily from the characteristics of the available data rather than from the learning algorithms themselves. Noise, sensor drift, asynchronous sampling, missing observations, and process heterogeneity all influence model performance and should be addressed before feature engineering and model training.

2.3. Implications of Data Characteristics for AI/ML Applications

The characteristics of datasets generated during AD directly influence the design and performance of AI models. Unlike many industrial applications operating under relatively stable conditions, AD datasets combine multiple process variables recorded at different sampling frequencies and with different levels of measurement uncertainty. Temporal dependencies further increase modelling complexity, requiring algorithms capable of capturing nonlinear relationships and delayed biological responses.
These characteristics also explain the growing role of data-driven approaches alongside mechanistic modelling. Mathematical frameworks, particularly ADM1, describe biochemical and physicochemical transformations in considerable detail but require extensive parameter calibration and significant computational resources [67]. Such requirements limit their use under continuously changing operating conditions. ML models learn directly from historical observations and can represent complex interactions without explicitly defining biochemical reaction pathways [68]. This flexibility has made them increasingly attractive for modelling industrial AD systems.
Data quality also determines the effectiveness of predictive models. Long-term industrial monitoring inevitably produces incomplete datasets because of sensor maintenance, calibration procedures, communication failures, and temporary process interruptions [69]. Missing observations reduce model reliability unless appropriate preprocessing methods are applied. Variational autoencoders have recently emerged as a promising approach for reconstructing incomplete datasets before model training [70]. Their performance, however, depends strongly on the quality and representativeness of the available training data, particularly when integrated with optimisation techniques such as genetic algorithms or reinforcement learning (RL) [71].
Differences between industrial digesters create another challenge for ML. Reactor configuration, substrate composition, operating conditions, and microbial communities vary considerably between facilities, limiting the transferability of models developed for individual plants. One response has been the development of hybrid modelling strategies that combine mechanistic process knowledge with the adaptability of ML. In these gray box frameworks, biochemical relationships described by models such as ADM1 provide the physical basis of the model, while data-driven algorithms adapt predictions to plant-specific operating conditions [72]. Meola and Weinrich demonstrated this concept by combining biomethane potential tests with a long short-term memory (LSTM) network to predict biogas production in a full-scale 188 m3 digester. Their hybrid framework maintained reliable prediction accuracy despite changes in feedstock composition, illustrating the robustness of combining mechanistic and data-driven approaches under industrial conditions [73].
In practice, the performance of AI models depends as much on the quality of the available data as on the learning algorithm itself. Missing observations, measurement uncertainty, asynchronous sampling, and reactor-specific variability must be addressed before reliable prediction and process control can be achieved. Table 1 summarizes the main data-related challenges identified in AD together with representative AI/ML approaches reported in the literature. The next section focuses on preprocessing methods developed to overcome these limitations.
Table 1. Key data-related challenges in AD and representative AI/ML approaches for their mitigation.

3. Data Processing and Pre-Modeling Workflows

The transition from empirical process management to predictive, data-driven control in AD is fundamentally limited by the quality and consistency of the input data [59]. Because ML and deep learning (DL) models operate without inherent mechanistic constraints derived from biological or thermodynamic laws, their predictive robustness depends entirely on the integrity of the ingested datasets. However, industrial-scale AD plants generate highly heterogeneous and multimodal data streams [97], characterized by a severe temporal mismatch: supervisory control and data acquisition (SCADA) systems log physical parameters continuously at a high frequency, whereas critical biological states are captured only via discrete, low-frequency laboratory analyses [98]. The corrosive and viscous operating environment accelerates sensor fouling, signal drift, and mechanical failures, introducing significant noise and temporal discrepancies. Implementing systematic preprocessing pipelines is an essential engineering step to transform these mismatched raw signals into a structured, mathematically consistent matrix suitable for reliable model training [99].
To systematically engineer raw, heterogeneous plant logs into high-fidelity predictive features, the preprocessing pipeline must address eight interrelated computational and biochemical challenges. First, systematic data cleaning and quality handling are required to isolate and correct hardware-induced anomalies, sensor fouling, and calibration drift without discarding critical process upsets [30]. Second, time-series structuring must synchronize asynchronous, multimodal data streams and capture the biological inertia of the methanogenic process using sliding windows and lag features [100]. Third, feature engineering incorporates domain-specific biochemical knowledge, such as thermodynamic thresholds and VFA-to-alkalinity ratios, providing models with contextualized operational surrogates [31]. Fourth, data fusion approaches establish the architectural strategies for merging diverse data streams, from macroscopic telemetry to microscopic optical spectra [101,102]. Fifth, robust scaling and normalization are deployed to accommodate the high variance of bioprocess parameters and mitigate the influence of sensor outliers. Sixth, to ensure realistic validation, preventing data leakage requires strict chronological boundaries during dataset splitting, preventing future information from corrupting historical training sets. Seventh, addressing data imbalance through synthetic sampling techniques allows the algorithmic representation of rare biological failures [102]. Finally, dimensionality reduction compresses ultra-high-dimensional biological or spectral inputs to prevent overfitting while preserving the core variances that govern system stability. Collectively, these preprocessing workflows ensure that downstream predictive models are trained on true biological signals rather than mechanical or sensor noise.

3.1. Data Cleaning and Quality Handling

While ML models offer advanced capabilities for AD monitoring and optimization, their transition from highly controlled, lab-scale environments to industrial applications is severely constrained by data quality challenges [103]. Unlike clean and symmetrically distributed laboratory datasets, real-world commercial AD telemetry is inherently noisy, highly non-linear, and compromised by the harsh operating conditions of full-scale reactors [97]. As a result, thorough data preprocessing is required to mitigate algorithmic bias and maintain the reliability of the predictive models.
The degradation of industrial AD data stems from a combination of hardware limitations and process dynamics, presenting three distinct methodological challenges. First, high total solids, corrosive biogas streams, and biofilm formation cause severe sensor fouling and calibration drift, which manifest as persistent signal noise, systematic measurement drift, or sudden hardware failures yielding missing values and extreme outliers [59]. Second, a severe temporal misalignment arises from the co-existence of asynchronous data streams: SCADA systems record macroscopic physical parameters (such as temperature, pH, and gas flow) continuously at high frequencies, whereas critical biochemical indicators of process stability (such as VFAs or COD) are analyzed in laboratories sporadically, often on a weekly basis [103]. Finally, because commercial digesters are operated near steady-state to maximize economic return, critical biological upset events, such as organic overloading or severe acidification, are highly underrepresented. This severe dataset imbalance systematically biases standard algorithms toward predicting the stable majority class, directly undermining their efficacy as proactive early-warning systems.
To address these industrial data complexities and ensure high-fidelity model inputs, several advanced data handling strategies must be employed, starting with robust outlier detection and noise filtering. Simple statistical thresholds are often insufficient for the dynamic AD process, making advanced anomaly detection algorithms, such as isolation forests or one-class support vector machines (SVMs), necessary to differentiate between genuine process disturbances (e.g., sudden temperature drops due to feeding) and erroneous sensor spikes. For example, in a recent full-scale study, Zhuang et al. (2025) developed a unified data cleaning pipeline utilizing algorithmic anomaly detection on high-dimensional SCADA telemetry [104]. By identifying abnormal multivariate states, this method allowed them to automatically detect and filter out sensor malfunctions and spurious noise, successfully isolating the true biological process signals before training their predictive models. Alongside these algorithms, signal smoothing techniques, such as moving average or Savitzky–Golay filters, are frequently applied to mitigate high-frequency electrical noise. As a practical application, when processing continuous NIR spectra to monitor critical parameters like total ammonia nitrogen (TAN) in active digesters, researchers have successfully utilized the Savitzky–Golay algorithm with a polynomial approximation [105]. This specific smoothing technique enabled them to eliminate baseline signal drift and significantly improve the precision of predictive models, without distorting the critical absorption peaks of the target chemical bonds.
Once the data is filtered and smoothed, another essential step involves imputation and resampling to synchronize asynchronous data streams. Because missing values are frequently caused by sensor downtime or the low frequency of laboratory tests, they are typically handled using advanced imputation techniques, such as k-nearest neighbors (k-NN) imputation or linear interpolation. This allows continuous features to be aligned with discrete target variables. For instance, to synchronize between high-frequency SCADA telemetry and weekly offline laboratory analyses, researchers have successfully implemented linear interpolation and forward-filling methods. This approach synchronized continuous physical parameters with discrete biochemical indicators (such as VFA), allowing predictive models to capture long-term process dependencies without suffering from discontinuous data gaps [106]. In scenarios involving severe industrial sensor downtime, algorithms like k-NN imputation have been utilized to estimate missing values based on the multidimensional similarity of surrounding operational states. By applying k-NN, researchers were able to mathematically reconstruct corrupted sensor logs and maintain high predictive accuracy for full-scale biogas yields, avoiding the need to discard historical datasets [107].
To address the problem of data imbalance and improve a model’s ability to predict digester failure, researchers increasingly utilize synthetic data generation and oversampling techniques. Algorithms like the synthetic minority over-sampling technique (SMOTE) are applied to artificially generate minority class instances (e.g., failure states), providing a more balanced training environment for predictive algorithms. For example, when modeling highly imbalanced full-scale industrial datasets, where rare biological failure states are vastly outnumbered by long periods of steady operation, researchers have implemented SMOTE to synthesize artificial failure data points based on the nearest neighbors of actual historical anomalies [102]. This strategic oversampling prevented the ML models from becoming heavily biased toward the majority class, significantly improving the algorithms’ sensitivity and precision in forecasting critical events, such as sudden organic overloading or acidification, before they reached irreversible stages.
In cases where physical sensor reliability is consistently poor, modern frameworks employ data fusion and soft sensors. By combining available raw sensor data with physicochemical principles, these virtual sensors can be constructed to cross-validate physical readings, mathematically estimate missing parameters, and filter out physically impossible data points before they reach the ML model. For instance, to overcome the limitations of expensive and easily fouled hardware sensors inside the digester, researchers have substantially developed data fusion frameworks that integrate continuous, low-cost physical measurements (such as pH, oxidation-reduction potential, and temperature) with established thermodynamic equations [101]. By deploying these physicochemical-driven soft sensors, they provided real-time, uninterrupted estimations of critical but difficult-to-measure biochemical parameters, such as VFAs and COD. This approach not only mitigated the impact of physical sensor drift but also drastically improved the overall predictive accuracy of downstream ML algorithms by supplying them with a continuous, cross-validated feature set.
Deploying predictive algorithms in commercial biogas plants requires shifting the focus from purely algorithmic tuning to rigorous, domain-aware data engineering [31,102]. In the absence of cleaning and validating incoming industrial data, even the most sophisticated neural networks will fail to deliver actionable operational insights [30,101].
In the harsh, corrosive, and highly viscous environment of an anaerobic digester, physical sensors are continuously exposed to conditions that compromise their integrity. Biofilm accumulation, chemical scaling, and electrical interference frequently lead to corrupted, drifting, or highly unstable data streams [30,50,101]. To prevent these erroneous signals from degrading the performance of process control systems or ML models, robust anomaly detection techniques must be implemented at the early stages of data preprocessing [104,108,109].
The most effective methodologies for identifying compromised sensor signals can be categorized into statistical, signal processing, and data-driven approaches. Among the statistical methods, rate of change and gradient analysis demonstrate high efficacy, as biological processes like AD are characterized by relatively slow dynamics. Rapid, instantaneous fluctuations in parameters such as temperature, pH, or digester volume are physically impossible under normal operating conditions. By calculating the first derivative of a time-series signal, engineers can set strict, physics-based thresholds. Any signal gradient exceeding these biological or physical limits indicates a sensor malfunction, such as a short circuit or sudden signal loss. For example, when processing high-frequency SCADA logs from full-scale digesters, researchers have applied first-derivative thresholds to parameters such as inline pH and reactor temperature [109,110]. By identifying instances where the rate of change exceeded the maximum thermodynamically possible cooling rate or the fastest known biological acidification rate, they isolated and removed erroneous spikes caused by routine sensor recalibrations and electrical shorts. This physics-based gradient filtering prevented downstream predictive models from misinterpreting standard hardware maintenance as catastrophic biological failures, thereby dramatically reducing false alarm rates in the final control system.
Beyond gradient analysis, rolling variance and statistical thresholding offer an additional layer of anomaly detection. While global statistical methods like Z-scores or the interquartile range are useful for identifying absolute outliers, they often fail to detect localized signal instability. Calculating the rolling standard deviation or moving variance over a defined time window is highly effective for identifying periods of excessive sensor noise. A sudden spike in signal variance, even if the mean value remains within an acceptable range, typically points to sensor fouling or poor electrical connections. For instance, during long-term monitoring of full-scale anaerobic digesters, researchers have utilized moving standard deviation windows to evaluate the reliability of continuous inline measurements, such as biogas flow rates and pH levels [50,109]. By continuously tracking signal variance, they were able to detect the onset of progressive sensor fouling, characterized by severe signal jitter, days before the actual mean values drifted out of acceptable operational limits. Flagging and isolating these high-variance data segments ensured that downstream ML algorithms were not trained on degraded inputs, thereby maintaining the predictive integrity of the models.
In data-driven approaches, evaluating multivariate consistency and cross-correlation is essential because process variables in biochemical systems are strongly coupled. Data-driven anomaly detection leverages these relationships to identify physical and biological discrepancies. For instance, a sudden drop in pH should theoretically correlate with an increase in VFAs or a change in biogas composition. If a pH sensor reports a decline but all other correlated sensors remain stable, the isolated pH signal is likely corrupted. ML algorithms, such as isolation forests or one-class SVMs, are frequently deployed to automatically detect these multivariate discrepancies. For instance, applying One-Class SVM to full-scale digester datasets has allowed researchers to map the multidimensional baseline of normal operational states. When an inline pH sensor indicated severe acidification, but the OC-SVM detected that correlated parameters, such as methane concentration and biogas flow rate, remained entirely stable, the system correctly flagged the event as a localized hardware calibration fault rather than a biological crisis [101,109]. In a different application, researchers utilized isolation forests to process complex SCADA telemetry by evaluating the joint distribution of feeding rates, reactor pressure, and liquid volume. Because isolation forests isolate anomalies based on multidimensional feature splits, the algorithm successfully identified multivariate discrepancies where sensor readings physically contradicted each other, such as continuous feeding with simultaneous, unexplained drops in digester volume. This enabled the automated removal of contradictory instances prior to neural network training, ensuring the model learned true biological dynamics rather than sensor logic errors [104].
Completing the triad of methodologies, signal processing techniques, specifically frequency domain analysis, play an important role in identifying compromised data, as unstable sensor signals often contain high-frequency noise that is distinct from the low-frequency trends of the actual biological process. Techniques such as Wavelet Transforms or Fast Fourier Transforms (FFT) allow researchers to decompose the raw signal into different frequency bands. If the power of the high-frequency components suddenly increases, it serves as a strong indicator of external interference or hardware degradation, allowing for targeted filtering before the data is utilized. For example, in the processing of highly volatile continuous gas flow and pressure measurements, researchers have successfully applied Wavelet Transforms to separate true biological dynamics from mechanical interference. Because vibrations from industrial feed pumps and compressors create distinct high-frequency signatures, wavelet analysis allowed engineers to isolate and subtract this mechanical noise without aggressively smoothing out the underlying, slower variations in actual biogas production [50]. Similarly, FFT algorithms have been deployed to monitor inline electrochemical probes. By analyzing the frequency domain, automated systems can detect the sudden emergence of specific electrical interference patterns, often caused by failing sensor shielding or ground loops in the harsh digester environment, and flag the hardware for maintenance before the corrupted signals propagate into the ML pipelines [50].
In consequence, by integrating these diverse validation techniques, ranging from dynamic statistical windows and gradient analysis to multivariate consistency checks and frequency domain filtering, researchers and biogas plant operators can establish an effective, multi-layered fault detection and diagnosis architecture. This essential data validation layer ensures that any unstable, drifting, or fundamentally corrupted signals generated within the harsh digester environment are automatically flagged and quarantined in real-time. Once isolated, these erroneous data streams can be strategically managed: either surgically discarded to prevent model bias or accurately reconstructed using context-aware imputation methods. Implementing this rigorous pre-modeling validation framework guarantees that deeply flawed operational telemetry cannot propagate through the data pipeline. It acts as the ultimate safeguard for downstream analytics, ensuring that advanced predictive algorithms, neural networks, and automated process control systems are trained exclusively on high-fidelity, biologically valid data. It is this foundational commitment to data integrity that ultimately reconciles theoretical ML models with reliable, AI-driven optimization in full-scale AD.

3.2. Time-Series Structuring

Restructuring asynchronous plant telemetry into consistent input-target matrices is a critical prerequisite for supervised ML in AD [111,112]. Full-scale AD facilities present a severe temporal mismatch between multimodal data streams: operational physical parameters (temperature, pressure, feed rates) are logged continuously via SCADA systems at high frequencies (1–5 min intervals), whereas critical biochemical indicators of process stability (VFAs, ratio of volatile organic acids to total inorganic carbon (FOS/TAC), COD) are analyzed sporadically via manual laboratory testing on a daily or weekly basis [113].
To reconcile this sampling disparity, high-frequency SCADA signals are typically downsampled to match the lower-frequency laboratory timeline [97]. Beyond enabling temporal alignment, systematic downsampling using hourly or daily medians functions as an implicit low-pass filter [112]. This aggregation suppresses high-frequency electrical noise, transient sensor spikes, and micro-fluctuations, thereby preventing downstream models, such as random forest (RF), from overfitting to irrelevant hardware anomalies and ensuring they capture the underlying macro-biological kinetics of the AD process.
To reconcile infrequent offline laboratory measurements with continuous SCADA telemetry, upsampling and imputation are required [114]. Although linear or spline interpolations estimate intermediate values, they introduce look-ahead bias by incorporating future data points, leading to artificial inflation of model accuracy during offline validation. In predictive operational scenarios, a ‘forward-fill’ strategy, carrying the last observed laboratory measurement constant until a subsequent sample is recorded, is methodologically preferred. This approach constrains model inputs to historically available information, accurately simulating the real-time decision-making environment of full-scale AD facilities.
AD is characterized by high biological and physical inertia, driven by prolonged hydraulic retention times (HRT) and the slow growth rates of anaerobic methanogenic consortia [59]. Instantaneous process states are insufficient for forecasting near-term bioreactor behavior. Because daily biogas yield is determined by historical feeding rates, organic loading, and substrate composition from preceding days or weeks, the biochemical system exhibits strong temporal autocorrelation. Developing reliable predictive models requires shifting from static, memoryless assessments to sequential time-series architectures that ingest extensive historical lag windows to capture these delayed, cascading biochemical responses.
To embed this historical context into the dataset, temporal structuring techniques such as lag features and historical trajectories are applied. Instead of restricting the predictive model’s input to only the most recently recorded value of a given parameter, the dataset is systematically expanded to incorporate a sequential window of its past observations. By structuring the data to include measurements from preceding hours or days, the algorithm is explicitly provided with the dynamic trajectory and momentum of the variable, rather than just a static, isolated data point. For instance, during the development of predictive software for an industrial-scale biogas facility, researchers implemented lag feature engineering to capture the delayed impact of feeding operations. Rather than attempting to predict daily biogas yield based solely on that specific day’s organic loading, they expanded the feature matrix to include the feeding volumes, temperatures, and pH levels from the preceding three to five days [107]. This historical restructuring allowed ML algorithms, such as RFs, to mathematically internalize the biological retention time, improving the model’s accuracy during periods of fluctuating or irregular feed schedules.
For sequential DL architectures capable of processing sequential data, such as recurrent neural networks (RNNs), LSTM networks, or convolutional neural networks (CNNs), the continuous operational timeline must be systematically segmented using sliding windows and sequence extraction. In this approach, a temporal window of a fixed duration (for instance, the preceding 72 or 168 h) advances along the time series by a defined step size. Each extracted window forms a multidimensional matrix of historical sensor data, serving as a single, cohesive input sequence to predict future states. For example, when deploying LSTM networks to forecast impending acidification events and biogas yields in commercial digesters, researchers have utilized sliding windows to encapsulate full weeks of high-frequency SCADA telemetry. By feeding the neural network these sequential matrices rather than isolated data points, the LSTM algorithm was able to model the gradual accumulation of inhibitory compounds and the fading impact of past feeding cycles [104,115]. This sequential learning approach allowed the model to deeply internalize the complex temporal dependencies and high inertia of the biological system, resulting in highly accurate, multi-day forecasts of digester health that traditional, memoryless algorithms simply could not achieve.
The culmination of temporal data structuring involves isolating the specific operational outcome the model is designed to forecast. While conventional regression architectures typically map synchronous input features to a simultaneous output, the primary ambition in data-driven AD management is proactive forecasting rather than simultaneous real-time estimation. The targeted performance metric, such as the anticipated daily biogas yield or an impending spike in VFAs, must be intentionally shifted forward along the chronological axis relative to the historical input matrix. By purposefully aligning past operational sequences with future target states, data engineers force the ML algorithm to resolve the temporal offset. This temporal alignment ensures that the model learns to predict future biological behaviors based strictly on preceding conditions, thereby providing biogas plant operators with genuine early-warning capabilities and a critical temporal window for corrective action.
While the biological rationale for incorporating ecological memory is clear, this restructuring must be rigorously formalized to construct the necessary feature and target matrices for ML training. Let x t R represent a multidimensional vector of m operational features (e.g., feeding rate, reactor temperature, pH) measured at a discrete time t. A conventional, memoryless predictive model attempts to map this instantaneous input directly to a synchronous target state, y t .
To mathematically account for process inertia and delayed biochemical responses, the input space is expanded using a temporal window. By defining a historical sequence length, w, which represents the required historical depth, the input at time t is transformed from a single vector into a two-dimensional sliding window matrix, X t w :
X t w = [ x t w + 1 ,   x t w + 2 , , x t ]
Concurrently, the predictive objective must be mathematically isolated. By defining a forecasting horizon, h (where h > 0), the target variable is temporally shifted forward to represent a future biological state. Depending on the operational goals of the biogas plant, structuring this target variable dictates the architectural complexity of the model, broadly categorized into single-step or multi-horizon forecasting [107].
In a single-step target approach, the objective is to predict one specific, isolated point in the future. For example, the advanced approximation function f might be optimized to map the historical input matrix exclusively to the expected biogas yield exactly 24 h later:
y ^ t + 24 = f ( X t w ) ,
where the circumflex notation y ^ explicitly denotes the estimated value predicted by the algorithm, distinguishing it from the true, physically measured future state (y).
However, predicting an impending biological collapse, such as a sudden acidification event caused by organic overloading, often requires a broader perspective. In these critical scenarios, the target is structured as a multi-horizon sequence. Instead of forecasting a single value, the algorithm is trained to output a continuous trajectory of the expected process behavior over several consecutive time steps:
[ y ^ t + 1 ,   y ^ t + 2 , , y ^ t + k ] = f ( X t w ) ,
where k explicitly defines the forecasting horizon length, representing the total number of future time steps to be predicted simultaneously.
Multi-horizon forecasting trains the model to anticipate the dynamic evolution of the reactor’s biochemistry over the next several days [104,115]. This approach is exceptionally valuable in full-scale operations, as it provides operators with a continuous predictive window. Visualizing the progressive trajectory of parameters like VFAs allows plant managers to implement proactive adjustments in feeding regimes, such as reducing the OLR, long before the biological imbalance becomes irreversible. Structuring the data pipeline using these time-series formulations enables the shift from real-time monitoring to predictive process control.
To illustrate the practical necessity of this architecture, consider a recent application of sequence-to-sequence modeling in a commercial anaerobic digester. Researchers deployed an LSTM network designed specifically for multi-horizon forecasting to predict the reactor’s stability index, such as the VFA-to-alkalinity ratio, over a 7-day future window (k = 7). By predicting the entire multi-day trajectory rather than just the immediate next day, the algorithm successfully identified impending acidification events up to 96 h before critical inhibition thresholds were breached [106]. This continuous predictive window provided plant operators with actionable foresight, allowing them to incrementally reduce the daily OLR and safely stabilize the biology without halting operations entirely. Such proactive, graded interventions are structurally impossible when relying on simple single-step forecasting, demonstrating why multi-horizon models are the cornerstone of modern decision-support systems in biogas plants.

3.3. Feature Engineering for AD Systems

While DL architectures can theoretically extract patterns from raw data streams, applying unconstrained black-box models to AD often yields suboptimal generalization due to complex thermodynamics and slow biokinetics [59,116]. Feature engineering addresses this limitation by embedding domain-specific biochemical knowledge into the input space, transforming raw plant logs into physically interpretable predictors.
These engineered inputs generally bifurcate into physical loading metrics and biochemical process surrogates. Physical loading parameters, such as OLR and HRT, reconcile raw feeding telemetry (e.g., feed pump runtimes or substrate mass) with reactor capacity and substrate characteristics. Utilizing OLR as a standardized indicator of biological stress, rather than raw feed mass, prevents model overfitting and enhances generalization across reactors of varying scales.
Conversely, biochemical surrogates and soft-sensor architectures represent the target biochemical state of the digester when direct online measurements are unavailable. Because continuous monitoring of VFAs or the FOS/TAC ratio is historically constrained by sensor fouling and high instrumentation costs, soft sensors mathematically reconstruct these hidden states. By integrating accessible, high-frequency physical SCADA inputs, including pH fluctuations, gas flow rates, and methane fractions, these virtual sensors provide continuous, real-time indicators of acidification risk without the delays associated with offline laboratory assays.
Lastly, feature engineering extends to extracting temporal derivatives and frequency-domain signatures to capture dynamic process shifts. While absolute pH is universally monitored, it is a lagging indicator of biological instability due to the high buffering capacity (alkalinity) of anaerobic digestate, which remains stable despite VFA accumulation [114]. To resolve this delay, feature engineering extracts the velocity of sensor signals via the first derivative or rolling variance, providing sensitive early-warning proxies that detect organic overloading before the system’s buffering capacity is exhausted.
Beyond time-domain derivatives, complex physical and hydrodynamic phenomena, such as foaming, changes in sludge rheology, and inadequate mixing, manifest as high-frequency oscillatory noise in SCADA signals, including gas flow rates or mechanical agitator torque. Converting these time-series measurements into the frequency domain via mathematical operations like the FFT or Wavelet Transforms decomposes the complex signals [59]. This frequency-domain representation isolates systemic bioprocess dynamics from stochastic mechanical interference, preventing downstream models from misinterpreting physical disturbances as biological upsets.
Rather than training models on raw, high-frequency telemetry, transforming signals into the frequency domain allows the extraction of distinct physical descriptors, such as dominant frequencies, spectral power distributions, and spectral entropy. For example, during bioreactor foaming, a critical operational failure that threatens gas infrastructure, thick foam traps expanding gas bubbles, dampening system hydrodynamics. Instead of relying on absolute pressure limits, ML models trained on high-frequency pressure or impedance signals can detect this structural state through an abrupt downward shift in the dominant frequency and a simultaneous reduction in spectral entropy. This frequency-domain transformation captures the physical rheological changes in the digestate, providing predictive features for gas holdup that remain entirely latent in standard time-domain averaging.
Beyond signal processing, a parallel feature engineering strategy involves embedding established biochemical and thermodynamic principles directly into the input matrix as explicit mathematical constructs. Rather than requiring purely data-driven algorithms to heuristically approximate complex, non-linear biokinetics from scratch, these domain-informed inputs pre-calculate key thermodynamic boundaries and kinetic rates. This framework integrates deterministic bioprocess modeling (e.g., ADM1) with statistical machine learning to establish effective, hybrid data-driven pipelines.
A prime illustration of this approach is the calculation of Free Ammonia Nitrogen (FAN). An industrial SCADA dataset typically contains independent telemetry columns for TAN, reactor temperature, and absolute pH. However, the actual biological toxicity to methanogenic archaea is not driven by the total ammonia, but specifically by the FAN. If left unengineered, an algorithm would have to implicitly learn the thermodynamic equilibrium laws governing these three distinct variables. Instead, data engineers actively employ standard chemical equilibrium equations to synthesize a new feature:
F A N = T A N 1 + 10 p K a p H
where the acid dissociation constant, p K a is a direct non-linear function of the reactor’s temperature. By engineering this specific inhibition index, the feature matrix is explicitly populated with the exact biological toxicity threshold. For instance, in the development of predictive frameworks for nitrogen-rich substrates (such as poultry litter or slaughterhouse waste), algorithms fed exclusively with raw TAN and pH data frequently failed to predict sudden biological failures. However, by explicitly engineering the FAN concentration as a distinct input vector, researchers drastically reduced the required architectural complexity of the neural networks. Because the non-linear thermodynamic relationship was already mathematically resolved during the preprocessing phase, the predictive models achieved significantly higher accuracy and stability in forecasting ammonia-induced methanogenic inhibition [30,115]. Integrating such deterministic indices ensures that the AI remains strictly anchored to the physical realities of the microbiological environment.

3.4. Data Fusion Approaches

To overcome the limitations of isolated, single-variable monitoring, full-scale AD systems are transitioning toward multimodal data fusion frameworks. A reliable characterization of the bioprocess requires integrating macroscopic physical telemetry with biochemical and metabolic indicators. Data fusion addresses this challenge by systematically combining disparate, asynchronous data streams into a unified, high-dimensional predictive space to enhance model robustness, suppress individual sensor noise, and improve diagnostic accuracy [59,117].
Relying on single-modality monitoring is insufficient to capture the dynamic biochemical state of a bioreactor. For example, standard SCADA telemetry (such as temperature, pH, hydrostatic pressure, and mixing torque) describes the physical and macroscopic environment, while inline biogas analyzers (monitoring methane, carbon dioxide, and hydrogen sulfide concentrations) reflect the terminal metabolic output. However, both modalities are blind to the intermediate biochemical state of the slurry, particularly VFA concentrations. Consequently, models relying solely on these traditional inputs can only react to biological upsets after they manifest in the gas phase, limiting the window for proactive intervention.
Integrating spectroscopic techniques like NIR or MIR helps address observational limitations. However, in situ spectroscopic monitoring of raw, undiluted anaerobic digestate is constrained by high turbidity and total solids, which preclude standard transmission measurements due to light scattering and high optical absorbance. To acquire meaningful spectra, systems must utilize attenuated total reflection probes with micro-scale path lengths. Furthermore, these optical windows are susceptible to biofouling, biofilm accumulation, and chemical scaling, which distort absorption signals and induce calibration drift. Implementing automated cleaning systems (e.g., ultrasonic or chemical backwashing) is required for continuous data fusion. While Raman spectroscopy can identify molecular structures, it is affected by intense background fluorescence when analyzing complex organic digestate matrices, requiring specialized NIR excitation or baseline-correction algorithms. Finally, fusing these modalities—the liquid-phase layer (biochemical state), the physical SCADA layer (environmental conditions), and the gas-phase layer (metabolic output)—must account for inherent biokinetic and physical phase lags. Biogas composition changes do not align instantaneously with liquid-phase disturbances due to dissolved gas supersaturation and gas-liquid mass transfer limitations. Incorporating these temporal offsets and mass transfer dynamics within the data fusion architecture is essential to prevent look-ahead bias and ensure the accurate cross-validation of physical and biochemical parameters.
Multimodal dataset structuring for ML in AD primarily utilizes either early (data-level) or late (decision-level) fusion architectures. Early fusion concatenates raw features from all modalities (e.g., SCADA, gas, and optical telemetry) into a single, high-dimensional feature matrix. While this enables the extraction of low-level cross-correlations across different sensor types, it is susceptible to the ‘curse of dimensionality’ and requires computationally intensive synchronization of asynchronous sampling rates. Conversely, late fusion trains independent, dedicated sub-models for each data modality and aggregates their predictions using dynamic weighted averaging or ensemble meta-classifiers. This modular design isolates sensor-specific errors, ensuring that localized sensor downtime or calibration drift does not corrupt the entire predictive pipeline.
The operational trade-offs of these architectures are illustrated by industrial co-digestion deployments. Standard early-fusion models that concatenate SCADA and biogas composition telemetry frequently output unstable feed-loading recommendations during routine gas-analyzer downtime (e.g., caused by moisture condensation). Restructuring the predictive framework into a late-fusion ensemble, where separate sub-models for SCADA and gas telemetry are integrated via a meta-classifier, overcomes this limitation. By weighting predictions based on real-time sensor confidence scores, the system automatically redirects reliance to the active physical sub-model during localized hardware failures, ensuring continuous process control [118].
As AD datasets grow in scale and complexity, static fusion methods are increasingly complemented by deep multimodal learning and adaptive neural architectures. These networks can automatically adjust their internal representations based on changing operating conditions and varying data availability within the reactor environment.
A practical application of this data fusion strategy is the framework proposed by Jeong et al. for full-scale municipal co-digestion, combining a dual-stage attention LSTM with a variable selection network [114]. Instead of using a single data stream, this hybrid architecture fuses continuous physical SCADA telemetry (e.g., slurry flow rates and reactor temperature) with asynchronous, discontinuous biochemical indicators analyzed sporadically in the laboratory (e.g., volatile solids, COD, and leachate characterization). The framework utilizes a variable selection network based on gated residual networks to dynamically assign weights to the sparse laboratory parameters, while a dual-stage attention LSTM captures the temporal dynamics of the continuous SCADA data. Fusing these heterogeneous, multimodal variables improved prediction robustness under real-world sensor downtime, increasing the out-of-sample coefficient of determination (R2) for biogas prediction from 0.38 (for a standard LSTM) to 0.76.
Traditional sequence networks degrade in performance due to volatile, non-linear feed fluctuations; advanced temporal architectures dynamically map variable interactions along the chronological axis. To capture these complex temporal dynamics in food waste AD, Han et al. implemented an Inverted Transformer, an iTransformer model [119]. Unlike standard Transformer networks that partition sequence steps, the iTransformer applies self-attention mechanisms to map multivariate correlations across inverted feature dimensions. Tested on high-dimensional food waste digestion datasets, this adaptive architecture achieved a methane yield prediction accuracy of 98.55% and an R2 of 0.9949. Integrating this predictive model into the plant’s operational framework enabled dynamic feedstock configuration adjustments over a 20-day horizon, successfully enhancing average biogas production while reducing overall carbon dioxide emissions by 12.3%. This demonstrates that adaptive neural architectures can resolve complex, non-linear dependencies to facilitate proactive and environmentally optimized bioprocess control.

3.5. Scaling and Normalization of Process Variables

A fundamental characteristic of AD monitoring datasets is the extreme variance in the magnitude of operational parameters. For instance, process pH typically fluctuates within a strictly narrow band (e.g., 6.8 to 8.0), whereas daily biogas production can span thousands of cubic meters, and VFA concentrations are measured in hundreds or thousands of milligrams per liter. Many ML algorithms, particularly distance-based models (e.g., SVM, k-NN), ANNs, and dimensionality reduction techniques like PCA, are inherently scale-sensitive. Utilizing unscaled data causes variables with larger absolute values to disproportionately dominate the objective function and gradient weight updates. In ANNs, this leads to slow convergence or trapped local minima, effectively causing the algorithm to ignore the subtle yet biologically critical dynamics of smaller-scale variables like pH.
While traditional normalization (Min-Max scaling) or standardization (Z-score scaling) are standard practices in data science, they frequently fail when applied to raw industrial AD data. These conventional methods rely on the global minimum, maximum, or mean of the dataset, making them highly vulnerable to the extreme outliers inherent in commercial plant logs (e.g., sensor malfunctions, pump cavitation, or electrical interference). For instance, applying a Min-Max scaler to the j-th feature of the input vector x t . If a brief electrical short causes the pH probe to falsely register a value of 14.0, the standard scaling formula:
x t , j n o r m = x t , j x m i n , j x m a x , j x m i n , j
where x t , j represents the raw scalar value of the j-th feature at discrete time t, and x m i n , j and x m a x , j denote the absolute minimum and maximum values of this specific feature across the entire dataset. In this scenario, the anomaly forces x m a x , j = 14.0 . As a result, the entire normal biological operational range (6.8 to 8.0) is mathematically compressed into an insignificantly small variance band. The ML model becomes virtually blind to natural fluctuations, destroying its predictive utility.
To address this, Robust Scaling techniques are strictly required for industrial bioprocess data. Instead of relying on the highly susceptible extremes or the statistical mean, robust scalers center and scale the data using statistics that are mathematically immune to extreme outliers: the median and the interquartile range. The transformation for the j-th feature is formally defined as:
x t , j r o b u s t = x t , j Q 2 , j Q 3 ,   j Q 1 , j
where Q 2 , j is the statistical median (the 50th percentile) of the j-th feature’s distribution, while Q 1 , j and Q 3 , j represent its 25th and 75th percentiles, respectively. By scaling features based exclusively on the I Q R ( Q 3 , j Q 1 , j ) , the true variance of the biological process is preserved, and hardware-induced outliers are structurally prevented from corrupting the feature matrix before it reaches the predictive or dimensionality reduction stages.
The importance of robust scaling is best illustrated in the deployment of anomaly detection systems for commercial co-digestion facilities. In one operational scenario, engineers initially utilized standard Z-score standardization on a dataset combining high-volume feeding rates and trace hydrogen sulfide (H2S) concentrations. Because the feeding pumps occasionally recorded massive artificial spikes due to telemetry packet loss, the Z-score standardizer heavily distorted the feature space. When this data was fed into an unsupervised PCA algorithm for fault detection, the principal components (PCs)were entirely skewed by the artificial feeding outliers, masking a genuine, gradual accumulation of H2S that preceded a biological upset. By replacing the preprocessing pipeline with an Interquartile range-based Robust Scaler, a methodology grounded in the foundational principles of outlier-resistant statistical estimation established by Rousseeuw and Leroy [120], the algorithm isolated the extreme pump outliers without compressing the natural variance of the gas quality. This was essential because, as demonstrated by the diagnostic frameworks developed by Dunia et al. [121], PCA is inherently sensitive to signal noise without the preprocessing step provided by the Robust Scaler. This allowed the PCA to accurately identify the underlying biochemical shift, detecting the acidification event several days before it catastrophically impacted biogas yields.
The need for advanced preprocessing is particularly evident when integrating online spectroscopy, such as NIR or MIR sensors, with data-driven process monitoring [117,122]. In spectroscopic telemetry of turbid anaerobic digestate, raw absorbance measurements do not reflect chemical concentrations alone; they are also distorted by physical light-scattering effects and instrumental baseline drift. If unsupervised dimensionality reduction techniques like PCA are applied to unscaled spectral matrices, the leading principal components tend to capture these scattering anomalies rather than underlying biochemical variables, such as VFAs or total alkalinity [123]. To mitigate this, advanced monitoring frameworks implement standard normal variate (SNV) transformation or multiplicative scatter correction (MSC) as specialized scaling operators. By normalizing each spectrum, these techniques isolate chemical absorption features from physical pathlength variations. Applying these preprocessing steps enables PCA and partial least squares (PLS) algorithms to accurately resolve target biochemical dynamics, preventing baseline artifacts from misrepresenting stable operations as imminent digester acidification.

3.6. Preventing Data Leakage: Chronological Time-Series Splitting

An important prerequisite for preparing AD datasets for ML is the selection of a powerful partition strategy for training, validation, and testing subsets [124]. While random k-fold cross-validation is the default approach in standard ML, applying random shuffling to continuous AD telemetry introduces severe data leakage [106]. Due to the high biological inertia and prolonged HRT of anaerobic bioreactors, AD time-series data exhibits strong temporal autocorrelation. Random partitioning inadvertently trains models on future system states to predict past observations, leading to artificially inflated accuracy metrics during offline validation that fail to replicate in real-world deployments where future telemetry is unavailable.
To maintain temporal order, pre-modeling pipelines must utilize chronological validation strategies, such as time-series splitting (forward chaining). In this approach, models are trained on continuous historical blocks and validated exclusively on subsequent, unseen chronological segments. To evaluate robustness against seasonal variations and sensor drift, walk-forward validation via sliding or expanding windows can be implemented. This iterative method continuously updates the training dataset, mimicking the operational conditions of real-time process control. Restricting evaluation to forward temporal boundaries ensures that predictive architectures are assessed solely on their ability to forecast bioprocess behavior from past observations.
This methodological necessity is clearly demonstrated in the industrial-scale study conducted by De Clercq et al., who developed an ML framework to forecast biomethane production at a commercial co-digestion facility using 1398 days of continuous plant telemetry [97]. To evaluate the predictive robustness of Elastic Net, RF, and XGBoost models across multiple forward horizons, specifically predicting biomethane output 1, 3, 5, 10, 20, 30, and 40 days into the future, the researchers rejected random partitioning in favor of a chronological train-test split. By systematically separating the training data from subsequent, unseen operational periods, they prevented temporal data leakage and ensured that the models were tested under conditions replicating real-time plant operation. This validation design allowed for an unbiased assessment of the out-of-sample coefficient of determination and root-mean-square error (RMSE), proving that strictly respecting temporal boundaries is essential for deploying reliable, AI-driven decision-support software in full-scale AD systems.

3.7. Addressing Data Imbalance: Synthetic Sampling Techniques

A defining characteristic of full-scale AD datasets is severe class imbalance, as commercial biogas plants are engineered for continuous, steady-state operation and spend most of their operational lifespan in stable states [102,125]. Critical process anomalies, such as severe acidification, toxic inhibition, or hydraulic overloading, are rare. In a supervised learning context, standard classifiers trained on such datasets tend to converge on a majority-class bias, optimizing overall accuracy while failing as early-warning systems [126,127].
To address this, synthetic data generation techniques like SMOTE are implemented to balance the feature space prior to model training [125]. Unlike simple random oversampling, which duplicates existing points and induces overfitting, SMOTE interpolates synthetic minority instances along line segments connecting a chosen sample to its k-nearest neighbors. This generates transitional states, such as early-stage process upsets, which broadens the model’s decision boundaries. However, evaluations show that SMOTE can exhibit spatial contraction, where the generated covariance matrix shrinks relative to the true distribution. This contraction worsens as the number of minority samples decreases or as dimensionality increases. For highly non-linear boundaries, adaptive synthetic sampling focuses generation on minority examples located near the decision border.
While standard SMOTE handles nominal class imbalances, predicting rare extreme continuous values, such as sudden VFA spikes, requires methods like SMOTE for regression [128]. In this approach, the dataset is partitioned using a relevance function and a threshold to separate typical operational states from rare extremes. Oversampling then generates synthetic continuous target values via a distance-based weighted average of selected seed cases.
Synthetic oversampling must be applied exclusively to the training set after chronological data splitting. Applying it to the entire dataset prior to splitting causes synthetic instances to leak into the validation and test folds [124]. This temporal data leakage artificially inflates performance metrics and biases the evaluation, as the model is tested on interpolated representations of the training samples.
Han et al. (2023) [125] demonstrated this balanced sampling approach in an industrial-scale study predicting methane production at a commercial food waste AD plant. Because extreme operational cases and anomalous logs were rare in the dataset, the researchers applied SMOTE to balance the minority class representations. Incorporating these synthetic instances into an LSTM network improved model robustness under data scarcity. The SMOTE-LSTM model achieved a prediction accuracy of 99.75% and an R2 of 0.9913, outperforming traditional backpropagation neural networks (BPNN), SVMs, and standard LSTM networks trained without synthetic data balancing.

3.8. Dimensionality Reduction for High-Dimensional Inputs

Integrating continuous SCADA signals, high-resolution spectroscopy, and microbial metagenomics into AD systems increases dataset dimensionality and complexity. Because high-dimensional feature matrices often contain multicollinearity and noise, they can lead to model overfitting and high computational costs [59,129]. Dimensionality reduction is therefore necessary to compress these matrices while retaining essential bioprocess variance and biological signals.
For datasets dominated by correlated, continuous operational variables, PCA serves as a linear projection technique to orthogonally transform the original feature space into uncorrelated PCs. However, because thermodynamic and biokinetic relationships in AD are highly non-linear [97], standard linear projections may fail to capture complex process dynamics. Kernel PCA addresses this limitation by applying kernel functions (e.g., an RBF) to implicitly map variables into higher-dimensional Hilbert spaces where non-linear boundaries become linearly separable [116].
For processing high-dimensional biological or spectroscopic features, deep learning architectures such as autoencoders offer an effective method for non-linear dimensionality reduction [130]. By training a neural network with a bottleneck layer to reconstruct its input, autoencoders compress sparse variables, such as NIR wavelengths or operational taxonomic units, into dense, low-dimensional latent vectors that capture the primary biochemical state of the bioreactor. In multimodal applications, these latent representations can integrate different data sources (e.g., SCADA time-series and daily spectral scans) into a unified feature set, which improves the training efficiency of downstream predictive models like random forests or feed-forward networks.
Li et al. [131] demonstrated the utility of PCA for compressing biological datasets when modeling methane yield and content across diverse biowastes. To integrate microbial community profiles (at the phylum and genus levels) with operational parameters, they applied PCA to reduce the dimensionality of the taxonomic features. Selecting the first eight principal components retained over 90% of the variance, simplifying the complex microbial input data into a practical set of predictors. Pearson correlation analysis between these components and process parameters identified important ecological relationships, such as the positive influence of pH and biochar on specific bacterial phyla (e.g., Synergistetes, Atribacteria, Cloacimonetes). This indicates that linear projection can effectively capture broad ecological patterns in AD environments.
Although ML applications in AD are expanding, current pre-modeling pipelines face methodological limitations, particularly regarding the treatment of missing or asynchronous data. In full-scale AD facilities, missing telemetry is rarely missing completely at random. Instead, sensor fouling, corrosion, and maintenance schedules typically result in data that is not missing at random. Applying conventional imputation methods, such as linear or spline interpolations, assumes smooth bioprocess transitions. This can artificially mask transient biokinetic shifts and introduce look-ahead bias during offline validation. Conversely, while the “forward-fill” strategy prevents data leakage by holding low-frequency laboratory measurements constant, it fails to capture the actual rate of biochemical changes, such as rapid VFA accumulation. Consequently, models relying on this approach may fail to detect sudden process imbalances.
Trade-offs exist within dimensionality reduction and class-balancing operations. Linear projection techniques like PCA are widely used to eliminate multicollinearity in SCADA datasets, but they do not capture the non-linear, thermodynamic, and microbial interactions that govern anaerobic ecosystems. PCA risks projecting low-variance yet biochemically important parameters, such as pH, into lower-order principal components, reducing the model’s sensitivity to biological operating thresholds.
Similarly, resolving class imbalance of process failures through synthetic oversampling (e.g., SMOTE or SMOTE for regression) introduces hidden spatial distortions. As mathematically demonstrated by Elreedy and Atiya [126], SMOTE exhibits a contractive nature that shrinks the generated covariance matrix, producing synthetic data points that may not respect the true boundary physics of the anaerobic environment. Generating synthetic continuous features in highly correlated spaces without thermodynamic constraints can interpolate biochemically impossible reactor states (e.g., pairing high VFA concentrations with high pH), which subsequently trains ML models on unrealistic conditions.
Finally, a main gap remains between statistical data preparation and deterministic bioprocess chemistry. Current feature engineering methods are highly heuristic. They seek to maximize statistical correlation with target variables (such as biomethane yield) but fail to incorporate fundamental conservation laws (such as mass, charge, and elemental balances represented in deterministic models like ADM1). This lack of physical grounding limits generalizability: while models may achieve high accuracy during offline validation on historical datasets, their predictive performance often degrades when deployed in real-time environments characterized by feedstock variability, microbial community shifts, and sensor calibration drift.

4. ML Models for Prediction and Diagnostics

The models discussed in this section are established mathematical and computational frameworks; therefore, their underlying principles are described conceptually, with emphasis on their adaptation to AD data and process-specific applications rather than on detailed mathematical derivations of individual algorithms.

4.1. Classical ML Models

When discussing modelling approaches in AD, it is important to distinguish between domain-agnostic data-driven algorithms and models developed specifically for AD. Most ML and DL approaches discussed in this review are general-purpose algorithms originally developed outside the AD domain. Their application to AD involves adapting these established approaches to process-specific datasets and variables, such as pH, VFAs, OLR, and biogas production. In contrast, mechanistic models such as ADM1 were developed specifically to describe the biochemical and physicochemical processes occurring during AD. Hybrid and physics-informed approaches combine these two perspectives by integrating general-purpose ML algorithms with AD-specific mechanistic knowledge and process constraints.
Classical ML methods remain the most frequently reported predictive algorithms in AD research, despite the rapid expansion of DL techniques. One reason is the nature of the available datasets. Measurements collected from full-scale digesters are usually limited in size, dominated by structured process variables, and affected by measurement uncertainty. Under such conditions, conventional ML algorithms often provide stable predictive performance without requiring extensive computational resources or large training datasets [132].
Among the models most adopted are RF, gradient boosting (GBM), SVM, and k-NN. Although they belong to the same family of supervised learning methods, each captures process behaviour differently. RF constructs numerous decision trees from randomly selected subsets of observations and combines their predictions into a single output. This strategy reduces sensitivity to noisy measurements and lowers the risk of overfitting [133]. GBM follows a different principle. Instead of building trees independently, successive models learn from the prediction errors generated by the previous ones, gradually improving the final estimate and allowing accurate representation of complex nonlinear relationships frequently encountered in AD datasets [134].
SVM approaches the problem from another perspective by searching for an optimal separating surface within a transformed feature space. The use of kernel functions allows nonlinear operating conditions to be represented without explicitly increasing model complexity, making SVM particularly attractive for biological processes where relationships between operational variables rarely remain linear [135]. The k-NN algorithm is conceptually simpler because predictions are obtained from historical observations located closest to a new operating point according to distance-based similarity measures [135].
Rather than competing directly, these algorithms have found different areas of application. Methane production forecasting remains one of the most common objectives, although prediction of VFA concentration, reactor stability indices, and disturbance detection have also received considerable attention [30]. Reported results indicate that the optimal model depends not only on the prediction target but also on substrate variability, measurement frequency, and data quality.
Several studies have identified SVM as one of the most reliable solutions for methane yield prediction. Gupta et al. and Ganeshan et al. reported that SVM maintained high predictive accuracy even when sensor measurements contained considerable variability, reflecting its strong generalization capability and relatively low tendency to overfit noisy datasets [136,137]. Different conclusions emerge when the composition of the feedstock changes substantially over time. In such cases, RF frequently achieves more robust predictions because averaging across multiple decision trees reduces the influence of atypical observations and local fluctuations within the training data [138].
Recent publications increasingly highlight eXtreme GBM as an effective alternative to conventional GBM. Owing to algorithmic improvements and regularization procedures, XGBoost has shown particularly promising results in early warning applications. Studies by Abubakar et al. and Choi et al. demonstrated accurate prediction of VFA accumulation, allowing deterioration of reactor performance to be anticipated before conventional monitoring parameters indicated process imbalance [139,140].
Performance comparisons also reveal an important practical limitation shared by many classical ML models. Predictive accuracy commonly decreases when reactors experience operating conditions that are poorly represented in the training dataset. Rutland et al. illustrated this phenomenon during prediction of the FOS/TAC ratio under abrupt changes in temperature and OLR. While the linear kernel SVM exhibited a noticeable decline in reliability, tree-based ensemble methods retained considerably better predictive performance throughout the disturbance period [141]. These observations suggest that ensemble algorithms may better accommodate the variability characteristic of industrial AD systems. More broadly, classical ML models generally have limited ability to extrapolate beyond the range of conditions represented in the training data, making their performance susceptible to degradation when operating conditions change in industrial AD systems.
Classical ML models’ simplicity can nevertheless constrain their practical deployment. The reviewed studies show differences in SVM performance depending on the selected kernel: while appropriate kernel functions can strengthen the model’s ability to generalize, the use of a linear kernel under nonlinear AD conditions may lead to substantial performance degradation compared with more adaptable ensemble methods.
Although more sophisticated neural architectures are increasingly being introduced into AD research, classical ML algorithms continue to represent an essential benchmark. Their relatively simple implementation, modest computational requirements, and consistent performance explain why they remain widely employed both as standalone predictive models and as reference methods for evaluating newer AI approaches.

4.2. Artificial Neural Networks

ANNs have become one of the most widely applied modelling tools in AD because they can approximate complex nonlinear relationships without requiring explicit biochemical equations. A typical feed-forward network consists of an input layer, one or more hidden layers, and an output layer. Information propagates only in the forward direction, where each neuron transforms the incoming signal through an activation function before passing it to the next layer [84]. Such architectures are well suited to describing nonlinear interactions among operational variables, including pH, temperature, OLR, and gas production.
The absence of internal memory, however, limits the ability of conventional feed-forward networks to represent delayed biological responses that are characteristic of AD. Changes in substrate composition or reactor loading often influence methane production only after several hours or even days, making current process conditions dependent on previous system states. One way to incorporate this temporal dependency is to supply the network with historical observations instead of single measurements. The simplest solution relies on sliding windows, where consecutive sequences of past measurements are treated as network inputs [86]. A more formal implementation is provided by the Time Delay Neural Network (TDNN), originally proposed by Waibel et al., which preserves the feed-forward structure while introducing delayed input connections to capture temporal relationships without recurrent feedback loops [85].
Feed-forward ANNs have been applied extensively to predict methane production, estimate reactor performance, and identify changes in process stability from routinely monitored variables [104,142,143,144,145,146]. Their popularity stems from relatively simple implementation, moderate computational requirements, and the ability to model nonlinear process behaviour using standard operational datasets. Nevertheless, studies employing TDNNs or other delay-based feed-forward architectures remain surprisingly limited, despite the inherently dynamic nature of AD. One of the few reported examples was presented by Schroer and Just, who demonstrated that incorporating delayed process information improved predictive performance compared with models relying exclusively on instantaneous measurements [106].
While conventional feed-forward ANNs offer computational simplicity and strong nonlinear mapping capabilities, their lack of internal memory restricts their ability to model delayed process responses. Intermediate approaches, such as TDNNs and sliding-window architectures, can partially address this limitation without the computational complexity of more advanced DL architectures, yet they remain underutilized in the current AD literature.
The limited use of time-aware feed-forward architectures has gradually shifted attention toward DL models specifically designed to analyse sequential data. These networks extend the capabilities of conventional ANNs by learning long-term temporal dependencies directly from process histories, reducing the need for manually selecting historical input windows.

4.3. DL for Sequential Data

The dynamic nature of AD makes prediction considerably more challenging than in many conventional industrial processes. Reactor performance depends not only on current operating conditions but also on biological responses that develop over extended periods. Microbial adaptation, substrate degradation, and inhibitor accumulation introduce temporal dependencies that cannot be fully represented by conventional feed-forward neural networks. DL architectures have therefore become increasingly important for analysing sequential process data and forecasting future reactor behaviour.
RNNs were developed to address this limitation by introducing internal memory capable of linking current predictions with previous system states [147,148]. Unlike feed-forward networks, recurrent architectures retain information from earlier observations, allowing temporal relationships to be incorporated directly into the learning process. Their application to AD is nevertheless constrained by the vanishing gradient problem, which reduces the influence of distant observations during training and limits the representation of long-term process dynamics [149].
To improve memory retention, gated recurrent architectures such as LSTM and gated recurrent units (GRU) regulate the flow of information through dedicated update mechanisms [87,88]. These networks selectively preserve relevant historical information while discarding less informative signals, making them well suited to biological systems where delayed responses frequently determine reactor performance. Applications reported in the literature include multistep methane production forecasting, prediction of process inhibition, and continuous assessment of digester stability.
An alternative strategy is based on CNNs adapted for one-dimensional time series. Instead of modelling observations sequentially, one-dimensional CNNs identify local temporal patterns by applying convolutional filters across consecutive sensor measurements [150]. This enables efficient detection of short-term changes associated with fluctuations in pH, VFAs, or OLR. Temporal Convolutional Networks (TCNs) extend this concept by combining causal and dilated convolutions, allowing information from longer time horizons to be incorporated without the sequential computations required by recurrent architectures [151]. As a result, TCNs often provide lower computational costs while maintaining competitive predictive performance.
Several comparative studies illustrate the strengths and limitations of different DL architectures in AD. Alrowais et al. applied a conventional RNN to model the co-digestion of wastewater sludge and mechanically pretreated wheat straw, obtaining an RMSE of 0.0038 and an R2 value close to one for biogas production [152]. McCormick and Villa compared LSTM and one-dimensional CNN models for predicting biogas yield from measurements of organic acids, ammonium concentration, and pH [153]. Although the LSTM achieved the lowest prediction error, the CNN exhibited better generalisation during transient operating conditions, reaching an accuracy of 89%. Oliveira et al. evaluated five DL architectures using the same prediction task and identified the GRU model as the most accurate, reporting the lowest RMSE of 139.1 together with the highest coefficient of determination (R2 = 0.871) [154]. The reported error values should be interpreted within the context of their respective datasets, target variables, and metric definitions rather than treated as directly comparable quantities across studies.
Recent work has also explored combinations of multiple DL architectures. Khan et al. investigated several CNN LSTM hybrid configurations for predicting the performance of a biochar-enhanced AD system using variables including the C/N ratio, TAN, and VFA concentration [155]. The parallel CNN LSTM model achieved predictive accuracy comparable to the individual architectures, while a stacking ensemble combining both sequential and parallel networks further increased the coefficient of determination to 0.94. More recently, Geng et al. demonstrated that a multiscale TCN outperformed conventional LSTM models during COD prediction in industrial wastewater treatment, reducing the RMSE by 30.41% while shortening training time by almost 45% [156].
Although no single DL architecture consistently outperforms all others, recent studies indicate that model selection should primarily reflect the temporal characteristics of the monitored process. Recurrent networks remain advantageous when long-term biological dependencies dominate system behaviour, whereas convolution-based models are often better suited to rapid process fluctuations and applications requiring fast computation.

4.4. Hybrid Gray-Box Models

Despite the rapid development of ML and DL, purely data-driven models remain dependent on the quality and representativeness of training data. Industrial anaerobic digesters rarely operate under constant conditions. Feedstock composition changes over time, microbial communities continuously adapt to new substrates, and sensor measurements are affected by uncertainty. Under such circumstances, models trained exclusively on historical observations may lose predictive accuracy when process conditions move beyond the range represented in the training dataset. Mechanistic models are less sensitive to this limitation because their predictions are constrained by biochemical and physicochemical relationships. Their practical application, however, is restricted by demanding parameter calibration and high computational cost, particularly for complex frameworks such as ADM1 [17,18,67,157].
Hybrid gray box models combine these complementary characteristics within a single predictive framework. Instead of replacing mechanistic descriptions, ML algorithms are integrated to improve selected elements of the simulation while preserving the underlying process representation. This strategy allows biological knowledge to remain embedded in the model while introducing the flexibility required to describe nonlinear behaviour observed under industrial operating conditions [72,93,157,158].
Several implementation strategies have emerged in recent years. Residual learning uses the output of a mechanistic model as a baseline, while a neural network is trained only to predict the remaining modelling error [159]. Surrogate models follow a different philosophy by replacing computationally expensive calculations with neural network approximations capable of reproducing the behaviour of the original model at a fraction of the computational cost [90,160]. Other studies focus on adaptive parameter estimation, where selected kinetic coefficients are continuously updated using incoming process measurements, allowing simulations to remain consistent despite changing reactor conditions [161,162,163]. More recently, physics-informed neural networks (PINNs) have attracted growing attention because physical and biochemical constraints are incorporated directly into the optimisation procedure rather than being imposed after model training [91,92].
The advantages of these approaches become most apparent during full-scale operation. Detailed mechanistic simulations often struggle with noisy industrial datasets, where measurement uncertainty may introduce numerical stiffness and reduce simulation stability [92]. Simplified kinetic models are computationally efficient but frequently fail to reproduce the complexity required for reliable optimisation and control. Gray box models attempt to balance these competing requirements by combining mechanistic consistency with data-driven adaptability [93].
Recent applications demonstrate the practical potential of this concept. Moradvandi et al. developed a surrogate-based Switched Box Jenkins framework for biological wastewater treatment, achieving prediction accuracies between 95% and 97% while reducing simulation time relative to the original mechanistic model [94]. Residual learning has also produced promising results for AD. Dong et al. proposed a two-stage framework for the co-digestion of straw and manure in which a neural network corrected the residual error generated by the mechanistic model, increasing prediction accuracy to an R2 of 0.97 [164]. Comparable improvements have been reported for Physics-Informed Neural Networks. Shaw et al. integrated ADM1 with a PINN architecture, improving prediction accuracy by approximately 25% while reducing model training time compared with a purely data-driven solution [95]. Wang et al. combined modified kinetic equations with PINNs and reported a 74% reduction in RMSE together with an R2 value of 0.994, outperforming a conventional ANN trained using the same experimental data [96].
Rather than representing an alternative to ML, gray box modelling extends its capabilities by incorporating established process knowledge into data-driven prediction. As increasing amounts of operational data become available from industrial digesters, this combination is expected to play an increasingly important role in predictive modelling, process optimisation, and DT development. The algorithms discussed in this chapter differ substantially in their learning mechanisms, computational requirements, interpretability, and suitability for specific AD applications. Table 2 summarizes the main AI and ML algorithm families currently applied in AD together with their principal advantages, limitations, and representative applications reported in the literature.
Table 2. Overview of AI and ML algorithm families applied in AD, including their main advantages, limitations, and representative applications.
While Table 2 outlines the theoretical strengths and inherent limitations of various algorithm families, their effectiveness varies depending on the operational scale and specific prediction targets. To illustrate these applications, Table 3 compiles a comparative overview of recent data-driven applications in AD, detailing their specific reactor scales, digested substrates, target variables, and reported performance metrics.
Table 3. Summary of AI and ML applications in AD process.

5. Soft Sensors

5.1. ML-Based Estimation of Key Variables

Reliable operation of anaerobic digesters depends on continuous information describing the biochemical state of the reactor. Although variables such as VFAs, alkalinity, FOS/TAC ratio, COD, or microbial activity provide valuable insight into process stability, many of them remain difficult to monitor continuously under industrial conditions. Conventional laboratory analyses are accurate but require sampling, sample preparation, and specialized analytical equipment, making them unsuitable for real-time process supervision [79,183]. Physical sensors are available for several operating parameters, yet they frequently require calibration, are susceptible to biofouling, and often exhibit delayed responses under changing reactor conditions [76,78]. These practical limitations have driven increasing interest in soft sensors capable of estimating critical variables indirectly from routinely measured process data.
Soft sensors are mathematical models that infer difficult-to-measure process variables using information collected from conventional sensors together with historical operational data. Instead of performing direct measurements, they establish relationships between easily accessible inputs, including temperature, pH, gas production, flow rate, or OLR, and variables that normally require laboratory analysis [77,78,184,185,186]. Their implementation reduces analytical costs, increases monitoring frequency, and enables continuous supervision without additional hardware installation [186].
Among the variables most frequently estimated by soft sensors, VFAs remain the primary target because their accumulation usually precedes reactor acidification and deterioration of methanogenic activity [66,79]. Early estimation of VFA concentrations allows operators to detect unstable operating conditions before substantial reductions in methane production become apparent [80]. Data-driven fault detection frameworks have demonstrated that deviations between estimated and measured VFA concentrations can successfully identify developing process disturbances, providing an additional decision support tool for reactor management [16,153].
Besides VFAs, soft sensors have also been developed for estimating alkalinity, FOS/TAC ratio, COD, total organic carbon, methane concentration, and indicators describing microbial activity [183,186]. Many of these variables remain impractical for continuous measurement because analytical procedures are expensive, require extensive maintenance, or cannot provide sufficiently rapid responses during rapidly changing operating conditions [76,183]. Indirect estimation therefore offers an attractive alternative for maintaining continuous process supervision while reducing instrumentation requirements.
From an economic and operational perspective, soft sensors can provide an alternative to extensive deployment of dedicated online instrumentation. Physical sensors for variables that are difficult to measure directly may require additional hardware, calibration, maintenance, and protection against biofouling, while laboratory analyses involve sampling and measurement delays [76,78,183]. In contrast, soft sensors use routinely collected process data to provide frequent or near-continuous estimates of otherwise difficult-to-measure variables without requiring a dedicated sensor for each target parameter [186]. This can reduce analytical and maintenance requirements while enabling earlier detection of process deviations and supporting more timely operational decisions [16,153,187]. However, these benefits depend on the availability of reliable input measurements and sufficiently representative data for model development and validation.
Therefore, although soft sensors offer new opportunities for process monitoring, their application remains subject to several important limitations. First, their development may be hindered by the insufficient availability of representative real-process data, which can make it difficult to anticipate and identify problems or disturbances that have not been adequately represented during model development [108]. Moreover, the predictive performance of conventional soft sensors is strongly associated with the selection of auxiliary variables, whose appropriate identification requires substantial process-specific knowledge [186]. Even when suitable input variables are available, data-driven models may still exhibit limited predictive accuracy and generalization, particularly when the training data do not adequately represent the variability of the process [79]. This issue is particularly relevant when models are expected to operate under conditions that differ from those encountered during training. In particular, tree-based approaches have limited extrapolation capability, which may restrict their reliability when operating conditions extend beyond the range represented in the training data [76]. However, the use of microbial community data also introduces additional challenges related to measurement frequency, data dimensionality, and the reproducibility of microbial profiles across reactors and operating conditions.
The range of input data used by soft sensors has expanded considerably during recent years. Conventional process measurements are increasingly complemented by spectroscopic observations obtained using infrared techniques, allowing estimation of variables such as COD, TOC, and VFA directly from spectral signatures without extensive sample preparation [183,188,189]. Electrochemical biosensors provide another promising data source. Microbial electrolysis cells have demonstrated an almost linear relationship between current density and VFA concentration across a broad measurement range, with reported coefficients of determination approaching 0.99 [190,191]. Similar results have been obtained using membraneless microbial fuel cells, where electrical potential exhibited a strong inverse correlation with VFA concentration regardless of substrate composition [192,193]. Such sensing platforms do not replace soft sensors but generate informative input signals that can substantially improve prediction accuracy.
Recent developments have also extended soft sensing beyond conventional physicochemical measurements. Information describing microbial communities, including taxonomic composition and functional characteristics, is increasingly incorporated into predictive frameworks. Feature selection performed using RF has identified bacterial groups such as Chloroflexi, Actinobacteria, Proteobacteria, Fibrobacteres, and Spirochaeta as important predictors of reactor performance despite their relatively low abundance within the microbial community [59,98]. These observations indicate that biologically derived datasets can complement conventional sensor measurements and improve estimation of variables directly associated with process stability.

5.2. Neural Soft Sensors

Although various ML algorithms have been applied as soft sensors, neural network architectures currently represent the dominant approach for estimating variables that cannot be monitored continuously in AD. Their widespread adoption results from the ability to learn complex nonlinear relationships directly from operational data without requiring explicit mathematical descriptions of biochemical pathways. This capability is particularly important in AD, where interactions between microbial populations, substrate composition, and operating conditions continuously modify process behaviour [81,194]. The general workflow of soft sensing in AD is illustrated in Figure 2, where readily available process measurements are used to estimate variables that cannot be monitored continuously.
Figure 2. Conceptual workflow of ML-based soft sensing in AD using easily measured process variables to estimate critical process indicators in real time.
Feed-forward ANNs were among the first neural models introduced for virtual sensing applications. Their ability to approximate nonlinear relationships enables estimation of variables including VFAs, methane concentration, alkalinity, and COD using measurements that are routinely available from conventional instrumentation [81]. Compared with traditional regression methods, neural networks generally provide higher predictive accuracy because they capture interactions that are difficult to describe analytically and remain relatively robust when process measurements contain moderate levels of noise [77,195].
As operational databases became larger and continuous monitoring systems generated longer time series, recurrent neural networks gained increasing attention. Among them, LSTM networks proved particularly suitable for soft sensing because their internal memory allows previous reactor states to influence subsequent predictions [196]. This characteristic reflects the behaviour of AD itself, where disturbances often evolve gradually and their consequences become visible only after prolonged biological adaptation. By incorporating historical observations into the prediction process, LSTM-based soft sensors achieve more reliable estimation of variables affected by delayed microbial responses [64,153].
More recent studies have explored hybrid neural architectures that combine convolutional and recurrent layers within a single framework. One-dimensional CNNs efficiently identify local patterns in continuously acquired sensor signals, whereas recurrent layers preserve longer temporal dependencies. Their integration enables simultaneous extraction of short-term fluctuations and long-term process trends, improving prediction of VFA concentration and overall reactor stability under dynamically changing operating conditions [153]. Similar concepts have also been employed in deep neural networks (DNNs) developed for estimating biogas composition directly from substrate characteristics, illustrating that neural soft sensors can support both process monitoring and production forecasting within the same computational framework [190,195].
Despite their high predictive capability, neural soft sensors remain strongly dependent on data quality. Industrial digesters rarely operate under constant conditions because feedstock composition, microbial activity, and environmental factors continuously evolve. Consequently, prediction accuracy gradually decreases when trained models are applied over extended operating periods without adaptation [77,79]. This phenomenon, commonly referred to as model drift, represents one of the principal obstacles limiting long-term deployment of neural soft sensors in industrial practice.
Several adaptive learning strategies have been proposed to reduce this problem. Stacked supervised autoencoders have been introduced to perform nonlinear feature extraction before the prediction stage, generating compact representations that improve model robustness against changing operating conditions [74]. These latent features can subsequently be combined with extreme learning machine algorithms, forming stacked supervised autoencoder–kernel extreme learning machine (SSAE KELM) frameworks that improve both computational efficiency and prediction accuracy. Reported improvements reached 14.31% for the training dataset and 9.67% for an independent testing dataset compared with hierarchical extreme learning machine (HELM) implementations [74]. Continuous retraining using newly acquired operational data further allows neural soft sensors to adapt to gradual process changes, thereby maintaining reliable performance during long-term reactor operation [64]. Table 4 summarizes representative ML and DL approaches developed for virtual sensing in AD, highlighting their input variables, predicted outputs, and reported predictive performance.
Table 4. Representative AI-based soft sensing applications in AD, including predicted variables, input data, modelling approaches, and reported performance.

5.3. Validation and Deployment Challenges

High prediction accuracy reported for soft sensors does not necessarily translate into reliable operation under industrial conditions. Most published models are developed using data collected from individual reactors operating under relatively narrow ranges of substrates and process conditions [76,83]. When transferred to another installation, differences in reactor design, hydraulic regime, feedstock composition, or operational strategy often reduce predictive performance, making additional model calibration necessary [77,82].
The dynamic character of AD further complicates model generalization. OLR, microbial community composition, substrate characteristics, and environmental conditions continuously evolve throughout reactor operation [28,81,82]. As these factors change, the statistical relationships learned during training gradually become less representative of the current process state. Retraining or periodic recalibration is therefore required to preserve prediction accuracy during long-term operation [64,79].
Reliable soft sensing also depends on the quality of the input data. Measurements acquired from physical sensors may be affected by biofouling, signal drift, mechanical degradation, or calibration errors, whereas laboratory analyses are usually performed at irregular intervals [183]. Missing observations and inconsistent sampling frequencies introduce additional uncertainty into the training data. Since soft sensors infer process variables directly from these measurements, errors originating from data acquisition propagate into the prediction model itself.
Validation procedures differ substantially between published studies. Model performance is frequently evaluated using datasets originating from the same experimental campaign that was used for training, whereas independent external validation remains relatively uncommon [113,182]. This makes objective comparison between different soft sensing approaches difficult and provides only limited information about their robustness under previously unseen operating conditions.
Dataset availability represents another practical limitation. Experimental studies often contain relatively few observations, while industrial datasets are rarely shared because of commercial restrictions or incompatible data acquisition systems [85,195]. As a result, many models are trained on datasets that capture only a small fraction of the variability encountered in full-scale digesters. Wider adoption of harmonized reporting protocols together with standardized data formats would facilitate data exchange and support the development of models with improved transferability between installations [59].
The practical value of a soft sensor ultimately depends on its integration into process control. Estimated variables should support operational decisions before process deterioration becomes visible through conventional monitoring. This concept is increasingly incorporated into advanced control frameworks, where outputs generated by soft sensors are combined with model predictive control or RL to adjust substrate feeding and improve reactor stability under changing operating conditions [64,77,166,168,201,202]. At present, the main challenge is no longer achieving high prediction accuracy under laboratory conditions but maintaining reliable model performance after deployment in continuously evolving industrial environments.

6. ML for Optimization and Control

6.1. Data-Driven Optimization Strategies

Optimization of AD is complicated by the strong interactions between biological, chemical, and operational variables. Evaluating different operating conditions often requires repeated simulations or experimental trials, both of which become increasingly expensive and time-consuming at industrial scale [158,173,203]. These limitations have encouraged the use of ML methods that approximate process behaviour directly from operational data instead of relying exclusively on detailed mechanistic calculations [204]. Figure 3 summarizes the overall architecture of an AI-assisted control framework for AD, linking data acquisition, model prediction, and process control within a continuous feedback loop.
Figure 3. Closed-loop AI framework for AD integrating data preprocessing, predictive modelling, and intelligent process control.
In AI- and ML-based AD optimization, the objective function is defined according to the specific process goal, such as maximizing biogas or methane production, improving methane yield, or maintaining process stability and energy efficiency [182]. The optimization seeks operating conditions that improve the selected objective while accounting for relevant process variables and operational constraints, such as pH, temperature, OLR, HRT, and substrate availability. Thus, the objective function and optimization criteria are application-specific rather than universal across AD systems.
Surrogate models represent one of the most widely adopted solutions. Rather than solving the complete set of biochemical equations, they learn the relationship between measured process variables and selected outputs from historical observations. Once trained, these models can reproduce reactor behaviour with considerably lower computational effort while maintaining satisfactory prediction accuracy within the operating range represented by the available data [16,146,158,205].
The reduction in computational demand makes surrogate models particularly useful for optimization problems that require repeated model evaluations. Trucchia et al. demonstrated that simplified ADM1-based representations enable rapid optimization while preserving sufficient prediction accuracy for engineering applications, thereby overcoming hardware limitations associated with full mechanistic simulations [158]. Similar observations were reported by Mougari et al., who showed that AI-based surrogate models can estimate reactor performance without complete knowledge of the underlying biochemical mechanisms, facilitating their implementation in industrial environments [146]. Bench-scale studies further indicate that these approaches reduce computational costs while supporting operation within process and safety constraints [206].
ML is increasingly applied not only to parameter optimization but also to routine plant operation. Statistical process monitoring combined with ML algorithms enables early detection of deviations that precede process deterioration [207,208]. Because these algorithms identify both linear and nonlinear relationships directly from historical datasets, they can recognize operational patterns that remain difficult to capture using conventional analytical methods [168]. This capability creates opportunities for predictive maintenance by identifying gradual equipment degradation before failures interrupt reactor operation [52,202,209].
Operational scheduling has also benefited from data-driven optimization. Instead of following predefined operating schedules, ML models can estimate how different feeding strategies or operating conditions are likely to influence future reactor performance. Such predictions support decisions regarding substrate loading, maintenance planning, and routine process management while accounting for stability indicators and inhibition phenomena [174,210]. However, despite these advantages, data-driven predictive algorithms remain subject to several important limitations. Practical implementation is heavily constrained by the lack of readily available, continuously monitored parameters in full-scale plants [211]. Additionally, when faced with high-dimensional data paired with small sample sizes, conventional regression and simple neural network models may provide insufficient estimation accuracy for reliable process optimization [195]. Furthermore, many existing models are trained on narrow datasets originating from specific reactors, which limits their transferability and cross-system applicability [212]. These limitations highlight the need to combine data-driven optimization with advanced control strategies, which are discussed in the following section.

6.2. Advanced Control Systems Using ML

Model predictive control (MPC) has become one of the principal frameworks for advanced control of AD because it continuously updates control actions using predictions of future reactor behaviour. Unlike conventional proportional integral derivative (PID) controllers, which respond only after process deviations occur, MPC anticipates future process states over a prediction horizon and adjusts operating parameters before instability develops [167,213,214]. Recent implementations increasingly incorporate ML models into the prediction layer, improving controller performance when reactor dynamics become highly nonlinear or substrate characteristics change over time [166,215].
Current applications extend beyond methane production alone and focus on maintaining stable reactor operation through continuous adjustment of key operating variables. OLR remains the most frequently optimized parameter because excessive substrate addition rapidly promotes VFA accumulation and inhibits methanogenic activity. ML-assisted MPC has also been applied to regulate feeding frequency, mixing intensity, and reactor temperature, allowing control actions to follow changing biological conditions rather than predefined operating schedules [174,216]. This adaptive behaviour is particularly valuable under variable feedstock composition, where fixed control strategies frequently lose effectiveness.
Several studies have demonstrated the practical feasibility of these approaches. Petre et al. reported stable biogas production under fluctuating operating conditions using an MPC framework designed for full-scale applications [210]. Earlier work by Blumensaat et al. showed that predictions generated by ADM1 can successfully support control decisions by forecasting reactor behaviour before process deterioration becomes apparent [167]. More recently, surrogate models have been incorporated into MPC architectures to reduce computational requirements while preserving sufficient prediction accuracy for online implementation [158].
ML also extends the scope of advanced control beyond conventional optimization. Kernel-based algorithms and other nonlinear modelling techniques improve the interpretation of multidimensional process data, enabling earlier recognition of changes associated with microbial inhibition or reactor instability [207]. Bio-inspired optimization algorithms have likewise demonstrated stable control under dynamically changing operating conditions without requiring frequent manual intervention [216]. Flexible modelling strategies that account for microbial dynamics and inhibition phenomena further improve controller robustness during transient operating states [174].
Comparisons with conventional control methods consistently favour data-assisted approaches when reactor conditions become highly variable. PID controllers remain attractive because of their simplicity, yet their performance deteriorates under nonlinear process dynamics and fluctuating substrate composition [203,213]. Advanced controllers such as MPC or non-linear extended prediction self-adaptive control (NEPSAC) maintain more stable operating conditions by continuously updating control actions from current process information [202,217]. The addition of ML further improves predictive capability by identifying operational relationships that are difficult to capture using mechanistic models alone, contributing to more consistent methane production and improved resistance to process disturbances [77,168,182].
However, despite these theoretical advantages, transitioning to industrial applications reveals practical limitations. While industrial-scale anaerobic digesters benefit from simple and robust PID control strategies, the slow biological dynamics, long retention times, process nonlinearities, and interactions between operating variables make fully autonomous control challenging [30,115]. Consequently, advanced computational models and data analytics currently provide value for predictive maintenance, process forecasting, and human-in-the-loop decision support rather than direct, fully automated closed-loop control [115]. A realistic pathway toward industrial implementation is therefore to introduce AI-assisted control progressively, beginning with process forecasting and decision support, followed by predictive maintenance and operator-assisted optimization, before considering fully automated closed-loop control. Such staged implementation allows models to be validated under real operating conditions while limiting the risks associated with direct automated intervention in complex biological systems.

6.3. Reinforcement Learning in AD Control

Unlike supervised learning, RL determines control policies through continuous interaction with the process rather than by learning from labelled datasets. An RL agent observes the current reactor state, performs a control action, and updates its policy according to the received reward. This framework is particularly suitable for AD because process behaviour changes over time and the effects of control actions often become visible only after considerable delays [168,170].
Recent studies have investigated RL for optimization of operational variables including OLR, feeding frequency, and co-digestion ratios. By continuously adapting decisions to changing reactor conditions, RL controllers can improve process stability without requiring explicit mechanistic descriptions of all biological interactions [166]. Their ability to learn directly from process responses also allows identification of nonlinear relationships that are difficult to represent using conventional deterministic models [28,169]. Similar concepts have been applied successfully to bioelectrochemical reactors, where RL-based models maintained reliable long-term predictive performance under variable operating conditions [165].
Despite these encouraging results, practical implementation remains strongly dependent on the quality of the simulation environment used during training. Since direct exploration on industrial digesters is unsafe, RL agents are commonly trained using mechanistic simulators or DTs before deployment [172]. The predictive accuracy of these environments therefore determines the quality of the learned control policy. Numerical simplifications and accumulated modelling errors within ADM1-based simulations may reduce controller performance when transferred to full-scale operation, indicating the need for more accurate surrogate and hybrid simulation frameworks [158].
Another important aspect is reward function design. Maximizing methane production alone may encourage aggressive operating strategies that increase the risk of acidification or ammonia inhibition. For this reason, recent studies formulate reward functions by combining energy production with indicators describing reactor stability, allowing the controller to balance productivity and operational safety [166,167]. Such multi-objective optimization better reflects the practical requirements of industrial biogas plants than single-target control strategies.
Although RL has demonstrated considerable potential, industrial applications remain limited. Reliable deployment requires realistic training environments, carefully designed reward functions, and validation under a wide range of operating conditions before autonomous control can be introduced in commercial digesters [158,168,172].

6.4. Integration with Digital Twins

DTs have become an important tool for developing and validating advanced control strategies in AD, particularly where experimental testing at full scale is limited by economic costs and operational risks [172,175]. Instead of performing potentially hazardous experiments on industrial reactors, operators can evaluate alternative operating conditions in a virtual environment before implementing them in practice. This capability is especially valuable for AD systems, whose behaviour is governed by nonlinear biological interactions that remain difficult to describe accurately using mechanistic models alone [173,176].
Most DTs combine process models with continuously acquired operational data to reproduce the current state of the reactor and predict its future behaviour. The availability of open-source implementations such as PyADM1 has facilitated the development of these platforms and simplified their integration with data-driven algorithms [179]. Incorporating ML models into DTs further improves predictive capabilities by compensating for modelling inaccuracies and adapting simulations to changing operating conditions [171,178]. As a result, virtual reactors can support both process monitoring and evaluation of alternative control strategies without interrupting plant operation [172].
One of the main advantages of DTs is the possibility of performing scenario analysis. Different feeding regimes, substrate compositions, or operating conditions can be evaluated before being introduced into the real reactor, reducing both technical risk and the cost of experimental optimization [171,180]. Model adaptation to specific substrates further increases the practical value of this approach because feeding strategies can be assessed under conditions representative of individual installations rather than idealized laboratory systems [174,181].
The integration of mechanistic knowledge with ML also improves the representation of biological processes that are difficult to capture using conventional simulation alone. Recent studies indicate that AI-assisted DTs can reproduce reactor dynamics with high predictive accuracy while supporting operational decision-making under changing process conditions [167,171]. As computational methods continue to mature, DTs are expected to become an important component of intelligent monitoring and control systems for full-scale AD facilities. The practical implementation of advanced dynamic models, such as LSTM and attention-based architectures, requires careful consideration of the trade-off between DL computational training costs and operational benefits [212,218]. Although these complex DL models can capture strong dynamics and highly nonlinear relationships that may be difficult to represent using simpler approaches, they require advanced optimization techniques, substantial training data, and greater computational resources [212]. Their use may therefore be constrained in facilities with limited monitoring infrastructure, computational capacity, or access to sufficiently representative datasets. At the same time, when appropriately trained and validated, advanced models can support real-time monitoring of key process indicators and earlier detection of process deviations, potentially reducing the delay associated with conventional offline laboratory measurements [187,218]. Therefore, the practical adoption of such models should be evaluated by balancing their predictive and operational benefits against data requirements, computational costs, implementation complexity, and the available monitoring infrastructure. The control strategies discussed throughout this chapter are summarized in Table 5, which compares their methodological foundations, practical advantages, and current implementation challenges.
Table 5. Comparison of conventional and AI-assisted control strategies for AD.

7. Explainable AI (XAI) in AD

The integration of ML and AI has significantly improved the management of complex industrial processes by providing advanced predictive capabilities. However, as these algorithms evolve from transparent, rule-based systems into complex, non-linear architectures, such as DNNs and ensemble methods, they become increasingly opaque, leading to the well-known “black-box” problem. While these advanced models excel at mapping highly intricate patterns within vast datasets, their internal decision-making mechanisms remain hidden from human users. In modern engineering and applied sciences, relying on models that provide high accuracy without underlying rationale poses a significant barrier to real-world deployment, as it obscures the physical and biological realities the models are attempting to represent.
To address the lack of interpretability in complex models, Explainable AI (XAI) has become an essential area of ML research [219,220,221,222,223]. XAI encompasses a suite of methodologies and algorithms designed to make the outputs of complex AI systems transparent, interpretable, and comprehensible to domain experts. The goal of XAI is to supplement model predictions and control actions with interpretable reasoning, allowing users to understand how specific conclusions are derived. Across various high-stakes domains, ranging from medical diagnostics to autonomous infrastructure control, the scientific community and regulatory bodies increasingly recognize explainability as a strict requirement for the safe and reliable application of AI.
In the specific context of AD, where process stability depends on complex microbial dynamics, interpretability is highly important. AI models in AD are frequently tasked with forecasting methane yield or detecting impending biochemical imbalances, such as sudden reactor acidification. Yet, an opaque warning of a system failure is fundamentally insufficient for plant engineers who must enact precise, targeted corrective measures. By integrating XAI frameworks, operators can use predictions to gain a quasi-mechanistic understanding of the data, identifying precisely which physicochemical parameters are driving the algorithm’s behavior.

7.1. Need for Interpretability in Industrial AD

Deploying AI in full-scale AD facilities marks a transition from traditional mechanistic modeling (such as ADM1) to data-driven operational control. Nevertheless, the industrial adoption of advanced ML algorithms is often limited by their inherent opacity. To move from academic simulations to full-scale applications, model interpretability has become a practical requirement driven by operator trust, process safety, and regulatory compliance.
In industrial AD plants, operational decisions, such as adjusting OLR or altering co-digestion ratios, are traditionally guided by the heuristic knowledge of plant operators and established mechanistic principles [224,225]. When a “black-box” ML model (e.g., a DNN or RF) predicts an imminent drop in methane yield, it provides a forecast without a biological or physical rationale [67,225]. If the prediction contradicts the operator’s intuition, the lack of explanatory power often leads to algorithm aversion, where operators dismiss the AI’s warning altogether. While early industrial deployments demonstrated the high predictive capacity of ML for biogas flow [107], subsequent research has emphasized that improving model transparency through XAI is required for establishing operational confidence and ensuring regulatory compliance in safety-critical environments [67].
Explainable AI addresses this issue by identifying which variables drive complex non-linear relationships [127]. For example, in a study predicting the biochemical methane potential of diverse solid substrates, an RF model interpreted via SHAP interaction analysis was used to identify complex variable interactions [226]. It revealed a critical, non-linear interplay between dry matter (moisture) and lignin content: while high moisture in low-lignin biomass acts as a physical pretreatment that facilitates rapid microbial colonization, this beneficial interaction is completely negated in high-lignin biomass, where the recalcitrant lignin matrix remains the absolute bottleneck regardless of hydration levels. Similarly, in a lab-scale sensor-data fusion platform, SHAP feature attribution decoded the complex, conjugate relationship between VFAs and alkalinity [101]. The model quantified that VFAs had a 10-fold stronger impact on methane predictions than pH, explaining to operators how a misleadingly stable pH can be maintained due to the system’s buffering capacity even during severe VFA accumulation. By identifying these specific input variables, the model provided plant engineers with an actionable diagnosis that aligned with their biochemical understanding, thereby establishing trust and enabling precise, targeted interventions rather than blind reliance.
AD is a highly sensitive multiphase biological process. Sudden shifts in reactor conditions can lead to severe process failures, such as severe acidification (souring) due to VFA accumulation, foaming, or ammonia toxicity. In these scenarios, the cost of a false negative (failing to predict a crash) or a false positive (unnecessarily halting feeding) can result in significant economic losses. Safety in industrial AD relies on early, accurate, and understandable warnings.
The necessity of XAI for process safety is evident in predictive maintenance and anomaly detection frameworks. For instance, in a full-scale industrial AD plant, a real-time predictive ANN was integrated into the plant’s SCADA system as an advisory soft sensor, forecasting biogas yield and VFA concentrations one hour ahead [104]. By anticipating short-term operational perturbations from feedstock changes, this XAI-assisted oversight allowed operators to proactively adjust feed rates and steam supply, reducing biogas yield fluctuations from ±18% to ±5% and improving overall operational stability by approximately 23%. The diagnostic value of feature attribution is demonstrated in soft-sensing frameworks designed to detect impending acidification before macroscopic reactor failure. In a pilot study predicting total VFAs, a critical early indicator of digester instability, SHAP analysis of an attention-based TabNet regressor revealed why conventional statistical metrics can be misleading. While global mutual information suggested a strong statistical correlation for pH, SHAP quantified that dissolved carbon dioxide partial pressure was the dominant, sensitive driver of transient acidification, as it directly reflects carbonate equilibrium and acidogenesis rates [101]. Similarly, on a hybrid physical-soft sensor platform, SHAP feature attribution decoded the conjugate buffering relationship between VFAs and alkalinity, explaining to operators how a misleadingly stable pH can be temporarily maintained due to the system’s buffering capacity even during severe, latent VFA accumulation [101]. In this context, XAI acts as a diagnostic safety layer, isolating the exact physiological stressors threatening the microbial consortium.
Beyond internal operations, biogas plants operate within strict environmental and safety regulatory frameworks. Facilities must adhere to stringent guidelines regarding digestate quality, H2S emissions, and overall environmental impact [227]. As AI systems begin to autonomously optimize feeding schedules and predict effluent qualities, regulatory bodies increasingly require model interpretability and explainability as prerequisites for normative compliance and operational safety in such critical infrastructures [127]. In this context, robust verification and auditing mechanisms are essential to ensure that autonomous optimization algorithms do not inadvertently violate environmental standards [225], aligning with the broader, cross-industry necessity for credible, independent AI auditing frameworks in safety-critical domains [228].
If an automated control system alters the feedstock mix resulting in a temporary spike in fugitive emissions or suboptimal digestate sanitation, operators must be able to explain the system’s reasoning to environmental auditors to ensure strict normative compliance [127,225]. Black-box models fail this regulatory scrutiny in safety-critical waste-to-energy environments because they cannot provide a trace of their decision logic, creating significant adoption barriers [22]. XAI frameworks solve this by creating transparent audit trails [22]. By utilizing surrogate models or global explainability metrics, plant managers can demonstrate to regulators that the AI’s control policies are bounded by physical safety limits and base their optimization strategies on valid, compliant operational parameters rather than exploiting unsafe data correlations [225].

7.2. Tools and Methods for Explainability

To address the opacity of advanced predictive models, researchers have developed various XAI methodologies. In the context of industrial AD, these tools are typically classified based on their scope, local explanations (interpreting a single prediction or event) versus global explanations (understanding the overall behavior of the model across all data), and their applicability, ranging from model-agnostic techniques to model-specific architectures [229]. Selecting an appropriate XAI method depends on the operational objective, such as diagnosing an isolated sensor anomaly or extracting generalized biochemical rules from the data.
SHAP is a mathematically rigorous and widely adopted model-agnostic framework in modern XAI [230]. Based on cooperative game theory, SHAP treats the input features (e.g., pH, temperature, OLR) as players in a coalition, and the model’s prediction as the payout [229]. It calculates the exact marginal contribution of each feature to the final prediction by evaluating all possible combinations of inputs. The importance of feature i, denoted as the Shapley value ϕ i , is formulated as equation:
ϕ i = S N { i } | S | ! ( | N | | S | 1 ) ! | N | ! [ υ ( S { i } ) υ ( S ) ]
where N is the set of all features, S is a subset of features without i, and υ ( S ) is the model output for subset S.
In AD applications, SHAP is highly effective because it provides both local and global interpretability [102]. Locally, it can decompose a specific day’s low methane yield prediction, showing operators precisely how much a sudden spike in VFAs negatively impacted the output, counterbalanced by a positive contribution from optimal reactor temperature [224]. Globally, aggregating SHAP values across an entire dataset allows researchers to map out non-linear relationships, such as identifying the exact threshold where an increasing C/N ratio transitions from being beneficial to detrimental [67].
While SHAP provides exact mathematical attribution, its computational cost can be prohibitive for highly complex models evaluated in real-time. Local interpretable model-agnostic explanations (LIME) offers a highly efficient alternative focused exclusively on local interpretability [231]. Instead of calculating all permutations, LIME perturbs the input data of a specific observation (creating slight variations in the operational parameters) and observes how the black-box model’s predictions change [219,231]. It then fits a simple, inherently interpretable model, such as a linear regression or a basic decision tree, around that specific instance to approximate the complex model’s local behavior.
For biogas plant operators, LIME acts as an on-demand diagnostic tool [225]. If an anomaly detection algorithm flags an impending reactor failure, LIME can rapidly generate a local explanation, indicating that the warning was triggered primarily because the current ammonia concentration and HRT cross-multiplied to create inhibitory operating conditions. This allows operators to understand the immediate localized risk without needing to comprehend the entire structure of the neural network.
Surrogate modeling is a global explanation strategy that approximates complex deep learning models using simpler, interpretable rules. In this approach, a highly complex, opaque model (such as an ensemble RF or a DNN) is trained to maximize predictive accuracy on AD sensor data. Subsequently, a simpler, transparent “surrogate” model (like a shallow decision tree or a rule-based algorithm) is trained not on the original dataset, but on the predictions generated by the complex model [219].
The primary objective is for the surrogate to mimic the decision boundaries of the black box as closely as possible. In industrial AD, this technique allows engineers to extract distinct, operational rule sets from complex AI [225]. For example, a surrogate decision tree might reveal that the underlying neural network predominantly divides its safe/unsafe reactor state predictions based on a strict FOS/TAC ratio threshold, followed by temperature stability. These extracted rules can then be integrated directly into a plant’s programmable logic controller or standard operating procedures [104].
Unlike SHAP, LIME, and surrogate models, which are applied post-hoc to already trained models, attention mechanisms are an intrinsic, model-specific form of explainability embedded within the DL architecture [219]. Originally designed for natural language processing, these architectures have become highly relevant in AD for analyzing sequential, time-series data without losing temporal resolution [225]. AD is characterized by high operational inertia; an overload event today might not manifest as a biological crash until several days later [117]. Intrinsic attention mechanisms address this delay by assigning dynamic mathematical weights to historical data points, effectively learning which past time steps are most relevant to predicting the current state. For instance, in forecasting a long-term biogas generation profile under anaerobic co-digestion, a Dual-Stage Attention LSTM integrated with a variable selection network developed by Jeong et al. achieved a 36% relative accuracy improvement over standard LSTMs by dynamically weighting both the temporal lags and the most influential input variables across discontinuous historical datasets [114]. Similarly, advanced transformer-based architectures, such as the iTransformer, have been successfully deployed to forecast daily AD yields and carbon emissions by using self-attention matrices to map long-term, non-linear temporal dependencies [119]. By visualizing these attention weights, operators can trace a predicted drop in biogas production back to its temporal origin. This allows them to confirm whether the model is focusing on yesterday’s feed composition or a temperature shock that occurred 72 h prior, providing a diagnostic window for preventative actions before irreversible acidification and process failure occur [117].

7.3. Interpreting ML Models in AD

The practical value of these XAI methodologies becomes apparent when applied to real-world AD datasets. In this domain, interpreting ML models provides novel biochemical insights and helps refine process control strategies [30]. AD is governed by a highly interdependent matrix of operational and environmental variables [101]. While a highly accurate black-box model can effectively forecast a drop in methane yield or predict an impending biological crash, model interpretation enables the use of these predictions for active process optimization.
By applying global explainability metrics, such as aggregated SHAP values or permutation feature importance, researchers and plant operators can rank the relative impact of various physicochemical parameters across the entire operational envelope [67]. For example, interpreting a methane yield prediction model can reveal non-linear dynamics that traditional linear regression might miss [97]. In an 8-year industrial-scale study, tree-based ensemble models optimized via automated ML significantly outperformed traditional linear models, showing that high-COD waste streams and specific time lags are the primary drivers of biogas yield [28]. SHAP analyses can reveal precise threshold effects; for instance, in a multitask prediction model of methane yield and content, SHAP and permutation importance identified soluble COD, volatile solids/total solids ratio, and OLR as the top three critical factors, demonstrating that while volatile solids/total solids has a positive impact, an OLR exceeding 4 g volatile solids/L/day triggers a steep VFA accumulation that inhibits methanogenic activity [131].
Interpreting models designed for process stability (e.g., early warning systems for reactor souring) is essential for identifying the specific causes of instability among interacting variables [232]. When an ML model flags an anomaly, feature attribution tools can isolate the dominant driver, distinguishing whether the predicted instability is caused by acute ammonia toxicity, sudden VFA accumulation, or a drop in alkalinity. On an integrated physical-soft sensor platform, SHAP analysis successfully decoupled these variables by quantifying that VFA concentration had a 10-fold stronger direct impact on methane predictions than pH, while also identifying the indirect, buffering effects of pH and alkalinity through their strong negative correlation with VFAs (r = −0.93) [101]. Similarly, in a pilot soft-sensing study predicting total VFAs, SHAP dependency curves resolved the limitations of global statistical correlation measures, such as Mutual Information, which mistakenly prioritized TAN, by proving that dissolved carbon dioxide partial pressure is actually the most sensitive, non-linear indicator of transient acidification [75]. By identifying these dominant drivers in real-time, operators can use this information for proactive interventions. For example, by integrating these interpretable ML models directly into a plant’s SCADA and programmable logic controller systems, operators can dynamically adjust feedstock feeding rates and steam settings to reduce gas-yield fluctuations from ±18% to ±5%, improving overall process stability by 23% [104]. Alternatively, multiobjective particle swarm optimization can be used to inversely design the operational parameters (OLR, HRT, temperature, and biochar dosage) for specific, real-world food wastes, generating highly optimized methane yields with experimental validation errors of less than 20% [131].
Beyond standard macroscopic parameters (like pH, temperature, and OLR), the modern digitalization of AD increasingly relies on advanced spectroscopic techniques, such as NIR, MIR, or Raman spectroscopy. These optical sensors are combined with ML algorithms (ranging from PLS to 1D CNN) to provide real-time, inline predictions of substrate composition, digestate quality, or intermediate VFA concentrations [22]. However, spectral models are inherently opaque. They utilize thousands of individual spectral wavelengths as inputs, creating a massive, highly collinear feature space. A major difficulty in deploying these optical sensors is ensuring that the model captures the actual chemical absorption features instead of overfitting to background noise, baseline shifts, or spurious correlations in the calibration data [117].
Interpreting these spectral models involves mapping the algorithm’s predictive logic back to specific wavenumbers or frequency bands. By applying XAI techniques, such as Variable Importance in Projection for PLS or attention mechanisms and SHAP for deep architectures, researchers can determine exactly which spectral peaks the model considers most important [127]. For instance, in the development of a MEMS-based MIR spectrometer utilizing a diamond attenuated total reflection probe, PLS regression successfully predicted individual VFA concentrations, achieving calibration R2 values of 0.965 for acetic acid and 0.956 for propionic acid [117]. Spectral mapping confirmed that the model’s predictive capacity was based on the physical chemistry of the system, aligning with the prominent carbonyl (C=O) absorption peak of the carboxyl group (-COOH) at approximately 5850 nm for total VFA estimation, and using the weaker, unique fingerprint region from 6800 to 9500 nm (corresponding to C-O stretch and O-H bend) to successfully differentiate individual acids. Similarly, when NIR spectroscopy was integrated into anaerobic reactors via diffuse reflectance probes for real-time TAN monitoring, DL and PLS models achieved high accuracy (R2 = 0.91, RMSE = 0.32 g N/L) over a range of 1.5–5.5 g N/L [105]. Model interpretation verified that the algorithms focused on the specific vibrational overtones of nitrogen-hydrogen (N-H) bonds rather than background scattering. NIR spectroscopy-based PLS models calibrated over a 520-day industrial campaign successfully predicted feedstock-specific biogas potential (R2 up to 0.84) by mapping spectral absorption back to crop-specific carbohydrate and lignocellulosic structures. This physical and chemical validation of the model’s learned parameters is essential for guaranteeing the robustness and generalizability of spectral monitoring systems in the highly dynamic environment of a full-scale biogas plant.
Despite the rapid integration of XAI in AD modeling, significant methodological and methodological limitations persist. A known limitation of post-hoc explainability tools like SHAP or LIME is their reliance on statistical correlations, meaning they cannot prove direct physical or biochemical causality. Because ML algorithms are intrinsically restricted to fitting input-output patterns within existing training datasets, they cannot offer a deeper mechanistic representation of the underlying biological pathways. XAI outputs in AD literature are frequently presented as secondary, descriptive interpretations of model accuracy rather than being treated as statistically robust hypotheses subjected to independent experimental or biological validation. This lack of independent verification creates significant adoption barriers in safety-critical environments where operators must thoroughly understand and trust AI-driven decisions.
The practical generalizability of XAI is limited by its heavy reliance on the quality, quantity, and distribution of input data. Models trained on narrow, homogeneous, or lab-scale datasets systematically underperform when deployed in highly heterogeneous and dynamic full-scale operational regimes, leading to biased predictions. In highly transient processes, relying on static global feature importance is structurally flawed, as the dominant process drivers change across different operational phases, such as during reactor start-up versus stable operation. Finally, because conventional data-driven models do not incorporate physical laws, such as the conservation of mass and energy, they may generate biochemically unrealistic or physically implausible predictions. To address these limitations, biokinetic principles or thermodynamic constraints should be integrated directly into the ML pipeline.

8. Challenges, Gaps, and Future Directions

8.1. Data-Related Limitations

The availability of process data has expanded considerably with the wider adoption of online monitoring systems in AD. Despite this progress, the quality of available datasets remains one of the main factors limiting the performance of AI algorithms. Most published studies rely on data collected from laboratory or pilot-scale experiments, where operating conditions are relatively stable and measurements are carefully controlled. Such datasets only partially represent the variability encountered in commercial biogas plants, where fluctuations in substrate composition, hydraulic loading, environmental conditions, and microbial activity continuously modify reactor behaviour [21,62,77].
Another limitation concerns the availability of informative process variables. Continuous measurements are routinely available for temperature, pH, gas flow, or methane concentration, whereas indicators that respond earlier to process disturbances, including individual VFAs, alkalinity, or the FOS/TAC ratio, still depend largely on laboratory analyses [21,47]. As a result, many datasets combine high-frequency sensor measurements with infrequent laboratory observations. These differences in sampling frequency complicate data integration, reduce temporal consistency, and limit the performance of predictive models, particularly when early fault detection is required [77,79,183].
Data quality presents an additional challenge. Measurements collected under anaerobic conditions are affected by sensor drift, biofouling, calibration errors, missing observations, and communication failures [66,76]. These problems introduce uncertainty into historical datasets and may gradually reduce prediction accuracy if they are not detected during preprocessing. Although imputation methods and automated quality-control procedures can partially compensate for incomplete observations, their effectiveness depends on the representativeness of the available data [69,70].
The limited availability of biological information further constrains model development. Microbial communities determine the efficiency and stability of AD, yet high-resolution microbiome datasets remain scarce because sequencing is expensive, technically demanding, and performed much less frequently than routine physicochemical measurements [56,57,59]. Consequently, most current AI models are trained primarily on operational variables and only indirectly account for microbial dynamics.
The absence of standardized datasets remains a broader obstacle to the development of reliable AI tools. Biogas plants differ in reactor design, feedstock composition, sensor configuration, sampling frequency, and data processing protocols, making direct comparison between studies difficult [21]. Publicly available benchmark datasets covering a wide range of operating conditions are still lacking, which limits objective comparison of ML algorithms and slows independent validation of newly proposed methods. Similar limitations have been recognised across recent studies on data-driven modelling and process control, where restricted data availability remains one limitation [20,158].
Improving dataset quality is therefore as important as developing more sophisticated learning algorithms. Greater standardisation of data acquisition, wider availability of long-term operational records, and systematic inclusion of biological measurements would provide a stronger basis for predictive modelling, soft sensing, and intelligent control of AD systems.

8.2. Transferability and Generalization of AI Models

A high prediction accuracy obtained for a single anaerobic digester does not necessarily translate into reliable performance under different operating conditions. ML algorithms learn statistical relationships from the data used during training, and these relationships often reflect the characteristics of one reactor, one feedstock, or one operating strategy rather than general features of AD. As a result, models developed for individual installations frequently lose accuracy when applied to reactors with different substrates, HRT, reactor configurations, or microbial communities [72,77,82].
This limited transferability remains one of the principal obstacles to wider implementation of AI in commercial biogas plants. Unlike laboratory experiments, full-scale facilities operate under continuously changing conditions, including seasonal variations in feedstock composition, fluctuations in OLR, and differences in operational practices. These factors gradually alter the relationships between process variables, making fixed prediction models progressively less reliable [64,79,201].
Several strategies have recently been proposed to improve model generalization. Domain adaptation seeks to adjust algorithms trained in one operating environment so that they remain applicable under different process conditions without requiring complete redevelopment. Transfer learning follows a similar principle by retaining knowledge acquired from previously analysed datasets and adapting only selected parts of the model to new operating conditions. Such approaches reduce the amount of additional training data required while preserving much of the information learned during earlier optimisation [64,72].
Federated learning represents another promising direction for industrial applications. Instead of transferring operational datasets between facilities, participating plants train a shared global model while retaining process data locally. This approach addresses concerns related to data ownership, confidentiality, and industrial privacy, while allowing algorithms to learn from a broader range of operating conditions than would be available at any individual site. Although federated learning has not yet been widely adopted in AD, its successful application in other industrial sectors suggests considerable potential for future biogas research.
Hybrid modelling strategies may further improve robustness across different reactors. Combining mechanistic process knowledge with ML allows the model to exploit both established biochemical relationships and patterns extracted directly from operational data. Rather than replacing mechanistic descriptions such as ADM1, hybrid approaches use them to constrain data-driven predictions, reducing unrealistic extrapolation outside the training domain [20,72,158]. Recent studies indicate that these combined models remain more stable under changing operating conditions than purely empirical approaches [73].
Improving generalization will require more than developing increasingly complex neural architectures. Broader validation across independent facilities, external testing on previously unseen datasets, and evaluation under realistic operating conditions are equally important for demonstrating practical reliability. These validation strategies remain considerably less common than single-site performance assessments but are likely to become an essential element of future AI studies in AD. Despite growing interest in transfer learning, federated learning, and domain adaptation, their practical validation in AD remains limited. Most studies still evaluate model performance using data collected from a single installation or under relatively homogeneous operating conditions, making it difficult to determine whether these approaches can generalize across independent biogas plants [77,82]. This limitation is further reinforced by the lack of standardized data collection protocols, publicly available benchmark datasets, and unified evaluation procedures, which restricts the reproducibility of reported results and objective comparison of AI models developed by different research groups [50,82]. Addressing these challenges is essential for translating promising laboratory-scale developments into reliable industrial applications.

8.3. Practical Challenges for Industrial Implementation

Accurate prediction alone is insufficient for practical implementation of AI in AD. Models intended for routine operation must remain reliable despite sensor degradation, process disturbances, equipment maintenance, and continuously changing reactor conditions. These factors rarely appear simultaneously in laboratory datasets but strongly influence the performance of AI systems operating in commercial facilities [66,76,77].
Sensor reliability represents one of the most important practical challenges. Instruments installed inside anaerobic digesters operate under chemically aggressive conditions and are continuously exposed to suspended solids, biofilm formation, and corrosion. These phenomena gradually reduce measurement quality, producing sensor drift, signal noise, and occasional data loss [66,76]. Since many ML algorithms assume that incoming data follow the same distribution as the training dataset, even moderate changes in sensor behaviour may reduce prediction accuracy and increase the risk of incorrect control decisions.
Maintaining long-term performance therefore requires continuous supervision of both data quality and model behaviour. Automated detection of abnormal measurements, online recalibration, and periodic model updating are becoming increasingly important components of AI-supported monitoring systems [64,77,79]. Instead of rebuilding prediction models after substantial performance deterioration, adaptive learning strategies allow algorithms to incorporate new operating data as reactor conditions evolve. This approach improves long-term stability while reducing the need for manual intervention.
Another important direction is the integration of AI with advanced SCADA systems. Recent studies increasingly combine soft sensors, surrogate models, MPC, RL, and DTs into unified decision-support platforms capable of predicting future process states and recommending corrective actions before operational instability develops [166,210,233,234]. Such architectures move AI beyond process monitoring towards active optimisation of AD.
Computational efficiency also becomes increasingly important as AI systems are integrated into real-time operation. Detailed mechanistic simulations and DNNs may provide excellent predictive performance but often require computational resources that limit their routine application. Reduced-order surrogate models and hybrid approaches address this limitation by preserving the dominant process dynamics while substantially reducing computational demand [20,158]. Faster calculations improve the feasibility of real-time optimisation and facilitate integration with DT environments.
Further progress will depend not only on algorithmic accuracy but also on practical reliability under everyday operating conditions. Future AI systems should be capable of identifying unreliable measurements, recognising operating conditions outside the training domain, estimating prediction uncertainty, and adapting automatically as new operational data become available. Achieving these capabilities will be an important step towards autonomous AD plants supported by AI of Things infrastructures, continuous learning, and standardized data exchange between monitoring, prediction, and control systems.

9. Conclusions

The rapid growth of operational data collected in AD systems has created new opportunities for applying AI to process monitoring, prediction, optimization, and control. Rather than replacing established biochemical knowledge, AI extends its practical use by extracting information from heterogeneous datasets that are difficult to interpret using conventional analytical approaches alone.
Available evidence indicates that successful AI applications depend on the entire data-processing workflow rather than on the learning algorithm itself. Reliable prediction requires high-quality data, appropriate preprocessing, effective feature engineering, and rigorous validation under realistic operating conditions. The combination of sensor measurements, spectroscopic signals, laboratory analyses, and microbial data has considerably expanded the information available for predictive modelling, allowing ML algorithms to identify relationships that remain difficult to capture using mechanistic models alone. Consequently, the greatest benefits are achieved when data acquisition, model development, and decision support are treated as parts of a single integrated framework.
Soft sensors represent one of the most mature applications currently available for industrial practice. Their ability to estimate variables that remain difficult, time-consuming, or expensive to measure online can substantially improve process monitoring without requiring additional analytical equipment. At the same time, advances in neural networks, ensemble learning, and hybrid modelling have improved the prediction of methane production, VFAs, and other indicators of reactor performance. These developments also support modern control strategies based on MPC, RL, and DTs, enabling a gradual transition from passive monitoring toward adaptive process management.
Industrial deployment nevertheless remains challenging. Most published models have been developed using data collected under relatively narrow operating conditions, which limits their transferability between reactors operating with different feedstocks, control strategies, and microbial communities. Long-term performance is further influenced by sensor drift, missing observations, changing operating conditions, and the limited availability of standardized datasets suitable for external validation and benchmarking. Furthermore, the absence of harmonized data collection protocols and publicly available benchmark datasets remains one of the major barriers to objective comparison, reproducibility, and broader industrial adoption of AI models in AD. Moreover, external validation of AI models across multiple full-scale biogas plants operating under diverse conditions is essential to improve their generalizability, reproducibility, and reliable deployment in industrial practice. Addressing these issues is likely to have a greater impact on practical implementation than further improvements in prediction accuracy alone.
The future of AI in AD will therefore depend not only on increasingly sophisticated algorithms, but also on their robustness, reproducibility, transferability across different industrial settings, and successful integration into full-scale operation. Hybrid frameworks integrating mechanistic process knowledge with data-driven learning provide a practical route towards reliable prediction and control under variable operating conditions while preserving the interpretability of established process models.
Future research should also investigate the potential of emerging generative AI, large language models, and agent-based AI systems for AD. These technologies could support the integration and interpretation of heterogeneous process data, facilitate knowledge-based decision-making, and enable more autonomous coordination of monitoring, prediction, optimization, and control tasks at biogas facilities. For example, AI agents could potentially integrate real-time sensor information with historical process data and model predictions to assist operators in identifying process disturbances and selecting appropriate corrective actions. However, their practical deployment will require domain-specific validation, reliable process data, and appropriate safety and operational constraints to ensure that AI-generated recommendations remain consistent with the biological and engineering requirements of AD systems.

Author Contributions

Writing—original draft, M.M., B.T., A.Z. and G.J.; Visualization, M.M., B.T. and A.Z.; Supervision, M.M. and P.J.; Writing—review and editing, M.M., B.T., A.Z., G.J. and P.J. All authors have read and agreed to the published version of the manuscript.

Funding

This work is supported by the National Science Centre Poland Miniatura 9 project number 2025/09/X/ST8/00630: “Investigation of the microbiological efficiency of biogas-to-biomethane conversion” and by the Statutory Funds of Electronics, Telecommunications and Informatics Faculty, Gdansk University of Technology.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

ADAnaerobic Digestion
ADM1Anaerobic Digestion Model No. 1
AIArtificial Intelligence
ANNArtificial Neural Network
BPNNBackpropagation Neural Network
CNNConvolutional Neural Network
CODChemical Oxygen Demand
C/NCarbon-to-Nitrogen ratio
DSTHELMDomain Space Transfer Hierarchical Extreme Learning Machine
DLDeep Learning
DNADeoxyribonucleic Acid
DNNDeep Neural Network
DTDigital Twin
FANFree Ammonia Nitrogen
FFTFast Fourier Transforms
FOS/TACRatio of Volatile Organic Acids to Total Inorganic Carbon
GBMGradient Boosting Machine
GCNGraph Convolutional Network
GRUGated Recurrent Unit
HELMHierarchical Extreme Learning Machine
HRTHydraulic Retention Times
k-NNk-Nearest Neighbors
LIMELocal Interpretable Model-agnostic Explanations
LSTMLong Short-Term Memory
MAPEMean Absolute Percentage Error
MIRMid-Infrared
MLMachine Learning
MPCModel Predictive Control
NEPSACNon-linear Extended Prediction Self-Adaptive Control
NIRNear-Infrared
OLROrganic Loading Rate
PCPrincipal Component
PCAPrincipal Component Analysis
PIDProportional-Integral-Derivative
PINNPhysics-Informed Neural Network
PLSPartial Least Squares
POD-RBFProper Orthogonal Decomposition-Radial Basis Function
RBF-NNRadial Basis Function Neural Networks
RFRandom Forest
RLReinforcement Learning
RMSERoot Mean Square Error
RNARibonucleic Acid
SCADASupervisory Control and Data Acquisition
SHAPShapley Additive Explanations
SMOTESynthetic Minority Over-Sampling Technique
SSAE-KELMStacked Supervised Autoencoder–Kernel Extreme Learning Machine
SVMSupport Vector Machine
TANTotal Ammonia Nitrogen
TCNTemporal Convolutional Networks
TDNNTime Delay Neural Network
RNNRecurrent Neural Network
VAEVariational Autoencoders
VFAVolatile Fatty Acid
XAIExplainable AI

References

  1. European Commission. REPowerEU Plan. Available online: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=COM%3A2022%3A230%3AFIN&qid=1653033742483 (accessed on 28 August 2025).
  2. European Commission. Directorate-General for Energy. Biomethane. Available online: https://energy.ec.europa.eu/topics/renewable-energy/bioenergy/biomethane_en (accessed on 28 August 2025).
  3. Gas for Climate; European Biogas Association & Guidehouse. Manual for National Biomethane Strategies; Gas for Climate: Brussels, Belgium, 2022. [Google Scholar]
  4. Hidalgo, D.; Martín-Marroquín, J.M.; Sánchez-Gatón, M.A.; Pérez-Zapatero, E.; Timmers, R.A. Future Horizons for Biomethane in the Context of the Energy Transition. Fermentation 2025, 11, 653. [Google Scholar] [CrossRef] [Scilit]
  5. Chomać-Pierzecka, E.; Zupok, S.; Ćwik, K.; Bykowski, P. Management Challenges in the Biogas Production Sector in Poland—Current Status, Potential and Perspectives. Energies 2025, 18, 6255. [Google Scholar] [CrossRef] [Scilit]
  6. Sica, D.; Esposito, B.; Supino, S.; Malandrino, O.; Sessa, M.R. Biogas-Based Systems: An Opportunity towards a Post-Fossil and Circular Economy Perspective in Italy. Energy Policy 2023, 182, 113719. [Google Scholar] [CrossRef] [Scilit]
  7. de Almeida, L.; van Zeben, J. Law in the EU’s Circular Energy System; Edward Elgar Publishing: Cheltenham, UK, 2023; pp. 1–292. [Google Scholar] [CrossRef] [Scilit]
  8. Motola, V.; Rejtharova, J.; Scarlat, N.; Hurtig, O.; Buffi, M.; Georgakaki, A.; Letout, S.; Mountraki, A.; Salvucci, R.; Rozsai, M.; et al. Clean Energy Technology Observatory: Advanced Biofuels in the European Union-2024 Status Report on Technology Development, Trends, Value Chains and Markets; Edward Elgar Publishing Limited: Cheltenham, UK, 2024. [Google Scholar]
  9. Uddin, M.M.; Wright, M.M. Anaerobic Digestion Fundamentals, Challenges, and Technological Advances. Phys. Sci. Rev. 2023, 8, 2819–2837. [Google Scholar] [CrossRef] [Scilit]
  10. Alengebawy, A.; Ran, Y.; Osman, A.I.; Jin, K.; Samer, M.; Ai, P. Anaerobic Digestion of Agricultural Waste for Biogas Production and Sustainable Bioenergy Recovery: A Review. Environ. Chem. Lett. 2024, 22, 2641–2668. [Google Scholar] [CrossRef] [Scilit]
  11. Hussain, Z.; Mishra, J.; Vanacore, E. Waste to Energy and Circular Economy: The Case of Anaerobic Digestion. J. Enterp. Inf. Manag. 2020, 33, 817–838. [Google Scholar] [CrossRef] [Scilit]
  12. Van, D.P.; Fujiwara, T.; Tho, B.L.; Toan, P.P.S.; Minh, G.H. A Review of Anaerobic Digestion Systems for Biodegradable Waste: Configurations, Operating Parameters, and Current Trends. Environ. Eng. Res. 2020, 25, 1–17. [Google Scholar] [CrossRef] [Scilit]
  13. Sevillano, C.A.; Pesantes, A.A.; Peña Carpio, E.; Martínez, E.J.; Gómez, X. Anaerobic Digestion for Producing Renewable Energy—The Evolution of This Technology in a New Uncertain Scenario. Entropy 2021, 23, 145. [Google Scholar] [CrossRef] [Scilit]
  14. Wang, W.; Shao, B.; Yuan, C.; Zhang, Z.; Wang, W. Thermophilic Anaerobic Digestion of Food Waste: A Review of Inhibitory Factors, Microbial Community Characteristics, and Optimization Strategies. Recycling 2026, 11, 10. [Google Scholar] [CrossRef] [Scilit]
  15. Kostopoulou, E.; Chioti, A.G.; Tsioni, V.; Sfetsas, T. Microbial Dynamics in Anaerobic Digestion: A Review of Operational and Environmental Factors Affecting Microbiome Composition and Function. Preprints 2023, 2023060299. [Google Scholar] [CrossRef] [Scilit]
  16. Batstone, D.J.; Keller, J.; Angelidaki, I.; Kalyuzhnyi, S.V.; Pavlostathis, S.G.; Rozzi, A.; Sanders, W.T.; Siegrist, H.; Vavilin, V.A. The IWA Anaerobic Digestion Model No 1 (ADM1). Water Sci. Technol. 2002, 45, 65–73. [Google Scholar] [CrossRef] [Scilit]
  17. Xu, Z. ADM1 Parameter Calibration Method Based on Partial Least Squares Regression Framework for Industrial-Scale Anaerobic Digestion Modelling. Doctoral Dissertation, Stellenbosch University, Stellenbosch, South Africa, 2019. [Google Scholar]
  18. Weinrich, S.; Mauky, E.; Schmidt, T.; Krebs, C.; Liebetrau, J.; Nelles, M. Systematic Simplification of the Anaerobic Digestion Model No. 1 (ADM1)—Laboratory Experiments and Model Application. Bioresour. Technol. 2021, 333, 125104. [Google Scholar] [CrossRef] [Scilit]
  19. Donoso-Bravo, A.; Sadino-Riquelme, M.C.; Valdebenito-Rolack, E.; Paulet, D.; Gómez, D.; Hansen, F. Comprehensive ADM1 Extensions to Tackle Some Operational and Metabolic Aspects in Anaerobic Digestion. Microorganisms 2022, 10, 948. [Google Scholar] [CrossRef] [Scilit]
  20. Wade, M.J. Not Just Numbers: Mathematical Modelling and Its Contribution to Anaerobic Digestion Processes. Processes 2020, 8, 888. [Google Scholar] [CrossRef] [Scilit]
  21. Ibarra-Esparza, J.; Ibarra-Esparza, F.E.; González-López, M.E.; Garcia-Gonzalez, A.; Gradilla-Hernández, M.S. Instrumentation and Continuous Monitoring for the Anaerobic Digestion Process: A Systematic Review. IEEE Access 2025, 13, 193820–193837. [Google Scholar] [CrossRef] [Scilit]
  22. Marycz, M.; Turowska, I.; Glazik, S.; Jasiński, P. Artificial Intelligence in Anaerobic Digestion: A Review of Sensors, Modeling Approaches, and Optimization Strategies. Sensors 2025, 25, 6961. [Google Scholar] [CrossRef] [Scilit]
  23. Pilarski, K.; Pilarska, A.A. Kinetics and Energy Yield in Anaerobic Digestion: Effects of Substrate Composition and Fundamental Operating Conditions. Energies 2025, 18, 6262. [Google Scholar] [CrossRef] [Scilit]
  24. Boe, K.; Batstone, D.J.; Angelidaki, I. An Innovative Online VFA Monitoring System for the Anaerobic Process, Based on Headspace Gas Chromatography. Biotechnol. Bioeng. 2007, 96, 712–721. [Google Scholar] [CrossRef] [Scilit]
  25. Yildirim, O.; Ozkaya, B. Prediction of Biogas Production of Industrial Scale Anaerobic Digestion Plant by Machine Learning Algorithms. Chemosphere 2023, 335, 138976. [Google Scholar] [CrossRef] [Scilit]
  26. Meola, A.; Weinrich, S. Full-Scale Dynamic Anaerobic Digestion Process Simulation with Machine and Deep Learning Algorithms at Intra-Day Resolution. Appl. Energy 2025, 390, 125781. [Google Scholar] [CrossRef] [Scilit]
  27. Xu, M.; Monson, C.; Kpodo, J.; Long, F.; Liu, H.; Liu, Y.; Liao, W. Artificial Intelligence for Anaerobic Digestion: Advancing Sustainable Biogas Production. Energy Adv. 2026, 5, 1021–1035. [Google Scholar] [CrossRef] [Scilit]
  28. Wang, Y.; Huntington, T.; Scown, C.D. Tree-Based Automated Machine Learning to Predict Biogas Production for Anaerobic Co-Digestion of Organic Waste. ACS Sustain. Chem. Eng. 2021, 9, 12990–13000. [Google Scholar] [CrossRef] [Scilit]
  29. Lamichhane, S.; Subedi, A.; Maharjan, R.; Jha, N.K.; Paudel, S.R. Applications of Machine Learning for Prediction and Optimization of Biogas Production: A Mini-Review. H2Open J. 2026, 9, 100048. [Google Scholar] [CrossRef] [Scilit]
  30. Rutland, H.; You, J.; Liu, H.; Bull, L.; Reynolds, D. A Systematic Review of Machine-Learning Solutions in Anaerobic Digestion. Bioengineering 2023, 10, 1410. [Google Scholar] [CrossRef] [Scilit]
  31. Galal, O.; Abdel-Daiem, M.; Alharbi, H.; Said, N. Mathematical Modeling and Machine Learning Approaches for Biogas Production from Anaerobic Digestion: A Review. Bioresources 2025, 20, 11237–11266. [Google Scholar] [CrossRef] [Scilit]
  32. Banerjee, S.; Lahiri, S.K. Digital Evolution of Bioreactor’s Trends and Frontiers in Artificial Intelligence/Machine Learning-Driven Process Intelligence. Syst. Microbiol. Biomanuf. 2026, 6, 71–93. [Google Scholar] [CrossRef] [Scilit]
  33. Butean, A.; Cutean, I.; Barbero, R.; Enriquez, J.; Matei, A. A Review of Artificial Intelligence Applications for Biorefineries and Bioprocessing: From Data-Driven Processes to Optimization Strategies and Real-Time Control. Processes 2025, 13, 2544. [Google Scholar] [CrossRef] [Scilit]
  34. Amirkhanov, B.; Kunelbayev, M.; Sabina, I.; Amirkhanova, G.; Nurgazy, T.; Zhumasheva, A.; Alipbeki, O. Enhancing Sustainable Biogas Generation Through a Real-Time Digital Twin of a Modular Bioreactor. J. Appl. Data Sci. 2025, 6, 2312–2331. [Google Scholar] [CrossRef] [Scilit]
  35. Abourida, M.; Short, M.; Klymenko, O.V.; Khamis, N.M.; Mahammedi, C.; Al-Mhdawi, M.K.S.; Sakr, A.H. The Evolving Landscape of AI-Driven Risk Management in the Biogas Production: A Systematic and Bibliometric Review. Waste Manag. Bull. 2026, 4, 100271. [Google Scholar] [CrossRef] [Scilit]
  36. Lohani, S.P.; Havukainen, J. Anaerobic Digestion: Factors Affecting Anaerobic Digestion Process. In Biogas: Fundamentals, Processes, and Technologies; Energy, Environment, and Sustainability; Springer: Singapore, 2018; pp. 343–359. [Google Scholar] [CrossRef] [Scilit]
  37. Korres, N.E.; Nizami, A.S. Variation in Anaerobic Digestion: Need for Process Monitoring. In Bioenergy Production by Anaerobic Digestion; Routledge: London, UK, 2013; pp. 194–230. [Google Scholar]
  38. Meegoda, J.N.; Li, B.; Patel, K.; Wang, L.B. A Review of the Processes, Parameters, and Optimization of Anaerobic Digestion. Int. J. Environ. Res. Public Health 2018, 15, 2224. [Google Scholar] [CrossRef] [Scilit]
  39. Náthia-Neves, G.; Berni, M.; Dragone, G.; Mussatto, S.I.; Forster-Carneiro, T. Anaerobic Digestion Process: Technological Aspects and Recent Developments. Int. J. Environ. Sci. Technol. 2018, 15, 2033–2046. [Google Scholar] [CrossRef] [Scilit]
  40. Lindner, J.; Zielonka, S.; Oechsner, H.; Lemmer, A. Effect of Different PH-Values on Process Parameters in Two-Phase Anaerobic Digestion of High-Solid Substrates. Environ. Technol. 2015, 36, 198–207. [Google Scholar] [CrossRef] [Scilit]
  41. Lin, Q.; De Vrieze, J.; Li, J.; Li, X. Temperature Affects Microbial Abundance, Activity and Interactions in Anaerobic Digestion. Bioresour. Technol. 2016, 209, 228–236. [Google Scholar] [CrossRef] [Scilit]
  42. Lin, Q.; He, G.; Rui, J.; Fang, X.; Tao, Y.; Li, J.; Li, X. Microorganism-Regulated Mechanisms of Temperature Effects on the Performance of Anaerobic Digestion. Microb. Cell Fact. 2016, 15, 96. [Google Scholar] [CrossRef] [Scilit]
  43. Donoso-Bravo, A.; Retamal, C.; Carballa, M.; Ruiz-Filippi, G.; Chamy, R. Influence of Temperature on the Hydrolysis, Acidogenesis and Methanogenesis in Mesophilic Anaerobic Digestion: Parameter Identification and Modeling Application. Water Sci. Technol. 2009, 60, 9–17. [Google Scholar] [CrossRef] [Scilit]
  44. Guštin, S.; Marinšek-Logar, R. Effect of PH, Temperature and Air Flow Rate on the Continuous Ammonia Stripping of the Anaerobic Digestion Effluent. Process Saf. Environ. Prot. 2011, 89, 61–66. [Google Scholar] [CrossRef] [Scilit]
  45. Chen, S.; Zhang, J.; Wang, X. Effects of Alkalinity Sources on the Stability of Anaerobic Digestion from Food Waste. Waste Manag. Res. 2015, 33, 1033–1040. [Google Scholar] [CrossRef] [Scilit]
  46. Björnsson, L.; Murto, M.; Mattiasson, B. Evaluation of Parameters for Monitoring an Anaerobic Co-Digestion Process. Appl. Microbiol. Biotechnol. 2000, 54, 844–849. [Google Scholar] [CrossRef] [Scilit]
  47. Boe, K.; Batstone, D.J.; Steyer, J.P.; Angelidaki, I. State Indicators for Monitoring the Anaerobic Digestion Process. Water Res. 2010, 44, 5973–5980. [Google Scholar] [CrossRef] [Scilit]
  48. Casallas-Ojeda, M.; Meneses-Bejarano, S.; Urueña-Argote, R.; Marmolejo-Rebellón, L.F.; Torres-Lozada, P. Techniques for Quantifying Methane Production Potential in the Anaerobic Digestion Process. Waste Biomass Valorization 2021, 13, 2493–2510. [Google Scholar] [CrossRef] [Scilit]
  49. Molina, F.; Castellano, M.; García, C.; Roca, E.; Lema, J.M. Selection of Variables for On-Line Monitoring, Diagnosis, and Control of Anaerobic Digestion Processes. Water Sci. Technol. 2009, 60, 615–622. [Google Scholar] [CrossRef] [Scilit]
  50. Jimenez, J.; Latrille, E.; Harmand, J.; Robles, A.; Ferrer, J.; Gaida, D.; Wolf, C.; Mairet, F.; Bernard, O.; Alcaraz-Gonzalez, V.; et al. Instrumentation and Control of Anaerobic Digestion Processes: A Review and Some Research Challenges. Rev. Environ. Sci. Biotechnol. 2015, 14, 615–648. [Google Scholar] [CrossRef] [Scilit]
  51. Lamb, J.J.; Bernard, O.; Sarker, S.; Lien, K.M.; Hjelme, D.R. Perspectives of Optical Colourimetric Sensors for Anaerobic Digestion. Renew. Sustain. Energy Rev. 2019, 111, 87–96. [Google Scholar] [CrossRef] [Scilit]
  52. Madsen, M.; Holm-Nielsen, J.B.; Esbensen, K.H. Monitoring of Anaerobic Digestion Processes: A Review Perspective. Renew. Sustain. Energy Rev. 2011, 15, 3141–3155. [Google Scholar] [CrossRef] [Scilit]
  53. Eccleston, R.; Wolf, C.; Balsam, M.; Schulte, F.; Bongards, M.; Rehorek, A. Mid-Infrared Spectroscopy for Monitoring of Anaerobic Digestion Processes-Prospects and Challenges. Chem. Eng. Technol. 2016, 39, 627–636. [Google Scholar] [CrossRef] [Scilit]
  54. Xu, Y.; Liu, J.; Sun, Y.; Chen, S.; Miao, X. Fast Detection of Volatile Fatty Acids in Biogas Slurry Using NIR Spectroscopy Combined with Feature Wavelength Selection. Sci. Total Environ. 2023, 857, 159282. [Google Scholar] [CrossRef] [Scilit]
  55. De Vrieze, J.; Verstraete, W. Perspectives for Microbial Community Composition in Anaerobic Digestion: From Abundance and Activity to Connectivity. Environ. Microbiol. 2016, 18, 2797–2809. [Google Scholar] [CrossRef] [Scilit]
  56. Nelson, M.C.; Morrison, M.; Yu, Z. A Meta-Analysis of the Microbial Diversity Observed in Anaerobic Digesters. Bioresour. Technol. 2011, 102, 3730–3739. [Google Scholar] [CrossRef] [Scilit]
  57. Upadhyay, A.; Kovalev, A.A.; Zhuravleva, E.A.; Kovalev, D.A.; Litti, Y.V.; Masakapalli, S.K.; Pareek, N.; Vivekanand, V. A Review of Basic Bioinformatic Techniques for Microbial Community Analysis in an Anaerobic Digester. Fermentation 2023, 9, 62. [Google Scholar] [CrossRef] [Scilit]
  58. De Vrieze, J.; Pinto, A.J.; Sloan, W.T.; Ijaz, U.Z. The Active Microbial Community More Accurately Reflects the Anaerobic Digestion Process: 16S RRNA (Gene) Sequencing as a Predictive Tool. Microbiome 2018, 6, 63. [Google Scholar] [CrossRef] [Scilit]
  59. Long, F.; Wang, L.; Cai, W.; Lesnik, K.; Liu, H. Predicting the Performance of Anaerobic Digestion Using Machine Learning Algorithms and Genomic Data. Water Res. 2021, 199, 117182. [Google Scholar] [CrossRef] [Scilit]
  60. Beale, D.J.; Karpe, A.V.; McLeod, J.D.; Gondalia, S.V.; Muster, T.H.; Othman, M.Z.; Palombo, E.A.; Joshi, D. An ‘Omics’ Approach towards the Characterisation of Laboratory Scale Anaerobic Digesters Treating Municipal Sewage Sludge. Water Res. 2016, 88, 346–357. [Google Scholar] [CrossRef] [Scilit]
  61. Salehi Jouzani, G.; Sharafi, R. New “Omics” Technologies and Biogas Production. In Biogas: Fundamentals, Process, and Operation; Springer: Cham, Switzerland, 2018; pp. 419–436. [Google Scholar] [CrossRef] [Scilit]
  62. Weinrich, S.; Nelles, M. Systematic Simplification of the Anaerobic Digestion Model No. 1 (ADM1)—Model Development and Stoichiometric Analysis. Bioresour. Technol. 2021, 333, 125124. [Google Scholar] [CrossRef] [Scilit]
  63. Rosén, C.; Jeppsson, U. Aspects on ADM1 Implementation Within the BSM2 Framework; Department of Industrial Electrical Engineering and Automation, Lund Institute of Technology: Lund, Sweden, 2016. [Google Scholar]
  64. Park, J.G.; Jun, H.B.; Heo, T.Y. Retraining Prior State Performances of Anaerobic Digestion Improves Prediction Accuracy of Methane Yield in Various Machine Learning Models. Appl. Energy 2021, 298, 117250. [Google Scholar] [CrossRef] [Scilit]
  65. Cruz, I.A.; Andrade, L.R.S.; Bharagava, R.N.; Nadda, A.K.; Bilal, M.; Figueiredo, R.T.; Ferreira, L.F.R. An Overview of Process Monitoring for Anaerobic Digestion. Biosyst. Eng. 2021, 207, 106–119. [Google Scholar] [CrossRef] [Scilit]
  66. Singh, A.; Kumar, V. Recent Developments in Monitoring Technology for Anaerobic Digesters: A Focus on Bio-Electrochemical Systems. Bioresour. Technol. 2021, 329, 124937. [Google Scholar] [CrossRef] [Scilit]
  67. Gupta, R.; Zhang, L.; Hou, J.; Zhang, Z.; Liu, H.; You, S.; Sik Ok, Y.; Li, W. Review of Explainable Machine Learning for Anaerobic Digestion. Bioresour. Technol. 2023, 369, 128468. [Google Scholar] [CrossRef] [Scilit]
  68. Simeonov, I.; Hubenov, V. Application of Artificial Intelligence for Prediction, Monitoring, Optimization and Control of Anaerobic Digestion Processes—A Review. Processes 2025, 13, 3812. [Google Scholar] [CrossRef] [Scilit]
  69. Seu, K.; Kang, M.S.; Lee, H. An Intelligent Missing Data Imputation Techniques: A Review. JOIV Int. J. Inform. Vis. 2022, 6, 278–283. [Google Scholar] [CrossRef] [Scilit]
  70. Sun, Y.; Li, J.; Xu, Y.; Zhang, T.; Wang, X. Deep Learning versus Conventional Methods for Missing Data Imputation: A Review and Comparative Study. Expert Syst. Appl. 2023, 227, 120201. [Google Scholar] [CrossRef] [Scilit]
  71. Kazadi Mbamba, C.; Keymer, P.; Alvi, M.; Topalian, S.O.N.; Ud Din, F.; Batstone, D.J. Enhancing Data Quality in Wastewater Processes: Missing Data Imputation with Deep Variational Autoencoders and Genetic Algorithms. Comput. Chem. Eng. 2025, 199, 109123. [Google Scholar] [CrossRef] [Scilit]
  72. De Buck, V.; Sbarciog, M.I.; Cras, J.; Bhonsale, S.S.; Polanska, M.; Van Impe, J.F.M. Critical Analysis of the Use of White-Box versus Black-Box Models for Multi-Objective Optimisation of Small-Scale Biorefineries. Front. Food Sci. Technol. 2023, 3, 1154305. [Google Scholar] [CrossRef] [Scilit]
  73. Meola, A.; Weinrich, S. Hybrid Modelling of Dynamic Anaerobic Digestion Process in Full-Scale with LSTM and BMP Measurements Prediction. In Proceedings of the 31st European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN 2023), Bruges, Belgium, 4–6 October 2023. [Google Scholar]
  74. Wang, Y.; Wang, S. Soft Sensor for VFA Concentration in Anaerobic Digestion Process for Treating Kitchen Waste Based on SSAE-KELM. IEEE Access 2021, 9, 36466–36474. [Google Scholar] [CrossRef] [Scilit]
  75. Amangeldy, B.; Boltaboyeva, A.; Tasmurzayev, N.; Baigarayeva, Z.; Imanbek, B.; Getahun, A.J.; Turmakhanbet, D.; Kuatova, M.; Wojcik, W. Explainable Machine Learning for Volatile Fatty Acid Soft-Sensing in Anaerobic Digestion: A Pilot Feasibility Study. Algorithms 2026, 19, 183. [Google Scholar] [CrossRef] [Scilit]
  76. Kazemi, P.; Steyer, J.P.; Bengoa, C.; Font, J.; Giralt, J. Robust Data-Driven Soft Sensors for Online Monitoring of Volatile Fatty Acids in Anaerobic Digestion Processes. Processes 2020, 8, 67. [Google Scholar] [CrossRef] [Scilit]
  77. Jiang, Y.; Yin, S.; Dong, J.; Kaynak, O. A Review on Soft Sensors for Monitoring, Control, and Optimization of Industrial Processes. IEEE Sens. J. 2021, 21, 12868–12881. [Google Scholar] [CrossRef] [Scilit]
  78. Kadlec, P.; Grbić, R.; Gabrys, B. Review of Adaptation Mechanisms for Data-Driven Soft Sensors. Comput. Chem. Eng. 2011, 35, 1–24. [Google Scholar] [CrossRef] [Scilit]
  79. Yan, P.; Shen, B.; Wang, Y. Soft Sensor for VFA Concentration in Anaerobic Digestion Process for Treating Kitchen Waste Based on DSTHELM. IEEE Access 2020, 8, 223618–223625. [Google Scholar] [CrossRef] [Scilit]
  80. Wang, Y.; Yan, P.; Gai, M. Dynamic Soft Sensor for Anaerobic Digestion of Kitchen Waste Based on SGSTGAT. IEEE Sens. J. 2021, 21, 19198–19208. [Google Scholar] [CrossRef] [Scilit]
  81. Pettigrew, L.; Delgado, A. Neural Network-Based Reinforcement Learning Control for Increased Methane Production in an Anaerobic Digestion System. In Proceedings of the 3rd IWA Specialized International Conference Ecotechnologies for Wastewater Treatment, Cambridge, UK, 27–30 June 2016; pp. 27–30. [Google Scholar]
  82. Lavergne, C.; Jeison, D.; Ortega, V.; Chamy, R.; Donoso-Bravo, A. A Need for a Standardization in Anaerobic Digestion Experiments? Let’s Get Some Insight from Meta-Analysis and Multivariate Analysis. J. Environ. Manag. 2018, 222, 141–147. [Google Scholar] [CrossRef] [Scilit]
  83. Dochain, D.; Vanrolleghem, P. Dynamical Modelling & Estimation in Wastewater Treatment Processes. Water Intell. Online 2015, 4, 9781780403045. [Google Scholar] [CrossRef] [Scilit]
  84. Rana, A.; Rawat, A.S.; Bijalwan, A.; Bahuguna, H. Application of Multi Layer (Perceptron) Artificial Neural Network in the Diagnosis System: A Systematic Review. In Proceedings of the 2018 3rd IEEE International Conference on Research in Intelligent and Computing in Engineering, RICE 2018, San Salvador, El Salvador, 22–24 August 2018. [Google Scholar] [CrossRef] [Scilit]
  85. Waibel, A.; Hanazawa, T.; Hinton, G.; Shikano, K.; Lang, K.J. Phoneme Recognition Using Time-Delay Neural Networks. In Backpropagation; Psychology Press: London, UK, 2013; pp. 35–61. [Google Scholar]
  86. Frank, R.J.; Davey, N.; Hunt, S.P. Time Series Prediction and Neural Networks. J. Intell. Robot. Syst. Theory Appl. 2001, 31, 91–103. [Google Scholar] [CrossRef] [Scilit]
  87. Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit]
  88. Chung, J.; Gulcehre, C.; Cho, K.; Bengio, Y. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. arXiv 2014, arXiv:1412.3555. [Google Scholar]
  89. Wold, S.; Esbensen, K.; Geladi, P. Principal Component Analysis. Chemom. Intell. Lab. Syst. 1987, 2, 37–52. [Google Scholar] [CrossRef] [Scilit]
  90. Hornik, K.; Stinchcombe, M.; White, H. Multilayer Feedforward Networks Are Universal Approximators. Neural Netw. 1989, 2, 359–366. [Google Scholar] [CrossRef] [Scilit]
  91. Raissi, M.; Perdikaris, P.; Karniadakis, G.E. Physics-Informed Neural Networks: A Deep Learning Framework for Solving Forward and Inverse Problems Involving Nonlinear Partial Differential Equations. J. Comput. Phys. 2019, 378, 686–707. [Google Scholar] [CrossRef] [Scilit]
  92. Weir, L.; Mathias, N.; Corbett, B.; Mhaskar, P. Noise Aware Parameter Estimation in Bioprocesses: Using Neural Network Surrogate Models with Nonuniform Data Sampling. AIChE J. 2025, 71, e18634. [Google Scholar] [CrossRef] [Scilit]
  93. Zhang, D.; Del Rio-Chanona, E.A.; Petsagkourakis, P.; Wagner, J. Hybrid Physics-Based and Data-Driven Modeling for Bioprocess Online Simulation and Optimization. Biotechnol. Bioeng. 2019, 116, 2919–2930. [Google Scholar] [CrossRef] [Scilit]
  94. Moradvandi, A.; Abraham, E.; Goudjil, A.; De Schutter, B.; Lindeboom, R.E.F. An Identification Algorithm of Switched Box-Jenkins Systems in the Presence of Bounded Disturbances: An Approach for Approximating Complex Biological Wastewater Treatment Models. J. Water Process Eng. 2024, 60, 105202. [Google Scholar] [CrossRef] [Scilit]
  95. Shaw, K.M.; Poh, P.E.; Ho, Y.K.; Chen, Z.Y.; Chew, I.M.L. Modeling the Anaerobic Digestion of Palm Oil Mill Effluent via Physics-Informed Deep Learning. Chem. Eng. J. 2024, 485, 149826. [Google Scholar] [CrossRef] [Scilit]
  96. Wang, Z.; Wang, S.; Zheng, X.; Liu, W.; Shen, Z. Integrating Kinetic Models with Physics-Informed Neural Networks (PINNs) for Predicting Methane Production from Anaerobic Co-Digestion of Enzyme-Modified Biodegradable Plastics and Food Waste Leachate. Water 2025, 17, 3411. [Google Scholar] [CrossRef] [Scilit]
  97. De Clercq, D.; Wen, Z.; Fei, F.; Caicedo, L.; Yuan, K.; Shang, R. Interpretable Machine Learning for Predicting Biomethane Production in Industrial-Scale Anaerobic Co-Digestion. Sci. Total Environ. 2020, 712, 134574. [Google Scholar] [CrossRef] [Scilit]
  98. Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
  99. Angelini, C. Regression Analysis. Encycl. Bioinform. Comput. Biol. ABC Bioinform. 2019, 1–3, 722–730. [Google Scholar] [CrossRef] [Scilit]
  100. Tisocco, S.; Weinrich, S.; Møller, H.B.; Ward, A.J.; Kilmartin, L.; Zhan, X.; Crosson, P. Machine Learning vs. ADM1: Reliable Biogas Prediction with Minimal Data Requirements in Full-Scale Plants. Environ. Sci. Ecotechnol. 2026, 29, 100662. [Google Scholar] [CrossRef] [Scilit]
  101. Wang, X.; Rashid, I.; Zhao, Z.; Oladele, M.; Xiang, W.; Huang, Y.; Wazer, E.; McCutcheon, J.; Bollas, G.; Contreras, J.; et al. Machine Learning Algorithm Integrated with Real-Time In Situ Sensors and Physiochemical Principle-Driven Soft Sensors toward an Anaerobic Digestion-Data Fusion Framework. ACS ES&T Water 2023, 4, 1061–1072. [Google Scholar] [CrossRef] [Scilit]
  102. Zhang, Y.; Li, L.; Ren, Z.; Yu, Y.; Li, Y.; Pan, J.; Lu, Y.; Feng, L.; Zhang, W.; Han, Y. Plant-Scale Biogas Production Prediction Based on Multiple Hybrid Machine Learning Technique. Bioresour. Technol. 2022, 363, 127899. [Google Scholar] [CrossRef] [Scilit]
  103. Bohutskyi, P.; Phan, D.; Kopachevsky, A.M.; Chow, S.; Bouwer, E.J.; Betenbaugh, M.J. Synergistic Co-Digestion of Wastewater Grown Algae-Bacteria Polyculture Biomass and Cellulose to Optimize Carbon-to-Nitrogen Ratio and Application of Kinetic Models to Predict Anaerobic Digestion Energy Balance. Bioresour. Technol. 2018, 269, 210–220. [Google Scholar] [CrossRef] [Scilit]
  104. Zhuang, Z.; Liu, X.; Jin, J.; Li, Z.; Liu, Y.; Tavares, A.; Li, D. Prediction, Uncertainty Quantification, and ANN-Assisted Operation of Anaerobic Digestion Guided by Entropy Using Machine Learning. Entropy 2025, 27, 1233. [Google Scholar] [CrossRef] [Scilit]
  105. Raju, C.S.; Løkke, M.M.; Sutaryo, S.; Ward, A.J.; Møller, H.B. NIR Monitoring of Ammonia in Anaerobic Digesters Using a Diffuse Reflectance Probe. Sensors 2012, 12, 2340–2350. [Google Scholar] [CrossRef] [Scilit]
  106. Schroer, H.W.; Just, C.L. Feature Engineering and Supervised Machine Learning to Forecast Biogas Production during Municipal Anaerobic Co-Digestion. ACS ES&T Eng. 2023, 4, 660–672. [Google Scholar] [CrossRef] [Scilit]
  107. De Clercq, D.; Jalota, D.; Shang, R.; Ni, K.; Zhang, Z.; Khan, A.; Wen, Z.; Caicedo, L.; Yuan, K. Machine Learning Powered Software for Accurate Prediction of Biogas Production: A Case Study on Industrial-Scale Chinese Production Data. J. Clean. Prod. 2019, 218, 390–399. [Google Scholar] [CrossRef] [Scilit]
  108. Kazemi, P.; Bengoa, C.; Steyer, J.P.; Giralt, J. Data-Driven Techniques for Fault Detection in Anaerobic Digestion Process. Process Saf. Environ. Prot. 2021, 146, 905–915. [Google Scholar] [CrossRef] [Scilit]
  109. Corominas, L.; Garrido-Baserba, M.; Villez, K.; Olsson, G.; Cortés, U.; Poch, M. Transforming Data into Knowledge for Improved Wastewater Treatment Operation: A Critical Review of Techniques. Environ. Model. Softw. 2018, 106, 89–103. [Google Scholar] [CrossRef] [Scilit]
  110. Aguado, D.; Alferes, J.; Arteaga, F.; Belia, L.; Copp, J.B.; Corominas, L.; Corona, F.; Ferrer, A.; Haimi, H.; Kazemi, P.; et al. Analytical Methods for Online Data Quality Assessment. In Metadata Collection and Organization in Wastewater Treatment and Wastewater Resource Recovery Systems; IWA Publishing: London, UK, 2024; pp. 163–240. [Google Scholar] [CrossRef] [Scilit]
  111. Mahanty, B.; Zafar, M.; Park, H.S. Characterization of Co-Digestion of Industrial Sludges for Biogas Production by Artificial Neural Network and Statistical Regression Models. Environ. Technol. 2013, 34, 2145–2153. [Google Scholar] [CrossRef] [Scilit]
  112. Goldstein, A.; Kapelner, A.; Bleich, J.; Pitkin, E. Peeking Inside the Black Box: Visualizing Statistical Learning with Plots of Individual Conditional Expectation. J. Comput. Graph. Stat. 2015, 24, 44–65. [Google Scholar] [CrossRef] [Scilit]
  113. Donoso-Bravo, A.; Mailier, J.; Martin, C.; Rodríguez, J.; Aceves-Lara, C.A.; Wouwer, A. Vande Model Selection, Identification and Validation in Anaerobic Digestion: A Review. Water Res. 2011, 45, 5347–5364. [Google Scholar] [CrossRef] [Scilit]
  114. Jeong, K.; Abbas, A.; Shin, J.; Son, M.; Kim, Y.M.; Cho, K.H. Prediction of Biogas Production in Anaerobic Co-Digestion of Organic Wastes Using Deep Learning Models. Water Res. 2021, 205, 117697. [Google Scholar] [CrossRef] [Scilit]
  115. Cinar, S.; Cinar, S.O.; Wieczorek, N.; Sohoo, I.; Kuchta, K. Integration of Artificial Intelligence into Biogas Plant Operation. Processes 2021, 9, 85. [Google Scholar] [CrossRef] [Scilit]
  116. Ekinci, E.; Özbay, B.; Omurca, S.İ.; Sayın, F.E.; Özbay, İ. Application of Machine Learning Algorithms and Feature Selection Methods for Better Prediction of Sludge Production in a Real Advanced Biological Wastewater Treatment Plant. J. Environ. Manag. 2023, 348, 119448. [Google Scholar] [CrossRef] [Scilit]
  117. Eccleston, R. Online Analysis and Optimisation of the Anaerobic Fermentation Process. Doctoral Dissertation, Universität Duisburg-Essen, Duisburg, Germany, 2020. [Google Scholar]
  118. Kimball, R.; Caserta, J. The Data Warehouse ETL Toolkit: Practical Techniques for Extracting, Cleaning, Conforming, and Delivering Data; John Wiley & Sons: Hoboken, NJ, USA, 2004; ISBN 978-0-7645-6757-5. [Google Scholar]
  119. Han, Y.; Zeng, C.; Ni, Q.; Wang, J.; Chu, Z.; Zhang, X.; Geng, Z.; Tan, L.; Liu, Y. Time Series Prediction of Anaerobic Digestion Yield and Carbon Emissions from Food Waste Based on ITransformer Model. Chem. Eng. J. 2025, 513, 163064. [Google Scholar] [CrossRef] [Scilit]
  120. Rousseeuw, P.J.; Leroy, A.M. Robust Regression and Outlier Detection; John Wiley & Sons: Hoboken, NJ, USA, 1987. [Google Scholar] [CrossRef] [Scilit]
  121. Dunia, R.; Qin, S.J.; Edgar, T.F.; McAvoy, T.J. Identification of Faulty Sensors Using Principal Component Analysis. AIChE J. 1996, 42, 2797–2812. [Google Scholar] [CrossRef] [Scilit]
  122. Wall, D.M.; O’Kiely, P.; Murphy, J.D. The Potential for Biomethane from Grass and Slurry to Satisfy Renewable Energy Targets. Bioresour. Technol. 2013, 149, 425–431. [Google Scholar] [CrossRef] [Scilit]
  123. Sousa, A.C.; Lucio, M.M.L.M.; Neto, O.F.B.; Marcone, G.P.S.; Pereira, A.F.C.; Dantas, E.O.; Fragoso, W.D.; Araujo, M.C.U.; Galvão, R.K.H. A Method for Determination of COD in a Domestic Wastewater Treatment Plant by Using Near-Infrared Reflectance Spectrometry of Seston. Anal. Chim. Acta 2007, 588, 231–236. [Google Scholar] [CrossRef] [Scilit]
  124. Kuhn, M.; Johnson, K. Applied Predictive Modeling; Springer: New York, NY, USA, 2013; ISBN 9781461468493. [Google Scholar]
  125. Han, Y.; Du, Z.; Hu, X.; Li, Y.; Cai, D.; Fan, J.; Geng, Z. Production Prediction Modeling of Food Waste Anaerobic Digestion for Resources Saving Based on SMOTE-LSTM. Appl. Energy 2023, 352, 122024. [Google Scholar] [CrossRef] [Scilit]
  126. Elreedy, D.; Atiya, A.F. A Comprehensive Analysis of Synthetic Minority Oversampling Technique (SMOTE) for Handling Class Imbalance. Inf. Sci. 2019, 505, 32–64. [Google Scholar] [CrossRef] [Scilit]
  127. Zamfir, F.S.; Carbureanu, M.; Mihalache, S.F. Application of Machine Learning Models in Optimizing Wastewater Treatment Processes: A Review. Appl. Sci. 2025, 15, 8360. [Google Scholar] [CrossRef] [Scilit]
  128. Torgo, L.; Branco, P.; Ribeiro, R.P.; Pfahringer, B. Resampling Strategies for Regression. Expert Syst. 2015, 32, 465–476. [Google Scholar] [CrossRef] [Scilit]
  129. Zhang, Y.; Jing, Z.; Feng, Y.; Chen, S.; Li, Y.; Han, Y.; Feng, L.; Pan, J.; Mazarji, M.; Zhou, H.; et al. Using Automated Machine Learning Techniques to Explore Key Factors in Anaerobic Digestion: At the Environmental Factor, Microorganisms and System Levels. Chem. Eng. J. 2023, 475, 146069. [Google Scholar] [CrossRef] [Scilit]
  130. Hinton, G.E.; Salakhutdinov, R.R. Reducing the Dimensionality of Data with Neural Networks. Science 2006, 313, 504–507. [Google Scholar] [CrossRef] [Scilit]
  131. Li, J.; Zhang, L.; Li, C.; Tian, H.; Ning, J.; Zhang, J.; Tong, Y.W.; Wang, X. Data-Driven Based In-Depth Interpretation and Inverse of Anaerobic Digestion for CH4-Rich Biogas. ACS ES&T Eng. 2022, 2, 642–652. [Google Scholar] [CrossRef] [Scilit]
  132. Sharifani, K.; Amini, M. Machine Learning and Deep Learning: A Review of Methods and Applications. World Inf. Technol. Eng. J. 2023, 10, 3897–3904. [Google Scholar]
  133. Mienye, I.D.; Sun, Y. A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects. IEEE Access 2022, 10, 99129–99149. [Google Scholar] [CrossRef] [Scilit]
  134. Nguyen, K.A.; Chen, W.; Lin, B.S.; Seeboonruang, U. Comparison of Ensemble Machine Learning Methods for Soil Erosion Pin Measurements. ISPRS Int. J. Geoinf. 2021, 10, 42. [Google Scholar] [CrossRef] [Scilit]
  135. Boateng, E.Y.; Otoo, J.; Abaye, D.A. Basic Tenets of Classification Algorithms K-Nearest-Neighbor, Support Vector Machine, Random Forest and Neural Network: A Review. J. Data Anal. Inf. Process. 2020, 8, 341–357. [Google Scholar] [CrossRef]
  136. Ganeshan, P.; Bose, A.; Lee, J.; Barathi, S.; Rajendran, K. Machine Learning for High Solid Anaerobic Digestion: Performance Prediction and Optimization. Bioresour. Technol. 2024, 400, 130665. [Google Scholar] [CrossRef] [Scilit]
  137. Gupta, R.; Murray, C.; Sloan, W.T.; You, S. Predicting the Methane Production of Microwave-Pretreated Anaerobic Digestion of Food Waste: A Machine Learning Approach. Energy 2025, 328, 136613. [Google Scholar] [CrossRef] [Scilit]
  138. Tryhuba, I.; Tryhuba, A.; Hutsol, T.; Cieszewska, A.; Andrushkiv, O.; Glowacki, S.; Bryś, A.; Slobodian, S.; Tulej, W.; Sojak, M. Prediction of Biogas Production Volumes from Household Organic Waste Based on Machine Learning. Energies 2024, 17, 1786. [Google Scholar] [CrossRef] [Scilit]
  139. Abubakar, U.A.; Lemar, G.S.; Bello, A.A.D.; Ishaq, A.; Dandajeh, A.A.; Jagun, Z.T.; Houmsi, M.R. Evaluation of Traditional and Machine Learning Approaches for Modeling Volatile Fatty Acid Concentrations in Anaerobic Digestion of Sludge: Potential and Challenges. Environ. Sci. Pollut. Res. 2024, 32, 28239–28252. [Google Scholar] [CrossRef] [Scilit]
  140. Choi, S.; Kim, S.I.; Yulisa, A.; Aghasa, A.; Hwang, S. Proactive Prediction of Total Volatile Fatty Acids Concentration in Multiple Full-Scale Food Waste Anaerobic Digestion Systems Using Substrate Characteristics with Machine Learning and Feature Analysis. Waste Biomass Valorization 2022, 14, 593–608. [Google Scholar] [CrossRef] [Scilit]
  141. Rutland, H.; You, J.; Liu, H.; Bowman, K. Application of Machine Learning for FOS/TAC Soft Sensing in Bio-Electrochemical Anaerobic Digestion. Molecules 2025, 30, 1092. [Google Scholar] [CrossRef] [Scilit]
  142. Quashie, F.K.; Fang, A.; Wei, L.; Kabutey, F.T.; Xing, D. Prediction of Biogas Production from Food Waste in a Continuous Stirred Microbial Electrolysis Cell (CSMEC) with Backpropagation Artificial Neural Network. Biomass Convers. Biorefin. 2021, 13, 287–298. [Google Scholar] [CrossRef] [Scilit]
  143. Chen, J.W.; Chan, Y.J.; Arumugasamy, S.K.; Yazdi, S.K. Process Modelling and Optimisation of Methane Yield from Palm Oil Mill Effluent Using Response Surface Methodology and Artificial Neural Network. J. Water Process Eng. 2023, 52, 103493. [Google Scholar] [CrossRef] [Scilit]
  144. Heydari, B.; Abdollahzadeh Sharghi, E.; Rafiee, S.; Mohtasebi, S.S. Use of Artificial Neural Network and Adaptive Neuro-Fuzzy Inference System for Prediction of Biogas Production from Spearmint Essential Oil Wastewater Treatment in up-Flow Anaerobic Sludge Blanket Reactor. Fuel 2021, 306, 121734. [Google Scholar] [CrossRef] [Scilit]
  145. Almomani, F. Prediction of Biogas Production from Chemically Treated Co-Digested Agricultural Waste Using Artificial Neural Network. Fuel 2020, 280, 118573. [Google Scholar] [CrossRef] [Scilit]
  146. Mougari, N.E.; Largeau, J.F.; Himrane, N.; Hachemi, M.; Tazerout, M. Application of Artificial Neural Network and Kinetic Modeling for the Prediction of Biogas and Methane Production in Anaerobic Digestion of Several Organic Wastes. Int. J. Green Energy 2021, 18, 1584–1596. [Google Scholar] [CrossRef] [Scilit]
  147. Jordan, M.I. Serial Order: A Parallel Distributed Processing Approach. Adv. Psychol. 1997, 121, 471–495. [Google Scholar] [CrossRef] [Scilit]
  148. Elman, J.L. Finding Structure in Time. Cogn. Sci. 1990, 14, 179–211. [Google Scholar] [CrossRef]
  149. Bengio, Y.; Simard, P.; Frasconi, P. Learning Long-Term Dependencies with Gradient Descent Is Difficult. IEEE Trans. Neural Netw. 1994, 5, 157–166. [Google Scholar] [CrossRef] [Scilit]
  150. Kiranyaz, S.; Avci, O.; Abdeljaber, O.; Ince, T.; Gabbouj, M.; Inman, D.J. 1D Convolutional Neural Networks and Applications: A Survey. Mech. Syst. Signal Process. 2021, 151, 107398. [Google Scholar] [CrossRef] [Scilit]
  151. Bai, S.; Kolter, J.Z.; Koltun, V. An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling. arXiv 2018, arXiv:1803.01271. [Google Scholar]
  152. Alrowais, R.; Abdel-Daiem, M.M.; Nasef, B.M.; Metwally, A.A.; Said, N. Sustainable Management of Wastewater Sludge Through Co-Digestion, Mechanical Pretreatment and Recurrent Neural Network (RNN) Modeling. Sustainability 2025, 17, 9323. [Google Scholar] [CrossRef] [Scilit]
  153. McCormick, M.; Villa, A.E.P. LSTM and 1-D Convolutional Neural Networks for Predictive Monitoring of the Anaerobic Digestion Process. In Artificial Neural Networks and Machine Learning–ICANN 2019: Deep Learning; Springer: Cham, Switzerland, 2019; pp. 725–736. [Google Scholar] [CrossRef] [Scilit]
  154. Oliveira, P.; Bessa, A.; Pereira, J.; Silva, S.; Duarte, M.S.; Durães, D.; Novais, P. Exploring Transfer Learning’s Impact on the Explainability of Deep Learning Models for Wastewater Treatment Plants’ Biogas Production. Expert Syst. 2026, 43, e70235. [Google Scholar] [CrossRef] [Scilit]
  155. Khan, M.; Surendra, K.C.; Baniya, S.; Rhymer, J.H.; Khanal, S.K. Biochar-Augmented Anaerobic Digestion System: Insights from an Interpretable Stacking Ensemble Deep Learning. Environ. Sci. Technol. 2025, 59, 15236–15250. [Google Scholar] [CrossRef] [Scilit]
  156. Geng, Y.; Zhang, F.; Liu, H. Multi-Scale Temporal Convolutional Networks for Effluent COD Prediction in Industrial Wastewater. Appl. Sci. 2024, 14, 5824. [Google Scholar] [CrossRef] [Scilit]
  157. Heiker, M.; Kraume, M.; Mertins, A.; Wawer, T.; Rosenberger, S. Biogas Plants in Renewable Energy Systems—A Systematic Review of Modeling Approaches of Biogas Production. Appl. Sci. 2021, 11, 3361. [Google Scholar] [CrossRef] [Scilit]
  158. Trucchia, A.; Frunzo, L. Surrogate Based Global Sensitivity Analysis of ADM1-Based Anaerobic Digestion Model. J. Environ. Manag. 2021, 282, 111456. [Google Scholar] [CrossRef] [Scilit]
  159. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
  160. Jin, R.; Chen, W.; Simpson, T.W. Comparative Studies of Metamodelling Techniques under Multiple Modelling Criteria. Struct. Multidiscip. Optim. 2001, 23, 1–13. [Google Scholar] [CrossRef] [Scilit]
  161. Kumpati, S.N.; Parthasarathy, K. Identification and Control of Dynamical Systems Using Neural Networks. IEEE Trans. Neural Netw. 1990, 1, 4–27. [Google Scholar] [CrossRef] [Scilit]
  162. Kalman, R.E. A New Approach to Linear Filtering and Prediction Problems. J. Basic Eng. 1960, 82, 35–45. [Google Scholar] [CrossRef] [Scilit]
  163. Rodríguez, A.; Quiroz, G.; Femat, R.; Méndez-Acosta, H.O.; de León, J. An Adaptive Observer for Operation Monitoring of Anaerobic Digestion Wastewater Treatment. Chem. Eng. J. 2015, 269, 186–193. [Google Scholar] [CrossRef] [Scilit]
  164. Dong, Z.; Liu, P.; Cao, S.; Zhao, B.; Wang, Y.; Wang, L.; Li, N. Data-Driven Prediction and Optimization of Straw-Manure Co-Anaerobic Digestion Based on a Two-Stage Residual Learning Framework. J. Environ. Chem. Eng. 2026, 14, 122869. [Google Scholar] [CrossRef] [Scilit]
  165. Xiao, J.; Liu, C.; Ju, B.; Xu, H.; Sun, D.; Dang, Y. Estimation of In-Situ Biogas Upgrading in Microbial Electrolysis Cells via Direct Electron Transfer: Two-Stage Machine Learning Modeling Based on a NARX-BP Hybrid Neural Network. Bioresour. Technol. 2021, 330, 124965. [Google Scholar] [CrossRef] [Scilit]
  166. Yesilevskyi, V.; Dyadun, S.; Kuznetsov, V. LSTM Networks for Anaerobic Digester Control. Sci. Bull. Natl. Min. Univ. 2019, 2019, 130. [Google Scholar] [CrossRef] [Scilit]
  167. Blumensaat, F.; Keller, J. Modelling of Two-Stage Anaerobic Digestion Using the IWA Anaerobic Digestion Model No. 1 (ADM1). Water Res. 2005, 39, 171–183. [Google Scholar] [CrossRef] [Scilit]
  168. Mahmoud, D.; Magolon, M.; Boer, J.; Elbestawi, M.A.; Mohammadi, M.G. Applications of Machine Learning in Process Monitoring and Controls of L-PBF Additive Manufacturing: A Review. Appl. Sci. 2021, 11, 11910. [Google Scholar] [CrossRef] [Scilit]
  169. Batstone, D.J. Mathematical Modelling of Anaerobic Reactors Treating Domestic Wastewater: Rational Criteria for Model Use. Rev. Environ. Sci. Biotechnol. 2006, 5, 57–71. [Google Scholar] [CrossRef] [Scilit]
  170. Spielberg, S.; Tulsyan, A.; Lawrence, N.P.; Loewen, P.D.; Gopaluni, R.B. Deep Reinforcement Learning for Process Control: A Primer for Beginners. AIChE J. 2020, 65, e16689. [Google Scholar] [CrossRef] [Scilit]
  171. Ahmed, M.; Khan, M.R. Artificial Intelligence-Enabled Digital Twins for Energy Efficiency in Smart Grids. Rev. Appl. Sci. Technol. 2025, 2, 580–615. [Google Scholar] [CrossRef] [Scilit]
  172. Sadrimajd, P.; Mannion, P.; Howley, E.; Lens, P.N.L. PyADM1: A Python Implementation of Anaerobic Digestion Model No. 1. bioRxiv 2021. [Google Scholar] [CrossRef] [Scilit]
  173. Parker, W.J. Application of the ADM1 Model to Advanced Anaerobic Digestion. Bioresour. Technol. 2005, 96, 1832–1842. [Google Scholar] [CrossRef] [Scilit]
  174. Rivera-Salvador, V.; Aranda-Barradas, J.S.; Espinosa-Solares, T.; Robles-Martínez, F.; Toledo, J.U. The Anaerobic Digestion Model IWA-ADM1: A Review of its Evolution. Ing. Agríc. Biosist. 2015, 1, 109–117. [Google Scholar] [CrossRef] [Scilit]
  175. Dalmau, J.; Rodríguez-Roda, I.; Steyer, J.-P.; Comas, J. Risk Assessment Module of the IWA/COST Simulation Benchmark: Validation and Extension Proposal. In Proceedings of the 3rd International Congress on Environmental Modelling and Software (iEMSs 2006); International Environmental Modelling and Software Society: Burlington, VT, USA, 2006. [Google Scholar]
  176. Bernard, O.; Hadj-Sadok, Z.; Dochain, D.; Genovesi, A.; Steyer, J.P. Dynamical Model Development and Parameter Identification for an Anaerobic Wastewater Treatment Process. Biotechnol. Bioeng. 2001, 75, 424–438. [Google Scholar] [CrossRef] [Scilit]
  177. Tao, F.; Zhang, H.; Liu, A.; Nee, A.Y.C. Digital Twin in Industry: State-of-the-Art. IEEE Trans. Ind. Inform. 2019, 15, 2405–2415. [Google Scholar] [CrossRef] [Scilit]
  178. Grieves, M. Digital Twin: Manufacturing Excellence through Virtual Factory Replication. White Pap. 2014, 1, 1–7. [Google Scholar]
  179. Rahimieh, A.; Gilavand, M. PyADM1-R2: A Python-Based Model for Simulating Intermediates in Anaerobic Digestion Using a Simplified ADM1 Approach. In Proceedings of the 5th International Conference on Chemistry and Chemical Engineering, Tehran, Iran, 12 September 2023. [Google Scholar]
  180. Kritzinger, W.; Karner, M.; Traar, G.; Henjes, J.; Sihn, W. Digital Twin in Manufacturing: A Categorical Literature Review and Classification. IFAC-Pap. 2018, 51, 1016–1022. [Google Scholar] [CrossRef] [Scilit]
  181. Kleerebezem, R.; Van Loosdrecht, M.C.M. Waste Characterization for Implementation in ADM1. Water Sci. Technol. 2006, 54, 167–174. [Google Scholar] [CrossRef] [Scilit]
  182. Asadi, M.; McPhedran, K. Biogas Maximization Using Data-Driven Modelling with Uncertainty Analysis and Genetic Algorithm for Municipal Wastewater Anaerobic Digestion. J. Environ. Manag. 2021, 293, 112875. [Google Scholar] [CrossRef] [Scilit]
  183. Steyer, J.P.; Bouvier, J.C.; Conte, T.; Gras, P.; Sousbie, P. Evaluation of a Four Year Experience with a Fully Instrumented Anaerobic Digestion Process. Water Sci. Technol. 2002, 45, 495–502. [Google Scholar] [CrossRef] [Scilit]
  184. Ma, W.; Situ, B.; Lv, W.; Li, B.; Yin, X.; Vadgama, P.; Zheng, L.; Wang, W. Electrochemical Determination of MicroRNAs Based on Isothermal Strand-Displacement Polymerase Reaction Coupled with Multienzyme Functionalized Magnetic Micro-Carriers. Biosens. Bioelectron. 2016, 80, 344–351. [Google Scholar] [CrossRef] [Scilit]
  185. Schneider, C.; Walker, S.; Phounglamcheik, A.; Umeki, K.; Kolb, T. Effect of Calcium Dispersion and Graphitization during High-Temperature Pyrolysis of Beech Wood Char on the Gasification Rate with CO2. Fuel 2021, 283, 118826. [Google Scholar] [CrossRef] [Scilit]
  186. Yan, P.; Gai, M.; Wang, Y.; Gao, X. Review of Soft Sensors in Anaerobic Digestion Process. Processes 2021, 9, 1434. [Google Scholar] [CrossRef] [Scilit]
  187. Pisa, I.; Santín, I.; Vicario, J.L.; Morell, A.; Vilanova, R. ANN-Based Soft Sensor to Predict Effluent Violations in Wastewater Treatment Plants. Sensors 2019, 19, 1280. [Google Scholar] [CrossRef] [Scilit]
  188. Smith, C.A.; Phiefer, C.B.; MacNaughton, S.J.; Peacock, A.; Burkhalter, R.S.; Kirkegaard, R.; White, D.C. Quantitative Lipid Biomarker Detection of Unculturable Microbes and Chlorine Exposure in Water Distribution System Biofilms. Water Res. 2000, 34, 2683–2688. [Google Scholar] [CrossRef] [Scilit]
  189. Steyer, J.P.; Bouvier, J.C.; Conte, T.; Gras, P.; Harmand, J.; Delgenes, J.P. On-Line Measurements of COD, TOC, VFA, Total and Partial Alkalinity in Anaerobic Digestion Processes Using Infra-Red Spectrometry. Water Sci. Technol. 2002, 45, 133–138. [Google Scholar] [CrossRef] [Scilit]
  190. Abiodun, O.I.; Jantan, A.; Omolara, A.E.; Dada, K.V.; Mohamed, N.A.E.; Arshad, H. State-of-the-Art in Artificial Neural Network Applications: A Survey. Heliyon 2018, 4, e00938. [Google Scholar] [CrossRef] [Scilit]
  191. Jin, X.; Li, X.; Zhao, N.; Angelidaki, I.; Zhang, Y. Bio-Electrolytic Sensor for Rapid Monitoring of Volatile Fatty Acids in Anaerobic Digestion Process. Water Res. 2017, 111, 74–80. [Google Scholar] [CrossRef] [Scilit]
  192. Schievano, A.; Colombo, A.; Cossettini, A.; Goglio, A.; D’Ardes, V.; Trasatti, S.; Cristiani, P. Single-Chamber Microbial Fuel Cells as on-Line Shock-Sensors for Volatile Fatty Acids in Anaerobic Digesters. Waste Manag. 2018, 71, 785–791. [Google Scholar] [CrossRef] [Scilit]
  193. Liu, H.; Logan, B.E. Electricity Generation Using an Air-Cathode Single Chamber Microbial Fuel Cell in the Presence and Absence of a Proton Exchange Membrane. Environ. Sci. Technol. 2004, 38, 4040–4046. [Google Scholar] [CrossRef] [Scilit]
  194. Moser, A.; Appl, C.; Brüning, S.; Hass, V.C. Mechanistic Mathematical Models as a Basis for Digital Twins. Adv. Biochem. Eng. Biotechnol. 2021, 176, 133–180. [Google Scholar] [CrossRef] [Scilit]
  195. Mahmoodi-Eshkaftaki, M.; Ebrahimi, R. Integrated Deep Learning Neural Network and Desirability Analysis in Biogas Plants: A Powerful Tool to Optimize Biogas Purification. Energy 2021, 231, 121073. [Google Scholar] [CrossRef] [Scilit]
  196. Waewsak, C.; Nopharatana, A.; Chaiprasert, P. Neural-Fuzzy Control System Application for Monitoring Process Response and Control of Anaerobic Hybrid Reactor in Wastewater Treatment and Biogas Production. J. Environ. Sci. 2010, 22, 1883–1890. [Google Scholar] [CrossRef] [Scilit]
  197. Jacobi, H.F.; Moschner, C.R.; Hartung, E. Use of near Infrared Spectroscopy in Monitoring of Volatile Fatty Acids in Anaerobic Digestion. Water Sci. Technol. 2009, 60, 339–346. [Google Scholar] [CrossRef] [Scilit]
  198. Li, X.Y.; Feng, Y.; Duan, J.L.; Feng, L.J.; Wang, Q.; Ma, J.Y.; Liu, W.Z.; Yuan, X.Z. Model-Based Mid-Infrared Spectroscopy for on-Line Monitoring of Volatile Fatty Acids in the Anaerobic Digester. Environ. Res. 2022, 206, 112607. [Google Scholar] [CrossRef] [Scilit]
  199. Kipf, T.N.; Welling, M. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of the 5th International Conference on Learning Representations (ICLR 2017); OpenReview.net: Amherst, MA, USA, 2017. [Google Scholar]
  200. Cho, K.; Van Merriënboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. Learning Phrase Representations Using RNN Encoder–Decoder for Statistical Machine Translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP); Association for Computational Linguistics: Doha, Qatar, 2014; pp. 1724–1734. [Google Scholar] [CrossRef] [Scilit]
  201. Gaida, D.; Wolf, C.; Bongards, M. Feed Control of Anaerobic Digestion Processes for Renewable Energy Production: A Review. Renew. Sustain. Energy Rev. 2017, 68, 869–875. [Google Scholar] [CrossRef] [Scilit]
  202. Fernandez, E.; Ipanaque, W.; Cajo, R.; Keyser, R. De Classical and Advanced Control Methods Applied to an Anaerobic Digestion Reactor Model. In Proceedings of the 2019 IEEE Chilean Conference on Electrical, Electronics Engineering, Information and Communication Technologies (CHILECON), Valparaíso, Chile, 13–27 November 2019. [Google Scholar] [CrossRef] [Scilit]
  203. Carroll, Z.; Long, S.C. Evaluation of Surrogates for Anaerobic Digestion Processes. Proc. Water Environ. Fed. 2011, 2011, 5685–5696. [Google Scholar] [CrossRef] [Scilit]
  204. Duda, R.O.; Hart, P.E.; Stork, D.G. Pattern Classification; Wiley: New York, NY, USA, 2001; Volume 2. [Google Scholar]
  205. Lauwers, J.; Appels, L.; Thompson, I.P.; Degrève, J.; Van Impe, J.F.; Dewil, R. Mathematical Modelling of Anaerobic Digestion of Biomass and Waste: Power and Limitations. Prog. Energy Combust. Sci. 2013, 4, 383–402. [Google Scholar] [CrossRef] [Scilit]
  206. Carroll, Z.S.; Long, S.C. Bench-scale Analysis of Surrogates for Anaerobic Digestion Processes. Water Environ. Res. 2016, 88, 458–467. [Google Scholar] [CrossRef] [Scilit]
  207. Apsemidis, A.; Psarakis, S.; Moguerza, J.M. A Review of Machine Learning Kernel Methods in Statistical Process Monitoring. Comput. Ind. Eng. 2020, 142, 106376. [Google Scholar] [CrossRef] [Scilit]
  208. Pan, S.J.; Yang, Q. A Survey on Transfer Learning. IEEE Trans. Knowl. Data Eng. 2010, 22, 1345–1359. [Google Scholar] [CrossRef] [Scilit]
  209. Bernard, O.; Chachuat, B.; Hélias, A.; Le Dantec, B.; Sialve, B.; Steyer, J.P.; Lardon, L.; Neveu, P.; Lambert, S.; Gallop, J.; et al. An Integrated System to Remote Monitor and Control Anaerobic Wastewater Treatment Plants through the Internet. Water Sci. Technol. 2005, 52, 457–464. [Google Scholar] [CrossRef] [Scilit]
  210. Fawzy, S.; Saeed, M.; Eladl, A.; El-Saadawi, M. Adaptive Control System for Biogas Power Plant Using Model Predictive Control. J. Mod. Power Syst. Clean Energy 2021, 9, 1193–1204. [Google Scholar] [CrossRef] [Scilit]
  211. Asadi, M.; Guo, H.; McPhedran, K. Biogas Production Estimation Using Data-Driven Approaches for Cold Region Municipal Wastewater Anaerobic Digestion. J. Environ. Manag. 2020, 253, 109708. [Google Scholar] [CrossRef] [Scilit]
  212. Xu, W.; Long, F.; Zhao, H.; Zhang, Y.; Liang, D.; Wang, L.; Lesnik, K.L.; Cao, H.; Zhang, Y.; Liu, H. Performance Prediction of ZVI-Based Anaerobic Digestion Reactor Using Machine Learning Algorithms. Waste Manag. 2021, 121, 59–66. [Google Scholar] [CrossRef] [Scilit]
  213. Åström, K.; Hägglund, T. PID Controllers: Theory, Design, and Tuning, 2nd ed.; International Society of Automation: Research Triangle Park, NC, USA, 1995. [Google Scholar]
  214. Ljung, L. System Identification. In Signal Analysis and Prediction; Birkhäuser: Boston, MA, USA, 1998; pp. 163–173. [Google Scholar]
  215. Schmidhuber, J. Deep Learning in Neural Networks: An Overview. Neural Netw. 2015, 61, 85–117. [Google Scholar] [CrossRef] [Scilit]
  216. Ramachandran, A.; Rustum, R.; Adeloye, A.J. Review of Anaerobic Digestion Modeling and Optimization Using Nature-Inspired Techniques. Processes 2019, 7, 953. [Google Scholar] [CrossRef] [Scilit]
  217. Dutta, A.; De Keyser, R.; Nopens, I. Robust Nonlinear Extended Prediction Self-Adaptive Control (NEPSAC) of Continuous Bioreactors. In Proceedings of the 2012 20th Mediterranean Conference on Control and Automation (MED), Barcelona, Spain, 3–6 July 2012; pp. 658–664. [Google Scholar] [CrossRef] [Scilit]
  218. Yuan, X.; Li, L.; Shardt, Y.A.W.; Wang, Y.; Yang, C. Deep Learning with Spatiotemporal Attention-Based LSTM for Industrial Soft Sensor Model Development. IEEE Trans. Ind. Electron. 2021, 68, 4404–4414. [Google Scholar] [CrossRef] [Scilit]
  219. Barredo Arrieta, A.; Díaz-Rodríguez, N.; Del Ser, J.; Bennetot, A.; Tabik, S.; Barbado, A.; Garcia, S.; Gil-Lopez, S.; Molina, D.; Benjamins, R.; et al. Explainable Artificial Intelligence (XAI): Concepts, Taxonomies, Opportunities and Challenges toward Responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef] [Scilit]
  220. Gunning, D.; Stefik, M.; Choi, J.; Miller, T.; Stumpf, S.; Yang, G.Z. XAI-Explainable Artificial Intelligence. Sci. Robot. 2019, 4, eaay7120. [Google Scholar] [CrossRef] [Scilit]
  221. Guidotti, R.; Monreale, A.; Ruggieri, S.; Turini, F.; Giannotti, F.; Pedreschi, D. A Survey of Methods for Explaining Black Box Models. ACM Comput. Surv. 2019, 51, 93. [Google Scholar] [CrossRef] [Scilit]
  222. Lundberg, S.M.; Allen, P.G.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems 30 (NeurIPS 2017); Curran Associates, Inc.: Red Hook, NY, USA, 2017. [Google Scholar]
  223. Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why Should i Trust You?” Explaining the Predictions of Any Classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD’), San Francisco, CA, USA, 13–17 August 2016; pp. 1135–1144. [Google Scholar] [CrossRef] [Scilit]
  224. Shirzad Gavabari, A.; Zamir, S.M.; Nosrati, M. Interpretable ANN-SHAP Framework for Multi-Objective Optimization of CH4 and H2S in Full-Scale Anaerobic Digestion. Renew. Energy 2026, 270, 125906. [Google Scholar] [CrossRef] [Scilit]
  225. Salma, A.; Laferte, J.M.; Fryda, L.; Djelal, H. Advancing Bioenergy: In-Situ and In-Silico Approach to Enhance Anaerobic Digestion. Waste Biomass Valorization 2025, 17, 1231–1250. [Google Scholar] [CrossRef] [Scilit]
  226. Fasheun, D.O.; Omage, F.B.; Ferreira-Leitão, V.S. An Explainable Machine Learning Framework for Hypothesis Generation in Biochemical Methane Potential Prediction. Bioenergy Res. 2026, 19, 109. [Google Scholar] [CrossRef] [Scilit]
  227. Korres, N.; O’Kiely, P.; Benzie, J.; West, J. Bioenergy Production by Anaerobic Digestion: Using Agricultural Biomass and Organic Wastes; Routledge: London, UK, 2013. [Google Scholar]
  228. Manheim, D.; Martin, S.; Bailey, M.; Samin, M.; Greutzmacher, R. The Necessity of AI Audit Standards Boards. AI Soc. 2025, 40, 6609–6624. [Google Scholar] [CrossRef] [Scilit]
  229. Al Azad, S.; Madadi, M.; Rahman, A.; Sun, C.; Sun, F. Machine Learning-Driven Optimization of Pretreatment and Enzymatic Hydrolysis of Sugarcane Bagasse: Analytical Insights for Industrial Scale-Up. Fuel 2025, 390, 134682. [Google Scholar] [CrossRef] [Scilit]
  230. Qi, J.; Wang, Y.; Xu, P.; Huhe, T.; Ling, X.; Yuan, H.; Chen, Y.; Li, J. Study on Biomass and Polymer Catalytic Co-Pyrolysis Product Characteristics Using Machine Learning and Shapley Additive Explanations (SHAP). Fuel 2025, 380, 133165. [Google Scholar] [CrossRef] [Scilit]
  231. Vishwarupe, V.; Joshi, P.M.; Mathias, N.; Maheshwari, S.; Mhaisalkar, S.; Pawar, V. Explainable AI and Interpretable Machine Learning: A Case Study in Perspective. Procedia Comput. Sci. 2022, 204, 869–876. [Google Scholar] [CrossRef] [Scilit]
  232. Chen, Y.; Huang, Z.; Ma, C.; Li, Z.; Zhang, Z.; Tan, T.; Chen, Y. A Newly Early Warning Model for Anaerobic Digestion Systems: Based on an Improved Sparrow Search Algorithm Combined with Least Square Support Vector Machine. Chem. Eng. J. 2024, 490, 151743. [Google Scholar] [CrossRef] [Scilit]
  233. Yoshida, K.; Shimizu, N. Biogas Production Management Systems with Model Predictive Control of Anaerobic Digestion Processes. Bioprocess Biosyst. Eng. 2020, 43, 2189–2200. [Google Scholar] [CrossRef] [Scilit]
  234. Şendrescu, D.; Petre, E.; Popescu, D.; Roman, M. Neural Network Model Predictive Control of a Wastewater Treatment Bioprocess. In Recent Researches in Automatic Control, Systems Science and Communications; Smart Innovation, Systems and Technologies; Springer: Berlin/Heidelberg, Germany, 2011; Volume 10, pp. 191–200. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.