Next Article in Journal
LC-MS/MS-Analysis and Biological Evaluation of Hop (Humulus lupulus): Antioxidant, Antidiabetic, Anticholinergic and Antiglaucoma Activities
Next Article in Special Issue
Non-Precious Electrocatalysts for Alkaline Oxygen Evolution: Transition Metal Compounds, Carbon Supports, and Metal-Free Systems
Previous Article in Journal
Study on Risk Analysis of a Rotary Kiln-Based Activated Carbon Manufacturing Process Using Fuzzy-FMEA
Previous Article in Special Issue
Current Trends and Innovations in CO2 Hydrogenation Processes
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Application of Machine Learning Models to Oil Refinery Programming

King Abdullah Petroleum Studies and Research Center, Riyadh 13415, Saudi Arabia
Processes 2026, 14(7), 1072; https://doi.org/10.3390/pr14071072
Submission received: 30 October 2025 / Revised: 4 March 2026 / Accepted: 6 March 2026 / Published: 27 March 2026
(This article belongs to the Special Issue Feature Review Papers in Section "Chemical Processes and Systems")

Simple Summary

Machine learning is used to model all crude oil refineries in the world. The model is rigorously tested and validated using existing plant data and independent process simulation results of the new Dangote Refinery located in Nigeria. Additionally, the model is used to investigate the readiness of oil refineries in the Middle East to adapt to global energy transition by refocusing their operations to produce more petrochemicals instead of fuels.

Abstract

Transparent and evidence-based representations of global crude oil refining systems remain limited in the public literature, constraining robust energy systems modeling and policy analysis. This study develops a comprehensive, configuration-based modeling framework for all operating crude oil refineries worldwide using plant-level process unit data. Forty unique refinery configurations are identified through an unsupervised decision tree-based clustering approach that accounts for process unit presence and relative conversion intensity. An extremely randomized trees (ETR) machine learning model is trained on approximately 11,000 refinery-year observations to predict refined product yields as a function of refinery configuration, capacity, and crude oil diet. The model achieves out-of-sample coefficients of determination exceeding 0.90 for all major products and outperforms multiple linear regression and other ensemble methods. The predictive model is integrated with a differential evolution optimization algorithm to enable refinery programming under operational and feedstock constraints. The application of this model to Gulf Cooperation Council (GCC) refineries shows that, under existing technologies, petrochemical feedstock yields are bounded at approximately 37%, significantly below announced long-term diversification targets of 70–85%. Yield improvements of up to 6 percentage points are feasible through operational optimization but are associated with capacity utilization adjustments and product trade-offs. The framework provides a scalable tool for refinery benchmarking, energy transition analysis, and strategic planning across facility, national, and global levels.

Graphical Abstract

1. Introduction and Literature Review

Crude oil refineries transform raw crude feedstocks into a spectrum of finished petroleum products through a coordinated sequence of physicochemical processing steps collectively referred to as refinery unit operations. Variations in process unit configuration, installed capacity, and operational parameters give rise to substantial heterogeneity across refinery systems. To characterize this heterogeneity, the refining industry has developed quantitative classification metrics that reflect capital intensity and upgrading capability. Prominent examples include equivalent distillation capacity, the Nelson complexity index, and the bottom-of-the-barrel index, which are widely used to benchmark structural complexity and conversion depth among refining facilities [1,2]. While these indices have proven useful for quantifying structural complexity and assessing refinery suitability for processing different crude oil qualities, they provide limited resolution for detailed refinery modeling applications [3]. Specifically, such aggregate metrics do not explicitly represent process-level interactions, operational flexibility, or product yield structures required for rigorous systems analysis.
Refinery models play a critical role in evaluating economic performance, energy efficiency, and environmental compliance across refining systems. From a process engineering perspective, they enable systematic exploration of optimal configuration design and operational strategies under prevailing technical, logistical, and market constraints at local, regional, and global scales. Insights derived from these models inform investment planning, downstream diversification strategies, and regulatory policy formulation. Consequently, accurate identification and representation of underlying refinery configurations constitute a foundational requirement for developing reliable, scalable, and policy-relevant refinery modeling frameworks.
Recent advances in machine learning and numerical simulation have increasingly been applied across energy extraction and transformation systems to enhance predictive performance and support decision-making under uncertainty. In subsurface energy applications, data-driven and hybrid modeling approaches have been used to assess carbon dioxide sequestration potential, hydrogen storage in depleted gas reservoirs, and geothermal resource exploitation under varying geological and operational conditions [4,5]. These studies illustrate the capability of machine learning techniques to capture complex, nonlinear system behavior where purely mechanistic models may be computationally intensive or constrained by incomplete system characterization. The broader adoption of such approaches across the energy sector underscores their suitability for representing high-dimensional, infrastructure-intensive systems.
Despite these advances in upstream and subsurface domains, comparatively limited work has focused on a comprehensive, data-driven representation of the global crude oil refining landscape. Refinery systems exhibit similarly intricate interactions among feedstock characteristics, process configurations, and product slates, rendering them well-suited to machine learning-based modeling frameworks. Extending these methodologies to refining enables scalable representation of heterogeneous facilities while preserving the structural granularity required for optimization, policy analysis, and energy transition assessment.
Abella et al. [6] employed a predominantly North American refinery dataset to delineate ten representative refinery configurations for modeling purposes. Their configuration-based framework enabled the development of models for estimating refinery energy consumption and associated greenhouse gas emissions, thereby facilitating systematic evaluation of efficiency improvement and emissions mitigation opportunities [7]. Such modeling approaches have also been applied to assess crude blending strategies and configuration adjustments aimed at maximizing refinery netback [8]. Furthermore, configuration-based models have been extended to quantify life cycle environmental impacts of refining systems [9]. Nevertheless, the reliability of such approaches depends critically on the accurate identification and characterization of refinery process units and their structural attributes.
In 2020, the global refining system comprised 867 crude oil refineries with a combined active capacity of approximately 103.32 million barrels per day (MMb/d), of which nearly 80% were located outside North America [10]. Systematic identification of distinct refinery configurations across this global landscape would substantially enhance modeling efforts at facility, national, regional, and global scales. By grouping structurally similar refineries, configuration-based approaches reduce the dimensionality of large-scale, multiregional, and multitemporal analyses, thereby streamlining model specification. Such structuring also limits the number of variables required to characterize refinery behavior while enabling aggregation of operational datasets from refineries with comparable process architectures. This aggregation improves statistical robustness and facilitates representation of variability in design capacities and operational practices.
Given the proprietary nature of detailed refinery data, knowledge of applicable configuration classes can also guide the development of representative process simulation models for scenario analysis. Moreover, configuration-based modeling supports benchmarking and dissemination of best practices among refineries operating under similar structural conditions through systematic performance comparisons across operators, countries, and regions. These capabilities are particularly important for integrated global energy system modeling, including assessments of energy transition pathways. In this context, configuration-level representation is essential for evaluating the adaptability of refining systems, especially with respect to increasing petrochemical diversification under scenarios in which electrification reduces long-term demand for transportation fuels.
As multiple jurisdictions advance net-zero emissions commitments—many of which prioritize reductions in greenhouse gas emissions from transportation—the global refining sector is increasingly exploring strategies for deeper petrochemical diversification. In response to anticipated structural changes in fuel demand, the International Energy Agency (IEA) projects that crude oil demand for transportation fuels will peak within the coming decade, with total oil demand stabilizing at approximately 106 MMb/d by 2030 [11]. In contrast, the Organization of Petroleum Exporting Countries (OPEC) projects continued growth in global oil demand, reaching approximately 116 MMb/d by 2045 [12]. Notwithstanding these differing outlooks, both organizations anticipate a sustained expansion in crude oil utilization within petrochemical value chains.
Within this evolving landscape, future refinery development is expected to prioritize configurations and operating strategies that enhance yields of petrochemical feedstocks and related products. Accordingly, refining–petrochemical integration can be conceptualized along a spectrum of increasing structural integration. As illustrated in Figure 1, five distinct levels can be identified, ranging from conventional refining systems that primarily produce transportation fuels with limited petrochemical co-production to advanced configurations and emerging technologies designed for direct crude-to-chemicals conversion.
At present, uncertainty persists in the literature regarding the attainable limits of petrochemical feedstock yields from existing refining technologies, particularly when assessed within the multiregional and global contexts required for energy transition modeling and policy analysis. The IEA employs a representative global petrochemical yield of approximately 16% in its modeling framework [11]. However, ongoing industry initiatives aimed at deeper refining–petrochemical integration involve both capacity expansion and operational upgrades intended to increase production of liquefied petroleum gas (LPG), naphtha, and key petrochemical intermediates such as ethylene, propylene, benzene, toluene, and xylenes.
For example, implementation of high-severity or high-olefin fluid catalytic cracking (FCC) technologies can increase propylene yields from roughly 4% to as high as 20%, depending on feedstock and operating conditions. Current global capacity of high-olefin FCC units is estimated at approximately 283 kb/d, primarily located in China, South Korea, and India [13]. When such upgrading opportunities are considered alongside the baseline global petrochemical yield estimate reported by the IEA, the implied upper bound could approach approximately 32% under favorable integration scenarios. Nevertheless, translating these aggregate estimates into actionable insights requires greater structural and geographic resolution. Petrochemical yield potential varies substantially across facilities and countries due to differences in refinery configuration, feedstock characteristics, and operational practices. Accordingly, configuration-level representation is essential to assess realistic yield limits at facility, national, and regional scales.
Industry projections suggest that increasing the level of refining–petrochemical integration through higher-complexity process configurations could raise chemical yields from approximately 10% to as much as 50% per barrel of crude processed [14,15]. Integrated refinery–petrochemical complexes also present opportunities to capture operational synergies, including improved energy efficiency, enhanced input availability, and greater process flexibility. For example, as illustrated in Figure 2, hydrogen generated as a by-product in petrochemical operations can be utilized as a critical input for refinery hydrotreating and hydrocracking processes, thereby improving overall system efficiency.
Moreover, integrated facilities can dynamically adjust product slates by shifting between transportation fuels, petrochemical feedstocks, and materials—depending on prevailing market conditions to optimize economic performance. Nevertheless, advancing to higher levels of petrochemical integration requires substantial capital investment in additional process units, retrofits, and technology licensing. Consequently, achieving targeted diversification levels necessitates careful optimization of capital allocation through integrated design strategies and operational flexibility with respect to process technologies, crude feed selection, and product slate configuration.
As petrochemical diversification advances, emerging technologies are being developed and progressively de-risked to enable direct, single-step conversion of crude oil into fundamental petrochemical intermediates, including ethylene, propylene, benzene, toluene, and xylenes. In parallel, industry stakeholders and policymakers have articulated long-term diversification targets, with recent announcements indicating ambitions to achieve petrochemical yield levels in the range of 70% to 85% by 2050 [13]. Achieving such targets represents a structural transformation of conventional refining systems and necessitates rigorous analytical evaluation.
Model-based analysis provides a systematic framework for assessing the implications of the energy transition on feasible pathways toward these diversification objectives. Beyond evaluating yield potential, integrated refinery models support analyses of efficiency improvements, emissions mitigation strategies, policy design, economic optimization, and investment competitiveness. They also enable the development of more realistic energy outlook scenarios and facilitate examination of future opportunities for co-processing renewable and hydrocarbon feedstocks within existing refinery infrastructures. Realizing these capabilities requires a comprehensive representation of the global refining landscape with sufficient structural and spatiotemporal resolution to inform decision-making at facility, national, and regional levels throughout the anticipated transition horizon.
In this study, a machine learning-based framework is developed to construct a configuration-level model of the global crude oil refining system. The predictive model is subsequently integrated with a mathematical optimization module to enable refinery programming under specified operational and feedstock constraints. The coupled modeling framework is applied to evaluate the attainable limits of petrochemical diversification across existing refinery configurations in the Gulf Cooperation Council (GCC) countries.
The remainder of the paper is organized as follows. Section 2 presents the methodology for classifying global refineries into distinct design configurations. Section 3 describes the data sources, machine learning formulation, and mathematical programming framework. Section 4 introduces the GCC case studies, and Section 5 discusses the results and their implications. Section 6 concludes with key findings and outlines directions for future research.

2. Identification of Global Refinery Configurations

Refinery configurations were identified using an unsupervised-decision-tree-based recursive partitioning approach applied to process unit–level structural data. Each refinery observation was represented by binary indicators of major process units (e.g., FCC, hydrocracker, delayed coker, reformer, and hydrotreaters). At each node of the tree, candidate splits on binary indicators were selected to maximize reduction in within-node multivariate variance, defined as the sum of Bernoulli variances of unit presence indicators. Recursive splitting proceeded until minimum node size and impurity reduction thresholds were satisfied, for a maximum tree depth corresponding to the number of major process units in global refineries. The terminal nodes of the pruned tree defined refinery configuration clusters, each representing a homogeneous combination of upgrading units and relative conversion depth. Cluster robustness was verified through repeated subsampling of the dataset. Finally, configurations were ordered on an ascending complexity scale using an index constructed from unit presence and conversion severity, resulting in forty distinct refinery configurations representing the global refining landscape over the study period.

Mathematical Formulation of the Decision-Tree Clustering (Unsupervised Recursive Partitioning)

Let the dataset of N refinery observations be D = g i i = 1 N . Each refinery observation g i is represented by a feature vector
g i = u i , 1 , , u i , K
where u i , k 0,1 are binary indicators for the presence of the major process units ( k = 1 , , K ) constituting the structural attributes used for clustering.
At any node t with subset of observations D t (cardinality n t ), a node impurity I ( D t ) is defined as the sum of per-feature variance (a multivariate generalization of within-cluster scatter):
I ( D t ) = k = 1 K B e r n V a r g D t ( u k )
where
B e r n V a r g D t u k = p t , k 1 p t , k p t , k = 1 n t g D t u k
Thus, this impurity measures structural heterogeneity across the presence indicators. For a candidate split s on feature f (as u k splits 0/1), D t is partitioned into its left/right children D L and D R with sizes n L , n R . The split quality is expressed in terms of impurity reduction as:
I f , D t = I D t n L n t I D L n R n t I D R
Therefore, the split is chosen on a feature ( f * ) that maximizes impurity reduction ( I ).
f * = arg m a x ( f ) I f , D t
where there are ties on reduction at a given node for multiple features, a preference for the split with a larger child-node minimum impurity is taken.
Recursive splitting is continued along each branch until the maximum tree depth ( D m a x ) or the last child nodes with zero impurity (yielding no impurity reduction) are reached. At full growth, the number of unique leaf nodes ( C ) corresponds to the unique refining clusters/configurations. Each leaf L c cluster c (c = 1 , , C ) corresponds to a unique unit presence pattern ( u g ) and is characterized empirically using a configuration complexity index.
L c = g c   : g c   f o l l o w s   s p l i t s   t o   l e a f   c
To order the resulting clusters by increasing processing complexity, each refinery is assigned a conversion-intensity complexity score constructed from the unit presence and unit-type severity indices s k as reported in Table 1. Processing severity is defined as an ordinal measure of conversion intensity, capturing the intrinsic tendency of a unit type to convert heavier hydrocarbons into lighter products (and petrochemical feedstocks). Specifically, if K denotes the set of refinery process units, the severity indices are normalized to [0, 1] as
s ~ k = s k s m i n s m a x s m i n ,                             k K
The configuration complexity score for cluster g is defined as
C g = k K a k u g , k + α k K a k u g , k s ~ k
where a k are unit weights (set to 1 unless otherwise specified) and α 0 controls the contribution of processing severity relative to structural breadth. Clusters are ordered in ascending C g and assigned configuration numbers accordingly ( C o n f i g 0 through C o n f i g C 1 , where C is the total number of unique unit-presence clusters—the leaf nodes). In the event of identical C g values, ties are broken by the number of conversion units present.
In this work, the S&P Global Platts World Refinery Database [10] is used to investigate global crude oil refinery configurations using process unit-level data for all operating refineries during 2005–2020. Figure 3 shows simplified process diagrams of the simplest and the most complex refining configurations identified. All the identified refinery configurations are presented in Supplementary Information, Section S1. Figure 4 shows the global refining capacity in the year 2020 by the identified refining configurations. At about 16.4 MMb/d, the Config_33 refinery has the highest total capacity in the world, whereas no Config_22 refinery capacity has existed since 2016 due to either shutdown or retrofitting to a different configuration class. The last Config_22 refinery was in operation in 2015 at a capacity of around 190.8 kb/d.

3. Refinery Modeling and Mathematical Programming

A machine learning-based (non-parametric) model is developed by implementing the extremely randomized trees regressor (ETR) method in python programming language. Details of the ETR algorithm are available in Geurts et al. [16]. However, the application of the algorithm in this context is that, for a refinery learning data sample of size N , expressed as:
l s N = x i ,   y i :     i = 1 , , N
where x i is the vector of explanatory variables (otherwise called features) and y i is the vector of corresponding output variables (otherwise called targets). For the case of n features and p targets, both can be denoted as
x i = x 1 i , , x n i
y i = y 1 i , , y p i
The sample values of the j t h feature organized in order of increasing value is x j ( 1 ) , , x j ( N ) for the features ( j = 1 , , n ) . Infinite ensemble of extremely random trees estimates the value of the target variable ( y ^ p ) in the following form:
y ^ p ( x ) = i 1 = 0 N i n = 0 N I i 1 , , i n x   X x 1 , , x n λ i 1 , , i n X x j X x j
where I i 1 , , i n x is the characteristic function of the hyper-interval as given by
[ x 1 i 1 , x 1 i 1 + 1 × ×   x n i n , x n i n + 1 .                                     i 1 , , i n 0 , , N n
And λ i 1 , , i n X are real-valued parameters that depend on the examples x i and y i , as well as the parameters for minimum sample size for splitting a node ( n m i n ) and the number of randomly selected features at each node ( K ).

3.1. Data Preprocessing and Quality Control

The refinery dataset obtained from the S&P Global Platts World Refinery Database contains approximately 15,000 refinery-year observations covering process unit capacities, crude oil runs by grade, and refined product yields between 2005 and 2020. Initial data screening revealed incomplete entries in which either crude feed composition or product yield distributions were missing. Observations lacking essential input–output pairs required for supervised learning were removed, resulting in a final modeling dataset of approximately 11,000 refinery-year observations.
Data consistency checks were performed to ensure material balance plausibility and logical bounds (e.g., non-negative crude runs and product yields). Observations violating physical feasibility constraints were excluded. Outlier detection was conducted using interquartile range (IQR) analysis within each refinery configuration class. Extreme statistical outliers were cross-checked against reported capacity and operational ranges to avoid removal of legitimate but infrequent operating conditions. No artificial truncation or winsorization was applied beyond physical feasibility screening.
Because tree-based ensemble methods are invariant to monotonic transformations and do not require feature normalization, no scaling was applied for the extremely randomized trees (ETR) model. However, for benchmarking with multiple linear regression (MLR), predictors were standardized to zero mean and unit variance to ensure comparability of coefficient magnitudes and numerical stability.

3.2. Feature Engineering and Encoding

The predictive model includes 47 explanatory variables representing refinery scale, feedstock characteristics, and structural configuration.
The feature set consists of:
  • Refinery Capacity—Crude distillation unit (CDU) capacity (continuous variable, kb/d).
  • Crude Oil Feed Blend—Volume of each of six categories of crude oil qualities, comprising Light Sweet (LSW), Light Sour (LSO), Medium Sweet (MSW), Medium Sour (MSO), Heavy Sweet (HSW), and Heavy Sour (HSO), defined according to density and sulfur content as shown in Table 2. The sum of crude diet volumes is subject to total refining capacity (CDU capacity).
3.
Refinery Configuration—Forty unique refinery configurations identified through clustering. The configurations are treated as categorical variables, which are encoded using one-hot encoding, producing 40 binary indicator variables.
After encoding, all predictors are numeric. The one-hot representation allows the ETR model to learn configuration-specific nonlinear relationships between crude diet, scale, and product yields without imposing parametric assumptions.

3.3. Model Benchmarking Across Algorithms

To evaluate predictive robustness, the extremely randomized trees (ETR) model was benchmarked against multiple alternative algorithms, including:
  • Multiple Linear Regression (MLR);
  • Random Forest (RF);
  • Extreme Gradient Boosting (XGB);
  • Other ensemble and nonlinear regression models (total of 15 candidate algorithms).
All models were trained using a 70/30 train-test split with identical feature sets. Performance metrics include:
  • Coefficient of determination ( R 2 );
  • Mean Absolute Error (MAE);
  • Root Mean Square Error (RMSE).
Across all refined products (LPG, naphtha, gasoline, kerosene/jet fuel, diesel/gas oil, and heavy fuel oil), the ETR model achieved the highest average out-of-sample R 2 values ( 0.90 ) and consistently lower MAE and RMSE compared to alternative models. Random Forest (RF) and Extreme Gradient Boosting (XGB) demonstrated competitive performance but exhibited slightly higher variance across product categories. The ETR model was selected for integration with the optimization framework due to its superior predictive stability and computational efficiency.
Figure 5 shows parity plots for all the refined products predicted with the developed ETR model. The refined products modeled include liquefied petroleum gas (LPG), naphtha (NAP), gasoline (GAS), kerosene/jet fuel (KJF), diesel/gas oil (DGO), and heavy fuel oil (HFO). The coefficient of determination ( R 2 ) is above 90% for all refinery products predicted with the ETR model, meaning that over 90% of the variation in the target variables is accounted for by the model. Table 3 presents all performance metrics used to compare the models. It is observed that the ETR model demonstrates higher predictive performance for the target refined products.
For a given refining configuration, potential sources of deviations between the predicted and observed values of refined products flows could be from differences at the process sub-unit level: design, process technology, operational expertise of plant operators (human factors), and model suitability. Consequently, 15 other training algorithms were applied to the datasets, and none outperformed the ETR model on all target refined products prediction. Based on R 2 results, the MLR model shows its poorest performance on naphtha prediction, whereas the ETR model has the poorest performance on heavy fuel oil prediction. Additional statistical analysis results for the models are available in the Supplementary Information, Sections S2 and S3.
Given the developed and tested machine learning-based global refinery model, it is possible to simulate operational behaviors of global refineries under various feedstock and product demand dynamics. In fact, the model has been used to assess the market supply impact of the recently commissioned 650 kb/d capacity Dangote Refinery in Nigeria, where it predicted the gasoline supply from the refinery to around 276,000 barrels per day [17]. Another independent study by Sahdev et al. [18], using the AVEVA USC Plan software (version 3.3.1.0) to develop an end-to-end flowsheet of the plant, estimated gasoline supply from the refinery to be about 270 kb/d. The gasoline production from the plant is of preference for comparison because the plant was designed for optimal gasoline yield due to its high demand in the Nigerian market.
With established, reliable performance, mathematical optimization models can leverage the predictive power of the developed machine learning model of global refineries to specify optimal performance conditions under given objectives of interest. Mathematical programming has been used extensively in crude oil refineries to represent plant behavior and optimize performance [19,20,21,22]. A major application area in the industry is production planning where, under dynamic or uncertain market conditions influencing the supply of crude oil feedstocks of various qualities to a refinery plant and/or conditions impacting the demand of refined products by intermediate and final consumers, the goal is to steer the operation of the refinery plant to fulfill all obligations of the business [21]. Production planning can be done at the individual facility level or coordinated across facilities owned or operated by the same entity. Under such an arrangement, a machine learning model can represent the production operations of the plant. Integrating production and optimization models furnishes the synergized benefits of high-fidelity predictions with an optimum performance seeking technique.

3.4. Differential Evolution Parameterization and Convergence

The optimization problem is formulated as a constrained minimization of the negated objective function, allowing maximization of petrochemical feedstock yields while satisfying:
  • Crude diet composition bounds
  • Capacity constraints
  • Historical operational limits
  • Non-negativity conditions
The DE algorithm’s population-based search and mutation mechanisms reduce the risk of convergence to local optima, making it suitable for the nonlinear, high-dimensional predictor space induced by the machine learning model.
Workflow of the combined model implementation is presented in Figure 6, and the following paragraphs provide further description of the differential evolution algorithm, which has been used to analyze the programming problem. All formulations of the problem are also implemented using the Python programming language.
The differential evolution algorithm was presented in a paper by Storn and Price [19]. In terms of its deployment to programming machine learning-based refinery models, one considers a refinery system whose attributes of interest are represented by real-valued predictors:
x j ; j = 0 ,   1 ,   2 ,   , M
For such refinery plants, realizations of normal operational conditions are governed by the design and operating experience, which specifies bounds on the predictors:
x j [ x j l ,   x j u ]
where x j l and x j u are the lower and upper bounds on the predictors. These bounds are included in the constraints of a mathematical programming problem. Other constraints of the system may comprise material balances, logistical, and/or energy balances, which are all real-valued. The optimization task is to determine the M-dimensional vector of predictors ( x ), using the differential evolution method, which minimizes an objective function or its negation (in the case of maximization). Differential evolution method generates new parameter vectors by modifying the current generation of predictors using the weighted difference between two population members as follows,
v = x r 1 , G + β · ( x r 2 , G x r 3 , G )
The vectors x r , G are randomly but not repeatedly selected from the population ( N P ) of predictor vectors ( x i , G ) in the generation G where
x i , G ;   i = 0,1 , 2 , , N P 1
where β is a factor that amplifies the differential variation. The new/trial predictor vector ( v ) must furnish a better value of the objective criterion before it can replace a current predictor in the population of the next generation. Like the crossover process in genetic algorithms and evolutionary strategies, differential evolution also uses a mutation mechanism to introduce diversity in the population of predictor vectors, which can speed up convergence and improve global optimality of results. Further details of the differential evolution algorithm have been presented in the paper by Storn and Price [19].
To integrate the ETR model of global refineries with the DE model for production optimization purposes, the predictors of the ETR model become the predictor vectors in the DE model. Hence, for the system of models, which takes a vector of predictors comprising crude oil feed by quality and refinery capacity as input, real-valued predictions of refined products from a specified type of refinery design are evaluated as,
h p x i , G ;   p = 0 ,   1 ,   2 , , P 1
The DE algorithm generates the best predictor vector ( x b e s t , G ) to optimize the properties h p , while satisfying all constraints on the system, including the bounds on the predictors. In the case of a refinery model, it is normally desirable to maximize the value of h p . Consequently, since the DE algorithm minimizes the objective function (i.e., the negation of h p ), this results in a min-max formulation of the optimization task. Storn and Price [19] observed that this case guarantees that all local minima and possible global minimum can be found using the differential evolution algorithm, irrespective of the feasible initialization points. Once the ETR model has been trained over the entire operational landscape of all world refineries in the database, it is possible to benchmark the performance of any refinery plant against the best-in-class of the same configuration globally, over every performance measure of interest.
The differential evolution (DE) optimization was implemented using the scipy.optimize.differential_evolution function in Python (version 3.11.15) with default parameter settings. The algorithm employed the “best1bin” strategy, with a population size equal to 15 times the number of decision variables, a mutation factor uniformly sampled in the interval [0.5, 1], a recombination rate of 0.7, a maximum iteration limit of 1000 generations, and Latin hypercube initialization to enhance diversity of the initial population. Upon termination, the best solution was further refined using the default local polishing step based on the L-BFGS-B algorithm. No manual tuning of DE hyperparameters was performed.
In the present formulation, the decision variables consist exclusively of crude oil diets processed by each refinery, while refinery configuration and capacity are treated as exogenous parameters. The optimization is therefore carried out over a bounded continuous space defined by crude feed shares, subject to inequality constraints enforcing that total feed volumes observe refining capacity limits and inequality constraints representing historical operating limits. The convergence of the DE algorithm follows SciPy’s default termination criterion, whereby optimization stops when the normalized standard deviation of the population objective values satisfies
σ f μ f < tol ,
where σ f and μ f denote the standard deviation and mean of population objective values, respectively, and the tolerance parameter is set to its default value of 0.01. For each refinery case, convergence was typically achieved within approximately two minutes on a standard workstation. The observed stability and runtime are consistent with the relatively low dimensionality of the decision space and the smooth response surface induced by the machine learning-based objective function.

4. Application: Optimal Petrochemical Feedstock Yields from GCC Region Refineries

The Gulf Cooperation Council (GCC) countries are among the top producers of crude oil products globally. The GCC comprises Bahrain, Qatar, Oman, Kuwait, the UAE, and Saudi Arabia. There are about 26 refineries in the GCC, of which six are condensate splitters, with a combined capacity of around 6 million barrels per day. As presented in the previous section on identification of refinery configurations, it was demonstrated that global refinery designs range from very simple to complex configurations—categorized on the ascending complexity scale of 0 to 39, where configuration 39 is the most complex deep conversion refinery design. GCC refineries comprise configurations 0, 1, 6, 7, 10, 16, 24, 30, 33, and 36, as illustrated in Table 4 with their respective total capacities in each country.
Due to the variety of refinery design compositions within each country in the region, the yield of petrochemical feedstocks can vary across countries and even facilities. In light of this, it can be useful to understand the operational and technical limits of petrochemical feedstock supply from all refineries in the region. For this purpose, two case studies are considered to benchmark the performance of each GCC refinery using, on the one hand, the average operating conditions of all global refineries of similar design, and on the other hand, the historical limits of the operational parameters as determined from the world refinery database. Such operational conditions are incorporated into the mathematical optimization problem formulation as discussed below.
The DE algorithm poses the objective function of the optimization task as a minimization problem. In this case, the objective is to maximize the yield of petrochemical feedstocks comprising liquified petroleum gas (LPG) and finished naphtha. Consequently, the objective value ( J ) is cast and computed as a negation of the objective criterion f ( x i , G ) which is to be maximized in the DE minimization problem formulation.
J = min [ f x i , G ]
The objective criterion can also be expressed in terms of the properties to optimize ( h p ) as,
f x i , G = p h p x i , G
This function is real-valued and represents the total volume of petrochemical feedstocks produced, as predicted by the machine learning model. The objective function is subjected to two different constraints representing the two benchmarking approaches investigated in this paper. However, a generalized form of the optimization constraints can be expressed as follows:
h p x i , G = 0                     p P r ,     P P r g q x i , G 0                     q Q r ,     Q Q r
where h q and g q are equality and inequality constraints representing historical crude flow conditions for the various types of crude oil qualities, at various geographical coverage levels of national, regional, or global refining landscapes. For each refinery configuration, the design capacity limits are also specified based on real industry data. The geographical coverage determines memberships of the sets K r , M r ,   P r , and   Q r , according to actual refinery data reported at the resolution ( r ) of focus (i.e., local, national or regional levels) for analysis. Supersets P and Q cover all constraints that enable benchmarking of individual refinery performance by the global best practices of other refineries of the same or similar design. Parameters of the constraint equations are obtained from the historical information in the refinery database. In the following sections, the specific constraints pertaining to each case study are presented, and their corresponding results are briefly discussed.

4.1. Case Study 1: Benchmarking Refinery Performance with Average Historical Operating Data

In this case, optimization constraint variables are bounded by their historical average values aggregated globally for each type of refinery design. Considering the sets of global refinery plants—each denoted by a unique facility name ( m ) and plant design configuration ( k )—the constraints to specify in this case include those on production capacities, crude oil feed diets, and nonnegative feedstock volumes, which are aggregated for each refinery configuration and averaged over a historical operating period. Each constraint is expressed in terms of the predictors as follows:
j x j , G , k , m , t X k , m , t 0                 G ,   k , m , t x i , G = x 0 , x 1 , x 2 , , x j , , x N G , k , m , t
where X k , m , t is the capacity of the refinery plant processing crude oil diet consisting of oil types x j during the time period t . For each of the crude oil types processed, the feed volume is bounded by the average historical minimum and maximum volumes evaluated across a set of historical time periods ( T ) for plants of configuration k which are found in locations situated in country r , as expressed in Equations (24) through (26). The cardinality operator ( φ ) computes the memberships of M r and T .
x j , k , a v e l b x j , G , k , m , t   x j , k , a v e u b               j , G ,   k , m , t
x j , k , a v e r = t T m M r x j , k , m , t φ ( M r ) φ ( T )               j , k ,   M M r
x j , k , a v e l b = m i n x j , k , a v e r                     j , k x j , k , a v e u b = m a x x j , k , a v e r                     j , k

4.2. Case Study 2: Benchmarking Refinery Performance with the Limits of Historical Operating Data

In this case, optimization constraint variables are bounded by their historical limit values reported for the marginal refinery among all refineries having the same type of refinery design. For the same groups of global refinery plants, Equation (23) above is retained, but the rest of the constraints to specify in this case include those on the historical limits of crude oil feed diets and nonnegative feedstock volumes for each plant design configuration over the operating period of interest, applied across all global refineries of the same configuration. Therefore, the bounds on predictor variables become:
x j , k , l m t l b x j , G , k , m , t   x j , k , l m t u b               j , G ,   k , m , t
x j , k , l m t l b = m i n x j , k , m , t m M , t T                     j , k x j , k , l m t u b = m a x x j , k , m , t m M , t T                     j , k
Unlike in the first case study, the limits of the predictor variables represent the evaluated absolute minimum and maximum historical operating values for the processed crude diets across global refineries that have the same plant design configuration k .

5. Results and Discussion

The mathematical formulation of the optimal petrochemical feedstock yields of GCC refineries under the two case studies presented above is implemented in the integrated model. In the first case study, crude oil feedstock blends are specified by the global average historical minimum and maximum feed compositions of each oil type supplied to refineries with the same design configuration as in a GCC country under consideration (as expressed in Equations (24) and (25)). Consequently, in this case, the lower bounds will tend to be always greater than zero and the upper bounds are always less than unity (so far as at least one of the global refineries on record, having the same configuration, has processed the same type of crude oil over the embedded history of analysis), which may not reflect actual operational conditions as can be observed in the baseline operational performance.
Figure 7 compares the baseline and optimal feedstock blends for each GCC country. The baseline compositions represent historical crude diets of the refineries in each country. For all the countries, baseline crude feed is composed of between 1 and 3 choices of crude oil types (qualities). This is understandable given that oil-producing countries are likely to tailor their refinery selection and operation to ingest locally sourced feedstocks. However, benchmarking local performance with global average feed compositions results in a more diverse crude diet, the components of which may not all be conveniently made available locally. Figure 8 also indicates that this may be counterproductive, particularly for facilities whose baseline performance has been tuned to maximize petrochemical feedstock yields, such as Qatar in this case. Therefore, imposing the constraints in Equations (24) and (25) can alter the solution space in a manner that generates suboptimal crude feed blending guidance for enhancing petrochemical feedstock yields. In other words, Qatar’s refineries seem to have been tuned to locally optimized crude diets; thus, imposing global average crude blend constraints (as in Case Study 1) alters feed flexibility, which can reduce achievable petrochemical yield.
Alternatively, the bounds on crude feed blend compositions can be specified based on the absolute limits of their historical operating values as represented in Equations (20) and (21) for the second case study. Figure 9 shows the baseline and optimal crude feed compositions for the refineries in each GCC country. Although the baseline and optimal compositions are similar in terms of the variety of crude oil qualities in the diet, the actual compositions of processed volumes are not the same in countries like KSA, UAE, Qatar, and Oman. For Bahrain, Kuwait, and Oman, the baseline and optimal crude diets are the same or very similar, but the call on capacity (utilization rate) is lower in the optimal case, even though the petrochemical feedstock yields and volumes are higher than for the baseline operation.
Figure 10 compares the yields for this case study across the GCC countries. In all the countries, optimal petrochemical feedstock yields are higher than the corresponding baseline values, unlike those in the first case study, where the yield for Qatar was lower. It can be observed that the UAE, KSA and Kuwait have the highest potential to improve petrochemical feedstock yields from the corresponding baselines of 15%, 11%, and 13% to optimal limits of 21%, 16%, and 17%, respectively, using the current refining configurations in the countries. Of these three countries, the UAE has the maximum petrochemical feedstock yield improvement opportunity of 6% relative to their baseline, whereas Qatar has the lowest at around 1%, although having the highest yield of about 37%, which is predominantly naphtha (87%). Consequently, under existing conventional refining configurations in the region, the scope for further improvement in petrochemical feedstock yields appears limited. Achieving the approximately 80% petrochemical yield targets articulated by several GCC operators would therefore require structural transformation beyond conventional operational optimization.
Apart from the petrochemical yield limitations of current refining technology, tuning existing refineries to maximize the yields can also result in major product volume shrinkages, particularly for refineries with low baseline petrochemical yields, since such refineries are most likely to have been tuned/optimized by design and operation to maximize yields of other products instead of petrochemicals. Figure 11 compares the relationships among product volume shrinkage in the refineries to petrochemical feedstock yields and product slate compositions. Qatar, with the biggest baseline petrochemical feedstock yield, posts the smallest product volume shrinkage, 1% below baseline product yields after petrochemical feedstock optimization, whereas Oman, with the smallest baseline yield, experiences the biggest product volume shrinkage of about 13%.
Since volume shrinkage was calculated as the difference in total product volumes between the optimized and baseline operational results, it is useful to understand the plausible reasons for the observed product volume shrinkage. Further analyses of the results are performed by investigating the calls on capacity, total product volumes, and yields of other products, which can be categorized as either lighter or heavier hydrocarbons, such as gasoline and heavy fuel oil, respectively. In comparing the baseline and optimized product yields, only slight differences are observed in the volumes of other products, whether the lighter or heavier ones. However, lower capacity utilization while optimizing for higher volumes of petrochemical feedstocks is obtained in optimized results. For the refineries in Qatar, baseline capacity call is 384 kb/d for a total product volume of 373 kb/d, while the values for optimal petrochemical yields are, respectively, 379 and 370 kb/d. For Oman, the baseline values are 228 and 214 kb/d for total feed and product volumes, respectively. The optimized petrochemical feedstock yields entail a crude oil feedstock of 193 kb/d and a total product volume of 187 kb/d.
In summary, petrochemical feedstock yields depend on refinery design and operational attributes. In other words, the types of refinery configurations in the country, along with the diet of crude oil feedstocks and operating parameters, determine the refined product slates. Some combinations of refinery configuration, type of oil feedstock, capacity, and other operational variables will lead to the highest yields of petrochemical feedstocks, which can be unraveled if all global refineries were to be investigated in a similar manner. Consequently, there is a geographical variability in refinery petrochemicals yields due to differences in the suite of refinery designs and diets of oil feedstock. Therefore, refinery retrofit efforts must consider the impact of design alterations on the potential for the plant to remain viable in the evolving scenarios of a future global energy system. Understanding the maximum yields of petrochemical feedstocks from existing national, regional, or global refineries can improve decision-making regarding the best approaches to closing future gaps when transitioning from the current refining landscape. This would highlight retrofitting opportunities to sustainably transition from maximizing transportation fuel yields to capturing more value from the petrochemical value chain.

5.1. Economic Interpretation of Petrochemical Yield Improvements

To contextualize the operational significance of the observed yield improvements, the economic implications of a representative six percentage-point increase in petrochemical feedstock yield are estimated. Considering the case of the UAE, where the baseline petrochemical feedstock yield of approximately 15% increases to 21% under the optimized scenario. Given an aggregate refining capacity of approximately 1232 kb/d across configurations present in the UAE (Table 4), a 6% increase corresponds to an incremental petrochemical feedstock volume of approximately 74 kb/d (27.0 million barrels per year). Assuming a conservative net margin differential of 10–20 USD per barrel between resulting petrochemical products and displaced fuel products, the incremental gross margin at 80% yield ranges between 216 and 432 million USD per year.
This estimate is illustrative and does not account for changes in operating costs, hydrogen demand, or capital expenditures required for sustained re-optimization. Nevertheless, it demonstrates that even single-digit percentage improvements in yield can translate into substantial revenue impacts at national refining scales. The economic magnitude reinforces the strategic importance of refinery configuration and operational flexibility in energy transition planning.

5.2. Sensitivity of Petrochemical Yield to Crude Diet Constraints

To evaluate the robustness of the optimization results, a sensitivity analysis was conducted by perturbing crude feed composition bounds within ± 10 % of the historical limits used in Case Study 2. Specifically, lower and upper bounds on crude grade proportions were relaxed and tightened symmetrically while maintaining feasibility constraints.
Results indicate that optimal petrochemical feedstock yields exhibit moderate elasticity with respect to crude diet flexibility. Across GCC configurations, a ± 10 % relaxation in crude composition bounds alters optimal petrochemical yields by approximately 0.5–1.2 percentage points, depending on refinery configuration. Configurations with deep conversion units (e.g., RH and coker) demonstrate lower sensitivity, reflecting structural upgrading capacity, whereas simpler configurations exhibit higher dependence on feedstock composition.
This sensitivity analysis suggests that while feedstock availability influences attainable petrochemical yields, structural configuration remains the dominant determinant of long-run yield potential.

6. Conclusions and Future Work

Until now, no comprehensive representation of the global crude oil refining landscape has been modeled or reported in the public literature. In this paper, real refinery design and operating information from over 800 existing crude oil refineries globally was used to establish the design configuration of world refineries, starting from the facility level to national, regional, and global levels. On this basis, the refinery plant data was used to train a machine learning model to predict refined product yields given plant capacity, configuration, and crude blend feedstock. The model was tested rigorously and compared to other modeling approaches to establish its superior predictive performance. It was further integrated with an optimization model to enable its use in the mathematical programming of oil refineries. Two application case studies were presented, evaluating the implications of global energy transition for the optimal limits of petrochemical diversification in refineries within the Gulf Cooperation Council countries. It was observed that the current refining technologies investigated could not meet existing transition targets. However, it is understood that, during the transition period, some retrofits of existing plants would be promising in enhancing petrochemical diversification—if the required capital expenditure was economically incurred.
Apart from operational planning and optimization use cases, refinery models are needed in energy outlook and energy transition scenarios analysis. They provide business and government decision-makers with information on market opportunities, environmental performance, and regulatory compliance under specified operating environments and conditions. This work addressed the production and mathematical programming aspects of refinery analysis. Extensions of the modeling approach to energy requirements and emissions are additional aspects for exploration. Future applications of the models can investigate the roles of local and global refined product demand constraints on crude flows and refining margin performance. Other relevant research questions include coprocessing of renewable feedstocks and understanding the effects of crude oil sanctions on refining economics and energy affordability in compliant and noncompliant countries.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/pr14071072/s1, Section S1: Summary process diagrams of global crude oil refineries; Section S2: Feature importance and error plots of the ETR model; Section S3: Specifications and estimation results of the MLR model.

Funding

This project was funded under the KOVA project at the King Abdullah Petroleum Studies and Research Center.

Data Availability Statement

S&P Global Platts has the proprietary right to the raw data used in this study. The author cannot share raw data, but interested parties can license the raw data directly from the owner.

Acknowledgments

The author thanks all KAPSARC colleagues and support staff for their constructive discussions on the KOVA project under which this research is funded and for their help with data acquisition procedures.

Conflicts of Interest

Evar Umeozor was employed by the company KAPSARC, which had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript, or in the decision to publish the results.

Abbreviations

CDUCrude distillation unit
VDUVacuum distillation unit
HDTRHydrotreater
LPGLiquefied petroleum gas
RFMRReformer
DSLRDesulfurizer
ISMRIsomerizer
RCCResid catalytic cracker
FCCFluid catalytic cracker
DHDistillate hydrocracker
RHResid hydrocracker
NAPNaphtha
GASGasoline
KJFKerosene/jet fuel
DGODiesel/gas oil
HFOHeavy fuel oil
HSOHeavy sour crude oil
HSWHeavy sweet crude oil
MSOMedium sour crude oil
MSWMedium sweet crude oil
LSOLight sour crude oil
LSWLight sweet crude oil
MMb/dMillion barrels per day
kb/dThousand barrels per day

References

  1. Kaiser, M. A review of refinery complexity applications. Pet. Sci. 2017, 14, 167–194. Available online: https://link.springer.com/article/10.1007/s12182-016-0137-y (accessed on 17 December 2023). [CrossRef] [Scilit]
  2. Oil & Gas Journal Research Center. Worldwide Refinery Survey and Complexity Analysis. Oil Gas J. 2022. Available online: https://ogjresearch.com/products/2025-worldwide-refinery-survey-with-complexity-analysis.html (accessed on 4 February 2025).
  3. Herce, C.; Martini, C.; Salvio, M.; Toro, C. Energy performance of Italian oil refineries based on mandatory energy audits. Energies 2022, 15, 532. [Google Scholar] [CrossRef] [Scilit]
  4. Wu, J.; Ansari, U. From CO2 Sequestration to Hydrogen Storage: Further Utilization of Depleted Gas Reservoirs. Reserv. Sci. 2025, 1, 19–35. [Google Scholar] [CrossRef] [Scilit]
  5. Wang, F.; Kobina, F. The influence of geological factors and transmission fluids on the exploitation of reservoir geothermal resources: Factor discussion and mechanism analysis. Reserv. Sci. 2025, 1, 3–18. [Google Scholar] [CrossRef] [Scilit]
  6. Abella, J.P.; Motazedi, K.; Guo, J.; Bergerson, J.A. Petroleum Refinery Life Cycle Inventory Model (PRELIM) PRELIM v1.5. University of Calgary Technical Documentation. Available online: https://www.ucalgary.ca/sites/default/files/teams/477/PRELIM-v1.5-Documentation.pdf (accessed on 1 April 2024).
  7. Motazedi, K.; Abella, J.P.; Bergerson, J.A. Techno–economic evaluation of technologies to mitigate green-house gas emissions at North American refineries. Environ. Sci. Technol. 2017, 51, 1918–1928. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Nduagu, E.; Umeozor, E.; Sow, A.; Millington, D. An Economic Assessment of the International Maritime Organization Sulphur Regulations on Markets for Canadian Crude Oil; Canadian Energy Research Institute: Calgary, AB, Canada, 2018. [Google Scholar]
  9. Young, B.; Hottle, T.; Hawkins, T.; Jamieson, M.; Cooney, G.; Motazedi, K.; Bergerson, J.A. Expansion of the petroleum refinery life cycle inventory model to support characterization of a full suite of commonly tracked impact potentials. Environ. Sci. Technol. 2019, 53, 2238–2248. [Google Scholar] [CrossRef] [Scilit]
  10. S&P Global Platts. Platts World Refinery Database 2022. (Proprietary Dataset); S&P Global Platts: London, UK, 2022. [Google Scholar]
  11. International Energy Agency. WEO Special Report: The Oil and Gas Industry in Net Zero Transitions; International Energy Agency: Paris, France, 2024; Available online: https://www.iea.org/reports/the-oil-and-gas-industry-in-net-zero-transitions (accessed on 22 February 2025).
  12. OPEC. World Oil Outlook. Annual Report. 2024. Available online: https://woo.opec.org/index.php (accessed on 2 November 2024).
  13. KAPSARC Oil Market Outlook—KOMO. Saudi Arabian Refineries and Refineries of the Future. Quarterly Editorial, 2022. KAPSARC Report. Available online: https://www.kapsarc.org/our-offerings/publications/kapsarc-oil-market-outlook-komo/ (accessed on 5 October 2023).
  14. S&P Global Platts. High Olefins Fluid Catalytic Cracking Processes. PEP Report 195B, 2016. Available online: https://cdn.ihs.com/www/pdf/RP195B-toc.pdf (accessed on 3 November 2024).
  15. Saudi Aramco. Crude Oil to Chemicals: Speeding up Transformation of Crude Oil to Chemicals. Public Reports. 2022. Available online: https://www.aramco.com/en/what-we-do/energy-innovation/advancing-energy-solutions/crude-oil-to-chemicals (accessed on 5 November 2024).
  16. Geurts, P.; Ernst, D.; Wehenkel, L. Extremely randomized trees. Mach. Learn. 2006, 63, 3–42. [Google Scholar] [CrossRef] [Scilit]
  17. Umeozor, E. Can the Dangote Refinery Deliver on Its Promise? Technical Report: Instant Insight; King Abdullah Petroleum Studies and Research Center: Riyadh, Saudi Arabia, 2024; Available online: https://www.kapsarc.org/our-offerings/publications/can-the-dangote-refinery-deliver-on-its-promise/ (accessed on 20 March 2025).
  18. Sahdev, M.; Srivastava, P.; Bakrewal, A.; Uthra, S.; Panopio, V.; Shah, J. Refinery Asset Report: Dangote, Nigeria. Technical Report: Oil Trading Analytics; Rystad Energy: Oslo, Norway, 2025. [Google Scholar]
  19. Storn, R.; Price, K. Differential evolution—A simple and efficient heuristic for global optimization over continuous spaces. J. Glob. Optim. 1997, 11, 341–359. [Google Scholar] [CrossRef] [Scilit]
  20. Symonds, G.H. Linear Programming: The Solution of Refinery Problems; Esso Standard Oil Company: New York, NY, USA, 1955; Available online: https://cir.nii.ac.jp/crid/1971993809698124811 (accessed on 10 October 2023).
  21. Catchpole, A.R. The application of linear programming to integrated supply problems in the oil industry. Oper. Res. 1962, 13, 161–169. [Google Scholar] [CrossRef] [Scilit]
  22. Alattas, A.M.; Grossmann, I.E.; Palou-Rivera, I. Integration of nonlinear crude distillation unit models in refinery planning optimization. Ind. Eng. Chem. Res. 2011, 50, 6860–6870. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Integration of crude oil refining with petrochemicals production.
Figure 1. Integration of crude oil refining with petrochemicals production.
Processes 14 01072 g001
Figure 2. Potential synergies from the integration of crude oil refining and petrochemical complexes.
Figure 2. Potential synergies from the integration of crude oil refining and petrochemical complexes.
Processes 14 01072 g002
Figure 3. Simplified process diagram of the simplest and most complex refining configurations, with arrows indicating interconnections between process units.
Figure 3. Simplified process diagram of the simplest and most complex refining configurations, with arrows indicating interconnections between process units.
Processes 14 01072 g003
Figure 4. Global crude oil refinery capacities by refining configuration (2020).
Figure 4. Global crude oil refinery capacities by refining configuration (2020).
Processes 14 01072 g004
Figure 5. Parity plot of the refinery production model test performance (out-of-sample performance on 30% test set with n ≈ 3300. R2 values exceed 0.90 for all products.).
Figure 5. Parity plot of the refinery production model test performance (out-of-sample performance on 30% test set with n ≈ 3300. R2 values exceed 0.90 for all products.).
Processes 14 01072 g005
Figure 6. Workflow for coupling the refinery production and mathematical programming models.
Figure 6. Workflow for coupling the refinery production and mathematical programming models.
Processes 14 01072 g006
Figure 7. Case study 1—baseline and optimal crude oil blend compositions by quality across the GCC.
Figure 7. Case study 1—baseline and optimal crude oil blend compositions by quality across the GCC.
Processes 14 01072 g007
Figure 8. Case study 1—baseline and optimal petrochemical feedstock yield of GCC refineries.
Figure 8. Case study 1—baseline and optimal petrochemical feedstock yield of GCC refineries.
Processes 14 01072 g008
Figure 9. Case study 2—baseline and optimal crude oil blend compositions by quality in the GCC.
Figure 9. Case study 2—baseline and optimal crude oil blend compositions by quality in the GCC.
Processes 14 01072 g009
Figure 10. Case study 2—baseline and optimal petrochemical feedstock yields of GCC refineries.
Figure 10. Case study 2—baseline and optimal petrochemical feedstock yields of GCC refineries.
Processes 14 01072 g010
Figure 11. Product volume shrinkage for corresponding petrochemical feedstock yields and product slate compositions.
Figure 11. Product volume shrinkage for corresponding petrochemical feedstock yields and product slate compositions.
Processes 14 01072 g011
Table 1. Refinery processing unit severity indices based on conversion intensity.
Table 1. Refinery processing unit severity indices based on conversion intensity.
Unit TypeCategorySeverity Index ( s _ k ) Rationale (Conversion-Intensity Definition)
Condensate Splitter/CDUPrimary separation0Baseline atmospheric separation; no conversion.
VDU (Vacuum Distillation Unit)Secondary separation1Enables downstream deep conversion of residue streams.
HydrotreaterTreating1Removes heteroatoms; minimal cracking or molecular restructuring.
DesulfurizerTreating1Quality upgrading; negligible hydrocarbon conversion intensity.
IsomerizerMild catalytic conversion2Skeletal rearrangement; limited cracking severity.
ReformerCatalytic upgrading2Aromatization and dehydrogenation; moderate structural transformation without deep cracking.
VisbreakerMild thermal cracking3Limited thermal cracking of residue; modest molecular bond breaking.
RCC (Resid Catalytic Cracker)Catalytic cracking4Catalytic cracking of heavier feeds; moderate conversion intensity relative to FCC.
FCC (Fluid Catalytic Cracker)Catalytic cracking5High intrinsic cracking intensity; strong conversion to gasoline and LPG/olefins.
Distillate Hydrocracker (DH)Hydrocracking6High hydrogen-assisted cracking intensity of distillates.
Resid Hydrocracker (RH)Deep hydrocracking7Very high hydrogen-assisted conversion intensity of heavy residues.
Coker (Delayed/Flexi)Deep thermal conversion8Maximum thermal cracking intensity; extensive molecular bond-breaking with carbon rejection to solid coke.
Table 2. Categories of crude oil qualities are based on density and sulfur characteristics.
Table 2. Categories of crude oil qualities are based on density and sulfur characteristics.
Crude GradeDensity (API)Sulfur (%)
LSW>34<0.6
LSO>34>0.6
MSW>25 & <34<0.6
MSO>25 & <34>0.6
HSW<25<0.6
HSO<25>0.6
Table 3. Prediction performance metrics from testing ETR, RF, XGB, and MLR models.
Table 3. Prediction performance metrics from testing ETR, RF, XGB, and MLR models.
TargetTest Performance—ETR ModelTest Performance—MLR Model
MAERMSER2MAERMSER2
LPG0.621.430.952.073.430.77
NAP1.122.720.964.648.530.57
GAS2.536.290.968.3214.710.84
KJF0.942.110.973.215.820.83
DGO2.325.340.986.4111.430.88
HFO1.934.540.937.4612.300.65
TargetTest Performance—RF ModelTest Performance—XGB Model
MAERMSER2MAERMSER2
LPG0.791.780.920.941.820.91
NAP1.924.120.921.624.230.91
GAS3.167.500.953.697.640.95
KJF1.272.930.961.512.980.96
DGO2.946.470.973.256.330.97
HFO2.525.780.912.965.900.90
Table 4. Designs and capacities of refineries in the GCC countries (capacities are per configuration and based on 2020 data).
Table 4. Designs and capacities of refineries in the GCC countries (capacities are per configuration and based on 2020 data).
CountryConfigurations of
Refineries
Respective
Capacities (kb/d)
Country Totals
(kb/d)
Bahrain16267267
Qatar0, 7278, 137415
Oman0, 30106, 197303
Kuwait24, 33362, 406768
United Arab Emirates0, 1, 36325, 90, 8171232
Saudi Arabia0, 6, 10, 24, 33240, 805, 885, 540, 4502920
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Umeozor, E. Application of Machine Learning Models to Oil Refinery Programming. Processes 2026, 14, 1072. https://doi.org/10.3390/pr14071072

AMA Style

Umeozor E. Application of Machine Learning Models to Oil Refinery Programming. Processes. 2026; 14(7):1072. https://doi.org/10.3390/pr14071072

Chicago/Turabian Style

Umeozor, Evar. 2026. "Application of Machine Learning Models to Oil Refinery Programming" Processes 14, no. 7: 1072. https://doi.org/10.3390/pr14071072

APA Style

Umeozor, E. (2026). Application of Machine Learning Models to Oil Refinery Programming. Processes, 14(7), 1072. https://doi.org/10.3390/pr14071072

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop