Next Article in Journal
Edge-Computing for the Early Detection of Falls by People and/or Animals in Reservoirs
Previous Article in Journal
Valorization of Cranberry Pomace Through Application in Probiotic Smoothies
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Intelligent Digital Rock Physics: Advances and Perspectives from Imaging Reconstruction to Pore-Scale Multiphase Flow Simulation

1
School of Civil Engineering and Architecture, Xi’an University of Technology, Xi’an 710048, China
2
State Key Laboratory for Geomechanics and Deep Underground Engineering, China University of Mining and Technology, Xuzhou 221116, China
3
International Joint Research Laboratory of Henan Province for Underground Space Development and Disaster Prevention, School of Civil Engineering, Henan Polytechnic University, Jiaozuo 454000, China
*
Authors to whom correspondence should be addressed.
Appl. Sci. 2026, 16(12), 6118; https://doi.org/10.3390/app16126118
Submission received: 27 May 2026 / Revised: 10 June 2026 / Accepted: 12 June 2026 / Published: 17 June 2026

Abstract

In characterizing unconventional reservoirs, conventional Digital Rock Physics (DRP) has long been constrained by three fundamental bottlenecks: the trade-off between imaging resolution and field of view, challenges in reconstructing multiscale pore topology, and the prohibitive computational cost of direct numerical simulation (DNS) at the pore scale. The deep integration of artificial intelligence and rock physics has given rise to a new paradigm—Intelligent Digital Rock Physics (IDRP). This paper provides a systematic review of the evolutionary trajectory of IDRP, with a focus on how machine learning is reshaping the end-to-end workflow from imaging and segmentation to reconstruction and simulation. First, we survey image super-resolution and 3D pore structure generation techniques based on convolutional neural networks (CNNs), generative adversarial networks (GANs), and diffusion models, elucidating their mechanisms for surpassing optical diffraction limits and incorporating macroscopic petrophysical constraints. Second, we outline algorithmic strategies for fusing multi-source heterogeneous data (e.g., Micro-CT and SEM) and representing dual-porosity or multi-continuum systems. Third, we critically examine the application of machine learning surrogates in single- and multiphase flow prediction, highlighting how physics-informed machine learning (PIML) and reinforcement learning (RL)—by embedding governing equations such as Navier–Stokes or Muskat–Leverett into loss functions—achieve both computational acceleration and physical consistency. We further identify key limitations of current IDRP approaches, including insufficient validation of generated topological realism, narrow generalization across lithologies, inadequate representation of dynamic wettability, and limited model interpretability. Finally, we propose a forward-looking roadmap centered on multimodal foundation models for rocks, coupled with neural operators and uncertainty quantification frameworks, emphasizing the critical pathways for translating IDRP into engineering digital twins for unconventional hydrocarbon development, coalbed methane production enhancement, Enhanced Geothermal Systems, and geological CO2 storage. This review offers a comprehensive reference for researchers at the intersection of geophysics, rock mechanics, and artificial intelligence.

1. Introduction

Digital Rock Physics (DRP), as a pivotal bridge linking pore-scale microstructural topology to macroscopic effective properties, has become an indispensable core characterization technology in the efficient development of unconventional hydrocarbons, coalbed methane reserve enhancement and production, geological carbon dioxide storage (CCUS), and subsurface energy storage projects [1]. Conventional DRP relies on high-resolution imaging—such as micro-computed tomography (Micro-CT) or scanning electron microscopy (SEM) (Figure 1A)—followed by mathematical segmentation to delineate multiphase boundaries (Figure 1B) and subsequent reconstruction of three-dimensional pore networks (Figure 1C). This enables direct numerical simulations using computational fluid dynamics (CFD) or the finite element method (FEM), thereby transitioning rock analysis from physical experimentation to virtual digital experimentation [2].
However, reservoir geological conditions are becoming increasingly complex—exemplified by tight sandstones, highly anisotropic carbonates, and fractured coal systems. Under this evolving context, the traditional serial workflow of imaging–segmentation–simulation is approaching fundamental physical and computational limits, manifesting primarily in three interrelated challenges: (1) the inherent trade-off between optical resolution and field of view (FOV); (2) the statistical breakdown of the representative elementary volume (REV) concept in heterogeneous media; and (3) the explosive computational cost associated with direct numerical simulation (DNS) methods such as the lattice Boltzmann method (LBM). These bottlenecks severely constrain the applicability of conventional DRP in large-scale, dynamic reservoir evaluation and real-time engineering decision-making, where timeliness and scalability are critical [3].
Figure 1. Micro-CT imaging and multiphase segmentation of Estaillade carbonate rock. (A) Original Micro-CT grayscale slice; (B) multiphase semantic segmentation result; (C) three-dimensional pore network visualization [4].
Figure 1. Micro-CT imaging and multiphase segmentation of Estaillade carbonate rock. (A) Original Micro-CT grayscale slice; (B) multiphase semantic segmentation result; (C) three-dimensional pore network visualization [4].
Applsci 16 06118 g001
In recent years, the deep integration of artificial intelligence with geophysics and rock mechanics has given rise to a new paradigm—Intelligent Digital Rock Physics (IDRP). IDRP is not merely an algorithmic augmentation of conventional workflows; rather, it systematically reconfigures the entire technical chain of porous media characterization and flow simulation by: (1) leveraging convolutional neural networks (CNNs) to surpass optical diffraction limits, (2) employing denoising diffusion probabilistic models (DDPM) for high-fidelity, controllable 3D topological reconstruction, and (3) developing physics-informed machine learning (PIML) surrogates that embed conservation laws into flow predictions.
This paper centers on the core scientific question: How can data-driven approaches overcome the fundamental limitations of traditional DRP without compromising physical consistency? We provide a systematic review of the technological evolution of IDRP—from image enhancement and 3D reconstruction to multiscale fusion and pore-scale multiphase flow surrogate modeling. Distinct from existing surveys that primarily catalog algorithms, this work critically analyzes the physical fidelity, computational cost, and applicability boundaries of each methodological stage. We further identify key unresolved challenges, including insufficient validation of generated topological realism, limited cross-lithology generalization, inadequate representation of dynamic wettability, and poor model interpretability. Finally, we propose a forward-looking technical roadmap toward engineering digital twins—one that emphasizes tight integration of data and physics, real-time inversion capabilities, and uncertainty-aware decision support. This review aims to serve as a comprehensive reference for researchers at the intersection of geophysics, rock mechanics, and artificial intelligence, and to lay a theoretical and methodological foundation for intelligent evaluation of unconventional reservoirs and safe utilization of subsurface space.
This paper is a critical narrative review, given the rapid and ongoing evolution of Intelligent Digital Rock Physics (IDRP)—where new architectures (diffusion models, neural operators, vision foundation models) emerge annually—a fixed PRISMA-style filtering would quickly become outdated and may exclude methodologically innovative but non-standard contributions. Instead, we adopt a transparent, reproducible literature selection strategy as follows.
Search databases: Web of Science Core Collection, Scopus, Google Scholar, and arXiv (for recent preprints not yet published in journals). Temporal scope: Primary focus on publications from 2018 to 2026. Keyword combinations used: (“digital rock physics” or “digital rock”) and (“deep learning” or “machine learning” or “neural network”); (“super-resolution” or “generative model” or “diffusion model”) and (“micro-CT” or “FIB-SEM”); (“pore-scale flow” or “multiphase flow”) and (“surrogate model” or “physics-informed” or “PINN”); Selection rationale: From the retrieved literature, we prioritized papers that (i) achieved measurable improvements over prior methods (e.g., higher super-resolution accuracy, faster inference, better physical consistency), (ii) explicitly addressed one of the three fundamental DRP bottlenecks (resolution-FOV trade-off, computational explosion, or single-scale statistical failure), or (iii) demonstrated engineering applicability through validation against laboratory measurements or field data. Supporting citations (e.g., for well-established methods like U-Net or LBM) were included only to provide necessary background. This narrative synthesis approach is appropriate for a field in rapid flux, where the goal is to identify emerging trends and critical gaps rather than to produce a static quantitative summary.

2. Technical Bottlenecks of Conventional DRP and the AI-Driven Paradigm

2.1. The Imaging Scale Paradox: Optical Trade-Offs Between Resolution and Field of View, and Statistical Failure of the Representative Elementary Volume

In digital rock physics characterization based on X-ray micro-computed tomography (Micro-CT) or focused on beam-scanning electron microscopy (FIB-SEM), there exists an inherent mutual exclusivity between spatial resolution and macroscopic field of view (FOV), dictated by the optical diffraction limit and detector physical constraints [5]. Accurately characterizing the transport properties of porous media—such as unconventional tight sandstones, highly anisotropic carbonates, and coal rocks—requires resolving intergranular pores, clay mineral interlayer pores, and microfracture networks at sub-micron to nanometer scales. However, constrained by source brightness, the numerical aperture of optical systems, and detector pixel size, achieving nanometer-scale voxel resolution inevitably comes at the cost of reducing the physical size of the scanned sample. Consequently, the FOV in typical high-resolution scans is often limited to the millimeter or even sub-millimeter range (as shown in Figure 2). This rigid trade-off inherently establishes the critical span of scales for multiscale characterization, defining the upper and lower analytical boundaries within which material architectures can be effectively resolved and quantified. Bridging this resolution-FOV gap remains a fundamental prerequisite for ensuring the continuity of physical laws across distinct observation windows [6].
This scale constraint directly leads to the statistical failure in establishing the Representative Elementary Volume (REV). According to the fundamental principles of continuum mechanics and statistical physics, the REV is defined as the minimum statistical volume required for macroscopic effective physical properties (such as absolute permeability, elastic modulus, and electrical conductivity) to converge to stable values [8]. When the scanned volume is excessively reduced to enhance resolution, its geometric configuration fails to capture the heterogeneity prevalent in macroscopic rock cores. Key structural features, such as diagenetic dissolution vugs, tectonic microfractures, and connectivity gradients, are absent. Consequently, the pore topology and phase distribution extracted from such samples exhibit significant statistical fluctuations. This results in systematic biases in the macroscopic physical properties calculated from isolated micro-scale samples, rendering them ineffective for extrapolation to core or reservoir scales.
Conversely, adopting a large field-of-view scanning strategy (such as medical-grade CT with millimeter-scale resolution) may cover a sufficiently macroscopic volume to approach statistical representativeness, but it results in the loss of high-frequency microscopic structural information due to spatial frequency cutoff. Microporosity is systematically underestimated, and nanoscale connected pore networks are artificially discretized, leading to structural distortions in the topological connectivity of fluid migration pathways. This binary dilemma between high resolution–small FOV and low resolution–large FOV has long forced traditional Digital Rock Physics (DRP) to rely on digital rock models with incomplete geometric information for subsequent flow simulations, fundamentally constraining the physical fidelity and engineering applicability of macroscopic seepage parameter predictions [9]. Intelligent Digital Rock Physics (IDRP) aims to break through this optical trade-off bottleneck at the algorithmic level by introducing deep learning-based super-resolution reconstruction and multi-scale data fusion mechanisms, providing a new paradigm for constructing digital rocks that possess both nanoscale geometric fidelity and macroscopic statistical representativeness.

2.2. The Computational Complexity Barrier: Computational Explosion in LBM and Physical Oversimplification in PNM

After acquiring and accurately segmenting the digital rock volume, the core step in traditional DRP lies in calculating the effective transport properties of the rock through numerical solvers. Current mainstream pore-scale flow simulation methods primarily rely on the Lattice Boltzmann Method (LBM) and Pore Network Models (PNM). Although these two approaches differ significantly in their technical routes, they are, respectively, constrained by the dual barriers of computational explosion and excessive simplification of physical mechanisms.
The Lattice Boltzmann Method (LBM) is based on mesoscopic kinetic theory, tracking the spatiotemporal evolution of fluid particle distribution functions by solving the discrete Boltzmann equation, thereby reconstructing the macroscopic Navier–Stokes equations [10]. This method possesses inherent advantages in handling complex, irregular pore boundaries and multiphase interfacial dynamics, strictly preserving the microscopic topographic features of solid-fluid interfaces to achieve high-fidelity hydrodynamic simulations. However, the computational cost of LBM increases non-linearly with voxel scale, with memory requirements and solution time typically positively correlated with high-order complexity relative to the number of grid cells. Particularly when critical throats are represented by only a few grid cells, the discretization of near-wall boundary conditions can easily trigger numerical oscillations and mass conservation errors. For complex heterogeneous rock cores containing billions of voxels, Direct Numerical Simulation (DNS) using LBM often requires continuous operation on large-scale parallel supercomputing clusters for days or even weeks to converge to a steady state, severely restricting its application in large-scale stochastic inversion, multi-parameter sensitivity analysis, and real-time engineering decision-making [11].
To overcome the computational bottleneck, Pore Network Models (PNM) employ a topological dimensionality reduction strategy, abstracting the three-dimensional continuous pore space into a mathematical graph structure composed of nodes (pore bodies) and edges (throats) [12]. Based on empirical capillary equilibrium displacement rules, this method assumes Poiseuille flow within idealized conduits and achieves rapid iteration by solving simplified flow-pressure network equations. PNM exhibits extremely low computational complexity, typically requiring only milliseconds to seconds for matrix inversion, which greatly expands the feasible domain for parameter scanning and uncertainty quantification. However, this geometric abstraction comes at the cost of sacrificing microscopic physical fidelity: PNM systematically smooths out pore wall roughness, local flow separation effects, and inertial force contributions at high flow rates, making it difficult to capture interface hysteresis and dynamic wettability reversal phenomena in complex multiphase flows. More critically, the hydraulic tortuosity calculated by PNM using geometric path algorithms deviates significantly from the actual physical flow field. Especially in highly anisotropic fracture-vug carbonate rocks, predictions of permeability and relative permeability based on idealized network models often exhibit systematic distortions [13].
In summary, LBM and PNM represent a typical trade-off dilemma in traditional DRP between physical fidelity and computational efficiency. The former is constrained by computational limits, making large-scale dynamic simulations difficult to achieve, while the latter suffers from distortion of seepage mechanisms in complex reservoirs due to physical dimensionality reduction. This barrier of computational complexity has directly spurred the demand for efficient surrogate models, providing clear theoretical motivation and engineering entry points for the subsequent introduction of Physics-Informed Machine Learning (PIML) and deep neural operators.

2.3. The Intrinsic Logic of AI Intervention: Super-Resolution Breaking Hardware Sampling Limits, Generative Models Reconstructing Geometric Topology, Surrogate Models Reducing Dimensionality for Accelerated Solutions, and Embedded Physical Constraints Ensuring Conservation Laws

In response to the dual barriers of optical characterization and computational solution in traditional DRP, the intervention of artificial intelligence is not a mere stacking of algorithms, but an inevitable evolution based on the deep integration of the physical essence of porous media and data-driven patterns. The IDRP framework systematically addresses the aforementioned resolution-FOV mutual exclusivity and computational explosion challenges by reconstructing the perception–reconstruction–solution–verification technical chain. Its intrinsic logic can be summarized into four mutually coupled core dimensions.
First, super-resolution reconstruction breaks through the hardware spatial sampling limit. Traditional Micro-CT imaging is constrained by the X-ray source focal spot size, detector pixel pitch, and geometric magnification, making it difficult to balance nanoscale pore details with macroscopic statistical representativeness in a single scan. Deep Single Image Super-Resolution (SISR) networks achieve algorithmic recovery of sub-voxel high-frequency structural information by learning the non-linear mapping relationship between low-resolution grayscale fields and high-resolution microscopic topology [14]. Mathematically, this mechanism is equivalent to blind deconvolution and regularized reconstruction of the undersampled spatial frequency spectrum. Consequently, it effectively alleviates the scale trade-off contradiction without increasing hardware scanning costs or radiation doses, providing high-fidelity geometric priors for the accurate construction of the Representative Elementary Volume (REV).
Second, generative models replace traditional geometric reconstruction and cross-scale fusion. Addressing the ill-posed inverse problem where 2D slices cannot uniquely map to 3D topology, as well as the registration and fusion challenges of multi-source heterogeneous data (such as large-FOV Micro-CT and high-resolution SEM), Variational Autoencoders (VAE), Generative Adversarial Networks (GAN), and Diffusion Probabilistic Models (DDPM) provide a statistical reconstruction path based on probability distributions. By learning the joint probability density of real rock core pore-throat distributions, these models enable the controlled generation of 3D continuous media from sparse priors or low-dimensional slices. Their core value lies in breaking the physical sampling limitation of “what you see is what you get,” allowing for the on-demand synthesis of digital rock ensembles that possess both statistical consistency and geological rationality, thereby providing a reproducible and quantifiable virtual experimental sample library for multi-scale characterization.
Third, surrogate models achieve dimensionality reduction and acceleration for microscale seepage solutions. The high computational complexity of Direct Numerical Simulation (DNS) in complex topological domains stems from its rigid dependence on grid-by-grid iteration of the Navier–Stokes equations. Deep neural networks (such as 3D CNNs, Graph Neural Networks, and Neural Operators) construct end-to-end surrogate solvers by internalizing the implicit mapping between geometric morphology fields and steady-state flow fields. This paradigm transforms the traditional discretized solution of partial differential equations into non-linear regression in high-dimensional feature space. Under the premise of ensuring prediction accuracy, it compresses computational time by several orders of magnitude, making large-scale parameter scanning, real-time history matching, and uncertainty quantification transition from theoretically feasible to engineering-ready.
Fourth, embedded physical constraints ensure thermodynamic consistency and conservation laws in algorithmic outputs. Purely data-driven models are prone to generating non-physical artifacts that violate mass conservation, momentum balance, or the second law of thermodynamics when operating outside the training distribution, severely limiting their credibility in engineering decision-making. Physics-Informed Machine Learning (PIML) embeds governing partial differential equations (such as the continuity equation and Darcy/Muskat-Leverett systems) directly into the loss function as soft or hard constraints, forcing network outputs to satisfy macroscopic physical property boundaries and local conservation laws [15]. Furthermore, the synergistic introduction of PIML and Reinforcement Learning (RL) allows the dynamic evolution of wettability and multiphase interface behavior during nano-displacement processes to be inverted within a framework driven by both physical constraints and reward mechanisms, significantly reducing the risk of non-physical extrapolation by pure data models in out-of-distribution scenarios.
In summary, the technical logic of IDRP is not a simple replacement of traditional physical models, but rather the construction of a dual-wheel drive mode combining data-driven characterization + physical mechanism constraints. Super-resolution and generative models address the scale discontinuity in geometric characterization, surrogate models break through the computational bottleneck of complexity, while physics-informed constraints ensure that algorithmic evolution remains anchored to the fundamental conservation laws of rock mechanics and seepage physics. This paper will sequentially analyze the specific technological evolutions and applicable boundaries of intelligent imaging and high-precision characterization, 3D pore structure generation and multi-scale fusion, and microscale multiphase flow surrogate models, thereby revealing the complete path of IDRP from laboratory digital rocks to oilfield digital twins.

2.4. Common Limitations of IDRP Methods

Despite the transformative potential of IDRP, all methods reviewed in this paper share several inherent limitations that must be explicitly acknowledged. First, data dependence: deep learning models require large, high-quality annotated datasets that are scarce in rock physics, and their performance degrades sharply when training data do not fully sample the geological heterogeneity space. Second, out-of-distribution (OOD) generalization failure: models trained on specific lithologies, imaging conditions, or pore topologies often fail when applied to unseen rock types or scanning parameters, producing physically implausible outputs without warning. Third, lack of hard physical guarantees: even with physics-informed soft constraints, most surrogate models do not strictly enforce conservation laws at every point, and prediction errors can accumulate in long-term or multiscale simulations. Fourth, interpretability deficit: black-box neural architectures obscure the causal relationships between pore geometry and flow response, limiting their acceptance in safety-critical engineering applications.

2.5. Three Levels of Physical Consistency: From Loss Penalty to True Conservation and OOD Reliability

In the IDRP literature, claims of physical consistency often conflate three fundamentally different concepts. To enable a rigorous assessment, we explicitly distinguish them as follows:
Level 1—Compliance with physical constraints in the loss function. This is the most common approach: a residual of the governing PDE (e.g., continuity equation or Navier–Stokes) is added as a soft regularization term to the total loss. The model is trained to reduce this residual, but no guarantee exists that the residual is zero at any point, especially outside the training distribution. Level 1 improves plausibility but does not ensure conservation.
Level 2—Actual conservation of mass/momentum. This requires that the predicted flow fields satisfy conservation laws exactly (up to numerical precision). Achievable via hard-constrained architectures (e.g., divergence-free layers), projection methods, or differentiable physics solvers. In practice, Level 2 is rarely demonstrated for 3D pore-scale multiphase flow due to the high computational cost and complex boundary conditions.
Level 3—Reliability of predictions outside the training set (OOD generalization). A model may achieve low PDE residuals on the test set drawn from the same distribution, yet fail catastrophically when applied to a different lithology, a different wettability condition, or a different capillary number. OOD reliability is the ultimate engineering requirement but remains largely unsolved in current IDRP. Most reported validation uses random splits of the same dataset, which does not test OOD performance.

3. Evolution of Intelligent Imaging and High-Precision Characterization Technologies

3.1. Super-Resolution Reconstruction: The Generational Leap from Linear Interpolation to Diffusion Models

Breaking through the physical diffraction limit of Micro-CT imaging and detector sampling constraints is the primary step in constructing high-fidelity digital rocks. Traditional image super-resolution techniques have long relied on bilinear or bicubic interpolation algorithms, which are essentially based on fixed linear weighted averaging of pixels in local spatial neighborhoods. However, while these deterministic interpolation methods perform adequately in smooth regions, they are highly prone to producing blurring and aliasing artifacts in edge and texture regions, severely limiting their super-resolution capabilities. To overcome this limitation, advanced prior constraints such as non-local means filtering and steering kernel regression were introduced into the super-resolution reconstruction framework. By exploiting self-similar redundancy and local gradient covariance structures in natural images, these methods significantly improved restoration quality and texture fidelity in edge regions [16]. Although computationally inexpensive, this approach severely lacks the ability to infer non-linear high-frequency features of geological structures, inevitably leading to blurred pore edges, topological discontinuity of microfractures, and catastrophic loss of key petrophysical textures. It is crucial to recognize that all super-resolution methods—including deep learning-based ones—do not physically overcome the Abbe diffraction limit; rather, they perform a regularized statistical inference of high-frequency spatial information based on learned priors from training data. Consequently, super-resolved outputs are probabilistic estimates, not unique ground-truth reconstructions, and multiple valid realizations may exist for the same low-resolution input.
Consequently, it fails to characterize sub-micron intragranular pores and complex dissolution morphologies. The introduction of early deep learning models (such as SRCNN) initially alleviated this issue (see Figure 3 for a visual comparison of the two). With the integration of deep learning architectures, rock image super-resolution technology has undergone a paradigm shift from deterministic pixel regression to probabilistic distribution generation. This generational leap can be divided into three key stages.
The first generation of technology centers on Convolutional Neural Networks (CNNs) and their residual learning frameworks. Early models, such as SRCNN, directly fitted the non-linear mapping between low-resolution (LR) and high-resolution (HR) image patches through multi-layer convolutions, achieving a Peak Signal-to-Noise Ratio (PSNR) improvement of approximately 2.5–2.8 dB over traditional bicubic interpolation on carbonate rock Micro-CT datasets [7]. However, optimization driven by Mean Squared Error (MSE) loss functions caused networks to focus excessively on pixel-level mathematical precision rather than the perceptual clarity of geological structures. This resulted in outputs exhibiting an overly smooth wax figure effect, failing to meet the stringent requirements of downstream seepage simulations for sharp phase interfaces. To overcome the bottlenecks of vanishing gradients and limited feature representation, residual networks (such as EDSR and VDSR) and channel attention mechanisms (such as RCAN) were widely introduced. The residual learning paradigm abandons the redundant computation of reconstructing the entire image from scratch, instead focusing on mapping high-frequency residuals (such as micropore boundaries and mineral dissolution morphologies) beyond the low-frequency features shared by LR and HR images (such as basic grain contours) [17]. Coupled with structural optimizations like the removal of batch normalization layers, these models achieved precise restoration of complex rock microscopic textures while significantly reducing memory consumption. Nevertheless, the optimization orientation based on MSE loss functions still predisposed first-generation technologies to produce the aforementioned overly smooth wax figure effect. Figure 4 compares the reconstruction results of models across different generations, clearly illustrating the differences in edge sharpness and geological structure fidelity among these technological iterations.
The second generation of technology is represented by Generative Adversarial Networks (GANs), aiming to address the perceptual distortion issues inherent in CNN architectures. By constructing a minimax game framework between a generator and a discriminator, GANs shift the optimization objective from minimizing pixel errors to adversarial approximation of data distributions. Architectures such as SRGAN and ESRGAN, through joint constraints of perceptual loss and adversarial loss, successfully recovered sub-pixel high-frequency geological details that traditional CNNs failed to generate [19]. Addressing the complexity of multi-mineral rocks, the Residual Dual-Channel Attention GAN (RDCA-SRGAN) further integrates attention mechanisms to dynamically enhance feature weights for critical throat edges and microfracture networks, effectively suppressing structural artifacts caused by mode collapse. More critically, physics-constrained generative networks (such as SCPGAN) explicitly embedded macroscopic physical property priors (e.g., bulk porosity, mineral phase ratios) into the discriminator loss function for the first time. This forces the generated structures to strictly adhere to petrophysical statistical laws while satisfying visual realism, thereby providing physically consistent geometric inputs for subsequent numerical simulations [20].
The third generation of technology is rapidly evolving towards Diffusion Probabilistic Models (DDPM) and Latent Diffusion Models (LDM). In contrast to the training instability inherent in the adversarial dynamics of GANs, diffusion models are based on a Markov process of forward noising and reverse denoising. By progressively iteratively recovering high-frequency structural information, they achieve a superior balance among generative diversity, training stability, and topological fidelity. In the field of digital rock physics, diffusion models have been successfully applied to the controlled reconstruction of 3D pore structures. By embedding multiple physical property conditions such as porosity and pore size distribution, they enable the synthesis of high-fidelity digital rocks in a geostatistical sense [21]. Combined with Low-Rank Adaptation (LoRA) fine-tuning and latent space dimensionality reduction techniques, diffusion models can achieve sub-voxel super-resolution reconstruction at a controllable computational cost while preserving the multi-scale connectivity of rocks. Currently, research on rock CT super-resolution is beginning to focus on the impact of reconstruction results on downstream petrophysical analysis, such as the potential effects of pore structure characterization and texture detail recovery on physical property calculations [22]. Future research needs to further address the generalization failure under cross-lithology domain shifts and establish a quantitative mapping relationship between geometric super-resolution and hydrodynamic fidelity, thereby completely closing the loop of algorithmic enhancement—physical verification.
The choice among interpolation, CNN, GAN, and diffusion-based super-resolution depends critically on the specific task constraints. Bicubic interpolation is applicable only when downstream tasks tolerate blurred pore boundaries (e.g., porosity estimation from statistical histograms), but fails for permeability prediction or multiphase flow where topological sharpness matters. CNN-based methods (SRCNN, EDSR, RCAN) offer a balanced trade-off between accuracy and training stability, yet they systematically over-smooth high-frequency texture and are unsuitable for reconstructing microfracture networks with sub-voxel apertures. GAN-based architectures recover sharper interfaces but suffer from mode collapse and training instability, making them unreliable when geological heterogeneity is high or when the training set is small. Diffusion models currently provide the best generative diversity and topological fidelity, but their iterative reverse process incurs much higher computational cost at inference time, and their performance in out-of-distribution lithologies remains unverified.

3.2. Multiphase/Multimineral Semantic Segmentation: From the Limitations of Thresholding to Generalization Breakthroughs with Foundation Models

In the digital rock physics workflow, multiphase/multimineral semantic segmentation is a critical prerequisite that determines the fidelity of subsequent physical simulations. Traditional segmentation methods rely heavily on global or local thresholding techniques (such as Otsu’s method and region growing), which are only applicable when mineral phases exhibit distinct multimodal distributions in the grayscale histogram [23]. However, in complex tight reservoirs (such as multimineral sandstones, highly anisotropic carbonates, and coal measures), the X-ray attenuation coefficients of quartz, feldspar, clay minerals, and organic matter overlap significantly (as shown in Figure 5), leading to severe crossover in grayscale value ranges. In such scenarios, thresholding methods are not only highly dependent on manual empirical settings but are also extremely susceptible to interference from image noise, beam hardening effects, and partial volume effects, resulting in numerous isolated noise points or phase boundary adhesion. More critically, traditional algorithms rely solely on local pixel intensity and fail to capture the topological contextual information of mineral spatial distribution. This leads to a systematic underestimation of microfracture network connectivity, causing segmentation results that may appear statistically reasonable but induce non-physical blockage of seepage channels or false connectivity in hydrodynamic simulations.
The introduction of Convolutional Neural Networks (CNNs) marks a paradigm shift in rock image segmentation from grayscale heuristic rules to end-to-end feature learning. Leveraging its symmetric encoder-decoder architecture and skip connections, U-Net enables the encoder to extract deep semantic features layer by layer through convolution and pooling operations, effectively distinguishing mineral phases with overlapping grayscale values. The decoder combines upsampling with cross-layer fusion of high-resolution spatial information to achieve precise localization of sub-voxel boundaries. To address the heterogeneity of complex geological media such as shale, researchers have optimized the U-Net architecture by improving loss function weighting mechanisms and introducing attention gating and residual connections [25]. Residual modules mitigate the gradient degradation problem in deep networks, while spatial and channel attention mechanisms dynamically weight features in critical throat and microfracture regions. This significantly enhances the segmentation integrity of thin-layered structures and Discrete Fracture Networks (DFN) (as shown in Figure 6a), while also ensuring high-precision identification of matrix pore structures (as shown in Figure 6b). Quantitatively, this semantic segmentation workflow enables a clear volumetric quantification of the void space, mathematically isolating the discrete macro-fracture network (Figure 6a)—which represents a lower volume fraction but dictates the bulk high-conductivity paths—from the hyper-dense matrix pore network (Figure 6b). Incorporating these explicit pore and fracture volume partitions into the analysis provides the foundational spatial closures required for initializing the dual-porosity fluid dynamic solvers.
Multi-task learning architectures (such as DMTNN) further integrate segmentation boundary optimization with macroscopic physical property prediction (e.g., porosity, elastic modulus), achieving synergistic improvements in pixel-level classification accuracy and physical parameter inversion through dynamic weight strategies [26]. However, the performance of such deep models is highly dependent on the specific distribution of training data. When there are shifts in scanning equipment, lithology, or imaging conditions, models are prone to encountering distribution shifts. This not only leads to a precipitous drop in segmentation accuracy but also severely deteriorates the accuracy of subsequent physical property predictions [27].
Figure 6. Semantic segmentation results of fractures and pores in complex rock cores based on the U-Net fully convolutional neural network. (a) Identification results of Discrete Fracture Networks (DFN); (b) Identification results of pore structures [28].
Figure 6. Semantic segmentation results of fractures and pores in complex rock cores based on the U-Net fully convolutional neural network. (a) Identification results of Discrete Fracture Networks (DFN); (b) Identification results of pore structures [28].
Applsci 16 06118 g006
To overcome the bottlenecks of scarce annotated data and cross-domain generalization, Few-Shot Learning (FSL) and Vision Foundation Models are reshaping the technical boundaries of semantic segmentation. FSL architectures (such as prototypical networks and cross-attention feature fusion modules) extract class-invariant features through metric learning, enabling high-precision phase boundary characterization on new lithology datasets with only a minimal number of annotated samples [29]. This mechanism effectively alleviates the traditional deep learning dependence on massive annotated datasets, endowing algorithms with the meta-learning capability to rapidly adapt to new geological structures. More disruptively, general-purpose vision foundation models represented by the Segment Anything Model (SAM) have internalized universal edge perception and semantic prior logic through pre-training on massive heterogeneous data. Combined with prompt learning and parameter-efficient fine-tuning techniques, SAM variants can precisely characterize pore–matrix interfaces, identify mineral inclusions with subtle density differences, and effectively suppress under-segmentation caused by class imbalance, all without the need for full retraining [30]. Although foundation models demonstrate exceptional generalization potential in zero-shot/few-shot segmentation, their deployment in rock physics scenarios still faces challenges: domain discrepancies between pre-trained feature spaces and micro-CT grayscale distributions, a tendency towards over-segmentation in complex topological structures, and the lack of hard constraints on macroscopic physical property conservation (such as porosity and mineral volume fractions) [31]. Future research needs to construct specialized fine-tuning protocols for digital rock physics and incorporate segmentation uncertainty quantification (UQ) into the error propagation chain of downstream flow simulations, thereby achieving a thorough transition from visually realistic segmentation to physically consistent characterization.
While U-Net and its variants achieve high pixel-wise accuracy on test sets from the same lithology and imaging conditions, their validity collapses under domain shifts (different scanners, resolutions, or mineral assemblages). Few-shot learning and foundation models (e.g., SAM) improve cross-domain generalization but introduce two new limitations: (i) they tend to over-segment complex pore-fracture networks, creating false topological connections; (ii) they lack hard constraints on global physical properties (e.g., total porosity, mineral volume fractions), often producing segmentations that look visually realistic but violate mass balance. Therefore, segmentation outputs must always be validated against macroscopic petrophysical measurements before being used in flow simulations.

3.3. Physical Consistency Constraints: Mechanisms and Effects of Embedding Macroscopic Porosity and Saturation Priors into Loss Functions

Although deep learning-based super-resolution and generative models have made significant progress in recovering microscopic geometric details and visual realism, purely data-driven architectures are still dominated by pixel-level errors or adversarial perceptual losses during optimization. This makes them highly prone to generating unreasonable structures that violate fundamental laws of rock physics. For instance, generative networks may artificially fill isolated micropores or sever critical throats, causing the local porosity distribution to deviate from measured ranges, or disrupting the mass conservation and topological connectivity essential for multiphase fluid seepage. Such geometric artifacts are difficult to detect using conventional image quality metrics (PSNR/SSIM), yet they can induce order-of-magnitude deviations in permeability or distortions in capillary pressure curves in subsequent LBM or PNM simulations. Therefore, explicitly embedding macroscopic physical priors into network optimization objectives has become a core mechanism for ensuring the physical self-consistency of digital rocks and the reliability of downstream simulations [32].
Constraint-based architectures, represented by Physics-Constrained Generative Adversarial Networks (SCPGAN), transform experimentally measured macroscopic physical parameters (such as bulk porosity, mineral phase volume fractions, and fluid saturation boundaries) into differentiable regularization terms, which are directly coupled into the loss functions of the discriminator or generator [33]. The core logic lies in constructing a bidirectional feedback loop of geometric generation—physical verification. During adversarial training, the discriminator not only evaluates the local texture realism and global structural rationality of the image but also calculates the deviation between the statistical moments of the generated samples (such as pore voxel proportion and two-phase correlation functions) and the target prior values, adding this deviation as a penalty term to the total loss. Mathematically, this mechanism is equivalent to imposing physical manifold projection constraints on the high-dimensional generative manifold, forcing the network output to strictly converge to the preset macroscopic physical property hyperplanes while satisfying adversarial equilibrium (for example, matching the distribution of porosity and permeability, as shown in Figure 7). For multiphase seepage scenarios, some studies further introduce saturation soft clamping and phase boundary curvature regularization to ensure that the generated distributions of wetting and non-wetting phases comply with thermodynamic equilibrium conditions, thereby avoiding the artificial construction of non-physical interface morphologies that violate the Young–Laplace equation.
The empirical effects of embedding physical constraints have significantly enhanced the engineering applicability of digital rocks. Studies indicate that three-dimensional porous media reconstruction methods based on Generative Adversarial Networks (GANs) can maintain key topological features of the pore-solid interface (with errors in porosity and specific surface area <5% and good consistency in Euler characteristics), while ensuring that the non-central covariance function of synthetic samples aligns well with the training images in terms of radial averages and along three Cartesian directions. Numerical simulation results for single-phase permeability show that the permeability distribution of synthetic samples basically matches that of the original micro-CT images in terms of magnitude and statistical variability (with differences in mean and variance within the range of statistical error) [35]. More importantly, this paradigm enables condition-controlled generation: researchers can drive the model to batch-generate statistically equivalent and physically consistent sets of 3D digital rocks by simply inputting conventional laboratory physical property parameters (such as helium porosity and mercury intrusion capillary pressure saturation curves). This effectively alleviates the constraints imposed by scarce core samples on parameter sensitivity analysis and uncertainty quantification, and provides high-quality, highly diverse labeled datasets for surrogate model training.
However, the mechanism of physical prior constraints still faces multiple challenges in terms of theoretical completeness and engineering deployment. First, the weight allocation of multiple constraint terms relies heavily on empirical hyperparameter tuning. Excessively strong physical penalties can suppress the network’s ability to express natural geological heterogeneity, leading to homogenized generated structures; conversely, excessively weak penalties fail to effectively curb non-physical artifacts outside the training domain. Second, current constraints are mostly limited to static scalar parameters (such as global porosity) and struggle to explicitly encode dynamic processes (such as pore compression under effective stress loading or time-varying wettability evolution) or complex tensor fields (such as anisotropic permeability tensors and elastic stiffness matrices) [36]. Future research needs to develop adaptive constraint weight optimization algorithms and frameworks coupled with differentiable physics solvers, achieving a transition from macroscopic statistical matching to microscopic dynamic conservation. This will lay the foundation for IDRP to build digital twins that truly possess multi-field coupling inference capabilities and strict error boundary control.
In summary, the physics-constrained generative models discussed in this section operate primarily at Level 1 (loss-function compliance). They successfully force global statistical moments (e.g., porosity, specific surface area) to match target priors, but they do not enforce pointwise conservation of mass or momentum, nor do they guarantee physically correct local interface curvatures. Actual conservation (Level 2) is not claimed, and OOD reliability (Level 3) has not been systematically evaluated for any super-resolution or segmentation method.

4. Three-Dimensional Pore Structure Generation and Multi-Scale Fusion

4.1. 3D 2D-to-3D Pore Structure Generative Modeling: The Paradigm Evolution from Probabilistic Ambiguity to Conditionally Controllable Diffusion

Reconstructing a three-dimensional continuous porous medium from limited two-dimensional slices or sparse statistical priors is essentially a typical ill-posed inverse problem. Mathematically, a single two-dimensional cross-section can map to infinite three-dimensional topological configurations. Traditional stochastic simulation methods based on Multiple-Point Statistics (MPS) rely heavily on the selection of training images and struggle to efficiently characterize the multi-scale connectivity of complex heterogeneous reservoirs [37]. The introduction of generative artificial intelligence provides a probabilistic solution path for this inverse problem based on data manifold learning. Its technological evolution has undergone statistical approximation via Variational Autoencoders (VAE), adversarial optimization via Generative Adversarial Networks (GAN), and is now rapidly transitioning towards Diffusion Probabilistic Models (DDPM) and conditionally controllable generation paradigms.
Early 3D reconstruction architectures predominantly employed Variational Autoencoders (VAEs). By encoding high-dimensional microscopic structures into a low-dimensional latent space and approximating the prior of latent variables with a mixture of Gaussian distributions, these models effectively quantify the structural ambiguity inherent in the 2D-to-3D mapping by outputting a probability distribution over plausible 3D structures. This probabilistic representation acknowledges that multiple distinct 3D pore topologies can be consistent with the same 2D slice—a fundamental ill-posedness that deterministic methods ignore. However, VAEs are optimized using Maximum Likelihood Estimation (MLE) and Mean Squared Error (MSE) loss, which essentially constitutes an expectation regression of the conditional probability distribution. Although this mechanism can generate structurally diverse samples, it inevitably leads to systematic over-smoothing of the outputs, such as blurred boundaries of critical throats and topological discontinuities in microfractures. This severely undermines the accuracy of digital rocks in characterizing hydraulic tortuosity and seepage connectivity, failing to meet the geometric fidelity requirements for downstream multiphase flow simulations.
To overcome the bottleneck of edge blurring, Generative Adversarial Networks (GANs) and their variants (such as WGAN-GP and CVAE-GAN) have been widely introduced into 3D pore reconstruction [38]. By constructing a minimax game framework between the generator and the discriminator, GANs shift the optimization objective from minimizing pixel-level errors to adversarially approximating data distributions, successfully recovering sharp solid–fluid interfaces and complex mineral dissolution morphologies. However, the inherent dynamic instability of adversarial training is highly prone to triggering Mode Collapse: to rapidly converge to a Nash equilibrium, the generator often tends to output only a few structural patterns that frequently succeed in deceiving the discriminator, leading to a severe lack of topological diversity and geological heterogeneity in the generated samples. In the modeling of complex carbonates or tight sandstones, this sample homogenization directly results in an underestimation of pore network connectivity probability, causing systematic biases in subsequent seepage parameter predictions.
In recent years, Denoising Diffusion Probabilistic Models (DDPM) have gradually replaced GANs and VAEs as the mainstream architecture for 3D digital rock synthesis, owing to their superior generative balance. Diffusion models are based on a rigorous Markov process of forward noising and reverse denoising: the forward process progressively injects Gaussian noise until a pure static distribution is reached, while the reverse process iteratively recovers coherent geometric structures by training a neural network to learn the gradient field (as shown in Figure 8). Mathematically, this mechanism is equivalent to estimating the data score function. It not only retains the high-frequency detail recovery capability comparable to GANs but also fundamentally avoids the oscillation risks and mode collapse defects inherent in adversarial training, ensuring that the generated ensemble possesses high statistical diversity and topological rationality. To overcome the bottlenecks of pixel-space diffusion models, such as the large number of iteration steps and high computational cost, Latent Diffusion Models (LDM) have been introduced into porous media generation tasks. This framework compresses 3D geological volumes into a low-dimensional latent space via a Variational Autoencoder, achieving a stepwise reduction in computational complexity while preserving key spatial topological correlations. Combined with parameter-efficient fine-tuning techniques such as Low-Rank Adaptation (LoRA), LDM significantly improves inference efficiency while maintaining generation quality, providing engineering feasibility for the construction of large-scale digital rock synthesis libraries [39].
The core breakthrough in current IDRP generative modeling lies in the mechanism of controllable generation constrained by macroscopic physical properties. Traditional unconditional generation can only ensure visual or local texture realism but fails to match macroscopic physical parameters measured in the laboratory. The conditional diffusion framework explicitly encodes target statistics (such as bulk porosity, specific surface area, or permeability priors) as conditional guidance vectors for the reverse denoising process, forcing the generative manifold to converge to preset physical hyperplanes. With macroscopic porosity as the sole constraint, Latent Diffusion Models can self-consistently preserve secondary physical indicators such as two-point correlation functions, hydraulic tortuosity, and permeability tensors at the implicit representation level. The generated 3D pore volumes align closely with real rock cores in terms of multidimensional statistical features, fully validating the physical self-consistency of the conditional generation framework [41]. This paradigm makes on-demand synthesis possible: researchers can batch-construct sets of digital rocks that are statistically equivalent, physically consistent, and topologically rational by simply inputting conventional core analysis or well logging inversion parameters. This completely breaks through the constraints imposed by scarce physical samples on parameter sensitivity analysis and uncertainty quantification.
Although diffusion models demonstrate significant advantages in geometric reconstruction and conditional controllability, their engineering deployment still faces challenges. The iterative nature of the forward-reverse processes results in generation times that remain higher than those of some traditional stochastic methods. The generation of four-dimensional (3D + time) pore structures under complex stress paths or dynamic wettability evolution is still in the exploratory stage. Furthermore, the physical consistency of the generated results highly depends on the measurement accuracy of the conditional priors and the configuration of loss weights. Future research needs to further integrate differentiable manifold optimization with multi-physics coupling constraints, promoting a transition in 3D generative modeling from static statistical matching to dynamic genetic-driven approaches.

4.2. Cross-Modal Data Fusion: Registration and Alignment of Micro-CT with SEM/FIB-SEM, and Voxel-Level Micropore Statistical Mapping

The pore systems of natural unconventional reservoirs (such as tight sandstones, highly anisotropic carbonates, and coal measures) rarely exhibit uniform unimodal distributions; instead, they are strongly heterogeneous multi-scale bodies spanning several orders of magnitude. Single imaging modalities face insurmountable physical blind spots in characterizing such media: while Micro-CT can acquire millimeter-scale fields of view to approach statistical representativeness, its voxel resolution is typically limited to the micrometer scale, systematically omitting intragranular pores in clay minerals, organic matter nanopores, and microfracture networks [42]. Conversely, although Scanning Electron Microscopy (SEM) or Focused Ion Beam-Scanning Electron Microscopy (FIB-SEM) can resolve nanoscale pore-throat topologies, they are constrained by extremely small fields of view and the inherent limitations of 2D slices or micrometer-scale 3D volumes, making it difficult to capture macroscopic heterogeneity and connectivity gradients. Traditional digital rock analysis has long relied on single-scale models. Limited by resolution and sample size, these models struggle to simultaneously capture nanoscale micropores and large-scale intergranular pores, leading to a systematic underestimation of total porosity (with deviations reaching several percentage points in some heterogeneous samples). Unresolved nanoscale pore networks may affect the accurate characterization of pore connectivity, thereby adversely impacting the accuracy of numerical simulations for transport properties such as permeability [43]. Cross-modal data fusion technologies aim to bridge macroscopic geometric structures with microscopic statistical priors through algorithms, constructing dual-porosity or multi-continuum digital rocks that possess both Representative Elementary Volume (REV) representativeness and nanoscale topological fidelity.
The primary technical bottleneck in cross-modal fusion lies in the spatial registration and scale alignment of multi-source heterogeneous data. Significant differences in imaging physical mechanisms, grayscale response functions, and spatial sampling rates between Micro-CT and SEM/FIB-SEM mean that direct superposition would introduce severe geometric distortions and feature misalignments. Current mainstream workflows adopt a hierarchical strategy of global coarse registration—local fine alignment. First, cross-modal invariant features such as mineral grain contours, fracture traces, and sedimentary bedding are extracted from both image types using Mutual Information maximization or Scale-Invariant Feature Transform (SIFT) algorithms to establish an initial spatial transformation matrix. Subsequently, non-rigid deformation field optimization (such as B-spline free-form deformation models) is introduced to compensate for local geometric distortions caused by mechanical damage during sample preparation, ion beam cutting taper effects, and thermal expansion [44]. To overcome the domain shift problem during registration, Cycle-Consistent Generative Adversarial Networks (Cycle-GAN) and style transfer architectures are widely used for domain adaptation of unpaired multi-resolution images. By implicitly learning the grayscale-texture mapping relationships between modalities, these methods achieve seamless projection of SEM microscopic textures into the CT macroscopic coordinate system [45].
Following spatial registration, voxel-level micropore statistical mapping becomes the core of the fusion algorithm. This process abandons simple image superposition in favor of establishing a data-driven multivariate response surface linking grayscale, mineralogy, and porosity. Deep learning models traverse high-resolution SEM/FIB-SEM annotated datasets to extract the non-linear correlations between sub-resolution porosity, CT grayscale intensity, and mineral proportions, thereby constructing a voxel-level transfer function [46]. This mapping mechanism enables the algorithm to quantitatively assign probability distributions of unresolved nanopore space to low-resolution CT voxels based on their local grayscale histograms and mineral classification results. Empirical studies demonstrate that multi-scale models corrected by machine learning can strictly converge helium porosity predictions to laboratory benchmark ranges. The prediction results align highly with core measurement data, while significantly compressing absolute errors to low levels, effectively validating the physical reliability of cross-modal statistical mapping.
The fused multi-scale volumetric data must be further mapped into dual-porosity or multi-continuum computational frameworks to support fluid–solid coupling and multiphase seepage simulations [47]. Addressing the typical topology of high-permeability discrete fracture networks—low-permeability matrix reservoirs in tight reservoirs, researchers have constructed Dual-Porosity FlowNet models. In these models, large-scale dissolution vugs and tectonic fractures are characterized as high-conductivity channels, while the tight matrix containing micropores serves as the domain for fluid storage and diffusion. Numerical solvers based on fused geometry enable strong coupling calculations between Stokes flow (in large pore/fracture zones) and Darcy/Brinkman flow (in micropore zones), accurately capturing the local anisotropy and dynamic transport behavior of karstified reservoirs. In terms of mechanical characterization, embedding multi-scale digital rocks into a parallel computational framework combining Differential Effective Medium (DEM) theory and the Finite Element Method (FEM) enables cross-scale elastic parameter inversion. For instance, the calculated multi-scale bulk modulus of 21.67 GPa deviates by less than 4% from the experimental baseline of 22.5 GPa, achieving high-precision alignment between digital twins and physical acoustic/mechanical measurements [36].
Although cross-modal fusion significantly enhances the physical self-consistency of multi-scale digital rocks, its engineering deployment faces three core challenges. First, grayscale domain shifts caused by differences in imaging conditions across devices (such as accelerating voltage, beam intensity, and sample tilt angle) can easily render statistical mapping functions ineffective on new samples, making generalization capability highly dependent on Domain Adaptation and meta-learning strategies. Second, multi-modal registration errors accumulate non-linearly during the voxel-level mapping process, potentially creating artificial connectivity paths or blocking real seepage networks; thus, Uncertainty Quantification (UQ) frameworks must be introduced to probabilistically assess the topological reliability of fusion results. Third, current fusion mechanisms largely remain at the level of static geometric statistical matching and do not explicitly encode diagenetic evolution history, pore compression responses under effective stress loading, or dynamic wettability migration. Future research needs to develop physics-guided differentiable multi-modal fusion architectures, embedding thermodynamic equilibrium constraints and genetic geological models into optimization objectives, thereby promoting a transition in cross-scale characterization from data patching to physically consistent reconstruction.

4.3. Dual-Porosity/Multi-Continuum Characterization: Complex Lithology Topological Modeling and Macro–Meso Parameter Inversion Validation

The pore systems of natural unconventional reservoirs rarely exhibit idealized unimodal distributions; instead, they are strongly heterogeneous bodies composed of multi-level pore networks and discrete fracture systems coupled across scales spanning several orders of magnitude [48]. In tight sandstones, interlayer pores in clay minerals coexist with intergranular dissolution pores in quartz/feldspar, exhibiting significant stress sensitivity due to diagenetic compaction. In carbonates, diagenetic dissolution and tectonic modification jointly shape a complex topology characterized by the triple coupling of vugs–fractures–microcrystalline matrix. Coal rock systems possess organic matter nanopores, tectonic shear fractures, and bedding-plane preferential flow channels. Traditional single-scale digital rock models inevitably face connectivity discontinuities and distorted transport mechanisms in such media, necessitating the construction of a Dual-Porosity/Multi-Continuum characterization framework.
Intelligent Digital Rock Physics provides a computable mathematical representation path for dual-porosity modeling of complex lithologies through multi-modal data fusion and generative topological reconstruction. At the geometric construction level, algorithms first separate the macroscopic pore-fracture network, which dominates seepage, from the microporous matrix, which serves as fluid storage, based on statistical mapping between high-resolution microscopic images and macroscopic CT voxels. Addressing the karstification characteristics of carbonates and the anisotropic fracture development in coal rocks, studies have introduced conditional diffusion models and Graph Attention Networks to explicitly couple the geometric occurrence and aperture distribution of Discrete Fracture Networks (DFN) with the topological connectivity of matrix pores. This generates 3D dual-phase volume models that possess both macroscopic heterogeneity and microscopic porosity gradients [39]. At the computational characterization level, this geometric structure is further abstracted into the solution domain for multi-continuum governing equations: the high-conductivity fracture system obeys the Navier–Stokes or Brinkman equations, while the low-permeability matrix domain adopts Darcy flow or modified models accounting for slip effects, with the two strongly coupled through interface mass and momentum exchange terms [49].
The physical credibility of dual-porosity/multi-continuum digital rocks must be ultimately confirmed through a closed-loop validation involving the inversion of macroscopic mechanical and acoustic parameters. Researchers embed the fused multi-scale volume models into a computational framework combining Differential Effective Medium (DEM) theory and parallel Finite Element Method (FEM). Under periodic boundary conditions and equivalent homogenization assumptions, this framework directly solves for the macroscopic elastic stiffness tensor and wave propagation characteristics of the digital rock. In terms of elastic wave velocity prediction, empirical studies demonstrate that an improved squirt flow model combined with a hybrid random medium model exhibits high accuracy in predicting velocities in carbonates. For limestone and dolomite core samples, this coupled model characterizes complex pore structures by inverting microfracture structures and accounts for local flow attenuation effects between microfractures of different aspect ratios. Simultaneously, it uses a random medium model to characterize matrix heterogeneity by adjusting autocorrelation lengths, rounding coefficients, and angular parameters. The relative error between predicted P-wave velocities and laboratory ultrasonic measurements remains consistently below 5% (less than 2% for some homogeneous samples), significantly outperforming traditional Biot theory (which has an error of approximately 10%). This fully validates the physical self-consistency and engineering applicability of the model in characterizing complex microfracture networks and matrix heterogeneity [50]. In coal rock and tight sandstone systems, by incorporating stress–seepage–damage coupled constitutive models, digital rocks successfully reproduce the stepwise decline in permeability caused by fracture closure under effective stress loading, as well as the evolution of elastic wave anisotropy, thereby verifying the physical fidelity of the multi-continuum framework in describing fluid–solid coupling processes.
Although dual-porosity/multi-continuum characterization significantly enhances the multi-dimensional representation capabilities of digital twins for complex reservoirs, its theoretical completeness and engineering applicability still face three challenges. First, current multi-scale topological fusion largely relies on static geometric statistical matching and does not explicitly encode diagenetic evolution history or pore compression responses under dynamic stress paths, leading to systematic biases in permeability predictions under varying confining pressures [51]. Second, the mass and momentum exchange mechanisms at multi-continuum interfaces heavily depend on empirical closure relationships, lacking universal derivations based on microscopic physical mechanisms. Third, simulations of digital rock dynamic evolution under mechanical–seepage–chemical multi-field coupling remain constrained by computational limits and uncertainties in cross-scale parameter transfer. Future research needs to develop differentiable multi-physics coupled solvers and data-mechanism hybrid-driven constitutive inversion frameworks, deeply integrating macroscopic field well logging data with microscopic digital rock dynamic updating mechanisms, thereby promoting a transition in dual-porosity characterization from static geometric assembly to dynamic physical evolution.

5. Microscale Multiphase Flow Surrogate Models and Physics-Informed Machine Learning (PIML)

5.1. Single-Phase Flow Prediction: Data-Driven Mapping from Geometric Morphology to Seepage Fields and Acceleration Mechanisms

Direct Numerical Simulation (DNS) of pore-scale single-phase seepage has long been constrained by the computational rigidity associated with the discretized solution of high-dimensional partial differential equations. To overcome the computational bottleneck of the Lattice Boltzmann Method (LBM) in domains containing billions of voxels, deep learning surrogate models have achieved a paradigm shift from physical iterative solution to high-dimensional feature regression by internalizing the non-linear mapping relationship between the geometric topology of porous media and steady-state flow fields. The core of this approach lies in constructing an end-to-end prediction architecture that maps morphological fields to velocity/pressure fields, thereby establishing strict engineering confidence between computational acceleration ratios and error boundaries.
Surrogate models based on Three-Dimensional Convolutional Neural Networks (3D CNNs) were the first to achieve breakthroughs in this field. Architectures represented by PoreFlow-Net take the binary pore structure or grayscale attenuation field of digital rocks as input, extract implicit features such as local pore curvature, throat contraction ratios, and spatial connectivity through multi-layer convolutional kernels, and directly output the 3D steady-state velocity field and absolute permeability tensor [52]. This mechanism bypasses the Markov process of time-step-by-time-step iteration in LBM, compressing the solution complexity from superlinear growth with grid scale to the linear order of forward propagation. On test sets of conventional sandstones and homogeneous carbonates, this multi-scale surrogate model compresses the residuals of absolute permeability predictions across different resolutions. As shown in Figure 9, the predicted absolute permeability systematically spans from 0.5 mD to 6.5 mD. Across all three spatial resolutions, the residual scatter points exhibit a highly symmetrical distribution centered around the dashed zero-error baseline (Residuals = 0.0). Quantitatively, for the high-resolution domain, the absolute residuals for the vast majority of test samples are strictly confined within a narrow band. As the spatial sampling frequency drops to middle- and low-resolution limits, the model maintains remarkable topological robustness; the maximum error boundaries expand only marginally. Furthermore, a distinct heteroscedastic pattern characteristic of rock porous media flow can be quantified: within the higher permeability range, the residual variance is remarkably low, whereas the localized variance slightly increases within the lower permeability bracket, where micro-throats dominate the transport physics.
Furthermore, the Pearson correlation coefficient between the spatial distribution of the velocity field and the benchmark solution remains highly consistent, demonstrating excellent fidelity in flow field reconstruction. Of greater engineering value, the predicted velocity field can serve as a physical initialization field for high-precision LBM solvers: by providing an initial momentum distribution close to the true steady state, it significantly reduces the number of computational steps required for subsequent iterative convergence, successfully constructing a hybrid solution loop of rapid estimation—precise correction.
To ensure the engineering reproducibility of the convolutional networks compiled in this section (such as the multi-scale baseline visualized in Figure 9), it is essential to deconstruct their underlying mathematical formulations, algorithmic topologies, and hyperparameter matrices. Mainstream 3D CNN architectures in IDRP generally fall into two categories based on their dimensional mapping and fluid physics.
For single-phase steady-state flow field predictions, the core physical assumption scales down to creeping, incompressible Stokes flow, where the local velocity vectors are uniquely governed by the rigid 3D pore boundaries. These models typically adopt a fully convolutional encoder-decoder topology (symmetrical 3D U-Net variants) with cross-layer skip connections to pass high-frequency spatial features. The architectural depth commonly ranges from 14 to 18 operational layers, utilizing 3 × 3 × 3 or 5 × 5 × 5 convolutional kernels with a stride of 1 or 2 to capture spatial dependencies. Hidden layers are activated via LeakyReLU to circumvent gradient vanishing, while the final output layer employs Linear or Tanh functions to output unrestricted velocity-pressure fields. The training optimization utilizes the Adam optimizer governed by cosine decay. Crucially, the objective function is formulated as a hybrid loss that couples pixel-level Mean Squared Error (MSE) with physics-informed soft regularization tracking mass conservation residuals.
Conversely, for multiphase transport behaviors, the dynamics are dictated by capillary-dominated displacement rules, where the relative permeability (kr) and capillary pressure (Pc) curves are non-linear reflections of pore scale, dynamic contact angles (θ), and saturation history. To handle varying core sizes, these frameworks integrate 3D convolution layers with Spatial Pyramid Pooling (SPP) modules across 12 to 15 layers. The inputs are multi-channel tensors that seamlessly fuse binarized structural geometry with localized wettability distributions and fluid phase fractions. The network constraints employ 3 × 3 × 3 filters alongside variable pooling windows (ranging from 1 × 1 × 1 to 4 × 4 × 4). Hidden units leverage regular ReLU functions, whereas the bounding outputs use Sigmoid or Softplus operations to mathematically restrict relative permeability within the physical domain. The training employs the Adam optimizer with automated step-wise plateau reduction, driven by Mean Absolute Error (MAE) or MSE losses tailored to minimize empirical curve deviations.
To address the limitation of CNNs in regular grid representations, which often overlook the long-range connectivity of pore networks, Graph Neural Networks (GNNs) have been introduced to reconstruct the topological essence of seepage channels. GNNs abstract the continuous pore space into an undirected graph structure: pore clusters are mapped as nodes, and throats as edges, with node and edge features encoding local geometric properties (such as hydraulic radius, coordination number, and Euler number) and spatial coordinates [54]. Through graph convolution and message-passing mechanisms, GNNs can explicitly capture the global path dependence of fluid flow in complex branching networks and the propagation laws of pressure gradients, significantly enhancing the characterization capability for highly heterogeneous, fracture-vug media. In tests on tight reservoirs, the GNN architecture demonstrated superior topological fidelity compared to grid-based CNNs in predicting the principal directions of anisotropic permeability tensors, effectively suppressing flow field distortions caused by local grid discretization, particularly in low-permeability end-members and high-tortuosity samples [55].
In terms of computational efficiency, surrogate models have achieved a performance leap of several orders of magnitude. Traditional LBM simulations for digital rocks of typical scales often face severe computational bottlenecks and prolonged calculation cycles on parallel computing clusters. In contrast, well-trained CNN/GNN surrogate models can achieve instantaneous flow field inference on standard GPU platforms, resulting in a leapfrog improvement in inference efficiency. This performance breakthrough completely shatters computational barriers, enabling large-scale Monte Carlo stochastic simulations, parameter sensitivity analyses, and real-time wellsite decision-making to transition from theoretical concepts to engineering practice [56].
However, purely data-driven surrogate models still face strict error boundary constraints in deployment. Their prediction accuracy heavily depends on the coverage of the training set distribution: when the porosity-tortuosity combinations, mineral phase ratios, or microfracture development states of test samples exceed the training domain, the models are prone to Out-of-Distribution (OOD) generalization failure, leading to overestimation of permeability or non-physical oscillations in the velocity field [57]. Furthermore, CNNs and GNNs essentially perform statistical interpolation rather than physical deduction, making it difficult to strictly guarantee local mass conservation and momentum balance. Future research needs to embed the continuity equation and Navier–Stokes residuals into the network loss function as soft or hard regularization terms, constructing Physics-Informed Surrogate models. This approach aims to compress systematic errors to engineering-acceptable thresholds while maintaining the advantage of thousand-fold acceleration, thereby laying a high-fidelity, interpretable algorithmic foundation for subsequent multiphase flow dynamics predictions.

5.2. Multiphase Flow Dynamics: End-to-End Prediction of Relative Permeability and Capillary Pressure Curves, and Spatiotemporal Evolution Modeling

Multiphase flow dynamics represents a core challenge in digital rock physics characterization. The physical processes are highly non-linear and strongly dependent on pore-scale wettability distribution, interfacial tension, capillary-force-dominated phase allocation mechanisms, and saturation history (hysteresis effects) (as shown in Figure 10). Traditional experimental acquisition of relative permeability (kr) and capillary pressure (Pc) curves is not only time-consuming and costly but also susceptible to interference from clay mineral hydration swelling and capillary end effects. To overcome this bottleneck, deep learning surrogate models have shifted towards an end-to-end direct prediction mode, bypassing time-consuming unsteady displacement simulations. Representative 3D convolutional architectures, such as 3DSPPConvNet, fuse multi-scale geometric features of digital rocks, wettability angle distributions, and fluid physical property parameters into high-dimensional input tensors. Through multi-scale receptive fields, they extract the non-linear correlations among throat contraction ratios, pore connectivity topology, and fluid allocation mechanisms. Empirical studies demonstrate that in the physics-informed 3DSPPConvNetPhy model, scale information is the most critical physical feature for predicting relative permeability curves. Furthermore, the trained model enables second-level batch inference, providing an efficient alternative for the rapid assessment of core-scale seepage parameters [58].
However, static curve prediction fails to capture the transient dynamic evolution of fluid front propagation, viscous fingering phenomena, and local saturation fields during the displacement process. To address this, deep learning surrogate models (such as Conditional Generative Adversarial Networks) have been applied to pore-scale multiphase flow modeling, enabling the prediction of saturation distribution and front migration patterns during fluid displacement [60]. ConvLSTM, by coupling spatial convolutional feature extraction with temporal sequence memory gating mechanisms, can accurately capture the transient front migration laws and saturation redistribution processes of the non-wetting phase displacing the wetting phase in digital rocks. In dynamic two-phase flow simulations, Physics-Informed Neural Networks (PINNs) further explicitly embed governing partial differential equations (such as the time-dependent Muskat-Leverett system or Navier–Stokes equations) into the network loss function as residual penalty terms [61]. This physical constraint mechanism forces the network outputs to strictly adhere to mass conservation, momentum balance, and thermodynamic boundary conditions, effectively suppressing the numerical divergence and non-physical oscillations common in long-term sequence inference by purely data-driven models, thereby achieving a unity of high fidelity and computational efficiency.
Although multiphase flow surrogate models have demonstrated significant achievements in prediction accuracy and computational acceleration, their engineering deployment still faces three core challenges. First, microscopic physical mechanisms such as dynamic contact angle evolution, interface roughness effects, and non-equilibrium capillary forces are often implicitly represented or oversimplified in existing neural networks, limiting the models’ Out-of-Distribution (OOD) generalization capability under extreme wettability conditions or high-speed displacement scenarios [62]. Second, ConvLSTM and PINNs remain constrained by high training computational costs and rigid time-step constraints when handling 3D unstructured grids and transient multiphase interface tracking. Third, there is no unified architecture for the explicit modeling of saturation history path dependence during imbibition/drainage processes. Future research needs to develop lightweight spatiotemporal convolutional architectures and adaptive time-step algorithms to break through the computational bottlenecks of transient interface tracking; construct long-term evolution models based on neural operators to achieve a transition from single-step snapshot prediction to continuous dynamic inference; and simultaneously develop explicit saturation history encoding mechanisms based on physical constraints, promoting a shift in multiphase flow prediction from static curve fitting to dynamic physical evolution inference.
End-to-end multiphase flow surrogates (3DSPPConvNet, ConvLSTM, PINNs) achieve remarkable speedups, but their applicability is constrained to saturation ranges, wettability conditions, and capillary numbers represented in the training data. Furthermore, PINNs that embed Muskat-Leverett equations as soft constraints still suffer from residual imbalances when the training data are noisy or sparse; hard constraint methods are computationally prohibitive for 3D transient problems. Thus, multiphase surrogates are currently suitable for rapid screening and sensitivity analysis but not yet for certifiable engineering decisions without experimental validation.

5.3. Physical Constraints and Reinforcement Learning: PIML Equation Residual Embedding and Dynamic Wettability Inversion Mechanisms

Purely data-driven neural networks often fall into a black box dilemma in pore-scale multiphase flow predictions due to the lack of physical priors, easily violating mass conservation, momentum balance, or the second law of thermodynamics when operating outside the training distribution. To overcome this limitation, Physics-Informed Machine Learning (PIML) explicitly embeds governing partial differential equations (PDEs) into network optimization objectives, constructing a dual-constraint architecture of data fitting and physical conservation. In single-phase flow field prediction, the continuity condition and momentum conservation residuals of the Navier–Stokes equations are directly added as regularization terms to the loss function. For multiphase seepage dynamics, the time-dependent Muskat-Leverett equations (coupling capillary pressure, relative permeability, and saturation evolution) participate in gradient backpropagation as soft constraints [63]. This mechanism forces network outputs to strictly adhere to fundamental hydrodynamic laws and, through saturation soft clamping and phase boundary curvature regularization, ensures that the volume fractions of wetting and non-wetting phases are always strictly constrained within physically consistent effective value ranges [64]. Relying on the regularization effect of physical residuals, the PIML framework can block the accumulation of non-physical artifacts at the optimization level during long-term displacement inference, effectively suppressing the numerical divergence and oscillation phenomena common in purely data-driven architectures, thereby achieving a unity of high-fidelity flow field reconstruction and long-range computational stability.
Addressing the highly non-linear hysteresis effects and dynamic wettability evolution in multiphase seepage, the synergistic framework of Reinforcement Learning (RL) and Physics-Informed Machine Learning (PIML) provides an adaptive optimization path for solving inverse problems. Traditional empirical models (such as LET or Corey constitutive models) rely heavily on manual parameter tuning for curve fitting, making it difficult to characterize the path dependence of imbibition/drainage and the time-varying features of contact angles. RL algorithms reconstruct the parameter optimization of and curves into a Markov Decision Process (MDP): the agent dynamically adjusts constitutive parameters to perform forward seepage simulations, converting the dynamic matching error between the predicted saturation field and core displacement experimental data into a negative reward signal, while introducing Muskat–Leverett equation residuals as physical constraint penalty terms. Through Proximal Policy Optimization (PPO) or Deep Deterministic Policy Gradient (DDPG) algorithms, the agent automatically explores the parameter space during trial-and-error iterations, causing the predicted curves to asymptotically converge to physical critical limits such as residual oil saturation and maximum water saturation [65]. This synergistic mechanism not only ensures that the inversion results satisfy thermodynamic boundary conditions but also, combined with Physics-Informed Long Short-Term Memory networks (PI-LSTM), successfully predicts fluid displacement signals in time-lapse seismic differential responses. This effectively eliminates the neglect of dynamic phase lag phenomena by traditional static curve representations, achieving a transition from static empirical fitting to dynamic physical inversion [66].
Although the integration of PIML and RL has significantly enhanced the physical self-consistency and engineering applicability of multiphase flow surrogate models, their deployment still faces three core challenges. First, there are difficulties in discretization errors and adaptive weight allocation for high-dimensional partial differential equation residuals on multi-scale unstructured grids; excessively strong physical penalties can suppress the network’s ability to represent natural geological heterogeneity, leading to homogenized generated flow fields [67]. Second, reinforcement learning suffers from insufficient sample efficiency and generalization capability under complex geological conditions; agents require extensive trial-and-error iterations to converge to optimal policies, failing to meet the timeliness requirements for real-time history matching in oilfields. Third, there is no unified architecture for the explicit topological encoding of imbibition/drainage saturation history paths, and Out-of-Distribution (OOD) generalization capabilities across lithologies remain bottlenecked under extreme wettability conditions or high-speed displacement scenarios [68]. Future research needs to develop adaptive optimization algorithms for physical constraint weights and Offline Reinforcement Learning (Offline RL) frameworks, explicitly embedding dynamic wettability migration, interfacial thermodynamic evolution, and stress-seepage coupling mechanisms into reward function design, thereby promoting a transition in multiphase flow inversion from single-condition parameter fitting to multi-scenario adaptive decision-making.
It is important to position the PIML and RL approaches within our three-level framework. Adding PDE residuals to the loss function (PIML) achieves Level 1—it nudges the solution toward physical plausibility but does not guarantee conservation. Some recent works have attempted Level 2 via hard constraints (e.g., divergence-free basis functions), but these remain restricted to 2D or simple geometries. The reinforcement learning inversion framework operates at Level 1 as well, as the reward function includes physical penalties. Level 3 (OOD reliability) is rarely addressed in the PIML literature, despite being the most critical for field applications.

5.4. Engineering Deployment of Surrogate Models: The Transition Path from Solver Replacement to Core Component of Digital Twins

The successful application of machine learning surrogate models in pore-scale seepage simulation marks a shift in their role from computational accelerators in the laboratory stage to core inference engines in digital twin systems for oil and gas field development and underground engineering. Traditional digital rock physics, constrained by the computational rigidity of LBM/PNM, struggles to support large-scale history matching, real-time production optimization, and multi-scenario uncertainty quantification. By establishing an end-to-end mapping between digital rock geometry and macroscopic seepage response, surrogate models enable near-real-time flow field inference on single GPU platforms, achieving an order-of-magnitude leap in computational efficiency. This lays a solid computational foundation for engineering-level real-time decision-making and dynamic optimization. However, realizing the transition from offline solver replacement to online digital twins requires overcoming four engineering deployment barriers: model lightweighting, multi-scale coupling, dynamic data assimilation, and uncertainty quantification (UQ).
At the architectural deployment level, surrogate models need to be upgraded from monolithic inference networks to differentiable, pluggable standardized computational modules. Through structured pruning, knowledge distillation, and INT8/FP16 mixed-precision quantization techniques, the parameter scale of 3D CNN or Graph Neural Network surrogate models can be significantly compressed. This enables lightweight local deployment on edge computing devices (such as downhole microprocessors or field industrial control computers) while strictly controlling the error boundaries of permeability predictions [69]. Meanwhile, cross-platform serialization formats based on ONNX or TorchScript allow surrogate solvers to be seamlessly embedded into mainstream commercial reservoir numerical simulators such as Eclipse, INTERSECT, and tNavigator. This replaces traditional static lookup table methods, achieving dynamic coupling between pore-scale physical mechanisms and macroscopic grid seepage equations [70]. This integration path not only preserves the computational efficiency of surrogate models but also compensates for the stability defects of pure AI architectures in large-scale field applications through the mature grid management and mass conservation verification mechanisms of commercial simulators.
Dynamic data assimilation and online update mechanisms are central to integrating surrogate models into the digital twin closed loop. Traditional static surrogate models are prone to Out-of-Distribution (OOD) failure during long-term reservoir injection and production, effective stress path evolution, or chemical displacement processes. The IDRP framework integrates Ensemble Kalman Filter (EnKF), variational data assimilation, and Online Fine-tuning algorithms to use real-time production dynamics (such as bottom-hole flowing pressure, water cut, and distributed fiber-optic temperature/acoustic monitoring) as observation operators, thereby inversely correcting the latent space feature weights or conditional guidance vectors of the surrogate model [71]. This mechanism enables digital rock characterization to adaptively update with field injection history and stress states, achieving a real-time sensing-inference-correction closed loop. Combined with Physics-Informed Long Short-Term Memory networks (PI-LSTM), the surrogate model can further output time-varying seepage parameter predictions with confidence intervals, providing probabilistic decision support for dynamic production allocation, injection-production optimization, and risk warning.
The construction of an Uncertainty Quantification (UQ) framework is the cornerstone for gaining engineering trust and ensuring compliant deployment of surrogate models. Pure point estimates fail to reflect the propagation effects of geological heterogeneity, imaging annotation noise, and model structural errors, which can easily lead to misjudgments in safety-critical scenarios. Current deployment frontiers adopt architectures such as Deep Ensembles, Monte Carlo Dropout, and Variational Bayesian Neural Networks (BNNs) to output prediction distributions rather than single scalars in a single forward pass. By constructing joint probability density fields for porosity–permeability–capillary pressure, the UQ module can quantitatively assess the prediction confidence of surrogate models in low-permeability end-members, fracture-developed zones, or high-water-cut periods. It directly feeds error boundaries back to upstream super-resolution reconstruction and segmentation networks, forming a self-optimizing workflow of error tracing—geometric correction—flow field re-evaluation. This mechanism effectively alleviates the credibility crisis of black box models in high-risk scenarios such as CO2 storage leakage assessment and coalbed methane water inrush warning, providing auditable mathematical bases for regulatory compliance and engineering acceptance.
Although surrogate models demonstrate significant potential in engineering deployment, their large-scale application still faces three practical challenges. First, the joint training of multi-physics coupled (thermal–hydro-mechanical–chemical) surrogate models requires massive amounts of high-fidelity labeled data, whereas real reservoir dynamic monitoring data are often sparse, asynchronous, and heavily noisy, constraining the model’s generalization capability and long-term stability under complex operating conditions. Second, interface standards between commercial simulators and open-source AI frameworks are not yet unified; the lack of industry norms for cross-platform data interaction, version management, and computational resource scheduling leads to high technology integration and operation/maintenance costs. Third, the computational bottlenecks of edge computing hardware and the extreme downhole environments (high temperature, high pressure, strong electromagnetic interference) impose stringent requirements on model robustness and hardware-level fault tolerance. Future research needs to promote the development of hybrid architectures combining rock physics foundation models + differentiable manifold optimization, establish cross-institutional digital rock benchmark datasets and standardized API protocols, and explore hardware-level acceleration paths for pore-scale flow field inference using neuromorphic computing and photonic chips, ultimately achieving the comprehensive transformation of IDRP surrogate models from laboratory algorithm verification to oilfield-level digital twin infrastructure.
Table 1 systematically summarizes typical performance for super-resolution, 3D generation, segmentation, single-phase surrogates, multiphase surrogates, and physics-informed/RL frameworks.

6. Conclusions

The emergence of Intelligent Digital Rock Physics (IDRP) marks a potential paradigm shift in the microscopic characterization of porous media and seepage mechanics research. Distinct from traditional DRP, which has long been constrained by the inherent mutual exclusivity between optical resolution and macroscopic field of view, the prohibitive computational cost of direct numerical simulation, and the statistical failures of single-scale geometric characterization, IDRP introduces three novel contributions: (i) break through the diffraction limit deep learning-based super-resolution reconstruction and diffusion probabilistic models for 3D generation, algorithmically mitigating the resolution–field-of-view trade-off via statistical inference of sub-resolution features without hardware upgrades; (ii) cross-modal data fusion that integrates multi-scale imaging and petrophysical measurements; (iii) physics-informed surrogate solvers and reinforcement learning frameworks that have demonstrated thousand- to ten-thousand-fold acceleration in controlled test cases while preserving mass and momentum conservation.
This systematic review profoundly confirms that the data-driven + physics-constrained dual-drive mechanism is the central design principle for IDRP to achieve engineering credibility and scientific self-consistency. While purely data-driven architectures possess excellent non-linear mapping and acceleration capabilities, they are prone to violating mass conservation, momentum balance, and thermodynamic laws outside the training distribution, falling into the dilemma of black-box extrapolation. Conversely, traditional physical models, although mechanistically clear, are constrained by the stiffness of high-dimensional solutions and parameter uncertainty. By embedding governing partial differential equations such as Navier–Stokes and Muskat–Leverett into network optimization objectives as soft or hard regularization terms, and integrating reinforcement learning with uncertainty quantification frameworks, IDRP can better respect the fundamental conservation laws of fluid dynamics and rock mechanics while maintaining massive computational acceleration. This fusion ensures that algorithmic outputs possess not only high fitting accuracy in a statistical sense but also physical interpretability, boundary controllability, and cross-domain extrapolation capabilities—a combination that has not been systematically achieved in previous DRP workflows, but whose practical realization requires careful handling of residual errors, weight balancing, and out-of-distribution robustness.
The in-depth development and large-scale implementation of IDRP are more than a linear progression of single algorithmic technologies, but the inevitable result of deep synergy among geology, computational mathematics [72], artificial intelligence, and rock mechanics/seepage physics [73]. From a practical significance perspective, IDRP has shown potential in fields directly aligned with national energy security and the Dual Carbon goals: efficient development of unconventional oil and gas (e.g., reducing core analysis time from weeks to hours), enhanced production of coalbed methane, geological carbon storage (CCUS) with real-time wettability adaptation, geothermal resource exploitation, and safe underground space utilization. By constructing high-precision, dynamically evolving digital rock twins, IDRP could eventually transform traditional core analysis workflows, achieving seamless integration from laboratory-scale microscopic characterization to field-scale macroscopic production optimization. A preliminary roadmap for industrial deployment includes: (i) establishing standardized benchmark datasets for heterogeneous rock types; (ii) developing rock multi-modal foundation models for zero-shot generalization; (iii) embedding uncertainty quantification into surrogate solvers for risk-aware decision-making.
In the future, with continued advances in data availability, model architectures, and validation protocols, IDRP may gradually move beyond the algorithm verification stage toward integration with subsurface digital twin infrastructure. However, it requires sustained interdisciplinary collaboration, open sharing of benchmark datasets, and—most critically—rigorous, reproducible demonstrations of IDRP’s predictive reliability under field-relevant conditions. Until then, IDRP should be viewed as a powerful complement to traditional physical experiments and numerical simulations.

Author Contributions

X.L. (Xue Li), writing—original draft and investigation; L.Z., investigation and data curation; F.G., conceptualization and writing—review and editing; X.L. (Xin Liang), conceptualization and methodology; Z.C., resources and writing—review and editing. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by Xi’an University of Technology (No. 451117008).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

References

  1. Liu, M.; Mukerji, T. Digital transformation in rock physics: Deep learning and data fusion. Lead. Edge 2022, 41, 591–598. [Google Scholar] [CrossRef]
  2. Andrä, H.; Combaret, N.; Dvorkin, J.; Glatt, E.; Han, J.; Kabel, M.; Keehm, Y.; Krzikalla, F.; Lee, M.; Madonna, C.; et al. Digital rock physics benchmarks—Part I: Imaging and segmentation. Comput. Geosci. 2013, 50, 25–32. [Google Scholar] [CrossRef]
  3. Liang, J.; He, L.; Chai, W.; Jia, N.; Liu, R. A new paradigm for physics-informed AI-driven reservoir research: From multiscale characterization to intelligent seepage simulation. Energies 2026, 19, 270. [Google Scholar] [CrossRef]
  4. Carrillo, F.J.; Soulaine, C.; Bourg, I.C. The impact of sub-resolution porosity on numerical simulations of multiphase flow. Adv. Water Resour. 2022, 161, 104094. [Google Scholar] [CrossRef]
  5. Saxena, N.; Hows, A.; Hofmann, R.; Alpak, F.O.; Dietderich, J.; Appel, M.; Freeman, J.; De Jong, H. Rock properties from micro-CT images: Digital rock transforms for resolution, pore volume, and field of view. Adv. Water Resour. 2019, 134, 103419. [Google Scholar] [CrossRef]
  6. Peta, K.; Kubiak, K.J.; Brown, C.A. Multiscale geometric analysis of dynamic wettability on complex, fractal-like, anisotropic surfaces. Measurement 2026, 264, 120328. [Google Scholar] [CrossRef]
  7. Soltanmohammadi, R.; Faroughi, S.A. A comparative analysis of super-resolution techniques for enhancing micro-CT images of carbonate rocks. Appl. Comput. Geosci. 2023, 20, 100143. [Google Scholar] [CrossRef]
  8. Keller, L.M. On the representative elementary volumes of clay rocks at the mesoscale. J. Geol. Min. Res. 2015, 7, 58–64. [Google Scholar] [CrossRef]
  9. Regaieg, M.; Varloteaux, C.; Farhana Faisal, T.; ElAbid, Z. Towards large-scale DRP simulations: Generation of large super-resolution images and extraction of large pore network models. Transp. Porous Media 2023, 147, 375–399. [Google Scholar] [CrossRef]
  10. Aidun, C.K.; Clausen, J.R. Lattice-Boltzmann method for complex flows. Annu. Rev. Fluid Mech. 2010, 42, 439–472. [Google Scholar] [CrossRef]
  11. Chen, T.; Ning, Y.; Amritkar, A.; Qin, G. Multi-gpu solution to the lattice Boltzmann method: An application in multiscale digital rock simulation for shale formation. Concurr. Comput. Pract. Exp. 2018, 30, e4530. [Google Scholar] [CrossRef]
  12. Zubov, A.S.; Murygin, D.A.; Gerke, K.M. Pore-network extraction using discrete Morse theory: Preserving the topology of the pore space. Phys. Rev. E 2022, 106, 055304. [Google Scholar] [CrossRef] [PubMed]
  13. Baychev, T.G.; Jivkov, A.P.; Rabbani, A.; Raeini, A.Q.; Xiong, Q.; Lowe, T.; Withers, P.J. Reliability of algorithms interpreting topological and geometric properties of porous media for pore network modelling. Transp. Porous Media 2019, 128, 271–301. [Google Scholar] [CrossRef]
  14. Yang, H.; Zhang, Y.; Cui, Z.; Xu, Y.; Yang, Y. DGRN: Image super-resolution with dual gradient regression guidance. Comput. Graph. 2023, 110, 141–150. [Google Scholar] [CrossRef]
  15. Molnar, S.M.; Godfrey, J.; Song, B. Balance equations for physics-informed machine learning. Heliyon 2024, 10, e38799. [Google Scholar] [CrossRef] [PubMed]
  16. Zhang, K.; Gao, X.; Tao, D.; Li, X. Single image super-resolution with non-local means and steering kernel regression. IEEE Trans. Image Process. 2012, 21, 4544–4556. [Google Scholar] [CrossRef] [PubMed]
  17. Yang, W.; Feng, J.; Yang, J.; Zhao, F.; Liu, J.; Guo, Z.; Yan, S. Deep edge guided recurrent residual learning for image super-resolution. IEEE Trans. Image Process. 2017, 26, 5895–5907. [Google Scholar] [CrossRef] [PubMed]
  18. Shan, L.; Liu, C.; Liu, Y.; Tu, Y.; Chilukoti, S.V.; Hei, X. Single image multi-scale enhancement for rock micro-CT super-resolution using residual U-Net. Appl. Comput. Geosci. 2024, 22, 100165. [Google Scholar] [CrossRef]
  19. Shan, L.; Liu, C.; Liu, Y.; Kong, W.; Hei, X. Rock CT image super-resolution using residual dual-channel attention generative adversarial network. Energies 2022, 15, 5115. [Google Scholar] [CrossRef]
  20. Hou, Z.; Cao, D.; Ji, S.; Cui, R.; Liu, Q. Enhancing digital rock image resolution with a GAN constrained by prior and perceptual information. Comput. Geosci. 2021, 157, 104939. [Google Scholar] [CrossRef]
  21. Luo, X.; Sun, J.; Zhang, R.; Chi, P.; Cui, R. A multi-condition denoising diffusion probabilistic model controls the reconstruction of 3D digital rocks. Comput. Geosci. 2024, 184, 105541. [Google Scholar] [CrossRef]
  22. de Castro Vargas Fernandes, J.; Duarte Vidal, A.; Carvalho Medeiros, L.; Menezes dos Anjos, C.E.; Surmas, R.; Gonçalves Evsukoff, A. Absolute permeability estimation from microtomography rock images through deep learning super-resolution and adversarial fine tuning. Sci. Rep. 2024, 14, 16704. [Google Scholar] [CrossRef] [PubMed]
  23. Andrew, M. A quantified study of segmentation techniques on synthetic geological XRM and FIB-SEM images. Comput. Geosci. 2018, 22, 1503–1512. [Google Scholar] [CrossRef]
  24. Sadeghnejad, S.; Hupfer, S.; Pingel, J.; Lanyon, B.; Schneeberger, R.; Blechschmidt, I.; Alonso, U.; Hauser, W.; Kraft, S.; Geckeis, H.; et al. Bentonite mass loss in fractured crystalline rock quantified from CT scans using digital rock physics and machine learning: Case study from the Grimsel Test Site (Switzerland). Appl. Clay Sci. 2025, 276, 107915. [Google Scholar] [CrossRef]
  25. Chen, Z.; Liu, X.; Yang, J.; Little, E.; Zhou, Y. Deep learning-based method for SEM image segmentation in mineral characterization, an example from Duvernay Shale samples in Western Canada Sedimentary Basin. Comput. Geosci. 2020, 138, 104450. [Google Scholar] [CrossRef]
  26. Cao, D.; Ji, S.; Cui, R.; Liu, Q. Multi-task learning for digital rock segmentation and characteristic parameters computation. J. Pet. Sci. Eng. 2022, 208, 109202. [Google Scholar] [CrossRef]
  27. Wu, Y.; An, S.; Tahmasebi, P.; Liu, K.; Lin, C.; Kamrava, S.; Liu, C.; Yu, C.; Zhang, T.; Sun, S.; et al. end-to-end approach to predict physical properties of heterogeneous porous media: Coupling deep learning and physics-based features. Fuel 2023, 352, 128753. [Google Scholar] [CrossRef]
  28. Kang, J.; Li, N.Y.; Zhao, L.Q.; Xiong, G.; Wang, D.C.; Xiong, Y.; Luo, Z.F. Construction of complex digital rock physics based on full convolution network. Pet. Sci. 2022, 19, 651–662. [Google Scholar] [CrossRef]
  29. Chen, Z.; Yuan, F.; Li, X.; Zhang, M.; Zheng, C. A novel few-shot learning framework for rock images dually driven by data and knowledge. Appl. Comput. Geosci. 2024, 21, 100155. [Google Scholar] [CrossRef]
  30. Wang, Z.; Hou, Z.; Cao, D. Enhancing SAM-based digital rock image segmentation via edge-semantics fusion. Appl. Comput. Geosci. 2025, 28, 100292. [Google Scholar] [CrossRef]
  31. Wang, Z.; Hou, Z.; Hou, S.; Cao, D. Generalizable Digital Rock Image Segmentation under Limited Data with the Segment Anything Model. Intell. Geoengin. 2026, 3, 1–10. [Google Scholar] [CrossRef]
  32. Geng, Z.; Johnson, D.; Fedkiw, R. Coercing machine learning to output physically accurate results. J. Comput. Phys. 2020, 406, 109099. [Google Scholar] [CrossRef]
  33. Li, J.; Meng, Y.; Xia, J.; Xu, K.; Sun, J. A physics-constrained two-stage GAN for reservoir data generation: Enhancing predictive accuracy. Eng. Res. Express 2025, 7, 035257. [Google Scholar] [CrossRef]
  34. Emeliana, Y.; Melkisedek, B.L.; Dharmawan, I.A. Integrating lattice Boltzmann and neural networks for modeling transport parameters in porous rocks: A systematic review. Results Eng. 2025, 29, 108688. [Google Scholar] [CrossRef]
  35. Mosser, L.; Dubrule, O.; Blunt, M.J. Reconstruction of three-dimensional porous media using generative adversarial neural networks. Phys. Rev. E 2017, 96, 043309. [Google Scholar] [CrossRef] [PubMed]
  36. Wang, H.; Yang, X.; Zhou, C.; Yan, J.; Yu, J.; Xie, K. Constructions of multi-scale 3D digital rocks by associated image segmentation method. Front. Earth Sci. 2025, 12, 1518561. [Google Scholar] [CrossRef]
  37. Abdollahifard, M.J. Fast multiple-point simulation using a data-driven path and an efficient gradient-based search. Comput. Geosci. 2016, 86, 64–74. [Google Scholar] [CrossRef]
  38. Zhang, T.; Liu, Q.; Wang, X.; Ji, X.; Du, Y. A 3D reconstruction method of porous media based on improved WGAN-GP. Comput. Geosci. 2022, 165, 105151. [Google Scholar] [CrossRef]
  39. Li, K.; Li, H.; Khudhair, A.; Yan, J.; Wang, B. Efficient generation of digital rock CT images using LoRA-enhanced stable diffusion models. Intell. Geoengin. 2025, 2, 96–108. [Google Scholar] [CrossRef]
  40. Naiff, D.; Schaeffer, B.P.; Pires, G.; Stojkovic, D.; Rapstine, T.; Ramos, F. Controlled latent diffusion models for 3D porous media reconstruction. Comput. Geosci. 2025, 206, 106038. [Google Scholar] [CrossRef]
  41. Zhu, L.; Bijeljic, B.; Blunt, M.J. Diffusion model-based generation of three-dimensional multiphase pore-scale images. Transp. Porous Media 2025, 152, 22. [Google Scholar] [CrossRef]
  42. Shah, S.M.; Gray, F.; Crawshaw, J.P.; Boek, E.S. Micro-computed tomography pore-scale study of flow in porous media: Effect of voxel resolution. Adv. Water Resour. 2016, 95, 276–287. [Google Scholar] [CrossRef]
  43. Xiong, T.; Chen, M.; Jin, Y.; Zhang, W.; Shao, H.; Wang, G.; Long, E.; Long, W. A new multi-scale method to evaluate the porosity and MICP curve for digital rock of complex reservoir. Energies 2023, 16, 7613. [Google Scholar] [CrossRef]
  44. Reimers, I.; Safonov, I.; Kornilov, A.; Yakimchuk, I. Two-stage alignment of FIB-SEM images of rock samples. J. Imaging 2020, 6, 107. [Google Scholar] [CrossRef] [PubMed]
  45. Yan, W.; Chi, P.; Golsanami, N.; Sun, J.; Xing, H.; Li, S.; Dong, H. Analysis of reconstructed multisource and multiscale 3-D digital rocks based on the cycle-consistent generative adversarial network method. Geophys. J. Int. 2023, 235, 736–749. [Google Scholar] [CrossRef]
  46. Kazak, A.; Simonov, K.; Kulikov, V. Machine-learning-assisted segmentation of focused ion beam-scanning electron microscopy images with artifacts for improved void-space characterization of tight reservoir rocks. SPE J. 2021, 26, 1739–1758. [Google Scholar] [CrossRef]
  47. Carrillo, F.J.; Bourg, I.C.; Soulaine, C. Multiphase flow modeling in multiscale porous media: An open-source micro-continuum approach. J. Comput. Phys. X 2020, 8, 100073. [Google Scholar] [CrossRef]
  48. Jiang, J.; Younis, R.M. Hybrid coupled discrete-fracture/matrix and multicontinuum models for unconventional-reservoir simulation. SPE J. 2016, 21, 1009–1027. [Google Scholar] [CrossRef]
  49. Morales, F.A.; Showalter, R.E. A Darcy–Brinkman model of fractures in porous media. J. Math. Anal. Appl. 2017, 452, 1332–1358. [Google Scholar] [CrossRef]
  50. Chen, Y.; Dong, P. Modeling of characteristics of complex microstructure and heterogeneity at the core scale. Appl. Sci. 2024, 14, 11385. [Google Scholar] [CrossRef]
  51. Xue, K.; Pu, H.; Li, M.; Luo, P.; Liu, D.; Yi, Q. Fractal-based analysis of stress-induced dynamic evolution in geometry and permeability of porous media. Phys. Fluids 2025, 37, 036630. [Google Scholar] [CrossRef]
  52. Santos, J.E.; Xu, D.; Jo, H.; Landry, C.J.; Prodanović, M.; Pyrcz, M.J. PoreFlow-Net: A 3D convolutional neural network to predict fluid flow through porous media. Adv. Water Resour. 2020, 138, 103539. [Google Scholar] [CrossRef]
  53. Nabipour, I.; Raoof, A.; Cnudde, V.; Aghaei, H.; Qajar, J. A computationally efficient modeling of flow in complex porous media by coupling multiscale digital rock physics and deep learning: Improving the tradeoff between resolution and field-of-view. Adv. Water Resour. 2024, 188, 104695. [Google Scholar] [CrossRef]
  54. Zhao, Q.; Han, X.; Guo, R.; Chen, C. A computationally efficient hybrid neural network architecture for porous media: Integrating convolutional and graph neural networks for improved property predictions. Adv. Water Resour. 2025, 195, 104881. [Google Scholar] [CrossRef]
  55. Alzahrani, M.K.; Shapoval, A.; Chen, Z.; Rahman, S.S. Pore-GNN: A graph neural network-based framework for predicting flow properties of porous media from micro-CT images. Adv. Geo-Energy Res. 2023, 10, 39–55. [Google Scholar] [CrossRef]
  56. Chandra, A.; Koch, M.; Pawar, S.; Panda, A.; Azizzadenesheli, K.; Snippe, J.; Alpak, F.O.; Hariri, F.; Etienam, C.; Devarakota, P.; et al. Accelerating porous media flow simulations with Fourier neural operators: An application to geologic storage of CO2. Adv. Theory Simul. 2026, 9, e00747. [Google Scholar] [CrossRef]
  57. Shmuel, A.; Glickman, O.; Lazebnik, T. Machine and deep learning performance in out-of-distribution regressions. Mach. Learn. Sci. Technol. 2024, 5, 045078. [Google Scholar] [CrossRef]
  58. Xie, C.; Zhu, J.; Yang, H.; Wang, J.; Liu, L.; Song, H. Relative permeability curve prediction from digital rocks with variable sizes using deep learning. Phys. Fluids 2023, 35, 096605. [Google Scholar] [CrossRef]
  59. Helland, J.O.; Jettestuen, E.; Aursjø, O. Pore-scale simulation of relative permeability hysteresis from a workflow of level-set and lattice-Boltzmann methods: The case of consolidated media with low pore-to-throat aspect ratio. Transp. Porous Media 2026, 153, 57. [Google Scholar] [CrossRef]
  60. Wang, Z.; Jeong, H.; Gan, Y.; Pereira, J.M.; Gu, Y.; Sauret, E. Pore-scale modeling of multiphase flow in porous media using a conditional generative adversarial network (cGAN). Phys. Fluids 2022, 34, 123325. [Google Scholar] [CrossRef]
  61. Imankulov, T.; Kuljabekov, A.; Bekele, S.D.; Zhantayev, Z.; Assilbekov, B.; Kenzhebek, Y. A systematic analysis of physics-informed neural networks for two-phase flow with capillarity: The Muskat–Leverett problem. Appl. Sci. 2025, 15, 13011. [Google Scholar] [CrossRef]
  62. Asadolahpour, S.R.; Jiang, Z.; Lewis, H.; Min, C. Deep learning for pore-scale two-phase flow: Modelling drainage in realistic porous media. Pet. Explor. Dev. 2024, 51, 1301–1315. [Google Scholar] [CrossRef]
  63. Wang, Y.; Lin, G. Efficient deep learning techniques for multiphase flow simulation in heterogeneous porousc media. J. Comput. Phys. 2020, 401, 108968. [Google Scholar] [CrossRef]
  64. Patel, H.V.; Panda, A.; Kuipers, J.A.M.; Peters, E.A.J.F. Computing interface curvature from volume fractions: A machine learning approach. Comput. Fluids 2019, 193, 104263. [Google Scholar] [CrossRef]
  65. Dong, P.; Liao, X.; Chen, Z. An approach for automatic parameters evaluation in unconventional oil reservoirs with deep reinforcement learning. J. Pet. Sci. Eng. 2022, 209, 109917. [Google Scholar] [CrossRef]
  66. Zhao, T.; Chen, G.; Pang, C.; Seenoi, P.; Papukdee, N.; Busababodhin, P. Time-lapse earthquake difference prediction based on physics-informed long short-term memory coupled with interpretability boosting. J. Seism. Explor. 2025, 34, 25. [Google Scholar] [CrossRef]
  67. Plankovskyy, S.; Tsegelnyk, Y.; Shyshko, N.; Litvinchev, I.; Romanova, T.; Velarde Cantú, J.M. Review of physics-informed neural networks: Challenges in loss function design and geometric integration. Mathematics 2025, 13, 3289. [Google Scholar] [CrossRef]
  68. Li, K.; Rubungo, A.N.; Lei, X.; Persaud, D.; Choudhary, K.; DeCost, B.; Dieng, A.B.; Hattrick-Simpers, J. Probing out-of-distribution generalization in machine learning for materials. Commun. Mater. 2025, 6, 9. [Google Scholar] [CrossRef]
  69. Daribayev, B.; Azatbekuly, N.; Mukhanbet, A. Optimization of neural networks for predicting oil recovery factor using quantization techniques. J. Probl. Comput. Sci. Inf. Technol. 2024, 2, 25–33. [Google Scholar] [CrossRef]
  70. Maldonado-Cruz, E.; Pyrcz, M.J. Fast evaluation of pressure and saturation predictions with a deep learning surrogate flow model. J. Pet. Sci. Eng. 2022, 212, 110244. [Google Scholar] [CrossRef]
  71. Wen, X.H.; Chen, W.H. Real-time reservoir model updating using ensemble Kalman filter with confirming option. SPE J. 2006, 11, 431–442. [Google Scholar] [CrossRef]
  72. Pachalieva, A.; O’Malley, D.; Harp, D.R.; Viswanathan, H. Physics-informed machine learning with differentiable programming for heterogeneous underground reservoir pressure management. Sci. Rep. 2022, 12, 18734. [Google Scholar] [CrossRef] [PubMed]
  73. Haghighat, E.; Raissi, M.; Moure, A.; Gomez, H.; Juanes, R. A physics-informed deep learning framework for inversion and surrogate modeling in solid mechanics. Comput. Methods Appl. Mech. Eng. 2021, 379, 113741. [Google Scholar] [CrossRef]
Figure 2. Schematic diagram illustrating the mutual exclusivity between resolution and field of view in Micro-CT imaging. (a) Low-resolution image; (b) High-resolution image [7].
Figure 2. Schematic diagram illustrating the mutual exclusivity between resolution and field of view in Micro-CT imaging. (a) Low-resolution image; (b) High-resolution image [7].
Applsci 16 06118 g002
Figure 3. Visual comparison of traditional bicubic interpolation and SRCNN super-resolution methods on carbonate rock Micro-CT images. (a) Results from traditional bicubic interpolation; (b) Results from SRCNN deep learning super-resolution [7].
Figure 3. Visual comparison of traditional bicubic interpolation and SRCNN super-resolution methods on carbonate rock Micro-CT images. (a) Results from traditional bicubic interpolation; (b) Results from SRCNN deep learning super-resolution [7].
Applsci 16 06118 g003
Figure 4. Comparison of reconstruction results of different generations of super-resolution models on carbonate rock Micro-CT images (where the red box indicates the region of interest (ROI)) [18].
Figure 4. Comparison of reconstruction results of different generations of super-resolution models on carbonate rock Micro-CT images (where the red box indicates the region of interest (ROI)) [18].
Applsci 16 06118 g004
Figure 5. Histogram of grayscale value distribution for multiple mineral phases in CT scan images and selection of regions of interest (x-axis: grayscale intensity, y-axis: frequency) [24].
Figure 5. Histogram of grayscale value distribution for multiple mineral phases in CT scan images and selection of regions of interest (x-axis: grayscale intensity, y-axis: frequency) [24].
Applsci 16 06118 g005
Figure 7. Comparison of the porosity frequency distribution of output samples from physics−constrained generative models with target prior values. (a) Histogram comparison of porosity frequency distributions between generated samples and target priors; (b) Analysis of the consistency in permeability distribution [34].
Figure 7. Comparison of the porosity frequency distribution of output samples from physics−constrained generative models with target prior values. (a) Histogram comparison of porosity frequency distributions between generated samples and target priors; (b) Analysis of the consistency in permeability distribution [34].
Applsci 16 06118 g007
Figure 8. Reverse generation pipeline for 3D pore structures based on Latent Diffusion Models (LDM) [40].
Figure 8. Reverse generation pipeline for 3D pore structures based on Latent Diffusion Models (LDM) [40].
Applsci 16 06118 g008
Figure 9. Residual analysis of the multi-scale CNN permeability prediction model [53].
Figure 9. Residual analysis of the multi-scale CNN permeability prediction model [53].
Applsci 16 06118 g009
Figure 10. Evolution characteristics of capillary pressure and interfacial area under different saturation histories. (a) Capillary pressure–saturation hysteresis curves during drainage and imbibition processes; (b) Evolution relationship of fluid interfacial area with saturation [59].
Figure 10. Evolution characteristics of capillary pressure and interfacial area under different saturation histories. (a) Capillary pressure–saturation hysteresis curves during drainage and imbibition processes; (b) Evolution relationship of fluid interfacial area with saturation [59].
Applsci 16 06118 g010
Table 1. Summary of typical performance of major IDRP methods.
Table 1. Summary of typical performance of major IDRP methods.
Method CategoryRepresentative TechniquesTypical Performance
Super-resolutionBicubic interpolationVery fast; blurry edges; no new detail
CNN (SRCNN, EDSR, RCAN)Fast; good PSNR; oversmoothed textures
GAN (SRGAN, ESRGAN)Moderate speed; sharp textures; risk of artifacts
Diffusion (DDPM, LDM)Slow (iterative); topologically faithful; diverse outputs
3D Pore GenerationVAEFast generation; smooth outputs
GAN (WGAN-GP, CVAE-GAN)Realistic texture; moderate speed
Diffusion (DDPM, LDM)High topological diversity; conditional control
Semantic SegmentationGlobal thresholding (Otsu)Very fast; fails on overlapping grayscale
U-Net (CNN)High pixel accuracy; needs large labeled data
Foundation models (SAM)Zero-shot capable; good edge detection
Single-Phase Flow Surrogate3D CNN (PoreFlow-Net)Very high permeability
Graph Neural NetworkExcellent on fractured/vuggy media
Fourier Neural OperatorResolution-invariant; fast
Multiphase Flow Surrogate3DSPPConvNetGood relative permeability; faster than PNM
ConvLSTMCaptures front migration; moderate speedup
PINN (Muskat-Leverett)PDE residuals reduced; better extrapolation
Physics-Informed + RLPIML + PPO/DDPGInverts dynamic wettability; reduces non-physical outputs
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Li, X.; Zhu, L.; Gao, F.; Liang, X.; Cao, Z. Intelligent Digital Rock Physics: Advances and Perspectives from Imaging Reconstruction to Pore-Scale Multiphase Flow Simulation. Appl. Sci. 2026, 16, 6118. https://doi.org/10.3390/app16126118

AMA Style

Li X, Zhu L, Gao F, Liang X, Cao Z. Intelligent Digital Rock Physics: Advances and Perspectives from Imaging Reconstruction to Pore-Scale Multiphase Flow Simulation. Applied Sciences. 2026; 16(12):6118. https://doi.org/10.3390/app16126118

Chicago/Turabian Style

Li, Xue, Lin Zhu, Feng Gao, Xin Liang, and Zhengzheng Cao. 2026. "Intelligent Digital Rock Physics: Advances and Perspectives from Imaging Reconstruction to Pore-Scale Multiphase Flow Simulation" Applied Sciences 16, no. 12: 6118. https://doi.org/10.3390/app16126118

APA Style

Li, X., Zhu, L., Gao, F., Liang, X., & Cao, Z. (2026). Intelligent Digital Rock Physics: Advances and Perspectives from Imaging Reconstruction to Pore-Scale Multiphase Flow Simulation. Applied Sciences, 16(12), 6118. https://doi.org/10.3390/app16126118

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop