A Hybrid Deep Learning Approach for Small-Sample TOC Prediction in Saline Lacustrine Shale
Abstract
1. Introduction
2. Convolutional Autoencoder-BP Neural Network Model
- (1)
- Data preparation stage: The dataset comprises labeled samples with TOC annotations and unlabeled logging data.
- (2)
- CAE training stage: The model is first initialized, after which the CAE is trained on unlabeled data to learn intrinsic feature relationships, reconstruct the input features, and minimize the reconstruction loss between the input and its reconstruction (Figure 1a).
- (3)
- TOC prediction stage: In the backbone, the pretrained CAE is fine-tuned on the labeled data, and the encoder outputs are fed into a BPNN to yield the backbone prediction; in the residual branch, a skip connection is applied to estimate the residual and generate the residual prediction. The final prediction is obtained by summing the backbone and residual outputs (Figure 1b).
2.1. Convolutional Autoencoder
- (1)
- Encoder: The encoder maps the input to a low-dimensional latent space that captures principal feature relationships and preserves the salient information of the original data.where X denotes the input data; Conv denotes the convolution; z denotes the potential space.
- (2)
- Decoder: The decoder maps the low-dimensional latent representation back to the original feature space to reconstruct the input data.wheredenotes the output of the reconstruction; DeConv denotes the transposed convolution.
- (3)
- Training objective: Minimize the MSE between the input and the reconstructed output, thereby enabling the model to learn intrinsic relationships among features.where denotes the MSE loss function.
2.2. BP Neural Network Model
3. An Application Example for Shale Oil
3.1. Field Background
3.2. Related Work
3.3. Building Convolutional Autoencoder-BP Neural Network Model
- (1)
- The input features are standardized as described in Equation (5).
- (2)
- The encoder applies four layers of 1D convolution:where x denotes the standardized features of the input; BatchNorm denotes the batch normalization; Dropout denotes the regularization; MaxPool denotes the maximum pooling; ReLU denotes the activation function; Conv denotes the convolution.
- (3)
- The decoder mirrors the encoder. Four layers of transposed convolution are applied, and the final layer uses Tanh to output the reconstructed feature:where DeConv denotes the transposed convolution; Linear denotes the linear processing.
- (4)
- The autoencoder is trained to minimize a weighted loss between the input and the reconstructed output. MSE is calculated as described in Equation (3).
- (5)
- The TOC predictor comprises a prediction backbone and a residual branch. The final prediction is obtained by summing the outputs of the two components.where denotes the main branch prediction; denotes the prediction of residual branch; Flatten denotes flattening; denotes the bias parameter; denotes the weight; denotes the final predicted value.
- (6)
- Finally, R2 and RMSE are used to evaluate the model; R2 was shown in Equation (4) and RMSE was shown in Equation (6).
3.4. Building BP Neural Network Model
3.5. Building Convolutional Neural Network Model
- (1)
- The input features are standardized as described in Equation (5).
- (2)
- Three layers of 1D convolution are applied:where X denotes the standardized features; BatchNorm denotes the batch normalization; Dropout denotes the regularization; Relu denotes the activation function; MaxPool denotes the maximum pooling.
- (3)
- The output from the final convolution and pooling is flattened, followed by a three-layer fully connected network for prediction:where Flatten denotes flattening; Linear denotes the linear processing; denotes the predicted value of TOC.
- (4)
- Finally, R2 and RMSE are used to evaluate the model; R2 was shown in Equation (4) and RMSE was shown in Equation (6).
3.6. Building Machine Learning Model
3.6.1. Building Random Forest Model
- (1)
- The input features are standardized as described in Equation (5).
- (2)
- The RF regressor consists of multiple decision trees. Each tree is trained on a subset drawn from the dataset by Bootstrap sampling. At each node split, a random subset of features is evaluated to find the best split, and the tree then produces a prediction. The final prediction is the average of all tree predictions:where denotes the prediction of a single tree; M denotes the total number of trees; denotes the predicted value of TOC; denotes the parameter space; denotes cross-validation; denotes the hyperparameter combination; arg min denotes the operation of the parameter when the minimum value is taken.
- (3)
- Finally, R2 and RMSE are used to evaluate the model; R2 was shown in Equation (4) and RMSE was shown in Equation (6).
3.6.2. Building Gradient Boosting Decision Tree Model
- (1)
- The input features are standardized as described in Equation (5).
- (2)
- The model is initialized.where denotes the initial predicted value; denotes the mean value of the training set target.
- (3)
- Iterative training includes residual computation, new tree fitting, model updating, and final prediction.where denotes the residual; denotes the true value of TOC; denotes the predicted value of the TOC; denotes the output of the tree; denotes the learning rate; denotes the final predicted value.
- (4)
- Finally, R2 and RMSE are used to evaluate the model; R2 was shown in Equation (4) and RMSE was shown in Equation (6).
4. Results
4.1. Geochemical Characteristics
4.2. Different Models’ Performance
4.2.1. BP Neural Network Model Performance
4.2.2. Random Forest Model Performance
4.2.3. Gradient Boosting Decision Tree Model Performance
4.2.4. Convolutional Neural Network Model Performance
4.2.5. Convolutional Autoencoder-BP Neural Network Model Performance
4.3. Validation of the Generalization Ability of the Model
4.4. Movability Evaluation of Shale Oil
4.5. Additional Robustness Checks and Methodological Limitations
5. Discussion
6. Conclusions
- (1)
- CAE-BPNN achieved the highest TOC prediction accuracy in P1f (R2 is 0.89; RMSE is 0.061). It outperforms GBDT, RF, CNN, and BPNN by leveraging unlabeled logs via CAE to learn expressive features, delivering strong performance under small-sample conditions.
- (2)
- Accurate TOC prediction and verified generalization provide a practical basis for shale-oil mobility evaluation. TOC-driven OSI assessment corroborates the low-TOC and high-mobility pattern in the Mahu Sag, further quantified by micro-migration analysis, which supports identification of favorable intervals for exploration and development.
- (3)
- CAE-BPNN has good scalability and can be applied to the prediction of geological sweet spot evaluation parameters such as S1, S2 and Tmax. Future research will focus on further study the evaluation and prediction model of shale oil geological sweet spots, and make the developed model have certain geological interpretation.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Nomenclature
| Abbreviation/Symbol | Full Name | Description |
| CAE | Convolutional autoencoder | An unsupervised neural network used for feature representation learning and input reconstruction |
| BPNN | Back-propagation neural network | A supervised neural network used for nonlinear regression prediction |
| CAE-BPNN | Convolutional autoencoder–back-propagation neural network | The hybrid deep learning framework proposed in this study for small-sample TOC prediction |
| RF | Random forest | An ensemble machine learning model based on multiple decision trees |
| GBDT | Gradient boosting decision tree | An ensemble learning model that improves prediction by sequentially fitting residual errors |
| CNN | Convolutional neural network | A deep learning model using convolutional operations for feature extraction |
| SVM | Support vector machine | A supervised machine learning method used for classification and regression |
| GPR | Gaussian process regression | A probabilistic regression method based on Gaussian processes |
| VAE | Variational autoencoder | A generative representation-learning model that learns probabilistic latent variables |
| LSTM | Long short-term memory | A recurrent neural network architecture designed for sequence modeling |
| RMSE | Root mean square error | A statistical metric used to measure prediction error |
| MSE | Mean squared error | A loss function commonly used in model training and reconstruction evaluation |
| MAE | Mean absolute error | A statistical metric used to measure the average absolute prediction error |
| AdamW | Adam optimizer with decoupled weight decay | An adaptive optimization algorithm with weight-decay regularization |
| BatchNorm | Batch normalization | A normalization operation used to stabilize neural network training |
| ReLU | Rectified linear unit | A nonlinear activation function commonly used in neural networks |
| ReLU | Leaky rectified linear unit | A modified ReLU activation function that allows a small gradient for negative inputs |
| Dropout | Dropout regularization | A regularization method used to reduce overfitting by randomly deactivating neurons during training |
| PCA | Principal component analysis | A dimensionality-reduction method used to extract major variance components |
| SHAP | Shapley additive explanations | An interpretability method used to evaluate feature contribution |
| VIF | Variance inflation factor | A statistical indicator used to assess multicollinearity among input variables |
References
- Sun, L.; Jia, C.; Zhang, J.; Cui, B.; Bai, J.; Huo, Q.; Xu, X.; Liu, W.; Zeng, H.; Liu, W. Resource potential of Gulong shale oil in the key areas of Songliao Basin. Acta Pet. Sin. 2024, 45, 1699–1714. [Google Scholar] [CrossRef]
- Hou, L.; Luo, X.; Lin, S.; Li, Y.; Zhang, L.; Ma, W. Assessment of recoverable oil and gas resources by in-situ conversion of shale-Case study of extracting the Chang 73 shale in the Ordos Basin. Pet. Sci. 2022, 19, 441–458. [Google Scholar] [CrossRef] [Scilit]
- Song, Y.; Song, Z.; Mo, Y.; Meng, Y.; Zhou, Q.; Jing, Y.; Tian, S.; Chen, Z. Maturity-dependent thermodynamic and flow characteristics in continental shale oils. Energy 2025, 318, 134867. [Google Scholar] [CrossRef] [Scilit]
- Wang, E.; Feng, Y.; Guo, T.; Li, M. Oil content and resource quality evaluation methods for lacustrine shale: A review and a novel three-dimensional quality evaluation model. Earth-Sci. Rev. 2022, 232, 104134. [Google Scholar] [CrossRef] [Scilit]
- Yu, M.; Gao, G.; Chen, S.; Li, J.; Liu, M.; Kang, J.; Xu, X.; Zhang, W. Hydrocarbon generation and expulsion differences of organic matter in saline lacustrine shale: Implications for shale oil exploration and development. Fuel 2026, 404, 136296. [Google Scholar] [CrossRef] [Scilit]
- Hu, T.; Liu, Y.; Jiang, F.; Pang, X.; Wang, Q.; Zhou, K.; Wu, G.; Jiang, Z.; Huang, L.; Jiang, S.; et al. A novel method for quantifying hydrocarbon micromigration in heterogeneous shale and the controlling mechanism. Energy 2024, 288, 129712. [Google Scholar] [CrossRef] [Scilit]
- Huang, W.; Hersi, O.S.; Lu, S.; Deng, S. Quantitative modelling of hydrocarbon expulsion and quality grading of tight oil lacustrine source rocks: Case study of Qingshankou 1 member, central depression, Southern Songliao Basin, China. Mar. Pet. Geol. 2017, 84, 34–48. [Google Scholar] [CrossRef] [Scilit]
- Romero-Sarmiento, M.F.; Ducros, M.; Carpentier, B.; Lorant, F.; Cacas, M.C.; Pegaz-Fiornet, S.; Wolf, S.; Rohais, S.; Moretti, I. Quantitative evaluation of TOC, organic porosity and gas retention distribution in a gas shale play using petroleum system modeling: Application to the Mississippian Barnett Shale. Mar. Pet. Geol. 2013, 45, 315–330. [Google Scholar] [CrossRef] [Scilit]
- Shan, X.; Chen, Z.; Fu, B.; Zhang, W.; Li, J.; Wu, K. Predicting total organic carbon from well logs based on deep spatial-sequential graph convolutional network. Geophysics 2023, 88, D193–D206. [Google Scholar] [CrossRef] [Scilit]
- Zhang, W.; Shan, X.; Fu, B.; Zou, X.; Fu, L.Y. A deep encoder-decoder neural network model for total organic carbon content prediction from well logs. J. Asian Earth Sci. 2022, 240, 105437. [Google Scholar] [CrossRef] [Scilit]
- Zheng, D.; Wu, S.; Hou, M. Fully connected deep network: An improved method to predict TOC of shale reservoirs from well logs. Mar. Pet. Geol. 2021, 132, 105205. [Google Scholar] [CrossRef] [Scilit]
- Schmoker, J.W. Determination of organic-matter content of Appalachian Devonian shales from gamma-ray logs. AAPG Bull. 1981, 65, 1285–1298. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Zhang, Y.; Li, J.; Hui, G.; Sun, Y.; Li, Y.; Chen, Y.; Zhang, D. Artificial intelligence large model for logging curve reconstruction. Pet. Explor. Dev. 2025, 52, 842–854. [Google Scholar] [CrossRef] [Scilit]
- Hui, G.; Chen, Z.; Yan, J.; Wang, M.; Wang, H.; Zhang, D.; Gu, F. Integrated evaluations of high-quality shale play using core experiments and logging interpretations. Fuel 2023, 341, 127679. [Google Scholar] [CrossRef] [Scilit]
- Lai, J.; Zhao, F.; Xia, Z.; Su, Y.; Zhang, C.; Tian, Y.; Wang, G.; Qin, Z. Well log prediction of total organic carbon: A comprehensive review. Earth-Sci. Rev. 2024, 258, 104913. [Google Scholar] [CrossRef] [Scilit]
- Passey, Q.R.; Creaney, S.; Kulla, J.B.; Moretti, F.J.; Stroud, J.D. A practical model for organic richness from porosity and resistivity logs. AAPG Bull. 1990, 74, 1777–1794. [Google Scholar] [CrossRef] [Scilit]
- Passey, Q.R.; Bohacs, K.M.; Esch, W.L.; Klimentidis, R.; Sinha, S. From oil-prone source rock to gas-producing shale reservoir–geologic and petrophysical characterization of unconventional shale-gas reservoirs. In SPE International Oil and Gas Conference and Exhibition in China; SPE: Beijing, China, 2010; p. SPE-131350-MS. [Google Scholar]
- Zhao, P.; Ma, H.; Rasouli, V.; Liu, W.; Cai, J.; Huang, Z. An improved model for estimating the TOC in shale formations. Mar. Pet. Geol. 2017, 83, 174–183. [Google Scholar] [CrossRef] [Scilit]
- Zhu, L.; Zhang, C.; Zhang, Z.; Zhou, X.; Liu, W. An improved method for evaluating the TOC content of a shale formation using the dual-difference ΔlogR method. Mar. Pet. Geol. 2019, 102, 800–816. [Google Scholar] [CrossRef] [Scilit]
- Zhou, C.; Wang, L.; Su, S.; Xue, K.; Wang, Q. The logging evaluation of organic carbon content based on ΔlogR-GR method: Case study of the first member of Maokou Formation in the southeastern Sichuan Basin. Nat. Gas Geosci. 2024, 35, 542–552. [Google Scholar] [CrossRef]
- Bolandi, V.; Kadkhodaie, A.; Farzi, R. Analyzing organic richness of source rocks from well log data by using SVM and ANN classifiers: A case study from the Kazhdumi formation, the Persian Gulf basin, offshore Iran. J. Pet. Sci. Eng. 2017, 151, 224–234. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Zhao, X.; Zhu, H.; Tang, Z.; Zhao, X.; Zhang, F.; Sepehrnoori, K. Engineering factor analysis and intelligent prediction of CO2 storage parameters in shale gas reservoirs based on deep learning. Appl. Energy 2025, 377, 124642. [Google Scholar] [CrossRef] [Scilit]
- Gul, S.; Eric, V.O. A machine learning approach to filtrate loss determination and test automation for drilling and completion fluids. J. Pet. Sci. Eng. 2020, 186, 106727. [Google Scholar] [CrossRef] [Scilit]
- Li, W.; Yue, D.; Wang, W.; Wang, W.; Wu, S.; Li, J.; Chen, D. Fusing multiple frequency-decomposed seismic attributes with machine learning for thickness prediction and sedimentary facies interpretation in fluvial reservoirs. J. Pet. Sci. Eng. 2019, 177, 1087–1102. [Google Scholar] [CrossRef] [Scilit]
- Wang, B.; Han, G.; Lu, X.; Ma, H.; Zhu, Z.; Liang, X. The integrated geosciences and engineering production prediction in tight reservoir based on deep learning. Geoenergy Sci. Eng. 2023, 223, 211571. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Liu, B.; Gao, Y.; Li, C.; Tang, R.; Kong, Y.; Xie, M.; Li, K.; Dan, S.; Qi, K.; et al. Mineral prospecting mapping with conditional generative adversarial network augmented data. Ore Geol. Rev. 2023, 163, 105787. [Google Scholar] [CrossRef] [Scilit]
- Kadri, M.M.; Ganguli, S.S.; Sen, S.; Hacini, M.; Kumar, P. Characterization and Feature Ranking of Well Log Variables Using Data-Driven Algorithms for Total Organic Carbon Estimation of Organic-Rich Shales. Energy Fuels 2023, 37, 19575–19589. [Google Scholar] [CrossRef] [Scilit]
- Macêdo, B.S.; Wayo, D.D.K.; Campos, D.; De Santis, R.B.; Martinho, A.D.; Yaseen, Z.M.; Saporetti, C.M.; Goliatt, L. Data-driven total organic carbon prediction using feature selection methods incorporated in an automated machine learning framework. Sci. Rep. 2025, 15, 10658. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Handhal, A.M.; Al-Abadi, A.M.; Chafeet, H.A.; Ismail, M. Prediction of total organic carbon at Rumaila oil field, Southern Iraq using conventional well logs and machine learning algorithms. Mar. Pet. Geol. 2020, 116, 104347. [Google Scholar] [CrossRef] [Scilit]
- Rui, J.; Zhang, H.; Ren, Q.; Yan, L.; Guo, Q.; Zhang, D. TOC content prediction based on a combined Gaussian process regression model. Mar. Pet. Geol. 2020, 118, 104429. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Wu, W.; Chen, T.; Dong, X.; Wang, G. An improved neural network for TOC, S1 and S2 estimation based on conventional well logs. J. Pet. Sci. Eng. 2019, 176, 664–678. [Google Scholar] [CrossRef] [Scilit]
- Yuan, Y.; Tan, D.; Yu, S.; Li, Y.; Han, B. A Prediction Model for Shale Gas Organic Carbon Content Based on Improved BP Neural Network Using Bayesian Regularization. Geol. Explor. 2019, 55, 1082–1091. [Google Scholar]
- Ahangari, D.; Daneshfar, R.; Zakeri, M.; Ashoori, S.; Soulgani, B.S. On the prediction of geochemical parameters (TOC, S1 and S2) by considering well log parameters using ANFIS and LSSVM strategies. Petroleum 2022, 8, 174–184. [Google Scholar] [CrossRef] [Scilit]
- Wang, P.; Peng, S.; Du, W.; Feng, F. Prediction model of total organic carbon content on hydrocarbon source rocks in coal measures based on geophysical well logging. J. China Coal Soc. 2017, 42, 1266–1276. [Google Scholar] [CrossRef]
- Liu, B.; Ma, Y.; Yasin, Q.; Wood, D.A.; Sun, M.; Gao, S.; Bai, Y. Characterization of lacustrine shale oil reservoirs based on a hybrid deep learning model: A data-driven approach to predict lithofacies, vitrinite reflectance, and TOC. Mar. Pet. Geol. 2025, 174, 107309. [Google Scholar] [CrossRef] [Scilit]
- Cheng, B.; Xu, T.; Luo, S.; Chen, T.; Li, Y.; Tang, J. Method and practice of deep favorable shale reservoirs prediction based on machine learning. Pet. Explor. Dev. 2022, 49, 1056–1068. [Google Scholar] [CrossRef] [Scilit]
- Heaton, J. An Empirical Analysis of Feature Engineering for Predictive Modeling. In SoutheastCon; IEEE: Norfolk, VA, USA, 2016. [Google Scholar]
- Khurana, U.; Samulowitz, H.; Turaga, D. Feature Engineering for Predictive Modeling Using Reinforcement Learning. In 32nd AAAI Conference on Artificial Intelligence/30th Innovative Applications of Artificial Intelligence Conference/8th AAAI Symposium on Educational Advances in Artificial Intelligence; AAAI Press: New Orleans, LA, USA, 2018; pp. 3407–3414. [Google Scholar] [CrossRef] [Scilit]
- Demir-Kavuk, O.; Kamada, M.; Akutsu, T.; Knapp, E.W. Prediction using step-wise L1, L2 regularization and feature selection for small data sets with large number of features. BMC Bioinform. 2011, 12, 412. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, L.; Zhang, C.; Zhang, C.; Zhang, Z.; Nie, X.; Zhou, X.; Liu, W.; Wang, X. Forming a new small sample deep learning model to predict total organic carbon content by combining unsupervised learning with semisupervised learning. Appl. Soft Comput. 2019, 83, 105596. [Google Scholar] [CrossRef] [Scilit]
- Masci, J.; Meier, U.; Cireşan, D.; Schmidhuber, J. Stacked Convolutional Auto-Encoders for Hierarchical Feature Extraction. In 21st International Conference on Artificial Neural Networks, ICANN 2011; Springer: Berlin/Heidelberg, Germany, 2011; pp. 52–59. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Chang, J.; Jiang, Z.; Gao, Z.; Zhang, C.; Wang, G.; Shao, X.; He, W. Visualization of dynamic micro-migration of shale oil and investigation of shale oil movability by NMRI combined oil charging/water flooding experiments: A novel approach. Mar. Pet. Geol. 2024, 165, 106907. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Li, W.; Tang, W. Tectonic Setting and Environment of Alkaline Lacustrine Source Rocks in the Lower Permian Fengcheng Formation of Mahu Sag. Xinjiang Pet. Geol. 2018, 39, 48–54. [Google Scholar]
- Hinton, G.E.; Salakhutdinov, R.R. Reducing the dimensionality of data with neural networks. Science 2006, 313, 504–507. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning representations by back-propagating errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Wang, G.; Zeng, L.; Liu, P.; Huang, Y.; Li, S.; Wang, Z.; Zhou, Y. New method for logging identification of natural fractures in shale reservoirs: The Fengcheng formation of the Mahu Sag, China. Mar. Pet. Geol. 2025, 176, 107346. [Google Scholar] [CrossRef] [Scilit]
- Hu, T.; Jiang, F.; Pang, X.; Liu, Y.; Wu, G.; Zhou, K.; Xiao, H.; Jiang, Z.; Li, M.; Jiang, S.; et al. Identification and evaluation of shale oil micro-migration and its petroleum geological significance. Pet. Explor. Dev. 2024, 51, 127–140. [Google Scholar] [CrossRef] [Scilit]
- Zhi, D.; Cao, J.; Xiang, B.; Qin, Z.; Wang, T. Fengcheng Alkaline Lacustrine Source Rocks of Lower Permian in Mahu Sag in Junggar Basin: Hydrocarbon Generation Mechanism and Petroleum Resources Reestimation. Xinjiang Pet. Geol. 2016, 37, 499–506. [Google Scholar]
- Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30; NeurIPS Proceedings: Long Beach, CA, USA, 2017; pp. 4765–4774. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2016; pp. 770–778. [Google Scholar]
- Teerapittayanon, S.; McDanel, B.; Kung, H.T. BranchyNet: Fast inference via early exiting from deep neural networks. In 2016 23rd International Conference on Pattern Recognition (ICPR); IEEE: New York, NY, USA, 2016; pp. 2464–2469. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Jiang, F.; Hu, T.; Xu, Y.; Guo, J.; Xu, T.; Xing, H.; Chen, D.; Pang, H.; Chen, J.; et al. Shale oil content evaluation and sweet spot prediction based on convolutional neural network. Mar. Pet. Geol. 2024, 167, 106997. [Google Scholar] [CrossRef] [Scilit]
- Bui, Q.-A.T.; Nguyen, D.D.; Le, H.V.; Prakash, I.; Pham, B.T. Prediction of Shear Bond Strength of Asphalt Concrete Pavement Using Machine Learning Models and Grid Search Optimization Technique. CMES-Comput. Model. Eng. Sci. 2025, 142, 691–712. [Google Scholar] [CrossRef] [Scilit]
- Demir, H.G.; Yesilyurt, I. A comparison of four machine learning techniques and continuous wavelet transform approach for detection and classification of tool breakage during milling process. Trans. Can. Soc. Mech. Eng. 2022, 47, 26–42. [Google Scholar]
- Jia, W.; Zong, Z.; Qin, D.; Lan, T. A method for predicting the TOC in source rocks using a machine learning-based joint analysis of seismic multi-attributes. J. Appl. Geophys. 2023, 216, 105143. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Wu, W.; Wu, H. TOC prediction using a gradient boosting decision tree method: A case study of shale reservoirs in Qinshui Basin. Geoenergy Sci. Eng. 2023, 221, 111271. [Google Scholar] [CrossRef] [Scilit]
- Xiao, H.; Hu, T.; Pang, X.; Xu, Y.; Hu, Y.; Li, C.; Xu, T.; Zheng, D.; Pu, T.; Ding, C.; et al. A new method for identification of effective hydrocarbon source rocks and evaluation of relative contributions to reservoirs. Mar. Pet. Geol. 2025, 178, 107424. [Google Scholar] [CrossRef] [Scilit]
- Jarvie, D.M. Shale Resource Systems for Oil and Gas Part 2; Shale-oil Resource Systems. In Shale Reservoirs—Giant Resources for the 21st Century; American Association of Petroleum Geologists: Tulsa, OK, USA, 2012. [Google Scholar]
- Efron, B.; Tibshirani, R.J. An Introduction to the Bootstrap; Chapman and Hall/CRC: New York, NY, USA, 1993. [Google Scholar]
- Kurokawa, H.; Mori, S. A local connected neural oscillator network for pattern segmentation. In Artificial Neural Networks—ICANN 96; von der Malsburg, C., von Seelen, W., Vorbrüggen, J.C., Sendhoff, B., Eds.; Springer: Berlin/Heidelberg, Germany, 1996; pp. 797–802. [Google Scholar]
- Kingma, D.P.; Welling, M. Auto-encoding variational Bayes. In International Conference on Learning Representations; University of Amsterdam: Amsterdam, The Netherlands, 2014. [Google Scholar]
- Lopez-Alvis, J.; Laloy, E.; Nguyen, F.; Hermans, T. Geophysical inversion using a variational autoencoder to model an assembled spatial prior uncertainty. J. Geophys. Res. Solid Earth 2022, 127, e2021JB022581. [Google Scholar] [CrossRef] [Scilit]
- Rodriguez, O.; Taylor, J.M.; Pardo, D. Multimodal variational autoencoder for inverse problems in geophysics: Application to a 1-D magnetotelluric problem. Geophys. J. Int. 2023, 235, 2598–2613. [Google Scholar] [CrossRef] [Scilit]














| n_estimators | max_depth | min_samples_split | min_samples_leaf | max_features | Subsample | |
|---|---|---|---|---|---|---|
| RF | 200 | 15 | 5 | 2 | ‘sqrt’ | / |
| GBDT | 300 | 5 | 10 | 4 | / | 0.8 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yuan, B.; Zhang, B.; Zhang, Y.; Zhang, J.; Hu, T.; Qiu, N.; Wang, Z.; Ma, M.; Xiong, Z.; Wang, M.; et al. A Hybrid Deep Learning Approach for Small-Sample TOC Prediction in Saline Lacustrine Shale. Energies 2026, 19, 3360. https://doi.org/10.3390/en19143360
Yuan B, Zhang B, Zhang Y, Zhang J, Hu T, Qiu N, Wang Z, Ma M, Xiong Z, Wang M, et al. A Hybrid Deep Learning Approach for Small-Sample TOC Prediction in Saline Lacustrine Shale. Energies. 2026; 19(14):3360. https://doi.org/10.3390/en19143360
Chicago/Turabian StyleYuan, Bo, Bolin Zhang, Yuanhao Zhang, Jun Zhang, Tao Hu, Nansheng Qiu, Zigen Wang, Mingming Ma, Zhiming Xiong, Miao Wang, and et al. 2026. "A Hybrid Deep Learning Approach for Small-Sample TOC Prediction in Saline Lacustrine Shale" Energies 19, no. 14: 3360. https://doi.org/10.3390/en19143360
APA StyleYuan, B., Zhang, B., Zhang, Y., Zhang, J., Hu, T., Qiu, N., Wang, Z., Ma, M., Xiong, Z., Wang, M., Jiang, Z., Li, M., & Pang, X. (2026). A Hybrid Deep Learning Approach for Small-Sample TOC Prediction in Saline Lacustrine Shale. Energies, 19(14), 3360. https://doi.org/10.3390/en19143360

