On the Classification–Causal Tradeoff in Neural Network Propensity Score Estimation
Abstract
1. Introduction
1.1. Causal Inference in Observational Research
1.2. Propensity Score Methods
1.3. Traditional and Emerging Estimation Architectures
1.4. Purpose of the Study
2. Materials and Methods
2.1. Population Model and Data Generation
2.2. Theoretical Framework: Classification–Causal Tradeoff
2.3. Stability-Aware Training and Hyperparameters
2.4. Model Architectures
2.5. Propensity Score Weighting and Analysis
3. Results
3.1. Accuracy and Loss Rate (For DNN and CNN)
3.2. Bias Reduction
3.3. Imbalance Reduction
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Johnson, R.B.; Christensen, L.B. Educational Research: Quantitative, Qualitative, and Mixed Approaches, 6th ed.; SAGE Publications: Thousand Oaks, CA, USA, 2016. [Google Scholar]
- Guo, S.; Fraser, M.W. Propensity Score Analysis: Statistical Methods and Applications, 2nd ed.; SAGE Publications: Thousand Oaks, CA, USA, 2014. [Google Scholar]
- Austin, P.C.; Stuart, E.A. Moving towards best practice when using inverse probability of treatment weighting (IPTW) using the propensity score to estimate causal treatment effects in observational studies. Stat. Med. 2015, 34, 3661–3679. [Google Scholar] [CrossRef] [Scilit]
- Austin, P.C. An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivar. Behav. Res. 2011, 46, 399–424. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Austin, P.C. A comparison of 12 algorithms for matching on the propensity score. Stat. Med. 2013, 33, 1057–1069. [Google Scholar] [CrossRef] [Scilit]
- Thoemmes, F.J.; Kim, E.S. A systematic review of propensity score methods in the social sciences. Multivar. Behav. Res. 2011, 46, 90–118. [Google Scholar] [CrossRef] [Scilit]
- Rosenbaum, P.R.; Rubin, D.B. The central role of the propensity score in observational studies for causal effects. Biometrika 1983, 70, 41–55. [Google Scholar] [CrossRef]
- Lee, J.; Little, T.D. A practical guide to propensity score analysis for applied clinical research. Behav. Res. Ther. 2017, 98, 76–90. [Google Scholar] [CrossRef] [Scilit]
- Rosenbaum, P.R. Observational Studies; Springer: Berlin/Heidelberg, Germany, 1995. [Google Scholar]
- Breiman, L. Statistical Modeling: The Two Cultures. Stat. Sci. 2001, 16, 199–215. [Google Scholar] [CrossRef] [Scilit]
- Westreich, D.; Lessler, J.; Funk, M.J. Propensity score estimation: Neural networks, support vector machines, decision trees (CART), and meta-classifiers as alternatives to logistic regression. J. Clin. Epidemiol. 2010, 63, 826–833. [Google Scholar] [CrossRef] [Scilit]
- McCaffrey, D.F.; Ridgeway, G.; Morral, A.R. Propensity score estimation with boosted regression for evaluating causal effects in observational studies. Psychol. Methods 2004, 9, 403–425. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zuur, A.F.; Ieno, E.N.; Walker, N.; Saveliev, A.A.; Smith, G. Mixed Effects Models and Extensions in Ecology with R; Springer: Berlin/Heidelberg, Germany, 2009. [Google Scholar] [CrossRef] [Scilit]
- Brand, J.E.; Zhou, X.; Xie, Y. Recent Developments in Causal Inference and Machine Learning. Annu. Rev. Sociol. 2023, 49, 81–110. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Guo, Y.; Strauss, V.Y.; Català, M.; Jödicke, A.M.; Khalid, S.; Prieto-Alhambra, D. Machine learning methods for propensity and disease risk score estimation in high-dimensional data: A plasmode simulation and real-world data cohort analysis. Front. Pharmacol. 2024, 15, 1395707. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Collier, Z.A.; Leite, W.L.; Zhang, H. Estimating propensity scores using neural networks and traditional methods: A comparative simulation study. Commun. Stat. Simul. Comput. 2022, 51, 5780–5795. [Google Scholar] [CrossRef] [Scilit]
- Borisov, V.; Leemann, T.; Seßler, K.; Haug, J.; Pawelczyk, M.; Kasneci, G. Deep neural networks and tabular data: A survey. IEEE Trans. Neural Netw. Learn. Syst. 2022, 35, 7499–7519. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Y.; Brettin, T.; Xia, F.; Partin, A.; Shukla, M.; Yoo, H.; Evrard, Y.A.; Doroshow, J.H.; Stevens, R. Converting tabular data into images for deep learning with convolutional neural networks. Sci. Rep. 2021, 11, 11325. [Google Scholar] [CrossRef] [Scilit]
- Sharma, A.; Vans, E.; Shigemizu, D.; Boroevich, K.A.; Tsunoda, T. DeepInsight: A methodology to transform non-image data to an image for convolution neural network architecture. Sci. Rep. 2019, 9, 11399. [Google Scholar] [CrossRef] [Scilit]
- Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press: Cambridge, MA, USA, 2016; Available online: https://www.deeplearningbook.org/ (accessed on 13 December 2018).
- Farrell, M.H.; Liang, T.; Misra, S. Deep neural networks for estimation and inference. Econometrica 2021, 89, 181–213. [Google Scholar] [CrossRef] [Scilit]
- Lee, B.R.; Lessler, J.; Stuart, E.A. Improving propensity score weighting using machine learning. Stat. Med. 2010, 29, 337–346. [Google Scholar] [CrossRef] [Scilit]
- D’Amour, A.; Ding, P.; Feller, A.; Lei, L.; Sekhon, J. Overlap in observational studies with high-dimensional covariates. J. Econom. 2021, 221, 644–654. [Google Scholar] [CrossRef] [Scilit]
- Couronné, R.; Probst, P.; Boulesteix, A. Random forest versus logistic regression: A large-scale benchmark experiment. BMC Bioinform. 2018, 19, 270. [Google Scholar] [CrossRef] [Scilit]
- Cohen, J. Statistical Power Analysis for the Behavioral Sciences, 2nd ed.; Routledge: New York, NY, USA, 1977. [Google Scholar] [CrossRef] [Scilit]
- R Core Team. R: A Language and Environment for Statistical Computing; R Foundation for Statistical Computing: Vienna, Austria, 2021; Available online: https://www.R-project.org/ (accessed on 13 April 2023).
- RStudio Team. RStudio: Integrated Development for R; RStudio (PBC): Boston, MA, USA, 2020; Available online: http://www.rstudio.com/ (accessed on 28 July 2022).
- Heaton, J. Introduction to Neural Networks with Java, 2nd ed.; Heaton Research: St. Louis, MO, USA, 2008. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 7–13 December 2015; pp. 1026–1034. [Google Scholar] [CrossRef] [Scilit]
- Hirano, K.; Imbens, G.W.; Ridder, G. Efficient estimation of average treatment effects using the estimated propensity score. Econometrica 2003, 71, 1161–1189. [Google Scholar] [CrossRef] [Scilit]
- Robins, J.M.; Hernán, M.N.; Brumback, B. Marginal structural models and causal inference in epidemiology. Epidemiology 2000, 11, 550–560. [Google Scholar] [CrossRef] [Scilit]
- Levy, J.M.; O’Malley, A.J. Don’t dismiss logistic regression: The case for sensible extraction of interactions in the era of machine learning. BMC Med. Res. Methodol. 2020, 20, 171. [Google Scholar] [CrossRef] [Scilit]
- Hill, J. Bayesian nonparametric modeling for causal inference. J. Comput. Graph. Stat. 2011, 20, 217–240. [Google Scholar] [CrossRef] [Scilit]
- Imai, K.; Ratkovic, M. Covariate balancing propensity score. J. R. Stat. Soc. Ser. B 2014, 76, 243–263. [Google Scholar] [CrossRef] [Scilit]
- Wager, S.; Athey, S. Estimation and inference of heterogeneous treatment effects using random forests. J. Am. Stat. Assoc. 2018, 113, 1228–1242. [Google Scholar] [CrossRef] [Scilit]





| Subject ID | TG | Y | T |
|---|---|---|---|
| 5 | 2.50 | 1.55 | 1 |
| 21 | 2.17 | 3.66 | 1 |
| 6 | 1.80 | 1.87 | 1 |
| ⋮ | ⋮ | ⋮ | ⋮ |
| 52 | –1.83 | –1.91 | 0 |
| 40 | –1.86 | –0.44 | 0 |
| 65 | –2.00 | –1.20 | 0 |
| Batch Size | rTY | DNN | CNN | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Training | Validation | Training | Validation | ||||||
| Acc. | Loss | Acc. | Loss | Acc. | Loss | Acc. | Loss | ||
| 100 | 0.1 | 0.863 | 0.346 | 0.863 | 0.346 | 0.862 | 0.347 | 0.859 | 0.351 |
| 0.3 | 0.868 | 0.337 | 0.868 | 0.338 | 0.867 | 0.339 | 0.864 | 0.343 | |
| 0.5 | 0.874 | 0.333 | 0.874 | 0.332 | 0.873 | 0.334 | 0.870 | 0.338 | |
| 1000 | 0.1 | 0.858 | 0.360 | 0.858 | 0.360 | 0.852 | 0.363 | 0.852 | 0.363 |
| 0.3 | 0.862 | 0.354 | 0.861 | 0.354 | 0.855 | 0.357 | 0.853 | 0.358 | |
| 0.5 | 0.868 | 0.350 | 0.868 | 0.349 | 0.860 | 0.353 | 0.859 | 0.355 | |
| 2000 | 0.1 | 0.855 | 0.370 | 0.854 | 0.370 | 0.844 | 0.374 | 0.845 | 0.373 |
| 0.3 | 0.858 | 0.365 | 0.857 | 0.365 | 0.847 | 0.369 | 0.848 | 0.368 | |
| 0.5 | 0.864 | 0.361 | 0.864 | 0.361 | 0.851 | 0.366 | 0.852 | 0.366 | |
| CNN (C) | DNN (D) | LR (L) | Overall Difference | Pairwise Comparison | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| M | SE | M | SE | M | SE | F | p | Cohen’s f | C vs. D | C vs. L | D vs. L | |
| Overall | 0.16 | 0.01 | 0.13 | 0.01 | –0.17 | 0.07 | 47.67 | <0.001 | 0.50 | 0.363 | <0.001 | <0.001 |
| rXX | ||||||||||||
| 0.1 | 0.14 | 0.01 | 0.10 | 0.02 | –0.34 | 0.12 | 38.08 | <0.001 | 0.79 | 0.503 | <0.001 | <0.001 |
| 0.3 | 0.14 | 0.01 | 0.12 | 0.02 | –0.19 | 0.17 | 9.94 | <0.001 | 0.40 | 0.884 | <0.001 | <0.001 |
| 0.5 | 0.21 | 0.00 | 0.18 | 0.01 | 0.01 | 0.08 | 15.08 | <0.001 | 0.50 | 0.431 | <0.001 | <0.001 |
| rTY | ||||||||||||
| 0.1 | 0.06 | 0.01 | 0.03 | 0.02 | –0.75 | 0.12 | 106.00 | <0.001 | 1.31 | 0.680 | <0.001 | <0.001 |
| 0.3 | 0.21 | 0.00 | 0.17 | 0.01 | –0.03 | 0.05 | 42.69 | <0.001 | 0.83 | 0.059 | <0.001 | <0.001 |
| 0.5 | 0.22 | 0.00 | 0.19 | 0.01 | 0.25 | 0.05 | 3.15 | 0.046 | 0.23 | 0.316 | 0.346 | 0.044 |
| CNN (C) | DNN (D) | LR (L) | Overall Difference | Pairwise Comparison | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| M | SE | M | SE | M | SE | F | p | Cohen’s f | C vs. D | C vs. L | D vs. L | |
| Overall | 0.21 | 0.00 | 0.18 | 0.00 | 0.05 | 0.01 | 714.60 | <0.001 | 0.46 | <0.001 | <0.001 | <0.001 |
| rTY | ||||||||||||
| 0.1 | 0.21 | 0.00 | 0.18 | 0.00 | 0.09 | 0.02 | 113.30 | <0.001 | 0.32 | <0.001 | <0.001 | <0.001 |
| 0.3 | 0.21 | 0.00 | 0.17 | 0.00 | 0.01 | 0.01 | 342.20 | <0.001 | 0.55 | <0.001 | <0.001 | <0.001 |
| 0.5 | 0.21 | 0.00 | 0.17 | 0.00 | 0.03 | 0.01 | 318.50 | <0.001 | 0.53 | <0.001 | <0.001 | <0.001 |
| rXX | ||||||||||||
| 0.1 | 0.21 | 0.00 | 0.17 | 0.00 | 0.06 | 0.02 | 136.30 | <0.001 | 0.35 | <0.001 | <0.001 | <0.001 |
| 0.3 | 0.21 | 0.00 | 0.18 | 0.00 | 0.07 | 0.01 | 182.80 | <0.001 | 0.40 | <0.001 | <0.001 | <0.001 |
| 0.5 | 0.22 | 0.00 | 0.18 | 0.00 | 0.00 | 0.01 | 543.80 | <0.001 | 0.69 | <0.001 | <0.001 | <0.001 |
| Covariate | ||||||||||||
| X1 | 0.18 | 0.00 | 0.12 | 0.01 | –0.17 | 0.05 | 100.20 | <0.001 | 0.73 | 0.004 | <0.001 | <0.001 |
| X2 | 0.21 | 0.00 | 0.17 | 0.01 | 0.06 | 0.03 | 45.02 | <0.001 | 0.49 | <0.001 | <0.001 | <0.001 |
| X3 | 0.22 | 0.00 | 0.18 | 0.01 | 0.10 | 0.03 | 38.25 | <0.001 | 0.45 | <0.001 | <0.001 | <0.001 |
| X4 | 0.23 | 0.00 | 0.19 | 0.01 | 0.11 | 0.03 | 39.66 | <0.001 | 0.46 | <0.001 | <0.001 | <0.001 |
| X5 | 0.23 | 0.00 | 0.16 | 0.00 | 0.11 | 0.03 | 38.98 | <0.001 | 0.46 | <0.001 | <0.001 | <0.001 |
| X6 | 0.23 | 0.00 | 0.20 | 0.00 | 0.12 | 0.03 | 41.71 | <0.001 | 0.47 | <0.001 | <0.001 | <0.001 |
| X7 | 0.15 | 0.00 | 0.12 | 0.01 | –0.21 | 0.05 | 100.40 | <0.001 | 0.73 | <0.001 | <0.001 | <0.001 |
| X8 | 0.21 | 0.00 | 0.17 | 0.01 | 0.06 | 0.03 | 49.60 | <0.001 | 0.51 | <0.001 | <0.001 | <0.001 |
| X9 | 0.22 | 0.00 | 0.18 | 0.01 | 0.11 | 0.03 | 38.65 | <0.001 | 0.45 | <0.001 | <0.001 | <0.001 |
| X10 | 0.22 | 0.00 | 0.19 | 0.01 | 0.11 | 0.02 | 36.09 | <0.001 | 0.44 | <0.001 | <0.001 | <0.001 |
| X11 | 0.23 | 0.00 | 0.19 | 0.00 | 0.11 | 0.03 | 39.06 | <0.001 | 0.46 | 0.001 | <0.001 | <0.001 |
| X12 | 0.23 | 0.00 | 0.19 | 0.00 | 0.12 | 0.03 | 42.05 | <0.001 | 0.47 | <0.001 | <0.001 | <0.001 |
| X13 | 0.13 | 0.01 | 0.11 | 0.01 | –0.26 | 0.06 | 83.76 | <0.001 | 0.67 | 0.45 | <0.001 | <0.001 |
| X14 | 0.21 | 0.00 | 0.17 | 0.01 | 0.04 | 0.03 | 58.99 | <0.001 | 0.56 | <0.001 | <0.001 | <0.001 |
| X15 | 0.22 | 0.00 | 0.19 | 0.01 | 0.09 | 0.03 | 47.59 | <0.001 | 0.50 | <0.001 | <0.001 | <0.001 |
| X16 | 0.23 | 0.00 | 0.19 | 0.01 | 0.11 | 0.03 | 44.09 | <0.001 | 0.48 | <0.001 | <0.001 | <0.001 |
| X17 | 0.23 | 0.00 | 0.19 | 0.00 | 0.11 | 0.03 | 40.54 | <0.001 | 0.47 | <0.001 | <0.001 | <0.001 |
| X18 | 0.23 | 0.00 | 0.20 | 0.00 | 0.12 | 0.03 | 42.38 | <0.001 | 0.48 | <0.001 | <0.001 | <0.001 |
| LR | DNN | CNN | |
|---|---|---|---|
| Bias reduction in outcome | 17% increased | 13% decreased | 16% decreased |
| Imbalance reduction in covariates | 5% decreased | 18% decreased | 21% decreased |
| Time (in minutes) | 6.5 | 2.1 | 31.8 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kim, S.; Lee, J.; Jung, K. On the Classification–Causal Tradeoff in Neural Network Propensity Score Estimation. Stats 2026, 9, 37. https://doi.org/10.3390/stats9020037
Kim S, Lee J, Jung K. On the Classification–Causal Tradeoff in Neural Network Propensity Score Estimation. Stats. 2026; 9(2):37. https://doi.org/10.3390/stats9020037
Chicago/Turabian StyleKim, Seungman, Jaehoon Lee, and Kwanghee Jung. 2026. "On the Classification–Causal Tradeoff in Neural Network Propensity Score Estimation" Stats 9, no. 2: 37. https://doi.org/10.3390/stats9020037
APA StyleKim, S., Lee, J., & Jung, K. (2026). On the Classification–Causal Tradeoff in Neural Network Propensity Score Estimation. Stats, 9(2), 37. https://doi.org/10.3390/stats9020037

