Next Article in Journal
A New Depth-Based Test for Multivariate Two-Sample Problems
Previous Article in Journal
On the Classification–Causal Tradeoff in Neural Network Propensity Score Estimation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Multiple Imputation of a Continuous Outcome with Fully Observed Predictors Using TabPFN

Center for Health, Policy and Economics, Faculty of Health Sciences and Medicine, University of Lucerne, 6002 Luzern, Switzerland
Stats 2026, 9(2), 38; https://doi.org/10.3390/stats9020038
Submission received: 24 February 2026 / Revised: 24 March 2026 / Accepted: 29 March 2026 / Published: 1 April 2026
(This article belongs to the Special Issue Statistical Methods for Hypothesis Testing)

Abstract

Handling missing data is a central challenge in quantitative research, particularly when datasets exhibit complex dependency structures, such as nonlinear relationships and interactions. Multiple imputation (MI) via fully conditional specification (FCS), as implemented in the MICE R package, is widely used but relies on user-specified models that may fail to capture complex dependency structures, especially in high-dimensional settings, or on more sophisticated algorithms that are considered data-hungry. This paper investigates the performance of TabPFN, a transformer-based, pretrained foundation model developed for tabular prediction tasks, for MI. TabPFN is pretrained on millions of synthetic datasets and approximates posterior predictive distributions without dataset-specific retraining, offering a compelling solution for imputing complex missing data in small to moderately sized samples. We conduct a simulation study focusing on univariate missingness in a continuous outcome with complete predictors, comparing TabPFN with standard MI methods. Performance is evaluated using bias, standard error, and coverage of the marginal mean estimand across a range of data-generating and missingness mechanisms. Our results show that TabPFN yields competitive or superior performance relative to Classification and Regression Trees and Predictive Mean Matching. These findings highlight TabPFN as a promising tool for missing data imputation, with particular relevance to health research.
Keywords: missing data; multiple imputation; TabPFN; pretrained models; in-context learning; simulation study; complex dependency structures missing data; multiple imputation; TabPFN; pretrained models; in-context learning; simulation study; complex dependency structures

Share and Cite

MDPI and ACS Style

Sepin, J. Multiple Imputation of a Continuous Outcome with Fully Observed Predictors Using TabPFN. Stats 2026, 9, 38. https://doi.org/10.3390/stats9020038

AMA Style

Sepin J. Multiple Imputation of a Continuous Outcome with Fully Observed Predictors Using TabPFN. Stats. 2026; 9(2):38. https://doi.org/10.3390/stats9020038

Chicago/Turabian Style

Sepin, Jerome. 2026. "Multiple Imputation of a Continuous Outcome with Fully Observed Predictors Using TabPFN" Stats 9, no. 2: 38. https://doi.org/10.3390/stats9020038

APA Style

Sepin, J. (2026). Multiple Imputation of a Continuous Outcome with Fully Observed Predictors Using TabPFN. Stats, 9(2), 38. https://doi.org/10.3390/stats9020038

Article Metrics

Back to TopTop