Improved Doubly Robust Inference with Nonprobability Survey Samples Using Finite Mixture Models: Application to Health Monitoring SMS Survey Data
Abstract
1. Introduction
2. Methods
2.1. Notations and Basic Settings
2.2. Overview of the Proposed NHDR Method
2.3. Stage 1: Identification of Heterogeneous Populations Using Finite Mixture Modelling
- (1)
- Set an upper limit for the number of latent classes based on the research question; if prior information is unavailable, can be set to 1–10 [33];
- (2)
- For each candidate number of classes from 1 to , fit the corresponding FMMs.
- (3)
- The optimal number of latent classes, , is then selected by minimizing the Bayesian Information Criterion (BIC), as its penalty term for model complexity is , which allows for asymptotically consistent selection of the true number of classes as the sample size increases [34]. This has been shown to provide a stable and reliable solution in mixture modeling [35,36].
2.4. Stage 2: Construction of Propensity Score Model and Outcome Projection Model Accounting for Population Heterogeneity
- (1)
- Propensity score model accounting for population heterogeneity: A logistic mixed effects model is used to estimate inclusion probabilities, treating the heterogeneous groups identified in stage 1 as a level-2 variable. The model is fitted with common covariates using the combined sample from and , with weighted by , where . This scaling adjustment is applied to stabilize the estimation process by mitigating the disproportionate influence that can arise from the disparity in effective sample size between the weighted and . Specifically, it prevents the composite likelihood from being dominated by a small number of units with extremely large design weights. As a result, this adjustment is crucial for achieving stable and efficient estimation of the model parameters [30]. We denote the model as , specified as:
- (2)
- Outcome projection model accounting for population heterogeneity: A generalized linear mixed effects model is constructed between common covariates and based on . This model is used to project outcomes in . Denote the model as with the following specification:
2.5. Stage 3: Construction of a Doubly Robust Estimator
3. Simulation Studies
3.1. Data-Generating Models
- (1)
- Under scenarios without population heterogeneity, continuous outcomes are generated as follows [20]:where comprises variables following Bernoulli, Uniform, Poisson, Chi-squared, and Normal distributions, respectively. This mixture reflects realistic data structures with diverse variable types. The covariance matrix is set to be an identity matrix, reflecting the conditional independence assumption common in FMMs. encodes the signs of coefficients. . We set to satisfy , where is the total variance of the covariates. The parameter reflects both and correlations. Binary outcomes are generated as: , where is chosen so that the overall prevalence is 50%.
- (2)
- Under scenarios with population heterogeneity, the population consists of at least two subpopulations. We define as the number of heterogeneous populations. Here we illustrate the data-generating models for . For individuals in subpopulation 1 () and subpopulation 2 (), continuous outcomes are generated as:
3.2. Simulated Scenarios
3.3. Evaluating Criteria
3.4. Results
4. Application to Health Monitoring SMS Survey Data
5. Discussion
6. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Software and Implementation Details
References
- Elliott, M.R.; Valliant, R. Inference for Nonprobability Samples. Stat. Sci. 2017, 32, 249–264. [Google Scholar] [CrossRef] [Scilit]
- Czajka, J.L.; Beyler, A. Declining Response Rates in Federal Surveys: Trends and Implications; Background Paper; Mathematica Policy Research: Washington, DC, USA, 2016. [Google Scholar]
- Barbier, S.; Loosveldt, G.; Carton, A. (Eds.) The flemish survey climate: An analysis based on the survey of social-cultural changes in flanders. In Proceedings of the International Workshop on Household Survey Nonresponse, Leuven, Belgium, 2–4 September 2015. [Google Scholar]
- Valliant, R.; Dever, J.A.; Kreuter, F. Practical Tools for Designing and Weighting Survey Samples; Springer International Publishing: Cham, Switzerland, 2018. [Google Scholar]
- Matias, J.; Leavitt, A. COVID-19 Social Science Research Tracker. 2020. Available online: https://github.com/natematias/covid-19-social-science-research (accessed on 21 December 2025).
- Green, R.K.; Nieser, K.J.; Jacobsohn, G.C.; Cochran, A.L.; Caprio, T.V.; Cushman, J.T.; Kind, A.J.; Lohmeier, M.; Shah, M.N. Differential effects of an emergency department-to-home care transitions intervention in an older adult population: A latent class analysis. Med. Care 2023, 61, 400–408. [Google Scholar] [CrossRef] [Scilit]
- del Mar Rueda, M.; Pasadas-del-Amo, S.; Rodríguez, B.C.; Castro-Martín, L.; Ferri-García, R. Enhancing estimation methods for integrating probability and nonprobability survey samples with machine-learning techniques. An application to a Survey on the impact of the COVID-19 pandemic in Spain. Biom. J. 2023, 65, 2200035. [Google Scholar] [CrossRef] [Scilit]
- Wu, C. Statistical inference with non-probability survey samples. Surv. Methodol. 2022, 48, 283–311. [Google Scholar]
- Baker, R.; Brick, J.M.; Bates, N.A.; Battaglia, M.; Couper, M.P.; Dever, J.A.; Gile, K.J.; Tourangeau, R. Summary report of the AAPOR task force on non-probability sampling. J. Surv. Stat. Methodol. 2013, 1, 90–143. [Google Scholar] [CrossRef] [Scilit]
- Bradley, V.C.; Kuriwaki, S.; Isakov, M.; Sejdinovic, D.; Meng, X.-L.; Flaxman, S. Unrepresentative big surveys significantly overestimated US vaccine uptake. Nature 2021, 600, 695–700. [Google Scholar] [CrossRef] [Scilit]
- Yang, S.; Kim, J.K. Statistical data integration in survey sampling: A review. Jpn. J. Stat. Data Sci. 2020, 3, 625–650. [Google Scholar] [CrossRef] [Scilit]
- Cornesse, C.; Blom, A.G.; Dutwin, D.; Krosnick, J.A.; De Leeuw, E.D.; Legleye, S.; Pasek, J.; Pennay, D.; Phillips, B.; Sakshaug, J.W.; et al. A Review of Conceptual Approaches and Empirical Evidence on Probability and Nonprobability Sample Survey Research. J. Surv. Stat. Methodol. 2020, 8, 4–36. [Google Scholar] [CrossRef] [Scilit]
- Valliant, R. Comparing Alternatives for Estimation from Nonprobability Samples. J. Surv. Stat. Methodol. 2020, 8, 231–263. [Google Scholar] [CrossRef] [Scilit]
- Valliant, R.; Dever, J.A. Estimating Propensity Adjustments for Volunteer Web Surveys. Sociol. Methods Res. 2011, 40, 105–137. [Google Scholar] [CrossRef] [Scilit]
- Castro-Martín, L.; del Mar Rueda, M.; Ferri-García, R. Combining statistical matching and propensity score adjustment for inference from non-probability surveys. J. Comput. Appl. Math. 2022, 404, 113414. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Graubard, B.I.; Katki, H.A.; Li, Y. Improving external validity of epidemiologic cohort analyses: A kernel weighting approach. J. R. Stat. Soc. Ser. A (Stat. Soc.) 2020, 183, 1293–1311. [Google Scholar] [CrossRef] [Scilit]
- Kern, C.; Li, Y.; Wang, L. Boosted Kernel Weighting—Using Statistical Learning to Improve Inference from Nonprobability Samples. J. Surv. Stat. Methodol. 2021, 9, 1088–1113. [Google Scholar] [CrossRef] [Scilit]
- Dagdoug, M.; Goga, C.; Haziza, D. Model-Assisted Estimation Through Random Forests in Finite Population Sampling. J. Am. Stat. Assoc. 2023, 118, 1234–1251. [Google Scholar] [CrossRef] [Scilit]
- Downes, M.; Gurrin, L.C.; English, D.R.; Pirkis, J.; Currier, D.; Spittal, M.J.; Carlin, J.B. Multilevel Regression and Poststratification: A Modeling Approach to Estimating Population Quantities From Highly Selected Survey Samples. Am. J. Epidemiol. 2018, 187, 1780–1790. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, Y.; Li, P.; Wu, C. Doubly robust inference with nonprobability survey samples. J. Amer. Stat. Assoc. 2020, 115, 2011–2021. [Google Scholar] [CrossRef] [Scilit]
- Chen, S.; Haziza, D. General purpose multiply robust data integration procedures for handling nonprobability samples. Scand. J. Stat. 2023, 50, 697–724. [Google Scholar] [CrossRef] [Scilit]
- Muthin, B.O. Latent variable modeling in heterogeneous populations. Psychometrika 1989, 54, 557–585. [Google Scholar] [CrossRef] [Scilit]
- Fitch, P.J.R.; Lovell, M.A.; Davies, S.J.; Pritchard, T.; Harvey, P.K. An integrated and quantitative approach to petrophysical heterogeneity. Mar. Pet. Geol. 2015, 63, 82–96. [Google Scholar] [CrossRef] [Scilit]
- Myint, P.K.; Luben, R.N.; Wareham, N.J.; Bingham, S.A.; Khaw, K.-T. Combined effect of health behaviours and risk of first ever stroke in 20 040 men and women over 11 years’ follow-up in norfolk cohort of european prospective investigation of cancer (EPIC norfolk): Prospective population study. BMJ 2009, 338, b349. [Google Scholar] [CrossRef] [Scilit]
- Le, L.T.H.; Hoang, T.N.A.; Nguyen, T.T.; Dao, T.D.; Do, B.N.; Pham, K.M.; Vu, V.H.; Pham, L.V.; Nguyen, L.T.H.; Nguyen, H.C.; et al. Sex Differences in Clustering Unhealthy Lifestyles Among Survivors of COVID-19: Latent Class Analysis. JMIR Public Health Surveill. 2024, 10, e50189. [Google Scholar] [CrossRef] [Scilit]
- Lv, J.; Liu, Q.; Ren, Y.; Gong, T.; Wang, S.; Li, L.; Community Interventions for Health (CIH) Collaboration. Socio-demographic association of multiple modifiable lifestyle risk factors and their clustering in a representative urban population of adults: A cross-sectional study in hangzhou, China. Int. J. Behav. Nutr. Phys. Act. 2011, 8, 40. [Google Scholar] [CrossRef] [Scilit]
- Mannering, F.L.; Shankar, V.; Bhat, C.R. Unobserved heterogeneity and the statistical analysis of highway accident data. Anal. Methods Accid. Res. 2016, 11, 1–16. [Google Scholar] [CrossRef] [Scilit]
- Kim, S.H. How heterogeneity has been examined in transportation safety analysis: A review of latent class modeling applications. Anal. Methods Accid. Res. 2023, 40, 100292. [Google Scholar] [CrossRef] [Scilit]
- Kelly, S.; Kaye, S.-A.; Oviedo-Trespalacios, O. What factors contribute to the acceptance of artificial intelligence? A systematic review. Telemat. Inform. 2023, 77, 101925. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Graubard, B.I.; Katki, H.A.; Li, Y. Efficient and robust propensity-score-based methods for population inference using epidemiologic cohorts. Int. Stat. Rev. 2022, 90, 146–164. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Valliant, R.; Li, Y. Adjusted logistic propensity weighting methods for population inference using nonprobability volunteer-based epidemiologic cohorts. Stat. Med. 2021, 40, 5237–5250. [Google Scholar] [CrossRef] [Scilit]
- Kim, S.H.; Mokhtarian, P.L. Finite mixture (or latent class) modeling in transportation: Trends, usage, potential, and future directions. Transp. Res. Part B Methodol. 2023, 172, 134–173. [Google Scholar] [CrossRef] [Scilit]
- Lezhnina, O.; Kismihók, G. Latent Class Cluster Analysis: Selecting the number of clusters. MethodsX 2022, 9, 101747. [Google Scholar] [CrossRef] [Scilit]
- Henson, J.M.; Reise, S.P.; Kim, K.H. Detecting Mixtures From Structural Model Differences Using Latent Variable Mixture Modeling: A Comparison of Relative Model Fit Statistics. Struct. Equ. Model. A Multidiscip. J. 2007, 14, 202–226. [Google Scholar] [CrossRef] [Scilit]
- McLachlan, G.J.; Lee, S.X.; Rathnayake, S.I. Finite mixture models. Annu. Rev. Stat. Its Appl. 2019, 6, 355–378. [Google Scholar] [CrossRef] [Scilit]
- Sinha, P.; Calfee, C.S.; Delucchi, K.L. Practitioner’s guide to latent class analysis: Methodological considerations and common pitfalls. Crit. Care Med. 2021, 49, e63–e79. [Google Scholar] [CrossRef] [Scilit]
- R Core Team. R: A Language and Environment for Statistical Computing; R Foundation for Statistical Computing: Vienna, Austria, 2020. [Google Scholar]
- Bates, D. Computational methods for mixed models. Vignette Lme4 2011, 1045, 1046. [Google Scholar]
- Harpole, J.K.; Woods, C.M.; Rodebaugh, T.L.; Levinson, C.A.; Lenze, E.J. How bandwidth selection algorithms impact exploratory data analysis using kernel density estimation. Psychol. Methods 2014, 19, 428–443. [Google Scholar] [CrossRef] [Scilit]
- Bretos-Azcona, P.E.; Sánchez-Iriso, E.; Cabasés Hita, J.M. Tailoring integrated care services for high-risk patients with multiple chronic conditions: A risk stratification approach using cluster analysis. BMC Health Serv. Res. 2020, 20, 806. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Abarda, A.; Contractor, A.A.; Wang, J.; Dayton, C.M. Exploring heterogeneity in clinical trials with latent class analysis. Ann. Transl. Med. 2018, 6, 119. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bourke, M.; Wang, H.F.W.; McNaughton, S.A.; Thomas, G.; Firth, J.; Trott, M.; Cairney, J. Clusters of healthy lifestyle behaviours are associated with symptoms of depression, anxiety, and psychological distress: A systematic review and meta-analysis of observational studies. Clin. Psychol. Rev. 2025, 118, 102585. [Google Scholar] [CrossRef] [Scilit]
- Morris, T.P.; White, I.R.; Crowther, M.J. Using simulation studies to evaluate statistical methods. Stat. Med. 2019, 38, 2074–2102. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Boyd, R.J.; Powney, G.D.; Pescott, O.L. We need to talk about nonprobability samples. Trends Ecol. Evol. 2023, 38, 521–531. [Google Scholar] [CrossRef] [Scilit]
- Van Lissa, C.J.; Garnier-Villarreal, M.; Anadria, D. Recommended Practices in Latent Class Analysis Using the Open-Source R-Package tidySEM. Struct. Equ. Model. A Multidiscip. J. 2024, 31, 526–534. [Google Scholar] [CrossRef] [Scilit]





| Notation | Description |
|---|---|
| Finite population of size | |
| Nonprobability sample of size | |
| Reference probability sample of size | |
| Sampling weight for individual in | |
| Target outcome for individual in | |
| Vector of covariates | |
| Inclusion indicator for the nonprobability sample (= 1 if ) | |
| Pre-specified upper limit for the number of latent classes | |
| Optimal number of latent classes selected via FMMs | |
| Latent class membership () | |
| Scaled sampling weights for individual in | |
| Logistic mixed effects model for the propensity score | |
| Generalized linear mixed effects for outcome projection | |
| Linear predictors from | |
| Kernel function (standard normal) | |
| Bandwidth parameter for kernel smoothing | |
| The kernel smoothed weights for individual in |
| ICC | NHDR | DR | IPW | SP | Naïve | |||
|---|---|---|---|---|---|---|---|---|
| NA | 1 | 0 | 0 | 95.66 | 95.27 | 100.00 | 95.27 | 0.23 |
| 0.3 | 2 | 1 | 1 | 96.20 | 93.83 | 100.00 | 93.87 | 0.60 |
| 2 | 2 | 1 | 95.69 | 93.67 | 100.00 | 93.73 | 0.67 | |
| 2 | 3 | 1 | 96.43 | 87.07 | 99.83 | 86.80 | 1.33 | |
| 2 | 4 | 1 | 96.06 | 77.63 | 98.27 | 77.63 | 2.37 | |
| 2 | 5 | 1 | 96.21 | 70.70 | 96.03 | 70.57 | 2.37 | |
| 3 | 2 | 1 | 95.61 | 93.07 | 100.00 | 93.07 | 0.63 | |
| 3 | 3 | 1 | 95.47 | 89.83 | 99.93 | 89.97 | 0.93 | |
| 3 | 4 | 1 | 94.86 | 79.10 | 99.80 | 79.33 | 1.07 | |
| 3 | 5 | 1 | 93.96 | 71.10 | 99.13 | 71.60 | 1.20 | |
| 4 | 3 | 1 | 94.89 | 90.10 | 99.93 | 90.20 | 0.60 | |
| 4 | 3 | 2 | 95.96 | 91.13 | 99.97 | 91.40 | 0.47 | |
| 4 | 4 | 1 | 92.79 | 60.47 | 96.33 | 60.63 | 1.10 | |
| 4 | 4 | 2 | 93.75 | 58.80 | 96.23 | 59.00 | 1.03 | |
| 4 | 5 | 1 | 92.96 | 56.90 | 95.13 | 57.13 | 1.03 | |
| 4 | 5 | 2 | 93.64 | 51.20 | 91.17 | 51.73 | 1.20 | |
| 5 | 4 | 1 | 85.75 | 63.50 | 97.97 | 64.37 | 0.77 | |
| 5 | 4 | 2 | 93.18 | 52.07 | 92.90 | 52.30 | 0.93 | |
| 5 | 5 | 1 | 86.21 | 62.17 | 96.97 | 62.90 | 0.67 | |
| 5 | 5 | 2 | 93.75 | 45.87 | 89.00 | 46.37 | 1.00 | |
| 0.5 | 2 | 1 | 1 | 95.49 | 94.83 | 100.00 | 94.97 | 1.07 |
| 2 | 2 | 1 | 95.22 | 94.40 | 100.00 | 94.43 | 0.97 | |
| 2 | 3 | 1 | 96.25 | 88.07 | 99.93 | 88.03 | 2.13 | |
| 2 | 4 | 1 | 97.16 | 79.40 | 99.87 | 79.60 | 3.37 | |
| 2 | 5 | 1 | 96.53 | 73.87 | 99.43 | 74.00 | 3.23 | |
| 3 | 2 | 1 | 95.91 | 93.27 | 100.00 | 93.47 | 1.20 | |
| 3 | 3 | 1 | 95.07 | 91.67 | 100.00 | 91.77 | 1.10 | |
| 3 | 4 | 1 | 95.41 | 82.57 | 100.00 | 82.80 | 1.53 | |
| 3 | 5 | 1 | 94.87 | 77.70 | 99.97 | 77.90 | 2.33 | |
| 4 | 3 | 1 | 94.86 | 91.77 | 100.00 | 92.10 | 1.33 | |
| 4 | 3 | 2 | 94.35 | 90.70 | 100.00 | 90.77 | 1.10 | |
| 4 | 4 | 1 | 94.24 | 68.47 | 99.67 | 68.37 | 1.03 | |
| 4 | 4 | 2 | 94.78 | 65.03 | 99.73 | 65.63 | 2.03 | |
| 4 | 5 | 1 | 93.96 | 62.37 | 99.33 | 62.47 | 1.70 | |
| 4 | 5 | 2 | 94.00 | 55.90 | 98.80 | 56.07 | 1.50 | |
| 5 | 4 | 1 | 87.92 | 70.10 | 99.90 | 70.70 | 1.33 | |
| 5 | 4 | 2 | 94.53 | 57.93 | 98.97 | 58.27 | 1.57 | |
| 5 | 5 | 1 | 87.74 | 67.67 | 99.83 | 68.27 | 1.20 | |
| 5 | 5 | 2 | 94.93 | 52.27 | 98.90 | 52.67 | 1.67 | |
| 0.8 | 2 | 1 | 1 | 95.35 | 94.03 | 100.00 | 94.07 | 3.47 |
| 2 | 2 | 1 | 96.09 | 95.40 | 100.00 | 95.37 | 3.60 | |
| 2 | 3 | 1 | 95.41 | 93.20 | 100.00 | 93.10 | 5.37 | |
| 2 | 4 | 1 | 95.97 | 88.40 | 100.00 | 88.23 | 5.57 | |
| 2 | 5 | 1 | 96.48 | 83.60 | 100.00 | 83.67 | 5.53 | |
| 3 | 2 | 1 | 95.44 | 94.73 | 100.00 | 94.70 | 3.93 | |
| 3 | 3 | 1 | 95.54 | 94.57 | 100.00 | 94.60 | 3.10 | |
| 3 | 4 | 1 | 94.78 | 90.23 | 100.00 | 90.37 | 4.33 | |
| 3 | 5 | 1 | 95.37 | 85.67 | 100.00 | 85.77 | 3.90 | |
| 4 | 3 | 1 | 95.55 | 93.90 | 100.00 | 93.97 | 3.40 | |
| 4 | 3 | 2 | 94.82 | 93.73 | 100.00 | 93.73 | 3.90 | |
| 4 | 4 | 1 | 94.76 | 80.47 | 100.00 | 80.43 | 4.30 | |
| 4 | 4 | 2 | 94.90 | 79.77 | 100.00 | 79.87 | 4.13 | |
| 4 | 5 | 1 | 94.41 | 78.77 | 100.00 | 78.63 | 4.17 | |
| 4 | 5 | 2 | 95.07 | 69.27 | 100.00 | 69.43 | 4.40 | |
| 5 | 4 | 1 | 91.28 | 82.20 | 100.00 | 82.30 | 3.20 | |
| 5 | 4 | 2 | 94.81 | 73.67 | 100.00 | 73.97 | 3.97 | |
| 5 | 5 | 1 | 91.34 | 79.33 | 100.00 | 79.50 | 3.97 | |
| 5 | 5 | 2 | 95.25 | 69.20 | 100.00 | 69.43 | 3.77 |
| ICC | NHDR | DR | IPW | SP | Naïve | |||
|---|---|---|---|---|---|---|---|---|
| NA | 1 | 0 | 0 | 2.09 | 2.00 | 2.12 | 2.00 | 24.47 |
| 0.3 | 2 | 1 | 1 | 2.23 | 2.47 | 2.63 | 2.47 | 24.68 |
| 2 | 2 | 1 | 2.26 | 2.54 | 2.68 | 2.55 | 25.04 | |
| 2 | 3 | 1 | 2.22 | 4.44 | 4.50 | 4.47 | 18.56 | |
| 2 | 4 | 1 | 2.34 | 7.55 | 7.52 | 7.56 | 12.76 | |
| 2 | 5 | 1 | 2.33 | 9.73 | 9.76 | 9.72 | 12.78 | |
| 3 | 2 | 1 | 2.37 | 2.84 | 3.09 | 2.84 | 25.98 | |
| 3 | 3 | 1 | 2.49 | 3.56 | 3.88 | 3.55 | 26.06 | |
| 3 | 4 | 1 | 2.62 | 6.20 | 6.06 | 6.16 | 21.95 | |
| 3 | 5 | 1 | 2.74 | 8.26 | 8.02 | 8.21 | 22.93 | |
| 4 | 3 | 1 | 2.84 | 3.99 | 4.36 | 3.98 | 30.47 | |
| 4 | 3 | 2 | 2.57 | 3.64 | 4.01 | 3.62 | 29.17 | |
| 4 | 4 | 1 | 3.75 | 12.83 | 13.83 | 12.70 | 25.47 | |
| 4 | 4 | 2 | 3.53 | 13.64 | 14.06 | 13.52 | 22.33 | |
| 4 | 5 | 1 | 3.60 | 14.31 | 15.25 | 14.14 | 25.17 | |
| 4 | 5 | 2 | 3.67 | 17.08 | 17.43 | 16.89 | 21.78 | |
| 5 | 4 | 1 | 7.05 | 12.46 | 12.90 | 12.25 | 29.14 | |
| 5 | 4 | 2 | 3.88 | 17.42 | 17.81 | 17.26 | 26.00 | |
| 5 | 5 | 1 | 7.09 | 13.32 | 13.82 | 13.11 | 29.55 | |
| 5 | 5 | 2 | 3.78 | 20.44 | 20.69 | 20.22 | 25.71 | |
| 0.5 | 2 | 1 | 1 | 2.71 | 2.68 | 2.91 | 2.68 | 26.90 |
| 2 | 2 | 1 | 2.67 | 2.89 | 3.08 | 2.90 | 26.78 | |
| 2 | 3 | 1 | 2.70 | 4.75 | 4.83 | 4.77 | 20.35 | |
| 2 | 4 | 1 | 2.64 | 7.92 | 7.88 | 7.93 | 14.12 | |
| 2 | 5 | 1 | 2.70 | 10.36 | 10.42 | 10.35 | 14.23 | |
| 3 | 2 | 1 | 2.80 | 3.22 | 3.53 | 3.23 | 27.41 | |
| 3 | 3 | 1 | 2.94 | 3.76 | 4.14 | 3.74 | 28.26 | |
| 3 | 4 | 1 | 3.01 | 6.29 | 5.99 | 6.26 | 23.99 | |
| 3 | 5 | 1 | 3.12 | 7.92 | 7.54 | 7.88 | 24.35 | |
| 4 | 3 | 1 | 3.27 | 4.05 | 4.48 | 4.03 | 31.46 | |
| 4 | 3 | 2 | 3.25 | 4.28 | 4.78 | 4.25 | 31.87 | |
| 4 | 4 | 1 | 3.95 | 12.45 | 13.93 | 12.31 | 27.10 | |
| 4 | 4 | 2 | 3.82 | 13.60 | 14.27 | 13.47 | 24.43 | |
| 4 | 5 | 1 | 4.07 | 14.49 | 15.78 | 14.32 | 27.09 | |
| 4 | 5 | 2 | 3.93 | 17.59 | 18.23 | 17.41 | 24.36 | |
| 5 | 4 | 1 | 7.23 | 12.12 | 12.93 | 11.93 | 31.97 | |
| 5 | 4 | 2 | 4.17 | 18.00 | 18.65 | 17.79 | 28.56 | |
| 5 | 5 | 1 | 7.38 | 13.23 | 13.98 | 13.02 | 30.98 | |
| 5 | 5 | 2 | 4.12 | 20.33 | 20.79 | 20.10 | 27.06 | |
| 0.8 | 2 | 1 | 1 | 4.89 | 5.28 | 5.75 | 5.27 | 33.71 |
| 2 | 2 | 1 | 4.69 | 5.19 | 5.59 | 5.19 | 34.62 | |
| 2 | 3 | 1 | 4.80 | 6.38 | 6.54 | 6.40 | 28.73 | |
| 2 | 4 | 1 | 4.91 | 9.40 | 9.38 | 9.43 | 23.54 | |
| 2 | 5 | 1 | 4.86 | 12.54 | 12.63 | 12.53 | 22.66 | |
| 3 | 2 | 1 | 5.17 | 5.32 | 5.79 | 5.31 | 36.36 | |
| 3 | 3 | 1 | 5.03 | 5.55 | 6.16 | 5.52 | 37.10 | |
| 3 | 4 | 1 | 5.30 | 7.75 | 7.50 | 7.74 | 32.09 | |
| 3 | 5 | 1 | 5.31 | 9.99 | 9.30 | 9.97 | 32.88 | |
| 4 | 3 | 1 | 5.63 | 6.35 | 7.07 | 6.33 | 41.48 | |
| 4 | 3 | 2 | 5.75 | 6.83 | 7.72 | 6.80 | 41.67 | |
| 4 | 4 | 1 | 6.45 | 14.03 | 16.77 | 13.87 | 37.86 | |
| 4 | 4 | 2 | 6.25 | 15.40 | 16.58 | 15.24 | 34.30 | |
| 4 | 5 | 1 | 6.24 | 15.90 | 18.70 | 15.69 | 37.61 | |
| 4 | 5 | 2 | 6.50 | 21.16 | 22.35 | 20.96 | 32.64 | |
| 5 | 4 | 1 | 9.42 | 13.86 | 15.37 | 13.67 | 43.20 | |
| 5 | 4 | 2 | 7.03 | 19.48 | 20.99 | 19.28 | 38.94 | |
| 5 | 5 | 1 | 9.60 | 15.12 | 16.77 | 14.90 | 41.73 | |
| 5 | 5 | 2 | 7.01 | 22.84 | 23.96 | 22.62 | 37.93 |
| ICC | NHDR | DR | IPW | SP | Naïve | |||
|---|---|---|---|---|---|---|---|---|
| NA | 1 | 0 | 0 | 0.59 | 0.55 | 1.06 | 0.56 | 0.03 |
| 0.3 | 2 | 1 | 1 | 0.62 | 0.59 | 1.38 | 0.59 | 0.04 |
| 2 | 2 | 1 | 0.62 | 0.60 | 1.35 | 0.60 | 0.04 | |
| 2 | 3 | 1 | 0.65 | 0.63 | 1.30 | 0.63 | 0.04 | |
| 2 | 4 | 1 | 0.70 | 0.68 | 1.22 | 0.68 | 0.04 | |
| 2 | 5 | 1 | 0.71 | 0.71 | 1.17 | 0.71 | 0.04 | |
| 3 | 2 | 1 | 0.64 | 0.61 | 1.37 | 0.61 | 0.04 | |
| 3 | 3 | 1 | 0.64 | 0.62 | 1.36 | 0.62 | 0.04 | |
| 3 | 4 | 1 | 0.66 | 0.64 | 1.33 | 0.65 | 0.04 | |
| 3 | 5 | 1 | 0.66 | 0.65 | 1.30 | 0.66 | 0.04 | |
| 4 | 3 | 1 | 0.66 | 0.65 | 1.39 | 0.65 | 0.04 | |
| 4 | 3 | 2 | 0.66 | 0.65 | 1.38 | 0.65 | 0.04 | |
| 4 | 4 | 1 | 0.70 | 0.71 | 1.32 | 0.71 | 0.04 | |
| 4 | 4 | 2 | 0.71 | 0.72 | 1.30 | 0.72 | 0.04 | |
| 4 | 5 | 1 | 0.71 | 0.72 | 1.30 | 0.72 | 0.04 | |
| 4 | 5 | 2 | 0.72 | 0.74 | 1.26 | 0.74 | 0.04 | |
| 5 | 4 | 1 | 0.82 | 0.73 | 1.38 | 0.73 | 0.04 | |
| 5 | 4 | 2 | 0.75 | 0.76 | 1.34 | 0.76 | 0.04 | |
| 5 | 5 | 1 | 0.81 | 0.73 | 1.37 | 0.73 | 0.04 | |
| 5 | 5 | 2 | 0.75 | 0.77 | 1.31 | 0.77 | 0.04 | |
| 0.5 | 2 | 1 | 1 | 0.68 | 0.64 | 1.76 | 0.64 | 0.04 |
| 2 | 2 | 1 | 0.67 | 0.65 | 1.71 | 0.66 | 0.04 | |
| 2 | 3 | 1 | 0.70 | 0.68 | 1.67 | 0.68 | 0.04 | |
| 2 | 4 | 1 | 0.75 | 0.73 | 1.6 | 0.73 | 0.04 | |
| 2 | 5 | 1 | 0.75 | 0.76 | 1.56 | 0.76 | 0.04 | |
| 3 | 2 | 1 | 0.69 | 0.66 | 1.75 | 0.66 | 0.04 | |
| 3 | 3 | 1 | 0.69 | 0.67 | 1.73 | 0.67 | 0.04 | |
| 3 | 4 | 1 | 0.71 | 0.69 | 1.72 | 0.69 | 0.04 | |
| 3 | 5 | 1 | 0.73 | 0.70 | 1.69 | 0.71 | 0.04 | |
| 4 | 3 | 1 | 0.72 | 0.70 | 1.78 | 0.70 | 0.05 | |
| 4 | 3 | 2 | 0.71 | 0.70 | 1.76 | 0.70 | 0.05 | |
| 4 | 4 | 1 | 0.76 | 0.76 | 1.73 | 0.76 | 0.05 | |
| 4 | 4 | 2 | 0.76 | 0.78 | 1.72 | 0.77 | 0.05 | |
| 4 | 5 | 1 | 0.75 | 0.77 | 1.72 | 0.77 | 0.05 | |
| 4 | 5 | 2 | 0.77 | 0.79 | 1.69 | 0.79 | 0.05 | |
| 5 | 4 | 1 | 0.86 | 0.78 | 1.81 | 0.78 | 0.05 | |
| 5 | 4 | 2 | 0.81 | 0.81 | 1.78 | 0.81 | 0.05 | |
| 5 | 5 | 1 | 0.87 | 0.79 | 1.80 | 0.79 | 0.05 | |
| 5 | 5 | 2 | 0.81 | 0.83 | 1.76 | 0.83 | 0.05 | |
| 0.8 | 2 | 1 | 1 | 0.90 | 0.88 | 2.99 | 0.88 | 0.07 |
| 2 | 2 | 1 | 0.89 | 0.91 | 2.91 | 0.91 | 0.07 | |
| 2 | 3 | 1 | 0.92 | 0.92 | 2.90 | 0.93 | 0.07 | |
| 2 | 4 | 1 | 0.95 | 0.97 | 2.85 | 0.97 | 0.07 | |
| 2 | 5 | 1 | 0.96 | 1.00 | 2.82 | 1.00 | 0.07 | |
| 3 | 2 | 1 | 0.91 | 0.90 | 3.02 | 0.90 | 0.07 | |
| 3 | 3 | 1 | 0.92 | 0.90 | 3.00 | 0.90 | 0.07 | |
| 3 | 4 | 1 | 0.94 | 0.91 | 3.01 | 0.92 | 0.07 | |
| 3 | 5 | 1 | 0.94 | 0.93 | 2.98 | 0.94 | 0.07 | |
| 4 | 3 | 1 | 0.96 | 0.95 | 3.09 | 0.95 | 0.07 | |
| 4 | 3 | 2 | 0.96 | 0.96 | 3.08 | 0.96 | 0.07 | |
| 4 | 4 | 1 | 0.98 | 1.00 | 3.10 | 1.00 | 0.07 | |
| 4 | 4 | 2 | 1.00 | 1.01 | 3.08 | 1.01 | 0.07 | |
| 4 | 5 | 1 | 0.98 | 1.01 | 3.09 | 1.00 | 0.07 | |
| 4 | 5 | 2 | 1.00 | 1.03 | 3.06 | 1.03 | 0.07 | |
| 5 | 4 | 1 | 1.08 | 1.02 | 3.22 | 1.02 | 0.08 | |
| 5 | 4 | 2 | 1.04 | 1.04 | 3.21 | 1.04 | 0.07 | |
| 5 | 5 | 1 | 1.09 | 1.03 | 3.21 | 1.02 | 0.08 | |
| 5 | 5 | 2 | 1.04 | 1.06 | 3.21 | 1.06 | 0.07 |
| HMSS | GHSS-GZ-7 | p | |
|---|---|---|---|
| N = 1527 | N = 9299 | ||
| Age (Mean ± SD) | 37.6 ± 9.16 | 42.7 ± 11.0 | <0.001 |
| Sex | 0.0110 | ||
| Male | 44.86 | 48.42 | |
| Female | 55.14 | 51.58 | |
| Employment status | <0.001 | ||
| Employed | 81.73 | 75.52 | |
| Student | 2.16 | 1.74 | |
| Retired | 4.98 | 9.18 | |
| Unemployed or others | 11.13 | 13.55 | |
| Marital status | <0.001 | ||
| Married | 62.67 | 79.29 | |
| Unmarried | 30.98 | 16.81 | |
| Divorced or others | 6.35 | 3.90 | |
| Educational level | <0.001 | ||
| Primary school or below | 0.52 | 10.64 | |
| Middle school | 4.32 | 27.10 | |
| High school or vocational high school | 12.18 | 20.08 | |
| college or above | 82.97 | 42.19 | |
| District | <0.001 | ||
| Central urban districts | 51.3 | 29.8 | |
| Peripheral districts | 48.7 | 70.2 |
| Naïve | DR | IPW | SP | NHDR | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Estimate | 95%CI | CIW | Estimate | 95%CI | CIW | Estimate | 95%CI | CIW | Estimate | 95%CI | CIW | Estimate | 95%CI | CIW | |
| Electronic screen use | 11.72 | (11.49, 11.95) | 0.460 | 9.62 | (9.10, 10.14) | 1.041 | 9.60 | (8.70, 10.50) | 1.803 | 9.58 | (9.04, 10.13) | 1.091 | 8.85 | (7.83, 9.87) | 2.044 |
| Poor/very poor self-rated health (%) | 7.53 | (6.26, 8.97) | 0.027 | 10.75 | (7.47, 14.03) | 0.066 | 10.59 | (6.28, 14.90) | 0.086 | 10.84 | (7.20, 14.47) | 0.073 | 8.04 | (3.29, 12.78) | 0.095 |
| Weekly alcohol consumption over 14 units (%) | 6.35 | (5.18, 7.69) | 0.025 | 6.27 | (4.12, 8.43) | 0.043 | 6.02 | (3.55, 8.48) | 0.049 | 6.84 | (4.61, 9.07) | 0.045 | 6.20 | (4.34, 8.06) | 0.037 |
| Exercise over 150 min per week (%) | 17.22 | (15.36, 19.21) | 0.039 | 19.24 | (15.59, 22.90) | 0.073 | 20.61 | (15.06, 26.16) | 0.111 | 18.35 | (13.93, 22.76) | 0.088 | 16.50 | (9.81, 23.2) | 0.134 |
| Daily sleep duration less than 5 h (%) | 8.38 | (7.04, 9.89) | 0.029 | 7.91 | (5.51, 10.32) | 0.048 | 7.86 | (4.69, 11.03) | 0.063 | 7.95 | (5.40, 10.50) | 0.051 | 7.40 | (5.01, 9.80) | 0.048 |
| Poor/very poor sleep quality (%) | 14.21 | (12.5, 16.06) | 0.036 | 15.83 | (12.43, 19.23) | 0.068 | 15.23 | (10.76, 19.69) | 0.089 | 16.50 | (12.66, 20.34) | 0.077 | 11.59 | (7.70, 15.48) | 0.078 |
| Depression (%) | 37.39 | (34.96, 39.88) | 0.050 | 34.77 | (30.44, 39.10) | 0.087 | 35.40 | (29.31, 41.5) | 0.122 | 35.78 | (31.13, 40.43) | 0.093 | 42.50 | (25.93, 59.07) | 0.331 |
| Anxiety (%) | 26.26 | (24.07, 28.54) | 0.045 | 24.67 | (20.79, 28.55) | 0.078 | 22.91 | (17.75, 28.06) | 0.103 | 25.62 | (21.56, 29.68) | 0.081 | 19.33 | (14.90, 23.76) | 0.089 |
| High/very high stress (%) | 64.51 | (62.05, 66.91) | 0.049 | 59.23 | (54.93, 63.54) | 0.086 | 57.76 | (51.36, 64.15) | 0.128 | 59.37 | (54.32, 64.42) | 0.101 | 62.57 | (48.61, 76.52) | 0.279 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yang, Z.; Wang, X.; Wu, W.; Gu, J. Improved Doubly Robust Inference with Nonprobability Survey Samples Using Finite Mixture Models: Application to Health Monitoring SMS Survey Data. Mathematics 2026, 14, 118. https://doi.org/10.3390/math14010118
Yang Z, Wang X, Wu W, Gu J. Improved Doubly Robust Inference with Nonprobability Survey Samples Using Finite Mixture Models: Application to Health Monitoring SMS Survey Data. Mathematics. 2026; 14(1):118. https://doi.org/10.3390/math14010118
Chicago/Turabian StyleYang, Ziying, Xu Wang, Wenjing Wu, and Jing Gu. 2026. "Improved Doubly Robust Inference with Nonprobability Survey Samples Using Finite Mixture Models: Application to Health Monitoring SMS Survey Data" Mathematics 14, no. 1: 118. https://doi.org/10.3390/math14010118
APA StyleYang, Z., Wang, X., Wu, W., & Gu, J. (2026). Improved Doubly Robust Inference with Nonprobability Survey Samples Using Finite Mixture Models: Application to Health Monitoring SMS Survey Data. Mathematics, 14(1), 118. https://doi.org/10.3390/math14010118

