1. Background
The overall rate of neonatal mortality in India is 17 per 1000 live births in 2024 as per the World Bank data, and more than one-fourth of neonates (26%) died in the first 24 hours of birth considering the neonatal deaths after live birth [
1,
2]. The neonatal mortality varies by province as per the report of National Family Health Survey 2019–2021 [
3]. States such as Bihar, Uttarakhand, Chhattisgarh, and Uttar Pradesh depict very high rates (above the rate of 30 per 1000 live births) [
3]. Furthermore, the inequity in access to maternal and child health services is prominent in states that are lagging in achieving health outcomes [
3]. Studies reflect less efficient resource allocation; traditional methods of generating insights from available data are the correlates of reduced programme effectiveness in achieving outcomes [
4].
1.1. Evidence from Classical Analytics
The current study explores both the classical and innovative schools of thought. From diverse classical empirical studies exploring the linkage between neonatal survival outcomes and social determinants found factors like previous history of failed pregnancy/childbirth, factors related to maternal healthcare access and lifestyle factors as major determinants. A mixed-effect multivariate analysis applying four consecutive Ethiopian DHS datasets explored the trend of infant mortality and its determinants [
5]. The study by [
6] tested the association between survival status of preceding children and risk of mortality in successive ones along with other determinants, using the Cox Proportional Hazard model. A systematic review and meta-analysis exploring the timing of neonatal mortality and severe neonatal morbidity covered 36 reports consisting of observational studies (cross-sectional, cohort, prevalence, prospective) in 51 study settings, including a total of 6,760,731 live births and 47,551 neonatal deaths [
7]. Another systematic review and meta-analysis evaluating the association between birth spacing, and the adverse pregnancy outcomes on a global scale, reviewed 129 observational studies (cross-sectional, cohort, case-control) and experimental (cohort) studies comprising 46,874,843 participants [
8]. Another study investigated the relation between seven ‘quality of diet’ scores in pregnancy, birth and child health outcomes in the form of a narrative review [
9]. A systematic review and meta-analysis of 81 maternal and birth cohort studies in Gulf countries (Bahrain, Kuwait, Oman, Qatar, Saudi Arabia, and the United Arab Emirates) analysed the association between maternal obesity and large-for-gestational-age newborns [
10]. A cross-sectional study by [
11] explored the association between ethnicity, teenage pregnancy, and adverse birth outcomes in Suriname using national level and a few facility level data. The study by [
12] explored the determinants of adverse birth outcomes in Sub-Saharan Africa using data from 76,853 children. Another research study conducted a scoping review of 10 studies that explored the preconception risk factors and interventions related to teenage pregnancy to reduce adverse pregnancy and birth outcomes [
13]. Different facility-based cross-sectional studies or community-based studies in different low-and middle-income countries worth mentioning here explored the correlates of adverse birth outcomes—stillbirth, preterm birth, low birth weight occurring after premature rupture of membrane—as PROM mothers are more likely to have adverse birth outcomes, which is evident in 3% to 8% of pregnancies globally, contributing to neonatal mortality [
14,
15,
16,
17]. Another study by [
18] investigated the association between exposure to malaria and the risk of adverse maternal health, pregnancy and birth outcomes from 253 studies.
In addition to this further, the evidence from a number of recent studies depict higher effectiveness of machine learning and deep learning techniques over classical modelling techniques in exploring the determinants of health outcomes, with manual model adjustments for the classification of outcomes [
19,
20,
21,
22,
23,
24,
25].
1.2. Evidence from the Application of Advanced Analytics—Machine Learning and Deep Learning
The studies grouped under the theme of innovative and translational science benefit approach tested novel methods to improve public health service delivery through enhancing the quality of data gathering, storage, analysis, and decision-making [
26]. Such literary works reflect that improvement in healthcare decision support focuses on quality improvement in clinical decision-support systems (CDSS) and has reached developing stage of digital health maturity from the initial ideation level in most of the applications. Some of the studies build a use case to optimise public health decision support and propose models to adopt the concept for faster and more accurate analysis of routine data using machine learning and deep learning. The process trains a system, then is tested to find the best model as per the evaluation metrics [
22,
23,
24].
These models are mainly used to access past decisions, identify contextual variables, predict and suggest better solutions to learning, and make decisions to improve implementation [
22,
26,
27]. In healthcare, CDSSs are mainly used for knowledge and information management applying generally rule-based logic [
28,
29,
30] or algorithms [
29,
31], and some use neural networks [
32,
33,
34]. Evidently, the most used algorithms are artificial neural networks (ANN), support vector machines (SVM), and decision trees (DTs) in proposing CDSSs in different settings [
35]. In low-and middle-income countries, DSSs are used to strengthen public health services in the areas of health surveillance, evaluation of programmes, and predictive modelling [
36]. Previous studies indicated that public health DSSs (PHDSSs) require an approach to ensure their effectiveness. In some of the better-performing states in India, PHDSSs have already shown their effectiveness in improving access to care through pilot interventions [
26,
37,
38].
India has started deploying state-of-the- art technology to create and test the feasibility of DSS in the public health sector to ensure equitable access to basic health services, including MCH services. Among the different initiatives, SMARThealth Pregnancy—a CDSS using a mobile application—was aimed at reducing the cardiovascular risks associated with pregnancy and implemented in rural areas of India covering primary health centres in Jhajjar district, Haryana, and Guntur district in Andhra Pradesh, as a pilot and feasibility study [
26]. The integrated mobile CDSS applied simple algorithmic rules, collecting data at point of care on anaemia, haemoglobin, oral glucose tolerance test, and heart rate, along with historical data on the pregnant women, to generate signals based on traffic colour coding. This mobile CDSS model was found to be a successful decision-support system applied from the comfort of the homes of the pregnant women and used by community health workers in their routine activities. It helped to identify high-risk pregnancies and tailor ANC requirements for them—depicting higher efficiency and effectiveness—and improving the quality of service [
26]. Reproducible algorithms are applied to explore the feasibility of smart analytic-driven PHDSS or CDSS in identifying factors behind high rate of caesarean deliveries or in detecting anomalies with higher accuracy [
39]. The study by [
40] proposed a DSS to identify COVID-19 risk factors with higher accuracy and precision. A study by [
41] demonstrated the effectiveness of a smart mobile application-based CDSS for community health workers in managing non-communicable diseases in rural India. In a related study, a case-control pilot evaluated an mHealth application-based PHDSS across two primary health centres in northern India [
42]. The intervention group was exposed to a decision tree algorithm-based mobile application as point-of-care use by community health workers to record patient demographics, history of morbidity, and examination and investigation details, to generate a birth plan after identifying the risks. The study shows significant differences in triage assessment, ANC uptake and counselling among the case and control groups, where the control group participants received a conventional service. This provides evidence of the higher effectiveness of PHDSSs in rural areas of India.
1.3. Application of ML and DL in Improving MCH Services
A systematic review explored the usability of ML and/or DL models in improving the predictability of neonatal, infant, perinatal mortality or stillbirth outcomes from earlier research, and found that 83% of the selected research used ML algorithms and 17% used DL [
43]. Different studies focused on neonatal mortality used ML or DL for classifying the occurrence of neonatal death and mostly considered birth weight, gestational age, child’s sex, mother’s age, educational level, ethnicity, heart rate, number of ANC visits, history of previous pregnancies (type of pregnancy, type of delivery, number of healthy and stillbirths), and birth order as predictors [
43,
44,
45,
46,
47,
48,
49,
50,
51]. Almost 60% of the studies focused on classifying the occurrence of neonatal outcomes used random forest, 50% used neural networks, 50% used SVM, 25% used k-NN, 17% used XGBoost or GBM, and less than 10% used ensemble methods [
43].
1.4. Feature-Augmented Models and Benefits over Basic Models
Different studies under classical analytics have shown how factors such as short birth interval increase the risk of repeated preterm birth with nutritional shortage, Again, the risk of preterm birth after pregnancy loss with a short birth interval is more likely and the higher occurrence of preterm birth increases the risk of more adolescent pregnancies with poor obstetric history, placing teenage married girls in a locus of early marriage, early pregnancy, preterm birth, failed pregnancy, and recurrent preterm births leading to adverse pregnancy and/or birth outcomes [
52,
53]. These studies raise concerns in the classical school regarding the need for mediation analysis. Different studies have shown that augmented models outperform single or ensembled basic models in advanced analytics. For example, [
54] inferred that an augmented model brought 99.67% accuracy compared to other models, such as SVM and RF, with less time in processing. A study using a deep convolutional neural network (CNN) with the incorporation of transfer learning and data augmentation improved the classification of progressive neurodegenerative disease [
55]. In this study, the authors gathered 504 images, and 360 images were used for data augmentation; they ultimately worked on 4200 images and built a model with 89% classification accuracy. In another study, the use of a sparse encoder with CNN improved the classification accuracy by 3% from conventional algorithms, where they achieved 91.3% accuracy in predicting diabetes among the population [
56]. Another study classified the diabetic population by applying four augmentation techniques with three neural network algorithms and achieved more than 95% accuracy compared to other models that used only blood glucose level as model input [
57].
1.5. Gap Analysis
Despite growing use of machine-learning and deep-learning methods for neonatal mortality prediction, important methodological gaps remain. Preprocessing, feature selection, clustering, and resampling are not always restricted to training data, creating a risk of data leakage and optimistic performance estimates. Evaluation also often relies too heavily on accuracy and ROC-AUC, although neonatal mortality is a rare outcome that requires PR-AUC, sensitivity, specificity, predictive values, calibration, and uncertainty estimates. In addition, conventional random splitting may place children from the same mother or household across development and test sets and ignore DHS sampling weights and clustered survey structure. Finally, complex combinations of clustering, PSO, SMOTE, feature augmentation, and neural networks are frequently presented without fair comparison against simpler baselines or the controlled ablation of each component. These limitations motivate a household-grouped, leakage-controlled, survey-aware evaluation framework that compares simple and complex models under identical validation conditions.
1.6. Research Questions
The research questions addressed under this study are as follows:
RQ1. How accurately can neonatal mortality be predicted using maternal, socioeconomic, healthcare access, pregnancy, and birth-related NFHS-5 variables under a household-grouped, leakage-controlled validation protocol?
RQ2. How do logistic regression, random forest, histogram gradient boosting, and artificial neural networks compare under an identical repeated grouped-validation framework?
RQ3. What are the individual contributions of socioeconomic clustering, PSO feature selection, SMOTE, and out-of-fold feature augmentation?
RQ4. How does predictive performance change when models are restricted to antenatal, delivery time, or retrospective information?
RQ5. How well does the selected model transport across the three study states?
RQ6. Which prespecified maternal, pregnancy, socioeconomic, and healthcare access variables are associated with neonatal mortality in a survey-weighted reference analysis?
1.7. Novelty of This Study
The novelty of the present study lies not in introducing a new algorithm, but in developing a rigorous, survey-aware, and leakage-controlled modelling framework for neonatal mortality prediction using NFHS-5 data. The principal methodological contributions are as follows:
Applying household-grouped development and test partitions to prevent records from the same household or mother from appearing in both datasets;
Using repeated grouped nested cross-validation to ensure that model selection and performance estimation are conducted under a consistent validation protocol;
Fitting all preprocessing, socioeconomic clustering, particle swarm optimisation, feature augmentation, calibration, and resampling procedures exclusively within the corresponding training folds;
Incorporating DHS sampling weights during the training and evaluation of eligible models and accounting for the clustered survey structure in a separate association analysis;
Comparing logistic regression, random forest, histogram gradient boosting, and artificial neural networks under the same validation framework;
Conducting a controlled ablation analysis to quantify the individual contributions of socioeconomic clustering, PSO feature selection, SMOTE, and out-of-fold feature augmentation;
Prioritising precision–recall area under the curve as the primary selection metric and reporting sensitivity, specificity, positive and negative predictive values, F1-score, calibration, and decision-curve net benefit;
Generating household-grouped bootstrap confidence intervals to quantify uncertainty in final-test performance;
Distinguishing antenatal, delivery time, and retrospective prediction settings according to the real-world availability of input variables;
Evaluating geographic transportability through leave-one-state-out validation across Bihar, Chhattisgarh, and Uttarakhand;
Clearly separating predictive modelling from non-causal survey-weighted association analysis and from the future development of an operational decision-support system.
1.8. Significance of Our Work
The significance of this study lies in providing a more credible and practically interpretable assessment of neonatal mortality prediction than would be obtained from a conventional random split or accuracy-centred evaluation. The final results demonstrate that histogram gradient boosting performed better than logistic regression, random forest, and artificial neural networks, whereas the addition of clustering, PSO, SMOTE, and feature augmentation produced limited independent improvement. This finding is important because it shows that increasing modelling complexity does not necessarily improve prediction in structured population-survey data. The study also demonstrates that predictive performance depends strongly on the time at which information becomes available: antenatal prediction remained comparatively weak, while delivery time variables substantially improved discrimination. Although the selected model achieved high sensitivity and negative predictive value, its low positive predictive value indicates that it should be considered only as a potential low-cost screening aid rather than a diagnostic or automatic referral system. The framework therefore provides an evidence base for future prospective validation, state-specific recalibration, and integration with HMIS or community health worker workflows, while avoiding unsupported claims of causal effects, improved neonatal outcomes, or an already implemented decision-support system.
2. Methods
2.1. Study Design and Analytical Scope
This study conducted a secondary analysis of the Indian DHS National Family Health Survey 2019–2021 (NFHS-5) to develop and evaluate machine-learning models for neonatal mortality risk prediction. The primary analytical objective was predictive rather than causal. Accordingly, the main analysis compared several machine-learning models under a household-grouped, leakage-controlled, and survey-aware validation framework. A separate survey-weighted regression analysis was performed to examine the adjusted associations between prespecified maternal, socioeconomic, pregnancy, healthcare access factors and neonatal mortality. This association analysis was not interpreted as evidence of causal effects. The study did not implement or prospectively evaluate an operational decision-support system; instead, it assessed whether the resulting prediction framework could provide a methodological basis for future low-cost neonatal risk-screening applications.
2.2. Data Source, Study Setting, and Study Population
Data were obtained from the Kids Recode file of the NFHS-5, collected in India between 2019 and 2021 [
3]. The analysis focused on Bihar, Chhattisgarh, and Uttarakhand because these states reported comparatively high neonatal mortality and represented distinct eastern, central, and northern Indian settings. Eligible records were restricted to children with sufficient information to define neonatal survival status and the required maternal, household, pregnancy, delivery, and healthcare access variables. After applying the study eligibility criteria and data-quality checks, the final analytical sample comprised 33,338 child records.
NFHS-5 used a stratified multistage cluster-sampling design. The analysis therefore retained the DHS sampling-weight variable, primary sampling unit identifier, survey strata identifier, household identifier, and mother identifier. These variables were used to support survey-aware estimation, household-grouped validation, and verification that related observations were not divided between the development and final-test datasets.
Figure 1 displays an excerpt of the NFHS-5 analytical dataset used in the study.
2.3. Outcome Definition
The primary outcome was neonatal mortality, defined as death within the first 28 completed days after birth. The outcome was encoded as a binary variable, where 1 represented neonatal death, and 0 represented survival beyond the neonatal period. Records with missing or internally inconsistent outcome information were excluded before model development. Because neonatal mortality represented a small proportion of the observations, it was treated as a rare-event classification problem throughout model development and evaluation.
2.4. Predictor Domains and Variable Construction
Candidate predictors were selected from established determinants of neonatal health and were organised into five domains: maternal and demographic characteristics, socioeconomic and household conditions, maternal health and behavioural characteristics, pregnancy and obstetric history, healthcare access, delivery, and neonatal characteristics. Derived variables included maternal-age categories, birth interval categories, maternal body mass index groups, anaemia status, history of failed pregnancy, antenatal-care uptake, institutional delivery, birth attendance indicators, postnatal care indicators, sanitation, water source, health insurance coverage, and Janani Suraksha Yojana assistance. Variables were retained only when they could be derived consistently from the NFHS-5 data dictionary and had a plausible temporal relationship with the prediction setting under consideration. Variables that directly determined an intermediate target were excluded from the corresponding first-stage prediction model to prevent target leakage.
2.5. Data Audit, Distribution Assessment, and Missingness
Before model development, the analytical dataset was examined for outcome prevalence, missingness, implausible values, category sparsity, and the distributions of continuous and categorical variables. Numerical predictors were imputed using statistics estimated from the training data, whereas categorical variables were imputed using training-derived category frequencies or most-frequent values. Missingness indicators were retained for selected clinically important variables when the absence of measurement could itself reflect healthcare access or data collection processes. Categorical variables were one-hot encoded, and numerical scaling was applied where required by the corresponding learner. All imputations, encoding, scaling, and missingness transformations were fitted exclusively on the training portion of each validation fold and then applied to the corresponding held-out observations. The data audit summary is illustrated in
Figure 2.
2.6. Complex Survey Design and Grouping Structure
The NFHS-5 sampling weight was normalised and used during the fitting of eligible non-resampled models and during weighted final-test evaluation. Primary sampling unit and survey strata identifiers were retained for the survey-weighted association analysis. Household and mother identifiers were used to prevent related records from appearing in both development and test partitions. This was necessary because children from the same household or mother may share socioeconomic, behavioural, environmental, and healthcare access characteristics. For the survey-weighted association analysis, uncertainty was estimated using stratified primary sampling unit bootstrap resampling. Regularisation was applied for numerical stability, and the resulting coefficients were interpreted only as adjusted associations.
2.7. Household-Grouped Development and Final-Test Split
The analytical sample was divided into development and final-test datasets using household as the grouping unit. All children belonging to the same household were assigned to the same partition. Mother identifiers were subsequently checked to confirm that no mother appeared in both datasets. The outcome distribution was approximately preserved across partitions where feasible. The final-test dataset was isolated before preprocessing, feature selection, clustering, resampling, calibration, threshold selection, or model tuning. It was used only once for final evaluation after the modelling pipeline had been selected using the development data.
2.8. Prediction Timepoint Definitions
Three prediction settings were evaluated to distinguish model performance according to the real-world availability of input information. The antenatal model included only maternal, socioeconomic, household, obstetric history, and healthcare access variables that could reasonably be known before delivery. The delivery time model additionally included information available during or immediately after childbirth, such as gestational duration, delivery setting, and birth-related characteristics. The retrospective model included the complete eligible predictor set, including postnatal information, and was treated as an upper-bound benchmark rather than an early-warning model. This separation prevented retrospective variables from being incorrectly presented as inputs for antenatal screening.
2.9. Socioeconomic Clustering
Socioeconomic clustering was evaluated as an optional contextual feature-engineering component. State, place of residence, wealth category, and maternal education were used to derive socioeconomic groupings. Nominal variables were one-hot encoded before K-means clustering, and the cluster model was fitted only on the training data within each validation fold. Cluster assignments for held-out observations were generated using the corresponding training-fitted cluster model. Clustering was included in the controlled ablation analysis to determine whether it provided predictive information beyond the original socioeconomic variables.
2.10. Particle Swarm Optimisation for Feature Selection
Binary particle swarm optimisation was used to evaluate candidate predictor subsets. Each particle represented a binary feature-inclusion mask. Binary particle swarm optimisation used 12 particles and 12 iterations, with inertia, cognitive, and social coefficients of 0.72, 1.49, and 1.49, respectively. Candidate subsets were evaluated using three-fold cross-validated precision–recall area under the curve within the corresponding training data. A feature-count penalty was included to discourage unnecessarily large subsets. PSO was rerun independently within each training fold. Validation and test observations were never used to select features. Convergence histories and selected feature sets were retained to support reproducibility.
2.11. Out-of-Fold Feature Augmentation
Feature augmentation was evaluated using predicted probabilities for four intermediate outcomes: preterm birth, low birth weight, failed pregnancy history, and teenage or late pregnancy. For each intermediate outcome, direct source or proxy variables were removed from the corresponding first-stage model. Household-grouped out-of-fold predictions were generated for the development data, while held-out predictions were obtained from models fitted only on the corresponding training data. These predicted probabilities were then evaluated as additional inputs to the neonatal mortality classifier. They were treated as predictive features rather than mediators, and no causal interpretation was assigned to them.
2.12. Class Imbalance Handling
Four training strategies were compared: no resampling, random undersampling, random oversampling, and synthetic minority oversampling. Resampling was performed only after preprocessing the training portion of each fold. Validation and final-test observations were never resampled or used to generate synthetic observations.
Because synthetic observations do not inherit natural DHS sampling weights, resampled models were treated as methodological sensitivity analyses. Survey weights were used during fitting for eligible non-resampled models and during final weighted evaluation.
2.13. Candidate Predictive Models
Four model families were evaluated under the same household-grouped validation framework:
The same development partitions and primary model selection criteria were used for all learners. Model complexity was not assumed to confer superior performance, and the final learner was selected empirically.
2.14. Controlled Ablation Analysis
A controlled ablation analysis was conducted to estimate the marginal contribution of each modelling component. The following specifications were evaluated sequentially:
Baseline logistic regression;
Logistic regression with socioeconomic clustering;
Logistic regression with PSO-selected features;
Logistic regression with clustering and PSO;
Logistic regression with clustering, PSO, and SMOTE;
Logistic regression with clustering, PSO, and out-of-fold feature augmentation;
The complete logistic regression framework;
Fair comparison of logistic regression, random forest, histogram gradient boosting, and artificial neural network models.
All comparisons used the same grouped-validation structure and primary performance criterion.
2.15. Repeated Household-Grouped Nested Cross-Validation
Model development used repeated household-grouped nested cross-validation. The outer folds estimated model performance, while the inner training process handled preprocessing, clustering, PSO feature selection, augmentation, resampling, and learner fitting. No information from the outer validation fold was used during any upstream modelling step. The same household grouping was maintained throughout the validation procedure. The primary model selection metric was precision–recall area under the curve because neonatal mortality was rare. The one-standard-error rule was applied to avoid selecting unnecessarily complex models when performance differences were small. A detailed workflow for the full mechanism is shown in
Figure 3.
2.16. Probability Calibration and Screening-Threshold Selection
Calibration was based on household-grouped out-of-fold development predictions generated through the complete nested pipeline. A probability calibrator was fitted using development predictions only. The final classification threshold was selected from the calibrated development predictions to target approximately 80% survey-weighted sensitivity. Neither calibration nor threshold selection used the final-test dataset. The selected threshold was subsequently applied unchanged to the final-test probabilities.
2.17. Performance Evaluation
Model discrimination was assessed using ROC-AUC and precision–recall AUC, with PR-AUC designated as the primary metric. Classification performance at the selected screening threshold was assessed using sensitivity, specificity, positive predictive value, negative predictive value, F1-score, and weighted confusion matrices. Probability accuracy and agreement were assessed using the Brier score, calibration intercept, calibration slope, expected calibration error, and calibration plots. Decision-curve analysis was used to estimate potential net benefit across clinically relevant probability thresholds. Operational screening burden was reported as alerts, true deaths detected, false alerts, and individuals flagged per true death detected per 1000 births. These metrics are displayed in
Figure 4.
2.18. Statistical Uncertainty and Model Comparison
Uncertainty in final-test performance was quantified using household-grouped bootstrap resampling, thereby preserving within-household dependence. Ninety-five per cent confidence intervals were calculated for ROC-AUC, PR-AUC, sensitivity, specificity, positive and negative predictive values, F1-score, Brier score, calibration intercept, calibration slope, and expected calibration error. Differences between model specifications across repeated cross-validation folds were evaluated using corrected repeated cross-validation confidence intervals and tests that accounted for dependence created by overlapping training sets.
2.19. Leave-One-State-Out Validation
Geographic transportability was examined through leave-one-state-out validation. In each analysis, models were trained using data from two states and evaluated in the third state. Preprocessing, clustering, PSO, augmentation, calibration, and threshold selection were repeated using only the two-state training data. The held-out state was not used during model development. This analysis was interpreted as internal geographic validation rather than independent external validation.
2.20. Model Interpretation
Permutation importance was calculated using development-validation data rather than the final-test dataset. Importance values reflected the reduction in predictive performance when individual features were permuted. These values were interpreted as measures of predictive contribution and not as causal effects or evidence of modifiable risk factors.
2.21. Survey-Weighted Association Analysis
A separate survey-weighted penalised logistic regression was fitted using a prespecified reduced set of maternal, socioeconomic, pregnancy, and healthcare access variables. DHS sampling weights were incorporated, and uncertainty was estimated through stratified primary-sampling-unit bootstrap resampling. Regularisation was used to improve numerical stability in the presence of rare outcomes and sparse categories. The resulting coefficients and confidence intervals were interpreted as adjusted associations only. They were not used to establish mediation, treatment effects, or causal mechanisms.
2.22. Software and Reproducibility
The analysis was conducted in Python (v3.11) using pandas (v2.2.2), NumPy (v2.0.2), scikit-learn (v1.5.2), imbalanced-learn (v0.12.4), SciPy (v1.14.1), statsmodels (0.14.2), Matplotlib (v3.9.2), and related scientific-computing libraries. Random seeds were fixed for reproducibility. The fully executed notebook, fitted final model pipeline, reproducibility manifest, and selected machine-readable result and audit files are provided as
Supplementary Materials. NFHS-5 individual-level microdata are not publicly distributed and must be obtained directly from the Demographic and Health Surveys Program under its applicable data-use conditions.
2.23. Ethical Considerations
The study used de-identified secondary data obtained through the Demographic and Health Surveys programme. No primary data were collected, and no participants were contacted by the authors. Access to the NFHS-5 microdata was subject to DHS data-use approval. The analysis therefore involved no additional intervention or direct risk to participants.
4. Discussion
4.1. Principal Findings
This study developed and evaluated a leakage-controlled, household-grouped, and survey-aware framework for neonatal mortality prediction using NFHS-5 data from Bihar, Chhattisgarh, and Uttarakhand. The principal finding was that histogram gradient boosting provided the strongest overall predictive performance among the evaluated learners. On the untouched final-test set, the model achieved moderate discrimination, with a ROC-AUC of 0.755 and a PR-AUC of 0.202, while retaining high sensitivity and negative predictive value. However, the positive predictive value remained low, reflecting the rarity of neonatal death and the substantial number of false-positive alerts generated at the selected screening threshold. The controlled ablation analysis further showed that socioeconomic clustering, PSO feature selection, SMOTE, and out-of-fold feature augmentation provided little independent improvement beyond simpler specifications. Performance also varied according to the time at which predictor information became available and across the three held-out states. These findings indicate that the main contribution of the study lies in rigorous validation and transparent assessment of model utility rather than in the superiority of a complex augmented architecture.
4.2. Selected Model’s Predictive Performance
The selected histogram gradient-boosting model demonstrated moderate ability to distinguish neonatal deaths from survivors. The ROC-AUC indicated acceptable ranking performance, while the PR-AUC provided a more conservative and relevant assessment given the low prevalence of neonatal death. The difference between these two measures is important. The ROC-AUC can remain relatively favourable in highly imbalanced datasets because it is influenced strongly by the large number of correctly ranked negative observations. The PR-AUC, in contrast, reflects the balance between detecting neonatal deaths and avoiding false-positive alerts, and therefore provides a more realistic indication of performance for rare-event screening.
The high sensitivity suggests that the selected threshold identified most neonatal deaths in the final-test set. Similarly, the very high negative predictive value indicates that observations classified as low risk were unlikely to experience neonatal death. These properties may be useful in a preliminary screening context where the primary aim is to avoid missing potentially high-risk cases. However, the low positive predictive value means that most high-risk classifications were false-positive alerts. Consequently, the model should not be interpreted as a diagnostic tool, nor should its output independently trigger treatment, referral, or resource-intensive intervention. Its most plausible role would be to support low-cost secondary review, additional counselling, or closer observation within an existing maternal and neonatal care pathway.
The calibration results were broadly acceptable, although not ideal. The calibration slope was reasonably close to one, and the expected calibration error was low, indicating that predicted probabilities were generally aligned with observed outcomes. Nevertheless, the calibration intercept was imprecisely estimated, and the confidence interval included values consistent with some degree of systematic overprediction or underprediction. This reinforces the need for recalibration before implementation in settings that differ from the development population.
4.3. Candidate Learners–Comparisons
Histogram gradient boosting outperformed logistic regression, random forest, and the artificial neural network under the same household-grouped validation protocol. This finding is consistent with the strengths of gradient-boosting methods for structured tabular data. Such models can represent nonlinear relationships, threshold effects, and interactions without requiring extensive manual specification, while generally being less data-demanding than deep neural networks.
The artificial neural network did not outperform the simpler alternatives. This is an important result because neural networks are sometimes assumed to be inherently superior due to their complexity. In the present study, the combination of a moderately sized tabular dataset, heterogeneous categorical and continuous variables, a rare outcome, and substantial missingness likely favoured tree-based boosting over a neural architecture. Neural networks can require larger effective sample sizes, careful tuning, and more positive outcome events to achieve stable performance. Their flexibility may also increase variance when the number of events is limited.
Random forest achieved a ROC-AUC comparable to histogram gradient boosting but a lower PR-AUC. This suggests that it ranked positive and negative observations reasonably well overall but was less effective at concentrating true neonatal deaths among the highest-risk predictions. Logistic regression provided a useful baseline and retained interpretability, but its linear specification was less able to capture the nonlinear patterns and interactions identified by boosting. The comparison therefore supports the use of histogram gradient boosting as the final predictive learner while also demonstrating why fair evaluation against simpler models is necessary.
4.4. Contribution of Clustering, PSO, SMOTE, and Feature Augmentation
The ablation analysis showed that the individual framework components contributed little additional predictive value. Socioeconomic clustering produced a negligible change in PR-AUC, suggesting that the original state, residence, wealth, and education variables already contained most of the information represented by the derived cluster assignments. Clustering may still offer descriptive or policy-oriented value by summarising socioeconomic profiles, but its contribution to predictive accuracy was minimal.
PSO feature selection produced a small numerical improvement, but the corrected confidence interval included zero. This finding suggests that PSO may have removed some redundant predictors without consistently improving out-of-sample performance. The result also illustrates that metaheuristic feature selection should not be assumed to improve prediction merely because it reduces dimensionality. Its value depends on whether the selected subset remains stable across folds and whether the removed variables are truly uninformative.
SMOTE did not improve performance after clustering and PSO. This result is plausible because synthetic oversampling can alter the empirical joint distribution of survey variables and may generate observations that do not correspond closely to realistic maternal or neonatal profiles. In addition, synthetic samples cannot be assigned natural DHS sampling weights. The limited benefit of SMOTE in this study suggests that survey-weighted non-resampled models may be preferable when the learner can already accommodate class imbalance.
Out-of-fold feature augmentation also failed to produce a clear gain. Although the predicted probabilities of preterm birth, low birth weight, failed pregnancy, and teenage or late pregnancy were generated without leakage, these intermediate predictions may not have added substantial information beyond the original variables. Their first-stage prediction errors may also have propagated into the final classifier. Overall, the ablation findings demonstrate that combining more components does not necessarily produce a better model. The strongest improvement came from the choice of learner rather than from additional clustering, feature selection, resampling, or augmentation.
4.5. Prediction at Antenatal, Delivery Time, and Retrospective Stages
The prediction-timepoint analysis showed that antenatal performance was notably weaker than delivery time and retrospective performance. This difference is clinically important because antenatal prediction is the most attractive setting for preventive action, yet it relies on a more restricted set of variables. Maternal characteristics, socioeconomic conditions, obstetric history, and antenatal care indicators provided some discriminatory information, but they were insufficient for strong rare-event prediction. Performance improved when delivery time variables became available. Gestational duration, birth weight, delivery setting, and related information are closely associated with neonatal survival and therefore increased discrimination. The delivery time model achieved performance close to the retrospective upper-bound model, suggesting that most of the useful predictive information was available by the time of childbirth or immediately thereafter. The retrospective model achieved the highest PR-AUC but should not be interpreted as an early-warning system because it included postnatal information. Its role is best understood as an upper-bound benchmark that quantifies the maximum performance achievable using the complete NFHS-5 feature set. These findings indicate that a practical implementation should distinguish clearly between antenatal risk screening and delivery time neonatal risk assessment. A delivery time model may be the more realistic candidate for initial operational testing, while antenatal prediction would require additional data sources or biomarkers to achieve stronger performance.
4.6. Cross-Geography Applicability
Leave-one-state-out validation demonstrated substantial heterogeneity in performance across Bihar, Chhattisgarh, and Uttarakhand. The strongest performance was observed when Chhattisgarh was held out, whereas the Bihar analysis showed particularly low specificity despite very high sensitivity. This indicates that the model generated a large number of false-positive alerts in Bihar. Several factors may explain this variation. The states differ in population size, socioeconomic composition, healthcare access, maternal characteristics, and outcome prevalence. Bihar also contributed the largest proportion of the analytical dataset and excluding it from the model development may have altered both the distribution of predictors and the calibration of predicted risks. The observed heterogeneity suggests that a single threshold is unlikely to perform equally well across all settings. These results do not demonstrate broad external generalisability. Rather, they provide internal geographic validation across three states. Before implementation, the model would require evaluation in additional Indian states and in more recent routine or facility-based data. State-specific recalibration and threshold selection may also be necessary to achieve an acceptable balance between sensitivity and false-positive burden.
4.7. Survey-Weighted Association Analysis
The survey-weighted association analysis identified positive associations between neonatal mortality and preterm birth, low birth weight, and facility delivery, while postnatal care for the child was negatively associated with mortality. The associations with preterm birth and low birth weight are consistent with their established roles as major indicators of neonatal vulnerability. Both conditions reflect biological immaturity, impaired foetal growth, and increased susceptibility to respiratory, other infectious, and metabolic complications. The positive association with facility delivery should not be interpreted as evidence that facility delivery increases neonatal mortality. Women with severe complications, high-risk pregnancies, obstructed labour, or foetal distress are more likely to deliver in health facilities. This process of referral and severity-related selection can create a positive association even when facility-based care is beneficial. The finding therefore likely reflects residual confounding by clinical severity and referral patterns.
The negative association with postnatal care is compatible with the role of early neonatal assessment in identifying complications and supporting timely management. However, the observational design prevents causal interpretation, and postnatal care uptake may also be associated with unmeasured factors such as maternal health literacy, facility access, and socioeconomic advantage. Other variables, including teenage or late pregnancy, failed pregnancy history, four or more antenatal care visits, and marriage before 18 years, had confidence intervals that included the null value after adjustment. Their lack of statistical precision does not imply that they are clinically unimportant. Rather, the associations may be mediated, confounded, measured imprecisely, or heterogeneous across subgroups. The regression findings should therefore be interpreted as adjusted associations within this specific survey population.
4.8. Implications for Neonatal Risk Screening and Decision Support
The findings support the potential use of machine learning as one component of neonatal risk screening but do not establish a deployable decision-support system. The selected model could theoretically be integrated into an HMIS or community health worker application to flag records for additional review. However, the low positive predictive value means that such integration would generate many false alerts. The downstream action must therefore be inexpensive, non-invasive, and unlikely to cause harm. A plausible workflow would involve automatic risk scoring at delivery, followed by human review of flagged cases, confirmation of key clinical information, and targeted observation or counselling. The model should not independently determine admission, treatment, or referral. Prospective implementation would also require monitoring of calibration drift, threshold performance, subgroup fairness, workflow burden, and unintended consequences.
4.9. Strengths of the Study
The study has several methodological strengths. Household-grouped splitting prevented records from the same household or mother from appearing in both development and test datasets. Repeated grouped nested cross-validation ensured that preprocessing, clustering, PSO, feature augmentation, resampling, calibration, and threshold selection were performed within the appropriate training data. The final-test set remained untouched until final evaluation. The study prioritised PR-AUC and reported a comprehensive set of rare-event metrics, including calibration and clinical utility measures. Household-grouped bootstrap confidence intervals quantified uncertainty while preserving within-household dependence. Controlled ablation allowed the contribution of each framework component to be assessed directly. Prediction timepoint analysis clarified the distinction between antenatal, delivery time, and retrospective use, while leave-one-state-out validation assessed geographic transportability. The survey-weighted association analysis also incorporated DHS weights, strata, and PSUs and was kept conceptually separate from predictive modelling.
4.10. Limitations
Several limitations should be acknowledged. First, NFHS-5 is an observational survey, and many variables rely on maternal recall or categorical reporting. Misclassification and recall error may therefore affect both predictors and outcomes. Second, neonatal death was rare, resulting in limited positive predictive value and wide uncertainty for some metrics. Third, although the model was evaluated using rigorous internal and geographic validation, no independent external dataset was available. The results therefore cannot establish generalizability to other Indian states, countries, time periods, or routine HMIS data. Fourth, antenatal prediction remained weak, limiting the usefulness of the model for early preventive intervention. Fifth, survey weights could be used directly in eligible non-resampled models, but synthetic resampling methods could not naturally preserve the complex survey design. Sixth, the penalised survey-weighted regression provided stable association estimates but is not equivalent to a conventional unpenalised design-based causal model. Seventh, households from the same PSU could appear across development and test partitions, leaving some potential contextual dependence. Finally, the study did not evaluate prospective implementation, cost-effectiveness, fairness across demographic groups, alert fatigue, workflow integration, or the effects on neonatal outcomes. These issues must be addressed before practical deployment.
4.11. Future Research
Future studies should validate the selected model prospectively using routine HMIS, facility, or community health worker data. Temporal validation and evaluation in additional Indian states are needed to assess transportability. State-specific recalibration and threshold selection should be examined, particularly for settings such as Bihar where specificity was low. Further work should also assess fairness across socioeconomic, geographic, ethnic, and maternal subgroups; evaluate the cost and workload associated with false-positive alerts; and involve healthcare workers in human-factors and usability testing. Additional predictors available in clinical settings, including laboratory measurements, physiological observations, and non-invasive approaches such as infant cry acoustic analysis, may improve antenatal or early neonatal prediction. Any future DSS should be evaluated prospectively for safety, clinical utility, and impact on service delivery before large-scale implementation.
5. Conclusions
This study developed and evaluated a household-grouped, leakage-controlled, and survey-aware framework for neonatal mortality risk prediction using NFHS-5 data from Bihar, Chhattisgarh, and Uttarakhand. Histogram gradient boosting achieved the strongest overall performance among the evaluated learners, with moderate discrimination, acceptable calibration, high sensitivity, and a very high negative predictive value on the untouched final-test set. However, the low positive predictive value indicated a substantial false-positive burden, limiting the model’s immediate use to low-cost risk screening rather than diagnosis, automatic referral, or treatment decision-making. The controlled ablation analysis showed that socioeconomic clustering, PSO feature selection, SMOTE, and out-of-fold feature augmentation provided little additional predictive benefit. Performance depended more strongly on the selected learner and on the time at which information became available. Delivery time prediction was more informative than antenatal prediction, while geographic validation showed meaningful variation across the three states, particularly in specificity. These findings support further development of a neonatal risk-screening framework but do not demonstrate causal effects, improved neonatal outcomes, or a fully implemented decision-support system. Prospective validation in routine clinical or HMIS settings, external evaluation across additional populations, state-specific recalibration, and assessment of workflow burden, fairness, safety, and cost-effectiveness are required before operational deployment.