Next Article in Journal
Global Fuzzy Adaptive Consensus for Uncertain Nonlinear Multi-Agent Systems with Unknown Control Directions
Previous Article in Journal
Constraint, Asymmetry, and Meaning: A Cybernetic Reinterpretation of Probabilistic Emergence Across Complex Systems
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Evaluation of a Non-Parametric Penalized Kaplan–Meier Estimator Under Interval-Censored Survival Data

by
Kayakazi Chophela
1,*,
Chioneso Show Marange
1 and
Akinwumi Sunday Odeyemi
1,2
1
Department of Computational Sciences, University of Fort Hare, Alice 5700, South Africa
2
SAMRC Microbial Water Quality Monitoring Centre, University of Fort Hare, Alice 5700, South Africa
*
Author to whom correspondence should be addressed.
Symmetry 2026, 18(3), 519; https://doi.org/10.3390/sym18030519
Submission received: 12 December 2025 / Revised: 23 January 2026 / Accepted: 13 February 2026 / Published: 18 March 2026
(This article belongs to the Section B: Mathematics)

Abstract

Interval-censored survival data arise frequently in biomedical and epidemiological studies where event times are observed only within observation intervals. Classical non-parametric estimators, such as the Kaplan–Meier (KM) estimator under imputation and the Turnbull estimator, often suffer from instability, irregular fluctuations, and overfitting when sample sizes are small or when the prevalence rate is low. Recent methodological developments, which include smoothed and penalized approaches, have been proposed to improve stability and reduce estimation error in such settings. This study evaluates and benchmarks the finite-sample performance of a nonparametric penalized likelihood KM estimator under interval-censored data. The method is compared with the classical KM estimator using four imputation strategies, that is, midpoint, regression, uniform, and multiple imputation. From a symmetry perspective, midpoint and uniform imputation preserve interval symmetry through deterministic and probabilistic mechanisms, respectively, whereas regression and multiple imputation intentionally introduce structural asymmetry to reflect data-driven risk heterogeneity and distributional uncertainty. To assess and benchmark the performance of the penalized KM estimator, an extensive Monte Carlo (MC) simulation study was conducted across varying sample sizes and prevalence rates using error-based metrics. The MC simulation results revealed that the nonparametric penalized KM estimator consistently outperforms the classical KM estimator in small samples across all prevalence rates. The gains are more pronounced under low prevalence rates where the penalized KM estimator is superior for small to relatively moderate samples of n 40–100. Among the imputation techniques, regression and multiple imputation generally exhibited superior performance. Real data application further confirms these findings, demonstrating that the nonparametric penalized KM estimator yields more stable and accurate survival curves than the classical KM estimator in small samples.

1. Introduction

Interval-censored data arise when a failure time, denoted by T, cannot be observed exactly, but is only known to lie within an interval determined by a sequence of examination times [1,2,3]. Thus, only the lower and upper bounds of the interval, L and U, are known, while the exact event time remains unknown. Mathematically, this is represented as L < T U , or in set notation, T ( L , U ] . A fundamental assumption in survival analysis is that the censoring mechanism is independent of, or non-informative about, the failure time of interest. Formally, for an interval I = ( L , U ] ,
P ( T t L = l , U = u , L < T U ) = P ( T t l < T u ) ,
which implies that the interval ( L , U ] (or equivalently, its endpoints) does not provide further information about T beyond the fact that the event occurs within the observed bounds. Such data often arise in biomedical and reliability studies, especially in clinical trials and longitudinal research with periodic follow-up visits [1,3,4]. In many practical settings, the analysis of interval-censored survival data requires estimation of the underlying survival function [1,5,6,7,8]. Treating interval-censored survival times as exact event times can lead to biased estimates and, consequently, results and conclusions that lack reliability [9,10,11]. Because the exact event times are unobserved, interval-censored data may naturally be viewed as a missing-data problem. As a result, a substantial body of research has focused on addressing this challenge through the development of imputation techniques and nonparametric procedures for the estimation of the survival function under interval censoring.
Classical nonparametric maximum likelihood estimators (NPMLEs), including the Kaplan-Meier (KM) estimator for right-censored data [12] and Turnbull (TB) for interval-censored data [5], form the basis of modern survival analysis. Despite their wide adoption, these methods face well-documented limitations when applied to interval-censored data under finite-sample conditions [13,14]. Since the exact time is unknown, applying the classical KM estimator on interval censored data is not straightforward, as it requires exact event times. To address this, several imputation-based approaches [15], such as midpoint [16], regression, uniform, and multiple imputation [17], are frequently used to convert interval-censored data into imputed event times before applying the KM estimator. Midpoint and uniform imputation preserve interval symmetry through deterministic and probabilistic mechanisms, respectively, while regression and multiple imputation incorporate structural asymmetry to account for data-driven risk heterogeneity and distributional uncertainty. However, these imputation-based KM methods exhibit considerable instability in small samples. They often produce irregular survival curves and suffer from overfitting, especially when the censoring intervals are wide or the prevalence of events is low [1]. Similarly, the classical TB, though appropriate for interval-censored data, tends to generate irregular multi-modal survival functions due to the inherent undersmoothing [18]. These limitations can result in inflated estimation errors and poor representation of the underlying survival process.
In general, NPMLEs, which includes the KM and the TB estimators for incomplete data [19], commonly suffer from overfitting when applied to small samples, which can lead to poor estimation performance [20]. Several approaches have been proposed to address this limitation, including semiparametric frameworks based on finite mixture models [21], spline-based methods [22,23,24], and smoothing techniques. Despite their flexibility, semiparametric procedures often face challenges related to parameter tuning, computational complexity, and nonconvergence, whereas smoothing methods are generally regarded as more practical. Kernel density estimation is a fundamental nonparametric smoothing technique that represents each observation as a localized “pump” of density with a common functional form [25,26,27]. Using these kernel densities, the extent of smoothing is primarily determined by the choice of kernel and the window width, commonly referred to as the bandwidth. In the context of the KM and TB estimators, applying smoothing provides a natural and straightforward approach for reducing excessive variability of the density estimate. More broadly, improving estimation efficiency under limited or incomplete data has received considerable attention across different statistical frameworks. For example, recent studies have shown that carefully designed sampling schemes, such as ranked set sampling and its extensions, can substantially enhance the finite-sample performance of parameter estimators for parametric distributions [28,29,30]. These developments highlight a growing emphasis on methodological strategies that mitigate small-sample inefficiency, which also motivates the present focus on smoothing and imputation-based approaches in nonparametric survival analysis.
Focusing on the classical KM estimator, several researchers have proposed smoothing techniques for KM estimation. Wang [31] considered kernel smoothing approaches to reduce variability in the KM curve, Efron [32] developed penalized likelihood formulations, while Babyak [33] proposed constrained estimators to control overfitting. On the other hand, recently, Wu and Kolassa [34] proposed a non-parametric approach to reduce the overestimation of the KM estimator when the event and censoring times are independent. These approaches share a common goal of reducing overfitting while preserving the nonparametric flexibility of the KM estimator. A notable contribution is the smoothed nonparametric penalized likelihood KM estimator proposed by Tubbs et al. [35]. This estimator applies kernel smoothing and a Bayesian Information Criterion (BIC)-type penalty to reduce overfitting by controlling curve complexity. The simulation results of Tubbs et al. [35] have shown that this estimator has demonstrated substantial improvements in reducing the estimation error and producing smoother and more reliable survival curves, especially when the sample sizes are small. Their proposed procedure advances nonparametric survival analysis by improving estimation accuracy and predictive performance, particularly in settings where survival data are constrained by censoring, small sample sizes, low prevalence, or a limited number of repeated measurements per subject. Tubbs et al. [35] demonstrated that smoothing NPMLEs through complexity-based penalized likelihood provides a simple yet effective strategy for enhancing survival and density estimation, as well as for improving exact time-to-event prediction through imputation.
Motivated by these developments, this study addresses a critical gap in the literature, namely, the absence of a systematic comparison between the improved nonparametric smoothed penalized KM estimator proposed by Tubbs et al. [35] and the classical KM estimator implemented using commonly applied imputation strategies under interval-censored settings. Through a comprehensive Monte Carlo (MC) simulation study, complemented by a convergence analysis, this work aims not only to provide clear evidence on the conditions under which smoothing enhances estimator performance, mitigates overfitting, and yields more reliable survival inference but also to identify the actual sample sizes at which the smoothed KM estimator demonstrably outperforms the classical KM. The findings are further validated using a real-world interval-censored dataset. Such insights are particularly valuable for practitioners analyzing interval-censored survival data, especially in applications where sample sizes are small and estimation stability is of primary importance. Accordingly, this study focuses exclusively on the KM estimator, given its widespread use and interpretability in survival analysis [12], as well as its central role in imputation-based and smoothing methodologies [31,32,35].
The remainder of this article is organized as follows. Section 2 presents the methodological framework, including the classical KM estimator, the smoothed penalized likelihood KM estimator, and the four imputation strategies examined in this study. Section 3 outlines the MC simulation design and describes the metrics used to evaluate the performance of the estimators under varying sample sizes and prevalence rates. It also presents the application of the proposed methods to a real interval-censored dataset. Finally, Section 4 provides concluding remarks and summarizes key insights from this study.

2. Materials and Methods

2.1. Kaplan–Meier Estimator

The KM estimator, introduced by Kaplan and Meier [12], is one of the most widely used non-parametric maximum likelihood estimators for the survival function from censored time-to-event data [36]. The KM estimator provides a product limit formulation that efficiently handles right-censored data without requiring strong distributional assumptions. The KM estimator was originally developed in the context of biomedical follow-up studies, where individuals can leave the study or remain event-free at the end of observation, resulting in right-censored event times. Although primarily intended for right-censored data, the same mathematical form also accommodates left-censored data with appropriate modifications. The KM estimator of the survival function is
S ^ ( t ) = t ( j ) t 1 d j n j ,
where t ( 1 ) < t ( 2 ) < < t ( J ) denote the ordered distinct event times, d j the number of events at t ( j ) , and n j the number at risk immediately prior to t ( j ) with S ^ ( 0 ) = 1 . This produces a step function that decreases only at observed event times while remaining constant between these times. The KM estimator is not, in its original form, designed for interval-censored data, in which the exact event time T i * is unknown but is known to lie within an interval ( L i , R i ] with 0 < L i < R i < (or R i = for right-censoring).
To apply KM estimation in the presence of interval censoring, analysts typically convert the interval-censored data into imputed event times via imputation. Common imputation methods include midpoint, uniform, regression and multiple imputation, each replacing ( L i , R i ] with a single representative value T i * and treating it as the actual event time. Once imputed, the KM estimator can be applied directly. This approach allows for comparison of KM-based curves across imputation methods, though it necessarily introduces approximation error relative to estimators specifically designed for interval-censored data, for example, the TB estimator proposed by Turnbull [5].

2.2. Non-Parametric Penalized Kaplan–Meier Estimator

To address the shortcomings of the classical KM estimator, several researchers have proposed smoothing techniques for KM estimation. Wang [31] considered kernel smoothing approaches to reduce variability in the KM curve, while Efron [32] developed penalized likelihood formulations. Babyak [33] further introduced constrained estimators to control overfitting. These approaches share a common goal of reducing overfitting while preserving the nonparametric flexibility of the KM estimator.
Building on this literature, Tubbs et al. [35] introduced the smoothed nonparametric penalized KM estimator, which applies a kernel-smoothing adjustment to the classical KM estimator and incorporates a penalized likelihood criterion to balance goodness-of-fit with curve smoothness. The smoothing step is implemented using Nadaraya–Watson normal kernel regression [37], with the degree of smoothing selected adaptively by minimizing a BIC-type penalty. Formally, the penalized objective is
BIC s ( d ) = 2 ln L ( d ) + k T ( d ) × ln N s ,
where L ( d ) is the likelihood of the smoothed survival curve with bandwidth d and k T ( d ) denotes the number of turning points in the estimated density function, which serves as a measure of the complexity of the curve. Specifically,
k T = card x R | sign f ^ n ( x + ) sign f ^ n ( x ) ,
where f ^ n ( x + ) and f ^ n ( x ) denote the right and left derivatives of the estimated density at x. The term N s reflects the effective information about the sample, with three alternative penalties considered: (i) ln N , based on the sample size; (ii) ln N m , the total number of repeated observations across individuals, and (iii) ln N e , the equivalent sample size adjusted for interval information. Simulation studies by Tubbs et al. [35] showed that although all three penalties performed comparably, the ln N e penalty generally provided marginally better results and is recommended for practice.
An important advantage of Tubbs et al.’s [35] estimator is its broad applicability across different censoring mechanisms. Thus, the smoothed nonparametric penalized KM method can be extended to accommodate left-censored, right-censored, and interval-censored observations. This makes it particularly valuable in biomedical and epidemiological studies, where mixed forms of censoring commonly arise due to staggered entry, delayed detection, or discrete follow-up schedules. Extensive simulations conducted by Tubbs et al. [35] demonstrated that their approach consistently reduced error between estimated and true survival functions by nearly half compared to classical KM estimators. Improvements were evident in both within-sample fit and out-of-sample prediction, with applications to breast cancer data further confirming its practical advantages. The smoothed nonparametric penalized KM method offers a powerful and stable alternative to the classical KM, particularly in small sample settings [35].

2.3. Imputation Methods for Interval Censoring

In this study, imputation is used to transform each interval-censored observation ( L i , R i ] into a pseudo-exact estimate T i * , upon which the KM estimator is applied. The imputation methods considered are midpoint, regression, uniform, and multiple imputation. These approaches differ in how they utilize the information contained in the censoring interval and auxiliary variables and thus influence the accuracy and stability of the resulting survival estimates. Midpoint and uniform imputation preserve the symmetry of the interval through deterministic and probabilistic mechanisms, respectively, while regression and multiple imputation incorporate structural asymmetry to account for data-driven risk heterogeneity and distributional uncertainty.

2.3.1. Midpoint Imputation

Midpoint imputation is a straightforward method in which missing or interval-censored values are imputed using the midpoint of the observed interval. It assumes that the event is equally likely to occur in the middle within the interval. This approach is simple to implement and is often used in preliminary survival analyses, reliability testing, and clinical studies when no prior information about the distribution of event times is available. This imputation method imposes interval symmetry by assigning equal weight to the lower and upper bounds, implicitly assuming a symmetric likelihood of event occurrence around the center of the interval. Due to its computational simplicity, midpoint imputation is frequently used in preliminary survival analysis, reliability studies, and clinical applications where there is no prior information on the underlying event-time distribution or potential asymmetry. Although easy to apply, midpoint imputation can introduce bias, especially when the width of censoring intervals varies significantly. Sun [1] noted that this method may oversimplify the data and affect the precision of survival estimates. Mathematically, for an interval ( L i , R i ] , the midpoint is given by
T i * ^ = L i + R i 2 ,
where L i and R i are the lower and upper bounds of the interval [1].

2.3.2. Regression Imputation

Regression imputation is a model-based technique in which unobserved event times are replaced by predictions from a fitted regression model. The method originates from early predictive imputation work by Buck [38], who first proposed using linear regression to replace missing survey values; it was later formalized in the modern missing-data framework by Little and Rubin [39]. In survival analysis with interval-censored data, regression imputation uses covariate information to generate interval-consistent event times that align with the censoring interval of each subject. Let T denote the event time and X a vector of p covariates. Suppose that the conditional distribution of T given X = x is characterized by density f ( t x ; θ ) and CDF F ( t x ; θ ) . Under a regression model, a principled approach is to draw the imputed time from the truncated conditional distribution as
T i * f ( t X i = x i ; θ ) truncated to ( L i , R i ] ,
with truncated density
f trunc ( t L i < t R i , x i ) = f ( t x i ; θ ) F ( R i x i ; θ ) F ( L i x i ; θ ) .
A deterministic alternative imputes the conditional expectation of T i , given the truncation
E T i * L i < T i * R i , X i = x i = L i R i t f ( t x i ; θ ) d t F ( R i x i ; θ ) F ( L i x i ; θ ) .
A widely used special case assumes a transformed linear regression model
g ( T i * ) = β 0 + β 1 X i 1 + + β p X i p + ε i , ε i N ( 0 , σ 2 ) ,
where g ( · ) is typically the identity or logarithm. Estimators β ^ and σ ^ 2 are obtained by regressing pseudo-exact times (e.g., midpoints) on the covariates. The deterministic regression imputed time is then given by
t ^ i = g 1 β ^ 0 + β ^ 1 X i 1 + + β ^ p X i p ,
while the stochastic version adds residual uncertainty,
T i * = g 1 β ^ 0 + β ^ 1 X i 1 + + β ^ p X i p + σ ^ Z i , Z i N ( 0 , 1 ) .
Finally, the imputed value is truncated to satisfy the censoring interval
T i Reg min { max ( T i * , L i ) , R i } .
In this study, regression imputation was implemented using the following procedure:
  • A pseudo-exact time T i * was initially created for each interval-censored observation using the midpoint of ( L i , R i ] .
  • The Weibull accelerated failure time (AFT) model was estimated using the survreg() function from the survival package in R. The survreg() function fits parametric survival regression models under the AFT framework by maximizing the full likelihood, accommodating right-, left-, and interval-censored survival data. Specifically, survreg() assumes that the logarithm of the survival time follows a linear regression model with an additive error term whose distribution determines the parametric form of the survival model. For the Weibull specification, the error term is assumed to follow a standard extreme value (Gumbel) distribution, which implies a Weibull distribution for the survival time T i . Model parameters, including the regression coefficients β , the intercept β 0 , and the scale parameter σ , are estimated via maximum likelihood.
    The Weibull shape parameter is given by k = 1 / σ and the scale parameter is λ = exp ( β 0 ) , linking the AFT parameterization used by survreg() to the conventional Weibull distribution. This flexibility, combined with native support for interval-censored data, makes survreg() well suited for regression-based imputation and survival modeling in interval-censored settings.
  • The fitted model produced predicted event times
    t ^ i = exp ( β ^ 0 + β ^ X i ) ,
    which were then truncated to the subject’s interval:
    t ^ i min { max ( t ^ i , L i ) , R i } .
  • In instances where the Weibull AFT model failed to converge, typically due to limited data variability, sparse failures, or instability in the likelihood, the midpoint of the interval was used as a robust fallback imputation:
    T i Reg = L i + R i 2 .
  • The final regression imputed times T i Reg were analyzed using both the classical KM estimator and the smoothed penalized-likelihood KM estimator.
Regression imputation is known to perform well when covariates are predictive of event time. This approach deliberately departs from left–right interval symmetry by allowing imputed values to be asymmetrically distributed within the censoring interval, reflecting covariate-driven risk heterogeneity and the underlying hazard structure. By accommodating systematic asymmetry in event–time dynamics, regression imputation provides a more realistic representation of the data-generating process, particularly in settings where proportional hazards or distributional skewness are present. In a comparison of imputation methods for doubly censored human immunodeficiency virus (HIV) data, ref. [40] found that regression imputation consistently produced the lowest mean squared error (MSE) for survival estimation, outperforming midpoint and even multiple imputation in several scenarios. Similarly, Zhang and Sun [41] reported that regression-based imputation was markedly more robust than simple random (uniform) or midpoint-based approaches, particularly as the width of the interval increased. These findings support the use of regression imputation as a strong model-based imputation strategy in interval-censored survival analysis and motivated its inclusion in this study.

2.3.3. Uniform Imputation

Uniform imputation is one of the simplest model-free approaches for handling interval-censored survival data. This approach enforces probabilistic allocation by assuming that all points within the interval are equally likely, thereby preserving symmetry with respect to the interval bounds. The idea of replacing an unobserved event time with a random draw from the subject’s censoring interval has been used for decades in simulation-based studies of interval-censoring and missing data and was formally described in the context of multiple imputation by Hsu et al. [42]. Under the uniform imputation approach, the event time is assumed to follow a continuous uniform distribution over the interval
T i * Uniform ( L i , R i ) .
The corresponding density is
f Uni ( t L i , R i ) = 1 R i L i 1 { L i < t R i } ,
which assigns equal probability to all possible event times within the observed interval.
In this study, for each individual with an observed event ( δ i = 1 ), the event time was imputed by sampling a random value from the interval ( L i , R i ] using (1). For right-censored individuals ( δ i = 0 , R i = ), no imputation was performed; thus, the observed censoring time L i was retained. A new uniform draw was generated for each bootstrap repetition and each MC dataset to ensure that the induced uncertainty was fully propagated through the performance evaluation metrics. The resulting interval-censored imputed event times were then analyzed using both the classical KM estimator and the improved smoothed penalized-likelihood KM estimator.
Uniform imputation is attractive because it does not assume any specific distribution for T i * beyond the observed interval, and its results are immediate and do not require numerical optimization. However, the method implicitly assumes that the event is equally likely to occur at any point within ( L i , R i ] , which may not hold when the true event–time distribution is skewed or when covariates influence the hazard. As a result, uniform imputation can introduce additional variability or bias, particularly when intervals are wide.
Uniform imputation has frequently been used as a neutral baseline method in studies of interval-censored and doubly censored data. Hsu et al. [42] showed that while uniform imputation performs reasonably well under narrow intervals, it is generally less efficient than probability-based or model-based imputation methods such as regression imputation. Uniform imputation remains valuable as a distribution-free approach, particularly in simulation work, because it isolates the effect of the censoring structure without relying on parametric or semiparametric modeling assumptions. For this reason, it provides an important benchmark against which more sophisticated methods, including regression imputation and multiple imputation, can be evaluated within the context of the improved smoothed penalized likelihood KM estimator considered in this study.

2.3.4. Multiple Imputation

Formally introduced by Rubin [43], multiple imputation is a statistical technique for handling missing or partially observed data by creating several completed datasets and combining their results to capture the uncertainty arising from missing data. In survival analysis, multiple imputation is particularly useful when event times are censored or incompletely observed, as it allows one to incorporate uncertainty about the true event time while enabling analysis with standard complete-data methods.
A total of m = 5 imputed datasets were generated, consistent with standard practice for simulation-based survival analysis and Rubin’s original recommendations. Each imputed dataset contained a complete set of pseudo-exact event times. Survival estimation was then performed separately on each imputed dataset using the KM estimator. The resulting survival estimates were combined across imputations by averaging the estimated survival functions at each time point, thereby accounting for both within-imputation variability and between-imputation uncertainty according to Rubin’s combining rules.
Let Y = ( Y obs , Y mis ) denote the complete dataset, where Y mis represents the missing or censored quantities. Multiple imputation proceeds in three main steps:
  • Generate M independent completed datasets:
    Y ( m ) = Y obs , Y mis ( m ) , m = 1 , , M ,
    where each Y mis ( m ) is drawn from an approximation to the posterior predictive distribution
    Y mis ( m ) p ( Y mis Y obs ) .
  • Fit the desired model separately to each imputed dataset, obtaining point estimates θ ^ ( m ) and the corresponding variance estimates U ^ ( m ) .
  • Combine M results using Rubin’s rules:
    θ ¯ = 1 M m = 1 M θ ^ ( m ) , U ¯ = 1 M m = 1 M U ^ ( m ) , B = 1 M 1 m = 1 M θ ^ ( m ) θ ¯ 2 , T = U ¯ + 1 + 1 M B ,
    where θ ¯ is the pooled point estimate, U ¯ represents the average within-imputation variance, B represents the between-imputation variance, and T represents the total variance incorporating both within- and between-imputation uncertainty.
Multiple imputation is widely recognized for its ability to improve efficiency and reduce bias when compared with single imputation methods. In interval-censored settings, Geskus [40] and Zhang and Sun [41] demonstrated that multiple imputation generally outperforms midpoint and uniform imputation, especially when intervals are wide or covariates strongly influence failure times. Although regression-based imputation may achieve lower MSE in some situations, multiple imputation remains uniquely advantageous in fully representing imputation uncertainty.
Multiple imputation relaxes the assumption of interval symmetry by allowing asymmetric allocations within the censoring interval, thereby accommodating data-driven risk heterogeneity and potential skewness in the underlying event-time distribution. By propagating both within-imputation and between-imputation variability, this approach provides more robust inference in the presence of structural asymmetry and uncertainty and is widely used in survival and epidemiological studies where a realistic representation of event-time dynamics is required.

2.4. Performance Evaluation Metrics

Let S ( t ) denote the true survival function used to generate the reference curve in the simulation study and let S ^ ( t ) represent the estimated survival function obtained after applying a given imputation method followed by KM estimation. Performance is assessed by comparing S ^ ( t ) and S ( t ) over a common discrete time grid { t i } i = 1 n . The pointwise estimation error at time t i is defined as
e i = S ( t i ) S ^ ( t i ) , i = 1 , , n .
This curve-based evaluation framework is appropriate when the interest lies in the overall accuracy of the estimated survival function rather than a single scalar parameter, and it is commonly adopted in simulation-based survival studies.

2.4.1. Mean Absolute Error (MAE)

The MAE measures the average absolute deviation between the true and predicted survival functions and is defined as
MAE = 1 n i = 1 n S ( t i ) S ^ ( t i ) .
MAE is the most frequently reported error-based metric because it offers direct measures of prediction error [44,45,46]. Zhou et al. [47] reported that MAE is among the most frequently used performance evaluation metrics for survival models (see [48,49,50,51], among others). It has been advocated as a meaningful accuracy measure for survival model evaluation [52], implemented in survival evaluation software [47], and been widely used in censored-data imputation and prediction studies [46,53]. In survival analysis, MAE directly measures the accuracy of predicted time-to-event outcomes [52], making it an important metric for evaluating the precision of survival models [53]. Compared with squared-error measures, MAE does not excessively penalize large deviations, making it less sensitive to isolated extreme errors. Thus, compared with squared-error measures, MAE does not excessively penalize large deviations, making it less sensitive to isolated extreme errors.

2.4.2. Normalized Mean Absolute Error (NMAE)

To facilitate scale-free comparison across simulation scenarios with different survival ranges or sample sizes, the NMAE is also used:
NMAE = 1 n range ( S ) i = 1 n S ( t i ) S ^ ( t i ) ,
where range ( S ) = max i S ( t i ) min i S ( t i ) . This normalization produces a dimensionless metric that enables meaningful comparison across datasets with differing scales or censoring configurations. The use of normalized absolute-error measures is motivated by the robustness and interpretability advantages of MAE over squared-error metrics [44] and extends the widespread application of MAE in survival model evaluation [46,52].

2.4.3. Maximum Absolute Error (MaxAE)

Although MAE summarizes average performance, it may fail to reflect extreme local discrepancies between the estimated and true survival functions. To explicitly capture the worst-case deviation, the MaxAE is considered and defined as
MaxAE = max i = 1 , , n S ^ ( t i ) S ( t i ) .
MaxAE measures the largest pointwise error over the evaluation grid and is therefore highly sensitive to extreme deviations. While it is not robust to outliers, MaxAE is useful as a diagnostic metric for identifying local instability or severe mismatches in estimated survival curves. Such worst-case error measures are particularly informative in censored-data imputation settings, where estimation uncertainty may be unevenly distributed across time [46].

2.4.4. Mean Squared Error (MSE)

The MSE averages the squared deviations between estimated and true survival probabilities and is defined as
MSE = 1 n i = 1 n S ( t i ) S ^ ( t i ) 2 .
By squaring the errors prior to averaging, MSE assigns greater weight to large discrepancies and is therefore sensitive to pronounced local instability in the estimated survival curve [44,45]. In survival analysis, MSE is typically reported alongside MAE to jointly assess calibration and precision, particularly in comparative studies of parametric and penalized survival models [53]. Similar to MAE, the lower the MSE, the higher the accuracy of prediction [47].

2.4.5. Normalized Root Mean Squared Error (NRMSE)

To retain the sensitivity of squared-error measures while enabling scale-independent interpretation, the NRMSE is defined as
NRMSE = 1 range ( S ) 1 n i = 1 n S ( t i ) S ^ ( t i ) 2 .
NRMSE generalizes the commonly used RMSE by scaling with the range of the true survival function. RMSE-based measures have been employed to assess estimator accuracy and stability in penalized nonparametric likelihood methods for arbitrarily censored data [22,35] as well as in interval-censored survival smoothing and estimation studies [22].

2.4.6. Normalized Average Root Median Error Deviation (NARMED)

Mean-based metrics such as MAE and MSE may still be influenced by extreme pointwise deviations, particularly in regions of sparse information such as the tail of the survival curve. To provide a robust complement, the NARMED is considered:
NARMED = 1 range ( S ) median S ( t i ) S ^ ( t i ) .
The NARMED builds on the robustness properties of median-based absolute error measures, which are well known to reduce sensitivity to outliers and heavy-tailed error distributions [44]. By combining median-based error assessment with scale normalization, NARMED provides a stable assessment of typical estimation error and is particularly suitable for interval-censored survival data [46].

3. Results

3.1. Simulation Study

3.1.1. Monte Carlo Simulation Design

To evaluate and benchmark the performance of the smoothed nonparametric penalized likelihood KM estimator under various imputation strategies for interval-censored data, a comprehensive MC simulation framework was conducted. As noted by Efron [54], as well as Robert and Casella [55], MC methods offer a flexible and reproducible framework for evaluating statistical procedures under controlled, yet realistic conditions. A key advantage of this simulation approach is the ability to compare estimated survival curves obtained using the classical KM and smoothed penalized likelihood KM estimates against the known true survival estimate. This allows for objective evaluation of the accuracy, bias, variance, and consistency of the estimators under different imputation methods. By performing more than thousands of replications, the simulation provides empirical evidence on the expected performance of each method in controlled settings, where the true event times are known but are treated as unobservable [1]. The MC simulation framework therefore supports both theoretical and practical evaluation of statistical methods for interval-censored data. The study systematically varied key design parameters, such as sample size and prevalence rate, to assess robustness and finite-sample behavior. The classical KM estimator served as a benchmark, providing a baseline for comparison with the smoothed nonparametric penalized likelihood KM estimator. Four imputation strategies, namely, midpoint, regression-based, uniform, and multiple imputation, were examined under both estimators to evaluate how imputation choice influences survival estimation performance.
To obtain accurate and stable performance assessments, each simulation scenario was evaluated using an MC bootstrap framework. For every combination of sample size ( n ) and prevalence rate ( p ) , m = 5000 bootstrap samples were drawn from a large synthetic population of size N = 500 , 000 . The use of a large synthetic population and a high number of bootstrap replications was motivated by the need to reduce MC variability and obtain stable estimates of error-based performance measures, particularly under small sample sizes and low prevalence rate. In such settings, performance evaluation metrics can exhibit substantial variability, and insufficient bootstrap replication may obscure meaningful differences between competing methods. The large synthetic population provides an empirical approximation to the underlying data-generating mechanism, thereby limiting finite-population effects when repeatedly sampling across ( n , p ) combinations. This is especially important for tail-sensitive measures such as MaxAE and robust, median-based metrics such as NARMED, whose convergence may be slower under interval censoring. Each bootstrap sample was subjected to the specified imputation procedure, followed by estimation of the KM survival curve. The resulting survival estimates were compared with the known true survival curve at multiple time points. The final performance scores for each method were obtained by averaging across all bootstrap replications, ensuring a reliable and consistent comparative evaluation. Performance was assessed using six complementary metrics: MAE, NMAE, MaxAE, MSE, NRMSE, and NARMED. Together, these measures provide a comprehensive assessment of bias, variability, robustness, and overall estimation efficiency.

3.1.2. Simulation of Interval-Censored Data

The simulation procedure involved generating synthetic event times and observation intervals for a large number of individuals, from which bootstrap samples were repeatedly drawn for analysis. The key components of the simulation setup were as follows:
  • The population size based on a total of N = 500 , 000 observations was initially generated to serve as the source population.
  • Random samples of size n { 10 , 20 , 40 , , 500 } were drawn from the population in each iteration.
  • For each sample size, m = 5000 bootstrap replicates were generated to ensure stability and precision of the performance estimates.
  • The different prevalence rates ( p ) { 0.2 , 0.3 , 0.4 , 0.6 , 0.8 , 1.0 } were applied.
The true event times, denoted by T i * , were independently simulated from a continuous uniform distribution over the interval ( 0 , 24 ] in a way that T i * U n i f o r m ( 0 , 24 ) , i = 1 , 2 , 3 , , N , where N = 500 , 000 determined a sufficiently large population size and was used to allow reliable bootstrapping and robust estimation of empirical survival distributions. The choice of a uniform distribution is common in interval-censored survival simulation studies, as it avoids temporal bias and ensures that all time points are equally likely [1,18].
Each individual was assumed to be observed at regularly spaced follow-up times (e.g., every month), though the specific time unit could vary depending on the context (e.g., daily, weekly, monthly, or annually), such that V = { 1 , 2 , 3 , , 24 } led to interval-censored data. For each simulated event time, the corresponding observation interval was determined such that L i < T i * R i , L i , R i V . For each subject, the actual event time T i * was not always directly observed but was captured through an interval ( L i , R i ] . Each observation was assigned a binary indicator that was then drawn from a Bernoulli distribution, i.e, δ i B e r n o u l l i ( p ) , where δ i = 1 denoted that the event occurred within the interval and δ i = 0 indicated right-censoring. In this study, the event times were subject to Case II interval censoring, where each participant was assessed at multiple scheduled visits, and it was known that the exact event time only fell between two successive observation points.

3.1.3. Generation of True Survival Estimates

In interval-censored settings, exact event times are unobserved, making analytical derivation of the true survival function infeasible. Therefore, a simulation-based benchmark approach was adopted. The procedure used to approximate the true survival function under interval censoring was carried out as follows:
  • A population of size N = 500 , 000 was used as the source population from which bootstrap samples were drawn.
  • For each bootstrap iteration m = 1 , 2 , , M , with M = 5000 , a random sample of size n was drawn from the population with a replacement.
  • For each bootstrap sample and each imputation method, a KM estimator S ^ m ( t ) was computed based on the imputed event times.
  • To reduce sampling variability and obtain a continuous survival estimate, each KM curve S ^ m ( t ) was smoothed using cubic spline interpolation.
  • The smoothed survival curves were evaluated on a dense and fixed time grid:
    T = { 0 , 0.25 , 0.5 , , 24 } .
  • The final reference (true) survival function was obtained by averaging the smoothed survival estimates across all bootstrap samples:
    S ¯ ( t ) = 1 M m = 1 M S ^ m ( t ) ,
    where S ^ m ( t ) denotes the smoothed KM survival probability at time t from the m th bootstrap replicate.
  • The resulting averaged survival function S ¯ ( t ) was treated as the true survival curve and used as the benchmark for evaluating estimator performance.
This bootstrap-based smoothing approach is widely used in survival simulation studies and has been shown to provide stable and reliable reference survival functions when assessing imputation strategies under interval censoring [35,54,56].

3.1.4. Computation of Performance Evaluation Metrics

To obtain accurate and stable performance assessments under a variety of data conditions, each imputation method was evaluated using an MC bootstrap framework. For every combination of sample size ( n ) and prevalence rate ( p ) , m = 5000 bootstrap samples were drawn from the full synthetic population of N = 500 , 000 . The samples were then subjected to the imputation procedure; this was followed by estimation of the KM survival curve. The resulting survival curves were compared with the known true survival curve at multiple time points, and the comparison was quantified using six standard performance evaluation metrics, which included MAE, NMAE, MaxAE, MSE, NRMSE, and NRMED. The final performance scores for each method were averaged over all bootstrap replicates, ensuring reliability and consistency in the comparative evaluation.

3.2. Monte Carlo Simulation Results

The results are presented in three stages. First, we provide a descriptive analysis of the simulated data. This is followed by an MC performance metrics evaluation of the KM-based estimators. Finally, a convergence validation analysis is conducted to quantitatively benchmark the performance of the smoothed KM estimator against the classical KM estimator.

3.2.1. Descriptive Analysis of Simulated Interval-Censored Data

Table 1 presents an illustrative example of a simulated dataset at a prevalence rate of p = 0.8 , showing the first ten subjects with their corresponding endpoints in the left and right intervals and the censoring status. As presented earlier, the simulation framework generated observation intervals with a constant width of two time units (e.g., weeks, months, or years) to mimic realistic follow-up schedules often encountered in longitudinal clinical studies. This structure ensured consistency and controlled variability across all simulated experiments, allowing direct comparison of estimator performance under similar conditions.
To illustrate the general behavior of the simulated survival data, Figure 1 presents the classical and smoothed penalized KM survival curves at a fixed sample size of n = 50 and prevalence rate p = 0.8 for the four imputation methods. Across all methods, survival probabilities decreased steadily over time, although the magnitude and smoothness of the decline differed by estimator. The classical KM curves exhibited visible stepwise discontinuities and small irregular jumps, reflecting the sensitivity of the estimator to random sampling variation within narrow censoring intervals. In contrast, the smoothed penalized KM curves showed smoother and more continuous transitions, indicating an effective reduction of overfitted stepwise fluctuations and sampling-induced noise. Overall, at this sample size, these plots demonstrate that the penalized smoothing procedure improves curve continuity and estimator stability without distorting the general shape of the survival function.

3.2.2. Monte Carlo Performance Evaluation of the KM-Based Estimators

This section evaluates how the smoothed penalized KM estimator compares with the classical KM estimator across varying prevalence rates and sample sizes. The performance of the smoothed and classical KM estimators was evaluated through extensive MC simulations ( m = 5000 ) across varying sample sizes ( n = 10 to 500) and fixed prevalence rates ( p = 0.2 , 0.4 , 0.6 , 0.8 , 1.0 ). The evaluation was conducted using four statistical imputation methods, that is, midpoint, regression, uniform, and multiple imputation, and performance was measured using MAE, NMAE, MSE, MaxAE, NRMSE, and NARMED. The results are presented in Table A2, Table A3, Table A4, Table A5, Table A6, Table A7, Table A8 and Table A9 and Figure A1, Figure A2, Figure A3, Figure A4, Figure A5, Figure A6, Figure A7 and Figure A8 included in the annexures.
Performance under Fixed Prevalence Rates with Varying Sample Sizes: A consistent pattern observed across all prevalence rates was the monotonic decrease in estimation error as the sample size increased. Both estimators demonstrated the statistical property of consistency, and all error metrics showed a systematic reduction with increasing n. However, this convergence behavior differed markedly between the two estimators. The classical KM estimator showed a steeper gradient of improvement, whereas the smoothed KM estimator demonstrated a more gradual improvement. In addition, the results reveal that the relative superiority of the estimators is fundamentally dependent on the sample size and the prevalence rate. For small sample sizes ( n 30 ), the smoothed KM estimator consistently outperformed the classical KM estimator across all prevalence rates and performance evaluation metrics. This advantage was most evident at the smallest sample sizes (i.e., n = 10 ).
A transition zone was identified where the performance advantage shifted from the smoothed to the classical KM estimator. The location of this transition was highly dependent on the prevalence rate. For a low prevalence rate ( p = 0.2 , see Table A2 and Figure A1), the transition occurred at a larger sample size ( n 80–100). For moderate prevalence rates ( p = 0.4 , 0.6 , see Table A3 and Figure A2 as well as Table A4 and Figure A3, respectively), the transition occurred at medium sample sizes ( n 40–80). For p = 1.0 (see Table A6 and Figure A5), the transition occurred at a small sample size ( n 30 ). Within this zone, the classical KM estimator generally became superior first when paired with regression and multiple imputation methods. For large sample sizes ( n 100 ), the classical KM estimator was clearly superior, exhibiting significantly lower error rates across all metrics and prevalence rates. The performance gap increased substantially as n increased, and the classical KM estimator often achieved less than half the MAE and MSE of the smoothed KM estimator at n = 500 .
Performance under Fixed Sample Sizes with Varying Prevalence Rates: For small sample sizes ( n = 10 and 30), according to Table A7 and Figure A6, the benefits of the smoothed penalized KM estimator are clear. The smoothed KM estimator consistently yielded lower errors across all prevalence rates. The largest visual separations between the smoothed and classical KM curves were observed at a small sample size ( n = 10 ). At moderate sample sizes ( n = 50 and 100), according to Table A8 and Figure A7, the differences between the estimators narrowed but did not disappear. The classical KM estimator produced slightly lower errors across all metrics, maintaining a small but consistent edge in the accuracy of the estimation at n = 50 . For large sample sizes ( n = 300 and 500), according to Table A9 and Figure A8, the classical KM estimator produced visibly lower errors across all metrics.

3.2.3. Convergence Validation for Simulated Data

To quantitatively assess convergence between the smoothed and classical KM estimators, we adopted a difference-based convergence criterion. For each combination of prevalence rate (p), imputation method, and performance metric, we defined the critical sample size n * as the smallest n for which the absolute difference in the corresponding metric between the two estimators fell below tolerance level δ . Formally,
n * = min n N : θ ^ smoothed ( n ) θ ^ classical ( n ) δ ,
where θ ^ ( n ) denotes the estimated value of a given performance metric evaluated at sample size n. In the main analysis, we set δ = 0.005 . This value was selected to represent a practically negligible difference in metric values, balancing two considerations: (i) the scale of the error-based metrics used in this study (which were typically small in magnitude once estimators stabilize) and (ii) the need to avoid declaring convergence too early due to MC noise at small n. In particular, a tolerance of 0.005 corresponded to a sub-percent level difference on the [ 0 , 1 ] survival-probability scale and provided an interpretable criterion for practical equivalence of estimator performance.
The analysis was performed across the entire range of sample sizes ( n = 10 to 500) and prevalence rates ( p = 0.2 to 1.0 ). Table 2 reports the primary convergence results for δ = 0.005 . The convergence validation results are presented in three parts. First, prevalence rate-dependent convergence is examined. This is followed by an assessment of convergence behavior across the different performance evaluation metrics. Finally, the convergence behavior is analyzed with respect to the imputation methods used.
Convergence Behavior Across Prevalence Rates: The convergence analysis reveals a strong inverse relationship between prevalence rate and critical sample size n. For lower prevalence rates ( p = 0.2 ), convergence typically required larger sample sizes ( n = 40–100 for most metrics). As prevalence increased to p = 0.6 , convergence occurred at moderate sample sizes ( n = 30–40), while, at p = 1.0 , convergence occurred at around n = 20–30.
Convergence Behavior Across Performance Evaluation Metrics: Different metrics exhibited different convergence patterns. The MSE-based metrics (MSE, NRMSE) converged most rapidly, with n 30 across nearly all simulation scenarios, indicating that squared error differences decreased rapidly with increasing sample size. Absolute error metrics (MAE, NMAE) showed intermediate convergence, which typically required n = 30–50 samples across prevalence rates. MaxAE demonstrated the slowest convergence, often requiring n 100 samples, particularly at lower prevalence rates.
Convergence Behavior Across Imputation Methods: The choice of imputation method significantly affected the convergence rates. The midpoint imputation consistently delayed convergence, requiring larger n values across all metrics and prevalence rates. Multiple imputation generally achieved the most rapid convergence, particularly for prevalence rates p 0.6 . On the other hand, regression and uniform imputation showed similar convergence patterns, typically intermediate between multiple and midpoint methods.

3.3. Empirical Study

3.3.1. Real Data Description

To complement and validate the simulation findings, the imputation techniques and simulation framework outlined earlier were applied to the acquired immunodeficiency syndrome (AIDS) Clinical Trials Group (ACTG) 181 dataset [57]. This allowed this study to assess whether the methodological performance observed in the simulated environment was consistent when applied to real-world data. The ACTG181 study was a clinical investigation involving 232 HIV-positive participants. The dataset includes 204 of the 232 participants who were tested at least once for cytomegalovirus (CMV) shedding and mycobacterium avium complex (MAC) colonization during the trial and who had no prior diagnosis of either condition at baseline. Testing was conducted during scheduled clinic visits, which occurred at regular monthly intervals. For participants who attended all scheduled visits, the event time was defined as the month in which the first positive test was recorded, generating discrete failure-time data. For participants who missed one or more visits and were subsequently found to be positive immediately after the missed visit(s), the event time was classified as having occurred within an interval, resulting in interval-censored failure-time data. For the present analysis, we focus on the CMV-shedding treatment group (hereafter referred to as Treatment 1). All individuals experienced the event of interest, and the event times were known to occur only between consecutive clinic visits. This design represents a classical instance of interval-censored data with no right-censoring, making the dataset particularly suitable for evaluating imputation procedures under conditions of complete event observation. The dataset is publicly available in the MLEcens package in RStudio version 2025.05.1+513 and has been widely applied in the statistical literature addressing interval-censored survival analysis methodologies (see [1,57,58], among others).

3.3.2. Real Data Bootstrapping Procedure

The ACTG181 Treatment 1 dataset was used to evaluate the performance of the smoothed nonparametric penalized likelihood KM estimator compared to that of the classical KM estimator. For analysis, each subject’s observation interval ( L i , R i ] was converted to a single event time using one of the imputation techniques, namely, midpoint, regression, uniform, and multiple imputation methods. For each imputation approach, the bootstrap algorithm was implemented in two layers, respectively, corresponding to the following:
(i) Construction of Reference Survival Function: To obtain a method-specific reference survival curve, m = 5000 bootstrap replicates were generated by sampling n observations with replacement from the full ACTG181 dataset. For each replicate, the event times were imputed using the chosen method and a KM estimator was fitted to Surv ( time , status ) . The KM curve was anchored at ( 0 , 1 ) and ( t max , S KM ( t max ) ) to ensure proper boundary behavior and subsequently smoothed using a cubic smoothing spline. Each smoothed curve was evaluated on a uniform time grid ranging from 0 to the dataset’s maximum right endpoint in unit increments. The m smoothed curves were then averaged point-by-point across replicates to produce a method-specific reference survival function, denoted by S ( t ) .
(ii) Finite-Sample Bootstrap Evaluation: A second bootstrap procedure, also with m = 5000 iterations, was carried out to assess the performance of a fixed sample size (for example, n = 150 ). In each replicate, the following steps were performed:
  • Randomly sample n observations with a replacement from the ACTG181 dataset;
  • Impute event times using the same imputation method under consideration;
  • Fit a KM survival curve on Surv ( time , status ) ;
  • Interpolate the resulting KM step function onto the common time grid.
This produced a bootstrap survival estimate S ^ b ( t ) for each replicate b = 1 , 2 , , m . Each estimated survival vector S ^ b ( t ) was then compared with its corresponding reference curve S ( t ) across the entire grid using six performance evaluation metrics: MAE, NMAE, MaxAE, MSE, NRMSE, and NARMED. For each metric, the replicate values m were summarized by the bootstrap mean, standard deviation, and a t-based 95% confidence interval.

3.4. Real Data Application Results

3.4.1. Descriptive Analysis for the ACTG181 Data

Table 3 presents the first ten observations of the ACTG181 dataset for Treatment 1. The columns show the subject ID, the left and right endpoints of the censoring intervals, and the censoring indicator. Unlike the simulated datasets, where intervals were generated with fixed widths for comparability, the ACTG181 intervals varied considerably in length between subjects.
Figure 2 presents the KM survival curves for the ACTG181 dataset under both the classical and smoothed penalized estimation frameworks for Treatment 1. The classical KM estimator exhibits stepwise and irregular survival declines, with pronounced early drops indicating sensitivity to limited event information. In contrast, the smoothed penalized KM curve follows a more continuous and statistically stable pattern, capturing the overall survival trend while suppressing overfitted fluctuations and sampling noise.

3.4.2. Real Data Bootstrap Resampling Results

To evaluate the robustness and applicability of the estimators on real data, a bootstrap resampling procedure with m = 5000 replicates was conducted on Treatment 1 of the ACTG181 dataset. Within each replicate, event times were imputed using the selected methods midpoint, regression, uniform, and multiple imputation followed by computation of both the classical and smoothed penalized KM curves. Performance was assessed using the same statistical metrics applied in the simulation study, namely, MAE, MaxAE, and MSE, along with their normalized counterparts (NMAE, NRMSE, and NARMED).
A consistent pattern emerged from the ACTG181 real-data analysis (Table A10 and Figure A9), closely mirroring the behavior observed in the simulated experiments. Across all imputation methods, estimation errors for both the classical and smoothed KM estimators decreased monotonically with increasing sample size, reflecting the expected consistency of both estimators. The rate and pattern of convergence differed between the two approaches: the classical KM estimator exhibited a steeper rate of improvement with increasing n, while the smoothed KM estimator improved more gradually, especially in the small-sample region.
For small samples ( n = 10 –30), the smoothed penalized KM estimator yielded visibly lower errors across all metrics, demonstrating its advantage in stabilizing survival estimates in small samples. Thus, for small samples ( n 30 ), the smoothed KM estimator consistently outperformed the classical estimator across all imputations and error metrics. The gains were most pronounced at n = 10 , where the smoothed KM reduced MAE by approximately 25–40% across imputation methods. Similar improvements were observed for NMAE, MSE, and NRMSE, confirming that smoothing effectively stabilizes the KM estimator in small samples.
A transition zone was observed between n = 20 and 30, where the classical KM estimator began to outperform the smoothed KM estimator for most imputation methods. The exact point of convergence varied slightly by method, such that regression and uniform imputations exhibited earlier stabilization ( n 20 –25), while midpoint and multiple imputations maintained a marginal smoothing advantage until about n = 30 .
At moderate sample sizes ( n = 80 and 100), the differences between estimators narrowed but did not disappear, that is, the classical KM produced slightly lower errors in most metrics, maintaining a small but consistent edge in the accuracy of the estimation. For these moderate sample sizes, the classical KM estimator became decisively superior across all imputations and metrics. This performance gap continued to widen with increasing n, confirming the asymptotic dominance of the classical KM estimator under high prevalence rates (i.e., p = 1).
For large sample sizes ( n = 150 and 200), according to Table A10 and Figure A9, the classical KM estimator achieved marginally better performance, confirming that smoothing benefits diminished as the sample size increased. It is evident that the results of the real data corroborate the simulation findings. The smoothed KM estimator exhibits clear advantages in small-sample scenarios. However, as the sample size increases, these benefits diminish rapidly and the classical KM estimator becomes consistently more efficient.

3.4.3. Convergence Validation for ACTG181 Real Data

To quantitatively assess the convergence between the smoothed and classical KM estimators for the ACTG181 dataset, a difference-based convergence criterion was used. For each imputation method and performance metric, the critical sample size n was defined as the smallest sample size at which the absolute difference between the smoothed and classical estimates fell below the prescribed convergence threshold. The analysis was conducted over the entire range of sample sizes ( n = 10 to 200) under a prevalence rate of p = 1.0 . Table 4 reports the resulting convergence sample sizes across all performance evaluation metrics and imputation methods.
Real Data Convergence Behavior Across Performance Evaluation Metrics: In all metrics evaluated, the convergence pattern reflected rapid alignment between the two estimators once n 25 –30. The MSE-based metrics (MSE, NRMSE) showed the fastest convergence, with n 15 –30 across all imputations, indicating that squared-error discrepancies diminished rapidly as the sample size grew. MAE and NMAE converged slightly later ( n 20 –25), while MaxAE required larger samples ( n = 30 –50), highlighting that local deviations in extreme observations persisted longer. NARMED achieved the earliest and most stable convergence overall ( n 20 ).
Real Data Convergence Behavior Across Imputation Methods: The convergence behavior varied significantly between the imputation methods, as reflected in Table 4. Midpoint imputation generally achieved the earliest convergence, with most metrics reaching the threshold δ = 0.005 at n = 15 –20, except MaxAE, which required n = 35 . Uniform imputation also converged relatively quickly, typically around n = 20 , with a slightly higher requirement for MaxAE ( n = 30 ). Regression imputation was the slowest to converge overall, consistently requiring the largest sample sizes across metrics (ranging from n = 25 to 50). Multiple imputation showed an intermediate convergence pattern, which required n = 20 –25 for most metrics, slightly slower than midpoint and uniform but notably faster than regression imputation.

4. Discussion and Conclusions

Given the limitations of the classical imputation-based KM method in small samples [59], a systematic evaluation of how the improved nonparametric penalized likelihood KM approach performs under varying samples and prevalence scenarios was needed. This study addressed this gap by conducting MC simulations and analyzing real interval-censored data to compare the performance of classical and nonparametric penalized KM estimators across selected imputation strategies. Multiple imputation strategies which include midpoint, uniform, regression-based, and multiple imputation were examined to capture differing imputation-induced representations of symmetry and asymmetry in event-time distributions and evaluate their impact on survival estimation. The TB estimator, while a general NPMLE for interval-censored data, does not readily accommodate standard imputation strategies, such as midpoint, uniform, regression-based, or multiple imputation, and it was therefore not considered in this study. Because no single metric fully captures all aspects of estimator performance in interval-censored settings, multiple complementary measures are required to provide a balanced evaluation. In particular, combining average, extreme, normalized, and robust error metrics allow the assessment of overall accuracy, sensitivity to outliers, scale effects, and stability under small samples and low prevalence rates.
Across extensive Monte Carlo simulations, the smoothed nonparametric penalized KM estimator consistently produced smaller error measures, particularly in small-sample and low prevalence settings. Penalization effectively reduced the inherent stepwise behavior of the classical KM estimator, supporting previous findings that smoothing lowers variance and leads to more reliable survival curve estimation [31,32,35]. In contrast, the classical KM estimator is superior in large-sample scenarios, with its advantage becoming more pronounced as the prevalence increases. The gains from penalization were especially pronounced when combined with regression and multiple imputation, which naturally accommodate data-driven and distributional asymmetry within censoring intervals. The midpoint imputation method was uniformly the worst-performing approach. Overall, the results suggest that explicitly accounting for the interplay between symmetry-preserving and asymmetry-inducing imputation strategies, together with penalized smoothing, leads to more reliable and interpretable survival estimates under interval censoring in small sample settings. The real-data application to the ACTG181 dataset further corroborated the simulation findings.
The convergence analysis provides quantitative support for some practical recommendations. The critical sample sizes n identified through this approach align closely with the transition zones observed in our performance analysis, validating the sample size guidelines of n 40–100 for low prevalence rates, n 30–40 for moderate prevalence rates, and n 20–30 for high prevalence rates. This convergence framework offers researchers an objective, metric-driven approach to estimator selection based on their specific sample size constraints and prevalence rate patterns. This study managed to extensively evaluate and benchmark the performance of an improved nonparametric penalized likelihood KM estimator; however, several areas remain open for further exploration. Firstly, the current study focused on sample size and prevalence rate. Future research should examine how interval width and observation frequency influence convergence rates, estimator bias, and estimation accuracy. Secondly, it is also important to note that the penalized likelihood estimation process is computationally intensive compared to the classical KM, especially under multiple imputation. Thirdly, further research could improve the efficiency of penalized estimation by using parallel computation or faster optimization routines such as quasi-Newton optimization techniques, which have demonstrated superiority in nonparametric procedures [60,61]. Further, in order to enhance precision, the use of alternative smoothing kernels, such as spline or Gaussian penalties, may also be a potential area of future research. In addition, although four imputation strategies were evaluated, additional or Bayesian interval-sampling approaches [62,63] may yield further improvements in estimator performance. Lastly, future studies should also consider benchmarking the nonparametric penalized likelihood KM estimator using the classical TB estimator as a baseline method under interval-censored data [64].

Author Contributions

Conceptualization, K.C., C.S.M. and A.S.O.; methodology, K.C. and C.S.M.; software, K.C. and C.S.M.; validation, K.C., C.S.M. and A.S.O.; formal analysis, K.C. and C.S.M.; investigation, K.C., C.S.M. and A.S.O.; resources, K.C. and C.S.M.; data curation, K.C. and C.S.M.; writing—original draft preparation, K.C. and C.S.M.; writing—review and editing, K.C., C.S.M. and A.S.O.; visualization, K.C. and C.S.M.; supervision, C.S.M. and A.S.O.; project administration, C.S.M.; funding acquisition, K.C. and C.S.M. All authors have read and agreed to the published version of the manuscript.

Funding

Services Sector Education and Training Authority (SETA) Bursary awarded to K.C. Also, thanks to the National Research Foundation (NRF) and the Department of Science, Technology, and Innovation (DSTI) of South Africa for sponsoring this study through the Research Developmental Grants for Y-Rated Researchers awarded to C.S.M. (Reference Number: CSRP240328211294).

Data Availability Statement

The data used in this article can be found in the respective cited source, MLEcens R package, https://rdrr.io/cran/MLEcens/man/actg181.html (accessed on 12 March 2025).

Acknowledgments

The authors would like to express their sincere gratitude to the reviewers and editor for their insightful comments and constructive feedback, which significantly contributed to improving the quality of this article.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Table A1. List of symbols and abbreviations used in this study.
Table A1. List of symbols and abbreviations used in this study.
Symbol/AbbreviationDescription
TTrue (unobserved) event or failure time
LLeft endpoint of the censoring interval
RRight endpoint of the censoring interval
IInterval within which the event time lies
nSample size
pPrevalence rate
t i Time point i at which survival is evaluated
S ( t ) True survival function at time t
S ^ ( t ) Estimated survival function at time t
S ^ i Estimated survival probability at time point i
R (normalization)Range of the true survival function, used for normalization
δ Convergence threshold or error tolerance
KMKaplan–Meier
TBTurnbull
NPMLENonparametric Maximum Likelihood Estimator
MLEMaximum Likelihood Estimation
MIMultiple Imputation
AFTAccelerated Failure Time
MCMonte Carlo
MAEMean Absolute Error
NMAENormalized Mean Absolute Error
MaxAEMaximum Absolute Error
MSEMean Squared Error
NRMSENormalized Root Mean Squared Error
NARMEDNormalized Root Median Error Deviation
MACMycobacterium Avium Complex
CMVCytomegalovirus
HIVHuman Immunodeficiency Virus
AIDSAcquired Immunodeficiency Syndrome
ACTGAIDS Clinical Trials Group
Table A2. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 0.2 .
Table A2. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 0.2 .
n MetricImputation Methods
MidpointRegressionUniformMultiple
SmoothedClassicalSmoothedClassicalSmoothedClassicalSmoothedClassical
10MAE0.04650.07210.04490.06960.04360.06870.07130.0627
NMAE0.13050.20250.12740.19760.12570.19820.20320.1775
MaxAE0.09200.16410.08770.16170.08550.15900.13980.1455
MSE0.00410.00860.00380.00800.00370.00770.00900.0064
NRMSE0.15070.23480.14670.23060.14490.23100.23600.2072
NARMED0.13120.19070.12830.18420.12670.18460.19580.1660
20MAE0.03280.04760.03290.04580.03280.04540.03600.0419
NMAE0.14020.20320.14350.19970.14180.19610.15640.1821
MaxAE0.07070.11160.07020.10940.07000.10830.07770.0988
MSE0.00220.00400.00220.00370.00220.00360.00270.0030
NRMSE0.16600.23950.17010.23640.16810.23180.18590.2149
NARMED0.13610.18620.14080.18310.13770.17860.14970.1672
30MAE0.03030.04010.03020.03750.03000.03730.02950.0339
NMAE0.14580.19300.14710.18280.14450.18000.14370.1650
MaxAE0.06530.09670.06390.09320.06350.09330.06330.0847
MSE0.00190.00300.00190.00260.00190.00260.00180.0021
NRMSE0.17270.23090.17410.22010.17100.21730.17090.1995
NARMED0.14370.17450.14620.16300.14350.16070.14130.1466
40MAE0.02660.03380.02710.03110.02700.03140.02750.0296
NMAE0.13190.16780.13360.15320.13440.15620.13790.1483
MaxAE0.05740.08490.05690.08120.05630.08150.05730.0755
MSE0.00150.00210.00150.00180.00150.00190.00150.0017
NRMSE0.15680.20360.15820.18670.15900.19050.16290.1807
NARMED0.12990.14830.13300.13420.13390.13580.13850.1298
50MAE0.02440.02940.02560.02820.02490.02740.02530.0259
NMAE0.12190.14680.12840.14130.12450.13710.12540.1279
MaxAE0.05200.07600.05200.07400.05100.07260.05170.0678
MSE0.00120.00160.00130.00150.00120.00140.00130.0013
NRMSE0.14450.17930.15120.17290.14670.16820.14810.1574
NARMED0.12100.12840.12980.12410.12510.11920.12660.1103
80MAE0.02130.02280.02190.02100.02240.02130.02140.0200
NMAE0.10660.11420.11030.10580.11170.10610.10610.0994
MaxAE0.04310.06070.04210.05750.04350.05850.04190.0540
MSE0.00090.00100.00090.00090.00090.00090.00090.0008
NRMSE0.12480.14070.12840.13150.13020.13190.12400.1239
NARMED0.10830.10030.11340.09080.11470.09150.10850.0848
100MAE0.01910.01960.02130.01880.02020.01800.02100.0179
NMAE0.09530.09760.10740.09440.10050.08980.10520.0899
MaxAE0.03820.05350.03960.05230.03850.05130.03940.0485
MSE0.00070.00070.00080.00070.00070.00060.00080.0006
NRMSE0.11120.12040.12330.11820.11630.11260.12140.1124
NARMED0.09720.08510.11240.08030.10390.07560.10910.0764
300MAE0.01630.01170.01800.01010.01790.01020.01760.0101
NMAE0.08180.05870.09050.05060.08960.05130.08740.0501
MaxAE0.02900.03450.02920.02940.02920.02980.02910.0280
MSE0.00040.00030.00050.00020.00050.00020.00050.0002
NRMSE0.09220.07380.10000.06410.09930.06510.09710.0634
NARMED0.08670.04980.09880.04310.09750.04340.09470.0427
500MAE0.01530.00940.01740.00800.01710.00780.01720.0078
NMAE0.07620.04670.08710.03990.08540.03880.08590.0388
MaxAE0.02600.02850.02730.02330.02730.02290.02690.0217
MSE0.00030.00020.00040.00010.00040.00010.00040.0001
NRMSE0.08510.05900.09550.05070.09390.04940.09430.0493
NARMED0.08170.03960.09580.03380.09340.03260.09480.0331
Note: MAE = Mean Absolute Error; NMAE = Normalized Mean Absolute Error; MaxAE = Maximum Absolute Error; MSE = Mean Squared Error; NRMSE = Normalized Root Mean Square Error; NARMED = Normalized Root Median Error Deviation.
Table A3. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 0.4 .
Table A3. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 0.4 .
n MetricImputation Methods
MidpointRegressionUniformMultiple
SmoothedClassicalSmoothedClassicalSmoothedClassicalSmoothedClassical
10MAE0.06010.08860.06080.08590.06000.08500.06510.0805
NMAE0.13320.19630.13770.19440.13440.19050.14610.1805
MaxAE0.12520.20510.12580.20060.12550.19990.13630.1868
MSE0.00690.01320.00730.01260.00710.01220.00840.0108
NRMSE0.15580.23010.16110.22820.15820.22390.17220.2121
NARMED0.13180.18120.13680.18010.13210.17650.14220.1674
20MAE0.04700.06200.04970.06020.04950.05970.04700.0561
NMAE0.11740.15490.12290.14890.12340.14880.11640.1391
MaxAE0.09960.15420.10060.15230.10130.15210.09650.1407
MSE0.00430.00680.00480.00660.00480.00650.00440.0057
NRMSE0.13850.18620.14430.18010.14500.18000.13730.1683
NARMED0.11670.13970.12350.13310.12410.13210.11570.1235
30MAE0.04080.04960.04400.04750.04310.04650.04310.0455
NMAE0.10180.12400.11040.11940.10790.11630.10720.1132
MaxAE0.08250.12660.08440.12510.08320.12340.08430.1178
MSE0.00320.00450.00360.00420.00350.00400.00350.0039
NRMSE0.11940.15010.12810.14620.12530.14260.12510.1389
NARMED0.10240.11160.11340.10530.11070.10150.10900.0990
40MAE0.03840.04260.04010.03980.04060.03990.03990.0391
NMAE0.09600.10650.09990.09920.10170.10000.09980.0978
MaxAE0.07460.11100.07520.10760.07530.10790.07450.1017
MSE0.00270.00340.00290.00300.00290.00300.00280.0029
NRMSE0.11130.12990.11520.12230.11680.12330.11490.1202
NARMED0.09810.09480.10290.08630.10540.08720.10280.0853
50MAE0.03660.03800.03870.03530.03860.03530.03840.0347
NMAE0.09180.09540.09660.08820.09630.08810.09570.0865
MaxAE0.06980.10130.06980.09660.07050.09740.06960.0914
MSE0.00230.00270.00250.00240.00260.00240.00250.0023
NRMSE0.10580.11680.11000.10900.11010.10940.10940.1069
NARMED0.09410.08430.10070.07690.09980.07630.09970.0756
80MAE0.03330.02980.03610.02760.03660.02760.03630.0271
NMAE0.08320.07450.09010.06880.09150.06890.09080.0677
MaxAE0.06000.08200.06160.07760.06200.07690.06120.0721
MSE0.00180.00170.00210.00150.00210.00150.00210.0014
NRMSE0.09430.09200.10110.08610.10250.08580.10160.0844
NARMED0.08710.06580.09570.05930.09790.05990.09680.0589
100MAE0.03230.02700.03580.02450.03560.02430.03530.0240
NMAE0.08070.06740.08970.06150.08930.06080.08830.0599
MaxAE0.05700.07490.05880.06900.05830.06860.05770.0642
MSE0.00160.00140.00190.00120.00190.00110.00190.0011
NRMSE0.09090.08320.09950.07700.09910.07610.09800.0748
NARMED0.08490.05960.09620.05350.09620.05220.09510.0521
300MAE0.02790.01620.03340.01360.03300.01370.03310.0136
NMAE0.06980.04060.08360.03410.08240.03410.08270.0339
MaxAE0.04660.04860.05110.03930.05080.03900.05060.0367
MSE0.00110.00050.00150.00040.00150.00040.00150.0004
NRMSE0.07770.05070.09150.04320.09010.04290.09040.0426
NARMED0.07440.03520.09110.02940.08970.02970.09020.0297
500MAE0.02730.01350.03340.01070.03280.01060.03330.0106
NMAE0.06800.03370.08330.02680.08210.02640.08310.0264
MaxAE0.04490.04130.05090.03080.05010.03030.05050.0286
MSE0.00100.00030.00150.00020.00140.00020.00150.0002
NRMSE0.07570.04210.09100.03390.08980.03340.09080.0332
NARMED0.07280.02910.09090.02310.08970.02290.09050.0230
Note: MAE = Mean Absolute Error; NMAE = Normalized Mean Absolute Error; MaxAE = Maximum Absolute Error; MSE = Mean Squared Error; NRMSE = Normalized Root Mean Square Error; NARMED = Normalized Root Median Error Deviation.
Table A4. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 0.6 .
Table A4. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 0.6 .
n MetricImputation Methods
MidpointRegressionUniformMultiple
SmoothedClassicalSmoothedClassicalSmoothedClassicalSmoothedClassical
10MAE0.07050.09840.06880.09680.06830.09730.07110.0954
NMAE0.11520.16080.11340.15960.11280.16050.11730.1574
MaxAE0.14630.23740.14340.23600.14180.23690.14800.2270
MSE0.00960.01630.00910.01580.00910.01610.00990.0155
NRMSE0.13500.19120.13320.18990.13230.19120.13790.1871
NARMED0.11440.14760.11270.14620.11200.14730.11590.1444
20MAE0.05470.07070.05710.06700.05830.06780.05590.0656
NMAE0.09110.11780.09540.11180.09710.11290.09330.1102
MaxAE0.10940.17790.10900.17320.11080.17510.10820.1642
MSE0.00550.00870.00580.00790.00600.00820.00560.0077
NRMSE0.10630.14150.10990.13550.11190.13680.10800.1334
NARMED0.09050.10700.09680.09990.09850.10110.09330.0989
30MAE0.04930.05680.05170.05290.05240.05300.05160.0519
NMAE0.08230.09470.08590.08790.08710.08800.08580.0863
MaxAE0.09440.14640.09310.14150.09400.14130.09370.1336
MSE0.00420.00570.00440.00510.00450.00510.00440.0049
NRMSE0.09500.11490.09750.10760.09890.10760.09770.1056
NARMED0.08180.08560.08750.07780.08930.07840.08700.0763
40MAE0.04580.04860.05110.04580.05020.04550.04980.0448
NMAE0.07620.08080.08520.07640.08370.07580.08290.0747
MaxAE0.08450.12890.08750.12410.08660.12300.08600.1158
MSE0.00340.00420.00410.00380.00400.00380.00390.0037
NRMSE0.08670.09830.09550.09400.09390.09330.09300.0916
NARMED0.07690.07270.08840.06730.08670.06720.08490.0661
50MAE0.04550.04420.04900.04080.04960.04090.04820.0399
NMAE0.07570.07350.08190.06810.08280.06820.08030.0665
MaxAE0.08100.11870.08260.11080.08300.11110.08130.1036
MSE0.00320.00350.00370.00300.00370.00310.00350.0029
NRMSE0.08550.08960.09130.08400.09220.08400.08960.0818
NARMED0.07680.06610.08460.06040.08580.06050.08260.0589
80MAE0.04200.03490.04610.03150.04680.03140.04690.0313
NMAE0.06990.05800.07670.05240.07790.05230.07820.0522
MaxAE0.07250.09640.07490.08720.07600.08690.07540.0816
MSE0.00260.00220.00310.00180.00310.00180.00310.0018
NRMSE0.07830.07110.08480.06500.08620.06480.08630.0644
NARMED0.07160.05190.07960.04610.08110.04610.08160.0466
100MAE0.04100.03170.04640.02840.04610.02840.04580.0278
NMAE0.06820.05270.07740.04720.07700.04730.07630.0464
MaxAE0.07010.08830.07370.07830.07360.07850.07310.0731
MSE0.00240.00180.00300.00150.00300.00150.00300.0014
NRMSE0.07620.06460.08520.05860.08480.05880.08410.0575
NARMED0.06970.04730.08110.04190.08030.04190.07960.0413
300MAE0.03630.02010.04530.01630.04510.01600.04470.0158
NMAE0.06050.03360.07540.02710.07500.02670.07460.0263
MaxAE0.06140.05890.07120.04510.07110.04480.07120.0421
MSE0.00180.00070.00270.00050.00270.00050.00260.0005
NRMSE0.06800.04140.08310.03380.08270.03330.08240.0328
NARMED0.06210.02990.07930.02410.07880.02360.07800.0233
500MAE0.03520.01710.04440.01250.04430.01240.04420.0122
NMAE0.05890.02870.07400.02090.07370.02060.07360.0203
MaxAE0.06080.05070.07090.03500.07070.03470.07060.0324
MSE0.00170.00050.00260.00030.00260.00030.00260.0003
NRMSE0.06690.03540.08190.02610.08160.02570.08150.0253
NARMED0.06090.02540.07780.01850.07750.01830.07770.0181
Note: MAE = Mean Absolute Error; NMAE = Normalized Mean Absolute Error; MaxAE = Maximum Absolute Error; MSE = Mean Squared Error; NRMSE = Normalized Root Mean Square Error; NARMED = Normalized Root Median Error Deviation.
Table A5. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 0.8 .
Table A5. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 0.8 .
n MetricImputation Methods
MidpointRegressionUniformMultiple
SmoothedClassicalSmoothedClassicalSmoothedClassicalSmoothedClassical
10MAE0.06620.10670.06730.10220.06710.10220.06670.1016
NMAE0.08280.13350.08440.12810.08400.12800.08360.1272
MaxAE0.13780.26160.13690.25610.13530.25620.13460.2451
MSE0.00840.01860.00850.01730.00850.01740.00830.0171
NRMSE0.09790.15830.09890.15280.09860.15260.09790.1509
NARMED0.08110.12360.08330.11730.08320.11710.08270.1173
20MAE0.05310.07380.05730.07020.05620.06930.05670.0697
NMAE0.06630.09210.07160.08760.07000.08650.07100.0872
MaxAE0.10540.18970.10780.18410.10580.18290.10630.1741
MSE0.00470.00920.00530.00850.00510.00830.00520.0084
NRMSE0.07690.11050.08190.10600.08020.10480.08090.1050
NARMED0.06410.08370.07000.07910.06840.07770.06990.0793
30MAE0.04960.06040.05400.05610.05370.05560.05380.0558
NMAE0.06210.07550.06760.07030.06710.06950.06740.0698
MaxAE0.09430.15800.09640.15080.09580.14960.09640.1414
MSE0.00390.00620.00430.00550.00430.00540.00430.0054
NRMSE0.07100.09100.07610.08560.07560.08470.07590.0846
NARMED0.06000.06880.06630.06290.06540.06220.06590.0628
40MAE0.04740.05230.05240.04820.05240.04850.05220.0477
NMAE0.05910.06520.06530.06010.06540.06050.06520.0596
MaxAE0.08750.13980.09200.13010.09160.13100.09150.1225
MSE0.00340.00470.00400.00410.00400.00410.00390.0040
NRMSE0.06720.07880.07340.07340.07340.07400.07310.0726
NARMED0.05720.05930.06360.05380.06390.05420.06380.0535
50MAE0.04540.04670.05180.04310.05130.04320.05020.0421
NMAE0.05680.05840.06480.05390.06420.05400.06270.0525
MaxAE0.08370.12580.08980.11720.08820.11710.08740.1080
MSE0.00300.00370.00380.00330.00370.00330.00360.0031
NRMSE0.06440.07060.07260.06600.07170.06600.07040.0641
NARMED0.05450.05320.06310.04840.06300.04840.06090.0473
80MAE0.04240.03760.04960.03370.04950.03360.04880.0333
NMAE0.05290.04700.06190.04200.06170.04190.06110.0416
MaxAE0.07760.10470.08640.09250.08690.09250.08510.0859
MSE0.00260.00240.00340.00200.00340.00200.00340.0020
NRMSE0.06040.05710.06980.05160.06970.05160.06880.0510
NARMED0.05050.04260.05870.03780.05890.03760.05860.0375
100MAE0.04110.03420.04830.02980.04930.03020.04880.0300
NMAE0.05140.04280.06040.03730.06170.03780.06110.0375
MaxAE0.07650.09610.08550.08250.08650.08280.08560.0771
MSE0.00240.00200.00330.00160.00340.00160.00330.0016
NRMSE0.05910.05200.06860.04590.06990.04660.06910.0459
NARMED0.04820.03870.05680.03330.05850.03400.05800.0341
300MAE0.03590.02300.04650.01710.04640.01750.04590.0170
NMAE0.04490.02880.05810.02140.05800.02190.05740.0213
MaxAE0.07180.06730.08630.04720.08580.04800.08510.0441
MSE0.00190.00090.00310.00050.00310.00050.00300.0005
NRMSE0.05340.03510.06740.02640.06720.02690.06650.0262
NARMED0.03980.02590.05340.01930.05320.01970.05250.0193
500MAE0.03430.02000.04540.01310.04480.01340.04520.0132
NMAE0.04290.02500.05680.01640.05600.01670.05650.0165
MaxAE0.07140.05880.08480.03650.08430.03710.08510.0343
MSE0.00180.00070.00300.00030.00290.00030.00300.0003
NRMSE0.05210.03060.06620.02030.06550.02070.06610.0203
NARMED0.03640.02250.05300.01470.05120.01490.05200.0150
Note: MAE = Mean Absolute Error; NMAE = Normalized Mean Absolute Error; MaxAE = Maximum Absolute Error; MSE = Mean Squared Error; NRMSE = Normalized Root Mean Square Error; NARMED = Normalized Root Median Error Deviation.
Table A6. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 1.0 .
Table A6. MC simulation results for mean performance of smoothed and classical KM methods under various imputation methods using m = 5000 for varying sample sizes at a fixed prevalence rate = 1.0 .
n MetricImputation Methods
MidpointRegressionUniformMultiple
SmoothedClassicalSmoothedClassicalSmoothedClassicalSmoothedClassical
10MAE0.06710.08460.06170.08290.06250.08240.06060.0802
NMAE0.06730.08490.06180.08310.06270.08260.06070.0803
MaxAE0.13790.26710.13010.26140.13200.26060.12710.2407
MSE0.00830.01350.00710.01270.00720.01260.00670.0119
NRMSE0.08080.10910.07460.10610.07560.10560.07300.1020
NARMED0.06320.06870.05710.06710.05780.06660.05630.0656
20MAE0.05430.06180.05380.05780.05240.05760.05210.0567
NMAE0.05470.06220.05410.05810.05280.05800.05250.0571
MaxAE0.11330.19870.11120.18600.10900.18470.10770.1701
MSE0.00540.00710.00500.00620.00480.00610.00470.0059
NRMSE0.06640.07940.06500.07430.06350.07400.06300.0724
NARMED0.04970.04960.04890.04580.04760.04580.04760.0459
30MAE0.05090.05150.05130.04670.05090.04730.04920.0459
NMAE0.05130.05190.05170.04710.05140.04770.04970.0463
MaxAE0.10460.16900.10470.15060.10380.15240.10010.1368
MSE0.00460.00490.00440.00400.00430.00420.00400.0039
NRMSE0.06180.06610.06170.06000.06130.06090.05910.0585
NARMED0.04680.04110.04700.03700.04680.03760.04570.0370
40MAE0.05120.04510.05040.04020.04950.04090.04940.0399
NMAE0.05160.04550.05090.04060.05000.04130.04990.0403
MaxAE0.10470.15120.10090.13070.09920.13150.09870.1189
MSE0.00460.00380.00410.00300.00400.00310.00400.0030
NRMSE0.06190.05810.06020.05180.05910.05260.05890.0510
NARMED0.04680.03600.04730.03180.04650.03260.04670.0323
50MAE0.05120.04110.04940.03620.04890.03650.04870.0355
NMAE0.05160.04140.04980.03660.04940.03690.04910.0359
MaxAE0.10290.13970.09750.11780.09740.11790.09620.1063
MSE0.00450.00310.00390.00240.00390.00250.00380.0023
NRMSE0.06150.05300.05850.04680.05820.04720.05780.0456
NARMED0.04750.03280.04690.02890.04640.02930.04640.0287
80MAE0.04990.03400.04800.02860.04840.02840.04730.0279
NMAE0.05040.03430.04850.02880.04890.02870.04780.0281
MaxAE0.09850.11900.09280.09230.09380.09230.09150.0838
MSE0.00420.00220.00360.00150.00370.00150.00350.0014
NRMSE0.05950.04420.05640.03690.05700.03670.05570.0359
NARMED0.04690.02690.04670.02310.04690.02290.04600.0227
100MAE0.05020.03170.04800.02540.04680.02330.04700.0249
NMAE0.05060.03200.04850.02570.04730.02350.04750.0252
MaxAE0.09750.11230.09230.08310.08980.07520.08930.0748
MSE0.00420.00190.00360.00120.00340.00100.00340.0012
NRMSE0.05940.04140.05620.03300.05480.03010.05490.0321
NARMED0.04740.02500.04700.02040.04600.01890.04650.0204
300MAE0.04610.02370.04400.01470.04400.01470.04470.0141
NMAE0.04650.02400.04450.01480.04450.01480.04520.0143
MaxAE0.08670.08750.08490.04770.08480.04750.08530.0427
MSE0.00350.00110.00310.00040.00310.00040.00310.0004
NRMSE0.05390.03150.05130.01900.05120.01900.05190.0182
NARMED0.04460.01800.04340.01190.04370.01190.04460.0115
500MAE0.04480.02180.04220.01130.04200.01130.04220.0111
NMAE0.04530.02200.04260.01150.04250.01140.04270.0112
MaxAE0.08360.08140.08260.03680.08190.03670.08230.0335
MSE0.00330.00090.00280.00020.00280.00020.00280.0002
NRMSE0.05210.02920.04920.01470.04890.01470.04920.0143
NARMED0.04400.01610.04150.00920.04150.00920.04180.0090
Note: MAE = Mean Absolute Error; NMAE = Normalized Mean Absolute Error; MaxAE = Maximum Absolute Error; MSE = Mean Squared Error; NRMSE = Normalized Root Mean Square Error; NARMED = Normalized Root Median Error Deviation.
Table A7. MC simulation results for mean performance evaluation metrics under various imputation methods for varying prevalence rates p at fixed sample sizes n = 10 and n = 30 .
Table A7. MC simulation results for mean performance evaluation metrics under various imputation methods for varying prevalence rates p at fixed sample sizes n = 10 and n = 30 .
n p MetricImputation Methods
MidpointRegressionUniformMultiple
SmoothedClassicalSmoothedClassicalSmoothedClassicalSmoothedClassical
0.2MAE0.04650.07210.04490.06960.04360.06870.07130.0627
NMAE0.13050.20250.12740.19760.12570.19820.20320.1775
MaxAE0.09200.16410.08770.16170.08550.15900.13980.1455
MSE0.00410.00860.00380.00800.00370.00770.00900.0064
NRMSE0.15070.23480.14670.23060.14490.23100.23600.2072
NARMED0.13120.19070.12830.18420.12670.18460.19580.1660
0.4MAE0.06010.08860.06080.08590.06000.08500.06510.0805
NMAE0.13320.19630.13770.19440.13440.19050.14610.1805
MaxAE0.12520.20510.12580.20060.12550.19990.13630.1868
MSE0.00690.01320.00730.01260.00710.01220.00840.0108
NRMSE0.15580.23010.16110.22820.15820.22390.17220.2121
NARMED0.13180.18120.13680.18010.13210.17650.14220.1674
100.6MAE0.07050.09840.06880.09680.06830.09730.07110.0954
NMAE0.11520.16080.11340.15960.11280.16050.11730.1574
MaxAE0.14630.23740.14340.23600.14180.23690.14800.2270
MSE0.00960.01630.00910.01580.00910.01610.00990.0155
NRMSE0.13500.19120.13320.18990.13230.19120.13790.1871
NARMED0.11440.14760.11270.14620.11200.14730.11590.1444
0.8MAE0.06620.10670.06730.10220.06710.10220.06670.1016
NMAE0.08280.13350.08440.12810.08400.12800.08360.1272
MaxAE0.13780.26160.13690.25610.13530.25620.13460.2451
MSE0.00840.01860.00850.01730.00850.01740.00830.0171
NRMSE0.09790.15830.09890.15280.09860.15260.09790.1509
NARMED0.08110.12360.08330.11730.08320.11710.08270.1173
1.0MAE0.06710.08460.06170.08290.06250.08240.06060.0802
NMAE0.06730.08490.06180.08310.06270.08260.06070.0803
MaxAE0.13790.26710.13010.26140.13200.26060.12710.2407
MSE0.00830.01350.00710.01270.00720.01260.00670.0119
NRMSE0.08080.10910.07460.10610.07560.10560.07300.1020
NARMED0.06320.06870.05710.06710.05780.06660.05630.0656
0.2MAE0.03030.04010.03020.03750.03000.03730.02950.0339
NMAE0.14580.19300.14710.18280.14450.18000.14370.1650
MaxAE0.06530.09670.06390.09320.06350.09330.06330.0847
MSE0.00190.00300.00190.00260.00190.00260.00180.0021
NRMSE0.17270.23090.17410.22010.17100.21730.17090.1995
NARMED0.14370.17450.14620.16300.14350.16070.14130.1466
0.4MAE0.04080.04960.04400.04750.04310.04650.04310.0455
NMAE0.10180.12400.11040.11940.10790.11630.10720.1132
MaxAE0.08250.12660.08440.12510.08320.12340.08430.1178
MSE0.00320.00450.00360.00420.00350.00400.00350.0039
NRMSE0.11940.15010.12810.14620.12530.14260.12510.1389
NARMED0.10240.11160.11340.10530.11070.10150.10900.0990
300.6MAE0.04930.05680.05170.05290.05240.05300.05160.0519
NMAE0.08230.09470.08590.08790.08710.08800.08580.0863
MaxAE0.09440.14640.09310.14150.09400.14130.09370.1336
MSE0.00420.00570.00440.00510.00450.00510.00440.0049
NRMSE0.09500.11490.09750.10760.09890.10760.09770.1056
NARMED0.08180.08560.08750.07780.08930.07840.08700.0763
0.8MAE0.04960.06040.05400.05610.05370.05560.05380.0558
NMAE0.06210.07550.06760.07030.06710.06950.06740.0698
MaxAE0.09430.15800.09640.15080.09580.14960.09640.1414
MSE0.00390.00620.00430.00550.00430.00540.00430.0054
NRMSE0.07100.09100.07610.08560.07560.08470.07590.0846
NARMED0.06000.06880.06630.06290.06540.06220.06590.0628
1.0MAE0.05090.05150.05130.04670.05090.04730.04920.0459
NMAE0.05130.05190.05170.04710.05140.04770.04970.0463
MaxAE0.10460.16900.10470.15060.10380.15240.10010.1368
MSE0.00460.00490.00440.00400.00430.00420.00400.0039
NRMSE0.06180.06610.06170.06000.06130.06090.05910.0585
NARMED0.04680.04110.04700.03700.04680.03760.04570.0370
Note: MAE = Mean Absolute Error; NMAE = Normalized Mean Absolute Error; MaxAE = Maximum Absolute Error; MSE = Mean Squared Error; NRMSE = Normalized Root Mean Square Error; NARMED = Normalized Root Median Error Deviation.
Table A8. MC simulation results for mean performance evaluation metrics under various imputation methods for varying prevalence rates p at fixed sample sizes n = 50 and n = 100 .
Table A8. MC simulation results for mean performance evaluation metrics under various imputation methods for varying prevalence rates p at fixed sample sizes n = 50 and n = 100 .
n p MetricImputation Methods
MidpointRegressionUniformMultiple
SmoothedClassicalSmoothedClassicalSmoothedClassicalSmoothedClassical
0.2MAE0.02440.02940.02560.02820.02490.02740.02530.0259
NMAE0.12190.14680.12840.14130.12450.13710.12540.1279
MaxAE0.05200.07600.05200.07400.05100.07260.05170.0678
MSE0.00120.00160.00130.00150.00120.00140.00130.0013
NRMSE0.14450.17930.15120.17290.14670.16820.14810.1574
NARMED0.12100.12840.12980.12410.12510.11920.12660.1103
0.4MAE0.03660.03800.03870.03530.03860.03530.03840.0347
NMAE0.09180.09540.09660.08820.09630.08810.09570.0865
MaxAE0.06980.10130.06980.09660.07050.09740.06960.0914
MSE0.00230.00270.00250.00240.00260.00240.00250.0023
NRMSE0.10580.11680.11000.10900.11010.10940.10940.1069
NARMED0.09410.08430.10070.07690.09980.07630.09970.0756
500.6MAE0.04550.04420.04900.04080.04960.04090.04820.0399
NMAE0.07570.07350.08190.06810.08280.06820.08030.0665
MaxAE0.08100.11870.08260.11080.08300.11110.08130.1036
MSE0.00320.00350.00370.00300.00370.00310.00350.0029
NRMSE0.08550.08960.09130.08400.09220.08400.08960.0818
NARMED0.07680.06610.08460.06040.08580.06050.08260.0589
0.8MAE0.04540.04670.05180.04310.05130.04320.05020.0421
NMAE0.05680.05840.06480.05390.06420.05400.06270.0525
MaxAE0.08370.12580.08980.11720.08820.11710.08740.1080
MSE0.00300.00370.00380.00330.00370.00330.00360.0031
NRMSE0.06440.07060.07260.06600.07170.06600.07040.0641
NARMED0.05450.05320.06310.04840.06300.04840.06090.0473
1.0MAE0.05120.04110.04940.03620.04890.03650.04870.0355
NMAE0.05160.04140.04980.03660.04940.03690.04910.0359
MaxAE0.10290.13970.09750.11780.09740.11790.09620.1063
MSE0.00450.00310.00390.00240.00390.00250.00380.0023
NRMSE0.06150.05300.05850.04680.05820.04720.05780.0456
NARMED0.04750.03280.04690.02890.04640.02930.04640.0287
0.2MAE0.01910.01960.02130.01880.02020.01800.02100.0179
NMAE0.09530.09760.10740.09440.10050.08980.10520.0899
MaxAE0.03820.05350.03960.05230.03850.05130.03940.0485
MSE0.00070.00070.00080.00070.00070.00060.00080.0006
NRMSE0.11120.12040.12330.11820.11630.11260.12140.1124
NARMED0.09720.08510.11240.08030.10390.07560.10910.0764
0.4MAE0.03230.02700.03580.02450.03560.02430.03530.0240
NMAE0.08070.06740.08970.06150.08930.06080.08830.0599
MaxAE0.05700.07490.05880.06900.05830.06860.05770.0642
MSE0.00160.00140.00190.00120.00190.00110.00190.0011
NRMSE0.09090.08320.09950.07700.09910.07610.09800.0748
NARMED0.08490.05960.09620.05350.09620.05220.09510.0521
1000.6MAE0.04100.03170.04640.02840.04610.02840.04580.0278
NMAE0.06820.05270.07740.04720.07700.04730.07630.0464
MaxAE0.07010.08830.07370.07830.07360.07850.07310.0731
MSE0.00240.00180.00300.00150.00300.00150.00300.0014
NRMSE0.07620.06460.08520.05860.08480.05880.08410.0575
NARMED0.06970.04730.08110.04190.08030.04190.07960.0413
0.8MAE0.04110.03420.04830.02980.04930.03020.04880.0300
NMAE0.05140.04280.06040.03730.06170.03780.06110.0375
MaxAE0.07650.09610.08550.08250.08650.08280.08560.0771
MSE0.00240.00200.00330.00160.00340.00160.00330.0016
NRMSE0.05910.05200.06860.04590.06990.04660.06910.0459
NARMED0.04820.03870.05680.03330.05850.03400.05800.0341
1.0MAE0.05020.03170.04800.02540.04680.02330.04700.0249
NMAE0.05060.03200.04850.02570.04730.02350.04750.0252
MaxAE0.09750.11230.09230.08310.08980.07520.08930.0748
MSE0.00420.00190.00360.00120.00340.00100.00340.0012
NRMSE0.05940.04140.05620.03300.05480.03010.05490.0321
NARMED0.04740.02500.04700.02040.04600.01890.04650.0204
Note: MAE = Mean Absolute Error; NMAE = Normalized Mean Absolute Error; MaxAE = Maximum Absolute Error; MSE = Mean Squared Error; NRMSE = Normalized Root Mean Square Error; NARMED = Normalized Root Median Error Deviation.
Table A9. MC simulation results for mean performance evaluation metrics under various imputation methods for varying prevalence rates p at fixed sample sizes n = 300 and n = 500 .
Table A9. MC simulation results for mean performance evaluation metrics under various imputation methods for varying prevalence rates p at fixed sample sizes n = 300 and n = 500 .
n p MetricImputation Methods
MidpointRegressionUniformMultiple
SmoothedClassicalSmoothedClassicalSmoothedClassicalSmoothedClassical
0.2MAE0.01630.01170.01800.01010.01790.01020.01760.0101
NMAE0.08180.05870.09050.05060.08960.05130.08740.0501
MaxAE0.02900.03450.02920.02940.02920.02980.02910.0280
MSE0.00040.00030.00050.00020.00050.00020.00050.0002
NRMSE0.09220.07380.10000.06410.09930.06510.09710.0634
NARMED0.08670.04980.09880.04310.09750.04340.09470.0427
0.4MAE0.02790.01620.03340.01360.03300.01370.03310.0136
NMAE0.06980.04060.08360.03410.08240.03410.08270.0339
MaxAE0.04660.04860.05110.03930.05080.03900.05060.0367
MSE0.00110.00050.00150.00040.00150.00040.00150.0004
NRMSE0.07770.05070.09150.04320.09010.04290.09040.0426
NARMED0.07440.03520.09110.02940.08970.02970.09020.0297
3000.6MAE0.03630.02010.04530.01630.04510.01600.04470.0158
NMAE0.06050.03360.07540.02710.07500.02670.07460.0263
MaxAE0.06140.05890.07120.04510.07110.04480.07120.0421
MSE0.00180.00070.00270.00050.00270.00050.00260.0005
NRMSE0.06800.04140.08310.03380.08270.03330.08240.0328
NARMED0.06210.02990.07930.02410.07880.02360.07800.0233
0.8MAE0.03590.02300.04650.01710.04640.01750.04590.0170
NMAE0.04490.02880.05810.02140.05800.02190.05740.0213
MaxAE0.07180.06730.08630.04720.08580.04800.08510.0441
MSE0.00190.00090.00310.00050.00310.00050.00300.0005
NRMSE0.05340.03510.06740.02640.06720.02690.06650.0262
NARMED0.03980.02590.05340.01930.05320.01970.05250.0193
1.0MAE0.04610.02370.04400.01470.04400.01470.04470.0141
NMAE0.04650.02400.04450.01480.04450.01480.04520.0143
MaxAE0.08670.08750.08490.04770.08480.04750.08530.0427
MSE0.00350.00110.00310.00040.00310.00040.00310.0004
NRMSE0.05390.03150.05130.01900.05120.01900.05190.0182
NARMED0.04460.01800.04340.01190.04370.01190.04460.0115
0.2MAE0.01530.00940.01740.00800.01710.00780.01720.0078
NMAE0.07620.04670.08710.03990.08540.03880.08590.0388
MaxAE0.02600.02850.02730.02330.02730.02290.02690.0217
MSE0.00030.00020.00040.00010.00040.00010.00040.0001
NRMSE0.08510.05900.09550.05070.09390.04940.09430.0493
NARMED0.08170.03960.09580.03380.09340.03260.09480.0331
0.4MAE0.02730.01350.03340.01070.03280.01060.03330.0106
NMAE0.06800.03370.08330.02680.08210.02640.08310.0264
MaxAE0.04490.04130.05090.03080.05010.03030.05050.0286
MSE0.00100.00030.00150.00020.00140.00020.00150.0002
NRMSE0.07570.04210.09100.03390.08980.03340.09080.0332
NARMED0.07280.02910.09090.02310.08970.02290.09050.0230
5000.6MAE0.03520.01710.04440.01250.04430.01240.04420.0122
NMAE0.05890.02870.07400.02090.07370.02060.07360.0203
MaxAE0.06080.05070.07090.03500.07070.03470.07060.0324
MSE0.00170.00050.00260.00030.00260.00030.00260.0003
NRMSE0.06690.03540.08190.02610.08160.02570.08150.0253
NARMED0.06090.02540.07780.01850.07750.01830.07770.0181
0.8MAE0.03430.02000.04540.01310.04480.01340.04520.0132
NMAE0.04290.02500.05680.01640.05600.01670.05650.0165
MaxAE0.07140.05880.08480.03650.08430.03710.08510.0343
MSE0.00180.00070.00300.00030.00290.00030.00300.0003
NRMSE0.05210.03060.06620.02030.06550.02070.06610.0203
NARMED0.03640.02250.05300.01470.05120.01490.05200.0150
1.0MAE0.04480.02180.04220.01130.04200.01130.04220.0111
NMAE0.04530.02200.04260.01150.04250.01140.04270.0112
MaxAE0.08360.08140.08260.03680.08190.03670.08230.0335
MSE0.00330.00090.00280.00020.00280.00020.00280.0002
NRMSE0.05210.02920.04920.01470.04890.01470.04920.0143
NARMED0.04400.01610.04150.00920.04150.00920.04180.0090
Note: MAE = Mean Absolute Error; NMAE = Normalized Mean Absolute Error; MaxAE = Maximum Absolute Error; MSE = Mean Squared Error; NRMSE = Normalized Root Mean Square Error; NARMED = Normalized Root Median Error Deviation.
Table A10. ACTG181 data results for performance evaluation metrics under various imputation methods for varying sample sizes n using m = 5000 at prevalence p = 1.0 .
Table A10. ACTG181 data results for performance evaluation metrics under various imputation methods for varying sample sizes n using m = 5000 at prevalence p = 1.0 .
n MetricImputation Methods
MidpointRegressionUniformMultiple
SmoothedClassicalSmoothedClassicalSmoothedClassicalSmoothedClassical
10MAE0.08500.09790.07220.11520.06340.08400.05830.0679
NMAE0.10300.11820.07750.11560.06440.10140.05160.0670
MaxAE0.59410.63320.16400.29720.38290.41260.14470.1564
MSE0.02640.02810.01000.02800.00650.01350.00530.0079
NRMSE0.14300.18400.08690.14550.09610.13630.06620.0806
NARMED0.10450.14220.07090.09670.04700.08010.04970.0625
20MAE0.07440.04140.06840.07740.05500.05440.05430.0548
NMAE0.09420.04140.06840.07750.05970.05910.05510.0550
MaxAE0.16100.20580.15680.19510.14430.16390.12490.1166
MSE0.00880.00300.00870.01160.00480.00500.00410.0051
NRMSE0.09210.05410.08020.09580.07030.07280.06140.0648
NARMED0.09230.03350.06630.06840.06410.05030.04800.0535
30MAE0.07950.02540.06370.06020.05720.03870.05370.0460
NMAE0.09170.02600.06370.06020.05780.03900.05480.0462
MaxAE0.14640.13210.13790.15350.13390.13070.12170.0978
MSE0.01000.00130.00700.00670.00530.00280.00390.0036
NRMSE0.10150.03640.07660.07450.07110.04990.06400.0543
NARMED0.08690.01920.06090.05390.07000.03130.05150.0445
40MAE0.08120.01900.06460.05220.06320.03460.05410.0360
NMAE0.09020.01900.06240.05220.05770.03460.05520.0361
MaxAE0.14240.13190.12830.13550.14410.11480.12100.0793
MSE0.01050.00070.00680.00520.00580.00210.00410.0023
NRMSE0.10340.02660.07380.06520.07170.04400.06400.0428
NARMED0.08510.01540.06140.04580.07160.02790.05340.0343
50MAE0.08300.01810.06340.04720.06520.03130.05410.0318
NMAE0.09300.01810.06030.04710.05880.03130.05410.0318
MaxAE0.16890.11280.11970.12430.14870.10360.11960.0697
MSE0.01190.00030.00750.00430.00600.00180.00440.0017
NRMSE0.10370.02680.07180.05940.07650.04000.06470.0378
NARMED0.08050.01320.06100.04030.07320.02510.05680.0304
80MAE0.08850.01050.05760.03720.06800.02340.05340.0259
NMAE0.09380.01450.06050.03720.06180.02340.05340.0260
MaxAE0.15470.04140.12180.09270.14790.07990.11640.0562
MSE0.01020.00020.00720.00260.00690.00100.00420.0012
NRMSE0.10940.01480.07060.04590.07410.03010.06610.0309
NARMED0.07510.00770.06040.03420.06900.01870.05610.0254
100MAE0.09180.01540.06610.03350.06960.02190.05440.0224
NMAE0.09210.01540.06610.03350.06960.02190.05400.0224
MaxAE0.20950.08770.11160.08550.15330.07350.11440.0487
MSE0.01320.00050.00700.00210.00730.00090.00450.0009
NRMSE0.10120.02160.08210.04160.08250.02810.06550.0266
NARMED0.07670.00980.05920.03000.07050.01750.05860.0219
150MAE0.08620.00870.06450.02740.07060.01730.05560.0187
NMAE0.08600.01360.06450.02740.07060.01730.05560.0187
MaxAE0.18160.02330.13940.06920.14910.05850.11430.0409
MSE0.01130.00010.00680.00140.00740.00060.00480.0006
NRMSE0.10610.01070.08090.03400.08350.02230.06750.0223
NARMED0.07280.00730.06050.02490.07420.01390.05870.0183
200MAE0.11340.01240.08070.02320.08160.01480.06750.0150
NMAE0.11360.01240.08070.02320.08160.01480.06750.0150
MaxAE0.25570.08160.17910.06100.17060.05150.13560.0337
MSE0.01990.00040.01060.00100.00990.00040.00700.0004
NRMSE0.14130.01910.10210.02910.09590.01930.08210.0180
NARMED0.08600.00680.06690.02040.08520.01160.07090.0145
Note: MAE = Mean Absolute Error; NMAE = Normalized Mean Absolute Error; MaxAE = Maximum Absolute Error; MSE = Mean Squared Error; NRMSE = Normalized Root Mean Squared Error; NARMED = Normalized Absolute Root Median Error.
Figure A1. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 0.2 .
Figure A1. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 0.2 .
Symmetry 18 00519 g0a1
Figure A2. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 0.4 .
Figure A2. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 0.4 .
Symmetry 18 00519 g0a2
Figure A3. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 0.6 .
Figure A3. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 0.6 .
Symmetry 18 00519 g0a3
Figure A4. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 0.8 .
Figure A4. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 0.8 .
Symmetry 18 00519 g0a4
Figure A5. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 1.0 .
Figure A5. Simulated data mean metrics vs. sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate, p = 1.0 .
Symmetry 18 00519 g0a5
Figure A6. MC simulation results for the performance evaluation metrics under various imputation methods for varying prevalence rates p at sample sizes n = 10 and n = 30 .
Figure A6. MC simulation results for the performance evaluation metrics under various imputation methods for varying prevalence rates p at sample sizes n = 10 and n = 30 .
Symmetry 18 00519 g0a6aSymmetry 18 00519 g0a6b
Figure A7. MC simulation results for the performance evaluation metrics under various imputation methods for varying prevalence rates p at sample sizes n = 50 and n = 100 .
Figure A7. MC simulation results for the performance evaluation metrics under various imputation methods for varying prevalence rates p at sample sizes n = 50 and n = 100 .
Symmetry 18 00519 g0a7aSymmetry 18 00519 g0a7b
Figure A8. MC simulation results for the performance evaluation metrics under various imputation methods for varying prevalence rates p at sample sizes n = 300 and n = 500 .
Figure A8. MC simulation results for the performance evaluation metrics under various imputation methods for varying prevalence rates p at sample sizes n = 300 and n = 500 .
Symmetry 18 00519 g0a8aSymmetry 18 00519 g0a8b
Figure A9. ACTG181 data mean metrics vs sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate.
Figure A9. ACTG181 data mean metrics vs sample size for performance comparisons of classical KM and smoothed KM for various imputation methods at fixed prevalence rate.
Symmetry 18 00519 g0a9

References

  1. Sun, J. The Statistical Analysis of Interval-Censored Failure Time Data; Springer: New York, NY, USA, 2006. [Google Scholar]
  2. Du, M.; Su, J. Statistical analysis of interval-censored failure time data. Chin. J. Appl. Probab. Stat. 2021, 37, 627–654. [Google Scholar]
  3. Sun, J.; Chen, D.G. (Eds.) Emerging Topics in Modeling Interval-Censored Survival Data; Springer: Cham, Switzerland, 2022. [Google Scholar]
  4. Finkelstein, D.M.; Wolfe, R.A. A semiparametric model for regression analysis of interval-censored failure time data. Biometrics 1985, 41, 933–945. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Turnbull, B.W. The empirical distribution function with arbitrarily grouped, censored, and truncated data. J. R. Stat. Soc. Ser. B (Methodol.) 1976, 38, 290–295. [Google Scholar] [CrossRef] [Scilit]
  6. Finkelstein, D.M. A proportional hazards model for interval-censored failure time data. Biometrics 1986, 42, 845–854. [Google Scholar] [CrossRef] [Scilit]
  7. Lindsey, J.C.; Ryan, L.M. Tutorial in biostatistics: Methods for interval-censored data. Stat. Med. 1998, 17, 219–238. [Google Scholar] [CrossRef] [Scilit]
  8. Jin, S.; Liu, Y.; Mao, G.; Sun, J.; Wu, Y. Nonparametric estimation of a survival function in the presence of measurement errors on the failure time of interest. Can. J. Stat. 2024, 52, 783–803. [Google Scholar] [CrossRef] [Scilit]
  9. Rücker, G.; Messerer, D. Remission duration: An example of interval-censored observations. Stat. Med. 1988, 7, 1139–1145. [Google Scholar] [CrossRef] [Scilit]
  10. Odell, P.M.; Anderson, K.M.; D’Agostino, R.B. Maximum likelihood estimation for interval-censored data using a Weibull-based accelerated failure time model. Biometrics 1992, 48, 951–959. [Google Scholar] [CrossRef] [Scilit]
  11. Dorey, F.J.; Little, R.J.; Schenker, N. Multiple imputation for threshold-crossing data with interval censoring. Stat. Med. 1993, 12, 1589–1603. [Google Scholar] [CrossRef] [Scilit]
  12. Kaplan, E.L.; Meier, P. Nonparametric estimation from incomplete observations. J. Am. Stat. Assoc. 1958, 53, 457–481. [Google Scholar] [CrossRef]
  13. Smith, P.J.; Thompson, T.J.; Jereb, J.A. A model for interval-censored tuberculosis outbreak data. Stat. Med. 1997, 16, 485–496. [Google Scholar] [CrossRef]
  14. Turkson, A.J.; Ayiah-Mensah, F.; Nimoh, V. Handling censoring and censored data in survival analysis. Int. J. Math. Math. Sci. 2021, 2021, 9307475. [Google Scholar] [CrossRef] [Scilit]
  15. Ali, A.K.; Arasan, J. Left, Right, Midpoint and Random Point Imputation Techniques for Weibull Regression Model with Right and Interval-Censored Data. Appl. Math. Comput. Intell. (AMCI) 2024, 13, 115–142. [Google Scholar] [CrossRef] [Scilit]
  16. Nakagawa, Y.; Sozu, T. Improvement of Midpoint Imputation for Estimation of Median Survival Time for Interval-Censored Time-to-Event Data. Ther. Innov. Regul. Sci. 2024, 58, 721–729. [Google Scholar] [CrossRef] [Scilit]
  17. Lou, Y.; Ma, Y.; Xiang, L.; Sun, J. A multiple imputation approach for flexible modelling of interval-censored data with missing and censored covariates. Comput. Stat. Data Anal. 2025, 209, 108177. [Google Scholar] [CrossRef] [Scilit]
  18. Huang, J.; Wellner, J.A. Efficient estimation for the proportional hazards model. Ann. Stat. 1995, 23, 749–768. [Google Scholar]
  19. Dempster, A.P.; Laird, N.M.; Rubin, D.B. Maximum likelihood from incomplete data via the EM algorithm. J. R. Stat. Soc. Ser. B (Methodol.) 1977, 39, 1–38. [Google Scholar] [CrossRef] [Scilit]
  20. Bishop, C.M. Pattern Recognition and Machine Learning; Springer: New York, NY, USA, 2016. [Google Scholar]
  21. Kuo, L.; Peng, F. A mixture-model approach to the analysis of survival data. In Generalized Linear Models: A Bayesian Perspective; Taylor and Francis: New York, NY, USA, 2000; pp. 255–270. [Google Scholar]
  22. Yavuz, A.Ç.; Lambert, P. Smooth estimation of survival functions and hazard ratios from interval-censored data using Bayesian penalized B-splines. Stat. Med. 2011, 30, 75–90. [Google Scholar] [CrossRef] [Scilit]
  23. Zhang, Y.; Hua, L.; Huang, J. A spline-based semiparametric maximum likelihood estimation method for the Cox model with interval-censored data. Scand. J. Stat. 2010, 37, 338–354. [Google Scholar] [CrossRef] [Scilit]
  24. Kooperberg, C.; Stone, C.J. Logspline density estimation for censored data. J. Comput. Graph. Stat. 1992, 1, 301–328. [Google Scholar] [CrossRef] [Scilit]
  25. Rosenblatt, M. Remarks on some nonparametric estimates of a density function. Ann. Math. Stat. 1956, 27, 832–837. [Google Scholar] [CrossRef] [Scilit]
  26. Whittle, P. On the smoothing of probability density functions. J. R. Stat. Soc. Ser. B (Methodol.) 1958, 20, 334–343. [Google Scholar] [CrossRef] [Scilit]
  27. Parzen, E. On estimation of a probability density function and mode. Ann. Math. Stat. 1962, 33, 1065–1076. [Google Scholar] [CrossRef] [Scilit]
  28. Wang, S.; Chen, W.; Chen, M.; Zhou, Y. Maximum likelihood estimation of the parameters of the inverse Gaussian distribution using maximum ranked set sampling with unequal samples. Math. Popul. Stud. 2023, 30, 1–21. [Google Scholar] [CrossRef] [Scilit]
  29. Chen, M.; Chen, W.X.; Yang, R. Double moving extremes ranked set sampling design. Acta Math. Appl. Sin. Engl. Ser. 2024, 40, 75–90. [Google Scholar] [CrossRef] [Scilit]
  30. Chen, M.; Chen, W.X.; Deng, C.H. Some new results on parameter estimation of the exponential–Poisson distribution in ranked set sampling. Appl. Math. J. Chin. Univ. 2025, 40, 413–428. [Google Scholar] [CrossRef] [Scilit]
  31. Wang, J.L. Smoothing hazard rates. Encycl. Biostat. 2005, 7, 4986–4997. [Google Scholar]
  32. Efron, B. The two-sample problem with censored data. In Proceedings of the Fifth Berkeley Symposium in Mathematical Statistics, IV; Prentice-Hall: New York, NY, USA, 1967; Volume 4, pp. 831–853. [Google Scholar]
  33. Babyak, M.A. What you see may not be what you get: A brief, nontechnical introduction to overfitting in regression-type models. Biopsychosoc. Med. 2004, 66, 411–421. [Google Scholar]
  34. Wu, Y.; Kolassa, J. Interval-specific censoring set adjusted Kaplan–Meier estimator. J. Appl. Stat. 2024, 51, 2436–2456. [Google Scholar] [CrossRef] [Scilit]
  35. Tubbs, J.D.; Chen, L.G.; Thach, T.Q.; Sham, P.C. Improved nonparametric penalized maximum likelihood estimation for arbitrarily censored survival data. Stat. Med. 2022, 41, 4006–4021. [Google Scholar] [CrossRef] [Scilit]
  36. Peterson, A.V., Jr. Expressing the Kaplan–Meier estimator as a function of empirical subsurvival functions. J. Am. Stat. Assoc. 1977, 72, 854–858. [Google Scholar]
  37. Nadaraya, E.A. On estimating regression. Theory Probab. Its Appl. 1964, 9, 141–142. [Google Scholar] [CrossRef] [Scilit]
  38. Buck, S.F. A method of estimation of missing values in multivariate data suitable for use with an electronic computer. J. R. Stat. Soc. Ser. B (Methodol.) 1960, 22, 302–306. [Google Scholar] [CrossRef] [Scilit]
  39. Little, R.J.; Rubin, D.B. Statistical Analysis with Missing Data; John Wiley & Sons: New York, NY, USA, 2019. [Google Scholar]
  40. Geskus, R.B. Cause-specific cumulative incidence estimation and the Fine–Gray model under both left truncation and right censoring. Biometrics 2011, 67, 39–49. [Google Scholar] [CrossRef] [Scilit]
  41. Zhang, Z.; Sun, J. Interval censoring. Stat. Methods Med. Res. 2010, 19, 53–70. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Hsu, C.H.; Taylor, J.M.; Murray, S.; Commenges, D. Multiple imputation for interval-censored data with auxiliary variables. Stat. Med. 2007, 26, 769–781. [Google Scholar] [CrossRef] [Scilit]
  43. Rubin, D.B. Multiple Imputation for Nonresponse in Surveys; Wiley: New York, NY, USA, 1987. [Google Scholar]
  44. Willmott, C.J.; Matsuura, K. Advantages of the mean absolute error (MAE) over the root mean square error (RMSE) in assessing average model performance. Clim. Res. 2005, 30, 79–82. [Google Scholar] [CrossRef] [Scilit]
  45. Hastie, T.; Tibshirani, R.; Friedman, J.H. The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed.; Springer: Berlin/Heidelberg, Germany, 2009. [Google Scholar] [CrossRef]
  46. Wang, Y.; Flowers, C.R.; Li, Z.; Huang, X. CondiS: A conditional survival distribution-based method for censored data imputation overcoming the hurdle in machine learning-based survival analysis. J. Biomed. Inform. 2022, 131, 104117. [Google Scholar] [CrossRef] [Scilit]
  47. Zhou, H.; Wang, H.; Wang, S.; Zou, Y. SurvMetrics: An R package for predictive evaluation metrics in survival analysis. R J. 2023, 14, 252–263. [Google Scholar] [CrossRef] [Scilit]
  48. Harrell, F.E.; Califf, R.M.; Pryor, D.B.; Lee, K.L.; Rosati, R.A. Evaluating the yield of medical tests. J. Am. Med. Assoc. 2022, 247, 2543–2546. [Google Scholar] [CrossRef]
  49. Schemper, M. The explained variation in proportional hazards regression. Biometrika 1990, 77, 216–218. [Google Scholar] [CrossRef]
  50. Graf, E.; Schmoor, C.; Sauerbrei, W.; Schumacher, M. Assessment and comparison of prognostic classification schemes for survival data. Stat. Med. 1999, 18, 2529–2545. [Google Scholar] [CrossRef]
  51. Moradian, H.; Larocque, D.; Bellavance, F. L1 splitting rules in survival forests. Lifetime Data Anal. 2017, 23, 671–691. [Google Scholar] [CrossRef] [Scilit]
  52. Qi, S.A.; Kumar, N.; Farrokh, M.; Sun, W.; Kuan, L.H.; Ranganath, R.; Greiner, R. An effective meaningful way to evaluate survival models. arXiv 2023, arXiv:2306.01196. [Google Scholar] [CrossRef] [Scilit]
  53. Haganawiga, K.J.; Pal, S.K.; Sirohi, A. A choice of performance metrics for evaluating predictive accuracy of survival models. Int. J. Stat. Med. Res. 2025, 14, 153–160. [Google Scholar] [CrossRef] [Scilit]
  54. Efron, B.; Tibshirani, R. An Introduction to the Bootstrap; Chapman & Hall/CRC: Boca Raton, FL, USA, 1993. [Google Scholar]
  55. Robert, C.P.; Casella, G. Monte Carlo Statistical Methods; Springer: New York, NY, USA, 1999. [Google Scholar]
  56. Pan, W. Smooth estimation of the survival function for interval-censored data. Stat. Med. 2000, 19, 2611–2624. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Betensky, R.A.; Finkelstein, D.M. A non-parametric maximum likelihood estimator for bivariate interval-censored data. Stat. Med. 1999, 18, 3089–3100. [Google Scholar] [CrossRef] [Scilit]
  58. Finkelstein, D.M.; Goggins, W.B.; Schoenfeld, D.A. Analysis of failure time data with dependent interval censoring. Biometrics 2002, 58, 298–304. [Google Scholar] [CrossRef] [Scilit]
  59. Hackmann, T.; Thomassen, D.; Stiggelbout, A.M.; le Cessie, S.; Putter, H.; de Wreede, L.C.; Steyerberg, E.W. The 4D PICTURE Consortium. Effective sample size for the Kaplan-Meier estimator: A valuable measure of uncertainty? Am. Stat. 2025, 80, 100–108. [Google Scholar] [CrossRef] [Scilit]
  60. Yang, D.; Small, D.S. An R package and a study of methods for computing empirical likelihood. J. Stat. Comput. Simul. 2013, 83, 1363–1372. [Google Scholar] [CrossRef] [Scilit]
  61. Marange, C.S. Finite-sample performance of convex optimization algorithms in empirical likelihood ratio-based goodness-of-fit test statistics. J. Stat. Comput. Simul. 2025, 95, 3330–3352. [Google Scholar] [CrossRef] [Scilit]
  62. Moghaddam, S.; Newell, J.; Hinde, J. A Bayesian approach for imputation of censored survival data. Stats 2022, 5, 89–107. [Google Scholar] [CrossRef] [Scilit]
  63. Shahmirzalou, P.; Rasekhi, A.; Jafari Khaledi, M.; Khayamzadeh, M. Comparison performance of the Bayesian Approach with the Weibull and Birnbaum-Saunders distributions in imputation of time-to-event censors. PLoS ONE 2024, 19, e0295977. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  64. Soutinho, G.; Meira-Machado, L. Imputation of the Response Variable in Survival Analysis with Interval-Censored Data. In International Conference on Computational Science and Its Applications; Springer: Cham, Switzerland, 2025; pp. 78–88. [Google Scholar]
Figure 1. MC simulation survival curves for classical and smoothed KM estimates under four imputation methods at n = 50 and p = 0.8 .
Figure 1. MC simulation survival curves for classical and smoothed KM estimates under four imputation methods at n = 50 and p = 0.8 .
Symmetry 18 00519 g001
Figure 2. Illustrative KM survival curves for the ACTG181 (Treatment 1) data using the midpoint imputation.
Figure 2. Illustrative KM survival curves for the ACTG181 (Treatment 1) data using the midpoint imputation.
Symmetry 18 00519 g002
Table 1. Illustrative simulated dataset with n = 10 at p = 0.8 .
Table 1. Illustrative simulated dataset with n = 10 at p = 0.8 .
IDLeft (L)Right (R)Censoring
123241
224NA0
314151
419201
522231
6561
710111
824NA0
913141
10561
Note: L and R denote the left and right endpoints that bound the unobserved event time T. For observed events (Censoring = 1), T ( L , R ] . For right-censored cases (Censoring = 0), R = NA indicates that T > L (the event was not observed within the study period).
Table 2. Critical sample sizes (n) where the absolute difference between smoothed and classical KM estimators fell below δ = 0.005 across metrics and prevalence rates.
Table 2. Critical sample sizes (n) where the absolute difference between smoothed and classical KM estimators fell below δ = 0.005 across metrics and prevalence rates.
MetricPrevalence RateMidpointRegressionUniformMultiple
0.260404040
0.440303030
MAE0.640303030
0.840303030
1.030203020
0.290606050
0.450404040
NMAE0.640303030
0.850303030
1.030203020
0.2100908070
0.4300140140120
MaxAE0.620010010090
0.8200909080
1.0180708060
0.230202020
0.420202010
MSE0.620202020
0.820202020
1.020202020
0.280606050
0.440303030
NRMSE0.640303030
0.840303030
1.030203020
0.280606050
0.440303030
NARMED0.640303030
0.840303030
1.030203020
Note: Observations at sample sizes n = 10 , 20 , 30 , 40 , 50 , 60 , 70 , 80 , 90 , 100 , 120 , 140 , 160 , 180 , 200 , 300 , 400 , and 500 were considered in determining the convergence thresholds.
Table 3. First ten observations of the ACTG181 Treatment 1 dataset.
Table 3. First ten observations of the ACTG181 Treatment 1 dataset.
IDLeft (L)Right (R)Censoring
11211
211241
33241
410241
56241
68241
710241
86201
98161
1010221
Note: L and R denote the left and right endpoints of the censoring interval for the unobserved event time T. A value of Censoring = 1 indicates that the event occurred within the interval ( L , R ] , whereas Censoring = 0 corresponds to right-censoring with T > L .
Table 4. Approximate critical sample sizes ( n * ) where the absolute difference between smoothed and classical KM estimators falls below δ = 0.005 across metrics for the ACTG181 dataset ( p = 1.0 ).
Table 4. Approximate critical sample sizes ( n * ) where the absolute difference between smoothed and classical KM estimators falls below δ = 0.005 across metrics for the ACTG181 dataset ( p = 1.0 ).
MetricMidpointRegressionUniformMultiple
MAE15252020
NMAE15252020
MaxAE35503025
MSE15302025
NRMSE20302020
NARMED15201520
Note: Sample sizes evaluated for convergence were n = 10 , 15 , 20 , 25 , 30 , 35 , 40 , 45 , 50 , 60 , 70 , 80 , 90 , 100 , 150 , and 200. The values reported represent the smallest n at which | Smoothed Classical | < 0.005 for each metric and imputation method.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Chophela, K.; Marange, C.S.; Odeyemi, A.S. Evaluation of a Non-Parametric Penalized Kaplan–Meier Estimator Under Interval-Censored Survival Data. Symmetry 2026, 18, 519. https://doi.org/10.3390/sym18030519

AMA Style

Chophela K, Marange CS, Odeyemi AS. Evaluation of a Non-Parametric Penalized Kaplan–Meier Estimator Under Interval-Censored Survival Data. Symmetry. 2026; 18(3):519. https://doi.org/10.3390/sym18030519

Chicago/Turabian Style

Chophela, Kayakazi, Chioneso Show Marange, and Akinwumi Sunday Odeyemi. 2026. "Evaluation of a Non-Parametric Penalized Kaplan–Meier Estimator Under Interval-Censored Survival Data" Symmetry 18, no. 3: 519. https://doi.org/10.3390/sym18030519

APA Style

Chophela, K., Marange, C. S., & Odeyemi, A. S. (2026). Evaluation of a Non-Parametric Penalized Kaplan–Meier Estimator Under Interval-Censored Survival Data. Symmetry, 18(3), 519. https://doi.org/10.3390/sym18030519

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop