Next Article in Journal
Lie Algebraic Homotopy 3-Types
Previous Article in Journal
Cost-Sensitive TSK for Recognition of Imbalanced Epilepsy EEG Signal
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Asymptotic Theory for Kernel Density Estimation Under Dependent Length-Biased Sampling

1
LMAC (Laboratory of Applied Mathematics of Compiègne), Université de Technologie de Compiègne, CS 60 319, 60 203 Compiègne, France
2
Department of Statistics and Operations Research, College of Sciences, Qassim University, P.O. Box 6688, Buraydah 51452, Saudi Arabia
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Symmetry 2026, 18(9), 1472; https://doi.org/10.3390/sym18091472
Submission received: 15 July 2026 / Revised: 25 August 2026 / Accepted: 30 August 2026 / Published: 31 August 2026
(This article belongs to the Section B: Mathematics)

Abstract

We establish an asymptotic theory for the Jones inverse-weighted kernel density estimator when length-biased observations form a strictly stationary short-range dependent sequence. The statistical difficulty is intrinsically composite: reciprocal weighting is singular at the origin, the normalizing mean is estimated from the same dependent sample, kernel localization shrinks with the bandwidth, and the centered summands form a row-wise stationary triangular array whose envelope diverges at rate h n 1 . Under a non-negative compactly supported Lipschitz kernel, an inverse-moment condition, geometric α -mixing, local regularity of the target density, and uniform local bounds on lagged bivariate densities, we prove strong uniform consistency on compact subsets of ( 0 , ) and, separately, the uniform stochastic bound O P { h n 2 + ( log n / ( n h n ) ) 1 / 2 } . A covariance-localization argument shows that the scaled serial-covariance contribution is O { h n log ( 1 / h n ) } = o ( 1 ) , so the first-order pointwise variance coincides with that of the corresponding independent length-biased estimator. Pointwise and finite-dimensional Gaussian limits are obtained by an explicit big-block/small-block argument with off-diagonal covariance control. The ratio normalization is treated directly: its variance contribution, its product with the localized fluctuation, and its cross-covariance with that fluctuation are all negligible at the n h n scale. We further derive first-order AMSE and AMISE criteria, their oracle bandwidths, and feasible pointwise studentization under undersmoothing. The numerical study separates oracle from data-driven bandwidth selection, evaluates full-ratio HAC and moving-block corrections, examines a Frank-copula Markov robustness design, and benchmarks the Jones estimator against an alternative length-biased estimator. The simulations support the first-order theory while demonstrating that persistent short-range dependence can remain consequential for finite-sample uncertainty.

1. Introduction

Length-biased sampling is a canonical instance of informative observation: the probability that an individual, object, duration, or lifetime enters the sample is proportional to its magnitude. The mechanism already appears in the stereological work of [1], where larger particles are more likely to intersect a sampling plane, and it was subsequently placed on a systematic statistical footing in [2,3,4]. More generally, weighted distributions arise naturally in renewal theory, survival analysis, reliability, epidemiology, biomedical studies, genetics, econometrics, and cross-sectional sampling; see, among others, refs. [5,6,7,8,9,10,11] and the survey [12]. In each of these settings, the observation law differs systematically from the scientific target so that the statistical problem is not one of ordinary smoothing under a misspecified marginal distribution but of recovering a latent target law through a structurally biased sampling operator.
Let X 0 denote the variable of scientific interest, with distribution function F, density f, and finite positive mean μ = 0 x d F ( x ) . Under length-biased sampling, the observed variable Y has density
g ( t ) = t f ( t ) μ , t > 0 .
Thus, the inferential object is f, whereas the observed process has one-dimensional marginal density g. The mapping f g is multiplicative in the state variable, and its inversion necessarily introduces the reciprocal weight t 1 . This inverse weighting is therefore not an ad hoc bias correction: it is the exact Radon–Nikodym mechanism by which the target law is reconstructed from its length-biased image.
The probabilistic relevance of dependence is particularly transparent in two classical examples. In renewal theory, inspection at a randomly selected epoch favors long renewal intervals and produces the inspection paradox together with the equilibrium relations among total, backward, and forward recurrence times; see [13]. In survival analysis, prevalent-cohort sampling over-represents long survival durations and is closely related, under stationarity, to random left truncation [9]. The corresponding censored setting has generated a substantial literature on unbiased survivor estimation, likelihood-based inference, regression, and semiparametric efficiency; see, for example, refs. [14,15,16,17,18]. Such examples also show why independence may be an unrealistic structural assumption: length-biased observations can inherit temporal, spatial, cluster, or longitudinal dependence from the underlying sampling mechanism.
This distinction has direct inferential consequences. In a prevalent-cohort study assembled in temporally clustered enrollment waves, or in a renewal/reliability monitor whose inspected durations are serially correlated, treating the length-biased observations as independent can leave the inverse-weighted point estimator qualitatively reasonable while materially understating finite-sample uncertainty. In particular, the lagged covariance of the localized inverse-weighted terms can remain numerically important even when it is asymptotically of smaller order. A principal purpose of the present analysis is therefore to identify exactly when the classical independent first-order variance is recovered and, equally importantly, why that first-order equivalence does not guarantee well-calibrated finite-sample studentization under stronger persistence.
The present paper concerns kernel density estimation in this dependent length-biased setting. For ordinary direct observations, kernel density estimation originates with [19,20], with foundational consistency results developed in [21,22,23,24,25,26,27,28]. If a conventional kernel estimator is applied directly to length-biased observations, however, it estimates g rather than f. The sampling distortion must therefore be inverted explicitly. Early kernel methods for length-biased density estimation include [29]; the estimator of [30] uses the natural reciprocal weighting and is closely connected with an inverse-weighted empirical reconstruction of the target distribution. Subsequent developments include the rejection-sampling approach of [31], the minimax analysis of [32], Bayesian bandwidth selection in [33], and strong uniform consistency and asymptotic normality for independent length-biased observations in [34].
Recent work has continued to develop adjacent parts of this literature. Borrajo, González–Manteiga, and Martínez–Miranda [35] provided dedicated bootstrap, cross-validation, and rule-of-thumb bandwidth selection procedures for kernel density estimation with length-biased data. Kakizawa [36] developed asymmetric-kernel density estimators for biased/length-biased non-negative data and established MISE, strong consistency, and asymptotic-normality results in the independent setting. Arvanitis [37] derived non-asymptotic concentration results for ordinary kernel density estimators under uniform ( ϕ -)mixing, a dependence framework distinct from the geometrically strong-mixing inverse-weighted array studied here. Recent work also includes biased-sampling copula-density estimation [38], kernel estimation of mean residual lifetime under length bias [39], a Berry–Esseen bound for a smoothed length-biased distribution estimator [40], kernel estimation of varextropy under length-biased sampling [41], non-parametric residual-extropy estimation under length-biased sampling with a kernel confidence interval [42], simulation from weighted distributions [43], and a recent general treatment of non-parametric function estimation under biased sampling [44]. Their results are complementary rather than substitutes for the present argument: within the literature reviewed here, we have not identified a result that supplies the specific combination used below: a Jones-type random ratio normalizer, geometric α -mixing, local covariance localization for the singular inverse-weighted kernel array, and a row-wise big-block/small-block central limit theorem. This is a literature-positioning statement, not a claim that no other related result can exist. This comparison is intentionally narrow; we do not claim that the present paper exhausts the broader literature on dependent or weighted non-parametric estimation, see Table 1.
The comparison is intentionally restricted to features that bear directly on the present contribution. The feasible selectors and dependence-aware finite-sample corrections are numerical procedures; no new optimality or coverage theorem for those procedures is asserted.
The estimator studied below is consequently not presented as a new estimator. It is a Jones-type inverse-weighted kernel estimator. The contribution of the present work is probabilistic: we establish a dependent-data asymptotic theory for this classical construction under a precisely specified short-range dependence regime. This distinction is essential. The novelty lies neither in the inverse-weighting identity nor in the elementary kernel representation but in the simultaneous control of a random normalizing factor, a reciprocal weight singular at the origin, a shrinking localization window, and serial dependence within a row-wise stationary triangular array.
The connection with symmetry/asymmetry is made at the level of the observation mechanism, not through the routine use of a symmetric smoothing kernel and not through any assumed symmetry of the target density. Length bias replaces the target law F by the tilted observation law G through
d G d F ( x ) = x μ , x > 0 ,
so inclusion is directionally unequal in magnitude: larger values receive systematically greater representation. The factor x 1 , together with the random normalization, is the exact de-tilting operation that removes this informative-sampling asymmetry. This is the precise scientific sense in which asymmetry enters the statistical problem. We do not claim a new group-theoretic symmetry principle, a symmetry property of f, or journal relevance merely because K is symmetric. The resulting asymmetry is therefore intrinsic to the observation operator rather than to the shape of the target density or to the routine symmetry of the smoothing kernel.
Given observations Y 1 , , Y n from a strictly stationary length-biased process, a kernel K, and a deterministic bandwidth h n , we consider
f n ( t ) = μ ^ n n h n i = 1 n 1 Y i K t Y i h n , t > 0 ,
where
μ ^ n = 1 n i = 1 n 1 Y i 1 .
As developed in Section 2, this estimator admits an exact factorization into the scalar normalization μ ^ n / μ and a localized empirical average. That representation is probabilistically decisive: the inverse-moment condition needed to control μ ^ n is global, whereas the singular factor Y i 1 becomes uniformly bounded inside the effective kernel window on every fixed compact subset of ( 0 , ) . The asymptotic analysis can therefore separate a global ergodic normalization problem from a local triangular-array problem.
This separation also identifies the correct scale of the stochastic difficulty. For fixed t > 0 , the localized summand has an L -envelope of order h n 1 and variance of the same order. Hence, the relevant triangular array becomes increasingly concentrated in space while its pointwise amplitude diverges. Uniform convergence and Gaussian approximation must consequently be proved in a regime where localization, dependence, and the bandwidth are coupled. In particular, neither a fixed-envelope empirical-process argument nor a formal appeal to an ordinary stationary central limit theorem is sufficient for the array arising from (2).
The dependent setting therefore requires substantially more than replacing independent observations by an abstract mixing sequence. Uniform control must be based on a concentration inequality whose assumptions agree with the declared dependence regime and on a Lipschitz modulus of continuity of the kernel; bounded variation does not provide the Lipschitz modulus used in the discretization argument. At the distributional level, the centered localized summands form a bandwidth-dependent triangular array with diverging envelope, so the big-block/small-block construction must be quantitatively compatible with both the effective sample size n h n and the rate at which dependence decays.
For these reasons, we work under geometric strong mixing rather than invoking a polynomial α -mixing condition unsupported by the concentration step. The geometric rate is matched to the variance-sensitive Bernstein inequality used in the uniform analysis and to the separated-block approximation entering the central limit theorem. We additionally impose uniform local bounds on the lagged bivariate densities of ( Y 0 , Y k ) . These bounds control short-lag covariances inside shrinking kernel windows and, for finite-dimensional convergence, are required not only near the diagonal ( t , t ) but also in neighborhoods of the off-diagonal pairs ( t j , t ) . The dependence assumptions are thus formulated at the level at which they enter the proofs, rather than through a generic notion of weak dependence.
A central feature of the argument is a covariance-localization principle. The local bivariate-density bound controls a finite set of relatively short lags, where a direct density calculation is sharper than a mixing inequality; the geometric mixing estimate controls sufficiently distant lags, where separation of sigma-fields becomes effective. Splitting the covariance series at a logarithmic lag then yields an absolute serial-covariance contribution of order
h n k 1 Cov { W n , 0 ( s ) , W n , k ( t ) } = O h n log 1 h n = o ( 1 )
uniformly on compact subsets of the positive half-line. This estimate is derived from primitive assumptions and is the mechanism by which the dependence contribution disappears from the first-order variance; no non-zero long-run variance constant is postulated and subsequently cancelled.
This point is important for the interpretation of the limiting distribution. The exact finite-sample variance contains the full lagged covariance sum of the localized inverse-weighted array, and that term may remain numerically substantial under persistent short-range dependence. Under the present assumptions, however, its contribution is asymptotically of smaller order than the diagonal kernel variance after normalization by n h n . The leading variance therefore coincides with that of the corresponding independent length-biased estimator while retaining the non-standard length-bias factor generated by reciprocal weighting. This is a first-order short-range phenomenon; it is not a statement that dependence is negligible at finite sample sizes, nor is it asserted for long-range regimes in which the serial covariance may survive at leading order.
The pointwise central limit theorem is established through an explicit triangular-array blocking argument. The proof controls the L 2 -mass of small blocks and of the terminal remainder, verifies convergence of the variance of the retained big blocks, removes residual inter-block dependence by a characteristic-function decoupling estimate, and proves the Lindeberg condition despite the diverging envelope. This architecture is also what makes the bandwidth–dependence condition transparent: the logarithmic gap length must suppress the mixing remainder while the big-block length must remain negligible relative to n h n . The finite-dimensional result is then obtained by a Cramér–Wold reduction supplemented with explicit off-diagonal covariance estimates, rather than by an unsupported assertion of asymptotic independence.
The support formulation is local rather than compact-support based. No finite right endpoint is imposed on F. Uniform statements are made on fixed compact intervals I = [ a , b ] ( 0 , ) , which separate the effective kernel support from the singularity of y 1 at the origin. This is the natural geometry of the problem: the lower boundary interacts with inverse weighting, whereas an unbounded upper tail creates no analogous obstruction to local smoothing. The theory therefore accommodates positive unbounded-support targets, including the Gamma, lognormal, and Weibull models used in the numerical study, provided the stated moment, local smoothness, and dependence conditions are satisfied.
Within this framework, the uniform and distributional results have deliberately different logical roles. Strong uniform consistency follows from localization, a Bernstein-based grid argument, and the consistency of the random normalization. Under local C 2 -regularity, the corresponding stochastic upper bound separates the deterministic order h n 2 from the empirical order { log n / ( n h n ) } 1 / 2 . Balancing these two proved upper-bound terms yields the reference order h n ( log n / n ) 1 / 5 ; no matching lower bound, exact limsup constant, or minimax statement is claimed. Pointwise Gaussian approximation is obtained under the separate block-compatible bandwidth regime, with either explicit second-order bias correction or undersmoothing. These distinctions are maintained throughout in order to avoid conflating consistency, concentration, and central-limit requirements.
The resulting theory yields strong uniform consistency on compact subsets of ( 0 , ) , a quantitative uniform stochastic upper bound under local twice differentiability, pointwise asymptotic normality, and finite-dimensional Gaussian convergence at pairwise distinct fixed evaluation points. The limiting covariance matrix in the latter result is diagonal because both the same-index and non-zero-lag cross-covariances are shown to be negligible at the kernel scale. The theory further yields first-order pointwise AMSE and integrated AMISE criteria, together with their oracle bandwidths and a feasible pointwise studentized limit under undersmoothing. The AMSE and AMISE quantities are used strictly as first-order bias–variance criteria; they are not identified with exact finite-sample risks without additional uniform-integrability arguments.
The random normalization plays a secondary but non-negligible role in this asymptotic decomposition. Under the inverse-moment and geometric-mixing conditions, μ ^ n μ = O P ( n 1 / 2 ) . Since n 1 / 2 = o { ( n h n ) 1 / 2 } whenever h n 0 , its stochastic contribution is asymptotically smaller than the localized kernel fluctuation at the pointwise scale. This establishes a useful separation principle: first-order non-parametric uncertainty is generated by the shrinking local array, whereas estimation of the normalizing mean enters only as a lower-order Slutsky perturbation under the present short-range regime.
The contribution should therefore be understood as a dependent-data extension of the classical length-biased kernel theory within a specific and verifiable probabilistic regime. No claim is made for arbitrary polynomial mixing, long-range dependence, or models for which the localized covariance sum remains of first order. Nor is the upper uniform rate described as exact or minimax optimal, and the oracle bandwidths are not presented as automatic data-driven selectors. These restrictions are substantive: relaxing them would require different concentration tools, a different covariance theory, or a different asymptotic normalization.
The theoretical analysis is complemented by a Monte Carlo study whose role is to assess, within the models considered, density recovery, integrated risk, Gaussian approximation, studentization, and confidence-interval calibration under increasing short-range dependence. The simulation is interpreted against the first-order theory rather than as a substitute for it: finite-sample variance inflation and imperfect coverage under stronger dependence are fully compatible with asymptotic negligibility of the scaled serial covariance.
The numerical analysis complements the asymptotic theory along four dimensions. First, the oracle AMISE bandwidth is separated from genuinely feasible rules, including a normal-reference construction, weighted cross-validation, two smoothed-bootstrap procedures, and a contiguous block cross-validation rule adapted to dependent observations. These selectors are assessed as finite-sample procedures; no consistency or optimality theorem for the resulting random bandwidth is asserted. Second, the finite-sample undercoverage of the first-order plug-in interval under stronger persistence is examined through a Newey–West/HAC long-run-variance estimator constructed from the complete ratio linearization and through a moving-block resampling experiment. Third, the Gaussian-copula AR(1) design is complemented by a first-order Frank-copula Markov construction calibrated to comparable Kendall dependence, thereby testing the robustness of the numerical findings to the copula mechanism. Fourth, the Jones estimator is compared with an alternative smooth-then-divide length-biased estimator, while an ordinary KDE of the observed length-biased marginal is retained only as a negative control for the sampling distortion. Recent adaptive-kernel methodology such as [45] is used only as general motivation for data-driven smoothing, and the copula-based reliability work of [46] is used only to motivate the second dependence design; neither reference supplies the asymptotic theory developed here.
The theoretical status of bandwidth selection is kept distinct from its numerical implementation. The AMSE/AMISE minimizers derived below are deterministic oracle benchmarks because their constants involve unknown functionals of f. A theorem establishing consistency or optimality of a fully data-driven bandwidth in the present dependent ratio model would require joint stochastic control of the selection criterion, reciprocal weighting, the Jones normalizer, and the bandwidth-indexed dependent empirical process. That random-bandwidth problem is a separate asymptotic question and is not conflated with the deterministic oracle analysis proved in this article.
The remainder of this paper is organized as follows. Section 2 develops the exact statistical reconstruction of the target law, introduces the localized inverse-weighted array, and states the assumptions in the form used by the proofs. Section 3 establishes the uniform results, the pointwise and finite-dimensional central limit theorems, the first-order AMSE/AMISE analysis, and feasible inference. Section 4 examines the finite-sample behavior of the estimator under the positive-support models and dependence regimes considered in the numerical study. Section 6 summarizes the conclusions and delineates the scope of the short-range theory. Finally, Section 7 contains the detailed probabilistic arguments, including the covariance localization, the Bernstein-based concentration analysis, and the abstract triangular-array blocking theorem.

2. Statistical Framework and Assumptions

This section fixes the probabilistic framework in which all subsequent asymptotic statements are proved. We first identify the target distribution from the length-biased law then introduce the inverse-weighted empirical reconstruction and the corresponding kernel estimator. We next isolate the localized triangular array that drives the stochastic asymptotics and, finally, state the assumptions in a form that matches exactly the probability tools used in Section 7. In particular, the assumptions required for the uniform exponential bound, for the covariance analysis, and for the big-block/small-block central limit theorem are stated separately rather than being hidden inside a generic “weak dependence” condition.

2.1. Length-Biased Sampling and Identification

Let X be a non-negative random variable with distribution function F, absolutely continuous with density f, and assume
0 < μ : = E X = 0 x d F ( x ) = 0 x f ( x ) d x < .
The length-biased distribution associated with F is
G ( t ) = 1 μ 0 t x d F ( x ) , t 0 .
Since F is absolutely continuous, G is absolutely continuous on ( 0 , ) with density
g ( t ) = t f ( t ) μ , t > 0 .
Thus, observations drawn from G over-represent large values relative to the target distribution F.
The transformation F G is invertible once the normalizing mean is recovered. Indeed, for every t 0 ,
F ( t ) = μ 0 t 1 y d G ( y ) ,
whereas
0 1 y d G ( y ) = 1 μ 0 1 y y d F ( y ) = 1 μ .
Consequently,
μ = 0 1 y d G ( y ) 1 .
The inverse factor y 1 is therefore not an auxiliary weighting device: it is the exact Radon–Nikodym correction required to recover the target law from the observed length-biased law.
No finite right endpoint is imposed on F. The target distribution may have support extending over the whole positive half-line. All asymptotic statements are local in the evaluation argument and are made either at a fixed t > 0 or uniformly over a fixed compact interval I = [ a , b ] ( 0 , ) . This formulation removes the former incompatibility between a compact-support assumption and the Gamma, lognormal, and Weibull models used in the numerical study.

2.2. Stationary Observations and Inverse-Weighted Empirical Reconstruction

Let { Y i : i Z } be a strictly stationary process with one-dimensional marginal distribution G, and suppose that the observed sample is Y 1 , , Y n . The bi-infinite formulation is adopted only to state the dependence coefficients and covariance identities in their standard stationary form; statistical inference is based solely on the observed finite stretch Y 1 , , Y n .
The empirical distribution function of the observed length-biased sample is
G n ( t ) = 1 n i = 1 n 1 { Y i t } , t 0 .
Replacing G by G n in (9) gives
μ ^ n = 0 1 y d G n ( y ) 1 = 1 n i = 1 n 1 Y i 1 ,
which coincides with (3). Likewise, substituting G n into (7) yields
F n ( t ) = μ ^ n 0 t 1 y d G n ( y ) = μ ^ n n i = 1 n 1 Y i 1 { Y i t } , t 0 .
Since
μ ^ n n i = 1 n 1 Y i = 1 ,
F n is a normalized inverse-weighted empirical distribution function.
For h > 0 , define
K h ( u ) = 1 h K u h .
The estimator (2) can then be written exactly as the kernel smoothing of F n :
f n ( t ) = 0 K h n ( t u ) d F n ( u ) = μ ^ n n i = 1 n 1 Y i K h n ( t Y i ) , t > 0 .
It is useful to remove the random normalizing factor from the localized kernel average. Define
A n ( t ) = 1 n i = 1 n 1 Y i K h n ( t Y i ) , t > 0 .
Then
f n ( t ) = μ ^ n A n ( t ) .
Equivalently, with 
W n , i ( t ) = μ Y i K h n ( t Y i ) = μ Y i h n K t Y i h n ,
one has
μ A n ( t ) = 1 n i = 1 n W n , i ( t ) , f n ( t ) = μ ^ n μ 1 n i = 1 n W n , i ( t ) .
The stochastic theory is therefore reduced to a localized, row-wise stationary triangular array together with the scalar random normalization μ ^ n / μ .

2.3. Canonical Bias and Variance Scales

The inverse weighting removes the length bias exactly at the expectation level. Using (6),
E A n ( t ) = 0 1 y K h n ( t y ) g ( y ) d y = 1 μ 0 K h n ( t y ) f ( y ) d y .
Thus
μ E A n ( t ) = 0 K h n ( t y ) f ( y ) d y .
Along any deterministic bandwidth sequence satisfying h n 0 , at every fixed interior point t > 0 , if f is twice continuously differentiable in a neighborhood of t, the moment conditions on K stated below give
μ E A n ( t ) f ( t ) = h n 2 2 m 2 ( K ) f ( t ) + o ( h n 2 ) ,
where
m 2 ( K ) : = R u 2 K ( u ) d u .
Hence, the leading deterministic bias has the classical second-order kernel form.
The leading stochastic scale is different because inverse weighting changes the local second moment. Along any deterministic bandwidth sequence satisfying h n 0 , at every fixed t > 0 where f is continuous,
h n Var { W n , 0 ( t ) } μ f ( t ) t R ( K ) , R ( K ) : = R K 2 ( u ) d u .
For dependent observations, strict stationarity yields the exact finite-sample identity
Var n h n 1 n i = 1 n W n , i ( t ) E W n , 0 ( t ) = h n Var { W n , 0 ( t ) } + 2 h n k = 1 n 1 1 k n Cov { W n , 0 ( t ) , W n , k ( t ) } .
The second term is retained as a finite-n serial-covariance contribution. It is not assumed to converge to a separate long-run-variance constant. Under the primitive short-range conditions below and the basic bandwidth requirement h n 0 , Section 7 proves the stronger absolute estimate
h n k = 1 Cov { W n , 0 ( s ) , W n , k ( t ) } = O h n log 1 h n = o ( 1 ) ,
uniformly for s , t in a fixed compact interval contained in ( 0 , ) . In particular, dependence can affect finite-sample variability while remaining asymptotically negligible at the n h n scale. No analogous conclusion is claimed for long-range dependence or for models violating the local bivariate-density condition.

2.4. Strong Mixing and Local Geometry

For integers a b , write
F a b = σ ( Y i : a i b ) .
The strong-mixing coefficients of the stationary process are
α ( k ) = sup m Z sup A F m B F m + k P ( A B ) P ( A ) P ( B ) , k 1 .
The estimation domain is local. Let I = [ a , b ] ( 0 , ) be fixed and let the deterministic bandwidth satisfy h n 0 . Because supp ( K ) [ 1 , 1 ] , for every sufficiently large n,
K h n ( t Y i ) 0 , t I Y i I n : = [ a h n , b + h n ] I : = [ a / 2 , b + 1 ] .
Consequently,
Y i 1 2 a
whenever the localized summand is non-zero. Thus, the singularity of y 1 at the origin is separated from the local kernel analysis by the fixed condition a > 0 . The global inverse-moment assumption below is needed for the random normalizing factor, not for the local boundedness of W n , i ( t ) .

2.5. Assumptions

The following assumptions are the primitive hypotheses used in the proofs. Their formulation is intentionally aligned with the exact probabilistic arguments in Section 7.
Assumption 1
(Kernel). The function K : R R is bounded, non-negative, symmetric, compactly supported on [ 1 , 1 ] , and globally Lipschitz. Thus, there exists L K < such that
| K ( u ) K ( v ) | L K | u v | , u , v R .
Moreover,
R K ( u ) d u = 1 , R u K ( u ) d u = 0 ,
and
m 2 ( K ) : = R u 2 K ( u ) d u ( 0 , ) , R ( K ) : = R K 2 ( u ) d u ( 0 , ) .
The positivity of R ( K ) is automatic from K = 1 : the kernel is not zero almost everywhere. It is recorded explicitly because σ L B 2 ( t ) must be strictly positive whenever f ( t ) > 0 . The non-degeneracy condition m 2 ( K ) > 0 is required only when the second-order AMSE/AMISE minimizers are written explicitly; if m 2 ( K ) = 0 , the leading bias order changes and the bandwidth calculation must be reformulated. The Lipschitz property (26) is an explicit assumption: bounded variation is not used as a substitute for Lipschitz continuity in the discretization argument.
Assumption 2
(Length-biased model). The target distribution F is absolutely continuous on [ 0 , ) , with density f, and 
0 < μ = 0 x f ( x ) d x < .
The stationary observations have marginal distribution G defined by (5), equivalently density
g ( t ) = t f ( t ) μ , t > 0 .
No finite upper endpoint of the support is imposed.
Assumption 3
(Local regularity of the target density). The density f is continuous on ( 0 , ) . For a fixed compact I = [ a , b ] ( 0 , ) , the following additional local conditions are invoked only when required:
(i) 
for uniform consistency, f is bounded and uniformly continuous on an open neighborhood of I;
(ii) 
for the uniform second-order bias bound, f C 2 on an open neighborhood of I and f is bounded there;
(iii) 
for a pointwise bias-corrected central limit theorem at t > 0 , f C 2 on a neighborhood of t; for a finite-dimensional bias-corrected central limit theorem at pairwise distinct points t 1 , , t m > 0 , f C 2 on a neighborhood of each t j ;
(iv) 
for the Hölder version, there exist β ( 0 , 1 ] , C t < , and  ε t > 0 such that
| f ( t + u ) f ( t ) | C t | u | β , | u | ε t .
Assumption 4
(Inverse moment). There exists q > 2 such that
E ( Y q ) < .
In particular,
E ( Y 1 ) = μ 1 < .
By the length-biased identity,
E ( Y q ) = 1 μ 0 y 1 q f ( y ) d y .
Thus, the restriction concerns the lower tail of the target distribution. A convenient sufficient condition is
f ( y ) = O ( y γ ) as y 0
for some γ > q 2 . No upper-tail inverse-moment condition is needed.
Assumption 5
(Geometric strong mixing). The process { Y i : i Z } is strictly stationary, and its strong-mixing coefficients (24) satisfy
α ( k ) C α e c α k , k 1 ,
for some constants C α < and c α > 0 . This geometric rate is the dependence hypothesis used both in the variance-sensitive Bernstein inequality and in the separated-block characteristic-function approximation. The main results are therefore not asserted under a merely polynomial mixing rate without a separate Fuk–Nagaev/Rio-type argument.
Assumption 6
(Uniform local bivariate densities). For every k 1 , the vector ( Y 0 , Y k ) admits a Lebesgue density g k on ( 0 , ) 2 . For every compact set C ( 0 , ) 2 ,
sup k 1 sup ( x , y ) C g k ( x , y ) < .
This condition is imposed uniformly in the lag k. At a single evaluation point, it controls the local covariance near ( t , t ) ; for the finite-dimensional central limit theorem, it also controls neighborhoods of the off-diagonal pairs ( t j , t ) , j . The latter off-diagonal control is essential and is not inferred from a diagonal density bound.
Assumption 7
(Bandwidth regimes). The bandwidth sequence { h n } is deterministic and
h n 0 , n h n .
Additional restrictions are imposed only for the corresponding conclusions. Every theorem-specific condition below is cumulative with (33); in particular, none of (34)(41) is to be read as a substitute for h n 0 and n h n .
(i) 
For the Bernstein-based almost-sure uniform fluctuation bound for the centered localized empirical term and the resulting strong uniform consistency of the full ratio estimator (without asserting an almost-sure rate for the full estimator),
n h n ( log n ) 5 + η f o r s o m e η > 0 .
(ii) 
For the pointwise and finite-dimensional big-block/small-block central limit theorems,
n h n ( log n ) 4 .
With
p n = n h n log n , q n = A log n ,
where A > A 0 and A 0 < is chosen sufficiently large as a function only of the geometric mixing constants in Assumption 5 so that the separated-block covariance and decoupling bounds used in Section 7 are summable. Then, (35), together with (33), implies
p n , q n , q n p n 0 , p n n h n 0 ,
and, writing
r n = n p n + q n ,
one also has
r n α ( q n ) 0 .
(iii) 
For the second-order bias-corrected pointwise central limit theorem,
n h n 5 = O ( 1 ) .
(iv) 
For the uncorrected undersmoothed central limit theorem,
n h n 5 0 .
(v) 
Under local Hölder smoothness of order β ( 0 , 1 ] , the uncorrected centered limit requires
n h n 1 + 2 β 0 ,
together with the central-limit condition (35).
Remark 1
(Cumulative bandwidth requirements). Whenever a result in Section 3, Section 4, Section 5, Section 6 and Section 7 invokes one of the specialized bandwidth restrictions (34)(41), the basic requirements (33) remain in force. The theorem-specific restrictions are therefore cumulative: none of the logarithmic or bias conditions replace the basic localization requirement h n 0 , see Table 2.

2.6. Scope and Logical Role of the Assumptions

Remark 2
(Localization and inverse-moment control). Two distinct mechanisms are involved. The compact-support property of K and the restriction I ( 0 , ) imply that Y i 1 is uniformly bounded whenever a localized kernel summand is non-zero. Assumption 4 is instead used to control the global normalizing statistic
n 1 i = 1 n Y i 1
and to obtain μ ^ n μ = O P ( n 1 / 2 ) .
Remark 3
(Hypotheses entering the uniform exponential bound). The uniform exponential argument uses Assumptions 1, 2, 3, 4, 5, and 6, together with the basic bandwidth requirement (33) and its uniform-concentration strengthening (34). The bivariate-density assumption enters through the uniform covariance bound required to obtain the variance proxy of order h n 1 in the Bernstein inequality. Thus, the local bivariate-density condition is an explicit hypothesis of the uniform theorem, entering precisely through the variance proxy required by the Bernstein bound.
Remark 4
(Short-range dependence and first-order variance). Assumptions 5 and 6 define the short-range regime considered here. They do not imply that serial dependence is numerically negligible at moderate sample sizes. Rather, they imply the asymptotic estimate (23); consequently, the serial covariance is of smaller order than the diagonal h n 1 -variance contribution. Hence, pronounced finite-sample dependence is entirely compatible with first-order asymptotic variance equivalence.
Remark 5
(Off-diagonal density control). For distinct evaluation points t j t , compact support of the kernel eventually gives
W n , 0 ( t j ) W n , 0 ( t ) = 0 .
This does not imply that the same-index covariance is zero, because 
Cov { W n , 0 ( t j ) , W n , 0 ( t ) } = E W n , 0 ( t j ) E W n , 0 ( t ) .
The scaled covariance nevertheless vanishes after multiplication by h n . For non-zero lags, the proof requires local bounds for g k near ( t j , t ) , which is precisely why Assumption 6 is formulated on arbitrary compact subsets of ( 0 , ) 2 rather than only near the diagonal.
Remark 6
(Separation of bandwidth roles). The basic condition (33) is logically prior to all the theorem-specific restrictions: h n 0 is required for localization and approximation of the identity, while n h n is the baseline effective-sample-size condition. Condition (34) controls the logarithmic linear term in the Bernstein denominator and the subsequent union bound over the discretization grid. Condition (35) guarantees the simultaneous small-block, Lindeberg, and block-separation relations in (37) and (38). Conditions (39) and (40) concern only the deterministic smoothing bias at the central-limit scale. Thus, the displayed bandwidth conditions have distinct logical roles and are cumulative whenever they are invoked together.
Remark 7
(Interpretation of the balanced uniform bound). The balance
h n 2 log n n h n
leads to the reference order h n ( log n / n ) 1 / 5 . This balances the two terms in the proved upper stochastic bound; it does not provide a lower bound, an exact limsup constant, or a minimax theorem. Accordingly, no “exact” or “optimal” uniform convergence-rate claim is made.
Remark 8
(Compatibility with the numerical target models). Assumption 2 allows unbounded positive support. Hence, Gamma, lognormal, and Weibull target densities are not excluded by the theoretical framework. For each numerical model, compatibility with the theory reduces to verifying the finite positive mean, the relevant local smoothness conditions, and the inverse-moment requirement (29), together with the dependence assumptions for the chosen stationary sampling mechanism. These compatibility checks are verified explicitly in Section 4; in particular, the numerical study uses a common admissible inverse-moment order for all three target models and a stationary short-range dependence construction.

3. Asymptotic Results

This section states the principal asymptotic results for the inverse-weighted kernel estimator (2) under the statistical framework of Section 2. The formulation is deliberately local: throughout,
I = [ a , b ] ( 0 , ) , 0 < a < b < ,
is a fixed compact estimation interval, while t > 0 denotes a fixed evaluation point. No finite right endpoint of the support of F is assumed.
For later reference, define the leading second-order smoothing bias
B n ( t ) : = h n 2 2 m 2 ( K ) f ( t ) ,
whenever f ( t ) exists, and define the length-bias-corrected variance constant
σ L B 2 ( t ) : = μ f ( t ) t R ( K ) , R ( K ) = R K 2 ( u ) d u .
The results are established under the short-range regime of Assumptions 5 and 6. The dependence contribution is not replaced by an assumed long-run variance. Instead, the proofs show directly that the scaled serial covariance is O { h n log ( 1 / h n ) } = o ( 1 ) . Thus, the first-order variance (43) is a theorem derived from the primitive assumptions, not an assumption imposed on the covariance structure.
Throughout this section, every theorem-specific bandwidth restriction is understood cumulatively with the basic bandwidth regime (33); in particular, h n 0 and n h n remain in force whenever (44), (50), or a bias-scale restriction is invoked. This convention is repeated in the principal theorem statements below whenever it is logically essential.
The section is organized as follows. Section 3.1 gives strong uniform consistency and a uniform stochastic upper rate. Section 3.2 establishes pointwise and finite-dimensional central limit theorems. Section 3.3 develops the corresponding first-order AMSE and AMISE criteria. Section 3.4 gives feasible studentization and pointwise confidence intervals. Section 3.5 records representative models and clarifies the scope of the assumptions.

3.1. Uniform Consistency and a Uniform Stochastic Upper Rate

The first theorem is qualitative and almost sure. The second gives a quantitative upper stochastic rate. The distinction is important: no matching lower bound, exact limsup constant, or minimax assertion is made.
Theorem 1
(Strong uniform consistency). Let I = [ a , b ] ( 0 , ) . Suppose Assumptions 1–6 hold, that f is bounded and uniformly continuous on an open neighborhood of I, and that the basic bandwidth condition (33) holds. Assume in addition the uniform-concentration bandwidth condition
n h n ( log n ) 5 + η for some η > 0 .
Then
sup t I | f n ( t ) f ( t ) | a . s . 0 .
Remark 9
(Strength of the uniform-concentration bandwidth condition). Condition (44) is not a generic kernel requirement and does not replace the basic condition (33). It is the additional compatibility condition needed by the variance-sensitive Bernstein inequality used in Section 7. For the localized summands, the envelope is O ( h n 1 ) , the variance proxy is O ( h n 1 ) , and the Bernstein denominator contains the term x M n ( log n ) 2 . At the target deviation
ε n log n n h n ,
condition (44) guarantees
ε n ( log n ) 2 = O ( log n ) 5 / 2 n h n 0 .
The condition is therefore intrinsic to the variance-sensitive exponential argument employed here.
Theorem 2
(Uniform stochastic upper rate). Under the assumptions of Theorem 1, suppose in addition that f C 2 on an open neighborhood of I and that f is bounded there. Then
sup t I | f n ( t ) f ( t ) | = O P h n 2 + log n n h n .
In particular, the deterministic order
h n log n n 1 / 5
balances the two terms on the right-hand side of (46), giving
sup t I | f n ( t ) f ( t ) | = O P log n n 2 / 5 .
The order in (48) is the consequence of the proved upper bound only; it is not asserted to be an exact, sharp, or minimax-optimal uniform rate. Moreover, the balancing choice (47) is admissible under the theorem’s concentration hypothesis, since
n h n ( log n ) 5 + η n 4 / 5 ( log n ) 24 / 5 + η .
Thus, the displayed “in particular” statement is obtained within, and not outside, the stated bandwidth regime.
Remark 10
(Almost-sure consistency and the stochastic upper bound). Theorems 1 and 2 have distinct logical content. The former proves only
sup t I | f n ( t ) f ( t ) | 0 almost surely .
The latter proves the quantitative statement
sup t I | f n ( t ) f ( t ) | = O P h n 2 + log n n h n .
The almost-sure bound in Lemma 9 concerns the centered localized empirical term A n E A n ; it is not promoted to an almost-sure rate for the full ratio estimator because the normalizing factor is only assigned the quantitative order μ ^ n μ = O P ( n 1 / 2 ) in the present paper. Accordingly, the expressions “strong uniform rate”, “exact almost-sure rate”, and equivalent terminology are not used for (46); throughout, it is called a uniform stochastic upper rate or a uniform stochastic upper bound.
Remark 11
(Logical scope of the uniform results). The condition (44) by itself does not imply h n 0 . The consistency and rate statements therefore require the basic bandwidth condition (33) in addition to the logarithmic concentration condition. The distinction is substantive because h n 0 controls the deterministic approximation bias, whereas (44) controls the stochastic concentration argument.
Remark 12
(Contribution of the random normalization). The decomposition used in the proof is
f n ( t ) f ( t ) = μ ^ n { A n ( t ) E A n ( t ) } + ( μ ^ n μ ) E A n ( t ) + { μ E A n ( t ) f ( t ) } .
The first term is of uniform stochastic order log n / ( n h n ) , the third is O ( h n 2 ) , and 
μ ^ n μ = O P ( n 1 / 2 ) .
Since
n 1 / 2 = o log n n h n ,
the random normalization is asymptotically of smaller order than the localized kernel fluctuation.
Remark 13
(Local estimation domain and unbounded target support). The condition I ( 0 , ) is imposed on the estimation region, not on the support of F. Its purpose is to separate the effective kernel support from the singularity of y 1 at zero. The target density may have unbounded support.

3.2. Pointwise and Finite-Dimensional Asymptotic Normality

For a fixed t > 0 , define the finite-n serial covariance contribution
L n ( t ) : = 2 h n k = 1 n 1 1 k n Cov { W n , 0 ( t ) , W n , k ( t ) } .
The notation L n ( t ) is intentionally finite-sample. No assumption is made that L n ( t ) converges to an unknown non-zero constant. Under the short-range conditions of Section 2, its absolute value is proved to converge to zero.
Theorem 3
(Pointwise asymptotic normality). Fix t > 0 with f ( t ) > 0 . Suppose Assumptions 1–6 hold, that f C 2 on a neighborhood of t, and that the basic bandwidth condition (33) holds. Assume in addition
n h n ( log n ) 4
and
n h n 5 = O ( 1 ) .
Then
n h n f n ( t ) f ( t ) B n ( t ) D N 0 , σ L B 2 ( t ) ,
where B n ( t ) and σ L B 2 ( t ) are given by (42) and (43), respectively.
Moreover,
| L n ( t ) | 2 h n k = 1 n 1 Cov { W n , 0 ( t ) , W n , k ( t ) } = O h n log 1 h n = o ( 1 ) .
Consequently, the limiting variance is
σ L B 2 ( t ) = μ f ( t ) t R ( K ) .
Proposition 1
(Explicit normalization–localization separation). Under the hypotheses of Theorem 3, define
D n : = μ ^ n μ 1 , U n ( t ) : = 1 n i = 1 n W n , i ( t ) E W n , 0 ( t ) .
Then
D n = O P ( n 1 / 2 ) , U n ( t ) = O P ( n h n ) 1 / 2 ,
and
n h n D n U n ( t ) = o P ( 1 ) , n h n D n E W n , 0 ( t ) = o P ( 1 ) .
Moreover, with  H n = n 1 i = 1 n Y i 1 and H = μ 1 ,
Cov { H n H , U n ( t ) } = O 1 n h n ,
so the covariance between the first-order linearization of the normalization error and the localized kernel fluctuation is negligible relative to the leading ( n h n ) 1 variance:
n h n Cov μ E W n , 0 ( t ) ( H n H ) , U n ( t ) = O ( h n ) 0 .
Moreover, the variance of the linearized pure-normalization term is itself strictly lower order:
n h n Var μ E W n , 0 ( t ) ( H n H ) = O ( h n ) 0 .
Hence, the diagonal normalization variance, the normalization–localization cross-covariance, and the non-linear product remainder all disappear at the first-order kernel scale; none are absorbed into an unspecified Slutsky term. Consequently,
n h n { f n ( t ) f ( t ) B n ( t ) } = n h n U n ( t ) + o P ( 1 ) .
Thus, both the magnitude of the random normalization and its first-order covariance with the localized stochastic term are explicitly negligible at the claimed Gaussian scale.
Remark 14
(First-order variance equivalence and finite-sample dependence). Theorem 3 makes a first-order asymptotic statement. The covariance sum in (49) can be substantial for moderate n, especially under persistent dependence. Under geometric α-mixing and uniform local bivariate-density bounds, the theorem establishes that this contribution is of smaller order than the diagonal h n 1 -variance term after the n h n normalization. This distinction is relevant for the interpretation of the simulation results and of finite-sample confidence-interval coverage.
Remark 15
(Block-size compatibility for the triangular-array CLT). The proof of Theorem 3 uses
p n = n h n log n , q n = A log n
with A > A 0 , where A 0 < is chosen sufficiently large as a function only of the geometric mixing constants in Assumption 5. Conditions (33) and (50) imply
q n p n = O ( log n ) 2 n h n 0 , p n n h n 0 ,
and, if 
r n = n p n + q n ,
then
r n α ( q n ) 0 .
Thus, the small-block negligibility, the Lindeberg bound, and the asymptotic decoupling of the big blocks are all consequences of an explicit bandwidth–dependence relation.
Theorem 4
(Joint asymptotic normality at distinct points). Let t 1 , , t m be pairwise distinct points contained in a fixed compact interval I ( 0 , ) , and assume f ( t j ) > 0 for every j. Suppose Assumptions 1–6 hold, the basic bandwidth condition (33) holds, and  f C 2 on a neighborhood of each t j . Assume further (50) and (51). Then
n h n f n ( t j ) f ( t j ) B n ( t j ) j = 1 m D N m ( 0 , Σ ) ,
where
Σ = diag σ L B 2 ( t 1 ) , , σ L B 2 ( t m ) .
Hence, the limiting Gaussian coordinates at pairwise distinct fixed points are independent. This is an asymptotic decorrelation statement at the kernel scale; it does not assert independence of the finite-sample estimators.
Remark 16
(Necessity of off-diagonal density control). For t j t , compact support of K implies, eventually,
W n , 0 ( t j ) W n , 0 ( t ) = 0 .
This does not imply zero same-index covariance:
Cov { W n , 0 ( t j ) , W n , 0 ( t ) } = E W n , 0 ( t j ) E W n , 0 ( t ) = O ( 1 ) .
Its contribution is negligible only after multiplication by h n . For non-zero lags, the proof requires the uniform bounds on g k in neighborhoods of the off-diagonal pairs ( t j , t ) . Accordingly, Assumption 6 is stated on arbitrary compact subsets of ( 0 , ) 2 , not merely in a neighborhood of the diagonal.
Corollary 1
(Undersmoothed central limit theorem). Under the hypotheses of Theorem 3, strengthen the bias-scale condition (51) to
n h n 5 0 .
Then
n h n { f n ( t ) f ( t ) } D N 0 , σ L B 2 ( t ) .
Corollary 2
(Bias-corrected central limit theorem). Under the hypotheses of Theorem 3, the second-order centered limit is
n h n f n ( t ) f ( t ) h n 2 2 m 2 ( K ) f ( t ) D N 0 , σ L B 2 ( t ) .
In particular, any bandwidth satisfying h n n 1 / 5 obeys (33), (50), and (51) and is therefore covered by this bias-corrected statement.
Corollary 3
(Local Hölder smoothness). Fix t > 0 with f ( t ) > 0 . Suppose Assumptions 1, 2, 4, 5, and 6 hold, the basic bandwidth condition (33) holds, and assume that for some β ( 0 , 1 ] , C t < , and  ε t > 0 ,
| f ( t + u ) f ( t ) | C t | u | β , | u | ε t .
If
n h n ( log n ) 4
and
n h n 1 + 2 β 0 ,
then
n h n { f n ( t ) f ( t ) } D N 0 , σ L B 2 ( t ) .
Remark 17
(Hölder bias and undersmoothing). Condition (66) is exactly the requirement that the local deterministic bias O ( h n β ) be negligible at the stochastic scale:
n h n h n β = n h n 1 + 2 β 0 .

3.3. First-Order Mean Squared Error and Oracle Bandwidths

The next statements concern the first-order asymptotic risk criterion implied by the bias and variance expansions above. They do not claim that the resulting bandwidth minimizes the exact finite-sample MSE.
Throughout this subsection, the first-order risk expansions are understood under Assumptions 1, 2, 4, 5, and 6, together with the basic bandwidth regime (33) and the local smoothness conditions stated for the corresponding point or compact interval. These assumptions are precisely those needed to make the random-normalization contribution and the scaled serial-covariance contribution lower order relative to the leading kernel variance.
Proposition 2
(Pointwise AMSE and oracle AMSE bandwidth). Let t > 0 satisfy f ( t ) > 0 and f ( t ) 0 , and suppose f C 2 on a neighborhood of t. Under the standing conditions of this subsection, the pointwise first-order AMSE criterion is
AMSE t ( h ) = h 4 4 m 2 ( K ) 2 f ( t ) 2 + 1 n h μ f ( t ) t R ( K ) .
Its unique positive minimizer is
h AMSE ( t ) = μ f ( t ) R ( K ) t m 2 ( K ) 2 f ( t ) 2 1 / 5 n 1 / 5 ,
and the minimum of the leading criterion is
AMSE t { h AMSE ( t ) } = 5 4 m 2 ( K ) 2 f ( t ) 2 1 / 5 μ f ( t ) t R ( K ) 4 / 5 n 4 / 5 .
For a common bandwidth on I = [ a , b ] ( 0 , ) , assume f C 2 on a neighborhood of I and
0 < I f ( t ) 2 d t < .
The corresponding first-order integrated criterion is
AMISE I ( h ) = h 4 4 m 2 ( K ) 2 I f ( t ) 2 d t + R ( K ) n h I μ f ( t ) t d t .
Its unique positive minimizer is
h AMISE ( I ) = R ( K ) I μ f ( t ) t 1 d t m 2 ( K ) 2 I f ( t ) 2 d t 1 / 5 n 1 / 5 .
Hence, the AMISE-minimizing order is n 1 / 5 , and the minimum first-order AMISE is of order n 4 / 5 .
Remark 18
(Oracle and feasible bandwidths). The quantities h AMSE ( t ) and h AMISE ( I ) depend on unknown functionals of f and are therefore oracle bandwidths. A practical rule having the form C n 1 / 5 , or a scale-adjusted version of that order, is not the theoretical AMSE minimizer unless its constant consistently estimates the functional appearing in (69) or (72). Accordingly, any numerical bandwidth used only as a practical reference rule must be described as such, rather than as the “optimal” bandwidth.
Remark 19
(Limits of the deterministic bandwidth theory). The oracle expressions above identify a first-order bias–variance benchmark, but they do not solve the practical random-bandwidth problem under dependent length-biased sampling. A full theory for a feasible selector would require uniform control of the bandwidth criterion and of the Jones ratio estimator evaluated at a random bandwidth under serial dependence. The feasible selectors investigated in Section 4 are therefore treated as computational procedures. A satisfactory consistency, rate, and optimality theory for such random selectors requires an independent theoretical treatment and is not asserted here.
Remark 20
(Interpretation of AMSE and AMISE). The criterion (68) is defined from the leading deterministic squared bias and the leading variance. The same convention applies to AMISE I ( h ) : it is the integrated first-order bias–variance criterion justified by the uniform bias, diagonal-variance, and serial-covariance expansions, not a claim of an exact finite-sample integrated MSE identity. The distinction does not alter the first-order minimization, but it is essential for interpreting AMSE/AMISE as asymptotic criteria rather than exact finite-sample risk identities.

3.4. Feasible Inference

For fixed t > 0 with f ( t ) > 0 , define
σ ^ L B 2 ( t ) : = μ ^ n f n ( t ) t R ( K ) .
The positivity issue is explicit. Assumption 1 now requires K 0 . Since Y i > 0 , μ ^ n > 0 whenever it is defined and therefore f n ( t ) 0 for every sample. Hence, σ ^ L B 2 ( t ) 0 identically. At an inferential point satisfying f ( t ) > 0 , pointwise consistency gives
P { f n ( t ) > f ( t ) / 2 } 1 ,
and, because  R ( K ) > 0 , it follows that P { σ ^ L B 2 ( t ) > 0 } 1 . Thus, the square root below is real-valued with probability tending to one; no sign-indefinite plug-in variance is hidden in the studentization argument. On the event { σ ^ L B 2 ( t ) > 0 } , write
σ ^ L B ( t ) : = { σ ^ L B 2 ( t ) } 1 / 2 .
Proposition 3
(Consistent variance estimation and studentized limit). Under the hypotheses of Theorem 3,
σ ^ L B 2 ( t ) P σ L B 2 ( t ) > 0 .
In fact, for every fixed constant c t satisfying 0 < c t < σ L B 2 ( t ) ,
P σ ^ L B 2 ( t ) c t 1 .
Thus, the denominator is not only positive with probability tending to one; it is asymptotically bounded away from zero in probability. If, in addition,
n h n 5 0 ,
define the everywhere-defined studentized statistic
T n ( t ) : = n h n { f n ( t ) f ( t ) } σ ^ L B ( t ) , σ ^ L B 2 ( t ) > 0 , 0 , σ ^ L B 2 ( t ) 0 .
Then
T n ( t ) D N ( 0 , 1 ) .
Moreover,
P { σ ^ L B 2 ( t ) > 0 } 1 .
Corollary 4
(Asymptotic pointwise confidence interval). Under the hypotheses of Proposition 3, with  n h n 5 0 , an asymptotic 100 ( 1 α ) % pointwise confidence interval for f ( t ) , defined on { σ ^ L B 2 ( t ) > 0 } , whose probability tends to one, is
f n ( t ) ± z 1 α / 2 σ ^ L B 2 ( t ) n h n 1 / 2 ,
where z 1 α / 2 is the ( 1 α / 2 ) -quantile of the standard normal distribution.
Remark 21
(Scope of first-order plug-in variance estimation). The estimator (73) is valid because (53) establishes that the scaled serial covariance term is asymptotically negligible under the stated short-range conditions. If the geometric mixing or local bivariate-density assumptions are relaxed, the present proposition is no longer justified by the arguments of this paper. A non-negligible long-run covariance term may then require a different asymptotic normalization and a separate variance estimator; its existence is not postulated here.
Remark 22
(Finite-sample limitation of first-order plug-in inference). The interval (78) is a first-order pointwise asymptotic interval, and its variance estimator retains only the first-order variance in (54). Although the scaled serial covariance vanishes asymptotically under Assumptions 5 and 6, the unscaled finite-n covariance sum can remain large under stronger short-range dependence. The resulting undercoverage is therefore a substantive finite-sample limitation of the present plug-in studentization, not merely a cosmetic numerical transient. In regimes such as those represented by ϕ = 0.6 and especially ϕ = 0.8 in Section 4, one should not infer adequate calibration from the first-order limit alone. The HAC and dependent-resampling procedures examined numerically in Section 4 address this finite-sample issue, while the formal asymptotic validity of such corrections under the present ratio model remains a separate theoretical problem.

3.5. Examples and Scope of Application

The following examples separate two issues that are sometimes conflated: the occurrence of length bias and the verification of the dependence assumptions. Length bias can arise naturally in many applications, but geometric α -mixing and the uniform local bivariate-density condition must still be justified for the particular data-generating mechanism.
Example 1
(Motivating example: equilibrium renewal sampling and the inspection paradox). Let T denote a generic renewal interval with density f T and 0 < μ T = E T < , and inspect an equilibrium renewal process at a single stationary epoch. If C is the total length of the renewal interval containing that epoch, then
g C ( c ) = c f T ( c ) μ T , c > 0 .
Thus, C, not the backward or forward recurrence time separately, has the length-biased total-lifetime law. Writing A for the backward recurrence time and R for the forward recurrence time, one has C = A + R , while under the classical equilibrium construction the marginal densities of A and R are proportional to the survival function of T, rather than to t f T ( t ) . The estimator of this paper therefore applies directly only when the observed marginal variable is genuinely a covering-interval length with density c f T ( c ) / μ T .
If covering-interval lengths are recorded at a sequence of inspection epochs, the resulting sequence may be dependent, but equilibrium renewal sampling by itself does not imply the geometric α-mixing condition (31). To avoid asserting an unproved primitive renewal theorem, this example is presented as motivation only. Application of Theorems 1–4 requires a separately verified stationary sampling scheme satisfying Assumptions 4, 5, and 6; otherwise, the renewal construction should not be advertised as a verified application of the main theorems.
Example 2
(Motivating example: prevalent-cohort survival sampling). Let T denote the total lifetime from disease onset to failure, with density f T and mean μ T . Under the ideal stationary-incidence prevalent-cohort mechanism, ascertainment at a cross-sectional time over-represents long total lifetimes; the total lifetime of an ascertained case has the length-biased density
g T ( t ) = t f T ( t ) μ T , t > 0 .
This statement must not be confused with the distribution of the time already elapsed from onset to examination. If A is that backward recurrence time (age at sampling) and R is the forward residual lifetime, then T = A + R ; in the classical stationary construction, conditional on T = t , the inspection position is uniform over the interval, whereas the marginal laws of A and R are equilibrium residual-life laws, not the length-biased density t f T ( t ) / μ T . Accordingly, the estimator (2) applies directly to a marginal variable only when that observed marginal genuinely has the form g ( t ) = t f ( t ) / μ . Additional truncation or censoring requires a separate observation model and is not silently absorbed into the present estimator.
Temporally ordered enrollment waves or clustered ascertainment can induce dependence, but the prevalent-cohort mechanism alone does not establish geometric α-mixing or the uniform local lagged-density bound. Therefore, this example is deliberately classified as motivating rather than fully verified. The present theorems apply only after stationarity, Assumption 4, geometric mixing, and Assumption 6 have been proved for the actual sampling mechanism.
Example 3
(Fully verified case: a Doeblin Markov model). Let { Y i : i Z } be a strictly stationary Markov chain on ( 0 , ) with invariant density
g ( y ) = y f ( y ) μ .
Assume that its transition kernel admits a density p ( x , y ) and that the following primitive conditions hold:
(i) 
there exist ε ( 0 , 1 ) and a probability density ν on ( 0 , ) such that
p ( x , y ) ε ν ( y ) f o r a l l x , y > 0 ;
(ii) 
the transition density is globally bounded,
sup x > 0 , y > 0 p ( x , y ) M < ;
(iii) 
for some q > 2 ,
E ( Y 0 q ) < .
Condition (i) is a Doeblin minorization. Writing P ( x , · ) = ε ν ( · ) + ( 1 ε ) Q x ( · ) , the usual contraction argument gives
sup x > 0 P k ( x , · ) G ( · ) TV ( 1 ε ) k .
For a stationary Markov chain, this implies geometric absolute regularity and hence geometric strong mixing; in particular,
α ( k ) β ( k ) C ( 1 ε ) k
for a finite constant C.
Next, if  p k is the k-step transition density, condition (ii) implies p k ( x , y ) M for every k 1 : the claim is immediate for k = 1 , and 
p k + 1 ( x , y ) = 0 p k ( x , z ) p ( z , y ) d z M 0 p k ( x , z ) d z = M .
Hence
g k ( x , y ) = g ( x ) p k ( x , y ) .
On every compact C 0 ( 0 , ) 2 , continuity of f implies that g is bounded on the first-coordinate projection of C 0 , and therefore
sup k 1 sup ( x , y ) C 0 g k ( x , y ) M sup x pr 1 ( C 0 ) g ( x ) < .
Thus, Assumptions 5 and 6 are derived from the primitive Markov conditions above, while (iii) is exactly Assumption 4. After imposing the theorem-specific local smoothness and bandwidth conditions, the strong uniform consistency, pointwise CLT, and finite-dimensional CLT apply directly. This example is therefore a genuinely verified application rather than a motivating dependence vignette.
Example 4
(Motivating example: size-biased durations and stereological measurements). Length or size bias also appears when sampling probabilities are proportional to a duration or size variable. Cross-sectional unemployment or sickness spells over-represent long spells; line- or plane-intercept stereology over-represents large fibers or particles; and equipment lifetimes observed at a fixed calendar time favor long survivors. Such observations may also be temporally, spatially, or cluster dependent.
These mechanisms motivate the model but do not themselves verify Assumptions 5 and 6. Whenever those assumptions are justified, estimation should be carried out on regions bounded away from zero. For point estimation, (69) and (72) provide oracle AMSE/AMISE benchmarks. A practical bandwidth must be described by the rule actually implemented, and the uncorrected normal interval (78) requires undersmoothing.
Remark 23
(Compatibility of the numerical models with the theorem hypotheses). The preceding theorems do not require compact support of F. Consequently, Gamma, lognormal, and Weibull target distributions are compatible with the support formulation of the theory. Section 4 records the corresponding finite-mean, local-smoothness, inverse-moment, and short-range dependence checks for the parameter values and stationary sampling mechanism used in the Monte Carlo study. Compatibility is therefore model- and parameter-specific, not a consequence of unbounded support alone.

4. Numerical Study

Organization of the numerical evidence.
The numerical section is organized so that the primary B = 2000 experiment, the computationally more expensive bandwidth and benchmark experiments, and the nested resampling study remain statistically distinct. Replication counts, Monte Carlo standard errors, numerical validation checks, and the exact provenance of every reported table and figure are stated explicitly. The complete cellwise tables needed for the discussion are retained in this article so that the numerical claims can be audited without relying on undocumented aggregation.
The Monte Carlo study has two complementary objectives. The first is to examine whether the finite-sample behavior of the Jones-type inverse-weighted kernel estimator is consistent with the first-order theory developed in Section 3. The second is to quantify a feature that is not visible from the limiting variance alone: the extent to which serial dependence can remain practically important at sample sizes for which the asymptotic covariance-localization mechanism has not yet become numerically dominant. This distinction is essential for interpreting the results. Under the assumptions of Theorem 3, the serial covariance contribution is asymptotically negligible after the n h n normalization, but this does not imply equality of finite-sample variances across dependence levels, nor does it imply exact Gaussianity or nominal coverage for a fixed n.
If X denotes the target random variable with density f on ( 0 , ) and finite mean
μ = 0 x f ( x ) d x ,
then the length-biased observation Y has density
g ( y ) = y f ( y ) μ , y > 0 .
For every simulated sample Y 1 , , Y n , we compute
μ ^ n = 1 n i = 1 n 1 Y i 1
and
f n ( t ) = μ ^ n n h n i = 1 n 1 Y i K t Y i h n , t > 0 .
The Epanechnikov kernel
K ( u ) = 3 4 ( 1 u 2 ) 1 { | u | 1 }
is used throughout. Hence
R ( K ) = K 2 ( u ) d u = 3 5 , m 2 ( K ) = u 2 K ( u ) d u = 1 5 .
These values were also checked numerically by quadrature. The simulation is therefore organized around the same quantities as the theoretical analysis: the second-order deterministic bias
B n ( t ) = h n 2 2 m 2 ( K ) f ( t ) ,
the first-order pointwise variance
σ L B 2 ( t ) = μ f ( t ) t R ( K ) ,
and the integrated AMISE criterion of Section 3.3.
All numerical results reported below come from one definitive publication configuration with B = 2000 replications for every scenario and master seed 20240517. The design comprises three target models, five sample sizes, and four dependence levels, giving 3 × 5 × 4 = 60 scenarios.

4.1. Target Distributions and Exact Length-Biased Sampling

We consider the three positive, unbounded-support target distributions
Γ ( 3 , 1 ) , Lognormal ( 0.5 , 0.4 ) , Weibull ( 2 , 2 ) .
Here, Γ ( a , θ ) denotes the Gamma law with shape a and scale θ , Lognormal ( m , s ) denotes the lognormal law with mean-log m and standard-deviation-log s, and  Weibull ( k , λ ) denotes the Weibull law with shape k and scale λ . These parameterizations are stated explicitly because the corresponding length-biased laws and inverse-moment properties depend on them.
Length-biased observations are generated from their exact distributions, not from an empirical resampling or approximate weighting device. If the target distribution is, respectively, Γ ( a , θ ) , Lognormal ( m , s ) , or  Weibull ( k , λ ) , then
Y Γ ( a + 1 , θ ) , Y Lognormal ( m + s 2 , s ) ,
and
Y = λ W 1 / k , W Γ ( 1 + 1 / k , 1 ) ,
respectively. These identities follow directly from g ( y ) = y f ( y ) / μ .
As an implementation check that is independent of the B = 2000 Monte Carlo experiment, we generated a diagnostic sample of size 200,000 from each length-biased law and compared the empirical and theoretical quantiles at probabilities
0.10 , 0.25 , 0.50 , 0.75 , 0.90 .
The maximum absolute discrepancy across the three models and five probabilities was 0.0055 , corresponding to a relative discrepancy below 0.3 % throughout and below 0.13 % at the median. This diagnostic sample is not included in, or pooled with, the  B = 2000 replications used for the reported Monte Carlo summaries.
The simulation models also satisfy the inverse-moment condition required by Assumption 4. Taking the common value q 0 = 2.5 , the length-biased Gamma ( 4 , 1 ) law satisfies
E ( Y q ) < , q < 4 ;
the length-biased lognormal law possesses inverse moments of every positive order; and, for the Weibull model,
E ( Y q ) < , q < k + 1 = 3 .
Thus, q 0 = 2.5 > 2 is admissible for all three target models. Moreover, the three target densities are C on every compact subset of ( 0 , ) . Consequently, the numerical design satisfies the same local smoothness, inverse-moment, and kernel non-degeneracy conditions that enter the bias, AMISE, and central-limit calculations.

4.2. Dependence Mechanism and Verification of the Short-Range Assumptions

Dependence is introduced through a stationary Gaussian-copula AR(1) construction. We consider
ϕ { 0 , 0.3 , 0.6 , 0.8 } ,
which we describe, respectively, as independence, mild dependence, moderate dependence, and strong short-range dependence. In particular, the case ϕ = 0.8 is not described as weak dependence in the numerical discussion.
For every replication,
Z 1 N ( 0 , 1 ) , Z i = ϕ Z i 1 + 1 ϕ 2 ε i , ε i iid N ( 0 , 1 ) ,
where Z 1 is independent of { ε i : i 2 } . We then set
U i = Φ ( Z i ) , Y i = G 1 ( U i ) ,
where G is the exact length-biased cdf corresponding to the target model. Because Z 1 is drawn directly from the invariant N ( 0 , 1 ) law, the Gaussian process is stationary from the first observation and no burn-in is required.
The dependence construction is deliberately chosen so that it lies inside the probabilistic regime used in the main theorems. For the three continuous target models, G is continuous and strictly increasing on ( 0 , ) , hence the map
z G 1 { Φ ( z ) }
is one-to-one. Therefore, the transformed process and the Gaussian AR(1) source generate the same sigma-fields. Since a stationary Gaussian AR(1) process with | ϕ | < 1 is geometrically strongly mixing, the sequence { Y i } is geometrically α -mixing for every fixed | ϕ | < 1 .
The uniform local bivariate-density condition can also be checked directly. At lag k, the pair ( Z 0 , Z k ) has Gaussian correlation
ρ k = ϕ k ,
and the joint density of ( Y 0 , Y k ) can be written as
g k ( x , y ) = c ρ k { G ( x ) , G ( y ) } g ( x ) g ( y ) ,
where c ρ denotes the Gaussian-copula density. If  C ( 0 , ) 2 is compact, then
{ ( G ( x ) , G ( y ) ) : ( x , y ) C }
is a compact subset of ( 0 , 1 ) 2 . On the simulated grid, | ρ k | 0.8 for every k 1 , so c ρ k is uniformly bounded on this compact image, while g is bounded on compact subsets of the positive half-line. It follows that
sup k 1 sup ( x , y ) C g k ( x , y ) < .
Thus, the Monte Carlo dependence mechanism verifies, rather than merely invokes, the geometric-mixing and local-bivariate-density assumptions used in the theory.

4.3. Evaluation Region and Bandwidth Regimes

For each target model, we use
I = [ q . 10 , q . 90 ] ,
where the quantiles are taken with respect to the target distribution F, not the length-biased observation law G. Pointwise summaries are reported at
q . 25 , q . 50 , q . 75 .
The sample sizes are
n { 100 , 250 , 500 , 1000 , 2000 } ,
and every sample size is combined with every dependence level. Within each replication, the same simulated sample is used for all numerical summaries: μ ^ n is computed once, and the two bandwidths defined below are then applied to that same sample.
For point-estimation performance, we use the oracle first-order AMISE reference bandwidth
h AMISE ( I ; n ) = C AMISE ( I ) n 1 / 5 ,
where
C AMISE ( I ) = R ( K ) I μ f ( t ) / t d t m 2 ( K ) 2 I f ( t ) 2 d t 1 / 5 .
The integrals are evaluated using the true target density and analytic f , with relative numerical tolerance 10 10 . The analytic second derivatives were independently checked against a five-point central finite-difference approximation at nine interior points of I for each model. The maximum absolute discrepancy was 2.15 × 10 8 and the maximum relative discrepancy was 3.09 × 10 7 .
As a further internal check, the closed-form AMISE bandwidth was compared with direct numerical minimization of the displayed AMISE criterion at n = 1000 . The relative discrepancies were 8.06 × 10 6 for the Gamma model, 4.03 × 10 5 for the Lognormal model, and 2.78 × 10 7 for the Weibull model.
For inference, we deliberately use the smaller bandwidth
h US ( I ; n ) = C AMISE ( I ) n 1 / 4 .
This is an undersmoothing bandwidth and is not called AMISE-optimal. It satisfies
h US ( I ; n ) 0 , n h US ( I ; n ) ,
together with
n h US ( I ; n ) ( log n ) 4 = C AMISE ( I ) n 3 / 4 ( log n ) 4
and
n h US ( I ; n ) 5 = C AMISE ( I ) 5 n 1 / 4 0 .
Hence, this bandwidth is asymptotically compatible with both the block-CLT condition and the undersmoothing requirement entering the feasible inference result.
The two bandwidth regimes are kept strictly separate in the numerical analysis. MISE and pointwise bias–variance–MSE summaries are reported under h AMISE , whereas the studentized statistic and confidence intervals are evaluated under h US . In particular, no confidence interval constructed at the AMISE bandwidth is presented as being covered by an asymptotic result that requires undersmoothing.
The interval I is bounded away from zero, in agreement with the local formulation of the theory. For the smallest sample sizes, however, an oracle kernel window can still interact with the left boundary. Behavior close to q . 10 is therefore treated as a finite-sample feature, and the main pointwise inferential diagnostics are reported at the more interior quartiles. Accordingly, h AMISE should be read as a first-order oracle reference bandwidth, not as an exact finite-sample optimizer, see Table 3.
Integrated squared error is evaluated under h AMISE ( I ; n ) on a deterministic grid of 1001 equally spaced points over I, using the trapezoidal rule. In a prespecified sensitivity check for the Gamma model with n = 500 and ϕ = 0 , evaluated on the same simulated data, increasing the grid from 1001 to 2001 points changed the ISE from 5.403311 × 10 4 to 5.403346 × 10 4 , a relative difference of 6.48 × 10 6 . Thus, the numerical integration error is negligible relative to Monte Carlo variation at the precision reported.

4.4. Monte Carlo Summaries, Studentization, and Monte Carlo Error

For every one of the 60 scenarios and every one of the B = 2000 replications, we record the harmonic-mean estimator μ ^ n , the density estimate on the full evaluation grid under h AMISE , the resulting integrated squared error, and the pointwise estimates
f n ( q . 25 ) , f n ( q . 50 ) , f n ( q . 75 )
under both bandwidths.
Under h US , the plug-in variance estimator is
σ ^ L B 2 ( t ) = μ ^ n f n ( t ) t R ( K ) .
Whenever σ ^ L B 2 ( t ) > 0 , we compute
Z n ( t ) = n h US { f n ( t ) f ( t ) } σ ^ L B ( t ) .
The corresponding nominal 95 % confidence interval is
f n ( t ) ± z 0.975 σ ^ L B 2 ( t ) n h US ,
and is reported without truncation at zero.
Positivity of the estimated variance is monitored explicitly rather than assumed numerically. Across all 60 scenarios, the 3 evaluation points, and the B = 2000 replications, there are 360,000 replication–point evaluations under h US . The definitive simulation recorded no non-positive plug-in variance estimates, no non-finite plug-in variance estimates, and no undefined studentized statistics. This finding is consistent with the use of the non-negative Epanechnikov kernel and with Proposition 3; it is reported as a diagnostic, not used as a replacement for the theoretical positivity argument.
Monte Carlo uncertainty is reported explicitly. For a cellwise empirical coverage proportion p ^ ,
MCSE cov = p ^ ( 1 p ^ ) B 1 / 2 , B = 2000 .
The calculation is performed separately for each ( model , n , ϕ , t ) cell; the three evaluation points are not pooled as if they represented 3 B independent Bernoulli observations. For MISE,
MCSE { MISE ^ } = sd ( ISE 1 , , ISE B ) B .
Pointwise Monte Carlo variance uses the explicit divide-by-B convention, and
MSE ^ = Bias ^ 2 + Var ^ MC
was verified numerically to floating-point tolerance for every reported cell. These Monte Carlo standard errors quantify simulation uncertainty in the reported summaries; they are not sampling standard errors of the statistical estimator f n .

4.5. Computational Reproducibility

The computational sequence is fixed throughout. Each replication begins with a stationary Gaussian AR(1) draw, which is transformed through the Gaussian cdf and then through the exact inverse length-biased cdf. The normalizing estimator μ ^ n is computed once from the resulting sample. The density is evaluated on the fixed 1001-point grid under h AMISE , from which ISE is computed by the trapezoidal rule. Pointwise bias, variance, and MSE are then recorded under h AMISE , whereas studentization and confidence-interval coverage are evaluated under h US .
Any non-positive or non-finite plug-in variance would be counted as a separate diagnostic event instead of being replaced by an arbitrary numerical value. After completion of the B = 2000 replications, the program computes Monte Carlo means, divide-by-B variances, MISE and its MCSE, empirical coverage and its binomial MCSE, and the first four empirical moments of the studentized statistic.
Every numerical entry in Tables and every figure in this section is generated from the same definitive publication workflow. The retained outputs include the MISE and MCSE summaries, coverage and MCSE summaries, pointwise bias–variance–MSE calculations, studentized-moment summaries, summaries of μ ^ n , the variance-positivity audit, variance-inflation calculations, the grid-sensitivity check, the analytic- f validation, the AMISE-bandwidth validation, and the length-biased-sampler diagnostic. The complete R code is available from the corresponding author upon reasonable request.

4.6. Finite-Sample Results

The numerical results are interpreted as finite-sample diagnostics of the first-order theory, rather than as empirical substitutes for it. In particular, asymptotic negligibility of the scaled serial covariance does not imply equality of the finite-sample variances across dependence levels, and pointwise asymptotic normality does not imply exact Gaussianity or nominal coverage for a fixed n.
Figure 1 compares the Monte Carlo mean density estimate with the target density at n = 2000 . The principal shape of each target density is recovered under all four dependence levels. The more consequential effect of increasing dependence is not a systematic failure of density recovery but an increase in stochastic variability, which is quantified below.
Figure 2 displays MISE against n on log–log axes, together with error bars of ± 2 MCSE { MISE ^ } . Under independence, the OLS slope of log ( MISE ) on log ( n ) , using n { 500 , 1000 , 2000 } , is 0.765 for Gamma, 0.714 for Lognormal, and 0.753 for Weibull. At  ϕ = 0.8 , the corresponding slopes are 0.834 , 0.831 , and 0.897 .
These finite-range slopes are close to the first-order AMISE benchmark 4 / 5 , but they are not interpreted as estimates of an asymptotic convergence exponent. The apparent steepening under stronger dependence can reflect the changing numerical importance of the serial covariance over the sample-size range considered. We therefore do not use these fitted slopes to claim a dependence-specific asymptotic rate.
Figure 3 examines the Gaussian approximation for the studentized statistic at q . 50 when n = 2000 . Under independence and mild dependence, the distribution is comparatively close to standard normal. Under ϕ = 0.6 , and more clearly under ϕ = 0.8 , residual variance inflation, heavier tails, and asymmetry remain visible. The empirical Q–Q points are plotted explicitly so that departures in both the center and the tails can be assessed directly. These plots are finite-sample diagnostics of the Gaussian approximation, not graphical proofs of asymptotic normality.
Figure 4 displays MISE directly as a function of ϕ . For every model and every simulated sample size, MISE increases as the dependence parameter increases, and the effect remains visible at n = 2000 . This is a finite-sample statement only. It is not extrapolated into a claim that the relative effect of dependence remains non-vanishing asymptotically.
The variance-inflation ratio provides a more direct measure of the finite-sample dependence effect:
VIR ( model , n , ϕ , t ) = Var ^ MC , ϕ { f n ( t ) } Var ^ MC , 0 { f n ( t ) } .
At n = 2000 and the grid point nearest q . 50 , the variance under ϕ = 0.8 , relative to independence at the same sample size and oracle bandwidth, is inflated by factors 2.464 for Gamma, 2.004 for Lognormal, and 2.645 for Weibull.
These are descriptive finite-sample variance ratios. They are not estimates of a non-zero asymptotic long-run variance multiplier. Their importance is interpretative: they show why the statement that the serial covariance is first-order negligible must not be paraphrased as saying that dependence has no finite-sample effect, see Figure 5.
The coverage results provide the clearest finite-sample qualification of the first-order inference theory. Averaging over models, sample sizes, and evaluation points, empirical coverage is 0.961 under independence and 0.948 under mild dependence ( ϕ = 0.3 ) . Under moderate dependence ( ϕ = 0.6 ) , average coverage falls to 0.909 , while under ϕ = 0.8 the corresponding average is 0.811 .
The same pattern remains evident at the largest sample size. At  n = 2000 and ϕ = 0.8 , average coverage over the three evaluation points is 0.831 for Gamma, 0.866 for Lognormal, and 0.818 for Weibull. The upper-quartile point is often the most difficult; at n = 2000 and ϕ = 0.8 , coverage at q . 75 is 0.782 for Gamma and 0.765 for Weibull.
Accordingly, we do not state that the intervals “approach the nominal coverage level” uniformly over the dependence regimes considered. Such a description is broadly supported over the simulated range under independence and mild dependence but not under ϕ = 0.6 and, especially, not under ϕ = 0.8 .
This undercoverage does not contradict Theorem 3. The theorem identifies the first-order limiting variance for each fixed dependence parameter satisfying the stated conditions; it does not supply a finite-n rate for the coverage error. At finite n, the exact variance still contains the lagged covariance sum in (22), whereas the first-order plug-in variance estimator does not estimate that full finite-sample covariance component. The observed undercoverage under stronger dependence should therefore be viewed as a genuine limitation of first-order plug-in inference in this setting, rather than as a negligible numerical irregularity, see Figure 6.

4.7. Interpretation of the Numerical Results

Three conclusions emerge consistently from the simulation.
First, the estimator behaves in the manner suggested by the first-order estimation theory. Pointwise bias under h AMISE decreases with n, MISE decreases systematically with sample size, and the empirical bias of μ ^ n is small in the reported scenarios, while its dispersion contracts as n increases. These findings are compatible with the second-order bias expansion and with the O P ( n 1 / 2 ) stochastic order established for the random normalization.
Second, short-range dependence can remain materially important at finite sample sizes even though its contribution disappears from the first-order limiting variance. The monotone deterioration of MISE and the variance inflation ratios make this distinction especially clear. The mathematically correct interpretation is therefore not that dependence is “irrelevant” but that its serial covariance contribution is asymptotically of smaller order after the localization and normalization governing the kernel estimator.
Third, the finite-sample distribution of the studentized statistic deteriorates systematically as persistence increases. Table 16 shows pronounced left-skewness and heavy tails at small n under ϕ = 0.8 . For example, at  n = 100 , the Gamma statistic at q . 25 has skewness 2.067 and excess kurtosis 13.212 . These distortions moderate as n grows, but they do not disappear uniformly over the simulated designs. Even at n = 2000 and ϕ = 0.8 , the reported empirical variances of the studentized statistics range from 1.413 to 2.732 , while all corresponding skewness coefficients remain negative, approximately between 0.147 and 0.308 . This residual variance inflation and asymmetry provide a direct numerical explanation for the persistent undercoverage of the first-order normal plug-in interval.
Taken together, the Monte Carlo evidence supports the asymptotic results while also identifying their finite-sample boundary. Under geometric short-range dependence, the serial covariance disappears from the first-order limiting variance; nevertheless, stronger persistence can still produce substantial variance inflation, non-Gaussian studentized distributions, and meaningful undercoverage at sample sizes as large as n = 2000 . In applications where such dependence is plausible, first-order plug-in studentization should therefore be interpreted with appropriate caution. Dependence-adaptive finite-sample variance correction, higher-order studentization, and simultaneous confidence procedures are beyond the scope of the present paper and provide natural directions for further work.

4.8. Extended Computational Design and Reproducibility

The additional finite-sample experiments are organized as separate designs rather than pooled into a single simulation label. The primary Gaussian-copula AR(1) experiment contains 60 model–sample-size–dependence cells with B = 2000 replications per cell. The Frank-copula Markov robustness experiment uses B = 2000 for each of 18 cells. The feasible-bandwidth experiment uses B = 60 per primary model cell and B = 30 for the additional Gamma-mixture shape-sensitivity design; the estimator benchmark uses B = 200 ; and the dependent-resampling experiment uses B outer = 100 Monte Carlo samples with B boot = 200 block resamples within each sample. The separation of these replication counts reflects their very different computational costs and prevents nested experiments from being confused with the primary Monte Carlo design, see Table 4.
The computational pipeline uses master seed 20240517 together with deterministic cell- and replication-level seeds. The definitive run was executed under R 4.3.3 on x86_64-pc-linux-gnu (Ubuntu 24.04.4 LTS). The recorded workflow includes the RNG configuration, package versions, per-file SHA-256 hashes, and a combined source hash. Numerical tables and figures are generated from the same checkpointed machine-readable objects, so the graphical and tabular summaries cannot diverge through manual transcription.

4.9. Feasible Bandwidth Selection for the Jones Estimator

The oracle bandwidth h AMISE isolates estimator behavior from bandwidth estimation error, but its constant depends on unknown features of f and is not available in applications. We therefore compare it with five feasible rules: a BGM-RT-style normal-reference selector, weighted cross-validation (WCV-Jones), SBoot1-Jones, SBoot2-Jones, and a contiguous leave-one-block-out cross-validation rule (Block-CV Jones). Every feasible bandwidth is recomputed from each replicated sample. Their empirical performance is evaluated directly; no asymptotic consistency or optimality property is inferred from the simulation.
The BGM-labeled procedures in the computational archive are independent, fully disclosed implementations of the corresponding published ideas rather than literal calls to the WData package. The terminology is therefore kept explicit—“BGM-RT-style”, “WCV-Jones”, “SBoot1-Jones”, and “SBoot2-Jones”—and the comparison estimator below is described as “Bhattacharyya-type”. This distinction concerns software provenance and does not alter the statistical definitions used in the numerical comparisons.
Table 5 and Figure 7 summarize the model-averaged selected-to-oracle bandwidth ratios over n { 100 , 500 , 1000 } and ϕ { 0 , 0.6 } for the three primary models. WCV-Jones and Block-CV remain comparatively close to the oracle benchmark, whereas SBoot1 systematically selects larger bandwidths and SBoot2 is diagnostically unstable over the chosen search range: its upper boundary is selected in approximately 97.8 100 % of the aggregated cells. The additional Gamma-mixture experiment preserves the same ranking, with selected-to-oracle ranges 0.999–1.006 (BGM-RT-style), 1.144–1.155 (WCV-Jones), 0.987–1.133 (Block-CV), 1.869–2.025 (SBoot1), and 3.998–4.025 (SBoot2). Boundary solutions are reported rather than removed or winsorized.
The selector comparison addresses implementation, not the unresolved random-bandwidth asymptotics. Replacing a deterministic h n by a data-dependent h ^ n in the dependent ratio estimator would require stochastic equicontinuity uniformly over a random bandwidth neighborhood, together with control of the selection criterion and the normalizer on the same probability space. Establishing such a result would constitute a distinct theoretical problem rather than a formal corollary of the oracle AMISE calculation.

4.10. Dependence-Robust Finite-Sample Inference

The principal finite-sample weakness of the first-order theory arises in studentization under strong persistence rather than in recovery of the density shape. We retain the plug-in interval as the asymptotic reference and examine two finite-sample corrections. For the HAC construction, write a i ( t ) = Y i 1 K h ( t Y i ) , b i = Y i 1 , A n = a ¯ , H n = b ¯ , μ ^ n = H n 1 , and  f n = A n / H n . The empirical influence sequence is
ψ ^ i ( t ) = μ ^ n { a i ( t ) A n ( t ) } μ ^ n f n ( t ) { b i H n } .
A Bartlett/Newey–West estimator is formed from ψ ^ i ( t ) using lag orders L n 1 / 4 and L n 1 / 3 . The normalization component is therefore retained in the finite-sample correction, in agreement with the complete ratio expansion proved in Proposition 1. The HAC procedure is evaluated numerically and is not assigned an additional asymptotic validity theorem in this article.
At the median point, averaging over the three primary models and n { 1000 , 2000 } , plug-in coverage equals 0.9566 under independence, 0.9394 at ϕ = 0.6 , and 0.8662 at ϕ = 0.8 . The quarter-order HAC values on the same cells are 0.9353, 0.9356, and 0.9250, while the third-order values are 0.9356, 0.9376, and 0.9334. The correction is consequently not uniformly preferable: when the plug-in interval is already well calibrated, it may widen the interval without improving coverage, whereas under strong persistence it recovers a substantial part of the coverage deficit.
At n = 2000 , ϕ = 0.8 , and  q . 50 , plug-in coverage is 0.8745, 0.8525, and 0.8610 for the Gamma, Lognormal, and Weibull targets. The quarter-order HAC rule raises these values to 0.9320, 0.9070, and 0.9280 and the third-order rule to 0.9405, 0.9180, and 0.9355, see Table 6. The corresponding mean interval-length ratios are 1.164, 1.105, and 1.230 for the quarter-order rule and 1.203, 1.148, and 1.279 for the third-order rule. Averaged over all three evaluation points at the same n and ϕ , the plug-in/quarter-order HAC coverages are 0.831/0.915 (Gamma), 0.828/0.900 (Lognormal), and 0.824/0.908 (Weibull). Thus, dependence-aware studentization materially improves the difficult cells, but it does not remove all coverage error.
The moving-block-bootstrap experiment provides an independent resampling check on a designated high-persistence subset, see Figure 8. It uses B outer = 100 Monte Carlo samples and B boot = 200 block resamples per sample. At  n = 1000 , ϕ = 0.8 , and  q . 50 , the Lognormal cell has coverage 0.78 for the plug-in interval and 0.83 for both HAC and the block bootstrap, see Figure 9. Because the outer experiment contains only 100 replications, its binomial Monte Carlo errors are necessarily larger than those of the primary design. The relevant conclusion is qualitative but important: dependence-aware corrections improve calibration in most difficult cells, yet no method examined here uniformly restores nominal coverage, see Table 7.

4.11. Robustness to a Second Dependence Mechanism

A second dependence design tests whether the preceding finite-sample pattern is specific to the Gaussian-copula AR(1) construction. Design B is a first-order Frank-copula Markov chain with uniform stationary marginal, transformed through the exact length-biased quantile function. Its Frank parameter is calibrated to the Gaussian design through Kendall’s τ = ( 2 / π ) arcsin ( ϕ ) . The values ϕ = 0.6 and 0.8 correspond to target Kendall coefficients 0.4097 and 0.5903 and Frank parameters 4.296 and 7.677, respectively. The marginal length-biased law is preserved by construction. This design is used as a robustness experiment; no claim is made that it constitutes an additional verified instance of every mixing assumption in the main theorems without a separate proof.
At n = 2000 , moving from the moderate to the stronger matched-dependence level increases MISE from 0.000583 to 0.000901 for Gamma, from 0.001476 to 0.002178 for Lognormal, and from 0.000870 to 0.001417 for Weibull. At the stronger level, median-point plug-in coverage is 0.916, 0.883, and 0.865; the quarter-order HAC correction raises these values to 0.934, 0.909, and 0.926. The same qualitative risk and coverage ordering therefore persists under a different copula mechanism, see Table 8 and Figure 10.

4.12. Estimator Benchmark

A separate B = 200 experiment compares the Jones estimator under the oracle, WCV-Jones, and Block-CV bandwidths with an independently implemented Bhattacharyya-type smooth-then-divide estimator. An ordinary KDE of the observed Y i ’s is retained only as a negative control for the biased marginal g and is never interpreted as an estimator of the target density f.
At n = 2000 under independence, the WCV/Block-CV Jones MISEs are 0.000868/0.000865 for Gamma, 0.002334/0.002229 for Lognormal, and 0.001539/0.001558 for Weibull, compared with 0.010941, 0.025066, and 0.018411 for the Bhattacharyya-type estimator, see Table 9 and Figure 11. The corresponding MISE advantage of the feasible Jones estimator is approximately 10.7–12.6-fold in these cells; at ϕ = 0.6 , the advantage is approximately 9.5–10.3-fold. These ratios are stated only for the reported n = 2000 comparisons and are not extrapolated to smaller sample sizes.

4.13. Computational Validation

Fifteen automated validation checks were applied to the computational pipeline. They include the analytical Epanechnikov moments, exact length-biased samplers, marginal preservation under both dependence generators, normalization of the inverse-weighted empirical distribution, numerical integration of f n , an independent numerical minimization of the oracle AMISE, bandwidth-selector finiteness and boundary diagnostics, the identity MSE ^ = Bias ^ 2 + Var ^ MC , the coverage-MCSE formula, HAC finiteness, block-bootstrap reconstruction, and deterministic seed reproducibility. All fifteen checks passed. Selected diagnostics are given in Table 10.
During construction of the aggregate selector and benchmark summaries, two file-pattern collisions were detected by the validation workflow. The aggregation rules were corrected before the reported tables and figures were generated, and the entire validation suite was rerun successfully. The final numerical summaries are therefore taken only from the post-correction output objects.

4.14. Synthesis of the Numerical Evidence

The extended experiments sharpen, rather than replace, the interpretation of the asymptotic results. Feasible bandwidth selection is practically viable in the designs considered, but the quality of individual selectors differs substantially; WCV-Jones and Block-CV track the oracle much more closely than the two bootstrap rules over the reported search region. The full-ratio HAC correction materially improves coverage when serial persistence is strong, but it is not uniformly superior under weak dependence and does not eliminate every coverage defect. The moving-block experiment reaches the same qualitative conclusion. The Frank-copula Markov design reproduces the principal dependence pattern under a distinct copula mechanism. Finally, at  n = 2000 , the feasible Jones estimator has substantially smaller MISE than the Bhattacharyya-type comparator in the tested cells. These are finite-sample statements. They neither modify the deterministic-bandwidth asymptotic theory nor imply new optimality, bootstrap, HAC, or Frank-process limit theorems.

4.15. Monte Carlo Tables

The following tables collect the numerical output from the primary B = 2000 experiment: model-specific evaluation points, pointwise bias–variance–MSE summaries under the oracle bandwidth h AMISE ( I ; n ) , MISE with Monte Carlo standard errors, empirical coverage of the studentized interval under the undersmoothing bandwidth h US ( I ; n ) , normality moments of the studentized statistic Z n ( t ) , and the behavior of μ ^ n . MISE and pointwise bias/variance/MSE are reported under h AMISE ( I ; n ) only, and coverage and the studentized moments are reported under h US ( I ; n ) only, consistent with Section 4: the harmonic-mean estimator μ ^ n does not depend on the bandwidth and is therefore reported once per scenario. Table 11, Table 12, Table 13, Table 14, Table 15 and Table 16 give the numerical details.

5. Empirical Illustration: Line-Transect Shrub Widths

A direct empirical setting for length-biased sampling is provided by the line-transect study of Cercocarpus montanus shrubs near Laramie, Wyoming, distributed as shrub.data in the WData package [47,48]. The field study used six transects arranged in two replicates. A shrub was observed when a transect intersected it; consequently, wider shrubs had proportionally larger inclusion probability. For the width variable, the sampling weight is therefore w ( x ) = x , so the observed width distribution has exactly the marginal form
g ( x ) = x f ( x ) μ , x > 0 ,
that motivates the estimator studied in this paper. This design is more directly aligned with the present uncensored density model than a prevalent-cohort data set involving left truncation or right censoring.
The WData implementation associated with these data includes the Jones estimator, the Bhattacharyya-type alternative, and several feasible length-biased bandwidth selectors. The ecological example therefore provides a concrete interpretation of the de-tilting operation: the ordinary empirical distribution of observed widths represents the size-biased law G, whereas reciprocal weighting reconstructs the target shrub-width distribution F by downweighting the larger widths that are preferentially intersected by the transects.
The empirical design also makes clear why marginal length bias and serial dependence must not be conflated. The recorded variables identify replicates, transects, individual shrubs or clumps, intercept length, width, height, and stem count, but the documentation does not establish Number as a temporal index. Consequently, no artificial AR or Markov ordering is imposed on the shrub observations. The data provide a genuine illustration of the length-biased observation mechanism and of the role of feasible bandwidth selection; empirical validation of the serial-dependence assumptions would require an ordered observation scheme for which the dependence structure is scientifically defined.

6. Discussion and Conclusions

The analysis gives a coherent dependent-data theory for the Jones inverse-weighted kernel estimator under length-biased sampling. The central probabilistic issue is not the kernel representation alone but the interaction of four mechanisms: the reciprocal weight, the random ratio normalizer, shrinking localization, and serial dependence. The exact factorization of the estimator isolates the global normalization problem from the local stationary triangular array, while the compactly localized kernel separates the singularity at the origin from interior estimation on fixed compact subsets of ( 0 , ) .
Under geometric strong mixing and uniform local control of lagged bivariate densities, short-lag covariances are handled by direct density calculations and distant lags by the mixing inequality. Splitting the covariance series at a logarithmic lag yields the key estimate
h n k 1 Cov { W n , 0 ( s ) , W n , k ( t ) } = O h n log 1 h n = o ( 1 ) ,
uniformly on compact sets. This establishes why the serial term disappears from the first-order variance without assuming that finite-sample dependence is negligible. Strong uniform consistency and the quantitative O P { h n 2 + ( log n / ( n h n ) ) 1 / 2 } bound are proved as logically distinct statements. Pointwise asymptotic normality is obtained from an explicit row-wise big-block/small-block argument, and finite-dimensional convergence follows from additional off-diagonal covariance control. The random normalizer is treated at the same level of explicitness: its linearized variance, its covariance with the localized fluctuation, and the non-linear product remainder are all negligible at the n h n scale.
The first-order AMSE and AMISE criteria therefore identify deterministic oracle bandwidths of order n 1 / 5 . They do not by themselves solve the random-bandwidth problem under dependent length-biased sampling. The numerical evidence shows that weighted cross-validation and Block-CV can remain comparatively close to the oracle in the designs considered, whereas the bootstrap selectors examined here are substantially more sensitive to the optimization range. A general theorem for a fully data-driven bandwidth would require a distinct stochastic analysis of the selection criterion jointly with inverse weighting, normalization, and dependence.
The finite-sample experiments also delimit the practical meaning of the first-order variance result. Under independence and mild dependence, the plug-in interval is well calibrated over the reported range. Under stronger persistence, the lagged covariance remains numerically important, even though it is asymptotically lower order. At  n = 2000 , ϕ = 0.8 , and the median, plug-in coverage equals 0.8745, 0.8525, and 0.8610 for Gamma, Lognormal, and Weibull, while the quarter-order full-ratio HAC correction raises these values to 0.9320, 0.9070, and 0.9280. The price is a corresponding increase in average interval length, and neither HAC nor moving-block resampling restores nominal calibration uniformly. This negative finding is substantively important: first-order asymptotic variance equivalence does not imply finite-sample equivalence to independence.
The Frank-copula Markov experiment reproduces the same qualitative risk and coverage ordering under a different dependence construction, while the estimator benchmark shows a substantial MISE advantage for the Jones procedure over the independently implemented smooth-then-divide comparator at the largest sample size. These experiments strengthen the empirical interpretation of the theory without enlarging the assumptions of the proved results.
The resulting picture is inherently two-scale. At first order, covariance localization removes short-range serial dependence from the limiting variance. At finite sample sizes, the same dependence may materially affect variance, Gaussian approximation, and coverage. Natural extensions include random-bandwidth theory, dependence-adaptive studentization accompanied by its own validity theorem, simultaneous confidence bands, long-range regimes in which the serial covariance survives at leading order, and observation schemes that incorporate censoring, truncation, multivariate weighting, or more general informative-sampling operators. From the symmetry/asymmetry perspective, the relevant structural object remains the size-tilting operator d G / d F = x / μ : it induces directional over-representation of large observations, and the inverse-weighted ratio estimator performs the exact statistical de-tilting. No symmetry of the target density is required.

7. Proofs

Throughout this section, C , C 1 , C 2 , denote finite positive constants whose values may change from one occurrence to the next. Unless explicitly stated otherwise, these constants are independent of n, of the lag index, and of the evaluation points in a fixed compact subset of ( 0 , ) . We write
K h ( u ) = 1 h K u h ,
and recall
A n ( t ) = 1 n i = 1 n 1 Y i K h n ( t Y i ) , W n , i ( t ) = μ Y i K h n ( t Y i ) = μ Y i h n K t Y i h n .
The estimator satisfies
f n ( t ) = μ ^ n A n ( t ) = μ ^ n μ 1 n i = 1 n W n , i ( t ) .
For fixed t > 0 , set
Z n , i ( t ) = W n , i ( t ) E W n , 0 ( t ) .
We first derive the local envelope and the exact covariance estimates for the localized triangular array. We then prove a uniform exponential bound using a Bernstein inequality whose hypotheses match the geometric strong-mixing assumption exactly. The central limit theorem is established separately as an abstract big-block/small-block theorem for row-wise stationary triangular arrays; only after that theorem has been proved do we apply it to Z n , i ( t ) . This separation is essential: it makes explicit every bandwidth–dependence compatibility condition and avoids importing an unverified central limit argument into the kernel calculation.

7.1. Localization and Probabilistic Input

Lemma 1
(Uniform localization and envelopes). Let I = [ a , b ] ( 0 , ) , 0 < a < b < . Under Assumption 1, and along a deterministic bandwidth sequence satisfying h n 0 , for all sufficiently large n,
K h n ( t Y i ) 0 , t I Y i I n : = [ a h n , b + h n ] I : = [ a / 2 , b + 1 ] .
Consequently,
sup t I 1 Y i K h n ( t Y i ) C h n , sup t I | W n , i ( t ) | C h n .
Moreover, if  L K is a Lipschitz constant of K, then for s , t I ,
1 Y i K h n ( t Y i ) 1 Y i K h n ( s Y i ) C h n 2 | t s | .
Proof. 
Since h n 0 , eventually h n < a / 2 and h n < 1 . Because  supp ( K ) [ 1 , 1 ] ,
K h n ( t Y i ) 0 | t Y i | h n .
Thus, for  t [ a , b ] ,
a h n Y i b + h n ,
which proves (84). In particular, Y i 1 2 / a on the effective kernel support. Hence
Y i 1 K h n ( t Y i ) 2 K a h n ,
and multiplication by μ gives the second bound in (85). Finally,
Y i 1 K h n ( t Y i ) Y i 1 K h n ( s Y i ) 2 a h n K t Y i h n K s Y i h n 2 L K a h n 2 | t s | .
This proves (86).    □
We next record the two dependence tools used below. If X and Y are measurable with respect to sigma-fields separated by k observations, then Davydov’s inequality gives, for  p , r > 1 with p 1 + r 1 < 1 ,
| Cov ( X , Y ) | 8 α ( k ) 1 p 1 r 1 X p Y r ,
and, for bounded X , Y ,
| Cov ( X , Y ) | 4 α ( k ) X Y .
Proposition 4
(Bernstein inequality used in the uniform proof). Let, for each n, { X n , i : i 1 } be a centered row whose strong mixing coefficients α n ( k ) satisfy
sup n 1 α n ( k ) e 2 c k , k 1 ,
for some c > 0 , and suppose
sup i 1 X n , i M n .
Define
v n 2 = sup i 1 Var ( X n , i ) + 2 j > i | Cov ( X n , i , X n , j ) | .
Then, there exists C c > 0 , depending only on c, such that for every n 2 and x 0 ,
P i = 1 n X n , i x exp C c x 2 n v n 2 + M n 2 + x M n ( log n ) 2 .
Proof. 
For a fixed n, apply Theorem 2 of [49] to the n-th row. The constants in that theorem depend only on the exponential mixing constant, not on M n , v n , or n. Because (89) is uniform over the rows, the same constant C c works for all n. The theorem does not require stationarity for the exponential inequality itself.
For completeness, the row-dependent envelope causes no hidden triangular-array issue. If  M n = 0 , the assertion is trivial. Otherwise, apply the cited theorem to X n , i / M n . Its envelope is one, its variance proxy is v n 2 / M n 2 , and the threshold is x / M n . Multiplying the resulting denominator by M n 2 gives exactly
n v n 2 + M n 2 + x M n ( log n ) 2 .
Thus, the bound is homogeneous and may be applied row by row with M n diverging, provided the geometric mixing constants are uniform in n, as assumed in (89). This gives (91).    □
Remark 24
(Equivalent parametrizations of geometric mixing). Assumption 5 is stated in the equivalent form α ( k ) C α e c α k . Since every strong-mixing coefficient is bounded by 1 / 4 , one may decrease the exponential rate and absorb both the multiplicative constant and the finitely many small lags into a new constant c > 0 , so that
α ( k ) e 2 c k , k 1 .
Hence, (89) is simply an equivalent normalization of the geometric rate already imposed in Assumption 5.
We also use the standard separated-block characteristic-function inequality. If U 1 , , U r are measurable with respect to successive sigma-fields separated by at least q indices, then
E exp i s j = 1 r U j j = 1 r E e i s U j 16 ( r 1 ) α ( q ) , s R .

7.2. The Random Normalizing Constant

Lemma 2
(Consistency and stochastic order of μ ^ n ). Under Assumptions 4 and 5,
μ ^ n a . s . μ
and
μ ^ n μ = O P ( n 1 / 2 ) .
More precisely, with 
H n = 1 n i = 1 n Y i 1 , H = E ( Y 1 ) = μ 1 ,
one has
μ ^ n μ = μ 2 ( H n H ) + o P ( n 1 / 2 ) .
Proof. 
Strong mixing implies ergodicity. Since E ( Y 1 ) = μ 1 < , Birkhoff’s theorem gives H n H almost surely; continuity of x x 1 at H > 0 gives (93).
Let q > 2 be as in Assumption 4. Applying (87) with p = r = q ,
Cov ( Y 0 1 , Y k 1 ) 8 Y 1 q 2 α ( k ) 1 2 / q .
Geometric mixing implies
k = 1 α ( k ) 1 2 / q < .
Therefore
Var ( H n ) = 1 n Var ( Y 0 1 ) + 2 n k = 1 n 1 1 k n Cov ( Y 0 1 , Y k 1 ) C n 1 + k = 1 α ( k ) 1 2 / q = O ( n 1 ) .
Hence, H n H = O P ( n 1 / 2 ) . Taylor expansion of x 1 at H gives
H n 1 H 1 = H 2 ( H n H ) + o P ( | H n H | ) ,
which is (95) and proves (94).    □

7.3. Deterministic Smoothing Bias

Lemma 3
(Expectation and local bias). Assume h n 0 . For every t > 0 ,
E A n ( t ) = 1 μ 0 K h n ( t y ) f ( y ) d y , E W n , 0 ( t ) = 0 K h n ( t y ) f ( y ) d y .
If f is uniformly continuous on a neighborhood of I = [ a , b ] ( 0 , ) , then
sup t I | μ E A n ( t ) f ( t ) | 0 .
If f C 2 on a neighborhood of I, then
sup t I μ E A n ( t ) f ( t ) h n 2 2 m 2 ( K ) f ( t ) = o ( h n 2 ) ,
and, in particular,
sup t I | μ E A n ( t ) f ( t ) | = O ( h n 2 ) .
At a fixed t > 0 where f is twice continuously differentiable,
E W n , 0 ( t ) f ( t ) = h n 2 2 m 2 ( K ) f ( t ) + o ( h n 2 ) .
Proof. 
From g ( y ) = y f ( y ) / μ ,
E A n ( t ) = 0 1 y K h n ( t y ) g ( y ) d y = 1 μ 0 K h n ( t y ) f ( y ) d y ,
which proves (98). For  t I and large n, the change of variables y = t h n u gives
μ E A n ( t ) = 1 1 K ( u ) f ( t h n u ) d u .
Consequently,
μ E A n ( t ) f ( t ) = 1 1 K ( u ) { f ( t h n u ) f ( t ) } d u .
Uniform continuity yields (99).
If f C 2 on a neighborhood of I, choose a compact interval I δ ( 0 , ) containing I in its interior and contained in that neighborhood. Since f is uniformly continuous on I δ , Taylor’s formula with integral remainder gives, uniformly in t I and | u | 1 ,
f ( t h n u ) = f ( t ) h n u f ( t ) + h n 2 u 2 2 f ( t ) + h n 2 u 2 r n ( t , u ) ,
where
sup t I sup | u | 1 | r n ( t , u ) | 0 .
Using K = 1 , u K ( u ) d u = 0 , and  u 2 K ( u ) d u = m 2 ( K ) in (104) yields (100), and (101) follows immediately.
At a fixed t > 0 where f is twice continuously differentiable, the same Taylor argument, now only locally at t, gives
f ( t h n u ) = f ( t ) h n u f ( t ) + h n 2 u 2 2 f ( t ) + h n 2 u 2 r n ( u ) , sup | u | 1 | r n ( u ) | 0 ,
and therefore proves (102).    □

7.4. Covariance Calculus for the Localized Array

Lemma 4
(Uniform diagonal and off-diagonal covariance bounds). Let I = [ a , b ] ( 0 , ) . Suppose Assumptions 1, 2, 5, and 6 hold, f is bounded on a neighborhood of I, and  h n 0 . Define
γ n , k ( s , t ) = Cov { W n , 0 ( s ) , W n , k ( t ) } .
Then, uniformly in s , t I ,
Var { W n , 0 ( t ) } C h n ,
| γ n , k ( s , t ) | C , k 1 ,
and
| γ n , k ( s , t ) | C h n 2 α ( k ) , k 1 .
Consequently,
sup s , t I k = 1 | γ n , k ( s , t ) | = O log 1 h n ,
and
sup s , t I h n k = 1 | γ n , k ( s , t ) | = O h n log 1 h n = o ( 1 ) .
Proof. 
First,
E W n , 0 2 ( t ) = 0 μ 2 y 2 h n 2 K 2 t y h n g ( y ) d y = μ h n f ( t h n u ) t h n u K 2 ( u ) d u .
Localization and boundedness of f on I imply (105).
For k 1 ,
E { W n , 0 ( s ) W n , k ( t ) } = μ 2 K ( ( s x ) / h n ) K ( ( t y ) / h n ) x y h n 2 g k ( x , y ) d x d y = μ 2 K ( u ) K ( v ) ( s h n u ) ( t h n v ) g k ( s h n u , t h n v ) d u d v .
For s , t I and | u | , | v | 1 , all arguments of g k lie, eventually, in the fixed compact set I × I . Assumption 6 thus implies that the right-hand side of (111) is bounded uniformly in n , k , s , t . The marginal expectations are uniformly bounded as well, and (106) follows.
On the other hand, the bounded covariance inequality (88) and the envelope (85) give
| γ n , k ( s , t ) | 4 α ( k ) W n , 0 ( s ) W n , k ( t ) C h n 2 α ( k ) ,
which is (107).
Let c > 0 be such that α ( k ) e 2 c k , and choose
d n = 3 2 c log 1 h n .
Then
k = 1 | γ n , k ( s , t ) | C d n + C h n 2 k > d n e 2 c k C log 1 h n + C h n 2 e 2 c d n C log 1 h n + C h n = O log 1 h n ,
uniformly in s , t I . This proves (108), and multiplication by h n yields (109).    □
Lemma 5
(Variance of arbitrary index subsets). Under the hypotheses of Lemma 4, for every finite J Z ,
sup t I Var i J Z n , i ( t ) C | J | h n .
Proof. 
For each lag k, the number of ordered pairs ( i , j ) J 2 with j i = k is at most | J | . Therefore, using stationarity,
Var i J Z n , i ( t ) | J | Var { W n , 0 ( t ) } + 2 | J | k = 1 | γ n , k ( t , t ) | C | J | 1 h n + log 1 h n C | J | h n ,
because h n log ( 1 / h n ) 0 .    □
Lemma 6
(Diagonal variance and vanishing serial covariance). Fix t > 0 . Suppose there exists a compact interval I t ( 0 , ) , with t in its interior, on which the hypotheses of Lemma 4 hold, and assume that f is continuous at t. Then
h n Var { W n , 0 ( t ) } μ f ( t ) t R ( K ) .
Moreover,
h n k = 1 n 1 1 k n Cov { W n , 0 ( t ) , W n , k ( t ) } = O h n log 1 h n = o ( 1 ) ,
and hence
Var h n n i = 1 n Z n , i ( t ) σ L B 2 ( t ) = μ f ( t ) t R ( K ) .
Proof. 
Multiplying (110) by h n gives
h n E W n , 0 2 ( t ) = μ f ( t h n u ) t h n u K 2 ( u ) d u .
Since t > 0 , the integrand is dominated by an integrable constant multiple of K 2 , and dominated convergence yields
h n E W n , 0 2 ( t ) μ f ( t ) t R ( K ) .
Also, E W n , 0 ( t ) f ( t ) , so
h n { E W n , 0 ( t ) } 2 0 .
This proves (114).
The absolute value of the left-hand side of (115) is bounded by
h n k = 1 n 1 | γ n , k ( t , t ) | ,
which is O ( h n log ( 1 / h n ) ) by Lemma 4. Finally, strict stationarity gives the exact identity
Var h n n i = 1 n Z n , i ( t ) = h n Var { W n , 0 ( t ) } + 2 h n k = 1 n 1 1 k n Cov { W n , 0 ( t ) , W n , k ( t ) } .
The first term converges to σ L B 2 ( t ) and the second tends to zero, proving (116).    □
Lemma 7
(Uniform diagonal variance on compacta). Let I = [ a , b ] ( 0 , ) . Suppose the hypotheses of Lemma 4 hold on I, and suppose that f is continuous on an open neighborhood of I. Then
sup t I h n Var { W n , 0 ( t ) } μ f ( t ) t R ( K ) 0 .
Proof. 
Choose δ > 0 such that
I δ : = [ a δ , b + δ ] ( 0 , )
is contained in the neighborhood on which f is continuous. Since h n 0 , for all sufficiently large n and all t I , | u | 1 implies t h n u I δ . The function x f ( x ) / x is uniformly continuous on the compact interval I δ . Hence, using (110),
sup t I h n E W n , 0 2 ( t ) μ f ( t ) t R ( K ) μ 1 1 K 2 ( u ) sup t I f ( t h n u ) t h n u f ( t ) t d u 0 .
Moreover,
sup t I | E W n , 0 ( t ) | C ,
so
sup t I h n { E W n , 0 ( t ) } 2 C h n 0 .
Subtracting the squared mean from the second moment proves (118).    □
Lemma 8
(Off-diagonal covariance at distinct evaluation points). Let s , t > 0 with s t . Let I s , t ( 0 , ) be a compact interval containing both s and t, and suppose the hypotheses of Lemma 4 hold on I s , t . Then, one has
h n k = 1 n 1 Cov { W n , 0 ( s ) , W n , k ( t ) } = O h n log 1 h n = o ( 1 ) .
Furthermore, for all sufficiently large n,
W n , 0 ( s ) W n , 0 ( t ) = 0 a . s . ,
and therefore
Cov { W n , 0 ( s ) , W n , 0 ( t ) } = E W n , 0 ( s ) E W n , 0 ( t ) = O ( 1 ) ,
so that
h n Cov { W n , 0 ( s ) , W n , 0 ( t ) } 0 .
Proof. 
If 2 h n < | s t | , the supports of K h n ( s · ) and K h n ( t · ) are disjoint, which proves (120). Subtracting the product of the means gives (121); the means converge to f ( s ) and f ( t ) , respectively, so the covariance is O ( 1 ) , proving (122).
For k 1 , formula (111) is localized near the off-diagonal point ( s , t ) . Assumption 6 gives the same uniform O ( 1 ) small-lag bound, while (88) gives the same C α ( k ) / h n 2 large-lag bound. Splitting the series at d n in (112) gives (119).    □

7.5. Uniform Exponential Control

Lemma 9
(Uniform stochastic fluctuation). Let I = [ a , b ] ( 0 , ) . Suppose Assumptions 1–6 hold, f is bounded on a neighborhood of I, the basic bandwidth condition (33) holds, and 
n h n ( log n ) 5 + η for some η > 0 .
Then
sup t I | A n ( t ) E A n ( t ) | = O a . s . log n n h n .
Proof. 
For t I , define
X n , i ( t ) = 1 Y i K h n ( t Y i ) , ξ n , i ( t ) = X n , i ( t ) E X n , i ( t ) .
The row { ξ n , i ( t ) : i 1 } is centered, strictly stationary, and its strong-mixing coefficients are bounded by those of { Y i } . By Lemma 1,
sup t I ξ n , i ( t ) M n : = C h n 1 .
The covariance estimates of Lemma 4 apply with μ absorbed into the constants. Hence, uniformly in t I ,
v n 2 ( t ) : = Var { ξ n , 0 ( t ) } + 2 k = 1 | Cov { ξ n , 0 ( t ) , ξ n , k ( t ) } | C h n + C log 1 h n C h n .
Thus, Proposition 4 applies with a variance proxy uniform in t.
Let
ε n = D log n n h n , x n = n ε n ,
where D > 0 will be chosen later. The denominator in (91) is bounded by
C n h n + 1 h n 2 + n ε n h n ( log n ) 2 .
Relative to n / h n , the second term has ratio
1 / h n 2 n / h n = 1 n h n 0 ,
and the third has ratio
ε n ( log n ) 2 = D ( log n ) 5 / 2 n h n 0
by (123). Therefore, for all sufficiently large n,
sup t I P | A n ( t ) E A n ( t ) | > ε n n c 1 D 2
for some c 1 > 0 .
We now discretize I. Let
Δ n = h n 2 n 3 ,
and let T n be a deterministic grid of I with mesh at most Δ n . Then
| T n | C Δ n 1 = C n 3 h n 2 .
For every t I , choose π n ( t ) T n with | t π n ( t ) | Δ n . By (86), and also after taking expectations,
| ξ n , i ( t ) ξ n , i ( π n ( t ) ) | C h n 2 | t π n ( t ) | C n 3 .
Therefore
sup t I 1 n i = 1 n { ξ n , i ( t ) ξ n , i ( π n ( t ) ) } C n 3 = o ( ε n ) .
Condition (123) implies h n 1 n eventually, and hence
| T n | C n 5 .
By the union bound and (128),
P max u T n | A n ( u ) E A n ( u ) | > ε n C n 5 c 1 D 2 .
Choose D such that c 1 D 2 > 7 . The right-hand side is summable. Borel–Cantelli therefore gives
max u T n | A n ( u ) E A n ( u ) | = O a . s . ( ε n ) .
Combining this with (131) proves (124).    □
Remark 25
(Role of the assumptions in the uniform proof). The proof uses three assumptions at three separate places: geometric strong mixing in Proposition 4, the local bivariate density bound in (126), and Lipschitz continuity of K in (130). The proof therefore uses each primitive assumption at the precise stage for which it is required.

A Triangular-Array Central Limit Theorem

Theorem 5
(Localized strongly mixing triangular-array CLT). For each n, let { Z n , i : i Z } be a centered strictly stationary row. Assume:
(i) 
there exist C , c > 0 , independent of n, such that the row mixing coefficients satisfy
α n ( k ) C e c k , k 1 ;
(ii) 
for a deterministic sequence h n 0 ,
Z n , 0 C h n ;
(iii) 
for every finite J Z ,
Var i J Z n , i C | J | h n ;
(iv) 
Var h n n i = 1 n Z n , i σ 2 ( 0 , ) ;
(v) 
n h n ( log n ) 4 .
Then
h n n i = 1 n Z n , i D N ( 0 , σ 2 ) .
Proof. 
Set
S n = h n n i = 1 n Z n , i .
Choose
p n = n h n log n , q n = A log n ,
where A > 0 will be fixed below. From (135),
n h n ( log n ) 2 .
Hence
p n , q n , q n p n = O ( log n ) 2 n h n 0 ,
and
p n n h n 1 log n 0 .
Also, p n + q n = o ( n ) .
Let
r n = n p n + q n .
Partition { 1 , , n } into r n consecutive big blocks B n , j of length p n , separated by small blocks C n , j of length q n , followed by a terminal remainder D n of length less than p n + q n . Define
U n , j = h n n i B n , j Z n , i , V n , j = h n n i C n , j Z n , i ,
and
R n = h n n i D n Z n , i .
Write
S n = P n + Q n + R n , P n = j = 1 r n U n , j , Q n = j = 1 r n V n , j .
Step 1: Q n and R n are negligible in L 2 . The union of all small blocks has cardinality r n q n . Assumption (133) gives
E Q n 2 h n n C r n q n h n C q n p n 0 .
Similarly,
E R n 2 C p n + q n n 0 .
Thus
Q n + R n 0 in L 2 .
Since sup n E S n 2 < by (134), Cauchy–Schwarz gives
| Cov ( S n , Q n + R n ) | { E S n 2 } 1 / 2 { E ( Q n + R n ) 2 } 1 / 2 0 .
Using P n = S n ( Q n + R n ) ,
Var ( P n ) σ 2 .
Step 2: covariance between distinct big blocks is negligible. By (132),
U n , j C p n n h n .
For j < , the distance between B n , j and B n , is at least
d j = q n + ( j 1 ) ( p n + q n ) .
The bounded covariance inequality yields
| Cov ( U n , j , U n , ) | C p n 2 n h n e c d j .
Summing first over m = j 1 and then over j,
1 j < r n | Cov ( U n , j , U n , ) | C r n p n 2 n h n e c q n m = 0 e c m ( p n + q n ) C p n h n e c q n .
Now
p n h n n h n h n log n = n / h n log n .
From (135), eventually h n 1 n , hence p n / h n n / log n . Since e c q n C n c A , choosing A such that c A > 2 gives
j < | Cov ( U n , j , U n , ) | 0 .
Combining
Var ( P n ) = j = 1 r n Var ( U n , j ) + 2 j < Cov ( U n , j , U n , )
with (144) and (147) yields
j = 1 r n Var ( U n , j ) σ 2 .
Step 3: decoupling of the big blocks. Let U n , 1 * , , U n , r n * be independent random variables with U n , j * = d U n , j . From (92),
E e i s P n j = 1 r n E e i s U n , j 16 r n α ( q n ) .
Because r n n / p n ,
r n C n log n n h n = C n h n log n .
Again, h n 1 n eventually, so r n C n log n . Therefore
r n α ( q n ) C n log n n c A 0
as soon as c A > 2 . Thus
E e i s P n E exp i s j = 1 r n U n , j * 0
for every fixed s R .
Step 4: Lindeberg condition for the independent blocks. By (145) and (139),
max 1 j r n | U n , j * | C log n 0 a . s .
for the independent copies as well. Hence, for every ε > 0 ,
j = 1 r n E ( U n , j * ) 2 1 { | U n , j * | > ε } = 0
for all sufficiently large n. Since
j = 1 r n Var ( U n , j * ) = j = 1 r n Var ( U n , j ) σ 2
by (148), the Lindeberg–Feller theorem yields
j = 1 r n U n , j * D N ( 0 , σ 2 ) .
Equation (151) transfers the same limit to P n . Finally, (143) and Slutsky’s theorem transfer it from P n to S n . This proves (136).    □

7.6. Application of the Abstract CLT to the Kernel Array

Lemma 10
(Central limit theorem for Z n , i ( t ) ). Fix t > 0 with f ( t ) > 0 . Suppose Assumptions 1, 2, 5, and 6 hold, and suppose that f is bounded on a neighborhood of t and continuous at t. Assume the basic bandwidth condition (33) and
n h n ( log n ) 4 .
Then
n h n 1 n i = 1 n W n , i ( t ) E W n , 0 ( t ) D N ( 0 , σ L B 2 ( t ) ) .
Proof. 
The array Z n , i ( t ) is centered and row-wise strictly stationary, and its mixing coefficients are bounded by those of { Y i } . Notice that no inverse-moment assumption is needed for this localized array CLT: inverse moments enter only through the separate random-normalization lemma. Local boundedness and continuity of f are precisely the marginal regularity used in the covariance and variance lemmas. The envelope condition (132) follows from Lemma 1. The subset-variance condition (133) is precisely Lemma 5. The total-variance convergence (134) is Lemma 6. Finally, (153) is (135). Hence, Theorem 5 applies and gives (154).    □

7.7. Proofs of the Main Results

Proof of Theorem 1.
For t I , write
f n ( t ) f ( t ) = μ ^ n { A n ( t ) E A n ( t ) } + ( μ ^ n μ ) E A n ( t ) + { μ E A n ( t ) f ( t ) } .
By Lemma 2, μ ^ n μ almost surely. By Lemma 9,
sup t I | A n ( t ) E A n ( t ) | = O a . s . log n n h n = o a . s . ( 1 ) .
Furthermore, sup t I | E A n ( t ) | C , and Lemma 3 gives
sup t I | μ E A n ( t ) f ( t ) | 0 .
Taking suprema in (155) proves (45).    □
Proof of Theorem 2.
The decomposition (155) is retained, but the conclusion is formulated as an upper stochastic rate rather than as an exact or minimax rate. By Lemma 9,
sup t I μ ^ n { A n ( t ) E A n ( t ) } = O P log n n h n ,
because μ ^ n = O P ( 1 ) . By Lemma 2,
sup t I | ( μ ^ n μ ) E A n ( t ) | = O P ( n 1 / 2 ) .
Since h n 0 ,
n 1 / 2 = o log n n h n .
Finally, Lemma 3 gives
sup t I | μ E A n ( t ) f ( t ) | = O ( h n 2 ) .
Hence
sup t I | f n ( t ) f ( t ) | = O P h n 2 + log n n h n .
Balancing the two proved upper-bound terms gives h n ( log n / n ) 1 / 5 and the resulting upper order ( log n / n ) 2 / 5 . No lower bound, exact limsup constant, or minimax claim is used.    □
Proof of Proposition 1.
Set
U n ( t ) = 1 n i = 1 n W n , i ( t ) E W n , 0 ( t ) .
Lemma 2 gives
D n = μ ^ n μ 1 = O P ( n 1 / 2 )
and
D n = μ ( H n H ) + o P ( n 1 / 2 ) .
By the diagonal variance calculation and the serial-covariance bound in Lemma 6,
Var { U n ( t ) } = O 1 n h n ,
hence
U n ( t ) = O P ( n h n ) 1 / 2
by Chebyshev’s inequality. Therefore
n h n D n U n ( t ) = D n { n h n U n ( t ) } = O P ( n 1 / 2 ) O P ( 1 ) = o P ( 1 ) .
Also, E W n , 0 ( t ) f ( t ) , so
n h n D n E W n , 0 ( t ) = O P ( h n ) = o P ( 1 ) .
It remains to make the normalization–localization covariance explicit. From (97),
Var ( H n ) = O ( n 1 ) ,
and the preceding bound gives
Var { U n ( t ) } = O ( ( n h n ) 1 ) .
Thus, Cauchy–Schwarz yields
Cov { H n H , U n ( t ) } { Var ( H n ) Var ( U n ( t ) ) } 1 / 2 = O 1 n h n ,
which is (56). Since E W n , 0 ( t ) = O ( 1 ) ,
n h n Cov μ E W n , 0 ( t ) ( H n H ) , U n ( t ) C n h n 1 n h n = O ( h n ) 0 ,
which proves (57). Since E W n , 0 ( t ) = O ( 1 ) and Var ( H n ) = O ( n 1 ) , the same calculation gives
n h n Var μ E W n , 0 ( t ) ( H n H ) C h n 0 ,
which is (58). Thus, the entire linearized ratio-normalization correction is negligible in the scaled variance decomposition, including its covariance with the local kernel term.
Finally, using the exact factorization,
f n ( t ) f ( t ) B n ( t ) = U n ( t ) + D n U n ( t ) + D n E W n , 0 ( t ) + { E W n , 0 ( t ) f ( t ) B n ( t ) } .
Lemma 3 gives
n h n { E W n , 0 ( t ) f ( t ) B n ( t ) } = o ( 1 )
under n h n 5 = O ( 1 ) . Combining the four terms proves (59).    □
Proof of Theorem 3.
Proposition 1 gives the complete ratio expansion
n h n { f n ( t ) f ( t ) B n ( t ) } = n h n U n ( t ) + o P ( 1 ) ,
where
U n ( t ) = 1 n i = 1 n W n , i ( t ) E W n , 0 ( t ) .
This expansion already includes, and separately controls, the pure normalization term, the product of the normalization error with the localized fluctuation, and the covariance of the first-order normalization linearization with that localized fluctuation.
By Lemma 10,
n h n U n ( t ) D N ( 0 , σ L B 2 ( t ) ) .
The same lemma uses the variance identity of Lemma 6, where the serial covariance is retained at finite n and then shown to satisfy
| L n ( t ) | = O h n log 1 h n = o ( 1 ) .
Therefore, Slutsky’s theorem applies to the fully audited expansion, not merely to the scalar ratio μ ^ n / μ , and yields (52) with σ L B 2 ( t ) = μ f ( t ) R ( K ) / t .    □
Proof of Theorem 4.
Let t 1 , , t m > 0 be pairwise distinct and let c = ( c 1 , , c m ) R m . If  c = 0 , the Cramér–Wold conclusion is immediate; hence, assume c 0 . Define
W ˜ n , i = = 1 m c W n , i ( t ) , Z ˜ n , i = W ˜ n , i E W ˜ n , 0 .
By Cramér–Wold, it suffices to prove a scalar central limit theorem for Z ˜ n , i .
First, consider its scaled variance. Expanding,
h n Var ( W ˜ n , 0 ) = = 1 m c 2 h n Var { W n , 0 ( t ) } + 2 1 < r m c c r h n Cov { W n , 0 ( t ) , W n , 0 ( t r ) } .
By Lemma 6, the diagonal terms converge to σ L B 2 ( t ) , whereas Lemma 8 makes every off-diagonal term tend to zero. Thus
h n Var ( W ˜ n , 0 ) = 1 m c 2 σ L B 2 ( t ) .
For the lagged terms,
h n k = 1 n 1 Cov ( W ˜ n , 0 , W ˜ n , k ) h n , r = 1 m | c c r | k = 1 n 1 Cov { W n , 0 ( t ) , W n , k ( t r ) } 0
by Lemmas 4 and 8. Therefore
Var h n n i = 1 n Z ˜ n , i = 1 m c 2 σ L B 2 ( t ) .
The row envelope remains O ( h n 1 ) . To verify the subset-variance condition explicitly, let J Z be finite and write
S , J : = i J Z n , i ( t ) .
By Cauchy–Schwarz and Lemma 5,
Var = 1 m c S , J = 1 m | c | Var ( S , J ) 2 C c | J | h n ,
where C c < depends only on the fixed vector c = ( c 1 , , c m ) . Thus, the subset-variance hypothesis of Theorem 5 holds for the Cramér–Wold array. Hence, Theorem 5 gives
h n n i = 1 n Z ˜ n , i D N 0 , = 1 m c 2 σ L B 2 ( t ) .
Applying Proposition 1 coordinatewise shows that the normalizing-factor and bias remainder terms are o P ( 1 ) after multiplication by n h n . The Cramér–Wold device therefore yields (60), with the diagonal limiting covariance matrix.    □
Proof of Corollary 1.
By Theorem 3,
n h n { f n ( t ) f ( t ) B n ( t ) } D N ( 0 , σ L B 2 ( t ) ) .
Since
n h n B n ( t ) = 1 2 m 2 ( K ) f ( t ) n h n 5 ,
the condition n h n 5 0 makes the bias negligible. Slutsky’s theorem gives (63).    □
Proof of Corollary 2.
The assertion is an immediate restatement of Theorem 3, because 
B n ( t ) = h n 2 2 m 2 ( K ) f ( t ) .
Hence, (64) holds for every bandwidth sequence satisfying the hypotheses of Theorem 3. In particular, if  h n n 1 / 5 , then
n h n 5 = O ( 1 ) , n h n ( log n ) 4 n 4 / 5 ( log n ) 4 ,
so this familiar second-order bandwidth is included as a special case.    □
Proof of Corollary 3.
The local Hölder condition and (104) imply
| E W n , 0 ( t ) f ( t ) | C h n β | u | β | K ( u ) | d u = O ( h n β ) .
Hence
n h n | E W n , 0 ( t ) f ( t ) | = O n h n 1 + 2 β 0 .
The local Hölder condition implies both continuity at t and boundedness of f on a sufficiently small compact neighborhood of t. Hence, the exact hypotheses of Lemma 10 are satisfied, so the stochastic term obeys the same localized array CLT. The normalization term is handled directly, without importing a step proved under the C 2 hypotheses of Theorem 3: by Lemma 2,
n h n μ ^ n μ 1 E W n , 0 ( t ) = O P ( h n ) = o P ( 1 ) ,
because E W n , 0 ( t ) f ( t ) . This proves (67).    □

7.8. AMSE, AMISE, and Feasible Inference

Proof of Proposition 2.
The quantity called AMSE t ( h ) is the first-order asymptotic criterion obtained by adding the square of the leading deterministic bias and the leading variance. By Lemma 3,
Bias lead { f n ( t ) } = h 2 2 m 2 ( K ) f ( t ) .
Lemma 6 and Theorem 3 identify the first-order stochastic variance scale of the localized estimator as
V lead ( t ; h ) = 1 n h μ f ( t ) t R ( K ) .
We use this scale to define the first-order AMSE criterion; no assertion that Var { f n ( t ) } itself admits this expansion in expectation is needed. Therefore
AMSE t ( h ) = h 4 4 m 2 ( K ) 2 f ( t ) 2 + 1 n h μ f ( t ) t R ( K ) ,
which is (68). This definition does not silently identify the displayed criterion with the exact finite-sample MSE without an additional uniform-integrability argument.
Put
a t = 1 4 m 2 ( K ) 2 f ( t ) 2 , b t = μ f ( t ) t R ( K ) .
When f ( t ) > 0 and f ( t ) 0 , a t , b t > 0 , and 
d d h a t h 4 + b t n h = 4 a t h 3 b t n h 2 .
The unique positive critical point satisfies
h 5 = b t 4 a t n ,
which is (69). The derivative is negative below and positive above this point, so the critical point is the unique positive minimizer. Substitution gives (70).
For a common bandwidth on I = [ a , b ] , the uniform bias expansion, Lemma 7, and the uniform serial-covariance bound (109) yield the integrated first-order criterion
AMISE I ( h ) = h 4 4 m 2 ( K ) 2 I f ( t ) 2 d t + R ( K ) n h I μ f ( t ) t d t .
Both integrals are finite because I is compact and bounded away from zero. If  I f ( t ) 2 d t > 0 , differentiation gives the unique positive minimizer (72).    □
Proof of Proposition 3.
Fix t > 0 with f ( t ) > 0 . Under the hypotheses of Theorem 3, that theorem implies
f n ( t ) f ( t ) = B n ( t ) + O P ( n h n ) 1 / 2 = o P ( 1 ) ,
because h n 0 and n h n . Together with Lemma 2, this gives
f n ( t ) P f ( t ) , μ ^ n P μ .
Hence, by the continuous mapping theorem,
σ ^ L B 2 ( t ) = μ ^ n f n ( t ) t R ( K ) P μ f ( t ) t R ( K ) = σ L B 2 ( t ) > 0 .
In addition, Assumption 1 gives K 0 , so f n ( t ) 0 and σ ^ L B 2 ( t ) 0 for every sample. Because f ( t ) > 0 and f n ( t ) f ( t ) in probability,
P f n ( t ) > f ( t ) 2 1 ,
which proves strict positivity of the plug-in variance with probability tending to one rather than merely inferring it abstractly from a signed quantity. The plug-in error can also be displayed explicitly:
σ ^ L B 2 ( t ) σ L B 2 ( t ) = R ( K ) t μ { f n ( t ) f ( t ) } + f ( t ) { μ ^ n μ } + { μ ^ n μ } { f n ( t ) f ( t ) } = o P ( 1 ) .
Since the limit is the strictly positive constant σ L B 2 ( t ) , for any 0 < c t < σ L B 2 ( t ) ,
P { σ ^ L B 2 ( t ) c t } P { | σ ^ L B 2 ( t ) σ L B 2 ( t ) | < σ L B 2 ( t ) c t } 1 .
This proves (75). In particular,
P { σ ^ L B 2 ( t ) > 0 } 1 .
Under n h n 5 0 , define
σ ˜ L B ( t ) : = σ ^ L B ( t ) , σ ^ L B 2 ( t ) > 0 , 1 , σ ^ L B 2 ( t ) 0 .
Because σ ^ L B 2 ( t ) P σ L B 2 ( t ) > 0 , the continuous mapping theorem gives
σ ˜ L B ( t ) P σ L B ( t ) .
Corollary 1 and Slutsky’s theorem therefore imply
n h n { f n ( t ) f ( t ) } σ ˜ L B ( t ) D N ( 0 , 1 ) .
The latter statistic and T n ( t ) differ only on E n c : = { σ ^ L B 2 ( t ) 0 } , and  P ( E n c ) 0 . Hence, T n ( t ) D N ( 0 , 1 ) , proving (77).    □
Proof of Corollary 4.
Let
E n : = { σ ^ L B 2 ( t ) > 0 } .
By Proposition 3, P ( E n c ) 0 , and on E n the statistic T n ( t ) equals the usual studentized ratio. Therefore
P E n , z 1 α / 2 n h n { f n ( t ) f ( t ) } σ ^ L B ( t ) z 1 α / 2 = P z 1 α / 2 T n ( t ) z 1 α / 2 + o ( 1 ) 1 α .
Rearranging the event on E n gives the interval (78). This is a pointwise first-order asymptotic coverage statement and does not imply exact finite-sample calibration.    □

Author Contributions

Conceptualization, S.D. and S.B.; methodology, S.D. and S.B.; validation, S.D. and S.B.; formal analysis, S.D. and S.B.; investigation, S.D. and S.B.; resources, S.D. and S.B.; writing—original draft preparation, S.D. and S.B.; writing—review and editing, S.D. and S.B. All authors have read and agreed to the published version of the manuscript.

Funding

Financial support was provided by the Deanship of Graduate Studies and Scientific Research at Qassim University under QU-APC-2026.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

All numerical summaries, tables, and figures reported in the article are generated from a reproducible R workflow using master seed 20240517. The complete R code used for exact length-biased sampling, dependence generation, bandwidth selection, density estimation, HAC and block-resampling calculations, Monte Carlo summaries, validation checks, and figure/table production is available from the corresponding author upon reasonable request. The computational record includes the R version and platform, RNG configuration, package versions, deterministic seed policy, validation outputs, and source-code hashes. The line-transect shrub-width data discussed in Section 5 are publicly distributed in the WData R package [47].

Acknowledgments

The authors gratefully acknowledge the Deanship of Graduate Studies and Scientific Research at Qassim University for its support under QU-APC-2026. The authors would like to thank the four anonymous referees for their careful reading of the manuscript and for their valuable comments and suggestions.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Wicksell, S.D. The corpuscle problem. A mathematical study of a biometric problem. Biometrika 1925, 17, 84–99. [Google Scholar] [CrossRef] [Scilit]
  2. McFadden, J.A. On the lengths of intervals in a stationary point process. J. R. Stat. Soc. Ser. B 1962, 24, 364–382. [Google Scholar] [CrossRef] [Scilit]
  3. Blumenthal, S. Proportional sampling in life length studies. Technometrics 1967, 9, 205–218. [Google Scholar] [CrossRef] [Scilit]
  4. Cox, D.R. Some sampling problems in technology. In New Developments in Survey Sampling; Johnson, N.L., Smith, H., Eds.; Wiley: New York, NY, USA, 1969. [Google Scholar]
  5. Patil, G.P.; Rao, C.R. The weighted distributions: A survey of their applications. In Applications of Statistics; Krishnaiah, P.R., Ed.; North Holland: Amsterdam, The Netherlands, 1977; pp. 383–405. [Google Scholar]
  6. Patil, G.P.; Rao, C.R. Weighted distributions and size-biased sampling with applications to wildlife populations and human families. Biometrics 1978, 34, 179–189. [Google Scholar] [CrossRef] [Scilit]
  7. Coleman, R. An Introduction to Mathematical Stereology; Memoirs No. 3; Department of Theoretical Statistics, University of Aarhus: Aarhus, Denmark, 1979. [Google Scholar]
  8. Vardi, Y. Nonparametric estimation in renewal processes. Ann. Stat. 1982, 10, 772–785. [Google Scholar] [CrossRef] [Scilit]
  9. Wang, M.C. Nonparametric estimation from cross-sectional survival data. J. Am. Stat. Assoc. 1991, 86, 130–143. [Google Scholar] [CrossRef]
  10. Huang, C.Y.; Qin, J. Nonparametric estimation for length-biased and right-censored data. Biometrika 2011, 98, 177–186. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Chan, K.C.; Chen, Y.Q.; Di, C.Z. Proportional mean residual life model for right-censored length-biased data. Biometrika 2012, 99, 995–1000. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Cristóbal, J.A.; Alcalá, J.T. An overview of nonparametric contributions to the problem of functional estimation from biased data. Test 2001, 10, 309–332. [Google Scholar] [CrossRef] [Scilit]
  13. Vardi, Y.; Shepp, L.A.; Logan, B.F. Distribution functions invariant under residual-lifetime and length-biased sampling. Z. Wahrscheinlichkeitstheorie Verwandte Geb. 1981, 56, 415–426. [Google Scholar] [CrossRef] [Scilit]
  14. Asgharian, M.; M’Lan, C.E.; Wolfson, D.B. Length-biased sampling with right censoring: An unconditional approach. J. Am. Stat. Assoc. 2002, 97, 201–209. [Google Scholar] [CrossRef] [Scilit]
  15. Bergeron, P.J.; Asgharian, M.; Wolfson, D.B. Covariate bias induced by length-biased sampling of failure times. J. Am. Stat. Assoc. 2008, 103, 737–742. [Google Scholar] [CrossRef] [Scilit]
  16. Huang, C.Y.; Qin, J. Composite partial likelihood estimation under length-biased sampling, with application to a prevalent cohort study of dementia. J. Am. Stat. Assoc. 2012, 107, 946–957. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Qiu, Z.; Qin, J.; Zhou, Y. Composite estimating equation method for the accelerated failure time model with length-biased sampling data. Scand. J. Stat. 2016, 43, 396–415. [Google Scholar] [CrossRef] [Scilit]
  18. Roy, P.; Fine, J.P.; Kosorok, M.R. Efficiency of naive estimators for accelerated failure time models under length-biased sampling. Scand. J. Stat. 2022, 49, 525–541. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Rosenblatt, M. Remarks on some non-parametric estimates of a density function. Ann. Math. Stat. 1956, 27, 832–837. [Google Scholar] [CrossRef] [Scilit]
  20. Parzen, E. On estimation of a probability density function and mode. Ann. Math. Stat. 1962, 33, 1065–1076. [Google Scholar] [CrossRef] [Scilit]
  21. Nadaraya, E.A. On non-parametric estimates of density functions and regression curves. Theory Probab. Its Appl. 1965, 10, 186–190. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Schuster, E.F. Estimation of a probability density and its derivatives. Ann. Math. Stat. 1969, 40, 1187–1196. [Google Scholar] [CrossRef] [Scilit]
  23. Van Ryzin, J. On strong consistency of density estimates. Ann. Math. Stat. 1969, 40, 1765–1772. [Google Scholar] [CrossRef] [Scilit]
  24. Bouzebda, S.; Didi, S. Some asymptotic properties of kernel regression estimators of the mode for stationary and ergodic continuous time processes. Rev. Mat. Complut. 2021, 34, 811–852. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Bouzebda, S. Weak convergence of the conditional single index U-statistics for locally stationary functional time series. AIMS Math. 2024, 9, 14807–14898. [Google Scholar] [CrossRef] [Scilit]
  26. Bouzebda, S.; El-hadjali, T. Uniform convergence rate of the kernel regression estimator adaptive to intrinsic dimension in presence of censored data. J. Nonparametr. Stat. 2020, 32, 864–914. [Google Scholar] [CrossRef] [Scilit]
  27. Bouzebda, S. On the weak convergence and the uniform-in-bandwidth consistency of the general conditional U-processes based on the copula representation: Multivariate setting. Hacet. J. Math. Stat. 2023, 52, 1303–1348. [Google Scholar] [CrossRef] [Scilit]
  28. Silverman, B.M. Weak and strong uniform consistency of the kernel estimate of a density and its derivatives. Ann. Stat. 1978, 6, 177–184. [Google Scholar] [CrossRef] [Scilit]
  29. Bhattacharyya, B.B.; Franklin, L.A.; Richardson, G.D. A comparison of nonparametric unweighted and length-biased density estimation of fibres. Commun. Stat.-Theory Methods 1988, 17, 3629–3644. [Google Scholar] [CrossRef] [Scilit]
  30. Jones, M.C. Kernel density estimation for length biased data. Biometrika 1991, 78, 511–519. [Google Scholar] [CrossRef] [Scilit]
  31. Guillamón, A.; Navarro, J.; Ruiz, J.M. Kernel density estimation using weighted data. Commun. Stat.-Theory Methods 1998, 27, 2123–2135. [Google Scholar] [CrossRef] [Scilit]
  32. Efromovich, S. Density estimation for biased data. Ann. Stat. 2004, 32, 1137–1161. [Google Scholar] [CrossRef] [Scilit]
  33. Ajami, M.; Fakoor, V.; Jomhoori, S. Bayesian and cross validation estimation bandwidth of kernel density function estimator for length-biased data. J. Stat. Sci. 2011, 5, 41–60. [Google Scholar]
  34. Ajami, M.; Fakoor, V.; Jomhoori, S. Some asymptotic results of kernel density estimator in length-biased sampling. J. Sci. Islam. Repub. Iran. 2013, 24, 55–62. [Google Scholar] [CrossRef] [Scilit]
  35. Borrajo, M.I.; González-Manteiga, W.; Martínez-Miranda, M.D. Bandwidth selection for kernel density estimation with length-biased data. J. Nonparametr. Stat. 2017, 29, 636–668. [Google Scholar] [CrossRef] [Scilit]
  36. Kakizawa, Y. Asymmetric kernel density estimation for biased data. J. Korean Stat. Soc. 2024, 53, 1110–1134. [Google Scholar] [CrossRef] [Scilit]
  37. Arvanitis, S. Concentration inequalities for Kernel density estimators under uniform mixing. J. Korean Stat. Soc. 2023, 52, 440–449. [Google Scholar] [CrossRef] [Scilit]
  38. Shirazi, E.; Ghanbari, B.; Yarmohammadi, M. Wavelet block thresholding for copula density estimation under biased sampling. J. Stat. Comput. Simul. 2023, 93, 2512–2533. [Google Scholar] [CrossRef] [Scilit]
  39. Zaminia, R.; Ajami, M.; Ghafouri, S. Kernel estimators for mean residual lifetime in length-biased sampling. Commun. Stat.-Theory Methods 2024, 53, 7927–7941. [Google Scholar] [CrossRef] [Scilit]
  40. Zaminia, R.; Ajami, M.; Fakoor, V. Berry–Esseen bound for smooth estimator of distribution function under length-biased data. Commun. Stat.-Theory Methods 2024, 53, 1800–1809. [Google Scholar] [CrossRef] [Scilit]
  41. Zaminia, R.; Goodarzi, F.; Hashemi, F. Some kernel estimators for varextropy function under length-biased sampling. Commun. Stat.-Theory Methods 2025, 54, 1557–1577. [Google Scholar] [CrossRef] [Scilit]
  42. Pavithradas, V.; Rajesh, G.; Rajesh, R. Nonparametric estimation of residual extropy function under length-biased sampling. Ric. Di Mat. 2025, 74, 3051–3067. [Google Scholar] [CrossRef] [Scilit]
  43. Bae, T. Rejection sampling for generating random numbers from weighted distributions. Commun. Stat.-Theory Methods 2025, 54, 5023–5034. [Google Scholar] [CrossRef] [Scilit]
  44. Abdel-Salam, A.S.G.; Ahmad, I.A. Nonparametric Functions Estimation Using Biased Data. Mathematics 2025, 13, 4037. [Google Scholar] [CrossRef] [Scilit]
  45. Ren, Y.; Zhang, J.; Xia, Y.; Wang, R.; Xie, F.; Guan, J.; Zhang, H.; Zhou, S. Regression-based conditional independence test with adaptive kernels. Artif. Intell. 2025, 347, 104391. [Google Scholar] [CrossRef] [Scilit]
  46. Chen, R.; Zhang, C.; Lu, Y.; Wang, S.; Du, S.; Zhang, Y. Bivariate-dependent remaining useful life prediction based on copulas and Tweedie exponential dispersion process. Reliab. Eng. Syst. Saf. 2026, 274, 112422. [Google Scholar] [CrossRef] [Scilit]
  47. Sánchez Martínez, N.; Borrajo García, M.I.; Conde Amboage, M. WData: Statistical Inference for Weighted Data, Version 0.1.1; CRAN: Vienna, Austria, 2025. [CrossRef] [Scilit]
  48. Muttlak, H.A. Some Aspects of Ranked Set Sampling with Size Biased Probability of Selection. Ph.D. Thesis, University of Wyoming, Laramie, WY, USA, 1988. [Google Scholar]
  49. Merlevède, F.; Peligrad, M.; Rio, E. Bernstein inequality and moderate deviations under strong mixing conditions. In High Dimensional Probability V: The Luminy Volume; Institute of Mathematical Statistics (IMS) Collections; Institute of Mathematical Statistics: Beachwood, OH, USA, 2009; Volume 5, pp. 273–292. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Mean length-biased kernel density estimate and true target density for n = 2000 , under the four dependence levels and h AMISE ( I ; n ) , based on B = 2000 replications. Dependence levels are distinguished by color and line type.
Figure 1. Mean length-biased kernel density estimate and true target density for n = 2000 , under the four dependence levels and h AMISE ( I ; n ) , based on B = 2000 replications. Dependence levels are distinguished by color and line type.
Symmetry 18 01472 g001
Figure 2. MISE versus n on log–log axes for the four values of ϕ , by model, under h AMISE ( I ; n ) , B = 2000 . Error bars are ± 2 MCSE { MISE ^ } . Panel-specific vertical scales are used to resolve within-model behavior; absolute MISE levels should therefore be compared from Table 13, rather than from panel height.
Figure 2. MISE versus n on log–log axes for the four values of ϕ , by model, under h AMISE ( I ; n ) , B = 2000 . Error bars are ± 2 MCSE { MISE ^ } . Panel-specific vertical scales are used to resolve within-model behavior; absolute MISE levels should therefore be compared from Table 13, rather than from panel height.
Symmetry 18 01472 g002
Figure 3. Finite-sample Gaussian diagnostics for Z n ( q . 50 ) under h US ( I ; n ) , with n = 2000 , for all three models and all four values of ϕ , based on B = 2000 replications. The two diagnostics are stacked at near-full text width so that model labels, axes, empirical Q–Q points, and tail departures remain legible at printed size.
Figure 3. Finite-sample Gaussian diagnostics for Z n ( q . 50 ) under h US ( I ; n ) , with n = 2000 , for all three models and all four values of ϕ , based on B = 2000 replications. The two diagnostics are stacked at near-full text width so that model labels, axes, empirical Q–Q points, and tail departures remain legible at printed size.
Symmetry 18 01472 g003aSymmetry 18 01472 g003b
Figure 4. MISE versus ϕ for n { 100 , 250 , 500 , 1000 , 2000 } , by model, under h AMISE ( I ; n ) , B = 2000 . Sample size is encoded by color and marker. Panel-specific vertical scales are used to preserve within-model resolution; absolute MISE magnitudes across models should therefore be compared from Table 13, not from relative panel height. Axis and strip labels are enlarged to remain legible at printed journal width.
Figure 4. MISE versus ϕ for n { 100 , 250 , 500 , 1000 , 2000 } , by model, under h AMISE ( I ; n ) , B = 2000 . Sample size is encoded by color and marker. Panel-specific vertical scales are used to preserve within-model resolution; absolute MISE magnitudes across models should therefore be compared from Table 13, not from relative panel height. Axis and strip labels are enlarged to remain legible at printed journal width.
Symmetry 18 01472 g004
Figure 5. Finite-sample variance-inflation ratio VIR ( t ) over the evaluation grid for ϕ { 0.3 , 0.6 , 0.8 } , faceted by sample size and model, under h AMISE ( I ; n ) , B = 2000 . The dotted reference line corresponds to VIR = 1 . Line type is used in addition to color to preserve readability in greyscale.
Figure 5. Finite-sample variance-inflation ratio VIR ( t ) over the evaluation grid for ϕ { 0.3 , 0.6 , 0.8 } , faceted by sample size and model, under h AMISE ( I ; n ) , B = 2000 . The dotted reference line corresponds to VIR = 1 . Line type is used in addition to color to preserve readability in greyscale.
Symmetry 18 01472 g005
Figure 6. Empirical coverage of nominal 95% intervals in Design A, comparing the original first-order plug-in interval with the quarter-order full-ratio HAC correction, across models, evaluation points, sample sizes, and dependence levels ( B = 2000 ). The dashed horizontal line denotes 0.95. Typography was regenerated at publication size, and line type/marker are used in addition to color.
Figure 6. Empirical coverage of nominal 95% intervals in Design A, comparing the original first-order plug-in interval with the quarter-order full-ratio HAC correction, across models, evaluation points, sample sizes, and dependence levels ( B = 2000 ). The dashed horizontal line denotes 0.95. Typography was regenerated at publication size, and line type/marker are used in addition to color.
Symmetry 18 01472 g006
Figure 7. Distribution of selected bandwidths relative to the oracle AMISE bandwidth. The four target columns include the Gamma-mixture shape-sensitivity experiment. A ratio of one marks the oracle reference; the unstable SBoot2-Jones selections are retained in the display.
Figure 7. Distribution of selected bandwidths relative to the oracle AMISE bandwidth. The four target columns include the Gamma-mixture shape-sensitivity experiment. A ratio of one marks the oracle reference; the unstable SBoot2-Jones selections are retained in the display.
Symmetry 18 01472 g007
Figure 8. Coverage of the first-order plug-in interval and the full-ratio HAC corrections across models, evaluation points, sample sizes, and dependence levels. Error bars report Monte Carlo uncertainty.
Figure 8. Coverage of the first-order plug-in interval and the full-ratio HAC corrections across models, evaluation points, sample sizes, and dependence levels. Error bars report Monte Carlo uncertainty.
Symmetry 18 01472 g008
Figure 9. Coverage in the moving-block-bootstrap validation experiment. The outer Monte Carlo size is 100 and the inner bootstrap size is 200; error bars reflect outer Monte Carlo uncertainty.
Figure 9. Coverage in the moving-block-bootstrap validation experiment. The outer Monte Carlo size is 100 and the inner bootstrap size is 200; error bars reflect outer Monte Carlo uncertainty.
Symmetry 18 01472 g009
Figure 10. Comparison of the Gaussian-copula AR(1) and Frank-copula Markov designs on their common robustness grid. The comparison concerns finite-sample behavior and does not enlarge the assumptions of the asymptotic theorems.
Figure 10. Comparison of the Gaussian-copula AR(1) and Frank-copula Markov designs on their common robustness grid. The comparison concerns finite-sample behavior and does not enlarge the assumptions of the asymptotic theorems.
Symmetry 18 01472 g010
Figure 11. Estimator benchmark on a common density scale. The ordinary KDE is a negative control for the observed length-biased marginal g, not an estimator of the target density f.
Figure 11. Estimator benchmark on a common density scale. The ordinary KDE is a negative control for the observed length-biased marginal g, not an estimator of the target density f.
Symmetry 18 01472 g011
Table 1. Positioning of the present contribution relative to closely related work. “Feasible BW” denotes a data-driven bandwidth rule; “dependent inference” denotes an inference procedure explicitly adjusted for serial dependence. Entries are intentionally conservative.
Table 1. Positioning of the present contribution relative to closely related work. “Feasible BW” denotes a data-driven bandwidth rule; “dependent inference” denotes an inference procedure explicitly adjusted for serial dependence. Entries are intentionally conservative.
ReferenceEstimator/SettingDependenceFeasible BWPrincipal Relevance Here
Jones (1991) [30]inverse-weighted length-biased KDEindependentnoclassical ratio estimator and random normalizer
Borrajo et al. (2017) [35]Jones-type length-biased KDEindependentyesrule-of-thumb, cross-validation and bootstrap bandwidth selection
Ajami et al. (2013) [34]length-biased KDEindependentnostrong consistency and pointwise asymptotic normality
Kakizawa (2024) [36]asymmetric-kernel biased-data KDEindependentnot centralrecent independent biased-data asymptotics
Arvanitis (2023) [37]ordinary KDE under mixingdependentnoconcentration under a dependence framework distinct from the present inverse-weighted array
Present paperJones-type length-biased KDEgeometric α -mixingnumerically studieddependent asymptotic theory, covariance localization, feasible bandwidth comparison, and finite-sample dependence-aware inference study
Table 2. Bandwidth conditions used by the principal conclusions. Every row is cumulative with h n 0 and n h n . The table separates almost-sure consistency from stochastic-rate and distributional assertions.
Table 2. Bandwidth conditions used by the principal conclusions. Every row is cumulative with h n 0 and n h n . The table separates almost-sure consistency from stochastic-rate and distributional assertions.
ConclusionAdditional Bandwidth ConditionLogical Status
Strong uniform consistency, Theorem 1 n h n / ( log n ) 5 + η sup t I | f n ( t ) f ( t ) | 0 almost surely; no quantitative almost-sure rate for the full ratio estimator is claimed.
Uniform quantitative bound, Theorem 2 n h n / ( log n ) 5 + η , plus local C 2 smoothness O P { h n 2 + log n / ( n h n ) } , i.e., a uniform stochastic upper rate only.
Bias-corrected pointwise/joint CLT n h n / ( log n ) 4 and n h n 5 = O ( 1 ) Gaussian limit after subtracting the second-order bias.
Uncorrected/studentized pointwise CLT and confidence interval n h n / ( log n ) 4 and n h n 5 0 Undersmoothing makes the h n 2 bias negligible at the n h n scale.
Local Hölder centered CLT n h n / ( log n ) 4 and n h n 1 + 2 β 0 The O ( h n β ) local bias is negligible at the Gaussian scale.
Table 3. Oracle first-order AMISE reference bandwidth h AMISE ( I ; n ) = C AMISE ( I ) n 1 / 5 and the undersmoothing bandwidth h US ( I ; n ) = C AMISE ( I ) n 1 / 4 , computed by numerical quadrature from the true f and analytic f on I = [ q . 10 , q . 90 ] , for every model and every simulated n. C AMISE ( I ) does not depend on n; it is repeated on each row for convenience.
Table 3. Oracle first-order AMISE reference bandwidth h AMISE ( I ; n ) = C AMISE ( I ) n 1 / 5 and the undersmoothing bandwidth h US ( I ; n ) = C AMISE ( I ) n 1 / 4 , computed by numerical quadrature from the true f and analytic f on I = [ q . 10 , q . 90 ] , for every model and every simulated n. C AMISE ( I ) does not depend on n; it is repeated on each row for convenience.
n C AMISE ( I ) h AMISE ( I ; n ) h US ( I ; n )
Gamma(3,1), C AMISE ( I ) = 3.30666
1003.306661.316411.04566
2503.306661.095980.83158
5003.306660.954100.69927
10003.306660.830600.58802
20003.306660.723080.49446
Lognormal(0.5,0.4), C AMISE ( I ) = 1.33964
1001.339640.533320.42363
2501.339640.444020.33690
5001.339640.386540.28330
10001.339640.336500.23823
20001.339640.292940.20032
Weibull(2,2), C AMISE ( I ) = 2.28059
1002.280590.907920.72118
2502.280590.755890.57354
5002.280590.658040.48229
10002.280590.572860.40555
20002.280590.498700.34103
Table 4. Computational experiments and replication counts. The primary, robustness, bandwidth, benchmark, and nested resampling designs are reported separately.
Table 4. Computational experiments and replication counts. The primary, robustness, bandwidth, benchmark, and nested resampling designs are reported separately.
ExperimentReplication CountRole
Design A: Gaussian-copula AR(1)60 cells, B = 2000 Primary estimation and inference design.
Design B: Frank-copula Markov18 cells, B = 2000 Dependence robustness experiment.
Feasible bandwidth selectors B = 60 ; Gamma-mixture B = 30 Feasible-selector comparison.
Estimator benchmark B = 200 Jones versus Bhattacharyya-type estimator on identical samples.
Dependent resampling B outer = 100 , B boot = 200 Dependent-resampling validation of finite-sample inference.
Table 5. Feasible bandwidth selection relative to the oracle AMISE bandwidth. The ratio range is computed after averaging the replication-level mean ratio over the three primary models within each ( n , ϕ ) cell then taking the range across n { 100 , 500 , 1000 } and ϕ { 0 , 0.6 } . Boundary frequencies are reported on the same aggregated grid.
Table 5. Feasible bandwidth selection relative to the oracle AMISE bandwidth. The ratio range is computed after averaging the replication-level mean ratio over the three primary models within each ( n , ϕ ) cell then taking the range across n { 100 , 500 , 1000 } and ϕ { 0 , 0.6 } . Boundary frequencies are reported on the same aggregated grid.
MethodRange of h ¯ / h oracle Boundary FrequencyFailure Frequency
BGM-RT-style0.936–0.984not applicable0.000
WCV-Jones0.892–1.0770.056–0.1220.000
Block-CV Jones0.905–1.1080.044–0.1440.000
SBoot1-Jones1.612–1.8230.000–0.0000.000
SBoot2-Jones3.713–3.9370.978–1.0000.000
Table 6. Strong-dependence median-point inference at n = 2000 , ϕ = 0.8 ( B = 2000 ). The last column is the ratio of mean interval length to the original plug-in interval.
Table 6. Strong-dependence median-point inference at n = 2000 , ϕ = 0.8 ( B = 2000 ). The last column is the ratio of mean interval length to the original plug-in interval.
ModelMethodCoverageMCSEMean LengthLength Ratio
GammaPlug-in0.87450.00740.050721.000
GammaHAC, L n 1 / 4 0.93200.00560.059031.164
GammaHAC, L n 1 / 3 0.94050.00530.061041.203
LognormalPlug-in0.85250.00790.104161.000
LognormalHAC, L n 1 / 4 0.90700.00650.115131.105
LognormalHAC, L n 1 / 3 0.91800.00610.119591.148
WeibullPlug-in0.86100.00770.077151.000
WeibullHAC, L n 1 / 4 0.92800.00580.094901.230
WeibullHAC, L n 1 / 3 0.93550.00550.098661.279
Table 7. Dependent-resampling validation at q . 50 and ϕ = 0.8 (outer B = 100 , inner block-bootstrap B = 200 ). The last column gives the mean block-bootstrap to plug-in interval-length ratio.
Table 7. Dependent-resampling validation at q . 50 and ϕ = 0.8 (outer B = 100 , inner block-bootstrap B = 200 ). The last column gives the mean block-bootstrap to plug-in interval-length ratio.
nModelPlug-InHACBlock BootstrapLength Ratio
1000Gamma0.830.920.921.200
1000Lognormal0.780.830.831.105
1000Weibull0.820.920.951.241
2000Gamma0.900.930.941.168
2000Lognormal0.880.910.911.113
2000Weibull0.860.940.941.264
Table 8. Frank-copula Markov robustness design at n = 2000 ( B = 2000 ). ϕ indexes the Gaussian-AR(1) benchmark used only to define the matched Kendall τ ; θ is the corresponding Frank parameter. Coverage is reported at q . 50 .
Table 8. Frank-copula Markov robustness design at n = 2000 ( B = 2000 ). ϕ indexes the Gaussian-AR(1) benchmark used only to define the matched Kendall τ ; θ is the corresponding Frank parameter. Coverage is reported at q . 50 .
ϕ Model θ τ Lag-1MISEMCSEPlug-InHAC
0.6Gamma4.2960.4100.5290.0005830.0000100.9520.946
0.6Lognormal4.2960.4100.5110.0014760.0000260.9570.944
0.6Weibull4.2960.4100.5540.0008700.0000150.9580.953
0.8Gamma7.6770.5900.7240.0009010.0000180.9160.934
0.8Lognormal7.6770.5900.7010.0021780.0000490.8830.909
0.8Weibull7.6770.5900.7510.0014170.0000300.8650.926
Table 9. MISE benchmark at n = 2000 ( B = 200 ). “Bhattacharyya-type” denotes the independently coded smooth-then-divide implementation used in the computational workflow.
Table 9. MISE benchmark at n = 2000 ( B = 200 ). “Bhattacharyya-type” denotes the independently coded smooth-then-divide implementation used in the computational workflow.
ϕ ModelJones OracleJones WCVJones Block-CVBhattacharyya-Type
0.0Gamma0.0004510.0008680.0008650.010941
0.0Lognormal0.0011570.0023340.0022290.025066
0.0Weibull0.0005940.0015390.0015580.018411
0.6Gamma0.0005520.0010960.0011160.011150
0.6Lognormal0.0014070.0025910.0026950.025512
0.6Weibull0.0008060.0018060.0019560.018620
Table 10. Selected diagnostics from the computational validation suite; all 15 prespecified checks passed.
Table 10. Selected diagnostics from the computational validation suite; all 15 prespecified checks passed.
CheckDefinitive Diagnostic
Epanechnikov constants K = 1 , R ( K ) = 0.600000 , m 2 ( K ) = 0.200000 .
Exact length-biased sampler ( n = 200,000)Maximum relative theoretical/empirical quantile discrepancy = 0.00329 .
Gaussian-AR(1) marginal checkKS p ( U ) = p ( Y ) = 0.2916 at ϕ = 0.6 , n = 20,000.
Frank-Markov marginal checkKS p ( U ) = p ( Y ) = 0.3634 , θ = 4.2957 , n = 20,000.
Inverse-weighted EDFTotal mass = 1.0000000000 .
Density integrationWide-grid integral of a test f n = 0.99874 .
Oracle AMISE validationMaximum relative closed-form/numerical-minimizer discrepancy 2.82 × 10 10 over four models.
HAC diagnosticBartlett LRV = 0.283187 , finite and non-negative in the validation cell.
Moving-block bootstrapLength preserved (500) and identical under repeated fixed-seed execution.
Serial reproducibilityIdentical ISE vectors under identical deterministic cell seeds.
Table 11. Model-specific evaluation points and true density values.
Table 11. Model-specific evaluation points and true density values.
Model q . 25 q . 50 q . 75
t f ( t ) t f ( t ) t f ( t )
Gamma1.727300.265182.674060.246593.920400.15241
Lognormal1.258860.631081.648720.604932.159330.36791
Weibull1.072720.402271.665110.416282.354820.29435
The three points are the first quartile, median, and third quartile of the corresponding target distribution. These fixed values are not repeated in the pointwise performance tables.
Table 12. Pointwise bias, variance, and mean squared error under the oracle AMISE bandwidth.
Table 12. Pointwise bias, variance, and mean squared error under the oracle AMISE bandwidth.
n ϕ q . 25 q . 50 q . 75
Bias Variance MSE Bias Variance MSE Bias Variance MSE
Gamma
1000−0.029530.000800.00167−0.009150.000770.000850.004260.000540.00056
1000.3−0.030070.001050.00196−0.009770.000850.000940.003550.000670.00068
1000.6−0.032150.001820.00285−0.009240.001080.001160.005280.001080.00111
1000.8−0.036360.003960.00528−0.007620.002050.002110.008890.001970.00205
2500−0.020050.000460.00086−0.006460.000400.000440.002480.000260.00026
2500.3−0.020690.000590.00102−0.005820.000440.000470.002450.000310.00032
2500.6−0.021620.000900.00137−0.006250.000620.000660.002920.000480.00049
2500.8−0.023980.001930.00251−0.005380.001110.001140.005370.000880.00091
5000−0.015510.000310.00055−0.005030.000260.000290.001760.000140.00014
5000.3−0.015710.000390.00063−0.005190.000260.000290.001580.000170.00018
5000.6−0.016210.000530.00080−0.005080.000360.000380.002130.000270.00028
5000.8−0.017400.001140.00144−0.004870.000670.000700.003180.000510.00052
10000−0.011530.000200.00033−0.004080.000160.000180.000890.000090.00009
10000.3−0.011920.000230.00037−0.003400.000160.000170.001050.000100.00010
10000.6−0.011250.000330.00046−0.003610.000220.000230.001170.000150.00015
10000.8−0.013020.000620.00079−0.003160.000370.000380.001490.000270.00027
20000−0.009100.000120.00020−0.002610.000090.000090.000890.000050.00005
20000.3−0.008770.000140.00022−0.002920.000100.000110.000950.000060.00006
20000.6−0.008910.000190.00027−0.002780.000120.000130.000810.000080.00008
20000.8−0.009750.000370.00047−0.003730.000210.000230.001400.000150.00016
Lognormal
1000−0.065520.003810.00810−0.029590.003840.004710.007490.002760.00281
1000.3−0.070850.005350.01037−0.029910.004020.004910.010280.003370.00348
1000.6−0.073130.009060.01441−0.029610.005600.006470.013190.005280.00545
1000.8−0.078530.018240.02441−0.027400.010230.010980.016800.009450.00974
2500−0.046370.002240.00439−0.019210.002040.002410.005070.001380.00141
2500.3−0.050260.002860.00539−0.018250.002340.002670.007640.001680.00174
2500.6−0.050810.004300.00688−0.020480.002940.003360.007660.002410.00247
2500.8−0.054460.008740.01170−0.020310.005200.005620.009680.004440.00453
5000−0.036650.001410.00276−0.014910.001260.001480.003600.000840.00085
5000.3−0.035510.001760.00302−0.015180.001340.001570.003380.000950.00097
5000.6−0.039670.002630.00420−0.014910.001820.002040.005690.001340.00137
5000.8−0.040920.004970.00665−0.012960.002980.003150.005440.002400.00243
10000−0.028870.000930.00176−0.010430.000850.000960.002930.000430.00044
10000.3−0.027230.001110.00186−0.011740.000870.001010.003060.000570.00058
10000.6−0.028590.001550.00237−0.009660.001110.001210.004310.000780.00080
10000.8−0.029250.002740.00359−0.011030.001720.001850.004300.001330.00135
20000−0.022100.000580.00107−0.007590.000490.000550.002910.000290.00030
20000.3−0.020780.000650.00108−0.008740.000540.000620.002630.000320.00032
20000.6−0.023310.000910.00146−0.008260.000620.000690.003590.000420.00043
20000.8−0.023110.001640.00217−0.008700.000990.001060.002970.000700.00071
Weibull
1000−0.037690.001670.00309−0.025200.001820.00246−0.002660.001500.00151
1000.3−0.040910.002260.00394−0.026330.001870.002560.000000.001940.00194
1000.6−0.040750.004110.00577−0.024950.002860.00348−0.000100.003320.00332
1000.8−0.048180.008090.01041−0.022060.004750.005240.008110.006190.00626
2500−0.026720.000970.00168−0.018710.000990.00134−0.001300.000690.00069
2500.3−0.027490.001240.00200−0.018150.001030.00136−0.001420.000900.00091
2500.6−0.029020.002070.00291−0.018160.001480.001810.000030.001550.00155
2500.8−0.030040.004220.00512−0.015940.002870.003120.002480.003080.00309
5000−0.020730.000660.00109−0.013130.000610.00078−0.000610.000400.00040
5000.3−0.021670.000800.00127−0.013760.000660.00085−0.000500.000500.00050
5000.6−0.021790.001260.00174−0.013750.000930.00112−0.000590.000840.00084
5000.8−0.022970.002660.00318−0.011120.001750.001880.002730.001860.00186
10000−0.015670.000410.00065−0.010670.000360.00047−0.001260.000240.00024
10000.3−0.016010.000500.00076−0.011350.000380.00050−0.001030.000280.00028
10000.6−0.016020.000730.00099−0.011190.000560.00068−0.000780.000450.00045
10000.8−0.017490.001410.00171−0.010940.000960.001080.001100.000910.00091
20000−0.011950.000250.00039−0.008530.000210.00028−0.001020.000140.00014
20000.3−0.011720.000310.00045−0.008070.000240.00030−0.000980.000160.00016
20000.6−0.012340.000430.00058−0.008560.000300.00037−0.000460.000250.00025
20000.8−0.012970.000800.00097−0.007730.000560.00062−0.000180.000450.00045
Rows are grouped by model. The entries are Monte Carlo estimates ( B = 2000 ) at the three evaluation points reported in Table 11 under the oracle bandwidth h AMISE ( I ; n ) . Variance uses the population (divide-by-B) convention, and  MSE = Bias 2 + Variance was verified numerically for every row.
Table 13. Mean integrated squared error under the oracle AMISE bandwidth.
Table 13. Mean integrated squared error under the oracle AMISE bandwidth.
n ϕ MISE se ( MISE )
Gamma
10000.0037190.000073
1000.30.0044160.000095
1000.60.0064550.000155
1000.80.0120280.000299
25000.0018940.000034
2500.30.0022080.000044
2500.60.0031990.000067
2500.80.0057680.000131
50000.0011850.000020
5000.30.0013240.000024
5000.60.0018380.000037
5000.80.0034060.000079
100000.0007090.000012
10000.30.0007860.000013
10000.60.0010520.000020
10000.80.0018270.000037
200000.0004110.000006
20000.30.0004620.000007
20000.60.0005900.000011
20000.80.0010710.000022
Lognormal
10000.0075490.000145
1000.30.0093260.000191
1000.60.0132930.000313
1000.80.0236540.000554
25000.0039310.000071
2500.30.0047940.000089
2500.60.0064030.000133
2500.80.0113070.000252
50000.0024030.000040
5000.30.0026900.000049
5000.60.0037910.000073
5000.80.0063170.000140
100000.0014890.000023
10000.30.0016590.000027
10000.60.0022020.000042
10000.80.0034890.000072
200000.0008940.000013
20000.30.0009720.000015
20000.60.0012790.000022
20000.80.0019960.000039
Weibull
10000.0051420.000112
1000.30.0063710.000137
1000.60.0097630.000223
1000.80.0173450.000411
25000.0026900.000051
2500.30.0031670.000062
2500.60.0048640.000114
2500.80.0088970.000208
50000.0016440.000030
5000.30.0019120.000038
5000.60.0028170.000063
5000.80.0054430.000131
100000.0009820.000017
10000.30.0011280.000021
10000.60.0016040.000031
10000.80.0028920.000061
200000.0005790.000009
20000.30.0006630.000011
20000.60.0009030.000016
20000.80.0015700.000032
se(MISE) denotes the Monte Carlo standard error of the reported MISE, sd ( ISE ) / B , B = 2000 . MISE is reported under the oracle AMISE bandwidth h AMISE ( I ; n ) only.
Table 14. Empirical coverage probabilities for nominal 95% studentized confidence intervals; Monte Carlo standard errors are shown in parentheses.
Table 14. Empirical coverage probabilities for nominal 95% studentized confidence intervals; Monte Carlo standard errors are shown in parentheses.
n ϕ q . 25 q . 50 q . 75
Gamma
10000.963 (0.004)0.965 (0.004)0.947 (0.005)
1000.30.949 (0.005)0.960 (0.004)0.923 (0.006)
1000.60.895 (0.007)0.932 (0.006)0.854 (0.008)
1000.80.781 (0.009)0.855 (0.008)0.726 (0.010)
25000.972 (0.004)0.966 (0.004)0.952 (0.005)
2500.30.950 (0.005)0.961 (0.004)0.930 (0.006)
2500.60.917 (0.006)0.928 (0.006)0.867 (0.008)
2500.80.805 (0.009)0.849 (0.008)0.750 (0.010)
50000.965 (0.004)0.965 (0.004)0.956 (0.005)
5000.30.951 (0.005)0.964 (0.004)0.930 (0.006)
5000.60.929 (0.006)0.941 (0.005)0.873 (0.007)
5000.80.813 (0.009)0.838 (0.008)0.762 (0.010)
100000.967 (0.004)0.960 (0.004)0.945 (0.005)
10000.30.952 (0.005)0.959 (0.004)0.925 (0.006)
10000.60.935 (0.005)0.941 (0.005)0.885 (0.007)
10000.80.832 (0.008)0.859 (0.008)0.772 (0.009)
200000.964 (0.004)0.966 (0.004)0.951 (0.005)
20000.30.957 (0.005)0.955 (0.005)0.938 (0.005)
20000.60.926 (0.006)0.949 (0.005)0.908 (0.006)
20000.80.849 (0.008)0.862 (0.008)0.782 (0.009)
Lognormal
10000.968 (0.004)0.970 (0.004)0.969 (0.004)
1000.30.938 (0.005)0.968 (0.004)0.955 (0.005)
1000.60.886 (0.007)0.940 (0.005)0.899 (0.007)
1000.80.777 (0.009)0.870 (0.008)0.786 (0.009)
25000.969 (0.004)0.972 (0.004)0.960 (0.004)
2500.30.950 (0.005)0.970 (0.004)0.949 (0.005)
2500.60.916 (0.006)0.943 (0.005)0.905 (0.007)
2500.80.818 (0.009)0.883 (0.007)0.792 (0.009)
50000.973 (0.004)0.973 (0.004)0.956 (0.005)
5000.30.961 (0.004)0.967 (0.004)0.952 (0.005)
5000.60.921 (0.006)0.945 (0.005)0.911 (0.006)
5000.80.835 (0.008)0.896 (0.007)0.814 (0.009)
100000.961 (0.004)0.963 (0.004)0.968 (0.004)
10000.30.953 (0.005)0.966 (0.004)0.942 (0.005)
10000.60.935 (0.005)0.946 (0.005)0.912 (0.006)
10000.80.859 (0.008)0.890 (0.007)0.824 (0.009)
200000.964 (0.004)0.963 (0.004)0.959 (0.004)
20000.30.963 (0.004)0.961 (0.004)0.951 (0.005)
20000.60.929 (0.006)0.946 (0.005)0.920 (0.006)
20000.80.858 (0.008)0.889 (0.007)0.851 (0.008)
Weibull
10000.967 (0.004)0.950 (0.005)0.937 (0.005)
1000.30.952 (0.005)0.946 (0.005)0.911 (0.006)
1000.60.901 (0.007)0.904 (0.007)0.804 (0.009)
1000.80.776 (0.009)0.844 (0.008)0.676 (0.010)
25000.970 (0.004)0.955 (0.005)0.949 (0.005)
2500.30.963 (0.004)0.946 (0.005)0.911 (0.006)
2500.60.907 (0.006)0.903 (0.007)0.835 (0.008)
2500.80.791 (0.009)0.814 (0.009)0.685 (0.010)
50000.964 (0.004)0.953 (0.005)0.951 (0.005)
5000.30.951 (0.005)0.948 (0.005)0.930 (0.006)
5000.60.911 (0.006)0.921 (0.006)0.848 (0.008)
5000.80.790 (0.009)0.828 (0.008)0.679 (0.010)
100000.963 (0.004)0.959 (0.004)0.947 (0.005)
10000.30.956 (0.005)0.952 (0.005)0.924 (0.006)
10000.60.923 (0.006)0.910 (0.006)0.860 (0.008)
10000.80.826 (0.008)0.827 (0.008)0.727 (0.010)
200000.973 (0.004)0.954 (0.005)0.949 (0.005)
20000.30.956 (0.005)0.949 (0.005)0.927 (0.006)
20000.60.929 (0.006)0.932 (0.006)0.868 (0.008)
20000.80.843 (0.008)0.846 (0.008)0.765 (0.009)
Each entry is empirical coverage with its binomial Monte Carlo standard error in parentheses, { p ^ ( 1 p ^ ) / 2000 } 1 / 2 . The nominal coverage level is 0.95 and all entries use the undersmoothing bandwidth h US ( I ; n ) . Persistent undercoverage under stronger dependence is interpreted as a finite-sample limitation of first-order plug-in inference.
Table 15. Monte Carlo behavior of the estimated mean parameter μ ^ n .
Table 15. Monte Carlo behavior of the estimated mean parameter μ ^ n .
n ϕ μ Mean ( μ ^ n ) SD ( μ ^ n )
Gamma
10003.00003.01570.2097
1000.33.00003.02390.2700
1000.63.00003.03820.3869
1000.83.00003.10170.5534
25003.00003.00250.1309
2500.33.00003.00960.1708
2500.63.00003.01730.2430
2500.83.00003.04380.3540
50003.00003.00560.0950
5000.33.00003.00150.1228
5000.63.00003.01150.1739
5000.83.00003.02950.2635
100003.00002.99740.0677
10000.33.00003.00230.0864
10000.63.00003.00090.1232
10000.83.00003.00630.1812
200003.00003.00060.0466
20000.33.00003.00060.0627
20000.63.00003.00140.0877
20000.83.00003.01000.1303
Lognormal
10001.78601.78800.0738
1000.31.78601.79560.1014
1000.61.78601.79860.1410
1000.81.78601.81020.2142
25001.78601.78690.0466
2500.31.78601.79060.0633
2500.61.78601.79450.0890
2500.81.78601.79830.1350
50001.78601.78550.0328
5000.31.78601.78650.0438
5000.61.78601.78990.0643
5000.81.78601.79110.0961
100001.78601.78600.0235
10000.31.78601.78570.0316
10000.61.78601.78830.0460
10000.81.78601.78760.0673
200001.78601.78640.0168
20000.31.78601.78660.0225
20000.61.78601.78880.0322
20000.81.78601.78710.0486
Weibull
10001.77251.77440.1288
1000.31.77251.78980.1599
1000.61.77251.79400.2237
1000.81.77251.82260.3064
25001.77251.77630.0822
2500.31.77251.77690.1005
2500.61.77251.78410.1464
2500.81.77251.79220.2082
50001.77251.77640.0582
5000.31.77251.77650.0732
5000.61.77251.77800.1011
5000.81.77251.78960.1535
100001.77251.77340.0417
10000.31.77251.77330.0520
10000.61.77251.77220.0722
10000.81.77251.78130.1077
200001.77251.77200.0297
20000.31.77251.77230.0374
20000.61.77251.77280.0509
20000.81.77251.77670.0756
μ is the true mean of the target distribution; Mean ( μ ^ ) and SD ( μ ^ ) are computed over the B = 2000 Monte Carlo replications. μ ^ n does not depend on the bandwidth and is computed once per replicated sample.
Table 16. Empirical moments of the studentized statistic Z n ( t ) .
Table 16. Empirical moments of the studentized statistic Z n ( t ) .
n ϕ PointMeanVarianceSkewnessExcess Kurtosis
Gamma
1000 q . 25 −0.4420.607−0.5110.584
1000 q . 50 −0.2060.768−0.3780.361
1000 q . 75 0.0291.007−0.328−0.031
1000.3 q . 25 −0.4670.741−0.5351.228
1000.3 q . 50 −0.2230.878−0.5030.645
1000.3 q . 75 −0.0341.275−0.5620.972
1000.6 q . 25 −0.5301.235−0.8171.374
1000.6 q . 50 −0.2261.104−0.7261.499
1000.6 q . 75 −0.0652.111−0.8341.600
1000.8 q . 25 −0.7503.530−2.06713.212
1000.8 q . 50 −0.2432.075−0.8641.438
1000.8 q . 75 −0.0953.888−1.0641.899
2500 q . 25 −0.3770.661−0.3650.610
2500 q . 50 −0.1700.792−0.2830.201
2500 q . 75 0.0330.987−0.215−0.024
2500.3 q . 25 −0.3970.776−0.3190.170
2500.3 q . 50 −0.1520.850−0.3340.317
2500.3 q . 75 0.0051.167−0.2230.046
2500.6 q . 25 −0.4401.089−0.240−0.008
2500.6 q . 50 −0.1861.150−0.3500.318
2500.6 q . 75 −0.0191.713−0.3150.123
2500.8 q . 25 −0.5652.348−0.5470.554
2500.8 q . 50 −0.1971.953−0.4160.142
2500.8 q . 75 0.0013.058−0.6030.752
5000 q . 25 −0.3360.704−0.266−0.045
5000 q . 50 −0.1660.836−0.1940.086
5000 q . 75 0.0240.956−0.085−0.125
5000.3 q . 25 −0.3560.819−0.141−0.085
5000.3 q . 50 −0.1720.865−0.1940.030
5000.3 q . 75 0.0011.152−0.2260.044
5000.6 q . 25 −0.3731.028−0.2530.305
5000.6 q . 50 −0.1751.081−0.188−0.103
5000.6 q . 75 0.0081.636−0.221−0.000
5000.8 q . 25 −0.4542.084−0.4080.326
5000.8 q . 50 −0.2241.956−0.4210.739
5000.8 q . 75 −0.0093.015−0.5570.988
10000 q . 25 −0.2970.764−0.2860.255
10000 q . 50 −0.1590.877−0.0980.027
10000 q . 75 −0.0131.037−0.2170.077
10000.3 q . 25 −0.3190.856−0.1410.256
10000.3 q . 50 −0.1120.877−0.130−0.018
10000.3 q . 75 −0.0061.167−0.084−0.129
10000.6 q . 25 −0.2791.081−0.0780.070
10000.6 q . 50 −0.1561.103−0.1700.028
10000.6 q . 75 −0.0231.571−0.1210.091
10000.8 q . 25 −0.4041.876−0.2120.045
10000.8 q . 50 −0.1221.751−0.177−0.008
10000.8 q . 75 −0.0362.701−0.3500.084
20000 q . 25 −0.3120.772−0.105−0.064
20000 q . 50 −0.1090.847−0.0550.076
20000 q . 75 0.0210.983−0.0880.056
20000.3 q . 25 −0.2720.857−0.1220.135
20000.3 q . 50 −0.1350.926−0.1420.009
20000.3 q . 75 0.0361.121−0.1490.145
20000.6 q . 25 −0.2951.058−0.1430.075
20000.6 q . 50 −0.1231.031−0.1190.079
20000.6 q . 75 −0.0021.426−0.1330.022
20000.8 q . 25 −0.3521.861−0.1470.248
20000.8 q . 50 −0.2171.714−0.2000.106
20000.8 q . 75 0.0252.580−0.1940.131
Lognormal
1000 q . 25 −0.4330.564−0.3440.186
1000 q . 50 −0.2640.684−0.2910.104
1000 q . 75 −0.0020.839−0.2360.229
1000.3 q . 25 −0.5030.777−0.4530.352
1000.3 q . 50 −0.2690.715−0.3480.144
1000.3 q . 75 0.0281.008−0.4630.285
1000.6 q . 25 −0.5621.255−0.6691.049
1000.6 q . 50 −0.2720.952−0.5020.371
1000.6 q . 75 0.0071.488−0.4190.212
1000.8 q . 25 −0.7052.745−0.9221.919
1000.8 q . 50 −0.3061.732−0.6661.249
1000.8 q . 75 −0.0722.903−1.0332.643
2500 q . 25 −0.3900.643−0.2380.159
2500 q . 50 −0.1960.704−0.2520.327
2500 q . 75 −0.0000.870−0.2280.240
2500.3 q . 25 −0.4520.770−0.197−0.145
2500.3 q . 50 −0.1820.778−0.2070.139
2500.3 q . 75 0.0551.016−0.3280.349
2500.6 q . 25 −0.4671.070−0.3190.251
2500.6 q . 50 −0.2320.963−0.2250.272
2500.6 q . 75 0.0211.374−0.2890.263
2500.8 q . 25 −0.5632.137−0.6351.224
2500.8 q . 50 −0.2631.573−0.3230.092
2500.8 q . 75 −0.0162.415−0.4020.206
5000 q . 25 −0.3690.666−0.131−0.012
5000 q . 50 −0.1850.768−0.218−0.101
5000 q . 75 0.0030.902−0.1030.217
5000.3 q . 25 −0.3460.780−0.190−0.079
5000.3 q . 50 −0.1920.803−0.2590.245
5000.3 q . 75 −0.0171.003−0.1130.109
5000.6 q . 25 −0.4291.075−0.2490.304
5000.6 q . 50 −0.1950.968−0.1190.104
5000.6 q . 75 0.0411.301−0.1490.005
5000.8 q . 25 −0.4961.934−0.4020.183
5000.8 q . 50 −0.1761.502−0.3420.279
5000.8 q . 75 −0.0292.261−0.3570.154
10000 q . 25 −0.3540.726−0.0920.020
10000 q . 50 −0.1350.851−0.2000.201
10000 q . 75 0.0120.843−0.1720.074
10000.3 q . 25 −0.3180.838−0.1650.188
10000.3 q . 50 −0.1900.851−0.103−0.154
10000.3 q . 75 0.0101.056−0.1910.264
10000.6 q . 25 −0.3411.058−0.1690.114
10000.6 q . 50 −0.1371.013−0.1740.096
10000.6 q . 75 0.0321.323−0.1950.391
10000.8 q . 25 −0.3921.712−0.2560.278
10000.8 q . 50 −0.1931.457−0.1890.000
10000.8 q . 75 0.0092.091−0.2300.067
20000 q . 25 −0.3400.781−0.0970.103
20000 q . 50 −0.1160.839−0.0970.105
20000 q . 75 0.0670.938−0.0910.020
20000.3 q . 25 −0.2830.814−0.072−0.095
20000.3 q . 50 −0.1610.908−0.1170.037
20000.3 q . 75 0.0300.999−0.072−0.088
20000.6 q . 25 −0.3621.049−0.004−0.005
20000.6 q . 50 −0.1520.987−0.1560.094
20000.6 q . 75 0.0771.264−0.1630.132
20000.8 q . 25 −0.3691.690−0.1690.101
20000.8 q . 50 −0.1691.413−0.1980.045
20000.8 q . 75 0.0291.917−0.3080.201
Weibull
1000 q . 25 −0.3920.639−0.7081.674
1000 q . 50 −0.3270.852−0.7002.149
1000 q . 75 −0.1311.112−0.4800.880
1000.3 q . 25 −0.4310.758−0.5600.679
1000.3 q . 50 −0.3520.875−0.5350.472
1000.3 q . 75 −0.0931.382−0.4660.354
1000.6 q . 25 −0.4591.350−0.7821.818
1000.6 q . 50 −0.3681.349−0.8021.149
1000.6 q . 75 −0.2092.494−0.7681.250
1000.8 q . 25 −0.6752.974−1.0682.770
1000.8 q . 50 −0.3792.477−1.3554.540
1000.8 q . 75 −0.2194.896−1.1062.185
2500 q . 25 −0.3320.649−0.3930.616
2500 q . 50 −0.3100.846−0.2840.140
2500 q . 75 −0.0681.019−0.3730.505
2500.3 q . 25 −0.3610.767−0.2860.113
2500.3 q . 50 −0.2960.900−0.3540.340
2500.3 q . 75 −0.0941.260−0.2260.021
2500.6 q . 25 −0.4041.200−0.4290.347
2500.6 q . 50 −0.3251.282−0.7832.105
2500.6 q . 75 −0.1172.140−0.5210.491
2500.8 q . 25 −0.4692.472−0.6641.051
2500.8 q . 50 −0.3302.425−0.7871.364
2500.8 q . 75 −0.1864.362−0.8962.524
5000 q . 25 −0.3160.732−0.2670.102
5000 q . 50 −0.2380.879−0.126−0.011
5000 q . 75 −0.0460.983−0.0490.149
5000.3 q . 25 −0.3480.818−0.198−0.096
5000.3 q . 50 −0.2580.952−0.3640.537
5000.3 q . 75 −0.0431.192−0.2970.377
5000.6 q . 25 −0.3641.194−0.4320.805
5000.6 q . 50 −0.2711.246−0.4901.166
5000.6 q . 75 −0.0931.894−0.2940.662
5000.8 q . 25 −0.4222.429−0.6521.688
5000.8 q . 50 −0.2402.317−0.6991.328
5000.8 q . 75 −0.0724.146−0.5830.927
10000 q . 25 −0.2810.730−0.1830.150
10000 q . 50 −0.2300.858−0.1610.028
10000 q . 75 −0.0791.008−0.087−0.033
10000.3 q . 25 −0.2840.863−0.1270.053
10000.3 q . 50 −0.2670.915−0.1740.260
10000.3 q . 75 −0.0691.166−0.132−0.018
10000.6 q . 25 −0.3061.137−0.2090.013
10000.6 q . 50 −0.2901.252−0.162−0.028
10000.6 q . 75 −0.0881.718−0.159−0.125
10000.8 q . 25 −0.3722.017−0.2160.054
10000.8 q . 50 −0.2972.029−0.286−0.073
10000.8 q . 75 −0.0383.350−0.4140.397
20000 q . 25 −0.2530.759−0.105−0.076
20000 q . 50 −0.2430.902−0.0270.052
20000 q . 75 −0.0741.043−0.1840.062
20000.3 q . 25 −0.2380.886−0.0860.045
20000.3 q . 50 −0.2100.951−0.1940.011
20000.3 q . 75 −0.0691.156−0.0870.163
20000.6 q . 25 −0.2981.113−0.1440.031
20000.6 q . 50 −0.2601.139−0.135−0.059
20000.6 q . 75 −0.0521.664−0.101−0.063
20000.8 q . 25 −0.3391.908−0.1920.033
20000.8 q . 50 −0.2291.912−0.2750.181
20000.8 q . 75 −0.0792.732−0.2460.018
Computed under the undersmoothing bandwidth h US ( I ; n ) , B = 2000 . Skewness and excess kurtosis use the population (method-of-moments, divisor-B) convention, not the bias-corrected sample convention. Under the usual normal approximation, the standardized statistic should have mean 0, variance 1, skewness 0, and excess kurtosis 0.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Bouzebda, S.; Didi, S. Asymptotic Theory for Kernel Density Estimation Under Dependent Length-Biased Sampling. Symmetry 2026, 18, 1472. https://doi.org/10.3390/sym18091472

AMA Style

Bouzebda S, Didi S. Asymptotic Theory for Kernel Density Estimation Under Dependent Length-Biased Sampling. Symmetry. 2026; 18(9):1472. https://doi.org/10.3390/sym18091472

Chicago/Turabian Style

Bouzebda, Salim, and Sultana Didi. 2026. "Asymptotic Theory for Kernel Density Estimation Under Dependent Length-Biased Sampling" Symmetry 18, no. 9: 1472. https://doi.org/10.3390/sym18091472

APA Style

Bouzebda, S., & Didi, S. (2026). Asymptotic Theory for Kernel Density Estimation Under Dependent Length-Biased Sampling. Symmetry, 18(9), 1472. https://doi.org/10.3390/sym18091472

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop