We develop a design-conditional limit theory for kernel estimators of conditional
U-functionals based on locally stationary functional random fields observed at irregular random locations and under incomplete response observation. The covariates take values in a separable Hilbert space, the responses are allowed
[...] Read more.
We develop a design-conditional limit theory for kernel estimators of conditional
U-functionals based on locally stationary functional random fields observed at irregular random locations and under incomplete response observation. The covariates take values in a separable Hilbert space, the responses are allowed to take values in a general Polish space, and the target is indexed by a class of symmetric kernels of a fixed order. Functional localization is induced by single-index semi-metrics, while spatial localization is performed on the rescaled observation domain. Missing responses are incorporated through a complete-case construction under a Missing At Random condition and a uniform-positivity assumption. The resulting estimator is a ratio of spatially weighted
U-statistics with random tuplewise observation indicators. The asymptotic analysis must account simultaneously for four sources of complexity: dependence within the spatial field, nonstationarity across an expanding domain, concentration in an infinite-dimensional covariate space, and the random thinning generated by missing responses. Conditioning on the sampling locations removes the randomness of the spatial design weights but does not eliminate dependence among the observations. We therefore derive a design-conditional projection decomposition adapted to the triangular-array structure of the model. The leading component is represented by a spatially dependent complete-case empirical process, whereas the higher-order canonical terms are controlled uniformly over the response kernels, functional-target points, single-index directions, and rescaled spatial locations. The proofs combine stationary tangent-field approximations for locally stationary random fields, large-block–small-block decompositions, coupling arguments under spatial absolute regularity, small-ball probability estimates, and entropy bounds for the joint indexing class. These arguments yield a uniform stochastic expansion in which the empirical fluctuation, the spatial–functional smoothing bias, and the local-stationarity approximation error appear as distinct contributions. In particular, the local-stationarity remainder has no counterpart in the strictly stationary theory and quantifies the cost of replacing the observed nonstationary field with its stationary tangent approximation. Under the MAR and positivity conditions, complete-case sampling reduces the effective local information and modifies the covariance structure, but it does not change the formal order of the uniform-convergence rate. Under strengthened moment, mixing, entropy, and negligibility conditions, we establish weak convergence of the normalized conditional
U-process in the corresponding supremum-norm function space to a tight centered Gaussian process. The limiting covariance is determined by the complete-case first-order projection and consequently retains the effect of the observation propensity and the spatial dependence structure. We also introduce a complete-case leave-tuple-out spatial prediction criterion for bandwidth selection and prove oracle optimality over admissible bandwidth families. The general theory applies to conditional rank association, discrimination probabilities, set-indexed conditional distribution functionals, and related pairwise statistical-learning criteria. Simulation experiments and applications to spatial environmental and epidemiological data illustrate the finite-sample implications of the theory and the stabilizing role of single-index localization. Viewed through the lens of data-driven science, the framework addresses a fundamental asymmetry between the information carried by irregular, locally heterogeneous functional covariates and the selectively observed response tuples. By combining design conditioning, complete-case normalization, tangent-field localization, and single-index dimension reduction, the proposed approach resolves this inferential asymmetry at the level of the model by matching estimation and uncertainty quantification to the information actually available locally, without imposing artificial stationarity or complete-data symmetry.
Full article