1. Introduction
Set-valued forecasts are increasingly used when a single point prediction is not sufficiently reliable for an operational decision. A maintenance platform may need a set of plausible fault classes before dispatching a technician, while a financial surveillance system may need a set of possible risk states before escalating an alert. From a probability-theoretic viewpoint, these systems generate dependent stochastic processes observed through a history-dependent labeling mechanism. Conformal risk control supplies a model-agnostic route from scores to sets whose loss is controlled at a user-specified level [
1]. A complementary stochastic-process perspective studies updating when information is restricted [
2], while online selective conformal procedures control false coverage statements after data-dependent reporting [
3]. Yet most existing guarantees presume that the calibration labels represent the population on which risk will be evaluated.
Thresholded event systems violate this premise in a particularly severe way. Labels are typically collected after an alarm score crosses a prescribed level, whereas the much larger collection of sub-threshold periods remains unannotated. The resulting sample is not merely imbalanced; it is selected by a history-dependent mechanism that can be correlated with both the prediction score and the future event. Online conformal methods adapt to distribution shift [
4], and multivariate time-series sets account for structured forecast errors [
5], but neither contribution alone identifies the population risk hidden in the silent region. Stochastic-process diagnostics in
Axioms likewise illustrate that inferential conclusions depend on which trajectory functionals are observed [
6].
Dependence creates a second difficulty. Time-series events may be driven by latent regimes, persistent covariates, or feedback from past alerts. Non-exchangeable conformal methods quantify penalties caused by dependence [
7], while Markov-specific analyses connect coverage gaps with mixing times [
8]. Multi-horizon procedures exploit structured residuals [
9]. These developments are important, but they usually assume that all calibration responses, or an exchangeability-preserving subset, are observed. In the present problem, temporal dependence and missing labels interact: the same history that raises the alert probability also changes the conditional event risk.
A natural remedy is to supplement alert-triggered labels with an auxiliary unselected stream. Examples include a small random audit of silent periods, complete historical sensor logs available for an earlier machine fleet, or a retrospective database in which outcomes were recorded independently of the alert rule. Partially observed point-process models illustrate how an auxiliary observation phase can restore information about hidden events [
10]. However, the statistical role of such a stream in finite-sample predictive risk control has not been characterized. In particular, it is unclear which support condition is indispensable, how the two streams should be combined, and which guarantee survives arbitrary temporal dependence.
Recent time-series conformal methods provide complementary ingredients. Adaptive bands can track heterogeneous trajectories [
11]; decomposition can separate components with different exchangeability properties [
12]; and multi-step online calibration can regulate aggregate error [
13]. Stochastic models with mean reversion further show how persistent latent dynamics alter stationary behavior [
14]. These methods motivate history-sensitive weighting, but they do not address a selection probability that may be exactly zero below an operational threshold. That zero-probability region is the central identification obstacle studied here.
The proposed framework, called selective-observation weighted risk control (SOWRC), treats alert and audit labels as one predictable inclusion process on a filtered probability space. It estimates candidate losses by a Horvitz–Thompson transform and calibrates a nested predictive set with a simultaneous martingale-mixture upper boundary. The resulting probability guarantee is driven by conditional unbiasedness and supermartingale concentration rather than exchangeability, so the construction accommodates arbitrary temporal dependence within the observed trajectory. It connects with automatically adaptive risk control [
15], label-robust conformal inference [
16], and recent distribution-generalization theory under hidden confounding [
17]. Its sequential layer follows modern confidence-sequence and e-value principles [
18,
19]. Covariate-shift systems [
20] motivate explicit treatment of support, while recent probabilistic studies of hydrological risk [
21] and Gaussian-product extremes [
22] emphasize the importance of calibrated tail decisions.
Our first contribution is a probability-theoretic impossibility result: if a silent region has zero inclusion probability, observationally equivalent stochastic laws can have different population risks. Our second contribution is a finite-sample calibration theorem on a threshold grid fixed before labels are inspected. It separates the pointwise validity of any bounded loss from the additional monotonicity and safe-terminal conditions required for data-dependent nested-set selection. Our third contribution is a prospective theorem with explicit error accounting: future-horizon risk is controlled when an external deployment-drift certificate links the future population to the certified calibration population, and the certificate’s own failure probability is added to the calibration error. Further results cover multiple losses, adaptive budgets, estimated propensities, and oracle efficiency. The framework complements application-oriented risk control [
23] and transfer conformal inference [
24].
The paper also clarifies the scope of its guarantees. Covariate-shift procedures require source–target support overlap [
20]; beyond-exchangeability methods reweight complete calibration samples [
25]; and adaptive conformal prediction tracks errors as labels arrive [
26]. Recent theory explains when split conformal inference can remain effective under temporal dependence [
27], while post-training adaptive conformal prediction addresses structured missingness in incomplete time series [
28]. Distribution-free risk-controlling sets [
29] and modern conformal tutorials [
30] provide the broader baseline. SOWRC addresses the different obstacle of zero observation support. Its arbitrary-dependence theorem certifies the weighted calibration population; prospective use additionally requires a certified drift envelope and its own error budget, which cannot be removed without further temporal structure.
The identification layer is also connected to classical unequal-probability sampling and missing-data theory. Horvitz–Thompson weighting recovers population means when inclusion probabilities are positive and known [
31], and inverse-probability methods formalize the same positivity/ignorability requirements under missingness [
32]. The contribution here is not a new claim that inverse weighting needs positivity. Rather, we embed that requirement in a history-adapted time-series prediction-set problem, prove an explicit zero-support impossibility result for the predictive risk target, and combine the recovered losses with a finite-sample martingale certificate that remains valid without exchangeability or mixing-rate assumptions, subject to the stated predictability and ignorability conditions.
The remainder is organized as follows.
Section 2 formulates selective observation and proves non-identifiability without support.
Section 3 develops SOWRC and establishes the core finite-sample theorem.
Section 4 treats time-uniform control, multiple losses, and adaptive budgets.
Section 5 studies robustness and efficiency.
Section 6 gives numerical implementation and reproducibility details.
Section 7 reports synthetic and complete-log experiments.
Section 8 concludes and presents three future directions.
Table 1 summarizes the notation used throughout the paper.
6. Numerical Implementation and Reproducibility
This section collects the numerical steps that were previously dispersed across the theoretical development. The theoretical results above define the target and validity conditions; the present section explains how the certificate is evaluated, how the threshold is selected, and what must be logged for replication.
Figure 2 summarizes the information flow. The two label streams are merged only after their inclusion probabilities are computed. History weights then determine the target temporal population, and the e-process converts weighted losses into a finite-sample boundary. The output is a predictive set, not a corrected point forecast.
The architecture deliberately separates prediction, observation, and certification. The event model may be replaced without changing the risk layer. The selective stream supplies many labels near operationally important alerts, while the audit stream supplies support in silent periods. Their known propensities enter a single weighted loss. A sequential e-process then accounts for dependence through conditional expectations rather than an independence approximation. This modularity is useful in practice because model retraining, audit design, and risk-policy updates can proceed on different schedules without invalidating the logical role of each component.
Algorithm 1 summarizes the selective-observation weighted risk-control procedure used in the numerical implementation.
| Algorithm 1 Selective-observation weighted risk control |
- Require:
Fixed grid , target , deployment buffer (externally certified only when prospective validity is claimed), calibration confidence , weights , propensities - 1.
Set for - 2.
for
do - 3.
Compute and - 4.
Invert Equation ( 46) to obtain - 5.
end for - 6.
- 7.
if
then - 8.
- 9.
else - 10.
- 11.
end if - 12.
return
|
Figure 3 illustrates one calibration path. The observed weighted risk closely follows the risk computed from the hidden complete labels, while the upper boundary remains above both. The selected threshold is the first nested-set parameter for which the simultaneous boundary falls below the buffered certification budget.
The dark curve is available only to the simulator and acts as an oracle diagnostic. The inverse-inclusion estimate tracks it despite highly nonuniform label availability. The red martingale-mixture boundary adds a nonasymptotic correction that is largest for small sets, where losses and inverse-inclusion exposure are high. The green vertical line occurs where the simultaneous boundary crosses the buffered 0.09 certification budget. Its location is more conservative than the empirical crossing, which explains the moderate set-size inflation observed later. Importantly, calibration uses no hidden labels; the oracle curve is displayed solely for verification.
Remark 3. SOWRC is conformalized risk control in the sense that it calibrates a nested family generated by any scoring model, but its proof is not an exchangeability proof. The essential mechanism is predictable inverse-inclusion weighting plus a martingale boundary. This distinction permits arbitrary temporal dependence subject to the predictability, ignorability, positivity, bounded-loss, and design-envelope conditions stated above, yet it also makes the positivity constant explicit: sparse audits widen the bound through , exposing the price of identification rather than hiding it in an asymptotic approximation.
For all reported calculations, the mixture contains
equally weighted betting fractions. Let
denote the largest positive solution margin for which
. We use a geometric grid
for
. Thus all 32 betting fractions and their weights are determined explicitly by Equation (
106); the supplementary implementation should emit this 32-entry vector and use the same vector when reproducing the reported tables and figures. The factor
keeps every component strictly inside the admissible region. Equation (
47) is inverted by bisection on
to an absolute tolerance of
. The threshold grid is fixed before labels are inspected. These choices are numerical design decisions rather than additional stochastic assumptions.
Computational Complexity
For classification, sort the observed true-label scores once. With
, preprocessing costs
Cumulative weighted losses over
J thresholds can then be computed in
Maintaining
K losses and
L mixture components online requires
per labeled time point and memory
The calibration layer is therefore typically cheaper than fitting the underlying event model.
Remark 4. Efficiency is governed by two separable quantities. The predictive model controls how rapidly risk falls as the set expands, represented by the margin κ. The observation design controls the boundary through . Better scores cannot repair zero support, and more audits cannot repair an uninformative score. This separation suggests a practical workflow: first guarantee positivity through randomized audits, then improve set efficiency through modeling and targeted audit allocation.
The reproducibility configuration used in
Section 7 is: Python 3.13.5, NumPy 2.3.5, SciPy 1.17.0, scikit-learn 1.8.0, and statsmodels 0.14.6 for the complete-log real-data replay. For the 20-replication synthetic application studies, the stream seed is set equal to the replication index (0–19); randomized audit-mask replays use seeds 0–99. Alert and audit decisions are generated before the corresponding protected label, and the logged quantity used by SOWRC is the declared inclusion probability
, not a minimum estimated from the realized path.
8. Conclusions and Future Work
This paper studied predictive sets when labels are observed selectively along a dependent stochastic process. Population risk is non-identifiable on any region with zero labeling probability. A randomized audit stream restores support, and SOWRC combines alert and audit labels through predictable inverse-inclusion losses. A martingale-mixture boundary then provides finite-sample control of the weighted calibration population on a deterministic threshold grid. The prospective result states the additional condition required for future use: an externally certified deployment-drift envelope, explicit error allocation, and a matching calibration buffer.
The theory now separates four logically distinct components. Positivity determines identification. Effective sample size determines statistical precision. Monotone nested losses and a safe terminal action justify data-dependent threshold selection. A deployment envelope determines whether a calibration certificate transfers to future events. The experiments use a deterministic grid, remove latent-state leakage, and name baselines according to their implementations. They now also include a complete-log real-data replay with 100 audit masks and Monte Carlo uncertainty, a decomposition of the main sources of conservativeness, a propensity-misspecification sensitivity analysis, and selection-level validation over 4000 dependent-process replications.
Three directions merit further study. First, audit policies should be optimized jointly with prediction under a long-run labeling budget while preserving minimum support. Second, deployment-drift envelopes should be learned from independent auxiliary trajectories with simultaneous uncertainty quantification. Third, doubly robust and distributionally robust extensions should control delayed outcomes, partially observed covariates, and regime-specific risks without sacrificing anytime validity.