Previous Article in Journal
Qualitative Analysis of a Three-Species Ratio-Dependent Lotka–Volterra Model with Cautious Effect and Feedback Control
Previous Article in Special Issue
Local-Time Sensitivity and Burst Instability for Threshold Functionals of One-Dimensional Diffusions
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Risk-Controlling Predictive Sets for Time-Series Events Under Selective Observation with Finite-Sample Guarantees

1
School of Statistics and Data Science, Nankai University, Tianjin 300071, China
2
College of Electronic Engineering, National University of Defense Technology, Changsha 410073, China
3
Department of Automation, University of Science and Technology of China, Hefei 230026, China
*
Authors to whom correspondence should be addressed.
Axioms 2026, 15(9), 706; https://doi.org/10.3390/axioms15090706 (registering DOI)
Submission received: 9 August 2026 / Revised: 18 September 2026 / Accepted: 20 September 2026 / Published: 21 September 2026
(This article belongs to the Special Issue Probability Theory and Stochastic Processes: Theory and Applications)

Abstract

Selective labels create a support failure for prediction along dependent stochastic processes: alert-triggered events are observed, whereas silent periods are usually unlabeled. We model this mechanism as predictable inclusion on a filtered probability space and show that population risk is non-identifiable when any silent region has zero labeling probability. Selective-observation weighted risk control (SOWRC) combines alert labels with randomized audits through Horvitz–Thompson losses and a martingale-mixture boundary. It provides finite-sample calibration-population control under arbitrary temporal dependence subject to predictable design choices, conditional ignorability, positivity, bounded losses, and deterministic design envelopes, together with a prospective guarantee under an externally certified deployment-drift envelope and explicit error allocation. Extensions cover anytime monitoring, multiple losses, adaptive budgets, and estimated propensities. Synthetic maintenance and financial studies, a complete-log replay on a real dependent sensor series with 100 audit-mask replications, and 4000 selection-level validation runs demonstrate support recovery and conservative probabilistic risk control on deterministic threshold grids.

1. Introduction

Set-valued forecasts are increasingly used when a single point prediction is not sufficiently reliable for an operational decision. A maintenance platform may need a set of plausible fault classes before dispatching a technician, while a financial surveillance system may need a set of possible risk states before escalating an alert. From a probability-theoretic viewpoint, these systems generate dependent stochastic processes observed through a history-dependent labeling mechanism. Conformal risk control supplies a model-agnostic route from scores to sets whose loss is controlled at a user-specified level [1]. A complementary stochastic-process perspective studies updating when information is restricted [2], while online selective conformal procedures control false coverage statements after data-dependent reporting [3]. Yet most existing guarantees presume that the calibration labels represent the population on which risk will be evaluated.
Thresholded event systems violate this premise in a particularly severe way. Labels are typically collected after an alarm score crosses a prescribed level, whereas the much larger collection of sub-threshold periods remains unannotated. The resulting sample is not merely imbalanced; it is selected by a history-dependent mechanism that can be correlated with both the prediction score and the future event. Online conformal methods adapt to distribution shift [4], and multivariate time-series sets account for structured forecast errors [5], but neither contribution alone identifies the population risk hidden in the silent region. Stochastic-process diagnostics in Axioms likewise illustrate that inferential conclusions depend on which trajectory functionals are observed [6].
Dependence creates a second difficulty. Time-series events may be driven by latent regimes, persistent covariates, or feedback from past alerts. Non-exchangeable conformal methods quantify penalties caused by dependence [7], while Markov-specific analyses connect coverage gaps with mixing times [8]. Multi-horizon procedures exploit structured residuals [9]. These developments are important, but they usually assume that all calibration responses, or an exchangeability-preserving subset, are observed. In the present problem, temporal dependence and missing labels interact: the same history that raises the alert probability also changes the conditional event risk.
A natural remedy is to supplement alert-triggered labels with an auxiliary unselected stream. Examples include a small random audit of silent periods, complete historical sensor logs available for an earlier machine fleet, or a retrospective database in which outcomes were recorded independently of the alert rule. Partially observed point-process models illustrate how an auxiliary observation phase can restore information about hidden events [10]. However, the statistical role of such a stream in finite-sample predictive risk control has not been characterized. In particular, it is unclear which support condition is indispensable, how the two streams should be combined, and which guarantee survives arbitrary temporal dependence.
Recent time-series conformal methods provide complementary ingredients. Adaptive bands can track heterogeneous trajectories [11]; decomposition can separate components with different exchangeability properties [12]; and multi-step online calibration can regulate aggregate error [13]. Stochastic models with mean reversion further show how persistent latent dynamics alter stationary behavior [14]. These methods motivate history-sensitive weighting, but they do not address a selection probability that may be exactly zero below an operational threshold. That zero-probability region is the central identification obstacle studied here.
The proposed framework, called selective-observation weighted risk control (SOWRC), treats alert and audit labels as one predictable inclusion process on a filtered probability space. It estimates candidate losses by a Horvitz–Thompson transform and calibrates a nested predictive set with a simultaneous martingale-mixture upper boundary. The resulting probability guarantee is driven by conditional unbiasedness and supermartingale concentration rather than exchangeability, so the construction accommodates arbitrary temporal dependence within the observed trajectory. It connects with automatically adaptive risk control [15], label-robust conformal inference [16], and recent distribution-generalization theory under hidden confounding [17]. Its sequential layer follows modern confidence-sequence and e-value principles [18,19]. Covariate-shift systems [20] motivate explicit treatment of support, while recent probabilistic studies of hydrological risk [21] and Gaussian-product extremes [22] emphasize the importance of calibrated tail decisions.
Our first contribution is a probability-theoretic impossibility result: if a silent region has zero inclusion probability, observationally equivalent stochastic laws can have different population risks. Our second contribution is a finite-sample calibration theorem on a threshold grid fixed before labels are inspected. It separates the pointwise validity of any bounded loss from the additional monotonicity and safe-terminal conditions required for data-dependent nested-set selection. Our third contribution is a prospective theorem with explicit error accounting: future-horizon risk is controlled when an external deployment-drift certificate links the future population to the certified calibration population, and the certificate’s own failure probability is added to the calibration error. Further results cover multiple losses, adaptive budgets, estimated propensities, and oracle efficiency. The framework complements application-oriented risk control [23] and transfer conformal inference [24].
The paper also clarifies the scope of its guarantees. Covariate-shift procedures require source–target support overlap [20]; beyond-exchangeability methods reweight complete calibration samples [25]; and adaptive conformal prediction tracks errors as labels arrive [26]. Recent theory explains when split conformal inference can remain effective under temporal dependence [27], while post-training adaptive conformal prediction addresses structured missingness in incomplete time series [28]. Distribution-free risk-controlling sets [29] and modern conformal tutorials [30] provide the broader baseline. SOWRC addresses the different obstacle of zero observation support. Its arbitrary-dependence theorem certifies the weighted calibration population; prospective use additionally requires a certified drift envelope and its own error budget, which cannot be removed without further temporal structure.
The identification layer is also connected to classical unequal-probability sampling and missing-data theory. Horvitz–Thompson weighting recovers population means when inclusion probabilities are positive and known [31], and inverse-probability methods formalize the same positivity/ignorability requirements under missingness [32]. The contribution here is not a new claim that inverse weighting needs positivity. Rather, we embed that requirement in a history-adapted time-series prediction-set problem, prove an explicit zero-support impossibility result for the predictive risk target, and combine the recovered losses with a finite-sample martingale certificate that remains valid without exchangeability or mixing-rate assumptions, subject to the stated predictability and ignorability conditions.
The remainder is organized as follows. Section 2 formulates selective observation and proves non-identifiability without support. Section 3 develops SOWRC and establishes the core finite-sample theorem. Section 4 treats time-uniform control, multiple losses, and adaptive budgets. Section 5 studies robustness and efficiency. Section 6 gives numerical implementation and reproducibility details. Section 7 reports synthetic and complete-log experiments. Section 8 concludes and presents three future directions. Table 1 summarizes the notation used throughout the paper.

2. Problem Formulation and Identification

2.1. Time-Series Prediction Under Selective Labels

Let ( Ω , A , P ) carry a filtered process { ( X t , Y t ) } t 1 , where X t X and Y t Y = { 1 , , m } . The post-label filtration and pre-label history are
F t = σ F t 1 , X t , S t , A t , I t , I t Y t ,
H t = σ ( F t 1 , X t ) .
A black-box predictor produces H t -measurable class probabilities p t ( y X t ) and a nonconformity score
r t ( x , y ) = 1 p t ( y x ) .
For a threshold λ [ 0 , 1 ] , define the nested set
C t , λ ( x ) = { y Y : r t ( x , y ) λ } .
Nestedness gives
λ 1 λ 2 C t , λ 1 ( x ) C t , λ 2 ( x ) .
The principal loss is event miscoverage,
L t ( λ ) = 1 { Y t C t , λ ( X t ) } = 1 { r t ( X t , Y t ) > λ } .
The pointwise concentration analysis also permits any H t σ ( Y t ) -measurable loss satisfying
0 L t ( k ) ( λ ) 1 , k = 1 , , K .
Data-dependent selection over a nested family additionally requires monotonicity and a certified safe terminal action, stated in Assumption 3.
The operational alert rule is represented by S t { 0 , 1 } . Its conditional probability may depend on the full observed history,
π t = P ( S t = 1 H t ) .
An auxiliary audit independently requests a label with conditional probability
ρ t = P ( A t = 1 H t ) .
A label is observed when either channel fires:
I t = 1 ( 1 S t ) ( 1 A t ) .
Under conditional independence of S t and A t given H t , the inclusion propensity is
q t = P ( I t = 1 H t ) = π t + ρ t π t ρ t .
The target is not alert-conditional risk. It is the weighted population risk over all periods,
R ¯ n ( λ ) = 1 A n t = 1 n a t R t ( λ ) , A n = t = 1 n a t ,
where a t > 0 is H t -measurable and
R t ( λ ) = E [ L t ( λ ) H t ] .
For a fixed terminal horizon, exponential forgetting is obtained with
a t , n ( γ ) = exp { γ ( n t ) } , γ 0 .
The second index emphasizes that these weights are redefined for a fixed-horizon target; the anytime result in Section 4.1 instead uses one predictable weight sequence satisfying global design bounds.
Four risk objects are kept distinct throughout the revision. The one-step conditional risk R t ( λ ) is a history-specific conditional expectation. The calibration-population risk R ¯ n ( λ ) is the weighted finite-horizon target certified by Theorem 2. The complete-label test risk used in Section 7 is an empirical diagnostic on a held-out, fully revealed segment and is not itself a theorem target. Finally, the operational or deployment risk R ¯ n , H dep ( λ ) concerns future periods and is covered only after the additional drift condition in Assumption 4. Table 2 summarizes these four risk quantities and their logical status.

2.2. Assumptions and an Impossibility Boundary

Assumption 1
(Conditional ignorability). For every t, inclusion is conditionally independent of the label given the pre-label history:
I t Y t H t .
Equivalently,
P ( I t = 1 H t , Y t ) = q t .
Assumption 2
(Audit positivity). There exists q min > 0 such that
q t q min a l m o s t s u r e l y f o r a l l t .
Assumption 1 allows arbitrary temporal dependence and selection based on observed history. It excludes label-dependent annotation after conditioning on that history. Assumption 2 is supplied by the audit stream; for a deterministic threshold alert, it is enough to choose ρ t ρ min > 0 .
The phrase “arbitrary temporal dependence” refers only to the absence of an exchangeability, independence, or mixing-rate requirement for the observed trajectory. It does not remove the structural conditions used by the proof: the alert/audit decision and any history weights must be predictable, inclusion must satisfy conditional ignorability, losses must be bounded, and a positive inclusion floor is required. Conditional ignorability is not nonparametrically testable from selectively observed labels alone. In practice, the most defensible design is to generate the audit request by external randomization before the current outcome is known. If an alerting or annotation system uses latent same-period information that is not recorded in H t , the assumption can fail and the theorem should not be invoked without enlarging the recorded history or performing a sensitivity analysis. Table 3 summarizes the role, guarantee, and assessability of the core assumptions and design conditions.
For an arbitrary predictable set-valued rule C 1 : n , write its weighted population miscoverage as
R n ( C 1 : n ) = 1 A n s = 1 n a s E [ 1 { Y s C s } H s ] .
Theorem 1
(Non-identifiability without silent-region support). Fix a horizon n and suppose that, for some t n , an event G H t satisfies P ( G ) > 0 and q t = 0 on G. Let C t be the set reported before observing Y t . If 
P { G , C t Y } > 0 ,
there exist two time-series laws P 0 and P 1 with the same distribution of the complete observed trajectory through time n but with R n , P 0 ( C 1 : n ) R n , P 1 ( C 1 : n ) . If instead P { G , C t = } > 0 , the rule incurs unit conditional miscoverage on that event under every compatible label law. Consequently, uniform control at arbitrarily small levels forces the full label space on every zero-inclusion region.
Proof. 
Because Y is finite, Equation (19) implies that there is a label y Y for which
p = P { G , C t , y C t } > 0 .
Let U denote the procedure’s independent random seed and augment the pre-label history by σ ( U ) ; equivalently, condition on U and average at the end. Construct P 0 and P 1 on a common covariate, selection, and randomization process. Outside G, let their conditional label laws coincide. On G, set Y t = y under P 1 . Under  P 0 , choose a measurable selector y 0 ( C t ) C t whenever C t is nonempty; on C t = , use the same fixed label under both laws. The output is measurable before Y t is generated, and the label is never revealed on G. The construction is existential: the model class permits arbitrary temporal dependence, but it does not require every admissible law to make future observables depend on the hidden Y t . We therefore choose, under both P 0 and P 1 , the same post-t transition kernel for the observable coordinates conditional on the recorded history, with that kernel not depending on the unrevealed value of Y t on G. This defines two legitimate members of the stated model class. The argument does not claim that one may alter Y t while holding fixed an arbitrary pre-specified data-generating mechanism in which Y t causally changes later covariates.
Since q t = 0 on G, the observed variable I t Y t is absent there. Under the paired laws just constructed, induction over the common observable transition kernels therefore gives, for every measurable function h of the observed trajectory,
E P 0 [ h ( F n ) ] = E P 1 [ h ( F n ) ] .
On the event in Equation (20), the time-t miscoverage under P 1 equals one and that under P 0 equals zero. Hence,
E P 1 [ 1 { Y t C t } 1 G ] E P 0 [ 1 { Y t C t } 1 G ] p > 0 .
All other time contributions can be chosen equal, so positive weighting yields
R n , P 1 ( C 1 : n ) R n , P 0 ( C 1 : n ) a t p A n > 0 .
Thus the observed law cannot determine the target risk. An empty output on a positive-probability subset of G has miscoverage of one regardless of the hidden label. Uniform control at every arbitrarily small level therefore rules out both empty and nonempty proper outputs, forcing C t = Y almost surely on every zero-inclusion region.    □
Figure 1 visualizes the support issue. Alert labels concentrate above a moving threshold, while the audit channel places sparse labels throughout the silent region. Those audited points are not intended to make the sample balanced. Their role is more fundamental: they make every history stratum observable with positive probability, which turns an unidentified target into an estimable population risk.
Figure 1 was redrawn with a more reactive, lagged-score threshold so that history dependence is visually apparent rather than nearly indistinguishable from a constant cutoff. Orange circles arise when the score exceeds the predictable adaptive threshold, whereas green crosses are randomized audits in silent periods. The threshold is computed only from information available before the current label, so its visible movement illustrates allowed adaptive selection without violating predictability. The purpose of the audits is to support recovery, not class balancing.
Remark 1.
The impossibility theorem is not a statement that selective data are unusable. It isolates the exact failure mode: an unobserved region with zero inclusion probability. A small randomized audit is therefore qualitatively different from merely collecting more alert labels. Additional selected labels reduce variance inside the alert region, whereas audits restore identification across the population support and make any finite-sample population guarantee logically possible.

3. Selective-Observation Weighted Risk Control

3.1. Unbiased Loss Recovery

Define the inverse-inclusion loss
Z t ( λ ) = I t q t L t ( λ ) .
Conditional ignorability gives the identity
E [ Z t ( λ ) H t , Y t ] = L t ( λ ) ,
and therefore
E [ Z t ( λ ) H t ] = R t ( λ ) .
The empirical weighted risk is
R ^ n ( λ ) = 1 A n t = 1 n a t Z t ( λ ) .
Its martingale difference is
D t ( λ ) = a t { Z t ( λ ) R t ( λ ) } .
By Assumption 2,
| D t ( λ ) | a t q min .
A predictable conditional variance proxy is
v t ( λ ) = E [ D t 2 ( λ ) H t ] a t 2 E L t ( λ ) q t | H t a t 2 q min .
Let
V n ( λ ) = t = 1 n v t ( λ ) .
When a variance-sensitive closed-form boundary is used, fix, before the labels are inspected, a deterministic envelope V n ( λ ) satisfying
V n ( λ ) V n ( λ ) almost surely .
For the fixed calibration horizon, choose deterministic design envelopes c n , b n > 0 such that
a t c n , a t q t b n almost surely for t n .
For example, if  a t a max , n by design, then c n = a max , n and b n = a max , n / q min are valid. Using deterministic envelopes, rather than maxima depending on future random histories, is essential for the sequential conditional-expectation argument below.
Lemma 1
(Conditional exponential bound). For 0 η < 3 / b n ,
E exp η D t ( λ ) η 2 v t ( λ ) 2 ( 1 η b n / 3 ) | H t 1 .
Proof. 
Condition on H t . By the definition of D t ( λ ) , Equation (29), and the deterministic design envelope a t / q min b n in Equation (33), we have | D t ( λ ) | b n . Expanding the exponential and using k ! 2 · 3 k 2 for k 2 gives
E [ e η D t H t ] 1 + η 2 v t 2 j = 0 ( η b n / 3 ) j .
The geometric series equals ( 1 η b n / 3 ) 1 . Since 1 + x e x ,
E [ e η D t H t ] exp η 2 v t 2 ( 1 η b n / 3 ) .
Rearranging proves the stated conditional inequality.    □
For a fixed threshold, Lemma 1 and the deterministic envelope in Equation (32) give
B n ( δ , λ ) = 2 V n ( λ ) log ( 1 / δ ) A n + b n log ( 1 / δ ) 3 A n .
If no refined deterministic variance envelope is available, Equation (30) gives the safe choice V n ( λ ) = t = 1 n a t 2 / q min and hence
B ¯ n ( δ ) = 2 log ( 1 / δ ) t = 1 n a t 2 / q min A n 2 + c n log ( 1 / δ ) 3 q min A n .
A plug-in expression based on the unknown realized V n ( λ ) is not used, because optimizing a fixed-time exponential bound with a random unaccounted variance term would require an additional self-normalization argument. The deterministic safe bound is nonasymptotic but can be conservative. We therefore derive an observable risk-adaptive boundary that avoids estimating V n .

3.2. Observable Martingale-Mixture Boundary

Write the cumulative target and its Horvitz–Thompson estimate as
μ n ( λ ) = A n R ¯ n ( λ ) , μ ^ n ( λ ) = A n R ^ n ( λ ) .
For the lower-deviation increment
H t ( λ ) = a t { R t ( λ ) Z t ( λ ) } = D t ( λ ) ,
we have
H t ( λ ) a t c n , E [ H t 2 ( λ ) H t ] a t 2 R t ( λ ) q t b n a t R t ( λ ) .
For 0 < η < 3 / c n , define
g c n ( η ) = η 2 2 ( 1 η c n / 3 ) , κ n ( η ) = η b n g c n ( η ) .
Whenever κ n ( η ) > 0 , keep the terminal-horizon envelopes c n , b n fixed and, for 0 k n , define the triangular process
E k n ( λ , η ) = exp κ n ( η ) t = 1 k a t R t ( λ ) η t = 1 k a t Z t ( λ ) .
Conditional Bernstein control makes { E k n : 0 k n } a nonnegative supermartingale. Its terminal value is exp { κ n ( η ) μ n η μ ^ n } . The one-step inequality is
E exp η H t ( λ ) g c n ( η ) b n a t R t ( λ ) | H t 1 .
Thus the unknown variance is replaced by the unknown risk itself, which can be eliminated by test inversion.
Let η 1 , , η L be fixed admissible betting fractions and w > 0 satisfy
= 1 L w = 1 .
For a candidate risk value r [ 0 , 1 ] , define the observable mixture
E n ( λ , r ) = = 1 L w exp κ n ( η ) A n r η A n R ^ n ( λ ) .
At r = R ¯ n ( λ ) , Equation (46) is a mixture of the terminal values of the supermartingales in Equation (43). Its inversion gives
U n ( λ ; δ ) = min 1 , inf r [ 0 , 1 ] : E n ( λ , r ) 1 δ ,
where the infimum of an empty set is one. The map r E n ( λ , r ) is increasing, so Equation (47) is computed by one-dimensional bisection.

3.3. Calibration over Nested Sets

Let Λ = { λ 1 < < λ J = 1 } be fixed independently of all calibration labels. For simultaneous reporting, choose δ j > 0 with
j = 1 J δ j δ .
The experiments use the deterministic grid λ j = ( j 1 ) / ( J 1 ) with J = 241 and δ j = δ / J .
Assumption 3
(Nested safety). For the selected operational loss, L t ( λ ) is nonincreasing in λ and the terminal threshold is safe:
L t ( λ J ) = 0 almost surely for every t .
For miscoverage, λ J = 1 returns the full label space and satisfies Equation (49).
Define
λ ^ n = min { λ j Λ : U n ( λ j ; δ j ) α } ,
where λ ^ n = λ J if the feasible set is empty. The reported rule is
C ^ n ( x ) = C n + 1 , λ ^ n ( x ) .
Theorem 2
(Finite-sample calibration-population control). Under Assumptions 1 and 2, for any temporally dependent process and predictable positive weights satisfying Equation (33),
P R ¯ n ( λ ) > U n ( λ ; δ ) δ
for every fixed λ and every bounded loss in Equation (7). For the nested selection rule under Assumption 3,
P λ j Λ : R ¯ n ( λ j ) U n ( λ j ; δ j ) 1 δ ,
and
P R ¯ n ( λ ^ n ) α 1 δ .
Proof. 
For fixed λ , Equation (44) and repeated conditioning show that each process in Equation (43) has expectation at most one. Their weighted sum also has expectation at most one. If  R ¯ n ( λ ) > U n ( λ ; δ ) , monotonicity in r implies
E n λ , R ¯ n ( λ ) 1 / δ .
Markov’s inequality proves Equation (52). Applying it at levels δ j and using a union bound gives Equation (53). On that event, every feasible grid point is valid. If no point is feasible, Assumption 3 makes the terminal fallback valid. Hence, Equation (54) follows.    □

3.4. Prospective Risk Under Bounded Deployment Drift

Theorem 2 controls the weighted calibration population and, by itself, cannot identify an unrestricted future law. For a deployment horizon H, let v s > 0 , V H = s = 1 H v s , and define
R ¯ n , H dep ( λ ) = 1 V H s = 1 H v s R n + s ( λ ) .
Assumption 4
(Certified deployment-drift envelope). Before deployment, an envelope Δ ^ n , H [ 0 , α ) and an error level δ dep [ 0 , 1 ) are fixed using information available before the protected outcomes. The external certification mechanism guarantees
P ( D n , H ) 1 δ dep ,
where
D n , H = λ Λ : R ¯ n , H dep ( λ ) R ¯ n ( λ ) + Δ ^ n , H .
A deterministic process bound is the special case δ dep = 0 . If an auxiliary log or validation stream is used, its uncertainty must be included in δ dep ; no independence from the calibration sample is required.
One concrete external construction is available when two complete-label logs are kept outside the SOWRC calibration sample: a reference log of m 0 periods intended to represent the calibration regime and a pre-deployment log of m 1 periods intended to represent the deployment regime. Let R 0 pop ( λ j ) and R 1 pop ( λ j ) denote the population risks represented by these two logs. To connect them to the theorem targets, assume certified coupling/representativeness radii γ 0 , γ 1 0 such that, uniformly on the fixed grid,
R 0 pop ( λ j ) R ¯ n ( λ j ) γ 0 , R ¯ n , H dep ( λ j ) R 1 pop ( λ j ) + γ 1 .
Thus, the reference log does not automatically estimate R ¯ n ; that connection is an additional representativeness requirement. On the fixed grid, let R ^ 0 ( λ j ) and R ^ 1 ( λ j ) be their empirical losses. If the entries within each log are independent and bounded in [ 0 , 1 ] , simultaneous Hoeffding radii
ϵ i = log ( 4 J / δ dep ) 2 m i , i { 0 , 1 } ,
lead to the conservative candidate envelope
Δ ^ n , H cand = min 1 , max j R ^ 1 ( λ j ) R ^ 0 ( λ j ) + + ϵ 0 + ϵ 1 + γ 0 + γ 1 .
A union bound gives the required calibration-to-deployment comparison when the stated coupling conditions hold. The displayed Hoeffding radii rely on the within-log independence assumption; if either log is temporally dependent, they must be replaced by a valid confidence sequence or process-specific concentration bound. The candidate is accepted as Δ ^ n , H only when Δ ^ n , H cand < α . If it is at least α , Assumption 4 is not certified for any nonterminal action, the prospective feasible set is declared empty, and the safe terminal fallback is used. This example is included to show how the abstract certificate can be produced; the numerical 0.01 buffer used later is deliberately not presented as such a certificate.
Use the buffered rule
λ ^ n , H dep = min { λ j Λ : U n ( λ j ; δ j ) α Δ ^ n , H } ,
with the safe terminal fallback. For a data-derived construction such as Equation (60), this rule is invoked only after the candidate envelope has passed the strict check Δ ^ n , H cand < α ; otherwise the safe terminal action is taken directly.
Corollary 1
(Finite-sample prospective control). Under the assumptions of Theorem 2 and Assumption 4, provided δ + δ dep < 1 ,
P R ¯ n , H dep ( λ ^ n , H dep ) α 1 δ δ dep .
Proof. 
On the intersection of the simultaneous calibration event in Equation (53) and D n , H , the buffered rule has calibration risk at most α Δ ^ n , H ; if its feasible set is empty, the safe terminal action has zero risk. Equation (58) then gives deployment risk at most α . A union bound shows that this intersection has probability at least 1 δ δ dep .    □

4. Temporal Dependence, Multiple Risks, and Adaptive Budgets

4.1. Time-Uniform Control

A fixed-horizon guarantee is insufficient when a system may recalibrate after any number of observations. In this subsection, { a t } t 1 is a single predictable sequence, not the horizon-reindexed family in Equation (14). Horizon-dependent forgetting may be handled by pre-specified restarts with separate confidence allocation. Assume known design constants
0 < a t c , a t q t b for all t ,
which follow from bounded normalized weights and positivity. Define g c ( η ) and κ ( η ) = η b g c ( η ) as in Equation (42), using the global constants. For the same admissible fractions and mixture weights as before,
E n ( λ ) = = 1 L w exp κ ( η ) A n R ¯ n ( λ ) η A n R ^ n ( λ )
is a nonnegative supermartingale. Ville’s inequality therefore gives
P sup n 1 E n ( λ ) 1 / δ δ .
Let U n ( λ ; δ ) denote Equation (47) evaluated with these global design bounds.
Theorem 3
(Anytime-valid calibration). For every fixed λ,
P n 1 : R ¯ n ( λ ) U n ( λ ; δ ) 1 δ .
For a finite grid with confidence allocations δ j ,
P n 1 , j : R ¯ n ( λ j ) U n ( λ j ; δ j ) 1 δ .
Proof. 
Equation (44) remains valid after replacing c n , b n by their global upper bounds. Repeated conditioning makes every component of Equation (64) a nonnegative supermartingale, and their weighted sum has initial value one. If the true risk exceeds the inverted bound at any time, the increasing mixture at that true risk crosses 1 / δ . Equation (65) bounds the probability of any crossing, which proves Equation (66). Applying the result at levels δ j and using j δ j δ proves Equation (67).    □

4.2. Simultaneous Control of Several Operational Losses

Suppose K losses encode miscoverage, dangerous false negatives, excessive set size, or group-specific errors. For loss k, define
R ¯ n ( k ) ( λ ) = 1 A n ( k ) t = 1 n a t ( k ) R t ( k ) ( λ ) ,
R ^ n ( k ) ( λ ) = 1 A n ( k ) t = 1 n a t ( k ) I t L t ( k ) ( λ ) q t .
Let U n ( k ) be the corresponding upper bound and choose confidence shares satisfying
k = 1 K j = 1 J δ k j δ .
The joint feasible set is
Λ n safe = { λ Λ : U n ( k ) ( λ ) α k , k = 1 , , K } .
Whenever Λ n safe is nonempty, an efficiency criterion g n ( λ ) selects
λ ^ n = arg min λ Λ n safe g n ( λ ) .
If the feasible set is empty, the system abstains and requests additional auditing rather than issuing an uncertified set.
Corollary 2
(Simultaneous multi-risk validity). Whenever Λ n safe is nonempty, with probability at least 1 δ ,
R ¯ n ( k ) ( λ ^ n ) α k , k = 1 , , K .
Proof. 
Apply Theorem 2 to every pair ( k , j ) with level δ k j . Equation (70) implies simultaneous validity of all bounds. On this event, every member of Λ n safe satisfies every population constraint, so the data-dependent minimizer in Equation (72) does as well.    □

4.3. Predictable Time-Varying Risk Budgets

Let α t ( 0 , 1 ) be H t -measurable. It may tighten during a high-cost regime and relax when review capacity is limited. For a predictable threshold sequence λ t H t , define
μ n ad = t = 1 n a t R t ( λ t ) , μ ^ n ad = t = 1 n a t Z t ( λ t ) ,
and the cumulative risk budget
B n = t = 1 n a t α t .
For any observed cumulative value m, let
Γ n ( m ; δ ) = min A n , inf u [ 0 , A n ] : = 1 L w exp κ ( η ) u η m 1 δ .
Thus, Γ n / A n is the anytime upper risk obtained from the adaptively indexed losses. For ordered set selection, assume the predictable envelope is nonincreasing in λ and that a designated terminal action has zero envelope; miscoverage with the full label set is the canonical example. A conservative pre-label projection uses an H t -measurable envelope satisfying L t ( λ ) t max ( λ ) 1 almost surely and
m t + ( λ ) = μ ^ t 1 ad + a t t max ( λ ) q t .
The budget-aware rule is
λ t = min λ Λ : Γ t m t + ( λ ) ; δ B t .
If no candidate, including the designated terminal action, is feasible, the system abstains and requests additional auditing; it does not claim budget validity for an uncertified action.
Theorem 4
(Adaptive-budget risk accounting). For any predictable sequences { α t , a t , λ t } satisfying the global bounds in Equation (63),
P n 1 : μ n ad Γ n ( μ ^ n ad ; δ ) 1 δ .
If the feasible set in Equation (78) is nonempty at every evaluated time and Z t ( λ ) t max ( λ ) / q t , then, with probability of at least 1 δ ,
t = 1 n a t R t ( λ t ) t = 1 n a t α t for all n .
Proof. 
Predictability of λ t preserves
E [ Z t ( λ t ) R t ( λ t ) H t ] = 0 .
The lower-deviation process therefore obeys the same conditional Bernstein inequality as Equation (44). Mixing and applying Ville’s inequality proves Equation (79). The map m Γ n ( m ; δ ) is nondecreasing. Since the realized update satisfies μ ^ t ad m t + ( λ t ) , Equation (78) implies Γ t ( μ ^ t ad ; δ ) B t . Combining this deterministic implication with Equation (79) yields Equation (80).    □
Remark 2.
Adaptive operation does not require exchangeable blocks or a fixed stopping time. What must remain predictable is the choice made before the current label is revealed: the audit probability, history weight, risk budget, and threshold. This ordering is operationally natural. A system may react to observed covariates and past errors, but it cannot use the current hidden outcome to choose the set that will be evaluated against that outcome.

5. Statistical Efficiency and Robustness

5.1. Effective Sample Size and Audit Design

The variance cost of selective observation can be summarized by
n eff = A n 2 t = 1 n a t 2 / q t .
For equal weights and constant inclusion probability q,
n eff = n q .
This quantity is more than a heuristic summary. From Equation (30),
V n ( λ ) A n t = 1 n a t 2 / q t A n = 1 n eff ,
so the leading concentration term contracts at the familiar n eff 1 / 2 rate. With equal weights and a deterministic alert, a crude worst-case design calculation replaces q t by the silent-region audit floor ρ min , giving a radius of order log ( J / δ ) / ( n ρ min ) plus a smaller log ( J / δ ) / ( n ρ min ) term. Thus, the minimum audit rate can be chosen by first specifying an acceptable certification radius or a maximum inverse weight and then solving for the required ρ min , subject to the labeling budget.
For a deterministic alert and an audit probability ρ t , Equation (11) becomes
q t = S t + ( 1 S t ) ρ t .
If auditing has per-period cost c t and a total budget B, a variance-oriented design solves
min ρ 1 : n t = 1 n a t 2 σ t 2 S t + ( 1 S t ) ρ t s . t . t = 1 n c t ρ t B ,
where σ t 2 is a predictable loss-variance proxy. Ignoring box constraints, the KKT conditions give
ρ t a t σ t c t on silent periods .
After clipping to [ ρ min , 1 ] , this rule prioritizes histories that are recent, uncertain, and inexpensive to audit. The variance proxy σ t 2 must be computed before the current label arrives. One admissible implementation is an exponentially weighted estimate based only on losses revealed through time t 1 , optionally stratified by the current observed covariates. Because the resulting ρ t is H t -measurable and is clipped below by ρ min , the adaptive audit policy preserves the predictability and positivity assumptions of the main theorems. A proxy that uses Y t or the current unrevealed loss would not be admissible.

5.2. Robustness to Estimated Propensities

In some archives, q t is not logged and must be estimated. Let q ^ t be H t -measurable and clipped as
q ˜ t = max { q ^ t , q clip } .
Define
Z ˜ t ( λ ) = I t L t ( λ ) q ˜ t , R ˜ n ( λ ) = 1 A n t = 1 n a t Z ˜ t ( λ ) .
Because predictable weights may themselves be random, Proposition 1 must not treat the realized sum t a t 2 as a deterministic variance envelope. Let s 2 , n be fixed before the protected labels are inspected and satisfy
t = 1 n a t 2 s 2 , n almost surely .
For example, a t c n implies the conservative choice s 2 , n = n c n 2 . A valid deterministic Bernstein radius is therefore
B ˜ n ( δ ) = 2 s 2 , n log ( 1 / δ ) / q clip 2 A n + c n log ( 1 / δ ) 3 q clip A n .
Assume a multiplicative calibration envelope
1 ε t q t q ˜ t 1 + ε t , 0 ε t < 1 .
By conditional ignorability, E [ Z ˜ t ( λ ) H t ] = R t ( λ ) q t / q ˜ t . Hence, the two-sided ratio envelope in Equation (92) implies the absolute bound
R t ( λ ) E [ Z ˜ t ( λ ) H t ] = R t ( λ ) 1 q t q ˜ t ε t R t ( λ ) ε t .
The weighted bias allowance is
E n = 1 A n t = 1 n a t ε t .
In practice, the envelope ε t should come from a validation exercise for the inclusion model rather than from the same outcomes used for risk calibration. For example, randomized audit assignments provide known Bernoulli probabilities and can be used to check whether the logged inclusion mechanism matches the declared policy. When alert propensities are estimated, cross-fitting or a held-out logging sample can supply a multiplicative calibration band for q t / q ˜ t . The complete-log experiment additionally perturbs the propensities by ± 10 % and ± 20 % on a complete-log replay to show the empirical effect of misspecification; such a sensitivity study is diagnostic unless the perturbation envelope is itself certified.
Proposition 1
(Propensity-robust upper bound). Replacing q t by q ˜ t and adding E n yields
R ¯ n ( λ ) R ˜ n ( λ ) + B ˜ n ( δ ) + E n
with probability at least 1 δ for each fixed λ.
Proof. 
Decompose
R ¯ n R ˜ n = 1 A n t = 1 n a t { R t E [ Z ˜ t H t ] } + 1 A n t = 1 n a t { E [ Z ˜ t H t ] Z ˜ t } .
For the first term, Equation (93) gives R t E [ Z ˜ t H t ] | R t E [ Z ˜ t H t ] | ε t , so its weighted contribution is at most E n . For the second martingale term, the increment bound is c n / q clip , while Equation (90) supplies the deterministic variance envelope s 2 , n / q clip 2 . Applying Lemma 1 with these deterministic quantities gives Equation (91); no realized predictable weight sequence is substituted into the deterministic envelope. Adding both terms proves Equation (95).    □

5.3. Oracle Efficiency

This subsection is a comparison-only efficiency analysis. None of the margin, Lipschitz, or oracle assumptions introduced here are required for the validity of Theorem 2; they are used only to quantify how far a conservative certified threshold may lie from an ideal full-information choice.
Let λ be the smallest grid threshold satisfying the true risk constraint,
λ = min { λ Λ : R ¯ n ( λ ) α } .
Assume a local margin condition: for λ λ ,
R ¯ n ( λ ) R ¯ n ( λ ) κ ( λ λ ) ,
and let expected set size s n ( λ ) be locally Lipschitz,
0 s n ( λ 2 ) s n ( λ 1 ) L s ( λ 2 λ 1 ) .
For comparison, suppose the deterministic variance envelope in Equation (32) is available and let the resulting analytic-boundary calibrator use
U n B ( λ ) = R ^ n ( λ ) + B n ( δ / | Λ | , λ ) ,
and let λ ^ n B be its smallest feasible threshold; if the feasible set is empty, the calibrator abstains. Under the grid-existence condition of Theorem 5, the feasible set is nonempty on the stated simultaneous event. Write
β n = max λ Λ B n ( δ / | Λ | , λ ) .
Theorem 5
(Finite-sample efficiency gap). On the two-sided simultaneous Bernstein event for Equation (100), suppose the grid contains a point at or above λ + 2 β n / κ (equivalently, λ + 2 β n / κ 1 for a grid containing 1). If the grid spacing is Δ Λ , then
0 λ ^ n B λ 2 β n κ + Δ Λ ,
and
0 s n ( λ ^ n B ) s n ( λ ) L s 2 β n κ + Δ Λ .
Proof. 
The two-sided Bernstein event based on the pre-specified variance envelopes gives R ¯ n ( λ ) U n B ( λ ) and U n B ( λ ) R ¯ n ( λ ) + 2 β n simultaneously over the grid. Consider the smallest grid point λ + satisfying
λ + λ + 2 β n / κ .
By Equation (98),
R ¯ n ( λ + ) R ¯ n ( λ ) 2 β n α 2 β n .
Hence, U n B ( λ + ) α , so the minimal feasible grid point cannot exceed λ + . Rounding contributes at most Δ Λ , proving (102). Applying Equation (99) gives (103).    □

6. Numerical Implementation and Reproducibility

This section collects the numerical steps that were previously dispersed across the theoretical development. The theoretical results above define the target and validity conditions; the present section explains how the certificate is evaluated, how the threshold is selected, and what must be logged for replication.
Figure 2 summarizes the information flow. The two label streams are merged only after their inclusion probabilities are computed. History weights then determine the target temporal population, and the e-process converts weighted losses into a finite-sample boundary. The output is a predictive set, not a corrected point forecast.
The architecture deliberately separates prediction, observation, and certification. The event model may be replaced without changing the risk layer. The selective stream supplies many labels near operationally important alerts, while the audit stream supplies support in silent periods. Their known propensities enter a single weighted loss. A sequential e-process then accounts for dependence through conditional expectations rather than an independence approximation. This modularity is useful in practice because model retraining, audit design, and risk-policy updates can proceed on different schedules without invalidating the logical role of each component.
Algorithm 1 summarizes the selective-observation weighted risk-control procedure used in the numerical implementation.
Algorithm 1 Selective-observation weighted risk control
Require: 
Fixed grid λ 1 < < λ J = 1 , target α , deployment buffer Δ ^ (externally certified only when prospective validity is claimed), calibration confidence δ , weights a 1 : n , propensities q 1 : n
  1.
Set δ j δ / J for j = 1 , , J
  2.
for  j = 1 , , J do
  3.
    Compute Z t ( λ j ) = I t L t ( λ j ) / q t and R ^ n ( λ j )
  4.
    Invert Equation (46) to obtain U n ( λ j ; δ j )
  5.
end for
  6.
F { λ j : U n ( λ j ; δ j ) α Δ ^ }
  7.
if  F then
  8.
     λ ^ min F
  9.
else
10.
     λ ^ λ J
11.
end if
12.
return  C n + 1 , λ ^ ( · )
Figure 3 illustrates one calibration path. The observed weighted risk closely follows the risk computed from the hidden complete labels, while the upper boundary remains above both. The selected threshold is the first nested-set parameter for which the simultaneous boundary falls below the buffered certification budget.
The dark curve is available only to the simulator and acts as an oracle diagnostic. The inverse-inclusion estimate tracks it despite highly nonuniform label availability. The red martingale-mixture boundary adds a nonasymptotic correction that is largest for small sets, where losses and inverse-inclusion exposure are high. The green vertical line occurs where the simultaneous boundary crosses the buffered 0.09 certification budget. Its location is more conservative than the empirical crossing, which explains the moderate set-size inflation observed later. Importantly, calibration uses no hidden labels; the oracle curve is displayed solely for verification.
Remark 3.
SOWRC is conformalized risk control in the sense that it calibrates a nested family generated by any scoring model, but its proof is not an exchangeability proof. The essential mechanism is predictable inverse-inclusion weighting plus a martingale boundary. This distinction permits arbitrary temporal dependence subject to the predictability, ignorability, positivity, bounded-loss, and design-envelope conditions stated above, yet it also makes the positivity constant explicit: sparse audits widen the bound through q min 1 , exposing the price of identification rather than hiding it in an asymptotic approximation.
The replication scripts, numerical settings, and random seeds are provided in the Supplementary Materials.
For all reported calculations, the mixture contains L = 32 equally weighted betting fractions. Let η max denote the largest positive solution margin for which κ n ( η ) > 0 . We use a geometric grid
η = 10 log 10 ( 10 5 ) + 1 31 { log 10 ( 0.98 η max ) log 10 ( 10 5 ) } , w = 1 / 32 ,
for = 1 , , 32 . Thus all 32 betting fractions and their weights are determined explicitly by Equation (106); the supplementary implementation should emit this 32-entry vector and use the same vector when reproducing the reported tables and figures. The factor 0.98 keeps every component strictly inside the admissible region. Equation (47) is inverted by bisection on [ 0 , 1 ] to an absolute tolerance of 10 8 . The threshold grid is fixed before labels are inspected. These choices are numerical design decisions rather than additional stochastic assumptions.

Computational Complexity

For classification, sort the observed true-label scores once. With  n o = t I t , preprocessing costs
T sort = O ( n o log n o ) .
Cumulative weighted losses over J thresholds can then be computed in
T grid = O ( n o + J ) .
Maintaining K losses and L mixture components online requires
T online = O ( K L )
per labeled time point and memory
M online = O ( K L + J ) .
The calibration layer is therefore typically cheaper than fitting the underlying event model.
Remark 4.
Efficiency is governed by two separable quantities. The predictive model controls how rapidly risk falls as the set expands, represented by the margin κ. The observation design controls the boundary through n eff . Better scores cannot repair zero support, and more audits cannot repair an uninformative score. This separation suggests a practical workflow: first guarantee positivity through randomized audits, then improve set efficiency through modeling and targeted audit allocation.
The reproducibility configuration used in Section 7 is: Python 3.13.5, NumPy 2.3.5, SciPy 1.17.0, scikit-learn 1.8.0, and statsmodels 0.14.6 for the complete-log real-data replay. For the 20-replication synthetic application studies, the stream seed is set equal to the replication index (0–19); randomized audit-mask replays use seeds 0–99. Alert and audit decisions are generated before the corresponding protected label, and the logged quantity used by SOWRC is the declared inclusion probability q t , not a minimum estimated from the realized path.

7. Experiments

7.1. Design, Data, and Baselines

We evaluated SOWRC on two synthetic dependent event streams designed to isolate selective-label bias. The maintenance stream contains vibration, temperature, acoustic, load, ambient, trend, and rolling-summary variables. The financial stream contains returns, volatility, momentum, drawdown, downside variation, spread, liquidity, and a noisy observable regime proxy; the latent Markov state is never supplied to the classifier. A logistic model is fitted on the initial selectively labeled segment with inverse-inclusion weights. Calibration labels arise from either the alert rule or randomized audits, whereas the final segment is fully revealed only for evaluation. To remove the synthetic-only limitation, Section 7.4 adds a complete-log replay using a real measured weekly sensor series. There, the original outcomes are retained, selective observation is imposed prospectively from recorded history, and 100 independent audit-mask realizations quantify Monte Carlo uncertainty.
Table 4 summarizes the controlled benchmarks. The split proportions are 0.22, 0.48, and 0.30 for training, calibration, and testing. Moderate AUC values make set calibration nontrivial. Alert rates below one third create substantial support imbalance, while audit rates of 0.10–0.12 prevent extreme inverse weights. The financial predictor is intentionally weaker after removing the latent regime state. Thus, the comparison tests the certification layer rather than relying on an unrealistically accurate classifier. All uncertainty summaries are computed across independent time-series seeds.
The operational target is α = 0.10 with calibration confidence level 1 δ = 0.90 . For a conservative sensitivity analysis, we reserve a numerical buffer of 0.01 and calibrate at 0.09 . This 0.01 value is only a numerical stress buffer. It is not an estimate of Δ ^ n , H and is never used to claim Corollary 1; prospective validity requires an externally constructed certificate such as Equations (59) and (60). The deterministic grid is Λ = { 0 , 1 / 240 , , 1 } , and each grid point receives δ / 241 . Baselines are alert-only calibration, unweighted observed-label calibration, a time-decayed observed-label quantile, an inverse-propensity-weighted (IPW) quantile, and an oracle quantile using complete calibration labels. Modern online, non-exchangeable, and adaptive conformal methods are not entered as nominally equivalent baselines because their published guarantees presuppose observed calibration outcomes or a support condition that the silent region violates. We therefore compare against their relevant ingredients—time adaptation and reweighting—without relabeling those implementations as methods whose assumptions are not met.

7.2. Predictive-Maintenance Results

Figure 4 displays a representative deployment interval. The upper panel shows the observed signal, alert-selected periods, and realized miscoverage, while the lower panel shows prediction-set size.
Figure 4 shows abrupt transitions between calm, degrading, and fault-prone regimes. The adaptive alert threshold is now drawn explicitly; unlike the previous version, its movement is visible on the same panel as the selected periods. Near-boundary observations are therefore seen relative to the moving cutoff rather than to an implicit constant line. The plotted trajectory is illustrative rather than evidence of average performance, and we do not use an adaptive-versus-constant visual comparison as a performance claim. Typicality is assessed by the across-replication summaries in Table 5; no performance claim relies on this single path. The aligned lower panel confirms that SOWRC expands to the two-label set during uncertain intervals and contracts after evidence stabilizes. Most remaining errors occur in silent periods near regime changes, where alert history is sparse. The set-size path is smoother than the raw vibration signal because calibration acts on weighted cumulative losses.
Table 5 isolates the selection effect. Alert-only calibration reaches 0.140 complete-label test risk because its apparent accuracy is concentrated in the alert region. Unweighted and time-decayed rules remain biased by the same observed-label composition. The IPW quantile corrects most of the bias but has no high-probability finite-sample certificate. SOWRC reaches 0.058 risk, with larger sets caused by the 0.01 sensitivity buffer and simultaneous grid correction. The mean calibration-to-test increase is 0.0035, but 3 of 20 seeds exceed 0.01 and the maximum increase is 0.018. Thus, these held-out results do not themselves certify Assumption 4. For the SOWRC test risk, the Monte Carlo standard error is 0.009 / 20 = 0.0020 and the normal-approximation 95% interval for the across-seed mean is approximately [ 0.054 , 0.062 ] . Reporting both the standard deviation and MCSE separates trajectory-to-trajectory variability from uncertainty in the reported mean.
Figure 5 compares complete-label test risk and average set size, with one-standard-deviation error bars. All means and bars are displayed inside the plotting window.
Figure 5 separates bias correction from statistical certification. Alert-only, unweighted, and time-decayed methods form a low-size but high-risk cluster. The IPW quantile lies close to the oracle, showing that inverse weighting removes most selection bias. SOWRC is more conservative because it also pays for simultaneous finite-sample validity and the numerical sensitivity buffer. Its entire one-standard-deviation risk bar remains below 0.10. The figure should not be read as claiming pointwise dominance: the additional set size is the explicit cost of support recovery, multiplicity control, and conservative buffering.

7.3. Financial Early-Warning Results

Figure 6 shows a financial deployment window with abrupt volatility and drawdown changes.
Figure 6 depicts prolonged calm periods interrupted by short stress episodes. The aligned lower panel separates discrete set expansion from the continuously scaled drawdown trajectory. The classifier observes only a noisy regime proxy, not the latent Markov state, so its discrimination is modest. SOWRC often enlarges the set before severe stress because volatility, liquidity, spread, and momentum jointly increase uncertainty. Remaining errors are dispersed across transitions rather than concentrated in one state. This behavior is consistent with aggregate population-risk control; it does not imply conditional coverage within every latent market regime, a stronger objective that would require additional structure and larger audit budgets. As with Figure 4, this trajectory is an illustrative realization rather than a hand-selected best case; the general behavior is summarized by the 20-seed statistics in Table 6.
Table 6 shows a smaller but persistent selection effect in complete-label test risk. The uncorrected methods exceed the 0.10 target despite compact sets. The IPW quantile and oracle are near 0.09, whereas SOWRC reaches 0.052 because the simultaneous boundary is costly for a weaker predictor and a 0.10 audit rate. The mean calibration-to-test increase is 0.0025; 1 of 20 seeds exceeds 0.01, with a maximum increase of 0.0118. The result demonstrates conservative held-out performance, not unconditional future validity or an empirical proof of the drift certificate. For SOWRC, the test-risk MCSE is 0.007 / 20 = 0.0016 , giving an approximate 95% interval of [ 0.049 , 0.055 ] for the across-seed mean.

7.4. Complete-Log Real-Data Replay and Conservativeness

To complement the controlled synthetic experiments, we added a complete-log replay based on the weekly Mauna Loa atmospheric CO2 series distributed with statsmodels [33]. The purpose is not to claim a domain-specific CO2 forecasting contribution, but to test selective observation on a genuine dependent sensor record for which every outcome is available before masking. The 59 missing weekly measurements in the distributed series are filled by one-sided forward filling before lagged features are constructed, so no future event label is used in preprocessing. After removing unavailable lags, 2231 weekly observations remain. We use 557 observations for model fitting, 1003 for calibration, and 671 for complete-label testing. An event is defined as an absolute next-week CO2 change exceeding the 80th percentile of the training changes. A logistic score uses lagged level/change, rolling mean and scale, and seasonal terms. The alert threshold is fixed at the 75th percentile of training scores, producing a calibration alert rate of 0.290. Silent labels are then independently audited with probability ρ = 0.30 . The threshold grid has 121 deterministic values, and each selective-observation method is replayed under 100 independently generated audit masks (seeds 0–99).
Table 7 shows the cost of finite-sample certification on a non-synthetic trajectory. SOWRC is markedly conservative, whereas the oracle and IPW quantile are much closer to the 0.10 target. For SOWRC the MCSE of mean test risk is 0.0023 / 100 = 0.00023 , yielding an approximate 95% Monte Carlo interval [ 0.0181 , 0.0190 ] . Because the complete labels are retained only for evaluation, this replay also makes clear which quantities are deployment diagnostics rather than inputs to the certificate. These Monte Carlo intervals should therefore be interpreted as conditional audit-mask variability for one fixed trajectory, not as a population-level interval over possible sensor series.
Table 8 separates four sources that were previously bundled together. Inverse weighting mainly repairs selective-observation bias; the martingale deviation term causes the largest enlargement; simultaneous Bonferroni protection adds a further visible cost; and the 0.01 numerical buffer is comparatively smaller. This decomposition does not imply that the pointwise row is a valid replacement for the simultaneous procedure: it is included only to quantify the multiplicity price. Sharper simultaneous boundaries that exploit monotonicity across λ are an important direction for reducing this price.
The misspecification experiment in Table 9 is deliberately reported as a sensitivity diagnostic rather than a validity theorem. Underestimating inclusion probabilities increases the inverse weights and therefore enlarges the sets; overestimation has the opposite effect. Formal validity with estimated propensities still requires the multiplicative calibration envelope stated in Proposition 1. In practice, this envelope can be assessed from randomized logging records, held-out propensity calibration, or cross-fitted policy models before outcomes from the calibrated decision are used.

7.5. Audit Ablation and Selection-Level Certificate Validation

The complete-log replay also permits the audit rate itself to be varied without changing the underlying trajectory or predictor. We use ρ { 0.10 , 0.20 , 0.30 , 0.40 , 0.50 } and 100 independent audit masks at every setting. Error bars in Figure 7 are 95% Monte Carlo intervals for the across-mask means.
At ρ = 0.10 and 0.20 , the simultaneous boundary is so wide that the safe terminal action is selected in all masks, giving the full two-label set and zero empirical test loss. At  ρ = 0.30 , the mean test risk is 0.0185 ± 0.0023 ; increasing the audit rate to 0.40 and 0.50 raises mean risk to approximately 0.0247 and 0.0278 while reducing set size, because the effective sample size grows and the certificate becomes less conservative. This behavior is the empirical counterpart of Equation (82): the audit floor controls both identifiability and the statistical price of inverse weighting.
We also replaced the earlier fixed-threshold diagnostic by a selection-level experiment that runs the final threshold-selection rule on a persistent two-state Markov score process. Conditional losses are generated from state-dependent beta distributions, the candidate grid contains J = 41 fixed thresholds, and the true selected conditional population risk is available analytically for checking. For each of the eight configurations below, 500 independent trajectories are generated. We vary horizon, audit rate, confidence level, target risk, and state persistence. A failure occurs when the risk of the threshold actually selected by SOWRC exceeds the declared target after the certificate is applied. Table 10 reports the resulting selection-level finite-sample validation results.
Across all 8 × 500 = 4000 selection-level replications, no failure was observed. With zero failures in 500 trials, the one-sided exact 95% binomial upper bound is 0.006 for each configuration. The result should not be read as evidence that the procedure is calibrated tightly to the nominal δ : the fallback rates and positive certificate gaps show substantial conservatism, especially for short horizons, sparse audits, or the smaller target risk. Its purpose is instead to verify that the data-dependent grid selection—the central operation of SOWRC—does not invalidate the finite-sample protection in these dependent simulations.

7.6. Implementation Checks and Limitations

All experiments use exact logged inclusion probabilities and a deterministic lower propensity bound supplied by the audit design; the code never substitutes a favorable minimum computed from the realized path. The candidate grid is fixed before labels are inspected, eliminating data reuse from empirical score quantiles. Predictable weights increase linearly from 0.95 to 1.05, and the mixture uses 32 admissible betting fractions. The implementation searches all 241 thresholds with Bonferroni confidence allocation. The code exports every replication, drift diagnostic, representative stream, summary table, validation experiment, and figure. The two application streams remain synthetic controlled demonstrations, but the complete-log CO2 replay now adds a real dependent sensor record with selective labels imposed only after the fact. The replay is intentionally a methodological stress test rather than a claim of domain-specific superiority; complete labels are used solely for post hoc evaluation.
Three limitations remain. First, conditional ignorability can fail when annotators use latent information absent from the recorded history. Randomized audits are preferable in that setting. Second, future-horizon control is not distribution-free under unrestricted drift; it requires an external certificate satisfying Assumption 4, with its error δ dep included, or continued online auditing through Theorem 3. Third, the guarantee is weighted population-level risk, not equal conditional risk in every regime or subgroup. These distinctions prevent calibration-population validity, prospective transfer, and conditional coverage from being conflated.
Remark 5.
The revised experiments distinguish three questions. Inverse weighting addresses bias from selective observation. The martingale boundary certifies the weighted calibration population. A separately certified drift envelope is required to connect that certificate to a future deployment horizon; the numerical buffer used here is only a sensitivity choice. None of these steps substitutes for another. This separation makes the source of conservatism visible and prevents low selected-sample error from being mistaken for future population-risk control.

8. Conclusions and Future Work

This paper studied predictive sets when labels are observed selectively along a dependent stochastic process. Population risk is non-identifiable on any region with zero labeling probability. A randomized audit stream restores support, and SOWRC combines alert and audit labels through predictable inverse-inclusion losses. A martingale-mixture boundary then provides finite-sample control of the weighted calibration population on a deterministic threshold grid. The prospective result states the additional condition required for future use: an externally certified deployment-drift envelope, explicit error allocation, and a matching calibration buffer.
The theory now separates four logically distinct components. Positivity determines identification. Effective sample size determines statistical precision. Monotone nested losses and a safe terminal action justify data-dependent threshold selection. A deployment envelope determines whether a calibration certificate transfers to future events. The experiments use a deterministic grid, remove latent-state leakage, and name baselines according to their implementations. They now also include a complete-log real-data replay with 100 audit masks and Monte Carlo uncertainty, a decomposition of the main sources of conservativeness, a propensity-misspecification sensitivity analysis, and selection-level validation over 4000 dependent-process replications.
Three directions merit further study. First, audit policies should be optimized jointly with prediction under a long-run labeling budget while preserving minimum support. Second, deployment-drift envelopes should be learned from independent auxiliary trajectories with simultaneous uncertainty quantification. Third, doubly robust and distributionally robust extensions should control delayed outcomes, partially observed covariates, and regime-specific risks without sacrificing anytime validity.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/axioms15090706/s1, Supplementary File S1: Replication materials, including experiment settings, random seeds, and figure-generation specifications.

Author Contributions

Conceptualization, methodology, formal analysis, software, visualization, and writing—original draft preparation, S.B.; validation and writing—review and editing, Z.F.; supervision, project administration, validation, and writing—review and editing, J.C. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the National Natural Science Foundation of China Youth Science Fund under Grant 62103178.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The synthetic generators, experiment settings, replication seeds, and figure-generation specifications are documented in the manuscript and are provided in the Supplementary Materials. The real-data replay uses the publicly distributed weekly Mauna Loa CO2 dataset in statsmodels; its preprocessing split, event definition, and audit-mask seeds are reported in the numerical-implementation and experiment sections.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Angelopoulos, A.N.; Bates, S.; Fisch, A.; Lei, L.; Schuster, T. Conformal risk control. In Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024; OpenReview: Amherst, MA, USA, 2024. [Google Scholar]
  2. Doria, S. Bayesian updating for stochastic processes in infinite-dimensional normed vector spaces. Axioms 2025, 14, 927. [Google Scholar] [CrossRef] [Scilit]
  3. Bao, Y.; Huo, Y.; Ren, H.; Zou, C. CAP: A general algorithm for online selective conformal prediction with FCR control. J. Mach. Learn. Res. 2025, 26, 1–74. [Google Scholar]
  4. Gibbs, I.; Candès, E.J. Conformal inference for online prediction with arbitrary distribution shifts. J. Mach. Learn. Res. 2024, 25, 1–36. [Google Scholar]
  5. Xu, C.; Jiang, H.; Xie, Y. Conformal prediction for multi-dimensional time series by ellipsoidal sets. In Proceedings of the 41st International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2024; Volume 235, pp. 55076–55099. [Google Scholar]
  6. Strang, A. A theoretical review of area production rates as test statistics for detecting nonequilibrium dynamics in Ornstein–Uhlenbeck processes. Axioms 2024, 13, 820. [Google Scholar] [CrossRef] [Scilit]
  7. Oliveira, R.I.; Orenstein, P.; Ramos, T.; Romano, J.V. Split conformal prediction and non-exchangeable data. J. Mach. Learn. Res. 2024, 25, 1–38. [Google Scholar]
  8. Zheng, F.; Proutière, A. Conformal predictions under Markovian data. In Proceedings of the 41st International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2024; Volume 235, pp. 61470–61497. [Google Scholar]
  9. Galvão Lopes, A.; Goubault, E.; Putot, S.; Pautet, L. ConForME: Multi-horizon conditional conformal time series forecasting. In Proceedings of the Thirteenth Symposium on Conformal and Probabilistic Prediction with Applications; PMLR: Cambridge, MA, USA, 2024; Volume 230, pp. 345–365. [Google Scholar]
  10. Jacquet, O.; Oscar, W.; Vaillant, J. Partially observed two-phase point processes. Axioms 2026, 15, 59. [Google Scholar] [CrossRef] [Scilit]
  11. Zhou, Y.; Lindemann, L.; Sesia, M. Conformalized adaptive forecasting of heterogeneous trajectories. In Proceedings of the 41st International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2024; Volume 235, pp. 62002–62056. [Google Scholar]
  12. Prinzhorn, D.; Nijdam, T.; van der Linden, P.; Timans, A. Conformal time series decomposition with component-wise exchangeability. In Proceedings of the Thirteenth Symposium on Conformal and Probabilistic Prediction with Applications; PMLR: Cambridge, MA, USA, 2024; Volume 230, pp. 432–465. [Google Scholar]
  13. Hallberg Szabadváry, J. Adaptive conformal inference for multi-step ahead time-series forecasting online. In Proceedings of the Thirteenth Symposium on Conformal and Probabilistic Prediction with Applications; PMLR: Cambridge, MA, USA, 2024; Volume 230, pp. 250–263. [Google Scholar]
  14. Zhang, H.; Sun, J.; Wen, X. Stationary distribution and density function for a high-dimensional stochastic SIS epidemic model with mean-reverting stochastic process. Axioms 2024, 13, 768. [Google Scholar] [CrossRef] [Scilit]
  15. Blot, V.; Angelopoulos, A.N.; Jordan, M.I.; Brunel, N.J.-B. Automatically adaptive conformal risk control. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2025; Volume 258, pp. 19–27. [Google Scholar]
  16. Einbinder, B.-S.; Feldman, S.; Bates, S.; Angelopoulos, A.N.; Gendler, A.; Romano, Y. Label noise robustness of conformal prediction. J. Mach. Learn. Res. 2024, 25, 1–66. [Google Scholar]
  17. Gnecco, N.; Peters, J.; Engelke, S.; Pfister, N. Boosted control functions: Distribution generalization and invariance in confounded models. J. Mach. Learn. Res. 2026, 27, 1–57. [Google Scholar]
  18. Howard, S.R.; Ramdas, A.; McAuliffe, J.; Sekhon, J. Time-uniform, nonparametric, nonasymptotic confidence sequences. Ann. Stat. 2021, 49, 1055–1080. [Google Scholar] [CrossRef] [Scilit]
  19. Vovk, V.; Wang, R. E-values: Calibration, combination, and applications. Ann. Stat. 2021, 49, 1736–1754. [Google Scholar] [CrossRef] [Scilit]
  20. Jonkers, J.; Van Wallendael, G.; Duchateau, L.; Van Hoecke, S. Conformal predictive systems under covariate shift. In Proceedings of the Thirteenth Symposium on Conformal and Probabilistic Prediction with Applications; PMLR: Cambridge, MA, USA, 2024; Volume 230, pp. 406–423. [Google Scholar]
  21. Hussain, T.; Villamor, E.; Shakil, M.; Ahsanullah, M.; Kibria, B.M.G. A novel probabilistic model for streamflow analysis and its role in risk management and environmental sustainability. Axioms 2026, 15, 113. [Google Scholar] [CrossRef] [Scilit]
  22. Chvoinikov, D.; Novikov, S.; Šiaulys, J. Extremes of product of Gaussian random variables. Axioms 2026, 15, 425. [Google Scholar] [CrossRef] [Scilit]
  23. Hulsman, R.; Comte, V.; Bertolini, L.; Wiesenthal, T.; Puertas Gallardo, A.; Ceresa, M. Conformal risk control for pulmonary nodule detection. In Proceedings of the Fourteenth Symposium on Conformal and Probabilistic Prediction with Applications; PMLR: Cambridge, MA, USA, 2025; Volume 266, pp. 445–463. [Google Scholar]
  24. Zhang, C.; Li, T.; Xie, J.; Kong, L.; Jiang, B. Transfer conformal predictive inference for regression. J. Mach. Learn. Res. 2026, 27, 1–68. [Google Scholar]
  25. Barber, R.F.; Candès, E.J.; Ramdas, A.; Tibshirani, R.J. Conformal prediction beyond exchangeability. Ann. Stat. 2023, 51, 816–845. [Google Scholar] [CrossRef] [Scilit]
  26. Zaffran, M.; Féron, O.; Goude, Y.; Josse, J.; Dieuleveut, A. Adaptive conformal predictions for time series. In Proceedings of the 39th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2022; Volume 162, pp. 25834–25866. [Google Scholar]
  27. Barber, R.F.; Pananjady, A. Predictive inference for time series: Why is split conformal effective despite temporal dependence? In Proceedings of the 37th International Conference on Algorithmic Learning Theory; PMLR: Cambridge, MA, USA, 2026; Volume 313, pp. 1–24. [Google Scholar]
  28. Chen, B.; Zhou, X.; Cheng, L. Post-training adaptive conformal prediction for incomplete time series. Trans. Mach. Learn. Res. 2026. Available online: https://openreview.net/forum?id=KMBU4wx79B (accessed on 19 September 2026).
  29. Bates, S.; Angelopoulos, A.N.; Lei, L.; Malik, J.; Jordan, M.I. Distribution-free, risk-controlling prediction sets. J. ACM 2021, 68, 43. [Google Scholar] [CrossRef] [Scilit]
  30. Angelopoulos, A.N.; Bates, S. Conformal prediction: A gentle introduction. Found. Trends Mach. Learn. 2023, 16, 494–591. [Google Scholar] [CrossRef] [Scilit]
  31. Horvitz, D.G.; Thompson, D.J. A generalization of sampling without replacement from a finite universe. J. Am. Stat. Assoc. 1952, 47, 663–685. [Google Scholar] [CrossRef]
  32. Robins, J.M.; Rotnitzky, A.; Zhao, L.P. Estimation of regression coefficients when some regressors are not always observed. J. Am. Stat. Assoc. 1994, 89, 846–866. [Google Scholar] [CrossRef]
  33. Seabold, S.; Perktold, J. Statsmodels: Econometric and statistical modeling with Python. In Proceedings of the 9th Python in Science Conference, Austin, TX, USA, 28 June–3 July 2010; SciPy: Austin, TX, USA, 2010; pp. 92–96. [Google Scholar]
Figure 1. Selective labels and silent-region audits under a visibly history-adaptive alert threshold. The threshold is updated from lagged observed scores only, so the change in the selection pattern is predictable rather than label-driven.
Figure 1. Selective labels and silent-region audits under a visibly history-adaptive alert threshold. The threshold is updated from lagged observed scores only, so the change in the selection pattern is predictable rather than label-driven.
Axioms 15 00706 g001
Figure 2. SOWRC information flow from selective labels and audits to risk-controlled predictive sets.
Figure 2. SOWRC information flow from selective labels and audits to risk-controlled predictive sets.
Axioms 15 00706 g002
Figure 3. Weighted risk estimate and finite-sample boundary across nested-set thresholds.
Figure 3. Weighted risk estimate and finite-sample boundary across nested-set thresholds.
Axioms 15 00706 g003
Figure 4. Maintenance signal, adaptive alert threshold, alert-selected periods, set size, and SOWRC miscoverage over time.
Figure 4. Maintenance signal, adaptive alert threshold, alert-selected periods, set size, and SOWRC miscoverage over time.
Axioms 15 00706 g004
Figure 5. Test-risk–set-size trade-off with one-standard-deviation error bars.
Figure 5. Test-risk–set-size trade-off with one-standard-deviation error bars.
Axioms 15 00706 g005
Figure 6. Financial drawdown, alert periods, set size, and SOWRC miscoverage over time.
Figure 6. Financial drawdown, alert periods, set size, and SOWRC miscoverage over time.
Axioms 15 00706 g006
Figure 7. Complete-log SOWRC replay as the silent-region audit probability varies. Points are means over 100 audit masks; bars are 95% Monte Carlo intervals. The bars quantify audit-mask variability conditional on the single replayed time series and do not include across-series or series-selection uncertainty.
Figure 7. Complete-log SOWRC replay as the silent-region audit probability varies. Points are means over 100 audit masks; bars are 95% Monte Carlo intervals. The bars quantify audit-mask variability conditional on the single replayed time series and do not include across-series or series-selection uncertainty.
Axioms 15 00706 g007
Table 1. Main symbols and definitions.
Table 1. Main symbols and definitions.
SymbolDefinition
tThe time index.
nThe fixed calibration horizon.
HThe prospective deployment horizon.
X t The observed covariate vector at time t.
Y t The event label at time t.
F t The post-label information available after time t.
H t The pre-label history σ ( F t 1 , X t ) .
S t The alert-selection indicator.
A t The auxiliary-audit indicator.
I t The label-inclusion indicator.
q t The conditional probability that the label is included.
Y The finite label space.
C t , λ ( x ) The nested predictive set at threshold λ .
λ The threshold controlling prediction-set size.
L t ( λ ) The bounded predictive loss.
R t ( λ ) The conditional population risk.
a t , n A horizon-indexed fixed-horizon weight, such as Equation (14).
a t A single predictable weight sequence used in the anytime setting.
c n , b n Deterministic design envelopes for a fixed terminal horizon.
c , b Global design envelopes used for time-uniform validity.
Z t ( λ ) The inverse-inclusion weighted observed loss.
A n The cumulative history weight.
U n ( λ ; δ ) The finite-sample upper risk boundary.
α t The predictable risk budget.
Λ The threshold grid fixed before labels are inspected.
Δ ^ n , H The certified calibration-to-deployment drift envelope.
δ dep The failure probability of the drift certificate.
ρ t The auxiliary-audit probability.
q min The deterministic lower bound on label-inclusion probability.
Table 2. Risk quantities used in the manuscript and their logical status.
Table 2. Risk quantities used in the manuscript and their logical status.
QuantityDefinition/LocationInterpretation
Conditional risk R t ( λ ) Equation (13)History-specific risk before the current label is revealed.
Calibration-population risk R ¯ n ( λ ) Equation (12)Finite-horizon weighted population target controlled by Theorem 2.
Complete-label test riskSection 7Empirical diagnostic computed only because the experimental test segment is fully revealed.
Deployment risk R ¯ n , H dep ( λ ) Equation (56)Future operational target; controlled only with a certified drift envelope.
Table 3. Core assumptions, associated guarantees, and empirical assessability.
Table 3. Core assumptions, associated guarantees, and empirical assessability.
Assumption/Design ConditionRoleMain ResultAssessability
Predictability of S t , A t , a t Choices are made before the current hidden label is revealed.Lemma 1, Theorems 2 and 3Checkable from system timing and logs.
Conditional ignorabilityMakes inverse-inclusion loss conditionally unbiased.Equations (25) and (26)Not fully testable from observed labels; strengthened by randomized audits.
Positivity q t q min > 0 Restores identification and bounds inverse weights.Theorem 1; Theorem 2Directly checkable when the audit policy is randomized and logged.
Bounded loss and design envelopesEnables finite-sample martingale concentration.Lemma 1Checkable by construction.
Nested safetyJustifies data-dependent threshold selection and safe fallback.Equation (54)Checkable from the loss and terminal set definition.
Certified drift envelopeLinks calibration risk to future deployment risk.Corollary 1Requires an external validation mechanism or process bound.
Table 4. Experimental design and representative time-series characteristics.
Table 4. Experimental design and representative time-series characteristics.
CharacteristicPredictive MaintenanceFinancial Risk
Total time points25,00026,000
Feature dimension78
Training/calibration/test5500/12,000/75005720/12,480/7800
Default silent audit probability0.120.10
Approximate alert rate0.250.26
Mean event prevalence0.250.14
Mean test AUC 0.728 ± 0.014 0.607 ± 0.019
Replications2020
Table 5. Predictive-maintenance performance over 20 dependent time-series replications. Entries are mean ± one standard deviation across independent trajectory seeds.
Table 5. Predictive-maintenance performance over 20 dependent time-series replications. Entries are mean ± one standard deviation across independent trajectory seeds.
MethodTest RiskSet SizeSilent RiskAlert Risk
Naive-selected 0.140 ± 0.012 1.180 ± 0.030 0.155 ± 0.013 0.092 ± 0.016
Observed-unweighted 0.129 ± 0.010 1.212 ± 0.033 0.149 ± 0.012 0.070 ± 0.015
Time-decayed quantile 0.130 ± 0.010 1.212 ± 0.036 0.149 ± 0.012 0.070 ± 0.016
IPW quantile 0.093 ± 0.012 1.355 ± 0.066 0.116 ± 0.014 0.023 ± 0.012
SOWRC 0.058 ± 0.009 1.537 ± 0.054 0.076 ± 0.011 0.003 ± 0.003
Oracle-full 0.092 ± 0.008 1.360 ± 0.043 0.116 ± 0.010 0.021 ± 0.009
Table 6. Financial-risk performance over 20 dependent time-series replications. Entries are mean ± one standard deviation across independent trajectory seeds.
Table 6. Financial-risk performance over 20 dependent time-series replications. Entries are mean ± one standard deviation across independent trajectory seeds.
MethodTest RiskSet SizeSilent RiskAlert Risk
Naive-selected 0.109 ± 0.006 1.083 ± 0.018 0.115 ± 0.006 0.092 ± 0.012
Observed-unweighted 0.108 ± 0.006 1.091 ± 0.020 0.115 ± 0.006 0.086 ± 0.010
Time-decayed quantile 0.108 ± 0.005 1.091 ± 0.020 0.115 ± 0.006 0.085 ± 0.013
IPW quantile 0.094 ± 0.009 1.169 ± 0.056 0.110 ± 0.008 0.047 ± 0.018
SOWRC 0.052 ± 0.007 1.524 ± 0.068 0.067 ± 0.008 0.008 ± 0.008
Oracle-full 0.092 ± 0.004 1.179 ± 0.026 0.111 ± 0.005 0.039 ± 0.011
Table 7. Complete-log CO2 replay under selective observation. Entries are mean ± standard deviation over 100 independent audit masks; “Oracle-full” uses all calibration labels before masking. The reported variability is conditional on this single observed CO2 time series and reflects only repeated audit-mask randomization; it does not quantify uncertainty across alternative real-world time series or uncertainty from selecting this particular series.
Table 7. Complete-log CO2 replay under selective observation. Entries are mean ± standard deviation over 100 independent audit masks; “Oracle-full” uses all calibration labels before masking. The reported variability is conditional on this single observed CO2 time series and reflects only repeated audit-mask randomization; it does not quantify uncertainty across alternative real-world time series or uncertainty from selecting this particular series.
MethodComplete-Label Test RiskMean Set Size
Alert-only 0.0402 ± 0.0000 1.8197 ± 0.0000
Observed-unweighted 0.0760 ± 0.0052 1.7459 ± 0.0091
IPW quantile 0.1076 ± 0.0070 1.6629 ± 0.0186
SOWRC 0.0185 ± 0.0023 1.9018 ± 0.0136
Oracle-full 0.1103 ± 0.0000 1.6587 ± 0.0000
Table 8. Conservativeness decomposition on the complete-log replay ( ρ = 0.30 , 100 audit masks). The pointwise martingale row is a diagnostic without grid-selection validity; Bonferroni restores simultaneous protection for the threshold sweep.
Table 8. Conservativeness decomposition on the complete-log replay ( ρ = 0.30 , 100 audit masks). The pointwise martingale row is a diagnostic without grid-selection validity; Bonferroni restores simultaneous protection for the threshold sweep.
Calibration RuleTest RiskMean Set Size
IPW empirical risk only 0.1076 ± 0.0070 1.6629 ± 0.0186
Martingale, pointwise δ = 0.10 0.0370 ± 0.0094 1.8298 ± 0.0207
Martingale + Bonferroni grid 0.0185 ± 0.0023 1.9018 ± 0.0136
Grid + numerical 0.01 buffer 0.0145 ± 0.0014 1.9214 ± 0.0077
Table 9. Propensity-misspecification diagnostic for the complete-log replay. The logged inclusion probability used by SOWRC is multiplied by the indicated factor before clipping; values are mean ± standard deviation over 100 audit masks.
Table 9. Propensity-misspecification diagnostic for the complete-log replay. The logged inclusion probability used by SOWRC is multiplied by the indicated factor before clipping; values are mean ± standard deviation over 100 audit masks.
Propensity ScaleTest RiskMean Set Size
0.8 0.0110 ± 0.0040 1.9390 ± 0.0199
0.9 0.0159 ± 0.0011 1.9173 ± 0.0058
1.0 0.0185 ± 0.0023 1.9018 ± 0.0136
1.1 0.0205 ± 0.0025 1.8904 ± 0.0128
1.2 0.0227 ± 0.0030 1.8791 ± 0.0134
Table 10. Selection-level finite-sample validation over 500 independent dependent-process replications per configuration. “Fallback” is the fraction of runs selecting the safe terminal action. This experiment is intended to confirm implementation correctness and conservativeness, not tight calibration to the nominal δ .
Table 10. Selection-level finite-sample validation over 500 independent dependent-process replications per configuration. “Fallback” is the fraction of runs selecting the safe terminal action. This experiment is intended to confirm implementation correctness and conservativeness, not tight calibration to the nominal δ .
n ρ δ α PersistenceFailures/50095% UpperFallback
10000.100.100.100.9500.0061.000
30000.100.100.100.9500.0060.000
30000.050.100.100.9500.0061.000
30000.200.100.100.9500.0060.000
30000.100.050.100.9500.0060.000
30000.100.200.100.9500.0060.000
30000.100.100.050.9500.0061.000
30000.100.100.100.8000.0060.000
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Bai, S.; Fang, Z.; Chen, J. Risk-Controlling Predictive Sets for Time-Series Events Under Selective Observation with Finite-Sample Guarantees. Axioms 2026, 15, 706. https://doi.org/10.3390/axioms15090706

AMA Style

Bai S, Fang Z, Chen J. Risk-Controlling Predictive Sets for Time-Series Events Under Selective Observation with Finite-Sample Guarantees. Axioms. 2026; 15(9):706. https://doi.org/10.3390/axioms15090706

Chicago/Turabian Style

Bai, Siyang, Zheng Fang, and Jie Chen. 2026. "Risk-Controlling Predictive Sets for Time-Series Events Under Selective Observation with Finite-Sample Guarantees" Axioms 15, no. 9: 706. https://doi.org/10.3390/axioms15090706

APA Style

Bai, S., Fang, Z., & Chen, J. (2026). Risk-Controlling Predictive Sets for Time-Series Events Under Selective Observation with Finite-Sample Guarantees. Axioms, 15(9), 706. https://doi.org/10.3390/axioms15090706

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Article metric data becomes available approximately 24 hours after publication online.
Back to TopTop