1. Introduction
Prediction intervals for a non-stationary commodity price must satisfy two demands at once. They need a defensible coverage guarantee, and they need to be sharp when volatility shifts. Regime-weighted conformal intervals appear to deliver both: when an interval is formed, past forecast errors drawn from volatility states resembling the forecast origin receive more weight than the rest, so the interval widens in turbulent regimes and tightens in calm ones.
On daily Ribbed Smoked Sheet No. 3 (RSS3) natural-rubber price changes, such a construction seems to work exactly as advertised. Weighting the split-conformal calibration residuals by a two-state Markov-switching regime probability, with a per-regime asymmetric tail allocation, yields intervals about 20% narrower than split-conformal and a significantly better interval score. This paper asks whether that advantage is real, and finds that it is not. The audit examines three implementation features associated with it: an oracle filtered state that conditions on the target-period shock, calibration-dependent estimation of the regime model, and omission of the finite-sample weighted-quantile correction. Once the correction and fixed symmetric tails are applied, no construction that withholds target-period information retains a statistically detectable improvement over split-conformal.
1.1. What the Conformal Guarantee Rests on
Because the argument turns on the internal mechanics of the method, it is worth stating plainly what a conformal interval is. Conformal prediction builds intervals from the empirical distribution of realized forecast errors rather than from an assumed error distribution [
1,
2,
3]; Stocker et al. [
4] give an introduction aimed at time-series practitioners. In its split, or inductive, variant [
5,
6], the sample is divided into disjoint blocks: a model is fitted on a training block, and a separate calibration block supplies the nonconformity scores, here the residuals of that fitted model. The interval for a new observation is the point forecast plus the appropriate empirical quantile of those scores, a step that conformalized quantile regression [
7] refines so that width tracks local variability.
The resulting interval carries a finite-sample marginal coverage guarantee of at least
, for any underlying model and any error distribution. That guarantee is what distinguishes conformal prediction from other distribution-free routes—resampling the standardized innovations of a state-space model [
8], for instance, yields intervals with no such property. It is not, however, assumption-free. It requires the nonconformity scores to be exchangeable, meaning that their joint distribution is invariant to reordering, and daily commodity price changes violate this through volatility clustering and abrupt regime shifts. Everything that follows concerns how far that violation can be repaired by re-weighting the calibration scores.
Three quantities compare intervals throughout: empirical coverage, mean interval width, and the Winkler, or interval, score [
9], which adds to the width a penalty of
times the distance by which a realization falls outside the interval. Lower is better, and a method is rewarded for narrowness only if it also contains the outcome. It is the metric on which the apparent 20% gain is claimed;
Section 2.5 gives the formula and the limitations that shape how the results should be read.
1.2. What the Literature Does and Does Not Check
Two lines of work relax exchangeability, and the conditions each imposes on the calibration weights are central to this study. The first re-weights the calibration scores. Weighted conformal prediction [
10] restores validity under covariate shift, and Barber et al. [
11] bound the coverage gap for arbitrary sequences by a weighted sum of total-variation distances—for weights that are fixed, chosen in advance and not functions of the data.
A fixed weight is not the same as a merely predictable weight, one measurable at the forecast origin. Both differ categorically from a weight the calibration outcomes themselves produce: a state inferred using the target residual, or a regime model fitted on the very residuals that serve as nonconformity scores. Weights of that third kind are neither fixed nor past-measurable, so the bound does not reach them. The same construction also demands a finite-sample correction: point masses at
added to the weighted calibration distribution, without which the guarantee fails at any finite sample size. Practitioners drop them easily, and including them makes intervals unbounded when the weights concentrate. A separate route sidesteps weighting entirely, restoring exact validity for hidden-Markov data by blocking [
12,
13], though for classification sets rather than intervals.
The second line adapts online. Adaptive Conformal Inference [
14] updates the miscoverage level from realized coverage errors, its dynamically-tuned extension aggregates several learning rates [
15], EnbPI [
16] recalibrates from a sliding window, and Lu et al. [
17] couple a deep switching state-space model with conformal inference. This family carries no explicit regime model yet survives drift on past miscoverage alone, so it sets the bar a regime-weighted construction must clear.
Table 1 places all of these constructions, and the two arms examined here, on the four dimensions this paper audits.
A third strand sharpens intervals by localizing calibration toward states resembling the forecast origin, and it is the one this paper audits. The idea predates conformal prediction: Zarnani et al. [
18] draw a new case’s interval from the error distribution of its fuzzy-clustered weather regime, with no coverage guarantee. Schmitt [
19] weights conformal risk-control scores by a time decay and a regime-similarity kernel over an observed volatility feature, and proves a weighted-exchangeability guarantee; RareCP [
20] retrieves similar calibration examples through a mixture of learned regime experts; and a parallel line refines the tail budget rather than the weighting [
21]. The closest instance in this journal, DG-TFT-CQR [
22], calibrates temporal-fusion-transformer quantiles with a single pooled, unweighted split-conformal step, and its authors note the absence of state-adaptive calibration as a limitation—the opening that regime-aware weighting seeks to exploit.
These constructions differ in how localization is obtained: an observed feature, cluster membership, or a latent volatility state inferred from a Markov-switching model, as studied here. None of them, however, examines systematically whether the localizing signal is available at the forecast origin, that is, filtered versus one-step predictive, nor whether the auxiliary model producing it is estimated on data disjoint from the calibration scores. These two questions are immaterial to the headline sharpness metrics yet decisive for admissibility. Two patterns in
Table 1 fix the scope of what follows: constructions with weights fixed in advance, and those restoring exact exchangeability by blocking, lie outside the diagnostic entirely, while it binds wherever a fitted auxiliary model supplies the signal—and in that group the timing of the signal and the provenance of the model are rarely stated explicitly.
1.3. Why Natural Rubber
Within commodity forecasting, latent regimes have a long modeling tradition. Markov-switching models [
23] formalize the alternation between calm and stress states, and regime dependence in commodity and energy price volatility is well documented [
24,
25]. GARCH and stochastic-volatility models describe the same heteroskedasticity through a continuous conditional variance; a discrete state is used here because the conformal weights need a similarity between the forecast origin and each calibration day, which a two-state probability vector supplies directly. The closest agricultural precedent is Kummaraka and Srisuradetchai [
26], whose Monte-Carlo-dropout intervals for Thai durian exports improve empirical coverage over ARIMA but carry no finite-sample guarantee.
Rubber itself has been forecast with econometric models [
27,
28], machine-learning regressors, and decomposition-augmented deep models [
29,
30,
31], while latent-regime forecasters [
17] tie the regime variable to the conditional mean. All of that work targets the point or volatility forecast. This study relocates the regime structure from the point layer to the uncertainty layer, where the calibration gap resides.
1.4. This Paper
The present study is diagnostic rather than constructive: it audits the regime-aware conformal family rather than adding a method to it. The construction it examines—a two-state Markov-switching regime driving a similarity-weighted, per-regime-asymmetric conformal interval—is a representative of that family, not a proposal.
Three results follow. A two-factor empirical design separates the timing of the regime signal from the reuse of calibration outcomes, and under the corrected construction the naive 20% sharpening disappears: the filtered, calibration-fitted cell falls to a paired Winkler difference of −0.22 with an interval spanning zero, and the forecast-feasible construction shows no detectable gain. A controlled Markov-regime simulation with an exact no-regime null then shows why. Replacing predictive with filtered probabilities manufactures a spurious advantage even where no regime information exists, while calibration-fitting adds none that is separately detectable under a stationary null. The controlled null attributes the manufactured gain to target-state leakage; calibration-dependent fitting remains associated with the empirical pattern, but its contribution cannot be separated from regime-model recency and transfer and is explicitly not test-data leakage. From these results the paper distils an audit protocol: information timing, calibration reuse, finite-sample correction and effective calibration size must be checked jointly.
Three limits belong at the outset. No finite-sample coverage guarantee is established for those weights; the fixed-weight correction is a benchmark, not a certificate. No published method is audited—the construction is assembled from components in the literature, and
Table 1 records where the diagnostic binds. And an interval spanning zero is not equivalence, since the sample sizes here would not resolve small effects.
Section 4.4 states each claim in full.
Section 2 describes the data, the base forecasters and the experimental design;
Section 3 reports the decomposition, its robustness and the simulation;
Section 4 interprets the two couplings and states the scope; and
Section 5 concludes with guidance for practitioners. The conclusions are limited to the implementations and regime conditions examined.
2. Data and Empirical Methodology
Figure 1 sets out the pipeline and the point at which its information properties are decided: a base forecast yields residuals, a Markov-switching model supplies the regime signal, and a regime-weighted conformal interval follows. The predictive, validation-fitted branch is forecast-feasible and keeps regime-parameter estimation separate from conformal calibration; the filtered branch is an oracle diagnostic when its origin state conditions on the target-period residual, and calibration-fitted regime parameters introduce calibration-dependent auxiliary estimation. These distinctions concern information availability and score reuse; formal coverage for the resulting data-dependent latent-state weights is treated separately in
Supplementary Section S2.
2.1. Data and Feature Construction
The target variable is the first-differenced price of Ribbed Smoked Sheet grade 3 (RSS3) rubber quoted free on board (FOB) in Bangkok,
(Baht/kg/day), where
is the near-month, first-position quote. The study window is 7 May 2018 to 13 July 2026, daily, and yields 1906 quoted trading days. Differencing on consecutive real observations leaves 1430 usable target observations.
Figure 2 plots both series against the partitions.
Figure 2.
The RSS3 FOB near-month price (
upper panel) and its first difference (
lower panel), 7 May 2018–13 July 2026, with the five chronological partitions of
Table 2 shaded alternately; “Val.” is the validation block and “BW” the bandwidth block. The level series drifts and shifts scale; the increments are centered but visibly heteroskedastic. The single largest decline of the sample,
Baht/kg on 26 June 2026, falls inside the test window and accounts for most of that window’s skew and excess kurtosis.
Figure 2.
The RSS3 FOB near-month price (
upper panel) and its first difference (
lower panel), 7 May 2018–13 July 2026, with the five chronological partitions of
Table 2 shaded alternately; “Val.” is the validation block and “BW” the bandwidth block. The level series drifts and shifts scale; the increments are centered but visibly heteroskedastic. The single largest decline of the sample,
Baht/kg on 26 June 2026, falls inside the test window and accounts for most of that window’s skew and excess kurtosis.
Table 2.
Chronological partitions of the sample, their spans and their roles.
Table 2.
Chronological partitions of the sample, their spans and their roles.
| Partition | Span | Role | n |
|---|
| Training | 7 May 2018–3 Jan 2023 | base-model estimation | 841 |
| Validation | 4 Jan–29 Sep 2023 | ridge penalty; regime fit in the forecast-feasible arm | 129 |
| Calibration | 3 Oct 2023–27 Jun 2025 | conformal residuals | 287 |
| Bandwidth (BW) | 1 Jul–30 Sep 2025 | selection of | 42 |
| Test | 1 Oct 2025–8 Jul 2026 | final evaluation, never used for any choice | 131 |
Three features of the data motivate the construction studied here. First, the level series is non-stationary and the differenced series is not. Augmented Dickey–Fuller and KPSS tests agree on both counts and at every conventional level, with the ADF statistic on
reaching
(
); Panel C of
Table 3 reports each test. The increment, rather than the level, is also the quantity whose uncertainty a one-step interval must express.
Second, the increments are far from Gaussian, and their departure is concentrated in the tails: excess kurtosis is 7.12 over the full sample and 10.83 over the test window (
Table 3, Panels A and B). One day accounts for almost all of the latter—excluding the
Baht/kg decline of 26 June 2026 returns the window to a skew of
and an excess kurtosis of 0.68—and it matters in advance, because a single exceedance enters the Winkler mean with a weight of twenty at
.
Third, and most directly relevant to regime weighting, the volatility of the increments is strongly clustered and shifts across the sample. Ljung–Box and ARCH-LM tests on
both reject at
, and the standard deviation of
ranges from 0.377 in the validation block to 0.996 in the calibration block (
Table 3, Panels B and C). An undifferentiated split-conformal interval cannot accommodate this heteroskedasticity, which is exactly what regime weighting is designed to exploit. The increments also carry first-order autocorrelation of 0.389 over the full sample and 0.325 in the test window, and 11.4% of days record no price change at all, reflecting the partly administered character of the FOB quote.
Because that one observation accounts for most of the higher moments, the dependence structure is reported in rank as well as linear form (
Table 3, Panel D), and the two cut in opposite directions. The persistence of
is real rather than an artefact: its rank autocorrelation stays between 0.24 and 0.36 out to five lags, significant at every one, so the volatility clustering that motivates regime weighting survives a statistic the outlier cannot dominate.
The linear statistic nonetheless overstates the first lag, 0.463 against a rank value of 0.364, and in the conditional mean the leverage runs the other way: the Pearson autocorrelation turns negative at lags four and five where the rank statistic stays near zero, so that apparent mean-reversion is a property of one day rather than of the series.
Section 3.4 returns to what the base forecaster does and does not recover from this structure.
All series are assembled into a single daily panel from primary sources: Thai physical-rubber prices from the Rubber Authority of Thailand [
32] and the Thai Rubber Association, exchange settlements from SHFE, TOCOM and SGX, Bank of Thailand reference rates, Brent and WTI crude, and four Chinese demand indicators.
Table 4 lists every series with its source, native frequency and alignment.
Monthly indicators are release-lagged: a value stamped in month m becomes effective on the first trading day of month and is held constant within it. Because several Chinese indicators are released mid-month and later revised, this uniform rule is a conservative approximation rather than an exact release calendar, and a series-specific audit of release dates and vintages remains a data-provenance limitation.
From the raw panel, 64 features are constructed under a single documented formula set: log returns of each price and exchange-rate series, trailing realized volatilities over 5, 10 and 22 trading days, cross-market spreads and futures–spot bases converted to Baht/kg, term-structure slopes, 252-day rolling z-scores, the Chinese demand signals, and calendar dummies for the wintering season and the COVID-19 period. Returns are computed only across consecutive real observations, so no artificial zero-return enters across market closures, and a machine-readable data dictionary records every formula.
Because the argument turns on what is knowable at a forecast origin, that provenance is part of the evidence rather than a technical appendix.
Table 5 collects the symbols used throughout, rather than introducing them piecemeal.
2.2. Base Point-Forecast Model
The conformal layer calibrates on whatever residuals a base forecaster leaves behind, so the diagnostic does not depend on which forecaster produces them. Two complementary forecasters of deliberately different complexity are used here, and the central decomposition is reported on both, so its conclusion does not depend on either. The simpler is a single-parameter autoregression,
which uses no covariates at all and attains a test mean absolute scaled error of 0.921. The richer is a ridge-regularized linear regression on 64 lagged economic features augmented with Variational Mode Decomposition components, following Pinitjitsamut [
31], which reaches 1.00 on the same window. Neither is offered as the recommended forecaster; reporting both is what makes the diagnostic separable from the choice of base model.
The regression stage that maps those features to a forecast is deliberately linear, chosen for transparency rather than point accuracy. Were the conditional mean the object of study, a more expressive model would be the tool of choice: the VMD-augmented bidirectional LSTM of Pinitjitsamut [
31], on the same target, attains substantially higher point accuracy while preserving the series’ dispersion. Here, a simple, well-understood regression keeps the regime-weighting question free of base-model tuning. It still produces residuals with realistic dispersion, at a predicted-to-realized ratio of roughly 0.5–0.9 and a validation-partition correlation near 0.41. That this forecast sits close to a no-change benchmark on the test window (
Section 3.4) is a property of the base model rather than of the conformal analysis, which turns on the regime structure of the residuals.
The ridge penalty is selected from
by validation mean-squared error, and the regime-similarity bandwidth
from
by the coverage–width criterion on the bandwidth block. Neither uses the test partition. The moving-block bootstrap is not a tuning choice but the inference tool used throughout;
Section 2.5 gives its settings. Its penalty selection on a validation partition disjoint from the calibration set prevents model-selection leakage into the calibration residuals, though it does not restore the exchangeability that dependent residuals lack (
Section 2.5,
Supplementary Section S2).
Because VMD is a global operator, applying it to the whole sample would leak future information; the decomposition therefore uses an expanding window, recomputed at every origin and lagged one trading day together with all 64 economic features.
Supplementary Section S1 sets out that treatment, and
Table A1 the VMD settings and every other reproducibility parameter, together with a check in which the VMD components are removed entirely—which leaves the decomposition of
Section 3.1 unchanged.
2.3. Regime Model and the Two Design Factors
This subsection defines the two experimental factors the paper varies. Both govern how much information the similarity weights may see: the first fixes
when the regime signal is measured, the second
which data the model producing it was estimated on. Every result in
Section 3 is a cell of the grid they define.
Volatility regimes are inferred from the base-model residuals by a two-state Gaussian Markov-switching model fitted to the log-squared residual. Writing
for that feature, with
guarding against a residual of exactly zero, and
for the latent state,
so each state carries its own mean and variance in
, and the transition matrix
P propagates the state one step ahead. Two design choices determine the information content of the weights this signal produces—analyzed in
Supplementary Section S2—and the study varies both.
- (i)
State.
Two probabilities are available. The filtered probability,
, conditions on the target day and therefore encodes its shock. The one-step predictive probability,
, conditions only on the past. The distinction is one of availability:
is known before the period-
t residual is observed, whereas
is not, so using
to build an interval for
is an oracle operation rather than a deployable forecast. Predictive timing does not, however, make the resulting similarity weights fixed, nor establish the coverage conditions of Barber et al. [
11], because the weight vector remains a function of the historical score path (
Supplementary Section S2.1). The filtered probability is retained only as an oracle comparison arm.
The Markov-switching parameters may be estimated on the calibration window, on a disjoint validation block, or on both. The first option is delicate because the calibration window supplies the conformal scores themselves. Estimating on it is operationally feasible at the final origin, since those observations are historical, but it makes the auxiliary regime model depend on the same scores later used for calibration. This calibration-dependent estimation, or score reuse, falls outside the fixed-weight analysis, and it may improve apparent localization through calibration-period adaptation. Estimation on a disjoint validation block removes that particular dependence. The recursively generated state probabilities remain random. The calibration-informed variants are retained as comparison arms.
Crossing the two factors yields the six regime-weighted constructions of
Section 3, which correspond to the forecast-feasible and coupled branches of
Figure 1. One cell is distinguished: the predictive-state, validation-fitted construction removes the two most direct couplings and is referred to throughout as the forecast-feasible construction. The term is deliberately narrow.
It records only that every quantity entering the interval is measurable at the forecast origin and that the regime model is estimated on a block disjoint from the calibration scores; it asserts neither causal inference nor finite-sample validity, since the similarity weights remain data-dependent functions of the realized score path (
Supplementary Section S2). The two states are named once and used consistently thereafter: the higher-mean state in
is the stress regime and the lower-mean state the calm regime. The naming is relative to each fitted model, and on this series the stress probability exceeds one-half on roughly 70% of sample days—a property of the fitted chain rather than a claim about turbulent days.
Table 6 reports the estimates, because their instability across blocks is itself part of the evidence. Expected stress durations range from 4.7 days when the model is fitted on validation to 20.0 days when fitted on validation and calibration together, and the mean stress probability over the test window ranges from 0.086 to 0.671. A regime model estimated on one block is therefore not the same object as one estimated on another—which is why the estimation sample enters as an experimental factor rather than an implementation detail, and why the transfer failure of
Section 3.5 is unsurprising in retrospect.
Panel B reports how well each signal tracks a volatility class the model cannot influence, and the two fitting blocks separate cleanly: the calibration-fitted signals reach an AUC of 0.63 to 0.64, the validation-fitted ones—those the forecast-feasible construction actually uses—only 0.54 to 0.57, with intervals including 0.50. In both cases, the predictive signal is no worse than the filtered one against next-day volatility, the quantity a one-step interval must anticipate. The signal is therefore weak where it is deployable, and its timing buys nothing for the target that matters.
2.4. Regime-Weighted Conformal Construction
At a forecast origin
, each calibration residual
i receives a weight reflecting how closely its regime vector
resembles that of the origin,
. The kernel is exponential in the
distance between the two,
and the weights are normalized over the calibration set together with the test point,
Here,
g is the chosen regime signal: the predictive
in the forecast-feasible construction, the filtered
in the oracle comparison arm. A larger
concentrates the weight on calibration days whose regime vector is closest to the origin’s;
returns uniform weights and hence ordinary split-conformal.
The interval bounds are the weighted lower- and upper-tail quantiles of the calibration residuals, added to the point forecast. Writing
for a point mass at
x, and augmenting the weighted calibration distribution with the
atoms that the weighted conformal quantile requires, the bounds are
where
is the
ith calibration residual. The miscoverage budget is split between the tails by regime,
, so that the heavier tail receives the smaller miscoverage level and hence the wider bound. Let
and
denote the weighted lower- and upper-tail spreads of the regime-
r residuals about their weighted median. The tail levels are set in inverse proportion to the spread each side carries,
which recovers symmetry when
.
Estimated from the calibration-tail spreads, these levels are data-dependent and fall outside the fixed-level coverage argument. The fixed symmetric allocation
is therefore the primary choice, and the asymmetric allocation is a sensitivity analysis. The bandwidth
is selected on the held-out bandwidth partition by a coverage–width criterion, separately for each construction. Algorithm 1 summarizes the procedure. It is applied identically across all factor cells, with only the regime signal
g and its estimation sample varying.
| Algorithm 1 Regime-weighted interval at forecast origin |
Input: calibration residuals with regime vectors ; origin regime vector ; point forecast ; level ; bandwidth . ( predictive in the forecast-feasible construction, filtered in the oracle comparison arm.)
Similarity weights. Form and by Equations ( 3) and ( 4). Tail budget. Identify the origin’s dominant regime r from ; from the weighted tail spreads set with as above. Weighted quantiles. Augment the calibration residuals with point masses at and carrying weight (the finite-sample correction of Barber et al. [ 11]). Compute , the weighted -quantile of the augmented set (with the atom), and , the weighted -quantile (with the atom). The primary specification uses the fixed symmetric levels ; the per-regime asymmetric levels of step 2 are a sensitivity variant. Interval. Return .
|
2.5. Evaluation Protocol and Experimental Design
The five chronological blocks of
Table 2 each serve one purpose, and the separation between them is what the diagnostic of this paper turns on. The base model is estimated on the training block; the validation block selects the ridge penalty and, in the forecast-feasible construction, estimates the regime parameters; the calibration block supplies the conformal residuals; the bandwidth block selects
; and the test block is held out.
Intervals target
(90% coverage). Four measures assess performance. The first is marginal coverage, reported with moving-block-bootstrap confidence intervals [
33] to account for serial dependence in the coverage indicators, the block resampling being what preserves the short-run dependence an i.i.d. bootstrap would destroy (block length
, 2000 resamples, percentile intervals, seed 0). The second is regime-conditional coverage, evaluated on a high-volatility mask that no construction can influence—the highest third of test-window days ranked by realized volatility—so that no construction is scored on its own regime signal. The third is mean interval width. The fourth is the Winkler (interval) score, which for a
interval
and realized
y equals
, with lower values sharper and better calibrated. Paired differences in width and Winkler score against split-conformal are reported with moving-block-bootstrap intervals.
The Winkler score is proper and standard in the interval-forecasting literature. Three limitations shape how the results below should be read; Zhou et al. [
34] survey them for load forecasting. First, the score implicitly favors unimodal predictive behavior and is ill-suited to heavy-tailed predictive distributions—a caution that matters for a target whose increments carry an excess kurtosis above seven (
Table 3). Second, it emphasizes average performance, so a favorable mean score is compatible with poor behavior in exactly the rare, high-impact episodes where an interval matters most. Third, it is not comparable across datasets of differing predictability, which is why the comparisons here are always paired against split-conformal on the same origins rather than quoted as absolute levels.
Two further properties follow from the definition. The score compresses calibration and sharpness into one number, so two methods can attain the same value for opposite reasons—one narrow and under-covering, the other wide and well-covered. And the exceedance penalty is large: at it multiplies each miss by twenty, so a handful of exceedances dominates the mean, the sampling distribution is heavy-tailed, and paired comparisons on a short test window have limited power. A fifth point is specific to this study: the score is undefined whenever an interval is unbounded, a case the finite-sample correction makes real.
None of these is hypothetical here: each surfaces in
Section 3.1,
Section 4.4 and
Section 3.6, respectively. Coverage, width, and the Winkler score are therefore reported jointly throughout, each with moving-block-bootstrap intervals, and no conclusion rests on the interval score alone.
The experimental design has five parts. A decomposition grid crosses split-conformal with the six regime-weighted constructions of
Section 2.3 on the test window, isolating each factor. The entire pipeline is then re-estimated at six rolling origins. The forecast-feasible construction is benchmarked against Adaptive Conformal Inference and its dynamically-tuned aggregate on a common rolling protocol. Point-forecast metrics locate the result in the forecastability of the residuals, and the grid is reported on both base forecasters, so that the null is not read off a single one. A controlled simulation (
Section 2.6 and
Section 3.9) supplies ground-truth evidence.
2.6. Controlled Simulation Design
The application shows that the apparent gain disappears but cannot show what a genuine gain would have looked like, because the regime structure of the RSS3 residuals is whatever it is. A simulation supplies that counterfactual: the construction of
Section 2.4 runs on synthetic scores whose predictability the design fixes in advance. Its parameters span that lever rather than replicate the rubber series, so the dispersion and shape of the simulated states deliberately differ from the empirical ones—a controlled experiment, not a resimulation.
Data-generating process. Scores are generated by a two-state Markov chain with symmetric persistence . Conditional on the state, the residual is normal in the calm state (standard deviation ) and skew-normal in the stress state (scale , shape ). Each state is centered to a zero conditional mean, so the regimes differ only in dispersion and shape, and no conditional-mean shift is confounded with the volatility regime. The skew-normal scale is not its standard deviation. At , the stress state has an actual standard deviation of about 2.3, so the state SD ratio is roughly 2.3 rather than 3.5. Its skewness is about , a mild negative skew in keeping with commodity down-moves.
Construction and partitions. The construction builds intervals with the finite-sample
correction and fixed symmetric tails of Algorithm 1, identically to the empirical analysis. The series is partitioned chronologically into validation, calibration, bandwidth, and test blocks as in
Table 2. One departure from the empirical specification must be recorded.
The Markov-switching model is fit to a causal local log-realized-variance, the trailing-window mean of the squared residual, because the one-point log-chi-square noise of a single residual prevents the filter from recovering even a strongly persistent state. That smoothing is a stated modeling choice, not a result, and it makes the latent state easier to recover than in the application, which fits the model to the log-squared residual (
Appendix A).
Section 3.9 quantifies the gap it creates.
Experimental lever. The design sweeps one lever, the persistence p, which governs how one-step-predictable the regime is. The separability , which governs how distinct the regimes are, is held fixed at 3.5. This is therefore a persistence experiment rather than a two-lever sweep, and a sweep of the separability lever is left to future work.
At each persistence, two constructions are compared with split-conformal by the paired Winkler difference: the forecast-feasible one, with predictive state and regime fit on validation, and the artifact one, with filtered state and regime fit on calibration. Differences are averaged over 250 Monte Carlo replications for the persistence experiment, 300 for the score-separation experiment, and 200 for the no-regime null, each with 95% Monte Carlo intervals, and no replication is discarded. Unbounded origins are retained for the unbounded-rate calculation but excluded from the conditional score, and every paired comparison is taken over the constructions’ common finite-origin set.
The no-regime null. To confirm that the pipeline does not itself manufacture an advantage, an exact no-regime null runs over 200 replications. Its two states share an identical emission, so the regime label carries no distributional information. The null runs as a full factorial, crossing state timing (filtered versus one-step predictive) with the regime-estimation sample (calibration-fit versus disjoint validation-fit). This lets the state-timing and estimation contributions be read off as direct paired contrasts.
Under the null, predictable weighting carries no population-level regime information and should show no systematic expected advantage over split-conformal, though finite-sample differences may occur. Two limits follow. Stationarity leaves the estimation-sample coupling little to transfer, so the design isolates the state coupling and does not reproduce the non-stationarity that activates the other in the application (
Section 3.1). And the score-separated arm, which fits the regime model on the first half of the calibration block and forms scores on the second, halves the effective calibration size to 150 of 300 scores, so any change it produces confounds score separation with a smaller calibration set.
3. Results
The evidence is presented in eight steps.
Section 3.1 exhibits the apparent advantage of regime weighting and examines how it varies with the two information couplings.
Section 3.2 shows the resulting null to be reproduced under a limited six-origin rolling-origin re-estimation,
Section 3.3 places the forecast-feasible construction against adaptive-conformal baselines, and
Section 3.4 reports the point-forecast context that locates the null in the residuals the conformal layer receives.
Section 3.5 reports the same decomposition on the AR(1) residuals, and
Section 3.6 evaluates two operational fallback rules for the unbounded intervals the finite-sample correction can produce.
Section 3.7 gathers the robustness evidence: the full bandwidth grid with weight-concentration diagnostics, a benchmark of the latent state against an observed volatility feature, and variation of the forecast origins, the nominal level, the benchmark family and the number of regime states.
Section 3.9 corroborates the mechanism in a controlled simulation with a known regime. Throughout this section, a negative ΔWinkler denotes an improvement over split-conformal, and a bootstrap interval that spans zero denotes no statistically detectable difference from it. Results are reported on the ridge residuals of
Section 2.2 unless the AR(1) residuals are named;
Section 3.5 gives the same decomposition on the latter. All constructions target 90% coverage (
). Single-window results use the test partition (
), and regime-conditional coverage is reported on the high-volatility mask defined in
Section 2.5.
Two qualifications attach to that mask. It is an ex-post evaluation device defined on realized test-window volatility, not a deployable regime classifier available at the forecast origin. And with only 44 days it is severely underpowered: a 90% target implies roughly four to five expected misses, so coverage estimates carry sampling uncertainty of order before serial dependence is considered. The high-volatility figures are therefore descriptive coverage on an ex-post subset, not evidence of conditional validity.
3.1. An Apparent Advantage and Its Decomposition
The first question is whether the headline gain belongs to regime weighting itself or to the information the weights are allowed to see. The two are separable only if the design choices that distinguish a regime-weighted interval from split-conformal are varied one at a time, rather than adopted together as a package.
Table 7 crosses those two choices. The first is the regime signal: the filtered probability, which conditions on the target day, or the one-step predictive probability, which conditions only on the past. The second is the estimation sample for the Markov-switching parameters: the calibration window, a disjoint validation block, or both.
Consider first the naive implementation: the plain weighted quantile without the finite-sample correction, using the filtered state and a regime model estimated on the calibration window. At the bandwidth this criterion selects,
, its
mean interval width is 2.434 Baht/kg against 3.047 for the corresponding split-conformal baseline, a reduction of 20.1%. The reduction is in width, not in coverage: coverage for every construction is reported separately in
Table 7 and changes by at most 0.02. The naive arm also attains a paired Winkler difference of
[
,
], excluding zero (Table S2, first row). Note that the naive arm uses the interpolating weighted quantile throughout, so its split-conformal baseline (3.047) differs from the order-statistic baseline of
Table 7 (3.118); the two are never mixed within a comparison. Reported in isolation, this reads as a clear methodological success, and it is the number such implementations report.
The naive figure does not survive the corrected specification. The correction does not merely widen the interval at a given bandwidth; it removes the bandwidths at which the apparent gain lives. Without the correction, no interval is ever unbounded, so the coverage–width criterion is free to select a highly concentrated kernel; the naive figure above is obtained at . Under the correction, that bandwidth is inadmissible: it leaves 31 of the 131 test origins unbounded, and already leaves 8. The selection is therefore forced down to , where the paired difference is rather than . Holding the bandwidth fixed at separates the two effects: the uncorrected construction gains there and the corrected one , so roughly half of the headline advantage is the atom and the remainder is the concentration the atom forbids.
The intervals of
Table 7 are accordingly built with the finite-sample
correction that the weighted-conformal quantile requires (
Supplementary Section S2.1) and with fixed symmetric tails. Under this specification, the coupled cell falls to a paired Winkler difference of
, with a bootstrap interval of [
,
] that spans zero. Its width, 3.189, is no longer distinguishable from split-conformal at 3.118.
Supplementary Section S3 reports the full six-cell grid with and without the correction. The finite-sample correction widens most exactly the cells whose weights are most concentrated—the coupled cells—because their effective sample is smallest; the apparent narrowing was therefore in large part a small-sample artifact, not only an information coupling.
Under the finite-sample-corrected specification, no cell significantly improves on split-conformal, and the two couplings remain separately identifiable in what survives. Replacing the filtered state with the predictive state is the only choice that removes the direct dependence on the target-day shock. It leaves a paired interval-score difference of [, ] for the calibration-fitted regime, and [, ] for the forecast-feasible construction. Both bootstrap intervals span zero. Independently, retaining the filtered state but estimating the regime out of sample, on the validation block alone, yields [, ], again with no statistically detectable difference from split-conformal. The forecast-feasible construction—predictive state, out-of-sample estimation—shows no statistically detectable difference from split-conformal in interval score; its mean width is 2.990 against 3.118 for split-conformal.
The intervals might still adapt in the right direction, widening on turbulent days and narrowing on calm ones, even where the average width is unchanged. They do not. On the highest third of days by realized volatility, the realized absolute residual is 0.964 against 0.513 on the remaining days, a ratio of 1.88. Split-conformal, which cannot adapt at all, has a width ratio of exactly 1.00 by construction. The filtered, calibration-fitted cell reaches only 1.10 (3.39 against 3.09), despite conditioning on the target day’s own shock. The forecast-feasible construction reaches 0.97 (2.93 against 3.02), which is to say it is marginally narrower on the volatile days. Regime weighting on this series therefore neither sharpens on average nor redistributes width towards the days where the risk lies.
Figure 3 displays the six cells with and without the correction. They locate the apparent advantage outside the regime-weighting idea itself: what narrows the interval is the information the weights may use, not the act of weighting. Whether that conclusion is an artifact of one arbitrary calibration–test split is the question
Section 3.2 takes up.
3.2. Robustness Under Rolling-Origin Re-Estimation
A single test window can flatter or penalize any method by accident of where it falls in the sample. The decomposition therefore repeats at six rolling origins, with every component re-estimated from scratch—expanding-window VMD, ridge with penalty selected on validation, the Markov-switching regime model, and the conformal calibration—comparing the three headline constructions per origin.
Across all six origins, the forecast-feasible construction tracks split-conformal in interval score, at a mean of 4.60 against 4.99 over the five finite origins, while the coupled filtered-and-calibration construction stays ahead at 4.07. Under the finite-sample correction, however, both weighted constructions return unbounded intervals at one of the six origins (
Table 8). Over the five origins at which every method stays finite, the ordering matches the single-window analysis. The generally low coverage at the earliest, most volatile origins (0.50–0.70 for every construction) is itself informative: the series is non-stationary enough that split-conformal itself under-covers materially in places, and regime weighting does not repair this.
The rolling evidence therefore reproduces the single-window null and adds a failure mode of its own: once correctly constructed, the regime-weighted interval is not reliably finite. That makes a comparison with methods which are always finite, and which adapt without any regime model at all, the natural next step.
3.3. Comparison with Adaptive Baselines
Failing to beat the simpler adaptive alternatives a practitioner would otherwise reach for is the more demanding test. Those alternatives carry no regime model and adapt only from realized miscoverage, so they set the bar the additional regime structure must clear. The benchmark asks whether the forecast-feasible construction improves on them: Adaptive Conformal Inference (ACI) and its dynamically-tuned aggregate (DtACI), evaluated on a common rolling protocol.
DtACI and ACI attain finite mean interval scores of 3.83 and 4.16, respectively, compared with 4.58 for rolling split-conformal. The finite-sample-corrected regime-weighted interval has no finite mean score in this protocol: under the finite-sample correction, it returns unbounded intervals at 8 of 131 origins (
Table 9). A direct comparison of mean interval scores is therefore not meaningful; ACI and DtACI remain finite throughout and offer substantially greater operational reliability. Concentrated weights drive this. Over the single test window, the weight sum stays above the threshold that
Section 3.6 states, with a minimum of 19.4 and an effective sample size that falls from 213 to 29 at the most concentrated origins. The shorter rolling windows breach it.
The additional structure thus fails to clear the bar and forfeits a property both baselines retain, namely an interval that is always finite. What remains to be explained is why the regime signal has so little to contribute to this series; the answer lies one layer down, in the point forecast.
3.4. Point-Forecast Context
The preceding subsections establish that regime weighting does not improve on split-conformal here, but not why. Because the conformal layer calibrates on the base model’s residuals, the answer lies in how much predictable structure that model leaves behind. On the test partition, the VMD-augmented ridge model attains a mean absolute error of 0.665 against 0.662 for a no-change benchmark—a mean absolute scaled error of 1.00—with a predicted-to-realized dispersion ratio of 0.37. The ridge regression therefore performs no better than a naive no-change forecast on this window.
That performance is a property of the fitted model rather than of the target. The differenced series is not a martingale. A moving-average specification is the natural alternative for a differenced price. An MA(1) fitted on the same expanding pre-test window attains a test mean absolute error of 0.613, against 0.602 for the AR(1) and 0.662 for the no-change benchmark—a mean absolute scaled error of 0.93 against 0.91. The autoregressive form is therefore retained as the simpler complementary forecaster, though the margin between the two is small enough that neither choice would change what the conformal layer receives. Its first-order autocorrelation is 0.389 over the full sample and 0.325 over the test window (
Table 3). A single-parameter AR(1), fitted on an expanding pre-test window, attains a test mean absolute error of 0.602 against 0.662 for the no-change benchmark, a mean absolute scaled error of 0.91. The ridge regression with 64 economic features and six VMD modes does not recover that dependence. Its residuals retain a rank autocorrelation of 0.276 at the first lag (
,
), against 0.388 for
itself (
Table 3, Panel D): under a third of the first-order dependence is removed, and none is left at lags two to five. The residuals passed to the conformal layer therefore stay close to the raw price increments, whereas an AR(1) removes that first-order dependence by construction.
Two consequences follow. The first leaves the diagnostic result untouched. The couplings of
Section 3.1 operate on how residuals are weighted by regime, not on how large the residuals are, and the controlled null of
Section 3.9 exhibits them with no base model at all. The second bounds the null. The empirical analysis establishes that the residuals of
this base model carry too little one-step-predictable volatility structure for regime weighting to exploit. That is why the forecast-feasible construction shows no detectable improvement on split-conformal, and why the filtered construction can appear to, by conditioning on the contemporaneous shock. Whether a base forecaster that does recover the conditional-mean dependence would leave exploitable regime structure in its residuals is the question
Section 3.5 takes up directly.
3.5. The Decomposition on the AR(1) Residuals
Section 3.1 reported the decomposition on the ridge residuals.
Table 10 reports the same decomposition on the AR(1) residuals of Equation (
1), which beat the ridge regression on point accuracy (test MASE 0.921 against 1.00). The comparison answers a natural objection—that the null is an artefact of a weak base forecaster—and it also shows where the two base models genuinely differ. The conformal layer, the partitions, the regime specification, the bandwidth grid, and the evaluation protocol are unchanged; only the residuals differ.
The pattern of
Table 7 is reproduced in full. In its naive form, the coupled construction attains a paired Winkler difference of
[
,
], whose bootstrap interval excludes zero—a headline gain of the same order as the one the VMD-augmented base produces. Every gain again belongs to a filtered cell with calibration-informed regime estimation. Coverage on the highest third of days by realized volatility is 0.841 for split-conformal and 0.864 for every regime-weighted cell except the forecast-feasible construction, which matches split-conformal exactly. Replacing the filtered state with the predictive state eliminates it at each estimation sample, and the forecast-feasible construction shows no statistically detectable improvement.
Two differences from
Section 3.1 both cut against the method. First, the finite-sample correction is a weaker defense here than for the ridge residuals. It reduces the coupled cell from
to
, but the bootstrap interval still excludes zero, because the selected bandwidth leaves the calibration weights less concentrated and the
atoms therefore bind at no origin. Applying the correction is thus necessary but not sufficient: on this base model only the removal of the filtered state disposes of the artifact.
Second, the forecast-feasible construction is numerically identical to split-conformal, not merely statistically indistinguishable from it. The regime model is estimated here on the validation block, the calmest of the five (
Table 3), and it labels almost the whole of the later sample as stress. Its predictive stress probability over the test window has a mean of 0.789 and a standard deviation of only 0.084, against 0.608 and 0.156 for the calibration-fitted model. The origin regime vector therefore barely moves from one day to the next, the similarity kernel is effectively flat at the selected bandwidth, and the weighted quantile selects the same order statistics as the unweighted one.
Figure 4 shows what the coupled construction does instead. That is the transfer failure of
Section 4 in its starkest form—an out-of-sample regime model that carries no usable information into the test period collapses the construction back onto the baseline it was meant to improve.
The base-model objection is therefore answered for the direction that matters. A base forecaster that does recover the conditional-mean dependence leaves the diagnostic conclusion intact: the apparent gain still belongs to information timing, and remains associated with calibration-dependent estimation, rather than to regime weighting as such. What a stronger base model changes is the magnitude of the artifact and the adequacy of the finite-sample correction against it, both in the unfavorable direction.
3.6. An Operational Fallback for Unbounded Intervals
The finite-sample correction is not optional, but it carries an operational cost that
Section 3.2 and
Section 3.3 record without resolving: when the calibration weights concentrate, the
atoms bind and the interval becomes unbounded. Such an interval is unusable, and a forecasting system that emits one has no answer to give.
The trigger is known in advance and costs nothing to evaluate. An interval is unbounded exactly when the calibration weights concentrate enough that
, which is nineteen at
; the sum is available at the origin, before the interval is formed. The remedy that follows is to abandon regime weighting at those origins and return the unweighted split-conformal interval.
Supplementary Section S3 evaluates that rule against the alternative of weakening the kernel instead, on a rolling protocol swept to drive the weights into concentration. Falling back to split-conformal dominates at every bandwidth, eliminates unbounded intervals entirely, and—this is the finding that matters for the paper’s thesis—confers no detectable advantage in exchange: against rolling split-conformal the paired Winkler difference spans zero at every bandwidth examined. The rule is deterministic, costs one summation, and is auditable after the fact from the fired-origin rate.
One caution attaches. Applied to the filtered state rather than the predictive one, the same hybrid appears to beat rolling split-conformal. That apparent gain is the information-timing artifact of
Section 3.1 operating inside the fallback rule: a fallback protocol inherits whatever couplings the construction it repairs already contains, and it repairs unboundedness only.
3.7. Robustness: Bandwidth, Signals, Origins, Levels and Regime Cardinality
Three questions remain open. How is the similarity bandwidth chosen, and how much do the results depend on it? How concentrated do the calibration weights become? And does the artifact belong to
latent state weighting specifically, or to any localizing signal that sees the target period?
Table S4 sweeps the full bandwidth grid to answer all three.
The selection rule governs which row of each block is reported elsewhere. On the bandwidth partition, the criterion minimizes
where
is the mean interval width and
the empirical coverage on that block: mean width with a steep penalty on any shortfall below nominal coverage, subject to the constraint that no bandwidth-block interval is unbounded, with ties broken toward the smaller
. Selection is made separately for each construction and never re-estimated within the test window. It returns
for the filtered, calibration-fitted cell and
for the forecast-feasible construction; the remaining cells are listed in
Table S4.
First, the unbounded rate is the operative constraint on the bandwidth. For every signal, the interval score improves as rises and the weights concentrate—and the region where it improves most is the region where the construction stops returning finite intervals. The coupled arm reaches its largest paired difference at , where 21.4% of origins are unbounded. Scoring only the finite origins would credit the method for exactly the configurations in which it fails, so no paired difference is quoted where any origin is infinite.
Second, the ranking is stable across the admissible range rather than an artifact of the selected bandwidth: wherever intervals remain finite, the forecast-feasible arm sits at or slightly above split-conformal at every , and the coupled arm below it, and no bootstrap interval excludes zero once the correction is in force.
Third, the artifact does not belong to latent-state weighting specifically. A parallel grid in
Supplementary Section S3 replaces the latent state with a standardized trailing realized volatility of the residuals, an observed feature of the kind used by Schmitt [
19], and varies only its timing. Lagged one day, and so measurable at the forecast origin, it produces no detectable gain at any bandwidth: paired differences run from
to
and every interval spans zero. Computed on the same day, it produces a detectable gain wherever intervals remain finite, reaching
[
,
] at
. The two blocks differ by one lag and nothing else.
The lag contrast settles a question the two-factor design of
Section 3.1 could not. What manufactures the apparent sharpening is the information timing of the localizing signal, not the latent character of the state that carries it. A construction weighting on an observed, pre-specified feature is therefore exempt only if that feature is also past-measurable, a scope
Section 4.4 states in full.
3.8. Origins, Nominal Levels, Benchmark Families and Regime Cardinality
The null of
Section 3.1 rests on one test window, one nominal level, one benchmark family, and a two-state regime specification, any of which could be carrying the result.
Table 11 varies all four.
Panel A extends the rolling evidence from six origins to 314 and adds two further nominal levels, and the pattern is unchanged. The coupled arm improves on split-conformal at every level, with its interval excluding zero wherever a paired difference can be computed; the forecast-feasible arm does not, being worse at and returning unbounded intervals at 10.2% and 1.6% of origins at the other two levels, so its unconditional interval score is undefined. Split-conformal itself under-covers at every level over this longer sequence, reaching 0.873 against a nominal 0.90 at —a property of the series rather than of any construction.
Panel B addresses whether recency alone accounts for the adaptivity that helps on this series. It does not. Rolling-window split-conformal, exponentially time-decayed conformal weighting at two decay rates, and an EnbPI-style bootstrap all attain a higher interval score than the fixed split-conformal baseline, with paired differences from
to
and every interval spanning zero. Taken with
Section 3.3, where methods driven by realized miscoverage do improve on the regime-weighted construction, this narrows the earlier reading: what helps here is error feedback, not recency as such.
Panel C answers whether a two-state specification is too coarse. Adding states reveals no structure a two-state model missed; it enlarges the artifact and leaves the deployable construction untouched. The filtered, calibration-fitted arm improves from
at
to
at
and
at
, with intervals excluding zero at both higher cardinalities, while the forecast-feasible arm stays at
,
and
, with every interval spanning zero. A richer state space gives the filtered probability more room to encode the target day’s own shock and the predictive probability nothing further to forecast with, so the audit protocol of
Section 5.2 applies unchanged as
K grows.
Bootstrap inference is also insensitive to the block length. For the filtered, calibration-fitted cell on the test window, the paired difference is at every block length, and the 95% interval runs from [, ] at length one to [, ] at length twenty-two. The conclusion drawn from it does not turn on the square-root rule.
3.9. Simulation: When Regime Weighting Can and Cannot Help
The controlled simulation of
Section 2.6 supplies the counterfactual the application cannot, since the design fixes how predictable the regime is. A falsification test comes first. Under its exact no-regime null, the regime label is pure noise, so no construction respecting the predictable-weight condition should show a systematic advantage over split-conformal.
Table 12 reports the exact no-regime null results. Neither predictive construction beats split-conformal; both are slightly worse (ΔWinkler
), because weighting on a noise signal only costs efficiency. Both filtered constructions record a lower finite-origin score (
) while leaving 11–12% of origins unbounded. The direct paired contrasts isolate which coupling is responsible: switching predictive to filtered timing at a fixed estimation sample produces a spurious
with a Monte-Carlo interval excluding zero, and of near-identical magnitude whether the regime is fit on validation or on calibration, while switching disjoint to calibration estimation at fixed timing produces no detectable change (
[
,
] predictive;
[
,
] filtered). The manufactured advantage therefore belongs to filtered state timing, while calibration-score reuse produces none under this null.
Figure 5 traces the paired Winkler difference as persistence varies. The forecast-feasible construction earns no detectable advantage where the regime is weakly predictable: it is slightly worse than split at low persistence (
at
) and shows a lower mean Winkler only once the regime is strongly persistent (
at
, over the common finite-origin set). Because it still returns unbounded intervals at about 5% of origins, that difference is conditional on the shared finite-origin set, not a claim of dominance. The filtered, in-sample artifact construction, by contrast, shows an apparent improvement at every persistence level, from
at
to
at
.
Table 13 fixes a genuinely predictable regime (
, separability 3.5, 300 replications) under the finite-sample-corrected construction, with every difference taken over the common finite-origin set. The oracle filtered state retains the largest apparent gain (
). Because the regime is genuinely predictable, the forecast-feasible construction fitted out of sample also gains (
). The predictive calibration-fitted and validation-fitted constructions attain similar conditional finite-origin scores under this stationary process (
versus
). That comparison provides no evidence that calibration-score reuse improves performance here, but it does not establish equivalence either, and without a direct paired factor contrast it cannot separate reuse from estimation-sample size.
Separating the regime-fit data from the scored data leaves no gain (
) in the calibration-split arm, whose effective calibration size is halved, so that change confounds score separation with the smaller calibration set. Taken together, these results make the apparent calibration-fit advantage in the non-stationary application (
Section 3.1) consistent with regime-model recency and transfer under distribution shift—which this stationary design deliberately does not reproduce—rather than with score reuse itself. The simulation establishes only that score reuse per se produces no detectable advantage here; identifying recency or transfer as the operative cause would require the non-stationary experiments left to future work.
One alignment check remains. The simulation fits its regime model to a trailing local log realized variance, whereas the empirical study fits the same model to the log squared residual (
Section 2.6); the smoothed feature is easier to recover, so the design is a favorable-case existence check rather than a like-for-like replica. Repeating the persistent-regime experiment with the empirical feature over 120 replications at
leaves the coupled construction unaffected, at
against
, while the genuine gain available to the forecast-feasible construction falls from
to
.
The contrast cuts both ways. The artifact survives the change of feature untouched, whereas the recoverable advantage depends on how legible the regime is in the signal actually used: against the one-point log-chi-square noise of a single squared residual, it is an order of magnitude smaller, and at
it falls below what a 131-observation test window can resolve (
Section 4.4). The empirical null is therefore what the aligned simulation predicts, not evidence against it.
The simulation thus supplies the counterfactual the application cannot. Under the finite-sample-corrected construction, regime weighting can lower the mean Winkler on finite origins—but only under the persistent, one-step-predictable regime examined here, with separability held fixed rather than swept. Even then, the advantage inflates when the state is a filtered oracle, and 5–7% of origins are unbounded, so it does not dominate split-conformal unconditionally. The RSS3 residuals carry only a weak and largely contemporaneous regime signal (
Section 3.4), placing them at the low-predictability end of this spectrum—which is what the absence of a detectable gain for the forecast-feasible construction reflects.
5. Conclusions and Guidance for Practitioners
5.1. Conclusions
On daily natural-rubber price changes, regime-aware conformal prediction appeared to sharpen intervals by a fifth relative to split-conformal, with a decisively better interval score. That improvement does not survive scrutiny, and this study identifies what produces it.
Three implementation features accompany the gain, and they do not have the same evidential status. A filtered origin state is genuine forecast-origin leakage: it is an oracle quantity, unavailable when the forecast is issued, and a controlled no-regime null reproduces its effect where no regime information exists at all. Omission of the finite-sample correction inflates sharpness wherever the calibration weights concentrate. Calibration-dependent regime estimation is operationally feasible score reuse, and its empirical contribution cannot be separated from regime-model recency and transfer; it is therefore reported as an association rather than a demonstrated cause.
The forecast-feasible construction—which removes the leakage and applies the correction—shows no statistically detectable advantage over split-conformal in the single-window analysis, under full rolling-origin re-estimation, across two to four regime states, or on either the VMD-augmented ridge or the AR(1) residuals. Across 314 rolling origins and three nominal levels, it provides no usable unconditional advantage and returns unbounded intervals at some origins, while adaptive baselines remain finite and attain lower interval scores.
The contribution is diagnostic rather than constructive. Past-measurability is identified as necessary for operational forecasting but not sufficient for a finite-sample guarantee, and the magnitude of the gain that a naive implementation manufactures is quantified. What generalizes is the mechanism, not the null.
5.2. Guidance for Practitioners
The mechanism reaches beyond rubber to any regime-aware conformal construction on dependent data. The protocol below separates genuine gains from artifacts.
Weight by predictive, not filtered, regime probabilities. The regime signal must condition only on information available at the forecast origin. A filtered state leaks the target-period shock and is an oracle quantity.
Separate regime-model estimation from conformal calibration. Fit the regime parameters on a block disjoint from the calibration and test sets; where that is impossible, treat the resulting weights as data-adaptive and validate them through a nested out-of-sample protocol.
Apply the finite-sample correction, and do not stop there. The
point masses are required for the weighted quantile, and omitting them inflates sharpness—but they are not a substitute for correct information timing, and
Section 3.5 exhibits an artifact that survives the correction unchanged.
Run the three-factor check before claiming an improvement. Cross-filtered against predictive state, calibration-informed against disjoint regime estimation, and omission against inclusion of the correction. A gain that survives only under oracle state information, calibration reuse, or an uncorrected quantile is not a deployable improvement.
Have a fallback ready for unbounded intervals. Compute
at each origin, and where it falls below
return the unweighted split-conformal interval;
Section 3.6 shows this to dominate the alternative of weakening the kernel. Report the fired-origin rate: a system that falls back most of the time is not a regime-weighted system.
Do not expect a richer state space to rescue the construction.
Section 3.7 shows the apparent gain growing monotonically from two to four states while the forecast-feasible construction stays at parity.
Benchmark against adaptive baselines. Methods driven by past miscoverage, such as ACI and DtACI, supply predictable and model-free adaptivity. A regime-weighted construction should beat them to justify its additional structure.
Assess the residuals’ forecastability, and diagnose why it is what it is. Where the base forecaster performs no better than a naive benchmark, little exploitable regime structure reaches the conformal layer once contemporaneous information is withheld. But a scaled error near one has two causes with different remedies—an unforecastable target, or an under-specified base model leaving recoverable structure in its residuals—and a simple benchmark such as an AR(1) separates them.
For the applied problem that motivated the study, the implication is direct. Once the couplings are removed, regime weighting confers no statistically detectable advantage over split-conformal on natural-rubber price changes, under the designs examined here. A plain split-conformal interval, or a lightweight adaptive one, is the appropriate default. The value of the regime-aware apparatus, where it exists, must be sought in markets whose volatility regimes are genuinely predictable one step ahead, and demonstrated under the protocol above.