Figure 1.
Five sequential stages—data assembly, availability control, feature construction, nested forecasting and forecast evaluation—feed a sixth diagnostic stage divided into signal, risk and policy branches.
Figure 1.
Five sequential stages—data assembly, availability control, feature construction, nested forecasting and forecast evaluation—feed a sixth diagnostic stage divided into signal, risk and policy branches.
Figure 2.
The national CEA price (left axis) and the European carbon price (right axis), 2021–2026, with the March 2025 sector-expansion date marked. The Chinese price rose through 2024, peaked near 105 CNY/t and then eased.The blue line shows the national CEA price on the left axis, and the green line shows the European carbon price on the right axis.
Figure 2.
The national CEA price (left axis) and the European carbon price (right axis), 2021–2026, with the March 2025 sector-expansion date marked. The Chinese price rose through 2024, peaked near 105 CNY/t and then eased.The blue line shows the national CEA price on the left axis, and the green line shows the European carbon price on the right axis.
Figure 3.
Daily returns (top) and GARCH(1,1)-t conditional volatility (bottom). Volatility is strongly clustered; the dashed line marks the March 2025 expansion.Daily returns (upper panel, blue) and GARCH(1,1)-t conditional volatility (lower panel, red). The orange dashed vertical line marks the March 2025 expansion.
Figure 3.
Daily returns (top) and GARCH(1,1)-t conditional volatility (bottom). Volatility is strongly clustered; the dashed line marks the March 2025 expansion.Daily returns (upper panel, blue) and GARCH(1,1)-t conditional volatility (lower panel, red). The orange dashed vertical line marks the March 2025 expansion.
Figure 4.
No model beats the random walk on out-of-sample RMSE (A), and active-day directional accuracy is near the 50% chance level (B). Red dashed vertical lines represent the 50% chance level.
Figure 4.
No model beats the random walk on out-of-sample RMSE (A), and active-day directional accuracy is near the 50% chance level (B). Red dashed vertical lines represent the 50% chance level.
Figure 5.
Out-of-sample one-step-ahead forecasts for the last 180 trading days: the random-walk and lasso forecasts essentially reproduce the previous close.
Figure 5.
Out-of-sample one-step-ahead forecasts for the last 180 trading days: the random-walk and lasso forecasts essentially reproduce the previous close.
Figure 6.
Holdout SHAP magnitudes: the price’s own short-term technical features receive the largest model-output attributions for next-day returns. The magnitudes are fitted-model associations, not causal or out-of-sample stability estimates.
Figure 6.
Holdout SHAP magnitudes: the price’s own short-term technical features receive the largest model-output attributions for next-day returns. The magnitudes are fitted-model associations, not causal or out-of-sample stability estimates.
Figure 7.
SHAP summary (beeswarm) for next-day returns, showing the direction and dispersion of each feature’s model-output attribution.
Figure 7.
SHAP summary (beeswarm) for next-day returns, showing the direction and dispersion of each feature’s model-output attribution.
Figure 8.
Daily -return correlations are near zero between the national market, the regional pilots and the European price, indicating weak cross-market co-movement at daily frequency.
Figure 8.
Daily -return correlations are near zero between the national market, the regional pilots and the European price, indicating weak cross-market co-movement at daily frequency.
Figure 9.
Partial -dependence profiles within the fitted forest for the two leading features, showing nonlinear model associations consistent with short-horizon mean reversion.
Figure 9.
Partial -dependence profiles within the fitted forest for the two leading features, showing nonlinear model associations consistent with short-horizon mean reversion.
Figure 10.
SHAP dependence for the five-day moving-average gap: the fitted forest maps large positive gaps to slightly negative model-predicted next-day returns.
Figure 10.
SHAP dependence for the five-day moving-average gap: the fitted forest maps large positive gaps to slightly negative model-predicted next-day returns.
Figure 11.
The variance ratio is below one but significant only at the two-day horizon (A); technical-only features are the strongest random-forest family but remain above random-walk RMSE (B).
Figure 11.
The variance ratio is below one but significant only at the two-day horizon (A); technical-only features are the strongest random-forest family but remain above random-walk RMSE (B).
Figure 12.
Out-of-sample price RMSE before and after the 2025 expansion date. Numerical post-period differences are small and are not treated as evidence of an abrupt expansion-period structural change.
Figure 12.
Out-of-sample price RMSE before and after the 2025 expansion date. Numerical post-period differences are small and are not treated as evidence of an abrupt expansion-period structural change.
Figure 13.
Rolling 252-day return AR(1) and annualized realised-volatility estimates, with BIC-selected log-squared-return break dates and the pre-specified expansion date marked. The paths are descriptive and do not identify policy causation.
Figure 13.
Rolling 252-day return AR(1) and annualized realised-volatility estimates, with BIC-selected log-squared-return break dates and the pre-specified expansion date marked. The paths are descriptive and do not identify policy causation.
Figure 14.
Symmetric EGARCH has the lowest full-sample AIC (A), while none of the fixed expansion-date or stability tests reaches the 5% threshold (B).
Figure 14.
Symmetric EGARCH has the lowest full-sample AIC (A), while none of the fixed expansion-date or stability tests reaches the 5% threshold (B).
Figure 15.
Selected robustness diagnostics: macro publication-timing and first-public-value sensitivity, MLP stabilization and common-sample one-step volatility loss. Lower values indicate smaller forecast loss within each panel; scales differ across panels.The dashed vertical line in Panels (a,b) marks the random-walk RMSE benchmark of 1.242 CNY/t.
Figure 15.
Selected robustness diagnostics: macro publication-timing and first-public-value sensitivity, MLP stabilization and common-sample one-step volatility loss. Lower values indicate smaller forecast loss within each panel; scales differ across panels.The dashed vertical line in Panels (a,b) marks the random-walk RMSE benchmark of 1.242 CNY/t.
Table 1.
Data sources, frequencies and forecasting roles.
Table 1.
Data sources, frequencies and forecasting roles.
| Data Block | Content/Frequency | Source | Availability Treatment and Role |
|---|
| National CEA | Daily open, high, low, close, volume and turnover | Wind Financial Terminal; underlying venue SEEE [45] | Same-day public market information; target and technical features |
| Regional pilots/EU | Daily closing prices | Wind Financial Terminal; official trading venues | Available after the respective day-t venue closes and before the next CEA session; short aligned gaps forward-filled |
| PPI and CPI | Monthly year-on-year rates | Wind final history; official NBS first releases [46,47] | Legacy 15-day/final-history baseline; actual timestamp and first-public rate in the real-time sensitivity |
| M0 and RMB loan balance | Monthly levels and year-on-year rates | Wind final history; official PBOC first releases [48,49] | Baseline growth from final levels; actual timestamp and officially reported first-public growth in the real-time sensitivity |
| Weather sensitivity | Daily HDD18, CDD18 and 2 m temperature range | NOAA/NCEP GFS 0.5-degree forecasts [50,51] | Archived forecast fields verified as available before the target CEA session; not realized weather |
| Public coal sensitivity | Ten-day Shanxi Blend (5500 kcal) circulation-market price, CNY/t | National Bureau of Statistics [47] | Last published level carried forward only after its displayed release time; not a daily delivered fuel cost |
| Generation sensitivity | Estimated total national electricity generation, GWh/day | CarbonMonitor-Power [52,53] | Assumed 21-calendar-day availability lag and 28-day check; generation proxy, not observed load |
| Restricted coal alternative | Daily Qinhuangdao Q5500 price, CNY/t | Wind Financial Terminal | One additional CEA-session lag; used only as a proprietary robustness measure because historical publication timestamps are unavailable |
Table 2.
Baseline predictor families and separately tested energy/weather sensitivity variables. Predictive rationales are not causally identified.
Table 2.
Baseline predictor families and separately tested energy/weather sensitivity variables. Predictive rationales are not causally identified.
| Family | Features | Predictive Rationale (Not Causally Identified) |
|---|
| Technical: return/price (9) | Lagged returns (1, 2, 3, 5 days); moving-average gaps (5, 10, 20); momentum (5, 10) | Proxies for reversal, continuation and deviation from recent reference prices |
| Technical: trading/risk (6) | Realized volatility (5, 10, 20); daily high–low range; log volume; volume anomaly | Proxies for trading intensity, concentration and short-horizon uncertainty |
| Cross-market: EU (2) | EU return; EU 20-day standardized price gap | Possible shared exposure to fossil-fuel, macroeconomic and climate-policy news; direct comparability is limited by separate products and institutions |
| Cross-market: Chinese pilots (2) | Pilot-average return; CEA-to-pilot price gap | Possible shared domestic information, qualified by different regional coverage, allocation, participation and liquidity |
| Macro: prices/activity (2) | PPI and CPI inflation | Public proxies for industrial input prices and aggregate-demand conditions |
| Macro: money/credit (2) | M0 growth; loan-balance growth | Public proxies for monetary and financing conditions |
| Calendar/regime (3) | Weekday; December compliance indicator; post-expansion indicator | Calendar and post-period indicators without a policy interpretation |
| Sensitivity only: energy/weather (5) | GFS HDD18, CDD18 and daily temperature range; last-published coal-price level; lagged national generation estimate | Forecast-available temperature, fuel-price and power-system-activity measures; the coal and generation variables retain their published measurement scope |
Table 3.
Descriptive statistics for representative target and predictor variables; the full predictor table is reported in
Appendix A (
Table A1).
Table 3.
Descriptive statistics for representative target and predictor variables; the full predictor table is reported in
Appendix A (
Table A1).
| Variable | n | Mean | Std. Dev. | Min | Max | Skew/Kurt. |
|---|
| CEA close (CNY/t) | 1202 | 70.223 | 16.272 | 41.460 | 105.650 | 0.37/ |
| Next-day log return | 1202 | 0.00043 | 0.01785 | | 0.09388 | 0.27/6.83 |
| 5-day MA gap | 1198 | 0.00082 | 0.01803 | | 0.11426 | 0.84/5.09 |
| Daily range | 1202 | 0.01653 | 0.02437 | 0.00000 | 0.18841 | 2.47/7.86 |
| EU return | 1159 | 0.00026 | 0.02906 | | 0.33863 | 0.24/27.72 |
| Pilot-average return | 1198 | 0.00010 | 0.04089 | | 0.28228 | 0.29/19.91 |
| PPI inflation (%) | 1191 | 0.513 | 4.997 | | 13.500 | 1.22/0.10 |
| M0 annual growth | 961 | 0.115 | 0.022 | 0.027 | 0.172 | /3.49 |
Table 4.
Nested expanding-window hyperparameter grids. No outer-test observation enters tuning or preprocessing.
Table 4.
Nested expanding-window hyperparameter grids. No outer-test observation enters tuning or preprocessing.
| Model | Configuration |
|---|
| Random walk/ARIMA | zero return; ARIMA order (1,0,1), both under the outer expanding window |
| Ridge/lasso | ; |
| Elastic net | ; mixing |
| k-nearest neighbors | |
| Support-vector regression | radial basis; ; |
| Random forest | 160 trees; depth ; minimum leaf |
| Gradient boosting | depth ; learning rate ; 200 iterations |
| Neural network | hidden units ; ; learning rate ; tolerance ; 2000 iterations |
| Ensemble | unweighted mean of the eight nested-tuned machine learning forecasts |
Table 5.
Nested expanding-window performance over 952 one-step-ahead forecasts. DM uses the HLN correction; Holm adjustment is applied separately to price- and return-loss comparisons.
Table 5.
Nested expanding-window performance over 952 one-step-ahead forecasts. DM uses the HLN correction; Holm adjustment is applied separately to price- and return-loss comparisons.
| Model | RMSE | Dir. Active (%) | Price DM p | Price Holm p | Return DM p | Return Holm p |
|---|
| Random walk | 1.242 | — | — | — | — | — |
| ARIMA | 1.250 | 47.8 | 0.479 | 0.479 | 0.491 | 0.491 |
| Neural network | 1.251 | 50.4 | 0.036 | 0.109 | 0.145 | 0.290 |
| Lasso | 1.257 | 46.9 | 0.065 | 0.130 | 0.042 | 0.136 |
| Elastic net | 1.266 | 46.5 | 0.016 | 0.063 | 0.011 | 0.063 |
| Ensemble | 1.268 | 47.1 | 0.009 | 0.054 | 0.027 | 0.136 |
| Random forest | 1.289 | 47.4 | 0.009 | 0.054 | 0.031 | 0.136 |
| Gradient boosting | 1.314 | 44.8 | <0.001 | <0.001 | <0.001 | 0.002 |
| k-nearest neighbors | 1.314 | 53.2 | <0.001 | <0.001 | <0.001 | <0.001 |
| Ridge | 1.335 | 49.6 | <0.001 | <0.001 | <0.001 | <0.001 |
| Support-vector regression | 1.474 | 50.5 | <0.001 | <0.001 | <0.001 | <0.001 |
Table 6.
MLP optimization and regularization diagnostics on the identical 952 outer-test origins. Warning counts cover all candidate-grid fits and final refits; the selected final fit did not reach its iteration cap in any block.
Table 6.
MLP optimization and regularization diagnostics on the identical 952 outer-test origins. Warning counts cover all candidate-grid fits and final refits; the selected final fit did not reach its iteration cap in any block.
| Specification | RMSE | MAE | Active Dir. (%) | Clipped n | Median Iter. | Warnings |
|---|
| Initial grid | 3.848 | 2.918 | 50.0 | 79 | 133 | 0 |
| Tight tolerance only | 1.278 | 0.825 | 49.7 | 0 | 507 | 31 |
| Strong only | 1.246 | 0.745 | 46.7 | 0 | 110 | 0 |
| Lower learning rate only | 7.332 | 6.752 | 48.2 | 606 | 406 | 40 |
| Compact width only | 4.691 | 3.680 | 52.4 | 187 | 111 | 0 |
| Combined stabilized Adam | 1.251 | 0.773 | 50.4 | 0 | 960 | 195 |
| Compact L-BFGS | 1.247 | 0.749 | 47.8 | 0 | 29 | 0 |
Table 7.
Formal efficiency diagnostics for daily CEA returns.
Table 7.
Formal efficiency diagnostics for daily CEA returns.
| Test | Setting | Estimate | Statistic | p-Value |
|---|
| Ljung–Box | lags 5/10/20 | — | 22.53/42.17/49.81 | <0.001/<0.001/<0.001 |
| Runs test | non-zero return signs | — | 0.484 | 0.628 |
| Variance ratio | /5/10 | 0.880/0.829/0.935 | // | 0.043/0.131/0.684 |
| BDS | dimension 2 | — | 10.05 | <0.001 |
Table 8.
Leading model-specific associations evaluated on an untouched final-quarter holdout.
Table 8.
Leading model-specific associations evaluated on an untouched final-quarter holdout.
| Feature | Mean Absolute SHAP () | Holdout Permutation Importance |
|---|
| 5-day moving-average gap | 2.127 | 0.0439 |
| Daily high–low range | 0.630 | 0.0426 |
| 5-day lagged return | 0.619 | 0.0018 |
| 20-day realized volatility | 0.464 | |
| 5-day momentum | 0.438 | 0.0020 |
| 20-day moving-average gap | 0.321 | |
| 10-day moving-average gap | 0.278 | 0.0069 |
Table 9.
Out-of-sample random-forest feature-family ablations. Each panel reuses its full specification’s block-specific settings; negative RMSE denotes lower error than that panel’s full model.
Table 9.
Out-of-sample random-forest feature-family ablations. Each panel reuses its full specification’s block-specific settings; negative RMSE denotes lower error than that panel’s full model.
| Specification | k | RMSE | MAE | Active Direction (%) | RMSE |
|---|
| Panel A: original fixed-delay/final-history baseline |
| Full baseline | 26 | 1.289 | 0.808 | 47.4 | 0.0000 |
| Technical only | 15 | 1.254 | 0.781 | 51.0 | |
| Without technical | 11 | 1.379 | 0.875 | 47.6 | 0.0902 |
| Without cross-market | 22 | 1.262 | 0.789 | 47.6 | |
| Without macro | 22 | 1.288 | 0.804 | 49.7 | |
| Without calendar/regime | 23 | 1.289 | 0.808 | 46.8 | 0.0003 |
| Panel B: public energy/weather sensitivity with exact-release macro values |
| Full public augmentation | 31 | 1.275 | 0.801 | 45.7 | 0.0000 |
| Technical only | 15 | 1.253 | 0.781 | 51.0 | |
| Without technical | 16 | 1.290 | 0.814 | 47.6 | 0.0150 |
| Without weather | 28 | 1.267 | 0.796 | 47.4 | |
| Without NBS ten-day coal | 30 | 1.274 | 0.799 | 46.0 | |
| Without generation proxy | 30 | 1.275 | 0.803 | 46.7 | 0.0008 |
| Original 26, exact-release macro | 26 | 1.268 | 0.795 | 46.8 | |
Table 10.
GARCH(1,1)-t estimates and unit-persistence test.
Table 10.
GARCH(1,1)-t estimates and unit-persistence test.
| Parameter | Estimate | p-Value |
|---|
| (mean) | | 0.172 |
| (constant) | 0.061 | 0.223 |
| (ARCH) | 0.297 | <0.001 |
| (GARCH) | 0.703 | <0.001 |
| (degrees of freedom) | 3.00 | <0.001 |
| Persistence | 1.000 | — |
| Wald test | 0.000 (z) | 1.000 |
Table 11.
Full-sample Student’s t volatility-specification comparison.
Table 11.
Full-sample Student’s t volatility-specification comparison.
| Model | Log Likelihood | AIC | BIC |
|---|
| EGARCH(1,1)-t | | 3857.73 | 3883.19 |
| Asymmetric EGARCH(1,1)-t | | 3858.90 | 3889.45 |
| GARCH(1,1)-t | | 3918.20 | 3943.66 |
| IGARCH(1,1)-t | | 3919.61 | 3939.98 |
| GJR-GARCH(1,1)-t | | 3920.13 | 3950.68 |
| GARCH-X expansion dummy | | 3920.86 | 3951.41 |
Table 12.
One-step-ahead volatility comparison on 702 common forecast dates. Absolute and squared returns are noisy proxies for latent volatility and variance.
Table 12.
One-step-ahead volatility comparison on 702 common forecast dates. Absolute and squared returns are noisy proxies for latent volatility and variance.
| Model | Volatility RMSE | Variance RMSE | QLIKE | Refit Failures |
|---|
| IGARCH-t | 1.416 | 8.148 | 1.833 | 0 |
| GJR-GARCH-t | 1.405 | 8.104 | 1.975 | 1 |
| GARCH-t | 1.413 | 8.179 | 1.983 | 0 |
| Asymmetric EGARCH-t | 4.826 | 55.054 | 3.049 | 0 |
| EGARCH-t | 4.831 | 54.556 | 3.057 | 0 |
Table 13.
Descriptive pre/post comparison and fixed-date break test.
Table 13.
Descriptive pre/post comparison and fixed-date break test.
| Period | n | Mean Price (CNY/t) | Mean |Return| (%) | Ann. Volatility (%) |
|---|
| Pre-expansion | 893 | 69.35 | 1.01 | 27.98 |
| Post-expansion | 310 | 72.80 | 1.11 | 29.40 |
| Chow test on the return AR(1): F , . |
Table 14.
Rolling, scanned and BIC-selected structural-instability diagnostics.
Table 14.
Rolling, scanned and BIC-selected structural-instability diagnostics.
| Diagnostic | Result | Expansion-Period Interpretation |
|---|
| HAC candidate-date scan | Minimum raw p: mean 0.151; variance proxy 0.146; all Holm | No candidate date survives multiplicity correction |
| Return AR(1) BIC segmentation | 0 breaks; BIC | No selected conditional-mean break |
| Log-squared-return BIC segmentation | 3 breaks; 18 Apr 2022, 21 Jul 2023, 18 Oct 2024; BIC | Selected volatility-proxy dates precede 26 Mar 2025 |
| 252-day rolling estimates | AR(1) ; annualized volatility | 25 Mar 2025: AR(1) ; volatility |
Table 15.
Out-of-sample predictability before and after the 2025 expansion.
Table 15.
Out-of-sample predictability before and after the 2025 expansion.
| Period | Model | n | RMSE (CNY/t) | Dir. Acc. Active (%) |
|---|
| Pre-expansion | Random walk | 642 | 1.199 | — |
| Pre-expansion | Lasso | 642 | 1.216 | 51.3 |
| Pre-expansion | k-nearest neighbors | 642 | 1.310 | 51.5 |
| Pre-expansion | Random forest | 642 | 1.271 | 49.1 |
| Pre-expansion | Gradient boosting | 642 | 1.275 | 47.1 |
| Post-expansion | Random walk | 310 | 1.328 | — |
| Post-expansion | Lasso | 310 | 1.339 | 39.4 |
| Post-expansion | k-nearest neighbors | 310 | 1.321 | 56.1 |
| Post-expansion | Random forest | 310 | 1.326 | 44.6 |
| Post-expansion | Gradient boosting | 310 | 1.390 | 40.8 |
| Active-direction sample: 493 pre-expansion and 289 post-expansion observations. |
Table 16.
Random-forest sensitivity to macro publication timing and value vintage.
Table 16.
Random-forest sensitivity to macro publication timing and value vintage.
| Availability Schedule | RMSE | MAE | Active Direction (%) | Macro Permutation MSE () |
|---|
| 15-day delay, final history | 1.289 | 0.808 | 47.4 | 0.159 |
| Exact release, final history | 1.265 | 0.794 | 46.0 | |
| 15-day delay, first public | 1.286 | 0.808 | 47.7 | 0.071 |
| Exact release, first public | 1.266 | 0.793 | 46.2 | |
| Macro omitted | 1.288 | 0.804 | 49.7 | — |
Table 17.
Lower-frequency aggregation of one-step daily random-forest forecasts. Period means are evaluated on interior target periods; these are aggregated forecast summaries, not direct weekly or monthly models.
Table 17.
Lower-frequency aggregation of one-step daily random-forest forecasts. Period means are evaluated on interior target periods; these are aggregated forecast summaries, not direct weekly or monthly models.
| Aggregation | Periods | Exact First-Public RMSE | Macro-Omitted RMSE | Random-Walk RMSE | Exact vs. Omitted HAC p |
|---|
| Weekly | 200 | 0.609 | 0.661 | 0.517 | 0.146 |
| Monthly | 47 | 0.382 | 0.392 | 0.293 | 0.389 |
Table 18.
Institutional comparison relevant to cross-market price comparability.
Table 18.
Institutional comparison relevant to cross-market price comparability.
| System | Product and Compliance Boundary | Allocation and Participation | Implication for Daily Comparison |
|---|
| National CEA | National allowance and registry; nationally covered key emitters no longer participate in the relevant local pilot | Predominantly free, output/benchmark-linked allocation in the study period; key emitters and eligible institutions/individuals | Separate national product prevents assuming direct pilot or EU arbitrage parity |
| Hubei, Guangdong, Shanghai and Shenzhen pilots | Each defines local allowances, accounts and surrender eligibility under its own rules | Allocation and participant-access provisions differ across pilots and over time | Local scarcity, compliance calendars and liquidity need not map one-for-one into national CEA returns |
| EU ETS | EU allowance in the Union Registry under a declining cap; no EU–China allowance link located in the cited official sources | Auctioning is the default, with harmonized free allocation for eligible sectors; covered operators and eligible market participants trade under the EU market-integrity framework | Common news may co-move prices, but direct unit interchangeability and identical market depth cannot be assumed |