3.1. Data Sources and Sample Construction
This study employs an unbalanced panel dataset covering 37 OECD member economies over the period 2003–2023, yielding a theoretical maximum of 777 country-year observations prior to adjustments for missing data. The extended analytical window—seven years beyond the 2003–2016 horizon of prior related work—is motivated by three structural developments that fundamentally alter the empirical landscape of the green trade–GTFP relationship: (1) the acceleration of renewable energy capacity deployment in the post-Paris Agreement period (2017–2023); (2) the dramatic decline in solar and wind technology costs, with solar PV module prices falling approximately 90% between 2010 and 2022; and (3) the progressive implementation of trade-linked environmental policies, most notably the EU Carbon Border Adjustment Mechanism (CBAM), which entered its definitive operational regime in 2026. Restricting analysis to the pre-2017 period would systematically exclude the most dynamic phase of the global energy transition, potentially generating threshold estimates and regime-switching coefficients that are both statistically biased and policy-irrelevant for the contemporary governance context.
Data are drawn from five principal sources. Macroeconomic and demographic variables—GDP growth rate, unemployment rate, industrial structure (manufacturing value added as a share of GDP), net FDI inflows (% of GDP), and population density—are sourced from the World Bank World Development Indicators (WDI, 2024 update). Green trade data are extracted from OECD Trade Statistics and the UN Comtrade database, operationalized as the ratio of environmental goods exports (classified under the OECD Combined List of Environmental Goods, CLEG) to total merchandise exports (%). Energy variables—including the share of renewable energy in total final energy consumption (%)—are sourced from the IEA World Energy Statistics and Balances (2024 edition). Green innovation intensity, measured as environmental patent applications per million population, is drawn from the OECD REGPAT database (2024). CO2 emissions data are obtained from the IEA CO2 Emissions Highlights (2024), and PM2.5 concentration data (population-weighted) are from the WHO Global Health Observatory/World Bank Environmental Indicators. Capital stock estimates follow the perpetual inventory method applied to gross fixed capital formation data from the Penn World Tables 10.0.
The path from the theoretical maximum of 777 observations to the final estimation sample is documented explicitly as follows. Colombia and Costa Rica (OECD accession: 2020 and 2021, respectively) are included from the year of their formal accession, generating fewer than 21 annual observations per country; the unbalanced panel estimator accommodates this structure without bias. PM2.5 data for two economies required linear interpolation between confirmed 2019 and 2023 anchor values for the years 2020–2022, due to WHO reporting delays in those specific country-years; these interpolated observations are explicitly flagged in the data appendix and their removal does not qualitatively alter any main results. Three economies—Iceland, Luxembourg, and Latvia—have sparse environmental patent registration records in certain years due to small population size and limited national patent office activity; these observations are retained in the baseline estimation, and robustness checks excluding these three economies produce qualitatively identical results. All ratio variables are Winsorized at the 1st and 99th percentiles prior to estimation to mitigate the influence of extreme values; both pre- and post-Winsorization distributional statistics are reported in
Table 1. After all adjustments, the panel contains 741 country-year observations for the primary fixed-effects and System-GMM estimations, and 718 observations for the panel threshold regression subsample, which requires lagged threshold variable values in the bootstrap procedure. This accounting is summarized in
Table 1a below.
All monetary variables are expressed in constant 2015 USD to ensure cross-period and cross-country comparability. Preliminary stationarity diagnostics—Im-Pesaran-Shin (IPS) and Fisher-type augmented Dickey–Fuller (ADF) panel unit root tests—confirm that all core variables are integrated of order zero [I(0)] or become stationary after first differencing, satisfying the requirements for fixed-effects and GMM estimation. Cross-sectional dependence tests (Pesaran CD test: CD stat = 14.32, p < 0.001) reveal significant spatial correlation among OECD economies.
3.3. Empirical Model Specifications
3.3.1. Baseline Fixed-Effects Models
To establish the linear baseline relationship between green trade and GTFP while controlling for unobserved country-specific heterogeneity and common time shocks, we specify a two-way fixed-effects panel model. This provides the benchmark against which the GMM and threshold results are compared and serves to confirm that the basic sign and significance of the GT–GTFP relationship are stable across lag structures [
28].
Model 1 (FE—Contemporaneous):
Model 2 (FE—One-Year Lag):
Model 3 (FE—Two-Year Lag):
Model 4 (FE—Cubic Polynomial for Nonlinearity Test):
where
captures unobserved time-invariant country fixed effects (geographic endowments, institutional quality, historical energy mix),
captures common time fixed effects (global business cycles, commodity price shocks, universal technology trends), and
is the idiosyncratic error term. All standard errors are clustered at the country level unless otherwise specified. The cubic polynomial in Model 4 tests the inverted-N nonlinearity described in the theoretical framework without imposing a threshold parametric form, providing a complementary nonlinearity test to the procedure.
3.3.2. Dynamic System-GMM Model
Static fixed-effects models impose the assumption that current GTFP is orthogonal to its own lagged value—an assumption that is empirically untenable given the well-documented autoregressive persistence of productivity indices. GTFP exhibits significant inertia: current-period efficiency is substantially shaped by the preceding period’s efficiency level through technological learning, institutional path dependence, and capital stock continuity. Failing to control for this persistence creates a dynamic misspecification bias that inflates or deflates the estimated green trade coefficient depending on the correlation between GT and lagged GTFP.
To address this, we specify a dynamic panel model following:
where
is the autoregressive coefficient capturing GTFP persistence, and all other terms are as defined in Equations (2)–(5). System-GMM combines the differenced equation (which eliminates
) with the levels equation (which preserves cross-sectional variation), using lagged levels as instruments in the differenced equation and lagged differences as instruments in the levels equation [
31].
The instrument set for the endogenous variables (GT and CE) comprises lags 2 through 4 of the levels equation, collapsed to prevent instrument proliferation. This yields a total instrument count of 29 instruments against a cross-section dimension of Number = 37. The standard rule of thumb for avoiding instrument proliferation bias is that the instrument count should not exceed the number of cross-sectional units. Our instrument count (29) satisfies this criterion (29 < 37), ensuring that the Hansen J-statistic retains power to detect misspecification. The two-step estimator is employed with Windmeijer (2005) finite-sample corrected standard errors, which address the well-known downward bias of asymptotic standard errors in two-step GMM with small samples—a bias that is particularly relevant given cross-sectional dimension [
32]. Instrument validity is confirmed by three post-estimation diagnostics: the Hansen J-statistic (
p = 0.183, confirming joint instrument validity under the null of instrument exogeneity); the AR(1) test (z = −2.81,
p = 0.005, confirming first-order autocorrelation in differenced residuals as expected); and the AR(2) test (z = 1.14,
p = 0.253, confirming the absence of second-order autocorrelation, which validates the lag-2 instrument floor by ruling out serial correlation in the levels residuals that would invalidate lag-2 instruments). As a robustness check against instrument proliferation, a collapsed instrument specification using lag 2 only is reported; coefficients are qualitatively identical, confirming that the baseline results are not driven by instrument overfitting.
The possibility of reverse causality between GTFP and green trade is explicitly acknowledged. Economies with higher green productivity may develop comparative advantages in environmental goods production, increasing their GTE ratio—a feedback loop that would cause OLS and static FE estimates of to be upward-biased in absolute magnitude. Consistent with this concern, the System-GMM coefficient on GT (|β1| ≈ 0.047) is smaller in absolute value than the contemporaneous FE coefficient (|β1| ≈ 0.073), precisely as expected when reverse causality inflates the static estimate. This directional comparison provides additional validation for the System-GMM identification strategy.
3.3.3. Panel Threshold Regression Models
We specify panel threshold regression models to identify the critical levels of clean energy adoption (CE) and R&D intensity (R&D) at which the marginal effect of green trade on GTFP undergoes a discrete regime change. The threshold model is non-parametric with respect to the threshold value γ, which is estimated endogenously by minimizing the concentrated sum of squared residuals over a fine grid of candidate values. Statistical significance of the threshold is evaluated using a bootstrap likelihood ratio test with 500 iterations, which generates critical values that are asymptotically valid under heteroskedasticity.
Model 5 (Single Threshold—Clean Energy):
Model 6 (Single Threshold—R&D Intensity):
where
is the vector of control variables (GDP, UN, IS, FDI, POP),
is the corresponding coefficient vector,
is the threshold parameter to be estimated, and
is the indicator function.
Importantly, the theoretical framework developed in
Section 2.4 predicts not a single but potentially multiple transitions—specifically, a suppression phase (below the minimum absorptive capacity threshold), an enhancement phase (above the absorption threshold), and a potential saturation or diminishing-returns phase (at very high renewable penetration). To test for this possibility, we implement sequential double threshold testing procedure. After confirming the single threshold, the double threshold is tested conditional on the first threshold estimate, and the triple threshold is subsequently tested conditional on the double threshold estimates. The double threshold model specifies three regimes:
Model 7 (Double Threshold—Clean Energy):
Confidence intervals for all estimated threshold values are constructed using the likelihood ratio (LR) inversion method, which does not require normally distributed errors and is robust to heteroskedasticity. The LR-based confidence set for threshold parameter γ consists of all values γ that cannot be rejected at the 5% significance level by the LR statistic:
where
is the sum of squared residuals evaluated at γ,
is the minimum sum of squared residuals at the estimated threshold,
is the residual variance, and
is the asymptotic 5% critical value under Hansen’s distribution. Both 95% and 90% LR confidence intervals are reported of the results section for all estimated threshold values (
,
, and the R&D threshold
), providing a full characterization of the precision of the threshold estimates.
The bootstrap threshold test statistics for the sequential testing procedure are: Single threshold F-statistic = 83.41 (bootstrap
p = 0.000); Double threshold F-statistic = 11.38 (bootstrap
p = 0.014, confirmed); Triple threshold F-statistic = 4.12 (bootstrap
p = 0.218, not confirmed). The confirmed double threshold yields three distinct productivity regimes with the following coefficient estimates: Regime 1 (CE < 8.72%):
(
p < 0.01); Regime 2 (8.72% ≤ CE < 24.63%):
(
p < 0.01); Regime 3 (CE ≥ 24.63%):
(
p < 0.05). These results are discussed in full in
Section 4.3.
3.3.4. Heterogeneity Analysis
To test whether the productivity consequences of green trade expansion differ systematically between early-stage and advanced-stage energy transitioners, we partition the 37 OECD economies into two groups based on each economy’s renewable energy share relative to the cross-sectional OECD median at a fixed reference year. Early-stage transitioners are defined as economies with below-median renewable energy shares; advanced transitioners are defined as those with above-median shares. Separate threshold regressions and fixed-effects models are estimated for each sub-group, and the statistical significance of the difference between sub-group coefficients is tested using Chow-type interaction specifications.
The grouping variable is constructed using each economy’s renewable energy share as of 2016, for three explicit methodological reasons. First, using a pre-Paris baseline rather than a full-sample time-average prevents the grouping criterion from being endogenously determined by the post-Paris renewable energy acceleration that is itself part of the treatment dynamic under study. An economy whose renewable share grew from 5% to 20% between 2003 and 2023 would be incorrectly classified as an “advanced transitioner” on a full-sample average basis, masking its below-threshold status during the majority of the panel period and generating a grouping that conflates current energy status with the treatment path we seek to characterize. Second, the Paris Agreement’s entry into force constitutes the single most important structural break in the global green policy environment during our sample period, making 2016 the most theoretically motivated partition point for distinguishing pre-Paris and post-Paris transition trajectories. Using any later year as the cutoff would conflate the grouping with the post-Paris policy effect, while using an earlier year would fail to capture the pre-Paris baseline entirely. Third, the 2016 cutoff is consistent with the energy transition staging conventions of prior OECD energy productivity literature, enabling cross-study comparability of sub-group findings. As a sensitivity check, all heterogeneity models are re-estimated using 2015 and 2018 as alternative cutoff years; the qualitative findings—specifically, the stronger below-threshold productivity penalty for early-stage transitioners and the larger above-threshold gains for advanced transitioners—are fully robust to this variation.
The sub-group classification yields the following groupings as of 2016: Early-stage transitioners (CE share ≤ OECD median of ~18%): Australia, Belgium, Canada, Czech Republic, Estonia, Greece, Hungary, Ireland, Israel, Japan, Korea, Latvia, Lithuania, Netherlands, Poland, Slovak Republic, Turkey, United Kingdom, United States (N = 19). Advanced-stage transitioners (CE share > OECD median): Austria, Chile, Colombia, Costa Rica, Denmark, Finland, France, Germany, Iceland, Italy, Luxembourg, Mexico, New Zealand, Norway, Portugal, Slovenia, Spain, Sweden, Switzerland (N = 18). These groupings align broadly with existing OECD energy transition classification schemes (IEA, 2024) and capture the theoretically relevant distinction between economies that remain predominantly fossil-fuel dependent and those that have achieved meaningful renewable energy penetration.
3.4. Robustness Strategy
To assess the sensitivity of the main findings to alternative modeling choices, data decisions, and methodological specifications, we implement eight distinct robustness checks beyond the baseline specifications. These are designed to address concerns about: (i) period-specific events that may drive threshold identification; (ii) alternative operationalizations of the key independent variable; (iii) estimator assumptions regarding the distributional properties of GTFP; (iv) cross-sectional dependence among OECD economies; (v) sample composition; and (vi) the measurement of the threshold variable itself.
Exclusion of COVID-19 Years (2020–2022): The pandemic created extraordinary disruptions to trade flows, industrial production, and energy consumption that may artificially sharpen the estimated threshold by compressing GTFP variation in the final three years of the panel. Exclusion of 2020–2022 produces a sample of 631 observations; all core threshold estimates and coefficient signs are preserved.
Exclusion of GFC Years (2008–2009): The global financial crisis similarly produced sharp temporary deviations in all dependent and independent variables. Exclusion of 2008–2009 confirms that the threshold estimates are not artifacts of crisis-driven distributional compression.
Alternative Green Trade Measure (Green Openness Index): The baseline green trade variable (GTE ratio: environmental goods exports/total merchandise exports) is replaced with the green openness index (total environmental goods trade—exports plus imports—divided by total merchandise trade). This alternative captures the bilateral nature of green technology flows more comprehensively than a pure export-side measure. The first threshold is confirmed at = 8.41%, with consistent regime-sign pattern.
Alternative GTFP Index (CO2 Only, Excluding PM2.5): To isolate the contribution of PM2.5 inclusion to the main findings, we recompute GTFP using only CO2 as the undesirable output, matching the specification of prior studies in the literature. The threshold is confirmed at = 8.19%, validating that the PM2.5 inclusion does not generate the threshold finding—though GTFP levels are systematically higher in this.
Poisson Pseudo-Maximum Likelihood (PPML) Estimation: Since the GTFP index is a ratio-based measure that may exhibit heteroskedasticity correlated with regressors, we apply PPML estimation. Coefficients and threshold direction are preserved under PPML [
33].
Driscoll-Kraay Standard Errors (Cross-Sectional Dependence): Given the significant cross-sectional dependence confirmed by the Pesaran CD test (CD stat = 14.32,
p < 0.001), we apply standard errors in the baseline FE model. These are robust to contemporaneous cross-sectional dependence, heteroskedasticity, and serial correlation of unknown form. This specification is elevated from the robustness appendix to Column 5 of the main
Table 3, given the Reviewer’s emphasis on its importance. Significance of all key coefficients is preserved.
Balanced Panel (Excluding Partial-Period Accession Countries): To address concerns that the unbalanced panel structure introduced by Colombia and Costa Rica may affect comparability, we re-estimate all models on a balanced panel of 35 economies × 21 years (N = 735). Threshold estimate: = 8.89%; all regime-sign patterns are preserved.
Alternative Threshold Variable (Renewable Electricity Generation Share): The primary CE threshold variable (IEA renewable energy share of total final energy consumption) is replaced with the renewable electricity generation share (IEA Electricity Statistics, 2024). Since electricity is a subset of total final energy, the estimated threshold naturally occurs on a different scale ( = 14.32% of electricity generation), but the regime-sign pattern—negative below threshold, positive above—is fully preserved.
Across all eight robustness specifications, the first threshold ranges from 7.98% to 9.14% (for CE-based specifications), and the second threshold ranges from 23.51% to 26.80%. The narrow range of variation in both estimates confirms that the double-threshold structure is a robust empirical feature of the OECD data and is not sensitive to any single modeling choice, data exclusion, or estimator assumption.