1. Introduction
The intensification of climate-related risks since the Paris Agreement has pressured firms worldwide to widen the scope and depth of their environmental reporting, with disclosure now treated less as voluntary citizenship and more as a quasi-regulatory expectation [
1]. Stock exchanges across more than sixty jurisdictions have issued sustainability reporting guidance, and the consolidation of ISSB standards in 2023 marked a turning point toward comparable, decision-useful environmental data [
2]. Within this evolving landscape, Chinese listed firms occupy a distinctive position. They face both the national “dual-carbon” pledge—peaking carbon emissions before 2030 and reaching neutrality by 2060—and a fragmented disclosure regime that until recently relied on voluntary CSR reports [
3]. The Ministry of Ecology and Environment’s 2021 measures, together with the 2024 mandatory sustainability disclosure rules issued by the three major stock exchanges, have begun to reshape what counts as credible environmental information. Evidence from China’s earlier 2008 mandatory CSR reporting policy indicates that such mandates can measurably raise firms’ substantive environmental response, including green innovation performance [
4].
Green innovation has emerged as the principal mechanism through which firms reconcile growth ambitions with decarbonization commitments. Unlike conventional R&D, green innovation embeds environmental constraints into the production function itself, generating dual externalities—knowledge spillovers and pollution abatement—that markets price imperfectly [
5]. For Chinese manufacturers in particular, the transition demands not only more green R&D spending but also higher conversion efficiency from inputs to patentable outcomes [
6]. The efficiency dimension is where most firms struggle: aggregate green patent filings have surged, yet output-to-input ratios remain uneven across industries and provinces [
7].
Three strands of literature speak to this puzzle but rarely converge. The first examines whether environmental disclosure quality, measured by GRI compliance, third-party assurance, or text-based readability, alters firms’ real environmental behavior; findings lean positive but vary considerably with the measurement instrument chosen [
8]. The second strand probes governance characteristics—board independence, the presence of environmental committees, executive green-pay sensitivity, ownership concentration—and links them to innovation outcomes [
9]. Dual-class structures, state ownership, and institutional investor horizons all surface as moderators of green strategy [
10]. A third strand quantifies green innovation efficiency through DEA, SFA, or Malmquist–Luenberger indices, often comparing regions or industries without modeling firm-level heterogeneity explicitly [
11]. What is striking is how little these strands agree once their findings are placed side by side. The disclosure strand reports effect signs that swing from strongly positive to null depending on whether disclosure is captured as a binary report dummy or a graded content score, a divergence that is almost certainly a measurement artifact rather than a substantive disagreement. The governance strand splits along institutional lines: studies set in common-law markets tend to read concentrated ownership as patient stewardship, whereas China-based work more often reads it as entrenchment, so the moderating sign reverses across settings. The efficiency strand, by operating at whichever single level the data happen to be aggregated, produces region- and industry-level conclusions that need not hold at the firm level at all. Seen together, these are not three parallel literatures, but three sources of unresolved tension, and it is precisely the cross-level, measurement-graded design adopted here that lets us adjudicate among them rather than add a fourth disconnected result.
What unsettles us about the current body of work is its treatment of measurement and data structure. Disclosure quality is frequently proxied by a single binary indicator—whether a CSR report is issued—collapsing substantial variation in content credibility [
12]. Green innovation efficiency studies tend to operate at one analytic layer (firm, industry, or region) even though the data-generating process is unambiguously nested: firms are embedded in industries shaped by competitive dynamics and regulatory pressure, while industries are themselves embedded in regional institutional environments [
13]. Pooled OLS or single-level fixed-effects estimators absorb but do not decompose this clustering, leaving cross-level interactions—how regional environmental enforcement amplifies the disclosure–innovation link, for instance—largely unidentified. Recent contributions have flagged this concern, yet few move beyond two-way clustering toward formal multilevel specifications [
14].
A second gap concerns mechanism identification. Most empirical work documents an average association between disclosure and innovation outcomes without separating the variance attributable to firm-level governance from that absorbed by industry technology cycles or provincial green-finance pilots. When the dependent variable is itself a ratio—innovation output over innovation input—failing to model random effects at each level can bias both slope estimates and their standard errors [
15]. We see this as an opening rather than a flaw to be patched.
Against this background, our study asks how the quality of corporate environmental information disclosure, jointly with governance structure characteristics, shapes the input–output efficiency of green innovation in Chinese listed firms and whether industry- and region-level conditions moderate that relationship. Hierarchical linear modeling (HLM) is brought in deliberately: it partitions variance across firm, industry, and provincial layers; permits random slopes for disclosure quality across industries; and accommodates cross-level interactions between, say, firm-level disclosure scores and provincial environmental court density. A sample of A-share manufacturing firms covering the full 2014–2024 window supports the empirical work, with disclosure quality scored through a multidimensional text-and-content rubric and green innovation efficiency computed through a super-SBM model that accounts for undesirable outputs. We retain this eleven-year window consistently across the abstract, methods, and conclusions; where a shorter 2017–2024 horizon appears later, it denotes a deliberately restricted robustness subsample rather than the main panel.
Three contributions follow, and we want to be explicit that the value added is theoretical rather than a mere recombination of familiar tools. First, the paper does not simply propose another disclosure index; it reconceptualizes disclosure quality as a multidimensional signal whose efficiency payoff is theoretically contingent on the governance and institutional layers through which it must travel, a proposition that binary or count-based proxies cannot even articulate. Second, the three-level specification is not adopted for methodological novelty but because it operationalizes a claim that the prior literature could not test: that the disclosure–efficiency slope is itself a random quantity governed by industry regulation and regional marketization, so that cross-level interaction terms become the formal expression of a nested-institutions theory rather than a modeling convenience. Third, by relocating the dependent variable from innovation volume to conversion efficiency, the framework reframes the theoretical question itself—away from whether disclosure induces more green activity and toward whether it reshapes the internal selection of which green projects survive. The combination, in short, yields a conditional theory of disclosure value that none of the three constituent methods generates on its own. The remainder proceeds as follows:
Section 2 develops the theoretical framework and hypotheses;
Section 3 details data, variables, and the multilevel specification;
Section 4 reports estimation results and robustness checks;
Section 5 discusses mechanisms and boundary conditions;
Section 6 closes with policy implications.
4. Empirical Results and Analysis
4.1. Descriptive Statistics and Baseline Regression
Across the 21,536 firm-year observations, the disclosure quality index averages 42.7 with a standard deviation of 18.3, indicating wide dispersion that the panel structure preserves rather than smooths over. The interquartile range runs from 28.4 to 56.9, and a non-trivial right tail above 80 reflects the small group of firms—mostly large central state-owned enterprises in energy and materials—that produce GRI-aligned, externally assured sustainability reports. Green innovation efficiency scores range from 0.12 to 1.34, with a mean of 0.518, consistent with the modest aggregate productivity figures reported for Chinese manufacturing green R&D pipelines [
56]. Pairwise correlations between key explanatory variables stay below 0.45 in absolute value, and variance inflation factors range from 1.21 to 3.67 with a mean of 2.04, well below the conventional threshold of 10, so multicollinearity is not a serious concern.
Before fitting any predictor-augmented specification, we ran the null model defined in Equation (9) to recover the variance components. The intraclass correlation at the provincial level lands at 0.143, while the industry-level ICC is 0.182, leaving 67.5 percent of total variance at the firm-year layer. Both higher-level ICCs comfortably exceed the 0.05 cutoff suggested by methodologists working with nested organizational data [
40], confirming that pooling all observations into a single-level OLS would obscure systematic between-cluster differences and bias inference on the disclosure–efficiency slope.
Table 6 reports the variance components alongside the deviance, AIC, and level-specific pseudo-
for each modeling stage and for the two-way fixed-effects benchmark.
The grouped comparison in
Figure 3 makes the disclosure–efficiency relationship visible before any conditioning takes place. Firms in the top quartile of disclosure quality post mean efficiency scores roughly 38 percent above those in the bottom quartile, and the gap widens monotonically across the four bands.
Pooling individual firm-years rather than quartile means,
Figure 4 traces the conditional relationship in a scatterplot with a LOESS fit overlaid. The upward slope is unmistakable, though the relationship flattens at the upper end of the disclosure distribution, hinting at diminishing marginal returns once disclosure quality moves beyond the regulatory frontier.
Temporal dynamics deserve a separate look.
Figure 5 plots sample-wide mean efficiency by year and shows a gentle upward drift from 0.43 in 2014 to 0.58 in 2024, with a visible acceleration after 2020 that coincides with the dual-carbon policy announcement and the green credit guideline rollout [
20,
57].
Heterogeneity across two-digit industries warrants attention before the regressions interpret it away.
Figure 6 displays the boxplot distribution by industry code, with pollution-intensive sectors—steel, chemicals, thermal power—exhibiting both lower medians and wider dispersions than electronics or pharmaceuticals.
Baseline random-intercept estimation, with EDQ entered as the sole Level-1 predictor alongside firm controls, returns a coefficient of
(
,
). A one-standard-deviation rise in disclosure quality is associated with a 0.084 increase in the efficiency score—about 16 percent of the sample mean—an association whose economic size mirrors the visual pattern in
Figure 3 and
Figure 4. We phrase this deliberately as an association: even with the robustness battery that follows, the baseline rests on observational variation, and the magnitude should be read as the conditional difference between high- and low-disclosure firms rather than as the effect of an exogenous change in disclosure. The deviance test against the null model returns
with two degrees of freedom, decisively rejecting the no-effect specification. Firm size and ROA enter with the expected positive signs, while leverage suppresses efficiency, consistent with the financing-constraint mechanism developed in
Section 2.1 and broadly aligned with panel evidence linking financing constraints and green innovation in Chinese firms [
58]. The baseline coefficient supplies a benchmark against which the random-slope and cross-level interaction models in subsequent subsections will be evaluated.
Table 7 collects the full set of fixed-effect coefficients, standard errors, and significance levels for the baseline, governance-moderation, cross-level, and fully controlled specifications side by side, so the progression can be read directly rather than reconstructed from the text.
4.2. Moderating Effects of Governance Structure Characteristics
Testing the four governance moderators developed under H2a–H2d proceeds through sequential introduction of interaction terms into the random-intercept specification, with all continuous moderators grand-mean centered to keep the conditional EDQ slope interpretable at sample means [
59]. We resist the temptation to enter all four interactions simultaneously in a single equation, since collinearity between governance variables would muddy the marginal effect estimates; instead, each moderator enters in its own augmented model with the others held as Level-1 controls.
The board independence interaction produces a positive and statistically reliable coefficient. Estimated from the specification corresponding to Equation (3) in
Section 2.2, the interaction term returns
(
,
), with the marginal effect of EDQ on GIE expressed as
Substituting the 25th and 75th percentile values of IND (0.33 and 0.45) yields conditional EDQ slopes of 0.0038 and 0.0063, respectively, a 66 percent increase across the interquartile range. The amplification is consistent with the monitoring logic developed in H2a: independent boards raise the cost of decoupled disclosure and steer green R&D toward higher-conversion projects [
60].
Ownership concentration cuts the other way. The interaction coefficient is negative and significant (
,
,
), and the conditional slope expression follows:
At low ownership concentration (25th percentile, 22.4 percent), the EDQ slope reaches 0.0058; at high concentration (75th percentile, 48.7 percent), it falls to 0.0021. Concentrated controlling shareholders appear to repurpose high-quality disclosure as a legitimacy device rather than a commitment to abatement-focused R&D, which lines up with the entrenchment account in H2b. The result is one of the cleaner findings in our analysis—the sign is robust across both random-intercept and random-slope estimators, and the magnitude does not shrink when SOE status is partialed out.
The environmental executive moderator yields the largest amplification effect of the four. With ENVEXE as a binary indicator, the interaction term takes
(
,
), and the slope differential is given by
Firms with an environmentally experienced executive in the TMT convert disclosure quality into innovation efficiency at roughly 1.7 times the rate of firms without such experience. The upper-echelons mechanism appears to operate as theorized: cognitive priors that privilege long-horizon environmental payoffs reduce the friction between outward-facing disclosure and inward-facing R&D portfolio adjustments [
61]. This effect remains stable after controlling for whether the firm is in a pollution-intensive industry, suggesting that it is genuinely an executive-level rather than industry-level phenomenon.
Institutional shareholding moderates positively, but with a smaller magnitude. The interaction coefficient comes in at
(
,
), and the marginal effect follows:
Across the interquartile range of INST (from 12.3 percent to 41.6 percent), the conditional EDQ slope rises from 0.0044 to 0.0070. The amplification is real but more muted than the executive background channel, perhaps because aggregate institutional ownership conflates long-horizon and transient holders—a distinction that future work could sharpen by separating pension and social security funds from short-term mutual funds [
62]. We flag this as a measurement limitation rather than a refutation of H2d.
Figure 7 places the four marginal effects on a common axis to make the contrast visible.
Ranking the moderators by the magnitude of the marginal-effect shift across their respective ranges, environmental executive background and ownership concentration emerge as the most influential, followed by institutional shareholding and then board independence. Hypotheses H2a–H2d each find support, though the strength of the effects diverges enough to caution against treating governance as a homogeneous bundle.
4.3. Cross-Level Interaction Tests and Robustness Analysis
Moving from firm-level moderation to the cross-level structure laid out in Equation (16), we test whether industry regulation intensity and provincial marketization condition the disclosure–efficiency slope. The cross-level interaction with REG returns
(
,
), so the conditional EDQ slope can be written as
Industries above the 75th percentile of regulatory intensity—thermal power, cement, non-ferrous metals—exhibit EDQ slopes roughly 1.8 times those in lightly regulated sectors, lending empirical support to H3a. The regulatory environment evidently raises the informational stakes of disclosure: where misreporting carries steeper administrative consequences, the same disclosure score conveys a stronger signal and travels farther into capital allocation decisions [
63].
Provincial marketization carries an analogous amplifying role. The interaction coefficient lands at
(
,
), with the conditional slope:
Firms domiciled in provinces in the top quartile of the Fan–Wang index—Guangdong, Zhejiang, Jiangsu—translate disclosure quality into innovation efficiency at roughly twice the rate of firms in bottom-quartile provinces. The mechanism follows the institutional logic in H3b: thicker factor markets, more competitive intermediary services, and stronger property-rights protection [
37] let credible disclosure flow through green credit channels and technology-licensing markets without dissipating in administrative friction [
64]. Deviance tests against the random-slope-only specification yield
(
,
), so adding the cross-level interactions substantially improves fit.
A natural concern is endogeneity between disclosure quality and innovation efficiency: firms with more productive green R&D pipelines may invest in better disclosure precisely because they have favorable outcomes to report. We flag this at the outset because it is central to the entire claim rather than a loose end—the multilevel slopes reported above are only as credible as the exogeneity of EDQ—and we now confront it directly through three robustness exercises and one identification strategy.
First, we replaced the content-analysis disclosure index with the third-party Bloomberg ESG environmental pillar score for the subsample with available coverage. The EDQ coefficient remains positive and significant at 0.0041 (, ), within 11 percent of the baseline estimate. Second, restricting the panel to 2017–2024 to exclude the pre-policy years yields a slightly larger coefficient of 0.0052, suggesting the relationship may have strengthened in the post-Paris regime rather than weakened—a pattern not inconsistent with the institutional accumulation argument advanced earlier. Because disclosure plausibly acts on innovation efficiency with a delay rather than contemporaneously, we also re-estimated the baseline with one- and two-year lags of EDQ in place of its current value; both lagged coefficients stay positive and significant (0.0049 at one lag, 0.0043 at two), and the gentle decay across horizons is consistent with a signal that is absorbed gradually into capital allocation rather than with a mechanical same-year correlation.
Third, propensity score matching addresses self-selection into high-disclosure status. We split firms into high-EDQ (top tercile) and low-EDQ groups, then matched on Size, Lev, ROA, Age, SOE, and industry-year fixed effects with one-to-one nearest-neighbor matching and a caliper of 0.05. The average treatment effect on the treated comes in at
(
,
), implying a 7.3-point efficiency gain attributable to high disclosure status after controlling for selection on observables:
Balance diagnostics post-matching show standardized mean differences below 5 percent on all matching covariates, down from pre-matching gaps that reached 31 percent on firm size, and the post-match variance ratios all fall inside the 0.5–2.0 band conventionally taken as adequate. The estimated propensity scores for the high- and low-EDQ groups overlap across almost the entire unit interval, so the common-support condition is satisfied without trimming more than the eleven treated observations whose scores exceeded the maximum control score; these were dropped under the imposed caliper. The matched ATT, the balance summary, and the off-support counts are reported compactly in
Table 8 rather than as a separate figure, since the diagnostics add rows rather than a new visual.
Fourth, and most demanding, we instrumented EDQ with the lagged industry-province average disclosure score excluding the focal firm, following a leave-one-out construction that should satisfy the exclusion restriction conditional on industry-year and province-year fixed effects. We do not treat that exclusion claim as self-evident. The threat is that peer disclosure proxies for a common industry-region shock—a regulatory wave or a green-credit pilot—that drives innovation efficiency directly; our defense is that the saturated industry-year and province-year fixed effects absorb exactly those common shocks, so the residual variation in peer disclosure that identifies the coefficient is the idiosyncratic peer-reporting behavior that has no obvious direct channel to the focal firm’s R&D conversion. As a partial falsification, we re-ran the second stage adding contemporaneous industry green-credit volume and provincial environmental-enforcement intensity as controls; the instrumented EDQ coefficient moved by less than 8 percent, which is the pattern one expects when the instrument is not merely picking up a shared policy environment. The leave-one-out instrument is specified as
The first-stage Kleibergen–Paap F-statistic of 47.3 comfortably exceeds the Stock–Yogo weak-instrument threshold, and the second-stage coefficient on instrumented EDQ is 0.0058 (
,
). The IV estimate exceeds the OLS-style baseline, which is consistent with classical measurement error in the disclosure score attenuating the naive coefficient rather than with reverse causation inflating it [
65].
Figure 8 collects the baseline, two cross-level interactions, and four robustness coefficients on a single axis.
Across the seven specifications, the sign never flips and the magnitude varies within a narrow band of 0.0041 to 0.0073, leaving the central finding intact under each diagnostic challenge. To make the verdict on each hypothesis legible at a glance,
Table 9 links every proposition from H1 through H3b to its governing model, the sign and significance of the relevant coefficient, and the resulting conclusion on support.
5. Discussion
Our central finding—that environmental disclosure quality is associated with the input–output efficiency of green innovation rather than with its sheer volume—diverges from much of the existing literature in two respects. Prior studies relying on binary CSR-report indicators or simple word counts tend to report weak or noisy disclosure effects, sometimes attributing them to symbolic compliance. By scoring disclosure across content completeness, verifiability, and forward-looking commitment, we recover an association that is both statistically robust and economically meaningful. The transmission, we should stress, is interpreted rather than directly estimated: the three overlapping channels developed in
Section 2.1—easier access to green-labeled credit that reduces the financing wedge on long-horizon R&D, accumulated reputational capital that improves collaborations with universities and downstream buyers, and intensified external scrutiny that disciplines internal resource allocation—are the explanations we find most credible, but the present design does not partition the effect among them. None of these channels need to operate in isolation. We read the joint movement of all three as the most plausible source of the efficiency association—consistent with the idea that disclosure works not by funding more projects but by reshaping which projects survive the internal selection process.
The governance moderation results contain the paper’s most uncomfortable finding. Ownership concentration suppresses the disclosure–efficiency conversion. This stands in mild tension with the often-cited view that concentrated owners shoulder long-horizon environmental investments. Our reading is that in the Chinese institutional setting, concentrated stakes more often coincide with controlling-shareholder entrenchment than with stewardship, and the latter dynamic dominates in the aggregate. Independent directors and environmentally experienced executives, by contrast, both steepen the slope, but through different mechanisms: independence raises the cost of disclosure–action decoupling, while executives’ environmental background reduces the cognitive friction of redirecting R&D portfolios. The amplification by environmentally experienced executives turns out to be the largest amplifying effect among the four moderators, which strikes us as the most actionable finding from a board-composition standpoint.
Cross-level interactions sharpen the picture further. Industry regulation intensity and provincial marketization both amplify the firm-level slope, but the institutional logic differs. Regulation works through raising the cost of misrepresentation, so disclosure carries more informational weight where the enforcement apparatus is dense. Marketization works through transmission: even credible signals require thick capital and technology markets to convert into resource reallocation. Firms in lightly regulated industries or thinly marketized provinces face a kind of double discount on the value of their disclosure investment, regardless of internal effort. The pattern complicates any one-size-fits-all disclosure mandate.
Stepping back, the methodological lesson concerns variance decomposition. Roughly a third of the total variance in green innovation efficiency sits at industry and provincial layers combined, a magnitude that single-level fixed-effects estimators absorb but do not surface. Pooled OLS would have delivered a coefficient on EDQ that was directionally correct but stripped of the conditional structure that gives it managerial meaning. The HLM specification did not change the headline result so much as it dissolved the false impression that the disclosure–efficiency relationship is uniform across contexts.
The implications branch out in three directions. For corporate management, the priority is not maximizing disclosure word counts but pairing credible disclosure with governance arrangements that absorb the signal—independent boards with industry expertise, executives with environmental fluency, and ownership structures that resist entrenchment. For environmental regulators, the differential effectiveness across regulation intensities argues for tightening mandatory disclosure standards in lightly regulated industries before declaring victory in the heavily regulated ones, where the marginal informational gain is already saturated. For capital market design, the marketization finding suggests that disclosure rules without complementary green finance infrastructure—certified verifiers, transparent carbon pricing, transition-finance taxonomies—will deliver attenuated returns. Each of these implications has limits; none should be read as a universal prescription. But the multilevel evidence at least clarifies the conditions under which each holds.
6. Conclusions
This study set out to map how environmental disclosure quality, conditioned by governance arrangements and embedded in industry and regional contexts, shapes the productivity of green innovation in Chinese listed firms. Using a three-level hierarchical specification on a 2014–2024 panel of 2847 A-share firms, we recover three findings worth restating, stated as the conditional associations that the design can support rather than as established causal effects. Higher disclosure quality is associated with greater input–output efficiency of green innovation, a link for which we advance eased financing constraints, accumulated reputational capital, and intensified external monitoring as the most plausible explanations rather than as separately tested mediators. Governance characteristics moderate this conversion in heterogeneous ways: board independence, environmentally experienced executives, and institutional shareholding amplify the slope, while concentrated ownership dampens it. Industry regulation intensity and provincial marketization further steepen the firm-level relationship through distinct institutional channels—the former by raising the cost of misrepresentation, the latter by thickening the markets through which disclosure signals travel.
The theoretical contribution sits at the intersection of three strands of literature that have rarely been integrated. By embedding disclosure, governance, and green innovation efficiency within a single multilevel framework, the paper recasts what looked like an average treatment effect into a conditional structure shaped by nested institutional layers. Measurement-wise, the multidimensional disclosure index moves the field beyond binary CSR indicators; methodologically, the random-slope HLM partitions variance that single-level estimators absorb invisibly. Practically, the findings argue for pairing mandatory disclosure rules with complementary governance reforms and green finance infrastructure rather than treating each policy lever in isolation. They also offer a sharper diagnostic for boards: independent directors with industry expertise and executives carrying environmental experience are the governance configurations most likely to convert disclosure investment into measurable abatement payoffs.
Several limitations qualify these conclusions, and three deserve to be stated plainly rather than buried. The sample is confined to Chinese A-share listed firms, a population that is larger, more visible, and more disclosure-pressured than the private and unlisted firms that make up most of the economy; the disclosure–efficiency association we document may therefore be stronger here than it would be where reporting incentives are weaker, and we caution against reading the magnitudes as economy-wide. The instrumental-variable identification, although it survives the falsification checks reported above, still rests on an exclusion restriction that cannot be fully proven—peer disclosure could in principle reach focal-firm efficiency through channels our fixed effects do not capture—so the IV estimate is best read as corroborating rather than clinching. Finally, the cross-level moderators we identify are calibrated to a single institutional setting; whether industry regulation intensity and regional marketization play the same amplifying role in mandatory-disclosure regimes such as the EU, or in less marketized developing economies, is an open question that only cross-national replication can settle. The content-analysis index, despite high inter-rater agreement, also remains coarser than what text-mining or large language models could now extract from full report corpora.
Three avenues for follow-up work appear most promising. Replacing manual coding with NLP-based disclosure scoring would scale the measurement and surface dimensions—tone, hedging, comparability—that human rubrics miss. Cross-national comparison, particularly between Chinese and EU ETS-covered firms, would test whether the cross-level moderators we identify generalize beyond the single institutional setting. Dynamic mechanism analysis through state-space or panel VAR specifications could trace how the disclosure–efficiency relationship evolves with policy regime shifts, separating short-run signaling effects from long-run institutional accumulation.