1. Introduction
The opening of government data has emerged as a key institutional arrangement in market-oriented reforms of data factors. In recent years, China has introduced a series of policies on data factors and digital government, prompting governments at various levels to launch public data platforms in phases and gradually consolidate and release data previously scattered across departments and administrative tiers. This institutional experiment raises an important empirical question: can government-led information disclosure genuinely improve resource allocation efficiency in markets? Existing research has accumulated findings along two relatively separate lines. Studies on open government data show that it enhances government transparency, fosters government–society interactions, and provides a data foundation for innovation [
1,
2,
3]. Meanwhile, the resource misallocation literature documents that factor allocation distortions across firms are a major source of aggregate total factor productivity (TFP) losses in China’s manufacturing sector and that eliminating such misallocation could substantially raise TFP [
4,
5,
6]. A small but growing literature has begun to connect these strands; Xu et al. [
7] directly examine open government data and firm-level allocative efficiency, but the systematic quantification of allocative distortions within a structural framework and the decomposition of the direction of the policy effect remain to be developed.
Open government data may improve factor allocation by reducing information asymmetry and transaction costs. For firms, publicly available government data help identify market demand, understand local industrial policies and regulatory rules, and assess regional infrastructure and labour supply, thereby informing investment, production, and hiring decisions. For financial institutions, such data enhance their understanding of firms’ operational conditions and regional risks, improving the efficiency of credit evaluation. In the labour market, public information on government services, employment opportunities, and environmental quality shapes labour mobility and occupational choice. These channels suggest that open government data is not merely a governance reform but may systematically affect factor markets by reshaping the information environment. Prior literature has examined the role of digital technologies—such as digitization and data analytics capabilities—in promoting resource allocation [
8,
9,
10]. Yet open government data, as a government-led data supply with strong externalities and broad coverage, may generate incremental effects beyond firms’ proprietary information systems. This possibility awaits rigorous empirical testing.
Building on these studies, we exploit the staggered rollout of city-level government data platforms in China as a quasi-natural experiment. Using panel data on listed manufacturing firms from 2011 to 2024, we employ a staggered difference-in-differences (DID) design to estimate the effect of open government data on firms’ resource allocation efficiency. Following the analytical framework of Hsieh and Klenow [
4], we construct a firm-level measure of resource misallocation based on the deviation of a firm’s actual scale from its optimal scale. We find that the launch of government data platforms significantly reduces the degree of firm-level resource misallocation. This result is robust to propensity score matching, exclusion of concurrent policy shocks, and double machine learning. Channel-level tests do not find statistically significant transmission through government transparency, firm-level digital innovation, or the reallocation of fiscal subsidy resources. Analysis of the direction of the effect reveals that the policy primarily curbs inefficient investment by over-allocated firms rather than alleviating financing constraints faced by under-allocated firms. Heterogeneity analysis indicates that the effect is more pronounced in industries with high digital intensity, in smaller cities, and in regions with less developed digital economies, suggesting that open government data helps bridge the “digital divide.”
This paper makes three main contributions. First, building on the direct evidence in Xu et al. [
7], we construct a firm-level measure of resource misallocation within the Hsieh–Klenow framework and provide causal evidence that open government data reduces allocative distortions, using a staggered DID design together with heterogeneity-robust estimators to address the bias of conventional two-way fixed effects under staggered adoption. Second, we decompose the average effect into a “corrective effect” on capital-overallocated firms and a significant worsening of allocative efficiency among expansion-needing firms, embedding the effective-investment channel documented by Wang et al. [
11] within a full allocative-efficiency framework. Third, we systematically test three potential channels—governance transparency, digital innovation, and subsidy allocation—and report unsupported channels as such, and we provide formal heterogeneity tests of group differences. From a sustainability perspective, reducing factor misallocation limits the wasteful use of capital, labour, and energy inputs, so evidence on how data-governance institutions reshape factor allocation speaks directly to the sustainable development agenda (e.g., SDG 8 and SDG 12).
6. Further Analysis
6.1. Mechanism Analysis
This section tests three potential channels through which open government data may affect firms’ resource allocation efficiency: governance transparency, digital innovation, and fiscal resource allocation. It is important to note that because the structural characteristics of the observable data vary across these mechanisms, we employ identification strategies tailored to each channel rather than applying a uniform mediation regression framework. Each mechanism’s analytical logic and regression approach are described below.
6.1.1. Test of the Government Governance Transparency Channel
Governance transparency is hypothesized to affect resource allocation efficiency through two pathways. First, it reduces information asymmetry between governments and market participants, enabling firms to make more accurate investment and operational decisions. Second, it strengthens external oversight of public resource allocation, improving the efficiency of fiscal spending. To examine this channel, we estimate the reduced-form effect of the data platform on the city-level government transparency index (GovTrasp) and, following the suggestion to treat transparency as a moderator rather than a mediator, test whether the policy effect on resource allocation efficiency varies with the level of government transparency.
Table 23 reports the results. Column (1) uses GovTrasp as the dependent variable: the DID coefficient is 2.861 but is not statistically significant once standard errors are clustered at the treatment (city) level (t = 1.35,
p = 0.180). Column (2) adds the contemporaneous interaction between DID and GovTrasp to the efficiency equation: the DID coefficient is −0.177 (t = −1.53), the interaction coefficient is 0.001 and statistically insignificant (t = 0.90), and the coefficient on GovTrasp itself is also insignificant (t = −0.57). The contemporaneous interaction is therefore not statistically significant.
Column (3) instead interacts DID with the pre-treatment baseline level of government transparency (GovTrasp_pre, the city average before platform launch; the full-sample average is used for never-treated cities). The interaction coefficient is 0.006, significant at the 5% level (t = 2.57, p = 0.011), indicating that the efficiency-improving effect of open government data is significantly stronger in cities with lower initial transparency; evaluated at the sample mean of the baseline index (49.6), the total DID effect is −0.103. Although the contemporaneous interaction is insignificant, the pre-treatment moderation test provides direct evidence that the transparency environment shapes the allocative effect of open data, consistent with the interpretation that open data delivers larger marginal benefits where pre-existing information channels are weaker, mirroring the heterogeneity results for regional digital development.
6.1.2. Test of the Digital Innovation Empowerment Channel
Digital technology is a potential micro-level transmission pathway through which open government data may affect firms’ resource allocation efficiency. Following the approach of the innovation channel, we use the number of digital invention patent applications (digpat) as a firm-level measure of digital innovation capability. This indicator is defined as ln(digital invention patent applications + 1), with digital categories identified by screening patent titles and IPC codes. Compared with text-mining-based digital transformation indices, digital invention patent data are subject to substantive examination and thus carry higher technical credibility.
To examine this channel, we estimate the reduced-form effect of open government data on firms’ digital innovation output without including the mechanism variable in the outcome equation.
Table 24 reports the results: the DID coefficient is 0.024 but is not statistically significant once standard errors are clustered at the treatment (city) level (t = 0.82,
p = 0.415). Although the point estimate is positive and consistent in sign with H3, the digital innovation channel is not supported by the data under treatment-level inference, and we report the channel as unsupported rather than claiming partial mediation.
6.1.3. Test of the Subsidy Resource Allocation Optimization Channel
The allocation of fiscal subsidies across firms is a key institutional factor influencing resource allocation efficiency. We examine whether the policy shifts subsidy allocation toward more innovative or higher-R&D-intensity firms, using two reduced-form interaction specifications.
The first specification uses the logarithm of subsidies received (lnsubsidy) as the dependent variable and interacts the DID indicator with lagged firm innovation (L.inno). Column (1) of
Table 25 reports the results: the coefficient on DID × L.inno is −0.037 and statistically insignificant (t = −0.91). The second specification uses subsidy intensity (subsidyratio1) as the dependent variable and interacts DID with lagged R&D intensity (L.rd_asset_ratio). Column (2) reports the results: the coefficient on DID × L.rd_asset_ratio is −0.0001 and statistically insignificant (t = −0.47). Neither specification indicates that the policy significantly redirects subsidy allocation toward more innovative or higher-R&D-intensity firms, and the subsidy allocation channel is therefore not supported by the data.
6.2. Direction Identification of Resource Allocation Efficiency
The baseline regression results indicate that open government data significantly improves resource allocation efficiency among manufacturing firms. However, these findings do not yet answer a more nuanced question: which type of resource allocation distortion does open government data primarily address? Specifically, does open government data improve allocative efficiency by curbing excessive investment by capital-overallocated firms—a “corrective effect”—or by alleviating financing constraints faced by capital-underallocated firms—a “relief effect”? The answer to this question is crucial for understanding the micro-level mechanisms through which open government data improves resource allocation efficiency.
Following the analytical framework of Hsieh and Klenow [
4], we decompose the direction of resource misallocation along two dimensions based on the distortion measures derived from the model: the capital distortion direction and the output expansion direction. Both classifications are directly constructed from the distortion measures defined in
Section 4.2.1, providing two complementary identification approaches.
First, the capital distortion direction. Recall from Equation (
7) the definition of capital distortion:
This indicator captures the deviation of the effective capital cost borne by the firm from the market equilibrium level. When , the firm faces an effective capital cost below the market average, implying that capital is overallocated to the firm (overK = 1). When , the firm faces an effective capital cost above the market average, implying that capital is underallocated (underK = 1, with overK = 0).
Second, the output expansion direction. Recall from Equation (
8) the ratio of the optimal output to actual output:
Define . When , the firm’s actual output is below its optimal level, suggesting the firm should expand its scale (expand = 1). When , the firm’s actual output exceeds its optimal level, suggesting the firm should contract its scale (contract = 1, with expand = 0). Note that the dependent variable in the baseline regression, , is derived from the absolute value of this ratio; the directional classification here splits the sign.
It is important to note that these two classification perspectives are complementary but distinct. The overK/underK classification directly identifies the direction of capital factor misallocation, whereas the expand/contract classification reflects the deviation direction of total output (including both capital and labour). The former is a factor-input identification, while the latter is an output-level identification; together they constitute the complete directional identification framework from factor inputs to output.
Because the directional dummies are mechanically constructed from the same optimal-to-actual output ratio that defines the dependent variable, we fix each firm’s classification at its pre-treatment value and hold it constant throughout the sample period. Specifically, for firms whose city launched a government data platform, the classification is taken from the year immediately before platform launch; for never-treated firms, and for treated firms without an observation in the year before launch, the classification is taken from the firm’s first sample year (and, if that value is also missing, the unit-level full-sample mean). This removes the contemporaneous conditioning on the outcome and preserves the direction as a pre-determined firm characteristic. Of the 3164 treated firms in the estimation sample, 1555 (49.1%) lack an observation in the year immediately before launch and are therefore classified using their first sample year, which may postdate platform launch; as a robustness check, we re-estimate the specification excluding these firms, which yields qualitatively identical results (see the note to
Table 26). The main effects of the resulting time-invariant dummies are absorbed by firm fixed effects, so the regressions identify the differential effect through the interaction with DID.
Table 26 presents the directional identification regression results on the full estimation sample (28,682 firm–year observations clustered in 249 cities). Column (1) uses the pre-treatment capital overallocation dummy (overK) as the grouping variable. The DID coefficient for capital-underallocated firms (overK = 0) is 0.000 and statistically insignificant (
); the interaction DID × overK is −0.127, significant at the 1% level (
). The formally tested linear combination for capital-overallocated firms (overK = 1) is −0.127, significant at the 1% level (
,
). This suggests that open government data primarily alleviates the efficiency loss from excessive investment by capital-overallocated firms. Column (2) uses the pre-treatment output expansion dummy (expand). For firms not classified as needing to expand (expand = 0), the DID coefficient is −0.126, significant at the 1% level (
); the interaction DID × expand is 0.512, significant at the 1% level (
). The linear combination for expansion-needing firms (expand = 1) is 0.386 and statistically significant (
,
). The policy therefore has a significant corrective effect on capital-overallocated firms, while it significantly increases the deviation from the optimal scale for firms classified as needing to expand.
These results carry clear economic implications: open government data improves resource allocation efficiency primarily by strengthening market discipline to curb excessive investment by capital-overallocated firms—a “corrective effect.” The output-expansion direction shows a parallel asymmetry: the policy significantly reduces misallocation among firms not classified as needing to expand, while the total effect for expansion-needing firms is positive and statistically significant (0.386,
), indicating that the policy significantly increases the deviation from optimal scale for firms that should be expanding rather than relieving their constraints. This finding is consistent with the theoretical logic that open government data reduces information asymmetry in factor markets: by making public information more accessible, the policy enables market participants and external monitors to better evaluate firms’ investment decisions, disciplining the investment behaviour of overcapitalized firms. It is worth noting that this mechanism operates at the level of the information environment in factor markets rather than through the specific intermediation channels examined in
Section 6.1, none of which exhibit a statistically significant direct effect; the directional evidence instead indicates that the efficiency gains materialize through the disciplining of overinvestment, rather than through a direct relaxation of financing constraints.
7. Conclusions
This study exploits the staggered rollout of city-level government data platforms in China as a quasi-natural experiment and draws on panel data of listed manufacturing firms from 2011 to 2024 to systematically examine the effect of open government data on firms’ resource allocation efficiency. The baseline regression results show that the launch of government data platforms significantly reduces the degree of firm-level resource misallocation, with the efficiency-improving effect concentrated among capital-overallocated firms identified by the directional analysis. This finding is robust to propensity score matching, exclusion of concurrent policy shocks, double machine learning, and an alternative investment-efficiency specification in which the dependent variable is the signed residual from the McNichols–Stubben investment expectation model (DID = −0.00190, t = −2.03, p = 0.043).
Mechanism analysis provides channel-level evidence on how open government data affects resource allocation efficiency. The governance transparency channel is not supported by the data: open government data does not exhibit a statistically significant effect on city-level government transparency once standard errors are clustered at the treatment level. A complementary moderation test indicates that the efficiency-improving effect is significantly stronger in cities with lower initial government transparency (interaction coefficient 0.006, t = 2.57); we interpret this pattern as reflecting larger marginal benefits of open data where pre-existing information channels are weaker, rather than as evidence that transparency is a transmission mechanism. The digital innovation channel is not supported: the reduced-form effect of open government data on firm-level digital invention patent applications is positive but not statistically significant under treatment-level clustering. The subsidy resource allocation channel is not supported: neither the responsiveness of subsidy allocation to firm innovation nor to R&D intensity changes significantly after the platform launch under treatment-level clustering.
Several considerations may account for the absence of statistically significant effects along these intermediate channels despite the significant baseline effect. First, the transparency index, digital patent counts, and subsidy measures capture only specific, observable aspects of the relevant mechanisms; the allocative benefits of open data may operate through broader, less readily measurable improvements in the information environment that are not fully reflected in these indicators. Second, firm-level adjustments along these channels may take longer to materialize than the sample period permits. Third, the reduced-form specifications and moderation tests we employ are designed to detect average changes, and they cannot rule out heterogeneous responses that offset one another at the mean. These considerations qualify, rather than overturn, the interpretation that the efficiency gains operate primarily through the disciplining of inefficient investment documented in the directional analysis.
Directional identification reveals a clear asymmetry in the resource allocation effect of open government data. Using directional classifications fixed at pre-treatment values, the policy primarily reduces efficiency losses among capital-overallocated firms through a “corrective effect” that curbs excessive investment, while significantly increasing the deviation from optimal scale for firms classified as needing to expand, suggesting that the mechanism operates through disciplining inefficient investment rather than relaxing financing constraints.
Heterogeneity analysis conducted across three dimensions—firm characteristics, industry characteristics, and regional environment—uses formal full-sample interaction tests to compare groups. Differences by ownership, subsidy dependence, labour quality, and analyst coverage are not statistically significant under city-level clustering. Heterogeneity is statistically significant only for industry digital intensity, city size, and regional digital economy development: the efficiency-improving effect is concentrated in high-digital-intensity industries, smaller cities, and regions with lower levels of digital economy development. Together, these patterns indicate that open government data helps bridge the regional “digital divide” rather than exacerbating it [
60], and that its allocative benefits are shaped by a combination of firm capabilities, industry conditions, and the surrounding information environment.
These findings carry several policy implications. First, governments should continue to advance the construction of open government data platforms, steadily improving dimensions such as data standardization and machine readability. Open government data directly enhances resource allocation efficiency in factor markets, and the directional evidence indicates that this improvement operates primarily by disciplining the investment behaviour of over-allocated firms, providing empirical support for deepening market-oriented reform of production factors. Second, the documented asymmetry between the “corrective effect” on over-allocated firms and the significant worsening of the deviation from optimal scale among expansion-needing firms implies that open data policies are better suited to disciplining inefficient investment than to easing financing constraints. Complementary policies—such as credit market reforms and targeted financial support—remain necessary to address the distinct frictions faced by firms that should expand but remain constrained. Third, local governments should recognize that the benefits of open government data depend on complementary conditions. Coordinating data platforms with talent development programs and information intermediaries can help firms make more effective use of open data. Fourth, firms should regard open government data as a strategic opportunity for digital transformation and innovation-driven development, increasing investment in data analysis capabilities and R&D talent to fully realize the potential benefits of open government data policies.
From the perspective of sustainable development, the findings of this study carry implications that go beyond firm-level efficiency. Resource misallocation wastes scarce capital, labour, and energy inputs and raises the resource and emission intensity of production, so improving the efficiency of factor allocation is itself a sustainability objective. By providing causal evidence that institutionalized data openness curbs inefficient investment and disciplines the allocation of public and private resources, this study suggests that open government data may serve as a low-cost, scalable instrument for sustainable resource governance. These results speak to the socio-economic dimension of sustainability and complement the United Nations Sustainable Development Goals, in particular decent work and economic growth (SDG 8); industry, innovation, and infrastructure (SDG 9); and responsible consumption and production (SDG 12). More broadly, the study illustrates an integrated approach in which digital governance, market efficiency, and sustainable development reinforce one another: data openness improves the information environment of factor markets, which in turn reduces the wasteful use of resources and supports inclusive, higher-quality growth. These insights are relevant not only for China but also for other developing economies designing data-governance reforms as part of their sustainability strategies.
Several limitations should be acknowledged. The Hsieh–Klenow industry benchmarks are computed from listed manufacturing firms only, a highly selected group, so the estimated level of misallocation may not generalize to the full manufacturing population. The output elasticities in Equation (
3) are estimated by industry–year OLS and are subject to the standard simultaneity problem; our control-function-based alternative yields a similar treatment effect (−0.115,
). Constant returns to scale are imposed through the normalization in Equation (
4), although the restriction cannot be rejected in most industry–year cells. VAT payable is estimated from cash-flow data using statutory rates. These concerns are addressed with the sensitivity analyses reported above.
Relative to closely related work on open government data and resource allocation, this study contributes by constructing a firm-level Hsieh–Klenow misallocation measure, addressing staggered-adoption bias with heterogeneity-robust estimators, and identifying the corrective rather than relief nature of the policy effect. In summary, this study provides causal evidence on the economic and social consequences of open government data from the perspective of resource allocation efficiency. As the market-oriented reform of production factors deepens, high-quality open government data has become an important institutional tool. Through well-designed institutional arrangements and enhanced coordination among open data policies, corporate strategies, and digital infrastructure, China can further improve firms’ resource allocation efficiency under resource constraints and promote higher-quality economic development.