4.1. Descriptive Statistics
Table 2 presents descriptive statistics for the full sample and the results of tests for differences between urban and rural groups. The treatment group (those with agricultural hukou) accounts for approximately 77.7% of the sample, which is higher than the urbanization rate of the permanent resident population reported by the National Bureau of Statistics. This discrepancy arises because the CHARLS survey is limited to individuals aged 45 and older, and the retention rate of agricultural hukou is significantly higher among this age group than across all age groups [
53]. Since eligibility for the Urban Resident Basic Medical Insurance prior to the integration was tied to household registration status, defining the treatment group based on agricultural hukou allows for the recreation of the actual constraints prior to the policy shock, consistent with the ITT principle.
A t-test comparing mean scores across groups revealed disparities between urban and rural populations across multiple dimensions prior to the integration. In terms of health outcomes, the rural elderly population faces dual vulnerabilities—both physical and psychological: their level of physical impairment (ADL score of 0.459) was higher than that of the urban population (0.317), and their level of psychological depression (CES-D score of 5.542) also showed a statistically significant disadvantage (T = −8.948, p < 0.001), confirming the cumulative effects of unequal distribution of urban-rural medical resources and differences in living environments on vulnerable groups. In terms of economic characteristics and socioeconomic status, the rural group’s average years of education (4.34 years) lagged behind that of the urban group (7.59 years). Yet, their proportion of the workforce (42.4%) exceeded that of the urban group (10.2%). This mismatch of low education and high labor participation reflects the weakness of the rural pension security system, making labor a necessary choice for sustaining livelihoods. In terms of healthcare expenditure, the logarithm of annual OOP expenditure for the rural group (0.362) was lower than that of the urban group (0.431) prior to the integration of the medical insurance systems. Combined with the relatively poorer baseline health status of the rural population, the coexistence of high health risks and low healthcare expenditures reflects the limitations in protection resulting from institutional disparities prior to the integration. Structural differences in baseline characteristics indicate that simple cross-sectional comparisons are prone to being distorted by heterogeneity among sample groups, highlighting the necessity of using a difference-in-differences model to control for pre-existing differences and identify the net effect.
4.2. Baseline Effect Analysis
Table 3 presents the baseline regression results regarding the impact of the urban–rural health insurance integration on health outcomes and financial burdens among the rural elderly population. After controlling for micro-level individual characteristics as well as two-way fixed effects for province and year, the empirical results reveal the dual dividend of the integration policy: financial safety nets and improvements in health.
In terms of the economic burden, the implementation of the policy has alleviated the pressure of medical expenses on rural households. As shown in Column (3), the estimated coefficient of the core explanatory variable Did for the logarithm of OOP medical expenditure (Ln
OOP) is −0.058 (
p = 0.034). After retransformation from the log form, this implies that the integration of medical insurance systems has reduced the absolute OOP medical expenditure of rural respondents by approximately 5.6%. The data confirm that by raising the pooling level, standardizing reimbursement rates, and expanding the scope of coverage, the integration of medical insurance successfully broke through the coverage bottleneck caused by the limited fund pool under the original New Rural Cooperative Medical Scheme (NRCMS), effectively fulfilling the core function of medical insurance in dispersing financial risks. At the same time, policy benefits have also become evident in terms of health outcomes. Since the measures of physical health (ADL) and mental health (CES-D) used in this study are both negative-scaled indicators, the results in columns (1) and (2) of
Table 3 are positive. The coefficient for ADL is −0.092 (
p = 0.003), and the coefficient for CES-D is −0.725 (
p < 0.001); both are negative. This indicates that as financial constraints on healthcare were eased, the rural elderly population gained access to more timely and higher-level healthcare services, thereby reducing the degree of functional impairment and alleviating psychological depression. These findings not only validate the effectiveness of health insurance system integration in narrowing the urban-rural gap in health and welfare but also support the study’s hypothesis H1, which posits that the policy of urban–rural health insurance integration can simultaneously reduce economic burdens and improve health outcomes.
4.3. Testing for Parallel Trends and an Examination of Endogeneity
To test the identification assumptions of the double-difference model and capture the dynamic characteristics of policy, this paper employs the event study approach, using 2015 as the base year to construct interaction terms for regression analysis (see Equation (2) for the model specification).
where
represents the dependent variable (ADL, CES-D, or LnOOP), given that the national-level integration policy, the “Opinions,” was issued in 2016, we set the most recent observation year prior to policy implementation (2015) as the base period (systematic multicollinearity in the model was omitted). We focus on the coefficients of the pre-policy interaction term
and the post-policy interaction term
. The former
is used to test the parallel trends hypothesis, while the latter
is used to capture the dynamic policy effects.
The results in
Table 4 indicate that the standard difference-in-differences model faces a pre-existing challenge related to parallel trends in this sample. In the economic burden dimension, the coefficient of the pre-policy interaction term (
) is 0.042 (
p = 0.009). Prior to the full implementation of the policy in 2016, the trajectories of OOP medical expenditure for the treatment group and the control group were not parallel (see
Figure 1), indicating structural differences in baseline levels between the two groups. Similarly, in the health outcomes dimension, the pre-policy interaction terms for physical and mental health showed marginally significant (−0.043,
p = 0.061) and highly significant (−0.288,
p = 0.015) coefficients, respectively. These pre-policy differences reveal a core endogeneity dilemma in causal identification: the rate of health decline and healthcare-seeking behavior among rural populations are constrained by long-standing socioeconomic disadvantages, resulting in development trajectories that differ from those of urban residents. Therefore, although the baseline regression captures positive signals, the baseline model confounds these with time-varying unobservables inherent to urban-rural heterogeneity.
To more rigorously assess the effectiveness of the PSM-DID identification strategy,
Table 5 and
Figure 2 present the event study estimates for the matched sample. After controlling for observable characteristic differences, the pre-policy coefficient for out-of-pocket costs (Ln
OOP) decreased to 0.034 (
p = 0.117), losing statistical significance and satisfying the parallel trends assumption.
In the health outcomes dimension, the pre-policy coefficient for mental health (CES-D) was −0.226 (p = 0.044), and the pre-policy coefficient for physical health (ADL) was −0.048 (p = 0.076). The pre-intervention differences across these two dimensions were not fully eliminated after matching; however, when comparing the relative magnitudes of the pre-intervention differences and post-intervention effects, the CES-D coefficient after the policy intervention (2018) was −0.746 (p < 0.01), with an absolute value approximately 3.3 times that of the pre-intervention difference (0.226); the post-intervention coefficient for ADL was −0.109 (p = 0.002), approximately 2.3 times the pre-intervention difference (0.048). The magnitude of the post-intervention effects was significantly greater than the baseline differences remaining after matching, and the post-intervention coefficients were consistent with the baseline regression results in terms of direction and significance, indicating that the causal identification results for the matched sample were generally robust.
4.4. Robustness Tests
To verify the reliability of the estimates from the baseline model, this paper conducted multiple robustness tests from various perspectives, including sample sensitivity, placebo tests, and PSM (the results are reported in
Table 6).
4.4.1. Sample Sensitivity and Policy Timing Tests
To rule out the influence of specific sample characteristics and differences in policy implementation timing, this study conducted two restriction tests: (1) Excluding samples from municipalities directly under the central government. Given the unique characteristics of these municipalities in terms of healthcare resource endowments and the size of their medical insurance funds, Panel A excluded samples from the four municipalities directly under the central government. Regression results show that the coefficient of Did on Ln OOP increased to −0.066 and was significant at the 5% level. This confirms that the conclusion regarding the reduction in financial burden is not driven by outliers in large cities and is broadly applicable nationwide. (2) Exclusion of samples with insufficient policy exposure. Given that the policy was coordinated and implemented by local governments, some provinces did not complete the integration until as late as 2018, resulting in insufficient policy exposure duration for respondents. In Panel C, after excluding the aforementioned provinces with delayed implementation, the coefficient of the DID interaction term remained stable (−0.048, p = 0.087). The data confirm that, after controlling for the interference of delayed implementation, the financial burden-reduction effect of health insurance integration remains valid, establishing the robustness of the policy’s long-term efficacy.
4.4.2. Placebo Test and Counterfactual Tests
To rule out the influence of omitted variables and pre-existing endogenous trends, this study conducted two falsification tests: (1) a counterfactual policy timing test. In Panel B, the policy implementation date was artificially moved forward to 2015, and only the actual pre-policy data (2013 and 2015) were retained in the full sample for a counterfactual regression. After excluding the years of the actual policy shock, the estimated coefficient of the dummy policy variable (Did Fake) on the economic burden was significantly negative (−0.040,
p = 0.013). This counterfactual result, together with the findings in
Section 4.3, demonstrates that spontaneous fine-tuning of the original NRCMS coverage level had already led to a certain baseline trend of reduced financial burden prior to the implementation of the national top-level policy. However, the absolute magnitude of this dummy variable coefficient (−0.040) is smaller than the long-term effect of the actual policy in the baseline model (−0.058). From the perspective of this difference in magnitude, this confirms that only the actual policy of integrating the two schemes possesses a breakthrough impact in reducing the financial burden. (2) Randomized placebo test (
Figure 3). To verify whether concurrent random macroeconomic shocks caused the effect, this study generated 500 “pseudo-treatment groups” from the full sample and repeated the regression analysis. The kernel density plot shows that the simulated dummy coefficients follow a standard normal distribution with a mean of 0, and the vast majority of
p-values are greater than 0.1. In contrast, the estimated Did coefficient (−0.058, shown by the red dashed line) deviates significantly from the core distribution range of the dummy coefficients, falling into the left-tail region. This robust spatial falsification rules out the possibility of random coincidence.
4.4.3. Cross-Model Validation and Selection Bias Elimination
The baseline model did not fully pass the parallel trends test. To address selection bias, this study employs PSM-DID for cross-model validation. Through 1:1 nearest-neighbor matching at the baseline period (2013), the standardized deviations of the vast majority of covariates were reduced to within 10% (
Figure 4), thereby constructing a balanced panel dataset with high comparability.
Table 6 (Panel D) shows that after controlling for group differences in observable characteristics, the positive effects of DID on physical health (−0.082) and mental health (−0.620) remain significant. Meanwhile, the coefficient for financial burden remained consistent in both magnitude and direction (−0.037). Although statistical significance decreased due to the loss of some samples during matching, the overall results, when combined with the aforementioned multidimensional test matrix, still corroborate the empirical conclusion that the integration of health insurance systems alleviates the financial burden of medical care.
4.4.4. Test of Excluding Household Income as a Control Variable
Table 5, Panel E reports the results of a robust regression after excluding the log of household income. As shown in the table, after excluding the income variable, the estimated coefficient of the core explanatory variable Did on economic burden (Ln
OOP) is −0.061 and is significant at the 5% level; the estimated coefficients on physical health (ADL) and mental health (CES-D) are −0.099 and −0.706, respectively, and both are significantly negative at the 1% level. Compared to the baseline model that includes household income (with corresponding coefficients of −0.058, ADL: −0.092, and CES-D: −0.725), the core explanatory variable exhibits high stability in terms of coefficient direction, magnitude, and statistical significance. The baseline estimates in this study are not sensitive to the specific settings of control variables, and potential “bad control” biases do not pose a substantial threat to the core causal inferences.
4.5. Mechanism Testing
To analyze the micro-level transmission mechanisms through which the integration of medical insurance systems reduces the burden of medical expenses, this paper further examines two dimensions: healthcare utilization behavior and economic risk mitigation (
Table 7).
The results of the mechanism tests reflect the robust characteristics of the policy implementation process. The data in Columns (1) and (2) show that the integration did not lead to significant changes in the probability of outpatient visits or hospital stays among rural residents, indicating that the policy expansion did not result in excessive utilization of medical resources by micro-level agents. Against this backdrop, the results in Column (3) show that the probability of rural households incurring CHE decreased by 1.9% (p = 0.004). These findings effectively address the competitive hypothesis proposed earlier. Without significantly increasing the total volume of services in the healthcare system, the price compensation effect of the health insurance integration successfully counteracted the potential demand-release effect. Rural households did not experience the retaliatory rebound in healthcare consumption feared in Hypothesis H2b, and the empirical results support the pure financial protection pathway described in Hypothesis H2a.
4.6. Heterogeneity Analysis
The distribution of the inclusive benefits resulting from institutional integration across different groups is a key factor in determining the fairness of the policy. This paper conducts a heterogeneous regression analysis based on educational attainment and age structure (
Table 8 and
Figure 5).
The analysis results reveal distinct patterns in the scope of policy impact. On the one hand, regarding socioeconomic status (SES), the reduction in healthcare costs for the low-education group and the high-education group was highly consistent in magnitude (–0.046 and –0.044, respectively). The universal nature of the cost-reduction effect across different educational groups fails to support the hypothesis in H3a that individuals with absolute essential needs would exhibit a stronger behavioral response. On the other hand, intergenerational heterogeneity was significant. The policy benefits were primarily evident among middle-aged and older adults (45–74 years old, Did = −0.067, p = 0.024). In contrast, the limited benefits for those aged 75 and older corroborate the core logic of Hypothesis H3b. That is, universal health insurance struggles to fully penetrate the non-medical financial barriers arising from aging and disability, leading to a decline in the effectiveness of policy benefits among the most vulnerable elderly population.