3.1. Model Specification and Identification Strategy
Before presenting the model, it is important to clarify the empirical objective and identification boundary of this paper. The present study does not define a treatment group and a control group in the DID sense, nor does it attempt to recover a counterfactual Female wage path in the absence of fertility-policy reform. Instead, it models how wave-level fertility-policy regime indicators enter a persistent wage process and how regime-stage wage associations evolve within a dynamic panel framework.
A central limitation is that the policy-stage variables are defined at the survey-wave level. As a result, the estimated twochild and threechild coefficients cannot be fully separated from secular wage growth, macroeconomic shocks, COVID-19-era labor-market disruption, or other institutional changes occurring during the same periods. Full-time fixed effects cannot be included because they would be perfectly collinear with the wave-level policy-stage indicators. Therefore, the coefficients should be interpreted as conditional regime-stage wage associations rather than causal fertility-policy effects.
The empirical strategy uses System GMM to address dynamic endogeneity arising from the lagged dependent variable and to examine whether regime-stage wage associations carry forward through the wage process. However, the dynamic specification does not solve the time-confounding problem created by wave-level policy coding. The exposure-based Female and CBW interaction specifications help examine differential wage associations among more plausibly exposed groups, but they also remain associational because exposure status is not randomly assigned. This identification boundary is maintained throughout the empirical analysis, discussion, and conclusion.
To avoid overstating the empirical design, the interpretation of the estimates follows a strict hierarchy throughout the paper. First, the aggregate twochild and threechild coefficients are interpreted only as broad wave-level regime-stage wage associations, not as identifiable fertility-policy effects. These coefficients may combine fertility-policy regime timing with secular wage growth, macroeconomic shocks, COVID-19-era disruption, sample-composition changes, and other contemporaneous institutional changes. Second, the Female and CBW interaction terms are interpreted as differential regime-stage wage associations for groups more plausibly exposed to fertility-related labor-market expectations, rather than as causal heterogeneous treatment effects. Third, the adjustment factor is used only as a diagnostic calculation to assess whether contemporaneous regime-stage wage associations carry forward through lagged wage dependence. It is not used to support large long-run causal policy claims. This interpretation hierarchy is maintained in the results, discussion, limitations, and conclusion.
3.1.1. Economic Motivation and Dynamic Wage Specification
China’s transition from strict fertility control to a pronatalist policy regime represents a major institutional shift with potential implications for wage determination. The central objective of this study is to examine how fertility-policy regime transitions enter the wage process and whether the associated wage shifts carry forward to a limited extent within a dynamic earnings structure.
From an economic perspective, fertility-policy transitions may be linked to wages through expectation-based adjustment mechanisms under demographic uncertainty. Changes in fertility policy may reshape expectations regarding childbearing, labor-force attachment, and career continuity, thereby affecting individual human-capital investment and labor-supply decisions. At the same time, firms may update expectations about workforce stability and anticipated labor costs, leading to adjustments in wage-setting behavior. These expectation-based responses provide a theoretical mechanism linking fertility-policy transitions to wage dynamics (
Becker, 1964;
Mincer & Polachek, 1974;
Phelps, 1972;
Arrow, 1973;
Goldin, 2014). When wages exhibit persistence, these wage adjustments are not necessarily confined to a single period but may persist to a limited extent over time within the dynamic wage process.
A key implication of this mechanism is that policy-stage wage associations may not be purely contemporaneous. If wages are persistent, wage shifts observed in one period may carry forward into subsequent periods through the wage process itself. To capture this dynamic structure, the analysis begins with a baseline dynamic wage specification:
where
denotes the logarithm of hourly wages,
captures wage persistence, and
is a vector of observed covariates. The parameter
measures the degree of persistence in wages, while
and
capture the average wage associations corresponding to the two-child and three-child policy stages. Because Equation (1) is dynamic, the wage associations linked to regime-stage indicators may remain partially persistent beyond the contemporaneous period through persistence in the wage process. Because the model includes a lagged dependent variable, a mechanical carry-forward calculation can be obtained from the estimated persistence parameter. For a policy-stage coefficient
, this diagnostic calculation is written as
where
denotes a mechanically adjusted association, not a substantively interpretable long-run policy effect. This calculation is reported only to assess whether the estimated lagged-wage dependence would materially change the contemporaneous regime-stage associations. Given the weak and unstable persistence documented in the empirical results, this multiplier should be interpreted as a diagnostic arithmetic transformation rather than as evidence of a meaningful long-run wage effect.
The motivation for employing a dynamic panel framework does not depend on the presence of strong persistence. Rather, the dynamic specification is warranted because current wages may remain partially dependent on past wage realizations, making static estimators potentially biased in the presence of lagged dependent variables and unobserved heterogeneity (
Nickell, 1981;
Arellano & Bond, 1991;
Blundell & Bond, 1998). From this perspective, the empirical contribution of the dynamic framework is diagnostic. It allows the analysis to evaluate whether the estimated wage process displays meaningful persistence. In the present empirical setting, the evidence indicates only modest persistence and limited intertemporal carry-forward.
While Equation (1) captures aggregate wage responses, it implicitly assumes that policy-stage associations are homogeneous across individuals. However, this assumption may be restrictive in the context of fertility-policy reforms. A large body of literature suggests that fertility-related policies may affect men and women differently due to differences in caregiving responsibilities, labor-force attachment, and employer expectations (
Becker, 1964;
Mincer & Polachek, 1974;
Blau & Kahn, 2017;
Albanesi et al., 2023).
Employers may update expectations about Female labor supply and potential career interruptions following fertility-policy transitions, leading to differential wage adjustments across gender groups. As a result, the aggregate associations captured by
and
may mask important heterogeneity in wage responses. To account for this, the baseline specification is extended to allow for gender-specific policy-stage associations through interaction terms:
Equation (3) extends the baseline model by introducing a broad gender-based exposure specification. In this model, interacts with the policy-stage indicators to examine whether wage associations during fertility-policy regime stages differ between women and other individuals. The interaction terms and capture additional wage associations for women during the two-child and three-child policy stages, respectively.
However, is a broad exposure definition. Not all women are equally likely to experience fertility-policy-related expectation shifts. Older women, unmarried women, or women outside the main childbearing and family-formation ages may not face the same employer expectations regarding childbirth, caregiving responsibilities, or career interruptions. Therefore, the Female-interaction specification is useful as a broad gender-based heterogeneity test, but it may not provide the most targeted measure of fertility-policy exposure.
To introduce a more targeted demographic-exposure proxy, this study further constructs a childbearing-women indicator,
. The preferred definition is married women aged 20–39, while a narrower definition, married women aged 20–35, is used as a robustness check. These variables are not intended to directly measure fertility intentions, employer expectations, maternity-related discrimination, or actual childbirth behavior. Rather, they identify demographic groups for whom fertility-policy-related labor-market conditions may be more relevant, based on age and marital status. The preferred
exposure-proxy specification is written as
In Equation (4), and capture additional regime-stage wage associations for the demographic-exposure proxy during the two-child and three-child policy stages, relative to other individuals observed in the same policy stages. Compared with the broad Female interaction model, the specification introduces a more targeted demographic grouping. However, it should be interpreted only as a proxy-based heterogeneity specification. It does not directly identify employer expectations, fertility intentions, discrimination, or causal exposure to fertility-policy reform.
This specification does not transform the analysis into a fully causal difference-in-differences design because the policy-stage indicators remain defined at the wave level, and exposure is not randomly assigned. However, it improves the empirical credibility of the analysis by introducing cross-sectional exposure variation through interaction terms. This helps distinguish aggregate regime-stage wage associations from differential wage associations among women more plausibly exposed to fertility-policy-related labor-market expectations.
3.1.2. Dynamic Endogeneity and Identification Strategy
The inclusion of the lagged dependent variable in Equations (1), (3) and (4) introduces dynamic endogeneity, as it is correlated with unobserved individual-specific effects. In short panels, fixed-effects estimators are biased (
Nickell, 1981), while pooled OLS fails to control unobserved heterogeneity. These issues are particularly relevant in the present context, where the panel is short and unbalanced.
Because the policy-stage indicators vary at the wave level, they do not generate a DID-style treatment contrast within each period. Accordingly, identification in this paper comes from disciplined dynamic-panel variation rather than from a parallel-trends counterfactual design. An important methodological constraint is that full-time fixed effects cannot be included in the specification because they would be perfectly collinear with wave-level fertility-policy regime indicators. This collinearity prevents the separate identification of full-time effects and policy-stage indicators.
Because the policy-stage variables are defined at the wave level, the empirical framework cannot fully disentangle fertility-policy-stage associations from broader macroeconomic and institutional changes occurring during the same periods. This is the central identification limitation of the study. The estimated twochild and threechild coefficients may partly reflect secular wage growth, COVID-19-era labor-market disruption, structural labor-market changes, and other contemporaneous institutional shifts. Accordingly, the estimates should be interpreted as dynamic regime-stage wage associations embedded within broader institutional transitions rather than isolated causal policy effects.
The purpose of the analysis is therefore not to establish a quasi-experimental treatment effect, but rather to examine how policy-stage institutional environments are associated with wage dynamics and exposure-based heterogeneity within a persistent wage process. This interpretation framework is maintained consistently throughout the empirical analysis and discussion sections.
To address concerns regarding omitted macroeconomic shocks, wage growth, COVID-19 impacts, and other time-varying aggregate factors, the robustness analysis later introduces linear time trends, regional time trends, and partial time controls that do not induce perfect collinearity. These exercises are used to assess whether the documented dynamic associations are sensitive to alternative ways of controlling for aggregate time variation. They do not fully eliminate concerns about unobserved macroeconomic shocks, and therefore, the estimates remain interpreted as regime-stage wage associations rather than causal policy effects.
To address dynamic endogeneity, the empirical strategy employs the System Generalized Method of Moments (System GMM), which combines equations in first differences and levels (
Arellano & Bond, 1991;
Blundell & Bond, 1998). The estimator exploits moment conditions based on lagged values of the dependent variable, under the assumption that the error term is not serially correlated beyond first order. These moment conditions require lagged wages to be correlated with current wages but orthogonal to the contemporaneous error term.
The first-differenced form of the model can be written as
where
denotes the exposure indicator used in the extended specifications, including
,
, and
. Because these exposure indicators are time-invariant, their main effects are eliminated in first differences, while their interactions with policy-stage indicators remain time-varying through changes in the policy-stage variables.
Because individual-specific effects are eliminated through differencing, identification of the persistence parameter relies on valid internal instruments derived from lagged values of the dependent variable. In this setting, the relevant variation comes from intertemporal changes in lagged wages across individuals, conditional on observed covariates and the maintained moment assumptions. The validity of this strategy depends on the absence of higher-order serial correlation in the idiosyncratic error term and on the appropriateness of the chosen instrument set.
This study relies on three maintained assumptions for System GMM estimation. First, the idiosyncratic error term should not exhibit second-order serial correlation, which is assessed using the Arellano–Bond AR(2) test (
Arellano & Bond, 1991). Second, the internal instruments should satisfy the overidentifying restrictions, which are assessed using the Hansen test (
Blundell & Bond, 1998). Third, severe omitted time-varying confounding should be limited after conditioning on individual controls and after examining alternative time-related specifications. This third assumption is assessed through robustness checks rather than conclusively verified.
These assumptions allow the dynamic panel model to estimate conditional intertemporal associations within the maintained System GMM framework. They do not by themselves establish causal identification. To reduce weak-instrument and overfitting concerns in the short and unbalanced panel, the preferred baseline specification uses collapsed GMM-style instruments based on lags 1–3 of the lagged dependent variable, whereas the exposure-based heterogeneity and weighting-comparison specifications use the more restrictive lag range of 2–3. This specification-specific instrument design is intended to limit instrument proliferation while retaining sufficient instrument relevance in the short-panel setting (
Roodman, 2009). In all reported specifications, the number of instruments remains below the number of panel individuals to reduce overfitting risk. All empirical analyses were conducted in Stata 18. Dynamic panel estimations were implemented using the xtabond2 command with collapsed instruments and Windmeijer-corrected standard errors.
The policy-stage variables ( and ) are defined at the wave level and treated as exogenous institutional variables. Their coefficients are therefore identified from variation across policy stages, conditional on observed covariates and individual-specific effects. Substantively, these coefficients should be interpreted as regime-stage wage associations within a dynamic framework rather than as individual-level causal treatment effects.
In the broader conceptual specification, interaction terms between gender and policy-stage indicators may also be considered to capture potential heterogeneous responses. Since gender is time-invariant, its main effect is eliminated in first differences, while interaction terms vary over time through the policy-stage variables. In the extended exposure-based specifications, interaction terms between policy-stage indicators and exposure-group indicators are included to examine differential regime-stage wage associations. The Female interaction provides a broad gender-based heterogeneity check, while the CBW interaction introduces a more targeted fertility-exposure definition. Because these interaction specifications remain based on wave-level policy indicators and non-random exposure status, they are interpreted as differential regime-stage associations rather than causal heterogeneous treatment effects.
To ensure reliable inference, instrument proliferation is controlled through collapsed instruments and restricted lag depth (
Roodman, 2009). This approach balances instrument relevance and parsimony, reducing the risk of overfitting and weak-instrument bias that can arise in short panel settings. Two-step estimation is implemented with finite-sample corrected standard errors (
Windmeijer, 2005), balancing efficiency and robustness. Overall, the System GMM framework is used as a disciplined dynamic-panel estimator under maintained moment assumptions, rather than as a design that solves the wave-level identification problem. Its role is to address dynamic endogeneity from the lagged dependent variable and to assess whether lagged wage dependence materially affects the interpretation of regime-stage wage associations. The estimator therefore supports a cautious dynamic interpretation, but it does not establish causal fertility-policy effects or eliminate time confounding.
3.1.3. Weighting, Wage Observability, and Composite Weights
The CFPS is a complex multi-stage survey, and wage observations are not available for all individuals in all periods. Wage data are observed only when individuals are employed and report valid labor income and working hours. As a result, the estimation sample may be subject to non-random selection, which can lead to biased estimates in the dynamic wage equation.
This issue is particularly important in the context of the dynamic and exposure-based specifications, where wage persistence, policy-stage associations, and interaction-based differential associations are estimated. If wage observability is systematically related to individual characteristics or labor-market conditions, ignoring selection may bias the estimated persistence parameter , as well as the policy coefficients and , and the interaction-based associations and .
To address these concerns, the empirical strategy combines survey weights with inverse-probability weighting (IPW). The purpose of this combined weighting strategy is to improve representativeness and to mitigate observed wage-observability differences associated with employment and reporting behavior. The final composite weight is defined as
where
denotes the CFPS survey weight and
is a stabilized inverse-probability weight.
The probability of observing a valid wage is modeled as where is an indicator for wage observability, and is a set of observed covariates affecting employment and reporting behavior. This selection mechanism reflects the fact that wage observations are conditional on employment and reporting behavior, which may be systematically related to individual characteristics. The predicted probability is given by , and the stabilized inverse-probability weight is constructed as where denotes the average predicted probability in period t, serving as a stabilizing factor.
This weighting scheme gives higher weight to observations with a lower predicted probability of wage observability and therefore mitigates observable selection associated with employment and reporting behavior. In the empirical implementation, the composite weights are incorporated into descriptive statistics and regression estimation to improve comparability between the descriptive sample and the estimated dynamic wage model (
Heckman, 1979;
Deaton, 1997;
Wooldridge, 2007,
2010). The IPW component should be interpreted as an adjustment for observed wage-observability differences, conditional on the covariates included in the selection model. It is not correct for unobserved selection, nor does it redefine the standard System GMM moment conditions. Accordingly, the weighting strategy is treated as a pragmatic sensitivity and representativeness adjustment rather than as a source of causal identification.
The IPW selection model is estimated using a survey-weighted logit including survey-wave indicators, such as Female, age, age squared, years of education, and urban residence, to predict the probability of observing valid wage information. Stabilized weights are constructed to reduce variance, and the final composite weight is defined as the product of the survey weight and the stabilized IPW weight. This approach mitigates observed wage-observability differences and improves representativeness, but it does not correct for unobserved selection. The weighting strategy is therefore implemented primarily as a pragmatic finite-sample adjustment for survey design and observed wage observability, rather than as a fully design-based GMM framework or a source of causal identification.
3.1.4. Evidence Interpretation: MBF-Based Inference
This study employs Minimum Bayes Factors (MBF) to evaluate statistical evidence, as proposed by
Goodman (
1999) and
Sellke et al. (
2001). MBF provides an upper bound on the Bayes factor associated with a given
p-value and offers a conservative interpretation of statistical strength, reducing the risk of overstating evidence against the null hypothesis. For practical interpretation, MBF values are calibrated as follows: MBF < 0.1 indicates strong evidence against the null; 0.1 ≤ MBF < 0.3 indicates moderate evidence; MBF ≥ 0.3 indicates weak or no evidence. This approach ensures transparent and disciplined inference and is reported alongside conventional diagnostic statistics for comparability with the existing literature. The notation MBF_SBB refers to the Minimum Bayes Factor calibrated following the Sellke–Bayarri–Berger approach (
Sellke et al., 2001).
3.2. Data, Sample Construction, and Variable Measurement
3.2.1. Data Source and Panel Structure
The empirical analysis is based on data from the China Family Panel Studies (CFPS), a nationally representative longitudinal survey covering multiple waves between 2010 and 2022. To ensure consistency in wage measurement and to accommodate the dynamic and exposure-based specifications, the estimation sample is constructed using the waves 2014, 2016, 2018, 2020, and 2022. This period is selected to capture the key fertility-policy regime transitions while ensuring consistency in wage measurements across survey waves.
The requirement of a lagged dependent variable implies that only individuals observed in consecutive waves can be included, resulting in a short and unbalanced panel. This panel structure directly motivates the use of a dynamic panel estimator with disciplined instrument selection, as discussed in
Section 3.1.
3.2.2. Wage Construction and Transformations
The dependent variable in all dynamic wage specifications is the logarithm of hourly wages. Hourly wages are constructed from annual labor income and working-time information, following standard survey-based procedures. Specifically, annual working hours are computed as weekly working hours multiplied by 52, and hourly wages are defined as annual labor income divided by annual working hours. This construction ensures comparability across individuals and time, providing a consistent measure of wage outcomes for dynamic analysis.
To ensure comparability across waves and consistency with the logarithmic transformation, observations with non-positive income or invalid working-time information are excluded. These restrictions ensure that the wage variable is well defined and comparable across the estimation sample.
3.2.3. Policy-Regime Coding
The key policy-stage variables used across the dynamic wage specifications are constructed as wave-level policy-stage indicators. These variables capture the institutional environment associated with fertility-policy regimes rather than individual treatment status. Specifically, it is defined as an indicator for observations corresponding to the stabilized two-child policy stage, while represents the three-child policy regime. This stage-based coding reflects the interpretation of fertility-policy transitions as macro-level institutional shifts that affect the wage-setting environment across all individuals.
This choice necessarily limits the interpretation of the policy-stage coefficients: they summarize wage shifts aligned with institutional regime stages but may also absorb broader time-related changes occurring in the same periods.
Because these policy variables are defined at the wave level, the empirical specification does not include a full set of time-fixed effects, which would otherwise be collinear with the regime-stage indicators. In this context, the regime-stage indicators are designed to capture institutional phase variation that would otherwise be absorbed by time fixed effects, allowing the analysis to focus on policy-aligned temporal shifts in the wage process. This approach is consistent with the identification strategy outlined in
Section 3.1, where policy-stage variation is used to capture institutional changes in the wage process.
Accordingly, the regime-stage indicators are interpreted as capturing institutional-phase effects rather than isolated causal impacts of specific policy changes. In the empirical results, their estimated coefficients are therefore discussed as regime-stage wage associations within a dynamic specification, rather than as clean individual-level causal policy associations. Policy-stage indicators are constructed according to the available CFPS survey waves and the observed timing of China’s fertility-policy regime transitions. In the empirical coding used in this study, the pre-two-child reference period covers the earlier CFPS waves before the stabilized two-child Phase II stage. The variable twochild equals one for the 2018 and 2020 survey waves and zero otherwise, capturing the latter two-child policy stage observed in the CFPS panel. The variable threechild equals one for the 2022 survey wave and zero otherwise, corresponding to the survey wave after the announcement of the three-child policy (
National Health Commission, 2021). Therefore, the policy-stage variables should be understood as wave-level regime-stage indicators in the CFPS panel rather than as individual-level treatment indicators.
Because these variables vary at the wave level, they are common to all individuals observed in the same survey wave. This coding captures institutional regime-stage timing but does not distinguish individual exposure by marital status, fertility intentions, age, occupation, or gender. Accordingly, the estimated coefficients are interpreted as aggregate regime-stage wage associations rather than individual-level causal treatment effects. The discussion of more exposed groups is used to clarify the theoretical mechanism and the interpretation of possible subgroup patterns, while more sharply identified causal exposure-specific estimates, such as DID/DDD designs with clearer treatment-control contrasts, are left for future research.
In addition to the wave-level policy-stage indicators, the revised analysis constructs exposure variables for heterogeneity analysis. The first exposure variable is , which captures broad gender-based exposure. The second and preferred exposure variable is , defined as married women aged 20–39. This variable is used as a demographic-exposure proxy based on age and marital status. It does not directly measure fertility intentions, actual childbirth plans, employer expectations, maternity-related discrimination, or caregiving responsibilities. Instead, it identifies a demographic group for whom fertility-policy-related labor-market conditions may be more relevant than for the full Female population.
To examine whether the results are sensitive to the age-band definition, a narrower exposure variable, , defined as married women aged 20–35, is also used as a robustness check. The empirical analysis, therefore, distinguishes four specifications: the baseline aggregate policy-stage model, the broad Female interaction model, the preferred CBW20–39 interaction model, and the narrower CBW20–35 robustness model. This structure allows the analysis to separate aggregate regime-stage wage associations from more targeted fertility-exposure patterns.
3.2.4. Core Covariates and Interaction Terms
The vector in Equation (1) includes standard demographic and socioeconomic controls, including age, age squared, years of education, and urban residence. These variables account for observable determinants of wages and improve estimation precision. In the exposure-based specifications, the analysis additionally constructs interaction terms between policy-stage indicators and exposure-group indicators. The Female interaction terms capture broad gender-based differential wage associations, while the CBW interaction terms provide a more targeted demographic-exposure proxy for married women of childbearing age.
Specifically, is defined as married women aged 20–39, and is used as a narrower robustness definition. The corresponding interaction terms, Female × , Female × , × , × , × , and × , are used to examine whether aggregate policy-stage wage associations conceal differential associations among women more plausibly exposed to fertility-policy-related labor-market expectations. These interaction coefficients are interpreted as differential regime-stage wage associations rather than causal heterogeneous treatment effects. They should not be interpreted as direct evidence of employer expectations, fertility intentions, or discrimination.
Table 1 reports the definitions and construction of all variables used in the empirical analysis. The notation in the model corresponds directly to the variables used in the dataset, thereby ensuring consistency between the econometric specification and the empirical implementation.
3.2.5. Sample Construction and Estimation-Sample Reconciliation
Because the descriptive analysis, preferred baseline model, exposure-based heterogeneity checks, and weighting-sensitivity analysis impose different data requirements, the reported samples are not a single mechanically nested sequence.
Table 2 therefore reconciles the analytical samples used across the reported tables. The descriptive sample requires valid wage and weighting information but does not require a lagged wage. The preferred composite-weighted System GMM model additionally requires valid lagged wages, controls, and composite survey-IPW weights. The exposure-based heterogeneity checks use a broader dynamic-panel sample with valid lagged wages, controls, and exposure indicators, while the weighting-scheme comparison uses a stricter common sample that is usable under all three weighting specifications. These model-specific requirements explain the differences in observations and panel groups across tables.
3.2.6. Conceptual Framework
Figure 1 presents the conceptual framework linking fertility-policy regime transitions to dynamic wage associations. Policy-stage changes are theoretically connected to labor-market behavior through expectation-based and institutional mechanisms, including fertility expectations, caregiving expectations, career-continuity concerns, and employer cost or risk expectations. These mechanisms may influence human-capital investment, labor-supply adjustment, statistical discrimination, and wage-setting behavior.
The framework also incorporates a dynamic wage process in which current wages depend partly on lagged wages. This structure allows policy-stage wage associations to operate not only contemporaneously but also through limited intertemporal persistence. Therefore, short-run regime-stage wage associations may exhibit limited carry-forward beyond the contemporaneous period.
In response to the concern that wave-level policy-stage indicators may absorb aggregate time variation, the revised framework explicitly incorporates exposure-based heterogeneity. The broad exposure definition is Female, while the preferred targeted exposure definition is CBW20–39, defined as married women aged 20–39. A narrower CBW20–35 group is used as a robustness exposure definition. These exposure channels motivate the Female × Policy and CBW × Policy interaction specifications used in the empirical analysis.
Finally, the framework accounts for selection into wage observability through survey weighting and inverse-probability weighting. This adjustment is used to mitigate observed wage-observability differences related to employment selection and reporting behavior.
3.2.7. Descriptive Statistics
The descriptive statistics reported in
Table 3 are constructed using the same estimation sample as the baseline regression analysis. The sample is restricted to individuals who are employed and report valid labor income and working hours, so that hourly wages can be consistently defined. Observations with missing or non-positive income or invalid working-time information are excluded before constructing the logarithmic wage measure. These restrictions ensure that the descriptive statistics correspond to the same estimation sample used in the dynamic wage analysis.
All descriptive statistics are computed using the composite weight , defined as , where denotes the CFPS survey weight and denotes the stabilized and controlled inverse-probability weight that adjusts for non-random wage observability and panel attrition. This composite weight combines the CFPS survey weight with the stabilized inverse-probability weight. The survey-weight component accounts for the CFPS sampling design, while the stabilized IPW component is used to mitigate observed wage-observability differences associated with employment and reporting behavior. The weighting procedure is therefore used to improve comparability between the descriptive sample and the estimated model, not to correct for unobserved selection.
This approach improves consistency between the descriptive statistics and the econometric specification. The reported means, standard deviations, and distributions reflect the same sample restrictions, variable definitions, and weighting structure used in the dynamic wage estimation. Accordingly,
Table 3 provides a weighted summary of the estimation sample underlying the empirical analysis.
The descriptive statistics in
Table 3 are computed using the same estimation sample and composite weighting scheme as the baseline dynamic specification. The sample includes individuals with valid labor income and working-time information, ensuring that the log hourly wage is well defined. All statistics are weighted using
which combines the survey weight with the stabilized IPW component to account for survey design and mitigate observed wage-observability differences.
The mean log hourly wage is approximately 2.16, with substantial dispersion (SD ≈ 1.06), indicating sufficient heterogeneity for identifying wage persistence and policy-stage associations. The policy-stage indicators show that about 14.5% of observations fall in the two-child regime and 12.1% in the three-child regime, providing the time variation needed to identify and . The sample has an average age of 42.6 years and an average education of 9.5 years, with 56.8% residing in urban areas. These characteristics reflect a diverse working-age population and support the inclusion of standard controls in the wage equation.