1. Introduction
The perception of a “welfare trap,” where receiving welfare causes an increased probability of receiving future welfare, has long pervaded discussions on welfare programs.
Plant (
1984) writes that some believe that “social welfare programs create their own dependence,” but finds that evidence pointing towards a welfare trap is weak at best.
Levine and Zimmerman (
1996) investigate the welfare trap as a mechanism to explain intergenerational correlation in the welfare program Aid to Families with Dependent Children but also finds limited evidence. Many authors, however, such as
Verho et al. (
2022), who argue that universal basic income can eliminate welfare traps, regard the “welfare trap” concept as a firmly established feature of the welfare system. We ultimately find evidence against this notion.
The idea that a welfare recipient becomes trapped and experiences more welfare spells begs the question: more relative to whom? In administrative program-use data, individuals enter the dataset only upon their first use of the program, so those who never experience first exposure are absent from the sample. This makes direct comparison to a “never-exposed” group impossible. Identification of the first-exposure effect (FEE) therefore requires a structural approach: we recover the unobserved counterfactual distribution from the observed data by imposing two key assumptions. The first is that there is no state dependence in the counterfactual, meaning that the history of visits would have no impact on future visits; outcomes are independent. The counterfactual individual that we envision is one who experiences unemployment and collects welfare but does not learn from the social program itself. The second key assumption is a first-only effect in the factual distribution, meaning that any behavioural change induced by the program occurs only due to the first visit. No new pertinent additional information that could impact the number of visits is presented to the individual.
We interpret “welfare trap” ideology to be that social programs alter the number of visits that would occur relative to the counterfactual. Individuals experience the program, learn, and gain new information. Behaviour is altered in a way that affects the number of visits that would have otherwise occurred in the absence of the learning event. We assume that information has a behavioural effect after the first visit only, and not subsequently, and is the cornerstone in the identification of the first-exposure effect (FEE). This assumption enables us to express the observed factual distribution as a mixture of the counterfactual distribution, recover both distributions, and estimate their difference.
We formally define the FEE as the difference between what factually happens in the presence of first-exposure, and what counterfactually would have happened if individuals were unresponsive to the information presented at first visit: , where is the observed factual data, and is the counterfactual outcome that would occur if the individual were impervious to any kind of learning or information from the first visit. We propose a general class of distributions for and and identify and estimate the FEE, as well as the associated marginal effects or policy impacts. Estimation of the FEE and marginal FEE, as well as their standard errors, is accomplished via an accompanying R package (fee version 1.0.0) fee, available on CRAN.
Analysis of welfare recipients has been a prominent research agenda with important policy implications. This research has primarily concentrated on the frequency and duration of welfare spells and the associated characteristics of welfare recipients, the state of the labour market, and the structure of welfare policy. The aim of this literature has been to develop causal inferences regarding the distribution of spells and their length that will assist in the development of more effective policies to shorten welfare spells and reduce welfare recidivism. We draw on this literature in a general way to illustrate our new approach using a large administrative data set on welfare recipients from Ontario.
The literature on duration modelling is vast and has important applications to the duration of welfare spells. A good survey of the literature is
Van den Berg (
2001). The emphasis is on the explanation of the empirical exit rate and its underlying determinants using mixed proportional hazards models. For example,
O’Neill et al. (
1987) analyze welfare spell duration for young women under the assumption that exits from welfare depend on superior utility outside welfare, where conditional exit probabilities are estimated by spell year, type of exit (marriage or other) and race for a large number of covariates. They find that most welfare spells end in the first year, but, like much of the literature, they do not analyze repeated spells. There are studies that combine analysis of welfare duration and recidivism (
Cao, 1996) and that combine spells to examine total time on welfare and the fraction of total welfare income over a fixed time interval, ignoring spell composition (
Gottschalk & Moffitt, 1994), largely to circumvent more complex analysis of duration models with repeated welfare spells. While we focus on the frequency of spells and not explicitly on duration, we employ regressions with a log offset, describing event counts per unit exposure (welfare spells per time off welfare). This creates an equivalence between duration and frequency analysis (
Laird & Olivier, 1981).
The literature on the frequency of spells is similarly immense and draws on the econometric modelling of count data. The focus in this literature is explicitly on whether individuals experience a welfare spell and, if so, whether and how often that spell is repeated. While our methodology borrows from some of the standard count data models used in this literature, our innovation is in the use of one-altered count data models, which, to our knowledge, have never been employed in this literature. One-altered count data models have been rather confined to the population size estimation literature (e.g., estimating the number of domestic violence victims, illegal immigrants, or animals that have been captured then recaptured; see
Böhning and Friedl (
2024) for an overview), where they are typically referred to as “one-inflated models” (although the inflation may be negative and thus one-deflating).
We estimate a FEE in social assistance usage, using monthly data on Ontario Works financial and employment assistance, provided by the Ontario Ministry of Community and Social Services in conjunction with Statistics Canada’s Research Data Centres (
Statistics Canada, 2025). Counts of social assistance spells, by individual, are constructed for the 142-month period of observation, resulting in a data set containing 234,320 individuals. The data includes many regressors at the individual level. Social assistance recipients obtain information upon their first exposure that may affect the probability of additional usage. The central tenet of our identification strategy is that all information that may cause a behavioural response is imparted upon the first spell. That is, there is only a first-exposure effect, and no subsequent effects, so the FEE manifests by altering the probability of a single spell. The type of information at first-exposure that we anticipate may invoke a behavioural response that includes (i) learning about one’s preferences towards unemployment, e.g., finding it enjoyable or not; (ii) learning the extent, or lack, of social stigma; (iii) sunk administrative costs, e.g., filling out forms; and (iv) job training that may prevent future unemployment.
We find that the FEE causes a decrease in social assistance usage on average, contrary to the prediction underlying the welfare trap. We estimate that, on average, individuals experience 0.251 fewer unemployment spells over the period, due to first exposure. This amounts to a reduction in unemployment spells of 12.3%. We estimate marginal effects, finding this reduction in unemployment to be much higher for men. We also find that the deterrence effect of first exposure diminishes as the individual becomes more highly educated, perhaps suggesting a higher ability to anticipate the information that is revealed at first visit.
While we apply our method to social assistance usage, more generally, we have developed a framework for evaluating social programs that are repeatedly used by individuals, where the efficacy of the program depends partly on recidivism, and where policy instruments take effect upon first use. In programs that influence their own use, and where only the first use produces an effect, our framework is appropriate.
3. Count Distributions
In order to obtain
(Equation (
11)), we can first specify the counterfactual distribution
. Everything follows from this choice: the factual distribution
and the specific mathematical form of
and the marginal effects. For
, we consider two extremely common distributional choices for count data; the positive Poisson (PP) and zero-truncated negative binomial (ZTNB) distributions. Both of these distributions arise by zero-truncating an underlying Poisson or negative binomial process, respectively.
3.1. Zero-Truncated (Positive) Count Distributions
Since the data is collected administratively, individuals enter the sample only during the first visit. The outcome variable is “zero-truncated” or “positive”: . Such is the case for the Ontario social assistance data that we examine later. Individuals who do not experience a bout of unemployment are not observed; hence, 0 counts are impossible. Administrative data tends to have this characteristic, as only those who come through the door enter the sample.
Zero-truncating the Poisson distribution leads to the positive Poisson (PP) distribution:
and the zero-truncated negative binomial (ZTNB) is
Either
or
, once regressors are incorporated, will be candidates to represent the counterfactual distribution
. That is, either
(and
) or
(and
).
3.2. One-Altered Count Distributions
By Assumption 2, the counterfactual distribution is one-inflated into , thus providing the factual distribution. In this section, we show how count distributions may be one-inflated and arrive at two candidate models for the factual distribution: the one-inflated positive Poisson (OIPP), and the one-inflated negative binomial (OIZTNB). Note that although these distributions are typically referred to as “inflated”, the inflation parameter may be negative and thus the 1s are deflated.
A count distribution may be one-altered with the introduction of an additional parameter
that allows for an extra (or decreased) probability of a 1-count occurring, relative to an underlying count distribution:
where
is the one-inflated distribution and
is the underlying count distribution.
One-inflating
(setting
) leads to the OIPP distribution (
Godwin & Böhning, 2017):
The OIPP distribution (
15) collapses to the PP distribution for
. The
parameter has the same interpretation in both PP and OIPP distributions, and so estimation of
under OIPP also provides the estimation of the PP distribution, which is a subtle but very important point, and is the crux of our identification strategy. A subset of the estimated parameters of the factual distribution are also the estimated parameters for the unobservable counterfactual distribution.
The OIPP distribution (
15) is the factual counterpart to the counterfactual PP distribution. That is, if the observable treatment group follows the OIPP distribution and then the unobservable control group would follow the PP distribution in the absence of the first-exposure effect. Upon the first social assistance spell, the probability of additional spells is altered by
due to information obtained from the social assistance experience.
For example, if is positive, then the first-exposure effect is “shifting” probability mass from the Poisson process to . For , the first-exposure effect leads to a decrease in the probability of subsequent counts. Some counts that would otherwise have been higher are instead 1, and subsequent counts become less probable by . This has the effect of reducing the mean compared to the counterfactual. If , then the first-exposure effect increases usage of the program (e.g., a welfare trap). For , individuals who would otherwise have experienced one visit instead experience more due to the first-exposure effect, and the mean of the distribution increases relative to what it would have been in the absence of first-exposure.
The PP and OIPP distribution may not be a good choice due to their restrictive assumption of homogeneity. If there is unobserved heterogeneity, then the zero-truncated negative binomial (ZTNB) is more appropriate than the PP distribution in order to represent the counterfactual distribution. As such, the observable one-inflated counterpart to the ZTNB distribution is required to represent the observable data. One-inflating
(setting
) leads to the one-inflated zero-truncated negative binomial (OIZTNB) distribution (
Godwin, 2017):
Again, if
suitably explains the observable data, and the assumptions under our identification strategy are satisfied, then the counterfactual distribution is estimated by evaluating
at
. We note again that while the distribution has been named “one-inflated” in the literature, it is actually a “one-altered” distribution, where deflation occurs when
, and where “one-deflation” actually corresponds to the welfare trap.
In
Section 2, we detailed the set of assumptions required in order for the observed data
to follow a one-altered count distribution,
. The validity of first-exposure leading to a one-inflated count distribution is empirically supported by other literature, although the first exposure effect has not yet been considered. A class of one-inflated count data models have been developed in order to address a preponderance of 1-counts often observed in data (one-deflation, as in the case of “encouragement”, is perfectly allowable). Models for one-inflation have recently been discussed by
Godwin and Böhning (
2017),
Godwin (
2017),
Godwin (
2019),
Böhning and van der Heijden (
2019),
Böhning and Friedl (
2021) and
Tajuddin et al. (
2022), for example. In this research, one-inflation is thought to occur due to the observational experience imparting information to the individual being observed. The concept is similar to the “observer effect,” the theory that a phenomenon cannot be observed without affecting the outcome. As soon as an individual is observed, they gain information, but subsequent observational experiences provide no additional information (consistent with Assumption 2).
The information gained at first exposure can lead to a reduction or increase in the probability of the individual being re-observed. If subsequent spells provide no new information to the individual, then the information causes a behavioural response that only perturbs the observed frequency of 1-counts.
Data from
Rossmo and Routledge (
1990) exemplifies the one-inflating phenomenon. Counts of individual prostitution arrests (the same prostitute can be arrested multiple times) are recorded over 23 months. The authors note that “the deterrence effect of a prostitution arrest is very limited” and that “prudent ones seem to learn from the experience, not to avoid street prostitution, but to avoid being arrested... This would lead to a relatively large number of single arrests.”
Godwin (
2017) estimates one-inflation for this data to be 35%. Prostitutes gain information upon first arrest that causes 35% of them not to be caught again. The counterfactual is the number of arrests that would have been made in the absence of the first-exposure learning effect.
Godwin and Böhning (
2017) and
Godwin (
2017) provide many other examples of where one-inflation is plausible.
3.3. Regression Modelling
The final issue is the incorporation of regressors into the models, that is, making
a function of
and
a function of
. This allows individual characteristics and policies to have an effect on the FEE and also allows for the development of marginal effects. Regression modelling of one-altered count data is developed in
Godwin (
2024). Regressors
are linked to
via the canonical link function:
where
is an additional vector of parameters to be estimated. The regressors
are linked to the one-inflation parameter
via a generalized logistic function:
where
is a lower bound that ensures probabilities are never negative and where
is an additional parameter vector to be estimated. For the OIPP model, the lower bound
takes on the form
and for the OIZTNB regression model, the lower bound
takes the form
These link functions and lower bounds ensure that
, and are important to note when deriving the marginal first-exposure effects.
The vector of parameters
and
are estimated via maximum likelihood in the oneinfl R package (
Godwin, 2024). The log-likelihoods are obtained by substituting
(from Equation (
19) or Equation (
20)) into
in Equation (
18) and
(Equation (
17)) and
(Equation (
18)) into the mass functions for OIPP (Equation (
15)) or OIZTNB (Equation (
16)). The specific functional form of the link functions merely ensures that the non-linear optimization routines can estimate the unbounded parameters (
and
) indirectly, rather than the bounded ones (
and
) directly. After obtaining the estimated
and
, the
and
are obtained from Equations (
17) and (
18).
3.4. Alternative Count Data Models
While Poisson and negative binomial models are the most popular, other count data models may be used to represent the factual distribution. The ZTNB and OIZTNB models presented in this paper are akin to the Negbin II model, but Negbin I and Negbin-k types may be more appropriate. Finite Poisson mixture models can also deal with unobserved heterogeneity. In some situations, the beta-binomial model may also be appropriate.
Choice of the count data model to represent the factual distribution is important but is a secondary task to our goal of illustrating how to estimate the FEE using a one-altered model. Whichever zero-truncated count model might be appropriate for the underlying counterfactual distribution, its one-altered counterpart is the model that must be estimated to represent the treatment group. Thus far, only positive Poisson and zero-truncated negative binomial (II) regression models have been extended to allow for one-inflation. Hence, we restrict our attention to these two models and leave the development of a richer suite of models for future work.
The appropriate one-altered count distribution may be discerned through the usual goodness-of-fit tests (such as Pearson’s chi-square distance statistic), information criteria (such as AIC and BIC), Wald, likelihood ratio, and Lagrange multiplier tests. In the results section, we show that the OIZTNB model fits the data much better than the OIPP model and that one-inflation is supported via a likelihood ratio test.
4. First-Exposure Effect (FEE) Under Poisson and Negative Binomial
The first-exposure effect (defined as ) takes on a specific form depending on the distribution for and (these distributions have been denoted f and , respectively). In the previous section, we proposed two distributions for : the positive Poisson (PP) and the zero-truncated negative binomial (ZTNB). These distributions lead to one-altered counterparts for the factual distribution f: one-inflated positive Poisson (OIPP) and one-inflated zero-truncated negative binomial (OIZTNB), respectively. In this section, we derive the first exposure effect under both pairs of these distributions, as well as their standard errors. In addition, we derive the marginal effects and their associated standard errors. These quantities are estimated by maximum likelihood, obtained using the estimates of the parameters of the factual distribution (either or ). Estimation of the FEE and marginal FEEs and their associated standard errors, is made available in an accompanying R package fee.
4.1. FEE Under Poisson
If the positive Poisson (PP) distribution is deemed suitable for the counterfactual distribution, then
and
. The mean of
from the PP distribution is
where
and
are as defined previously. The mean of
from the factual OIPP distribution is
so that the FEE under Poisson is
When
Equation (
23) is negative, the first-exposure effect causes individuals to never return to the program with probability
. Equation (
23) expresses the mean reduction in visits due to first exposure.
By estimating the parameters and in the factual OIPP distribution, the counterfactual PP distribution is also estimated. is the same in both distributions. The first-exposure effect manifests solely in the parameter. The difference between the factual and counterfactual distributions is identified through .
4.2. FEE Under Negative Binomial
By setting the counterfactual distribution to ZTNB, the factual distribution to OIZTNB, and taking the difference in means between the two distributions gives the first-exposure effect under the negative binomial assumption:
4.3. Marginal First-Exposure Effects
Of particular importance is the marginal effect that characteristics and policies in X and Z have on the FEE. That is, we would like to know how much the FEE changes due to changes in the regressors. For example, this would allow one to say that an extra year of education leads to a change of y visits due to first-exposure or that men are y visits more sensitive to first-exposure than women.
In order to determine the impact of a regressor, the derivative of the first-exposure effect can be taken for the continuous regressors: , where is a regressor in X and/or Z. For dummy variables, the difference in predicted first-exposure effects can be evaluated at both values of the dummy.
For the Poisson model, the first derivative of the first-exposure effect with respect to a regressor
in
X and/or
Z is
where
For the negative binomial model, the first derivative of the first-exposure effect with respect to a regressor
in
X and/or
Z is
where
where
and
are as defined previously, and where
is the probability of a 1 count according to OIZTNB distribution.
It is worthwhile to note that when appears only in Z, and not in X, that the marginal FEE and the marginal effects in OIZTNB (or OIPP as the case may be) coincide. If the variable only influences one-inflation (), then it has no bearing on the counterfactual distribution, and the marginal first-exposure effect is simply the effect that the variable has on the factual distribution.
4.4. Estimation of FEE and Marginal FEEs
The FEE and the marginal FEEs are estimated via maximum likelihood estimation (MLE) in the R package fee, which accompanies this paper. First, the factual distribution is estimated via the oneinfl package (
Godwin, 2024). This provides the MLEs
and
(and
if appropriate). MLEs for the FEEs and marginal FEEs derived in this paper are then obtained by evaluating them at
,
, and
.
Provided that the data are i.i.d., the regularity conditions for the OIPP and OIZTNB models hold, establishing the usual desirable maximum likelihood properties for those models. Due to the invariance of MLEs, these properties transfer to the MLEs of the FEE and marginal FEEs. That is, the estimated FEE and marginal FEEs are consistent, asymptotically efficient, and asymptotically normal.
As is common in non-linear models, the FEE and marginal FEEs themselves depend on the values in and . A choice must be made in terms of where to evaluate and . In similar contexts, it is common to evaluate estimators at all data points, giving n estimates, and then take the average. Other common approaches are to evaluate the estimators once at the sample means of the data, or for some specific case. For example, Stata allows for “average effects” (calculating the effect at every data point and taking the average), “effect at means” (calculating the marginal effect at the sample means), and marginal effects at specified values (i.e., a hypothetical or representative situation for the regressors). We follow suit and allow for calculation of an average FEE, an FEE at means, and an FEE at specified values for .
In order to determine whether the estimated FEE is significant, to construct confidence intervals, and to allow for hypothesis testing in general, estimated standard errors for the FEE and marginal FEEs are needed. We opt to use the delta method, since estimation of the one-altered models can be computationally burdensome, making the bootstrap too slow. The Hessian matrix is derived numerically in oneinfl (
Godwin, 2024), and fee obtains the desired transformation via a numerically approximated Jacobian using the numericDeriv function in base R.
6. Summary
We present a framework for estimating a first-exposure effect (FEE) in a social program. The first exposure occurs due to information gained by the individual during their initial experience with the social program, which affects subsequent use of the program. The first-exposure effect is identified by linking the factual distribution to a counterfactual distribution through a parameter . The unobserved counterfactual distribution represents the outcomes that the individual would have experienced in the absence of the information gained during first-exposure. Comparison of the two distributions allows for causal inferences, and we focus on the difference in mean between the two distributions.
The FEE manifests by altering the probability of a 1 count. A central assumption of the identification strategy is that any information that induces a behavioural response occurs during the first visit. Additional observational experiences provide no new information to the individual that would subsequently affect the number of visits. This suggests the use of a one-altered count data model for the factual distribution. The counterfactual distribution is nested inside the one-altered count model and occurs when one-inflation (or deflation) is zero. After stating assumptions that identify the link between factual and counterfactual distributions, we derive the FEE and marginal FEEs under Poisson and negative binomial type models. We estimate the effects through an accompanying R package fee.
We apply the method to Ontario Social Assistance, using data on the number of spells experienced by individuals over a 142-month period of observation. We find that the average first-exposure effect deters recipients from future use by −0.251 on average. That is, we find that recipients would have experienced more social assistance spells in the absence of the information gained from their first spell. We find this effect to be much stronger for men than for women and for those currently attending school. Education has a positive effect on the FEE; the first experience is less of a deterrent for more highly educated people. Importantly, we find the opposite of a welfare trap (with p-value 0).
A policy implication arising from the estimated marginal effects is that the welfare impact of the FEE may be maximized by focusing efforts on deterring the first exposure of groups with specific characteristics. Policies aimed at preventing a first unemployment spell would be more productive for women, for those who are divorced, for those not in school, and for those who are more highly educated.