Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,184 characters · 11 sections · 33 citation commands
Bounds for Treatment Effects in the Presence of Anticipatory Behavior
\thispagestyle{empty} \setcounter{page}{1}
This paper accommodates the matter of anticipation in the analysis of treatment effects, by employing a potential-outcomes framework in a difference-in-differences (DID) model. The concept of anticipation is familiar to researchers in economics and social sciences, as seen in the work of malani2015interpreting, as well as bovskovic2018much, for example. When anticipation occurs, forward-looking units change their behavior in reaction to the possibility of a new policy, and thus, a treatment has an impact before its implementation. Therefore, considering the role of anticipation is crucial when evaluating an economic process and its outcome. However, despite its importance, most available published studies do not formally consider anticipation. People usually make a “no anticipation” assumption, combined with a procedure of dropping data closely before the treatment if this assumption is possibly violated, based on the argument that anticipation occurs only within a fixed time period prior to the introduction of a policy. Even in the few cases in which anticipation is taken into account, the anticipatory behavior is accounted for in a restricted manner, such as an ad-hoc restriction on units' forward-looking behavior like rational or adaptive expectations.
When anticipatory behavior takes place, the identification strategies commonly used with multiple periods, such as the DID model, fall apart. Consider an early retirement incentive (ERI) program for teachers nearing the age of retirement. If those teachers foresee the possibility of retiring early, their behavior might change before the program is introduced. Due to the effect of future treatment status on pre-treatment outcomes, the observable pre-treatment outcomes are no longer drawn from the distribution of potential outcomes, if the treatment never takes place. Individuals will change their responses according to how they expect to be treated in the future. Thus, further information about units' anticipatory behavior is required. However, such information is usually unattainable, because it is generally impossible to observe.
This paper provides novel strategies to build identified sets for treatment effects under assumptions restricting the anticipatory behavior. Easy-to-implement estimation and inference strategies are also provided. I start from a DID model with two time periods, and then generalize it to incorporate more complex models. I provide conditions for partial identification results of causal parameters when the anticipation status is unknown, and I incorporate anticipation in many widely used empirical designs. Employing a potential-outcomes framework, I analyze the treatment and the effects of anticipation based on the treatment rules, the anticipation assignments, and outcomes.
The departure from point identification starts with formulating restrictions on anticipatory behavior. In most cases, I do not have additional information, such as proxy variables, that helps us identify which participants have anticipated the policy change. As a result, I can say nothing about the pre-treatment distortion caused by anticipation. In this paper, I introduce a two-period DID model where anticipation occurs in the first period and the treatment occurs in the second period. Further, I introduce two natural assumptions to help construct bounds for the treatment effect in the absence of such additional information. The first is a bound for the proportion of anticipators within the treatment group. This bound should be available from observed data. It can be a constant, or a parameter that can be estimated. This practice is common in the literature, such as the work of manski2013deterrence. The selection of this bound can vary from application to application, with one possible example being the treatment ratio. I provide models to motivate specific choices of the bound under various circumstances. The second assumption restricts the magnitude of the anticipatory effect. It requires that the absolute value of the anticipatory effect is no larger than that of the actual treatment effect. By doing so, I build a link between the magnitude of an anticipator's reaction and the response to the implementation of the policy. Based on this relationship, an inequality between the treatment effect and the average pre-treatment bias caused by anticipation can be constructed with the help of the proportion of anticipators discussed above. Therefore, I can find a corresponding treatment effect range for the anticipators by characterizing how they react and linking that anticipatory effect to the actual treatment effect. Under these two assumptions, the fraction of units that anticipate the policy change may vary, but the average distortion caused by anticipation is bounded and the parameter of interest is set identified.
As for the implementation purpose, I propose estimation and inference strategies based on easy-to-implement modifications to existing methods. The identification strategy provides an identified set with perfectly correlated and proportional upper and lower bounds. I propose a uniformly valid confidence set for my estimators with some modifications to imbens2004confidence under this specific setup. In their method, the upper and lower bounds of the confidence set are found by extending both sides of the identified set. The extended lengths are proportional to the standard errors of the bound estimators and differ between upper and lower bounds. However, suppose this method is applied directly here. In that case, it may run into a counterintuitive situation in which the confidence set for the treatment effect is shorter when the parameter is partially identified than when it is point-identified. I propose modifying this approach by extending both sides of the identified set by the same length proportional to the larger standard error of the two, which is a natural way to ensure the uniform validity. Analyzing this confidence set also provides researchers with further empirical implications. When the treatment and the anticipatory effect go in different directions, I find a specific range of t-statistics for the zero treatment effect null hypothesis. If the t-statistic obtained when anticipation is ignored falls within this range, the conclusion of whether to reject the null hypothesis does not change when considering anticipation. This confidence set also helps build a framework for sensitivity analysis on certain conclusions of interest by choosing different bounds for anticipation possibility.
I apply the results of this paper to examine the effect of an early retirement incentive program on student achievement. This program is aimed at teachers near the age of retirement and offers them financial incentives to retire before becoming eligible for full pension benefits. If the program is anticipated, eligible teachers may react in advance of its introduction, and such behavior might affect students' grades. The empirical results illustrate the potential pitfalls of failing to consider anticipation in program evaluation: the effect can be greatly overestimated in the worst case. I also conduct a sensitivity check by analyzing the level of anticipation probability one is willing to tolerate while maintaining the consistency of the original conclusion. It shows the conclusion is robust even when about three fourths of target units anticipate.
To permit the incorporation of anticipation in other common empirical setups, I provide several modifications. Instead of focusing only on the pre-treatment effect of anticipatory behavior in the treated group, I discuss the anticipatory behavior in the control group by introducing an imperfect anticipation setup, where individuals make mistakes while anticipating. Post-treatment effects of anticipatory behavior are also discussed. To be consistent with common empirical approaches, generalizations to include covariates, multiple periods, and nonlinear potential outcomes are provided and analyzed in the appendix.
This paper contributes to the literature on causal inference and program evaluation; see abadie2018econometric, athey2017econometrics for example. My paper is most closely related to the work of malani2015interpreting, who discuss anticipatory behavior by interpreting the pre-trend phenomenon as a result of anticipation. The authors propose a parametric time series model in which anticipation is an expectation of the future treatment for everybody, by relying on the rational or adaptive expectation assumption. I incorporate the idea of anticipation in a DID framework with potential outcomes to remove parametric restrictions and allow heterogeneous anticipatory behavior among units. heckman2007dynamic, under a different scenario, present a reduced form dynamic treatment effect model that also permits anticipation, but at the price of imposing further assumptions on the functional structure of the outcome equation.
This paper also contributes to the literature on DID and event-study designs by considering anticipatory behavior. The additional anticipation could have an impact prior to the introduction of a policy. Therefore, the present research is related to the literature aiming at more robust inference and identification strategies that allow for non-parallel trends assumptions, and to papers focusing on pre-trend analysis.
To interpret and deal with observed changes in outcomes prior to a treatment, manski2018right propose a result on partial identification for the average treatment effect under “bounded variation” assumptions. These assumptions relax the parallel trends assumption by allowing for differences within a certain magnitude. rambachan2022more follow the idea that pre-treatment differences in trends are informative about counterfactual post-treatment differences and provide identification and inference results based on several common restrictions of this relationship. freyaldenhoven2019pre propose a method that includes an additional covariate that is correlated with the outcomes through confounds only, and not treatments. ye2021negative propose a partial identification method for treatment effects with two groups of control units whose outcomes exhibit a negative correlation relative to the treated units. In this paper, I interpret pre-trends as a result of unobservable anticipation activities. People may change their behavior because of their anticipation of future treatment. If people have information and may benefit by acting on it before a treatment, anticipation is a reasonable explanation for an observed pre-treatment effect, even when the parallel trends assumption is valid.
This paper is also complementary to the causal interpretation of event-study coefficients; see borusyak2017revisiting, sun2020estimating, de2020two, and goodman2021difference. With a generalization to the longitudinal data, this paper can be regarded as relaxation of the “no anticipation” assumption of these papers. This paper is also related more generally to the partial identification literature. In the present study, partial identification is obtained through moment inequalities, a method that is discussed in molinari2020microeconometrics.
The rest of this paper is organized as follows: Section 2 generalizes the commonly used DID model and introduces the basic setup about anticipation; Section 3 provides extra assumptions and shows readers how to build the identified sets; Section 4 describes estimation and inference; Section 5 provides an empirical application; Section 6 offers further discussion; and Section 7 concludes. The mathematical proofs, together with some additional results, discussions, and generalizations, are collected in the supplemental appendix.
To illustrate anticipation in program evaluation, I consider an early retirement incentive program available for teachers that offers experienced teachers financial incentives to retire before they would be eligible for full pension benefits. Suppose one is interested in the effect of this early retirement incentive program on students' grades. Anticipation from teachers, whether treated or not, can be expected for several reasons in this program. Teachers who anticipate may have received inside information from others, and they can also speculate based on changes that have already happened. Younger teachers ineligible for the program won't react to it regardless of the anticipation status in both cases. However, teachers who anticipate the program and decide to retire early may put in less effort than younger teachers. Such behavior may harm students' grades before the implementation of the early retirement incentive program, and ignoring anticipation can lead to a bias while analyzing the effect of this program. The fact that teachers can anticipate based on unobservable information and adjust their behavior accordingly to gain benefits implies the future treatment will have an effect before its adoption and distort the treatment effect estimation if ignored. An accurate assessment of anticipation is therefore essential for the program evaluation.
First I briefly describe the “canonical” two-period DID model in this section. As a well-understood starting point, this simple setting serves as a good baseline for understanding the approach I use.
Consider a model with two periods $t\in\{0,1\}$ and $n$ units, $i\in\{1,\dots,n\}$. Each unit is assigned an observable binary treatment $D_i$ that takes value $d\in\{0,1\}$ in the second period. The key identifying assumption requires that the treated and control group follow parallel trends in the absence of treatment, and the parameter of interest is the average treatment effect for treated (ATT).
Potential outcomes, defined below, depend on the time period and binary treatment status. The potential outcome for unit $i$ in period $t$ is denoted by the random variable $Y_{it}(d)$. Given a value of the implemented treatment $d$, the observed outcome of unit $i$ at period $t$, $Y_{it}$ can be written as
and the parameter of interest $\mu=\mathbb{E}[Y_{i1}(1)-Y_{i1}(0)|D_i=1]$. For identifying purpose, I need to assume “Parallel Trends” and “No Anticipation”, which require
and
Under these two assumptions, the parameter of interest $\mu$ is identified, and for estimation purposes, I need to further put independent restrictions on the sampling process. The key idea here is that following the parallel trends assumption, one can use the change in the control group to mimic that in the treated group in the absence of treatment and get information about the unobservable potential outcomes for the treated group in the absence of treatment in the post-treatment period. However, as pointed out, this approach requires no anticipatory behavior, which assumes that the post-treatment status should have no impact on pre-treatment outcomes. This assumption may be too restrictive in some situations, for example, the early retirement incentive program mentioned above. Thus, figuring out a way to accommodate anticipatory behavior under this DID setup is important.
To deal with the unobservable anticipatory behavior, I introduce another indicator for anticipation status. Suppose that in the first period, each unit has an unobservable binary anticipation status $A_i$ that takes value $a\in\{0,1\}$. Here, $a=1$ means this unit anticipates the future, and $a=0$ means this unit does not anticipate the future. The potential outcomes can now be written as $Y_{it}(a,d)$. Introducing a second index in the expression of potential outcomes is common when analyzing indirect effects, for example, analysis of spillover effects in vazquez2021identification. By implicitly assuming perfect anticipation, which means the anticipated treatment status should be the same as the actual treatment, I focus only on the pre-treatment anticipatory behavior in the treated group at this time. Later, I discuss anticipatory behavior in the control group and the post-treatment effect of the anticipatory behavior. After introducing another index for the anticipatory behavior, potential outcomes, defined below, can now depend on the binary treatment and anticipation status. I refer to the existence of the latent treatment in the first period as anticipation and the effect of the anticipatory behavior on unit $i$'s potential outcome before the treatment occurs as the anticipatory effect.
As stated above, I focus only on the pre-treatment anticipatory behavior in the treated group now, which means I allow $Y_{i0}(0,1)$ and $Y_{i0}(1,1)$ to be different from each other with no changes on $Y_{i0}(0)$, $Y_{i1}(0)$ and $Y_{i1}(1)$. With a little abuse of the notation, I use the single index expression $Y_{it}(d)$ for the potential outcomes when the anticipation status makes no difference. Because of the existence of anticipatory behavior, the observed pre-treatment outcomes of the treated group are now a mixture of those who anticipate and those who don't. This mixture brings extra difficulties in identification, because one cannot distinguish the anticipators from those who do not anticipate, and further assumptions are needed. To start with, I consider those assumptions that come from the canonical DID model.
Assumption (ref) models the sampling process and states the potential outcomes, treatments, and unobservable anticipation status to be independent and identically distributed across units so that expectations are not indexed by $i$.
Recall that $Y_{i0}(0)$ represents the pre-treatment potential outcome if one will not get treated. This assumption requires future treatment does not make a difference for the pre-treatment outcome in the absence of anticipatory behavior. One can interpret this assumption in a way that anticipation of a future treatment is the only channel through which future events affect the present. In the early retirement incentive program example, this assumption implies there is no difference in the pre-treatment grades for students taught by the same teacher regardless of the teacher's decision about early retirement if he has no anticipation of the program.
The table below shows potential outcomes in the DID model with two periods.
Compared with the commonly used DID framework, the critical difference is that the observed pre-treatment outcome $Y_{i0}$ for the treated group is a mixture of the potential outcomes for those who do not anticipate $Y_{i0}(0,1)$ and those who do anticipate $Y_{i0}(1,1)$ in the treated group in the pre-treatment period. $\mathbb{E}[Y_{i0}|D_i=1]$ is no longer a good measure for the first-period potential outcome without treatment for the treated group.
Under this setup, the parameter of interest I focus on is still the average treatment effect for treated (ATT) with a slight modification: \[\mu_g=\mathbb{E}[g(Y_{i1}(1))-g(Y_{i1}(0))|D_i=1],\] where $g(.)$ is a known measurable real function with $\mathbb{E}\lvert g(Y)\rvert<\infty$. Define the corresponding anticipatory effect for anticipators as \[\tau_g=\mathbb{E}[g(Y_{i0}(1,1))-g(Y_{i0}(0,1))|D_i=1,A_i=1].\] The $g(.)$ function is slightly generalized from the commonly defined ATT. When $g(.)$ is the identity function, $\mu_g$ is the widely used ATT. If $g(.)$ is an indicator function such as $g_{u}(Y)=\mathbb{I}(Y\le u)$, then $\mu_g$ can be interpreted as the change in the probability of the outcomes being no more than a specific cutoff $u$ and can be used to help identify the distribution of potential outcomes. Different choices of this $g(.)$ function lead to different estimators. Introducing the $g$ function enables handling of some nonlinear structures for parameters I am interested in. For simplicity of notation, I write $\mu_g$ as $\mu$ and $\tau_g$ as $\tau$ when $g(.)$ is the identity function.
Although the expression seems to be the same as the parallel trends assumption in the canonical DID model, Assumption (ref) requires that the treatment and control group change following parallel trends before and after the treatment in the absence of both anticipation and the treatment.
Then, I briefly discuss what the commonly used DID estimator estimates without considering anticipation and compare the finding with the parameter that I am interested in. Throughout this section I choose $g(.)$ to be the identity function.
In the DID regression model with two periods, \[Y_{it}=\beta_0+\beta_1 t+\beta_2 D_i+\beta_3 tD_i+\varepsilon_{it}.\] Under Assumptions (ref)-(ref), the coefficient of interest, $\beta_3$, can be written as
If the DID estimator is used directly, it will suffer from a bias equal to the average distortion caused by anticipation. This bias arises from the fact that the observable pre-treatment outcomes for the treated group do not reflect the potential outcomes for them without the treatment. Those who anticipate have already reacted in the first period and deviated from the parallel-trends benchmark. Thus, applying the DID estimator directly suffers from a bias determined by both the proportion of those who anticipate and the magnitude of anticipatory effects. The last equality points out that this parameter can also be written as a weighted average of the treatment effect $\mu$ for those who do not anticipate and the net treatment effect after the adoption of the policy for those who do anticipate $\mu-\tau$. In general, the relationship between this estimand and the treatment effect depends on the sign of the anticipatory effects. Suppose the treatment and anticipatory effects have the same sign. In that case, anticipation will drive the DID estimator toward zero relative to the treatment effect because of contamination. The graph below captures the idea of this distortion. Figure (ref) shows the result obtained by applying the DID estimator directly, whereas Figure (ref) describes the situation that considers anticipation. Anticipation causes the distortion between $\beta_3$ and $\mu$.
Anticipation makes the commonly used DID estimator a mixture of anticipatory and treatment effects. The fundamental difficulty in obtaining identification is distinguishing between those who anticipate and those who do not. This section introduces several assumptions to build upper and lower bounds for treatment effects under different circumstances. Motivations for specific assumptions are provided. The following results link observed outcomes, potential outcomes, treatment assignments, and anticipation status and are used in further discussions.
The analysis above shows that two unobservable variables are contributing to the pre-treatment distortion. One is the possibility of treated units anticipating, $\mathbb{P}[A_i=1|D_i=1]$, and the other is the anticipatory effect for anticipators $\tau_g$. These two variables both need to be analyzed to recover the treatment effect. If a reasonable proxy is available for the anticipation treatment, one can use this proxy variable to measure the anticipation status for each unit. However, such a proxy variable is not always available. To overcome the difficulty of not being able to distinguish people who anticipate from others, I introduce a bounding parameter $\pi\in(0,1)$ that summarizes how the anticipation probability can be bounded. Here, $\pi$ is a parameter that can be obtained from available information, including the treatment assignment and outcomes of units, and it can be either constant or at least estimated from observable terms. Researchers can choose a $\pi$ based on their empirical setups, and I discuss possible choices of $\pi$ in the next section.
Although I cannot observe the anticipatory effect $\tau_g$ directly, I can build a relationship between it and the treatment effect for treated $\mu_g$. $\tau_g$ is caused by people's anticipation of a possible future treatment and behavior before the treatment to gain benefit. People's reactions and behavior are guided by their own guesses of the future policy. On the other hand, $\mu_g$ measures the treatment effect treated people receive when the policy is adopted. This effect happens based on revealed policy and treatment status. For example, suppose someone is going to sell his property at a lower-than-usual price because of anticipation of a possible negative price shock in the future. In this case, he has no reason to accept a price that is even lower than the price when the shock comes. Therefore, one might reasonably expects that the magnitude of treatment effects should be no smaller than that of the anticipatory effect, because the former is a reaction based on known information, whereas the latter one is based on uncertainty. In the example of the early retirement incentive program, this statement requires that the effect caused by teachers who exert less effort because they anticipate a possible early-retirement opportunity is no larger than the treatment effect when the early retirement incentive program is implemented. Because the anticipatory effect and treatment effect may not be in the same direction, I impose restrictions only on the magnitude.
Assumptions (ref) and (ref) help build bounds for the two unobservable terms, the proportion of anticipators among treated and the magnitude of the anticipatory effect separately based on the above assumptions. The parameter of interest $\mu_g$ is partially identified using observed variables, especially with the help of the commonly used DID estimand.
Theorem (ref) points out that the treatment effect is located in an interval where the DID parameter without anticipation is one of its bounds. The other bound is obtained by enlarging or reducing it by a specific ratio depending on the bounding parameter $\pi$ and signs of treatment and anticipation effects. As shown in Figures (ref) and (ref), the distortion happens only within the group of treated and anticipate units, so once the sign of the anticipatory effect is determined, the sign of the bias is also determined. The DID parameter without anticipation is by design one side of the interval. If anticipatory behavior happens in both the control and the treated group, distortions happen in both groups, and the sign of the bias is ambiguous. The distortion bias has a limited magnitude restricted by both the bounding parameter $\pi$ and the treatment effect magnitude, so I can build partial identification results for the parameter of interest based on observables.
Although I impose the bounds for anticipation probability and magnitude restrictions, these assumptions are not the only way to build partial identification results for the parameter of interest with anticipation. Empirical setups may exist in which these assumptions are not reasonable, and researchers would like to impose alternative assumptions, such as bounded outcomes or further conditional independent restrictions. These assumptions are also reasonable under specific situations, such as when the $g(.)$ function I am interested in is bounded by itself. I do not argue that the identified set under one assumption is tighter than the other, so I should choose one of them; rather, different sets of assumptions may be reasonable under different empirical circumstances. Incorporating more combinations of alternative assumptions and providing identification results allow us to incorporate anticipation in more situations and give researchers the freedom to modify assumptions based on the empirical setup. I propose several different combinations of assumptions as well as corresponding upper and lower bounds expressions of the treatment effects in the appendix.
This section discusses several possible choices of the bounding parameter $\pi$ for the anticipation probability among treated units under different setups.
Example 1 $\pi=\pi_0$, where $\pi_0$ is a constant number. If $\pi$ is a constant number, a common upper bound exists for the possibility of anticipation. This choice of $\pi$ may be consistent with the setup where people get treated randomly and receive private information that helps with anticipation. Then, the overall anticipation possibility should be no more than the proportion of people who have access to this private information.
Example 2 $\pi=\mathbb{P}[D_i=1]$. This example states that the possibility of people within the treatment group anticipating does not exceed the proportion of people treated at last. This argument follows the idea that anticipation will happen when a future treatment sends some signals and unobservable information in advance. These signals are the bases for someone anticipating a treatment. Suppose the density of these signals caused by future adoptions of policies is related to the overall scope of the treatment. In that case, using the treated probability to help bound the proportion of people who anticipate is reasonable. In the appendix, I explain this choice and corresponding assumptions in a model where people anticipate from public information.
Example 3 The univariate bound can be modified to incorporate the idea of stratification. Suppose researchers are willing to divide units into several subgroups and allow anticipation behavior to differ among subgroups. In that case, I can choose $\pi$ as a $k$-dimensional vector if I have $k$ subgroups in total. For instance, the vector of assignment can be summarized according to genders or geographical areas, and researchers can get a bound separately for each subgroup. This choice of $\pi$ can also be regarded as a bound conditional on a discrete variable that divides the group based on several categories, and can link to the case with covariates.
Example 4 Suppose anticipation behavior happens among known reference groups for each unit, as mentioned in manski2013identification. In that case, I can choose $\pi$ based on subgroup information. For example, if researchers would like to use the treatment ratio to capture the density of information, and on the other hand, they also believe this kind of interaction only happens among units within a specific geographical distance, they can choose $\pi$ as the treatment ratio for each subgroup defined by the given geographical distance.
Additionally the choice of $\pi$ can also play the role of sensitivity analysis. The expression of the bounds should be monotonic in $\pi$, and researchers can use different choices of $\pi$ to explore the robustness of obtained conclusions by checking the specific cutoff under which the consistency of conclusion can be maintained. This sensitivity analysis also helps us understand to what extent the conclusion depends on the choice of bounds, and researchers can report the range of anticipation probability that rejects a particular null hypothesis.
The previous section illustrates that by using a DID approach, the treatment effect for treated with anticipation is partially identified under certain assumptions. The population average expressions of the interval bounds lead to straightforward estimators using sample means under independent assumptions. This section builds uniformly effective confidence sets for the partially identified parameters.
Assume researchers observe data from a distribution $\text{P}\in\mathbf{P}$ with the unobservable parameter, $\mathbb{P}[A_i=1|D_i=1]\in[0,\pi]$. $\mathbf{P}$ refers to the family of distributions that satisfy the sampling, potential outcomes restrictions. For the inferential goal under partial identification, I build a confidence set that is uniformly consistent in level $\alpha$, namely \[\lim_{n\to\infty}\inf_{\text{P}\in\mathbf{P},\mathbb{P}[A_i=1|D_i=1]\in[0,\pi]}\mathbb{P}[\mu_g\in CS_{\alpha}^{\mu}]\ge\alpha,\] where $CS_{\alpha}^{\mu}$ is the $\alpha$ level confidence set for the parameter of interest $\mu_g$.
Here, I provide confidence sets based on imbens2004confidence, stoye2009more, and stoye2020simple. The upper and lower bounds for the identified set are estimated using the same sample and are thus highly correlated. I can build an easy-to-implement confidence set by modifying the method addressed above. For notation simplicity, refer to the upper and lower bounds of parameter $\mu_g$ as $\mu_{g,u}$ and $\mu_{g,l}$. A uniformly effective confidence set can be built if the corresponding estimators $\hat{\mu}_{g,u}$ and $\hat{\mu}_{g,l}$ exist and satisfy the following assumptions
To be consistent with the setup in the empirical application, I analyze the case $\tau_g\le0\le\mu_g$ as an example, and I have $\mu_g\in m_g\left[\frac{1}{1+\pi},1\right]$. Corresponding bound estimators will be
If one would like to choose $\pi=\mathbb{P}[D_i=1]$, a straightforward $\hat{\pi}$ will be $\frac{1}{n}\sum_{i=1}^{n}D_i$. The standard errors can be found in the supplemental appendix.
The assumptions and results mainly follow imbens2004confidence and stoye2009more. When compared with the imbens2004confidence approach, the confidence set I construct is slightly different in that I choose to extend along with the upper and lower bounds by the same length $C_n\frac{\hat{\sigma}}{\sqrt{n}}$ where imbens2004confidence choose the same critical value but the standard errors are different. An intuitive explanation is that although the estimators of the upper and lower bounds are ordered by construction, which is an important assumption mentioned in stoye2009more, the upper and lower bounds can have the reverse order and the interval changes from $[\frac{m_g}{1+\pi},m_g]$ to $[m_g,\frac{m_g}{1+\pi}]$ when the confidence set contains both positive and negative values. Therefore, the corresponding variances for the estimators of upper and lower bounds need to be accommodated to use the larger one for both bounds. This modification works for the construction of confidence sets with perfectly correlated and proportional upper and lower bounds, especially when the confidence set contains 0. The proof is discussed in the appendix.
The change in the expression of confidence sets changes the significance level of rejecting the specific null hypothesis, $H_0: \mu_g=0$, in many cases compared with the situation without anticipation. One interesting case worth mentioning happens when $\mu_g$ and $\tau_g$ have different signs. For the specific null hypothesis, $H_0: \mu_g=0$, I can calculate the values of t-statistics that guarantee the conclusion of whether rejecting it or not unchanged regardless of anticipation.
Corollary (ref) gives empirical researchers a specific cutoff $t^{*}$ for the most common case of testing $H_0: \mu_g=0$. If the absolute value of the t-statistic exceeds $t^{*}$ for the case of different signs, taking anticipation into consideration will not change the conclusion of rejecting the null hypothesis. For example, when $\alpha=0.95$, the corresponding $t^{*}$ is 3.3. Thus, if the absolute value of the t-statistic is larger than 3.3 without anticipation, the zero hypothesis for the treatment effect can still be rejected regardless of the anticipation probability when the treatment and anticipatory effects have different signs. This corollary gives empirical researchers a cutoff where they can claim the effectiveness of their conclusions even with anticipation as long as the t-statistic is large enough.
In this section, I illustrate the results of this paper in the environment established by fitzpatrick2014early, which analyzes the effects of an early retirement incentive program on students' achievement. The authors conducted a DID analysis using exogenous variations from the early retirement incentive (ERI) program targeting teachers in Illinois during the mid-1990s to evaluate the effect of large-scale teacher retirements on student achievement. The Teacher Retirement System in Illinois requires retired members who are at least 55 years old and have 20 years of service experience to collect pension benefits at a 6% discount rate below age 60. If both the employer and employee pay a one-time fee, an Early Retirement Option allows eligible members to collect their full benefits. In 1992-1993 and 1993-1994, an early retirement incentive (ERI) program was offered as an alternative to ERO, which allowed employees to buy five extra years of age and experience as long as they retired immediately. This alternative allowed those at least 50 years old and with 15 years of service credit to increase their retirement benefits.
Notably, the ERI programs may affect students' learning, because such programs might lead to a change in teachers' experience and age structure, which will eventually influence students' grades. In their paper, the authors used a DID approach to analyze how promoting the ERI program affected students' grades. They found no evidence of an adverse effect and even found a positive effect on grades in some circumstances. I analyze the average treatment effect for treated by taking anticipation into consideration. The outcome of interest is students' grades, and the major difference is teachers might now anticipate the program in advance and benefit from it.
The authors collected data from several sources. The Teacher Service Record is an administrative dataset that contains information on employees from Illinois Public Schools. The second set of data provides school-level information on test scores for given subjects and grades. The third source contains demographic information of students in schools. The analysis is restricted to teachers of third, sixth and eighth grades, because standardized testing in Illinois focuses on these grades. One major issue for the data is the ERI take-up is not observed directly. The authors exploited the fact that teachers with 15 or more years of experience were most likely to take up the program, and used it as a proxy for the intensity of treatment by the ERI program.
Consider the restrictions I impose on potential outcomes for the case with anticipation. I require that the students' grades in classes whose teachers are ineligible or choose not to retire early should not be affected, and I also require if teachers are not aware of this program in advance, we should see no change in the grades. Further, I require that once the ERI program is implemented, whether teachers anticipate or not should no longer affect the students' grades. The parallel trends assumption requires that trends in students' grades among schools with fewer treated teachers are precise counterfactuals for trends among schools with more treated teachers without anticipation.
Following the idea that teachers, regardless of eligibility, may have some information from a third party before the implementation of the ERI program so that they may have anticipated something, I choose $\pi=\mathbb{P}[D_i=1]$, and the probability of getting treated is estimated by calculating the proportion of experienced teachers with more than 15 years of service credit, and I get correpsonding $\hat{\pi}$. Recall this bound is used to capture the intensity of potential unobservable information, which is proportional to the intensity of treatment, and I can use the proportion of teachers with more than 15 years of teaching experience to bound the anticipation probability. The magnitude effect assumption requires that the anticipatory effect, $\tau$, which is the result of potential behavior changes of teachers who think they can retire early, has a smaller magnitude than the treatment effect, $\mu$, which is the change in students' grades caused by the ERI program after its implementation. Even under a perfect anticipation setup from econometricians' perspective, the teachers are not confident they would be eligible for the policy and take it up; thus the expectation that the anticipatory effect does not have a larger magnitude than the treatment effect when the policy occurs is reasonable. Further, the expectation that the treatment effect $\mu$ has the same sign as the non-negative DID estimator from fitzpatrick2014early is reasonable. On the other hand, I follow the argument in the same paper that claims teachers near the retirement age and who anticipate the possibility of early retirement may exert less effort than younger teachers. Therefore, one can reasonably argue the anticipatory effect $\tau$ is non-positive.
With all the assumptions discussed above, I can analyze the treatment effect with anticipation starting from the following equation in fitzpatrick2014early that estimates the DID estimator \[Y_{igt}^{s}=\beta_0+\beta_1(Teachers\ge 15)_{ig}\times Post_t+\beta_2Teachers_{ig}\times Post_t+\gamma\bm{X}_{it}+\delta_{ig}+\varphi_{tg}+\varepsilon_{itg}^{s}.\] $Y_{igt}^{s}$ is the test score of grade $g$ for subject $s$ in school $i$ and year $t$. $Teachers\ge15$ is the number of teachers with at least 15 years of experience before 1994 and who are thus eligible for the program. $Teachers$ is the average total numbers of teachers. $Post$ serves as the period, an indicator variable that equals 1 after the school year of 1993. Vector $\bm{X}$ contains demographic information, and $\delta$ and $\varphi$ are corresponding fixed effect terms. Although covariates are included here, the parametric assumption that it enters the outcome linearly implies the treatment effect is homogeneous across different values of controls. The intensity of information related to the choice of $\pi$ has already been captured by the proportion of experienced teachers. In the absence of anticipation, $\beta_1$ from this equation estimates the effect of the ERI program on students' grades. Based on the estimator for $\beta_1$ and $\pi$ that I choose, I can analyze the results with anticipation. I check the results for different grades and subjects and compare the cases for all teachers. Results are shown in Table (ref). Similar results using data from subject-specific teachers are listed in the supplemental appendix.
For the partial identification results, I provide identified sets as well as the 95% confidence sets. From the initial results in fitzpatrick2014early, I notice the estimator is at time negative. Because these negative estimates are all insignificant at the $95\%$ level, I conclude this distortion error is due to the finite sample bias. For these estimators, I obtain the identified sets and confidence sets by changing the sign restriction. I find the confidence sets, after I incorporate anticipation, still cannot reject the null hypothesis $\mu=0$, regardless of the sign I choose. The result changes mainly in two aspects. On the one hand, the results with anticipation suggest the treatment effect can be smaller than the one we get directly from the DID approach, because the DID estimator also captures the pre-treatment negative effect caused by anticipation. The effect can be overestimated up to about 30% because of anticipation. On the other hand, the confidence sets, compared with the DID approach, are slightly shifted leftwards, and this result also reminds people to be more careful when interpreting the non-negative treatment effect. Despite these differences, results incorporating anticipation still support the conclusion that the ERI programs have a non-negative effect on students' grades. These results imply incorporating anticipation can make the result more robust and still support our idea of the non-negative effect of ERI programs on student achievement.
I conduct a robustness check to see the range of choices for $\pi$ that keeps the significance of the estimator at a 95% level, and show the result in Figure (ref). I focus on the effect of the early retirement incentive program on the reading grade in grade 8. I present the identified set as well as the 95% confidence set for a sequence of $\pi$, including 0.1, 0.25, 0.5, $\mathbb{P}[D_i=1]$, 0.75, and 0.9. The shorter interval represents the identified set, and the longer one represents the confidence set. I observe that at an anticipation probability of 0.75, more precisely around 0.7, the confidence set marginally contains point 0, which means this positive treatment effect is quite robust even when taking anticipation into consideration. The null hypothesis will only be rejected when about three fourths of the target teachers anticipate it.
This section discusses modifications of the two-period DID model that only considers the pre-treatment anticipatory behavior in the treated group to incorporate anticipation in broader setups. These generalizations build the anticipation framework on more empirical related assumptions and cover problems researchers encounter in applied work.
First, we focus on the restrctions on pre-treatment anticipatory behavior in the treated group only. In our discussion about Assumption (ref), I noted one can understand the focus only on anticipatory behavior within the treated group as an implicit assumption of “perfect anticipation” which implies units that anticipate will get the anticipated treatment status in the future. However, this assumption might be too strong in some circumstances. For example, people may anticipate the existence of a specific policy, but they are not sure they will get treated. In the early retirement incentive program example, teachers can anticipate the possibility of early retirement, but they are not sure about the amount of service credit they can buy and thus cannot perfectly anticipate their future treatment status. I consider the consequences if a mistake is made when anticipating a future treatment in this section and successfully incorporate the anticipatory behavior in the control group. Furthermore, exploring the robustness of our conclusion by checking the error rate under which consistent conclusions can still be obtained is essential. If researchers aim to test a particular null hypothesis, they can also report the lowest error rate at which the null hypothesis is no longer rejected.
To distinguish between anticipated treatment status and the actual treatment units received and to incorporate imperfect anticipation, I now define the anticipation status variable $A_i$ as a random variable that takes three values $\{-1,0,1\}$. The difference between $A_{i}=0$ and $A_{i}\ne0$ distinguishes between those who anticipate and those who don't. However, among those who anticipate, $A_{i}=1$ indicates this anticipates correctly, whereas $A_{i}=-1$ indicates incorrect anticipation. The potential outcome for unit $i$ in period $t$ still depends on both anticipation and treatment $(a,d)$ and is denoted by the random variable $Y_{it}(a,d)$. The only difference is now $a\in\{-1,0,1\}$ and $A_i$ is no longer a binary treatment. Compared with the benchmark model, some modifications need to be made on the assumptions to incorporate imperfect anticipation.
The random sampling assumption remains unchanged, and the only modification is that anticipatory behavior also happens within the control group, so we have more potential outcomes now and we need another index in the expression of pre-treatment potential outcomes for the control group as well.
The assumption that requires anticipation to be the only channel for the future to affect the present is still needed. However, the fact that now one's pre-treatment behavior is affected by his anticipated status, which is likely to be different from the treatment status he receives in the future, needs to be addressed.
Assumption (ref) mainly describes two groups of units. The first group either does not anticipate or they anticipate they will not get treated in the future so they will behave the same way. People in the second group expect that they will get treated in the future and they will behave the other way. Because anticipation is the only way I allow the future to affect the present, one's anticipated treatment status determines their pre-treatment behavior. A treated person with incorrect anticipation should have the same anticipated treatment status as an untreated person who anticipates correctly, and thus, they should behave the same way as those who do not anticipate. However, those who will get treated and anticipate correctly should behave the same way as those who won't be treated but anticipate incorrectly, because they all think they will be covered. This assumption points out that what drives people's pre-treatment behavior is their beliefs about the treatment status. Trying to distinguish between anticipated treatment status and real treatment received is important in the case in which people make mistakes while anticipating.
One may argue that in the situation of imperfect anticipation, people may no longer clearly anticipate the future treatment and may believe they will get treated with a probability. Different people hold different beliefs about their treatment possibilities and behave differently. This situation can be regarded as a case in which anticipation is a multivalue treatment, and people with different beliefs receive different levels of anticipation treatment. Because all things related to anticipation are unobservable, introducing more levels of different anticipation treatments also requires more assumptions regarding each group. Therefore, I still focus on the case in which people's anticipation about the future concerns whether they will be treated.
The parameter of interest $\mu_g$ is still \[\mu_g=\mathbb{E}[g(Y_{i1}(1))-g(Y_{i1}(0))|D_i=1],\] with similar restrictions on $g(.)$ function. Because people's anticipation status does not affect their post-treatment outcomes, I still use one index to represent the potential outcomes $Y_{it}(d)$. The anticipatory effect is modified as
where I implicitly require that the anticipatory effects for those who anticipate correctly and incorrectly are the same.
The key idea of the parallel trend is to require those who get treated and those who do not behave in the same way without the treatment. Following this idea, I need to choose those who have an anticipated untreated status when compared with the outcome in the first period.
The first part of Assumption (ref) now requires us to choose a $\pi$ as the bound for the probability of anticipation among treated and control groups. This modification on restriction is straightforward because under the “perfect anticipation” situation, units that won't get treated will not react to the anticipation, but now they may react to it because of the incorrect anticipation. If one would like to argue a specific relationship exists between the possibility of anticipation within treated and control groups, this assumption might be relaxed. For now, I assume a common bound $\pi$ for two probabilities is chosen. In the second part, I assume the fraction of units that anticipate incorrectly is known as $\varepsilon$ across treated and control groups. Recall that in the discussion about the bias caused by anticipation, I point out that the bias is driven by those who anticipate and react to it in the first period. Under the setup of imperfect anticipation, the proportion of units that cause the bias is determined by anticipated treatment status and thus related to both the proportion of those who anticipate and the accuracy rate among anticipators.
Assumption (ref) is the same magnitude restriction as before. Based on the assumptions above, we are now able to partially identify the parameter of interest $\mu_g$.
If $\varepsilon=0$, which represents the situation of perfect anticipation, this interval degenerates to the interval we get in the benchmark model for the treatment effect. Compared with the interval I get for the treatment effect above, one significant difference is that now the treatment effect is not bounded by the DID estimator from one side. That difference derives from the fact that both the control group and the treated group deviate in the first period. For a perfect anticipation setup, those in the control group will not react to the anticipation, and only the treated group is moving either upward or downward depending on the signs of anticipatory effect. When imperfect anticipation is allowed and both control and treated groups react to it, the distortion in the first period can be either positive or negative depending on the anticipation possibility in different groups. To accommodate this framework in more empirical settings, I don't impose specific assumptions on the relationship between the possibility of anticipating in different groups. If in specific situations, for example, a case in which those who get treated receive private information, researchers can decide the relationship between these two possibilities, improving the identified set based on further assumptions is possible.
Although I start with the case in which $\varepsilon$ is known, a better explanation for including the error rate of anticipation is to understand this procedure as a sensitivity check. Researchers can choose different possible error rates $\varepsilon$ and analyze the region where their conclusions are robust to the choice of error rate. Further, if a specific null hypothesis is tested, the error rate among which the conclusion holds consistently can also be reported. This robustness check procedure helps people understand to what extent the conclusion is affected by the assumption of perfect anticipation.
One might also be interested in what happens if the anticipatory behavior in the treated group has an effect on the post-treatment behavior and the post-treatment potential outcomes in the treated group are also different between those who anticipate and those who do not. Starting from the benchmark model, now let us assume $Y_{i1}(0,1)$ and $Y_{i1}(1,1)$ are different. Then, we have two effects related to anticipation: $\tau_1=\mathbb{E}[g(Y_{i0}(1,1))-g(Y_{i0}(0,1))|D_i=1]$ and $\tau_2=\mathbb{E}[g(Y_{i1}(1,1))-g(Y_{i1}(0,1))|D_i=1]$. Following logic similar to that above, one can find $\mu_g=m_g+\mathbb{P}[A_i=1|D_i=1](\tau_1-\tau_2)$, which implies that if no more assumptions about the relationship between the pre- and post-treatment effect caused by anticipation are imposed, nothing more can be said. Further, if one would like to assume the effect caused by anticipation remains unchanged before and after the treatment, this equation points out that the existence of anticipation will have no effect on the identification and estimation of the parameter of interest here. This assertion is a generalization of the no anticipation assumption in the canonical DID model where both effects are assumed to be zero.
Further generalizations including the model incorporates anticipation with covariates in multiple periods as well as nonlinear outcomes that involves the change-in-changes model are also provided in the appendix.
This paper proposes a potential-outcomes framework for analyzing treatment effects in the presence of anticipatory behavior. Based on a two-period DID model, the findings of this paper show how the standard estimator can be biased, and provide a weighted average of the treatment effect and the anticipatory effect. I also provide conditions under which I can obtain upper and lower bounds for the treatment effects. The motivation and implications of each assumption are discussed to accommodate empirical research backgrounds. This paper contributes to the empirical research by introducing anticipation in a practical and easy-to-generalize way starting from the classical DID model, which makes it robust to the existence of this kind of forward looking behavior. An easy-to-implement estimation and inference strategy is also provided. Based on this inference procedure, I propose a sensitivity analysis approach that discusses the validity of conclusions under different restrictions on the anticipation possibility, and this approach suggests a specific range of t-statistics that guarantee the effectiveness of the conclusion without anticipation when the treatment and anticipatory effect have different signs. I illustrate the results in this paper by examining the effect of early retirement incentive programs on student achievement while considering anticipation, and show potential pitfalls if anticipation is ignored. To make this framework more general and less restrictive, I provide several modifications based on the two-period DID model to be consistent with common empirical setups.
The analysis for this paper still leaves open questions, some of which are discussed in the appendix. I provide several alternative combinations of assumptions that can be used to obtain partial identification results for treatment effects with anticipation, for example, bounded outcomes assumptions and further conditional independence restrictions on potential outcomes and treatments. The choice of these assumptions depends on the empirical backgrounds researchers are working on. Alternative assumptions combined with available bounds provide applied workers with more choices that fit into broad applied circumstances. Further work can focus on some frequent issues in empirical studies, for example, anticipation effects related to instrumental variables. When a time gap exists between the instrumental variable and the treatment, the instrumental variable can cause people to anticipate future treatment and thus react to it before the treatment occurs. This setup is also related to cases with imperfect compliance and situations where anticipation will affect people's future selections into treatments. Another possible extension is to incorporate the anticipation phenomenon in the synthetic control framework. This extension makes sense because the treatments in typical synthetic control applications are often big policy changes that would naturally be anticipated. Following ferman2019synthetic, the anticipation treatment can be regarded as an unobservable confounder that is correlated with treatment, because only those who get treated in the future will react to anticipation. In that case, the pre-treatment weight that fits well may not construct good counterfactual post-treatment outcomes for the treated unit, and thus causes problem. Analyzing the behavior of the synthetic control estimator and DID estimator with anticipation and comparing the performance of these two approaches will be of interest for applied work. By considering these situations, I am more likely to incorporate anticipation in more diverse empirically relevant situations and introduce it into more applied models. \newline
Acknowledgements. I am deeply grateful to Matias Cattaneo for continued advice and encouragement. I thank Yuehao Bai, Ming Fang, Max Farrell, Yingjie Feng, Zheng Gong, Florian Gunsillius, Andreas Hagemann, Xuming He, Michael Jansson, Shaowei Ke, Ziteng Lei, Xinwei Ma, Kenichi Nagasawa, Mel Stephens, Gonzalo Vazquez-Bare, Yian Yin and seminar participants at many institutions for their valuable feedbacks. All errors are my own.