Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,689 characters · 16 sections · 51 citation commands
Correcting Attrition Bias using Changes-in-Changes$^*$
\newtheorem{theorem}{Theorem}
\newtheorem{assumption}{Assumption}
\newtheorem{proposition}{Proposition} \newtheorem{corollary}{Corollary} \newtheorem{definition}{Definition} \newtheorem{lemma}{Lemma} \newtheorem{example}{Example} \newtheorem{remark}{Remark} \newenvironment{proof}[1][Proof]{
} \setcounter{page}{0} \thispagestyle{empty} \makeatletter \makeatother
\thispagestyle{empty}
Attrition is a common and potentially important source of selection bias in a range of treatment effect studies. Attrition has long been recognized as a concern in settings that rely on panel data.\footnote{See, for example, FGM1998, vdBL1998 and ZK1998.} In addition, as randomized experiments become a widely implemented methodology in applied economics, attrition tests and corrections are increasingly relevant to empirical practice MM2017,GHO2022. The current empirical literature relies on a wide range of approaches to correct for attrition bias. None of them, however, are specifically tailored to take advantage of the panel data available in many randomized experiments with baseline (pre-treatment) outcome data as well as in quasi-experimental difference-in-difference designs.
We propose a novel attrition correction based on the changes-in-changes (CiC) approach AI2006. The correction is suitable for treatment effect settings where baseline outcome data are available. Extending the CiC framework to correct for attrition bias requires that the outcome is monotonic in a scalar unobservable and that the distribution of this unobservable conditional on treatment and response status is stable over time. While these assumptions are restrictive, they still allow the distribution of the outcome to vary across time, since that outcome can be a time-varying function of the unobservable.
The proposed method relies on a key insight: Under the extended CiC conditions, there are two transformations that relate the baseline outcome distribution to the distributions of the treated and untreated potential outcomes in the post-treatment period (Lemma (ref)). These two transformations are identical for all treatment-response subpopulations and can be identified using the treatment and control respondents. Using these transformations, we can identify not only the counterfactual distribution for the treatment and control respondents, but also the distribution of both potential outcomes for the treatment and control attritors. The identification of the average treatment effects for the respondents (ATE-R) as well as the entire study population (ATE) then follows immediately. For outcomes that are continuous and strictly monotonic in the unobservable, the parameters of interest are point-identified. For discrete outcomes, bounds can still be obtained for these objects under weak monotonicity restrictions (see Section (ref) of the online appendix).
Since the CiC assumptions do not require random assignment, our approach is not only suitable for randomized experiments with baseline outcome data, but also quasi-experimental difference-in-difference designs. An advantage of random assignment, however, is that the CiC assumptions have an intuitive testable implication (without additional pre-treatment periods), which consists of the equality of the “CiC-extrapolated” average treatment effect on the treated and the untreated.\footnote{Without random assignment, the CiC assumptions can be tested in the presence of additional pre-treatment periods.}
We formally compare the assumptions required for CiC with those required for the inverse probability weighting approach (IPW) as well as widely used bounding approaches such as L2007 and BCGB2015.\footnote{It is important to point out that, unlike CiC, all of these approaches require (conditional) random assignment of the treatment.} In particular, the CiC assumptions can accommodate response models that depend on treatment without requiring monotonicity restrictions. By contrast, we show that the IPW assumptions do not allow response to depend on treatment status, whereas several bounding approaches such as L2007 and BCGB2015 require response to be monotonic in treatment status. The CiC identification approach instead exploits structural restrictions on the outcome model, whereas IPW and the aforementioned bounding approaches do not. We further note the particular challenges of applying approaches aside from the CiC when the object of interest is the ATE. L2007 and other related approaches do not recover this object, while the IPW approach requires a conditional missing-at-random assumption. We then consider the practical implications of these assumptions for a range of reasons for attrition in order to demonstrate how practitioners can assess the suitability of the CiC assumptions in empirical settings.
Finally, we illustrate the CiC attrition correction and compare it to existing approaches in an empirical example revisiting the randomized evaluation of the Progresa cash transfer program. We find the CiC-corrected estimate of the ATE is significantly different from both the uncorrected treatment effect estimate as well as the IPW-corrected estimates. We use several methods to consider the plausibility of the assumptions underlying these two approaches. First, a test for attrition bias applied to this outcome finds that internal validity for the population is violated.\footnote{We implement the test of attrition bias proposed in GHO2022.} Second, we do not find evidence against the CiC identifying assumptions relying on their testable implication under random assignment (Remark (ref)). Third, we conduct an analysis of correlates of attrition to consider plausible drivers of nonresponse and its implications for the corrections in this empirical example.
This paper contributes to the literature on attrition corrections in treatment effect models which build on seminal work on sample selection H1976,H1979. It provides a tractable, nonparametric approach that exploits the presence of baseline outcome data through restrictions on the outcome model. The standard Heckman correction (Heckit) approach assumes a parametric model where the treatment effect is homogeneous across individuals and the joint distribution of the errors in the outcome and response models is normal. MM2017 survey the field experiment literature and find that the most widely-used methods include IPW and various bounding approaches.\footnote{MM2017 also propose a modified version of the inverse probability weighting approach that uses additional data from an intense tracking phase.} The former approach does not restrict the outcome model, but requires unconfoundedness and rules out the possibility that response depends on treatment status to identify the ATE-R. It further requires a conditional missing-at-random assumption to identify the ATE. There are several bounding approaches in the literature that do not impose any restrictions on the potential outcomes. Manski1989 and HM1995,HM2000 provide bounds on treatment effect parameters that require minimal assumptions, however the bounds are typically wide in practice. L2007 proposes bounds on the average treatment effect for the always-responders, a subset of the respondent subpopulation, assuming monotonicity restrictions on potential response. While our method imposes a monotonicity condition on the outcome variable, it does not impose monotonicity of response. BCGB2015 show that tighter bounds on the average treatment effect for a subpopulation of respondents can be obtained by exploiting additional data, specifically the number of calls required to obtain a response. Their proposed bounds specifically exploit the monotonicity of the response in the number of maximal attempts to reach an individual to tighten L2007's (L2007) bounds.\footnote{To do so, the authors assume that the maximum number of attempts to reach an individual is randomly assigned and excluded from the outcome. Relatedly, in the context of survey design, DiNardoetal2021 propose to randomly assign the probability of observation across survey participants.}
This paper proceeds as follows. Section (ref) introduces the model, discusses the implications of the time-invariance assumption for various response models, and provides the identification results with and without random assignment. Section (ref) conducts a formal comparison of the CiC correction with IPW and bounding corrections. Section (ref) applies the CiC and IPW corrections to the Progresa randomized experiment. Section (ref) concludes.
Let $Y_t$ and $D_t$ denote the observed outcome and treatment status in period $t$. The treatment path is denoted by $D=(D_0,D_1)$. To simplify notation, we denote the group membership by $G$, where $G=1$ for the treatment group which receives the treatment path $D=(0,1)$, and $G=0$ for the control group which receives the treatment path $D=(0,0)$, noting that $G=D_1$. We consider the following setting:
$Y_{0}(0)$ denotes the untreated potential outcome in the baseline period ($t=0$) and $Y_1(d)$ denotes the potential outcome for $d=0,1$ for the follow-up period ($t=1$). The outcome variable in the baseline period ($t=0$) is always observed, whereas the outcome variable in the follow-up period ($t=1$) is observed only if $R=1$. $R$ specifically denotes response in the follow-up period with $R(d)$ denoting the potential response given treatment status $d=0,1$. We assume there are no response issues in the baseline period $(t=0)$.
In this paper, we are interested in identifying the average treatment effects for the treated and untreated respondents (ATT-R and ATU-R), for the respondents (ATE-R), and for the study population\footnote{The study population is the population that the sample is drawn from.} (ATE), defined as follows:
We obtain a random sample of the vector $(G, R, Y_{0}, Y_{1}^*)$, where all random variables are observed except $Y_1^*$ which is only observed when $R=1$.\footnote{In our setting, we observe $Y_0$ for all units. As a result, $R$ equals to one for units with observations of both $Y_0$ and $Y_1^*$, whereas it equals zero for units with observed $Y_0$ but missing $Y_1^*$. Thus, $R$ is a response-group indicator, whereas $G$ is a treatment-group indicator.}
In order to identify the above parameters, we assume that the potential outcomes are given by the following model,
The variable $U_{t}$ denotes the unobserved heterogeneity in the outcome, which is assumed to be scalar. In the following, we present our identifying assumptions which impose specific restrictions on $\mu_t(\cdot)$ and $U_t$. Let $F_Y$ denote the cumulative distribution function of a random variable $Y$.
Assumption (ref).(ref) states that the distribution of unobservables that affect the outcome ($U_t$) is stable over time within each treatment-response subgroup. It is similar to Assumption 3.3 in AI2006, except that we condition on the observed response status. This assumption rules out time variability in the distribution of unobservables within each treatment-response subpopulations, but it admits selection into the program and survey response as it allows for differences in the distribution of $U_t$ and $Y_t(d)$ across the four subgroups. For a more detailed discussion of the plausibility of this assumption considering both the unobservable determinants of the outcome and response, see Section (ref). We further impose Assumption (ref).(ref) to ensure that the distribution function of $Y_t|G=g,R=1$ is invertible for $g=0,1$.
Assumption (ref).(ref) implies that for each period, the untreated potential outcome is strictly increasing in the unobserved heterogeneity. It is the same as Assumption 3.2 in AI2006. Assumption (ref).(ref) requires that the treated potential outcome is strictly increasing in the unobserved heterogeneity.\footnote{For instance, if $U_t$ is a single characteristic such as ability and the outcome is profits, Assumption (ref) implies that higher levels of ability correspond to higher potential profits.} These two monotonicity assumptions are the main driver of our identification results as we show in Lemma (ref). Assumption (ref) is automatically satisfied when the structural function is additively separable, such as $Y_t(d)=\gamma_t d+\lambda_t(1-d)+U_t$, where $U_t=\alpha+\varepsilon_t$ includes a time-invariant and time-varying component. It holds, however, for the broader class of potentially nonlinear, monotonic transformations, $Y_t(d)=\mu_t(d,U_t)$. We provide further examples in Section (ref).
The strict monotonicity conditions imposed in Assumption (ref) have consequences for the interpretation of $U_t$ and the time-invariance condition in Assumption (ref). These conditions specifically imply that $U_t$ may be viewed as the normalized outcome, specifically $U_t=\mu_t^{-1}(0;Y_t(0))$, where $\mu_t^{-1}(0;y)$ denotes the inverse of $\mu_t(0,u)$. In light of Assumption (ref), Assumption (ref) requires that the change in the potential outcome distribution across time is only driven by the change in the monotonic structural function, $\mu_t(d,\cdot)$. This allows for the outcome distribution to change across time as long as it admits a normalization that renders its distribution stable across time.\footnote{To illustrate this point, consider a setting where the potential outcome is given by the location-scale model: $\mu_t(d,U_t)=\sigma_t^dU_t+\alpha_t^d$. Then, the conditional time-invariance assumption requires that the change in the potential outcome distribution across time is solely due to the change in $\mu_t(d,\cdot)$ as follows, $F_{Y_t(d)|G,R}(y)=P(\sigma_t^dU_t+\alpha_t^d\leq y|G,R)=P\left(U_t\leq \frac{y-\alpha_t^d}{\sigma_t^d}\mid G,R\right)=F_{U_0|G,R}\left(\frac{y-\alpha_t^d}{\sigma_t^d}\right)$.}
Since we will provide attrition corrections in randomized experiments, we formally define the assumption of random assignment.
Assumption (ref) states that the individuals are randomly assigned to the treatment $(G=1)$ and control $(G=0)$ groups. This assumption applies to randomized experiments with simple and cluster randomization designs. Below we provide identification results with and without random assignment.
In Section (ref), we rely on the strict monotonicity of the structural function (Assumption (ref)) together with Assumption (ref) to provide point-identification results for continuously distributed random variables with strictly increasing distributions. Before we do so, we examine the conditions that are necessary and sufficient for a given response model to satisfy the conditional time-invariance assumption.
In this section, we provide necessary and sufficient conditions for the conditional time-invariance assumption (Assumption (ref).(ref)) required for our identification result. These conditions are useful in applications where researchers may have a priori information on the sources of non-response in their setting. To illustrate the conditional time-invariance assumption further, we discuss its plausibility in the context of examples of unobservable determinants of the outcome and response.
We first consider some examples of unobservable determinants of the outcome. In some settings, the conditional time-invariance assumption is natural. For instance, consider a setting in which the treatment is a microcredit program, the outcome of interest is profits, and $U_t$ is an unobserved determinant of profits. If the relevant unobserved heterogeneity is a trait that is not typically viewed as changing over time ($U_0=U_1$), such as ability or risk preferences, then the conditional time-invariance assumption is trivially satisfied.
Alternatively, if profits are determined by a time-varying unobservable ($U_0\neq U_1$), the conditional time-invariance assumption can still be satisfied. To determine this, however, it is crucial to relate the interpretation of $U_t$ to the structure imposed on the potential outcomes. To fix ideas, let $\tilde{U}_t^d=U_t\sigma_t^d+\alpha_t^d$ denote a health shock, which may have a different mean and variance across time as well as by treatment status. For instance, if the follow-up period coincides with a season during which malaria is endemic, then the health shock can have a negative mean and lower standard deviation to signify the higher likelihood of receiving a bad health shock, specifically malaria. That is, the mean and standard deviation of the health shock can vary by treatment status as well such that receiving the treatment can change the conditional distribution of health shocks faced by the individuals. More generally, if profits, given by $Y_t(d)=\tilde{\mu}_t(d,\tilde{U}_t^d)$, are a strictly monotonic transformation of $\tilde{U}_t^d$, then they are also a strictly monotonic transformation of $U_t$, the normalized health shock.\footnote{To see this, recall that $U_t=(\tilde{U}_t^d-\alpha_t^d)/\sigma_t^d$ and $\tilde{U}_t^d=\tilde{\mu}_t^{-1}(d;Y_t(d))$, so $\mu_t(d,U_t)=\tilde{\mu}_t(d,\tilde{U}_t^d)$ is a composition of two monotonic transformations. To normalize the outcome and obtain $U_t$, we invert the composition of the two functions, $U_t=\left(\tilde{\mu}_t^{-1}(d;Y_t(d))-\alpha_t^d\right)/\sigma_t^d$.} Thus, while the normalized health shock has to obey the conditional time-invariance assumption (Assumption (ref).(ref)), the conditional distribution of $\tilde{U}_t^d$ is allowed to change over time and by treatment status.
To further analyze the plausibility of this conditional time-invariance assumption, it is helpful to understand how it relates to the response model. Thus, we consider unobservable determinants of response and characterize the restrictions that are necessary and sufficient for the conditional time-invariance assumption.
All proofs are in the appendix. This proposition provides a condition on the unobservables determining response that holds iff the distribution of $U_t|G,R$ is time-invariant. The proposition specifically states that if response is determined by a vector of unobservables $V$ as well as treatment, then Assumption (ref).(ref) would hold iff the joint distribution of $(U_t,V)$ is time invariant conditional on $G$.\footnote{It is important to note that this condition allows for $U_t$ and $V$ to be dependent conditional on $G$. To see this, note that $F_{U_t,V|G}(u,v)=C_{U_t,V|G}(F_{U_t|G}(u),F_{V|G}(u))$ (Sklar theorem), where $C_{U_t,V \vert G}$ denotes the copula between $U_t$ and $V$ conditional on G. As a result, while the time invariance of the joint distribution requires the time invariance of the copula $C_{U_t,V|G}$ and the distribution of $U_t|G$, it does not restrict the type of copula that governs the dependence between $U_t$ and $V$.} Note that under random assignment, which is given by $(V,U_0,U_1)\perp G$ in our context, the condition in Proposition (ref) would be replaced by its unconditional version, $(U_0,V)\overset{d}{=}(U_1,V)$.
Next, we consider an example of a response model and discuss the condition in Proposition (ref) in the context of this example.
The above example demonstrates that the conditions in Proposition (ref) do not impose restrictions on the functional form of $R$ and are consistent with a multidimensional $V$. This feature is especially attractive in settings where response is determined by multiple factors.
Proposition (ref) provides a general necessary and sufficient condition that allows for dependence between $V$, $U_0$ and $U_1$ and obeys time-invariance restrictions as illustrated in the above example. The proposition is not explicit, however, on what precise conditions would imply the time invariance of $(V,U_t)|G$ if $V$ is a function of $U_0$ and $U_1$. The following corollary addresses this issue by examining a special case of Proposition (ref) where response is determined by a function of $U_0$ and $U_1$.\footnote{It is worth noting that if response is solely determined by baseline outcome, $R=f(Y_0)$, random assignment ($(Y_0(0),Y_1(0),Y_1(1),R(0),R(1))\perp G$) implies $(Y_1(0),Y_1(1))\perp G|R$, which would yield a case where no correction would be warranted and the ATE-R would be identified from the simple difference in means between treatment and control respondents. However, while this is a theoretically interesting case, it is not very relevant from a practical perspective since response at follow-up is also likely affected by unobservable factors in the follow-up period and the treatment status itself.} We emphasize that unlike Proposition (ref), the condition in the following corollary is merely sufficient, and not necessary, for Assumption (ref).(ref) to hold.
The above corollary establishes that if response is determined by $G$, $U_0$, and $U_1$, then for the time-invariance condition to hold conditional on $G$ and $R$, it is sufficient for response to be symmetric in $U_0$ and $U_1$ and the distribution of $(U_0,U_1)$ conditional on $G$ to be exchangeable in $U_0$ and $U_1$.\footnote{It is worth noting that the exchangeability condition implies the time invariance of the distribution of $U_t$ conditional on $G$, specifically $U_0|G\overset{d}{=}U_1|G$.} The following example provides an example of a response model that obeys these conditions.
Finally, it is important to consider an example where response depends on the unobservable determinant of the follow-up outcome, $U_1$. This example neither obeys the conditions in Corollary (ref) nor Assumption (ref).(ref).
In this section, we outline how the CiC identification approach can be applied to point-identify our objects of interest. We provide results both for the respondent subpopulation and study population. Let $\mathbb Y$ denote the support of the random variable $Y$, and $\mathbb{Y}_{g,r}^{d,t}$ denote the support of $Y_{t}(d)|G=g,R=r$. Define $F_Y^{-1}(q)=\inf\{y\in \mathbb Y \vert F_Y(y) \geq q\}$.
Before we proceed to our main identification results, the following lemma helps us understand how Assumptions (ref) and (ref) can allow us to “extrapolate” not only to the respondent subpopulations but also to the attritor subpopulations.
Lemma (ref).(ref)(i) shows that under the time-invariance assumption (Assumption (ref).(ref)) and the strict monotonicity of the untreated potential outcome (Assumption (ref).(ref)), the distribution of the untreated potential outcome for any treatment-response subpopulation in the follow-up period at a given $y$ equals the distribution of the untreated potential outcome of that subpopulation in the baseline period evaluated at $T_0(y)$, where the transformation, $T_0(\cdot)$, is the same for all treatment-response subpopulations. Since we observe the distribution of the untreated potential outcome of the control respondents in both baseline and follow-up periods, Lemma (ref).(ref)(ii) shows that we can identify $T_0(y)$ for $y\in\mathbb{Y}_{0,1}^{0,1}$ using the control respondents by the continuity and strict monotonicity of the outcome distribution (Assumption (ref).(ref)).
If we also impose the strict monotonicity assumption on the treated potential outcome (Assumption (ref).(ref)), Lemma (ref).(ref)(i) shows that the treated potential outcome distribution for any treatment-response subpopulation at a given value $y$ equals the distribution of the untreated potential outcome of that subpopulation in the baseline period evaluated at $T_1(y)$, where the transformation, $T_1(\cdot)$, is the same for all treatment-response subpopulations. Since we observe the untreated potential outcome in the baseline period and the treated potential outcome in the follow-up period for the treatment respondents, we can use them to identify $T_1(y)$ for $y\in\mathbb{Y}_{1,1}^{1,1}$ (Lemma (ref).(ref)(ii)).
In sum, Lemma (ref) shows that we can use the control and treatment respondents to identify $T_0(y)$ and $T_1(y)$ on their respective support. Since $T_0(y)$ and $T_1(y)$ are the same for all subpopulations, the identification of the distribution of an unobserved potential outcome for a given treatment-response subpopulation follows immediately assuming that we can observe the baseline outcome distribution for this subpopulation and that additional support conditions hold. We finally note that Lemma (ref) does not require random assignment. This allows us to provide identification results for our parameters of interest without random assignment (Assumption (ref)).
In the following, we provide identification results for the average treatment effects for the respondent subpopulations. Note that since individuals choose to respond or not, treatment is no longer randomly assigned conditional on response without further restrictions. As a result, the results in this section do not require random assignment. They instead exploit the conditional time-invariance assumption as well as the structural assumptions imposed by the CiC conditions. The identification results in this section constitute a direct application of CiC to the respondent subpopulation.
We first establish the identification of the ATT-R, since it requires the strict monotonicity condition on the untreated potential outcome only (Assumption (ref).(ref)) in addition to the assumptions on the unobservables. Let $\mathbb{U}_{g,r}$ denote the support of $U_{0}|G=g, R=r$ for $g=0,1$, and $r=0,1$. For two sets $\mathbb{A}$ and $\mathbb{B}$, $\mathbb{A}\subseteq \mathbb{B}$ denotes that $\mathbb{A}$ is contained in $\mathbb{B}$.
This proposition establishes that the counterfactual distribution of the treatment respondents is identified by evaluating the distribution of the untreated potential outcome of that subpopulation at baseline at the transformation $T_0(y)$ identified from Lemma (ref). The ATT-R is then identified from the counterfactual distribution.
Next, we provide the identification result for the ATE-R, which requires the strict monotonicity of both treated and untreated potential outcomes in $U_t$.
The proof of the above proposition follows from Lemma (ref). Since the ATE-R is a probability-weighted average of the ATT-R and ATU-R, the identification result in Proposition (ref) builds on the identification of the ATT-R in Proposition (ref). It then establishes the identification of the ATU-R, which requires identifying the treated potential outcome distribution for the control respondents. That distribution is obtained by evaluating the baseline distribution of control respondents at the transformation $T_1(y)$ identified from Lemma (ref).
In this section, we present identification results for the study population. Since the random assignment of treatment simplifies the identification of the ATE, we provide identification results with and without that assumption. Under random assignment, the identification of the ATE only requires identifying the treated (untreated) potential outcome distributions for treatment (control) attritors. Thus, researchers analyzing data from a randomized controlled trial can implement the correction indicated by Proposition (ref). In contrast, researchers using other research designs should implement the correction indicated in Proposition (ref) as the identification of the ATE relies on separately identifying the counterfactuals for the ATT and ATU for respondents and attritors.
We first examine the attrition correction without assuming random assignment. The law of iterated expectations allows us to write $E[Y_1(d)]$ as follows: {{
}}
For $d=0$, the only terms that are observable on the right-hand side are the probabilities as well as the expected potential outcome without the treatment for the control respondents, $E[Y_1(0)|G=0,R=1]$. Therefore, in order to identify $E[Y_1(0)]$, it remains to identify the distributions of the untreated potential outcome for all remaining subpopulations, $F_{Y_1(0)|G=1,R=0}$, $F_{Y_1(0)|G=1,R=1}$, and $F_{Y_1(0)|G=0,R=0}$. Similarly, for $d=1$, the only terms that are observable on the right-hand side are the probabilities as well as the expected potential outcome with the treatment for the treatment respondents, $E[Y_1(1)|G=1,R=1]$. As a result, in order to identify $E[Y_1(1)]$, it remains to identify the distribution of the treated potential outcome for all remaining subpopulations, $F_{Y_1(1)|G=1,R=0}$, $F_{Y_1(1)|G=0,R=1}$, and $F_{Y_1(1)|G=0,R=0}$.
The next proposition provides sufficient conditions such that we can apply Lemma (ref).(ref) to identify $F_{Y_1(0)|G=1,R=0}$, $F_{Y_1(0)|G=1,R=1}$, and $F_{Y_1(0)|G=0,R=0}$ as well as Lemma (ref).(ref) to identify $F_{Y_1(1)|G=1,R=0}$, $F_{Y_1(1)|G=0,R=1}$, and $F_{Y_1(1)|G=0,R=0}$. The identification of the ATE follows.
Proposition (ref) has two main practical implications. First, it demonstrates that the CiC approach can identify the ATE in settings without (simple) random assignment, such as quasi-experimental difference-in-difference designs. We specifically have to obtain the average treatment effect for each treatment-response subgroup. For the treatment (control) respondents, we obtain the ATT-R (ATU-R) by applying the CiC approach to identify their average outcome without (with) the treatment. Furthermore, since we do not observe either potential outcome for the attritors, we have to apply the CiC approach to identify the average potential outcome with and without the treatment. The ATE is then obtained as a probability-weighted average of the group-specific average treatment effects.\\
Next, we examine the identification of the ATE under random assignment. Under this assumption, we have $ATE=E[Y_1(1)|G=1]-E[Y_1(0)|G=0]$. Using the law of iterated expectations, we have {{
}} The only unobservable objects on the right-hand side of the above equations are the average outcomes of the control and treatment attritors, $E[Y_{1}(0)|G=0,R=0]$ and $E[Y_{1}(1)|G=1,R=0]$. The following proposition provides sufficient conditions such that we can apply Lemma (ref).(ref) and (ref).(ref) to identify $F_{Y_{1}(0)|G=0,R=0}$ and $F_{Y_{1}(1)|G=1,R=0}$, respectively, and thereby their expectations.
This proposition recovers the ATE in the case of random assignment by identifying the outcome distribution at follow-up of control attritors and respondents using $T_0(y)$ and $T_1(y)$ from Lemma (ref), respectively. As a result, random assignment simplifies the identification of the ATE. The following remark demonstrates that it also provides a testable implication of the CiC assumptions.
In this section, we describe the most widely-used corrections in practice, which are IPW and L2007 bounds, and compare their assumptions to the CiC assumptions.\footnote{See MM2017 for a review of approaches to correcting for attrition bias in the field experiment literature.} While the CiC correction exploits restrictions on the outcome model, the existing approaches exploit restrictions on how response depends on treatment.
Unlike the CiC approach, the IPW corrections rely on the assumption of selection on observables (i.e, unconfoundedness). In particular, to identify the average treatment effect on the respondents, it is required that treatment assignment is independent of potential outcomes and potential response, once we condition on baseline covariates $X_0$ ($G \perp (Y_1(0),Y_1(1),R(0),R(1)) \vert X_0$).\footnote{For the special case where $X_0=Y_0$, $G\perp (Y_1(0),Y_1(1),R(0),R(1))|Y_0$.} A second assumption required for the identification of the ATE-R is that potential response does not depend on treatment status, $R(0)=R(1)$.\footnote{While unconfoundedness by itself is not testable, combining it with $R(0)=R(1)$ implies $G\perp (Y_1(0),Y_1(1),R)\vert X_0$, which further implies the testable restriction, $R\perp G\vert X_0$. This is therefore a testable restriction of the IPW identifying assumptions of the ATE-R. We emphasize, however, that the additional restriction required for the identification of the ATE relying on IPW, specifically $(R(0),R(1))\perp (Y_1(0),Y_1(1))|X_0$, is not testable.} This condition rules out individuals who would only respond if assigned to the treatment or control group, the so-called treatment-only or control-only responders, and implies that the respondents ($R=1$) solely consist of always-responders. In addition to unconfoundedness, the identification of the ATE requires that potential response is independent of potential outcome conditional on $X_0$, $(R(0),R(1))\perp (Y_1(0),Y_1(1))|X_0$. See Section (ref) in the online appendix for the definitions of the ATE-R and ATE using these corrections.
There are two main differences between the IPW and CiC assumptions for the identification of the ATE-R. First, while the CiC approach exploits monotonicity of the outcome model, IPW restricts the response model and rules out the possibility of differential attrition rates across treatment and control groups, prevalent in the empirical literature.\footnote{In a detailed review of published field experiments GHO2022 find that 37% (11%) of field experiments have differential attrition rates higher than 2% (5%). This proportion is likely a lower bound on the proportion among all field experiments, given the publication bias towards field experiments with smaller differential attrition rates. It is important to note, however, that if treatment-only and control-only responders exist in the population, then differential attrition rates estimate the difference in proportion between these two responder subpopulations and, therefore, should not be used as an indication of whether response depends on treatment status or not. } In addition, the IPW approach requires treatment status to be conditionally randomly assigned, while the CiC correction allows for selection into response and treatment status under the condition that the distribution of unobservables $U_t$ within each treatment-response subgroup does not change between baseline and follow-up. Furthermore, in order to identify the ATE, IPW requires a conditional independence assumption between potential responses and outcomes, which states that missingness (i.e. attrition) is random conditional on $X_0$.
In sum, while the IPW and CiC approaches are non-nested in general, the main advantages of the CiC approach are twofold. First, it allows response to depend on treatment, a likely concern in practice. In addition, it does not require missingness to be conditionally at random to identify the ATE. Thus, there are a number of settings in which CiC can be applied where it would not be appropriate to apply IPW. In settings where the assumptions of both approaches are suitable, however, we note that CiC requires the availability of non-degenerate baseline outcome data while IPW can be implemented when only baseline covariates are available. Furthermore, we note that IPW delivers point-identification of the ATE-R and ATE regardless of the outcome distribution. In contrast, the CiC approach provides point-identification for continuous outcomes, and provides bounds on the ATT-R, ATE-R and ATE for discrete outcomes.
When comparing CiC and IPW, it is helpful to consider specific response models such as those described in Section (ref). Example (ref) provides a simple response function in the context of random assignment that can illuminate the implications of CiC conditions: $R=1\{V\geq c_0+c_1G\}$. Here, we consider settings in which this type of response function may apply. Let $V$ be the opportunity cost of time, since that is likely an important reason that participants are reluctant to respond in many settings. We return to our example in which the treatment is a microcredit program, the goal of which is to increase profits. As in the discussion of the high-level assumptions above, $U_t$ is an unobservable that is likely to affect a participant's work in the business and thus profits. Since we are now also imposing structure on the response function, we note that the (normalized) health shock ($U_t$) would also affect the opportunity cost of time and thus the likelihood of responding to a survey. That is, $V$ has a dependence relationship with $U_t$. In this case, the CiC assumptions allow the realization of the health shock and the opportunity cost of time to be related, as long as the relationship is stable over time. In contrast, the IPW assumptions for the ATE-R require that the opportunity cost of time at follow-up only depends on the health status at baseline ($U_{0}$) rather than the health status at the time of the actual follow-up survey ($U_1$).
Alternatively, now let us consider two cases of the same example in which profits are instead largely determined by an unobservable, such as ability, which is constant over time ($U_0=U_1$). First, consider the case in which $c_1>0$. This may describe the relevant response function if the microcredit program allows beneficiaries to expand their businesses, and thus treated individuals are busier and less likely to respond. Of course, in these cases, it would not be appropriate to apply IPW, since the IPW assumptions do not allow response to vary with the treatment status. This case would meet the conditions for CiC, however, as the conditional time-invariance assumption holds trivially if $U_0=U_1$.
Next, let us consider a case in which migration determines response, and $V$ is a one-time shock at endline caused by a conflict. The conflict shock could differentially affect participants based on ability, if, for example, higher ability individuals are better able to adapt and are less likely to leave. In that case, $U$ and $V$ would be related. But, if the response function is the same for the treatment and control ($c_1=0$) and ability is constant over time, both the CiC and IPW assumptions (for the ATE-R and ATE) would hold.\footnote{We emphasize that this conclusion relies on random assignment. Furthermore, note that if we condition on $Y_0$ in IPW, then the conditional missing-at-random assumption required for the IPW correction for the ATE trivially holds when $U_0=U_1$ under Assumption (ref). To see this, note that under this assumption conditioning on $Y_0$ is equivalent to conditioning on $U$, which is assumed to be time-invariant in this example. As a result, $(R(0),R(1))\perp (Y_1(0),Y_1(1))|Y_0$ holds trivially, since the potential outcomes are fixed in this case once we condition on $U$.}
Now suppose that the unobservables that affect the outcome are the same unobservables that affect response ($U=V$). Corollary (ref) establishes a sufficient condition under which the time-invariance assumption and this restriction holds. For example, consider a case in which the treatment is a matching grant program for charitable donations and the outcome is donations, and thus the preference for reciprocity is the (time-invariant) unobservable that determines the outcome as well as response. In addition, treatment status may affect whether an individual will affect response to a survey ($c_1>0$). Since the preference for reciprocity is constant over time, the CiC assumptions would hold, whereas the IPW assumptions would not because response depends on treatment status.
To further examine the sufficient condition established by this corollary, we consider the response equation given in Example (ref), $R=1\{U_0+U_1\leq c_0+c_1G\}$. This function describes a situation in which response depends on the flow of an unobservable that accumulates in each period. Let $U$, for example, be confidence. Once a participant's confidence reaches a certain threshold, they migrate and are not available to respond to the survey. In addition, returning to our microcredit example, increasing confidence leads to greater business success and higher profits. It would make sense for confidence in the baseline period to matter, if the migrant needs to begin reaching out through their networks or otherwise planning for their migration in the baseline period. If the accumulated confidence across both the baseline and follow-up periods matter, then the CiC condition will hold and the IPW assumption will not. By contrast, if only the confidence accumulated in the baseline period matters, then the CiC assumptions would not hold in general, whereas the IPW assumptions for the ATE-R would only hold if $c_1=0$. Finally, if only the additional confidence accumulated in the follow-up period matters, then both the CiC and IPW assumptions would fail in general.
These examples highlight a range of possible examples of response functions, and how the CiC and IPW assumptions apply.\footnote{Although each of the examples above focuses on a single possible unobservable determining response for illustrative purposes, our approach allows for more complex response models as indicated in Example (ref).} We focus on comparing the CiC and IPW assumptions since they both can recover the ATE-R and ATE. It is important to emphasize that to identify the ATE, IPW requires an additional assumption, specifically that attrition is random conditional on the covariates as we describe in Section (ref).
L2007, aware of the likely possibility that some individuals are induced to respond due to their treatment status, exploits an assumption of monotonicity of response in treatment status, $R(0)\leq R(1)$, while maintaining unconfoundedness, $G \perp (Y_1(0),Y_1(1), R(0), R(1))|X_0$. This monotonicity assumption states that treatment assignment only affects response in one direction, and implies that the control respondents consist solely of always-responders $((R(0),R(1))=(1,1))$. Thus, as a result, it allows for the partial identification of the average treatment effect for the always-responders, a subset of the respondent subpopulation.\footnote{See Section (ref) in the online appendix for the equations with the bounds for continuous and binary outcomes.}
The Lee bounds and the CiC corrections we propose in this paper differ in terms of the identifiable objects and the required conditions. First, L2007 bounds the average treatment effect for always-responders, which is a subpopulation of the respondents, whereas the CiC corrections we propose can identify the average treatment effects for the respondents as well as the entire study population. Second, L2007 requires random assignment of the treatment unconditionally or conditional on some covariates, whereas the CiC corrections can be extended to settings where treatment is not randomly assigned assuming that the distribution of the unobservable determinant of the outcome is stable over time for each treatment-response subgroup. Third, our approach requires monotonicity of the outcome in a scalar unobservable, while the L2007 (L2007) approach requires instead monotonicity of the response in the treatment. This assumption is unlikely to hold in settings where nonresponse is determined by multiple factors, such as reciprocity or cost of time, since they are likely to lead to treatment-only responders and control-only responders, respectively. One key advantage of the Lee bounds, however, is that they do not require baseline data.
While Lee bounds are worst-case scenario bounds in the spirit of HM1995, more recently, BCGB2015 exploit an insight that additional data on reluctance to respond to surveys combined with the same assumption of monotonicity of response can provide bounds that are tighter than L2007's (L2007). To do so, the authors assume that the maximum number of attempts to reach an individual is randomly assigned and excluded from the outcome. In addition to monotonicity of response in treatment, they further require response to be monotone in the survey effort.
We apply our proposed CiC corrections and other common corrections to two outcomes from a large-scale randomized evaluation of the impact of Progresa, a conditional cash transfer program in Mexico. The Progresa evaluation, which was implemented in 1997, randomized 506 villages into a treatment group and a control group. These villages were designed to be representative of a larger group of 6,396 eligible villages in Mexico. Thus, both the average treatment effect for the respondent subpopulation (ATE-R) and for the study population (ATE) are likely to be of interest in this setting. In the 320 treatment villages, families received a conditional cash transfer if they were below the given threshold on a poverty index and engaged in specific education and health-seeking behaviors. In the control villages, no households were offered a transfer. There is a vast literature that has studied a range of outcomes from the Progresa evaluation, with some of the most studied outcomes focusing on education and health Skoufias2001,Schultz2004,AMS2012, PT2017.
The goal of this application is to demonstrate the implementation of the CiC correction on a continuous outcome with baseline data, as well as the comparison of the CiC and IPW corrections for that outcome.\footnote{We only include baseline outcome in both the CiC and IPW corrections to simplify the direct comparison of approaches since the CiC correction requires baseline outcome data. Both corrections allow covariates, however. The use of covariates in the CiC correction is discussed in Remark (ref).} Both CiC and IPW provide point estimates for the ATE-R and the ATE for continuous outcomes, and IPW is the most widely-used correction in the literature.\footnote{Another widely used attrition correction are the bounds proposed by L2007. Since this approach only focuses on the average treatment effect on the always-responders, however, it is not directly comparable to our method that focuses on the ATE-R and the ATE.} In Section (ref) of the online appendix, we also apply both the Manski and Lee bounding approaches to this continuous outcome, and implement all four corrections for a related binary variable.
The continuous outcome we examine is the value of a productive asset, specifically farm animals. Close to 90% of the households in the population targeted by Progresa engage in agricultural activities and make investments in productive assets for their farms.\footnote{This outcome is first proposed in GMR2012. Our findings are not directly relevant to that paper since we only focus on the third follow-up, while GMR2012 focus on the outcome pooled across all three follow-ups. We also restrict our sample to those who appear in the baseline survey.} At baseline, the average value of farm animals in these households was 1,819 Mexican pesos (denoted \$), which is equivalent to approximately 102 USD today.
The potential for attrition bias in treatment effect estimates from the original Progresa follow-ups has often been discussed in the literature Bobonis2011, PRT2007, BT1999. Thus, attrition in this setting is significant and warrants the implementation of corrections. We focus here on the final follow-up that takes place 18 months after the program began, and that has an attrition rate of 12.2%.\footnote{This attrition rate is close to the average attrition rate for field experiments where the unit of observation is the individual or household MM2017. We focus on this final follow-up since examining a single follow-up allows us to more clearly outline reasons for attrition, and the final follow-up is often seen as definitive. Furthermore, since assets are not likely to adjust quickly, it makes sense to focus on impacts in the final follow-up. That said, for researchers who would like to generate pooled estimates of the CiC corrected estimates, they can simply average across the corrected estimates for the individual follow-ups.} For the purposes of this analysis, the attrition rate is conditional on appearing in the baseline survey, which included a total of 12,299 households. Thus, the outcome is observed for more than 10,000 households in this follow-up survey.
We first examine the CiC-corrected estimates for the ATE-R and ATE in relation to the na\"{i}ve (uncorrected) estimate of the treatment effect for productive assets, $\widehat{\Delta}_R$, which is simply the difference in the mean outcome between treatment and control respondents at follow-up. The na\"{i}ve estimate of the treatment effect on the value of production animals is \$351, which is significant at the 1% level and is relative to a control mean of \$1,096 (see Panel A of Table (ref)). The CiC-corrected estimate for the ATE-R is \$284 while the CiC-corrected estimate for the ATE is \$289. The similarity of these two estimates suggests that, if the CiC assumptions hold, there is relatively little treatment heterogeneity. A key consideration, however, is whether these coefficients differ from the na\"{i}ve estimate of the treatment effect (Panel B of Table (ref)). On the one hand, the CiC-corrected ATE-R is not significantly different from the na\"{i}ve treatment effect estimate, since the difference in the two estimates is \$67 with a standard error of \$125. On the other hand, for the ATE, the difference is \$62 with a standard error of \$26, which is significant at the 5% level. A difference in the power of these tests may explain this result.
Meanwhile, the IPW approach does not suggest that a correction is required for either the ATE-R or the ATE. The IPW-corrected estimates, which are \$349 for the ATE-R and \$342 for the ATE, are nearly identical in magnitude to the na\"{i}ve estimate and are not close to being significantly different from it even at the 10% level. Thus, the CiC-corrected estimates for the ATE imply that the treatment effect is a 26% increase relative to the control while the IPW-corrected estimates imply that the treatment effect is a 31% increase. This difference in two corrected estimates is modest, but statistically significant at the 5% level and potentially meaningful in considering the returns to programs such as Progresa.
Since the different corrections we implement rely on different identifying assumptions, it is crucial to discuss which are more plausible in this setting. In order to shed some light on this question, we first consider whether attrition is likely to be causing a violation of internal validity that would necessitate a correction. Thus, we implement the tests proposed in GHO2022 to assess the impact of attrition on the internal validity of the estimated treatment effects in the Progresa study (see Panel C of Table (ref)). These tests rely on a panel framework, and are based on testable restrictions of the relevant identifying assumptions in the presence of attrition that ensure internal validity for the respondents (IVal-R), which identifies the ATE-R, and the study population (IVal-P), which identifies the ATE. The testable restriction for IVal-R consists of the equality of the mean baseline outcome across treatment and control groups conditional on response status, that is, $E[Y_{0}|G=0,R=r]=E[Y_{0}|G=1,R=r]$ for $r=0,1$. In contrast, the testable restriction for IVal-P consists of the equality of the mean baseline outcome across the four treatment-response subgroups, $E[Y_{0}|G=g,R=r]= E[Y_{0}]$ for $(g,r)\in\{0,1\}^2$.
In applying these tests to the productive asset outcome, we find that the test of internal validity for the respondents (IVal-R) is not rejected, whereas the test of interval validity for the study population (IVal-P) is rejected. Indeed, we find substantially higher mean baseline value of production animals for respondents relative to attritors. Thus, consistent with the findings of the CiC correction, the results of these tests for attrition bias indicate that a correction is required for the ATE but not the ATE-R.\footnote{We emphasize that while these tests are helpful in interpreting the corrections, they should not be used as pre-tests to decide whether to implement a correction or not due to the resulting pre-test bias issue.} A caveat, however, is that it is possible that the IVal-R test is unable to detect a violation of internal validity for respondents.
Next, we consider whether the CiC assumptions hold. The CiC approach under attrition relies on two high-level assumptions for the four treatment-response subgroups: a monotonic relationship between the outcome and its unobservable determinant as well as a conditional time-invariant distribution for the unobservable in question. Specifically, we apply the test of the implication of CiC assumptions under random assignment, $ATT_{CiC}=ATU_{CiC}$, (see Remark (ref)). As shown in Panel D of Table (ref), the CiC estimates of the ATT and ATU are almost exactly equal with a difference of $0.1$ and this difference is not statistically significant. Thus, since we do not find evidence that the CiC assumptions are violated here, the results of this test are consistent with interpreting the CiC corrections as estimates of causal impacts.
By contrast, the IPW conditions do not allow response to depend on treatment status, which implies equal attrition rates. In this setting, the differential rate of 2.6 percentage points, but is marginally not significant with a p-value 0.117.\footnote{In Section (ref) of the online appendix, we test the restriction of the two IPW assumptions required for the identification of the ATE-R. There, we marginally reject the equality of attrition rates across treatment and control groups, but do not reject the joint null that response is independent of treatment conditional on $Y_0$.} Furthermore, equal attrition rates are not sufficient to ensure that response does not depend on treatment when monotonicity of response is violated. For the ATE, the IPW further requires the assumption that attrition is random conditional on variable(s) that are included in the correction, which is not testable.
To complement the results of these tests of identifying assumptions, we also heuristically consider the plausibility of the CiC or the IPW assumptions in this setting by considering likely reasons for attrition. The IPW assumptions do not restrict the outcome model, but do restrict response. Meanwhile the CiC conditions restrict the outcome model, but they are compatible with a wide range of response models. If researchers have a specific response model in mind, they can apply that understanding in considering whether the conditional time-invariance assumption holds.\footnote{For more general comparisons of the CiC and IPW approaches under various response models, see Section (ref) and Section (ref).}
Other studies that propose corrections generally focus on one of two main reasons for attrition: participants may simply be reluctant to respond to interviews or migration may hamper the enumerators' ability to find and interview participants BCGB2015, MM2017. There are several factors that influence each of these reasons for attrition, however, and thus neither reason maps directly into a particular set of unobservables that determine response. For example, migration can be triggered primarily by a high tolerance for risk driving a willingness to search for better economic opportunities or by covariate shocks such as conflict and droughts. Likewise, the main determinant of the reluctance to respond to surveys in any given setting could be one of a range of different types of factors: (lack of) reciprocity, the sensitivity of the questions, and the opportunity cost of time.
It is not common practice for researchers to report analysis on reasons for attrition in field experiments, and thus it is not surprising that specific data on the reasons for attrition in the Progresa evaluation is not publicly available.\footnote{In conducting a review of attrition in 96 published field experiments, GHO2022 did not find that such studies discuss data on reasons for attrition in general. That is not surprising since collecting data on why it is not possible to find a specific respondent, when they cannot be found, may not be possible by construction. Of course, in some cases it is possible to collect such data from a respondent's associates. There may also be occasional circumstances in which general reasons for attrition may be understood without additional data collection, such as when civil unrest or a natural disaster drives displacement, even if such a reason cannot be linked to specific respondents. } Although stated reasons for attrition are not typically available, authors commonly examine drivers of attrition by implementing a determinants of attrition test that examines how respondents and attritors differ in terms of baseline covariates, which can help one infer what are the likely underlying unobservables determining response. In GHO2022, we find that authors conduct a determinants of attrition test for 29% of experiments where there is attrition. Thus, we implement such a test for this application using available baseline data (see Table (ref) in the online appendix).
While examining covariates of attrition cannot reveal a specific underlying response function, it can suggest potential patterns of attrition that are consistent with the results of the application and the tests described in Section (ref). For example, the findings from the determinants of attrition test are consistent with the idea that reluctance in the form of opportunity cost of time is a key reason for survey non-response. In particular, response is significantly correlated with the likelihood of having an adult at home or being a farm household, which is unsurprising given that households that engage in agricultural production are more likely to work at home and thus have a lower opportunity cost of time during the day to respond to a survey.\footnote{We define a farm household as one that used agricultural land or owned production animals at baseline, which differs from the outcome of value of production animals. Of course, alternative drivers of attrition are consistent with the findings from the determinants of attrition test as is discussed in Section (ref) of the online appendix.} Response here then is plausibly related to an unobservable determinant of the outcome in this setting, $U_t$, which could represent the skills to succeed in accumulating agricultural assets, for example. This would be consistent with the CiC assumptions if the joint distribution between agricultural skills and the opportunity cost of time is stable across time. Furthermore, since being a farm household likely depends on several unobservables that affect both response and potential outcomes at follow-up, we expect attrition to be nonrandom, even after conditioning on the baseline covariates. If that is the case, IPW assumptions for the ATE would not be satisfied.
Of course, relying on determinants of attrition tests have limitations, however, in fully capturing potential reasons for attrition. In particular, baseline data is less likely to be informative about stochastic factors, such as health shocks. It is plausible that such a factor is an unobservable determinant of agricultural assets, and also explains a differential reluctance to respond across treatment and control groups in this setting, even if the test of differential attrition rates here is marginally insignificant. As discussed in Section (ref), such a case can be accommodated by our model, but cannot be accommodated by the IPW assumptions. As we also discuss in that section, the IPW assumptions and the CiC assumptions would both hold in this setting because it is a randomized experiment, if: (i) $U$ is some constant factor across time such as ability, and (ii) response is not affected by treatment status. In general, however, the most realistic models of response are likely to be those in which more than one factor influences response. As discussed in Example (ref), such models can be accommodated by the CiC assumptions.
In this paper, we propose an attrition correction method for the average treatment effects on the respondents as well as the study population that relies on the CiC framework. We achieve identification through two main assumptions: continuity and strict monotonicity of each potential outcome in a scalar unobservable, and time invariance of the distribution of the unobservable conditional on treatment and response status. We then show that these assumptions are likely to hold in a range of typical settings, and can accommodate a variety of different response functions.
We further compare these assumptions with other widely-used approaches. We focus in particular on the comparison with the IPW correction, since it provides point-identification for the average treatment effect for respondents (ATE-R) and the average treatment effect from the study population (ATE). The IPW correction relies on the assumptions of response independent of treatment status and unconfoundness for the ATE-R as well as conditionally random attrition for the ATE. In contrast to the CiC approach, Lee bounds rely on monotonicity of response, and are not designed to provide bounds for the ATE. We illustrate the performance and plausibility of these corrections for an application to an outcome from the randomized evaluation of the Progresa conditional cash transfer program in Mexico.
The CiC approach provides point-estimates for continuous outcomes and bounds for discrete outcomes. Given researchers commonly consider both continuous and discrete or binary versions of the same outcome, this study highlights that there is potential value in focusing on continuous versions of outcomes when correcting for attrition. The CiC corrections proposed here do not require random assignment, but do require that baseline outcome is available. Thus, they can be applied to quasi-experimental difference-in-difference designs. There is an ongoing debate, however, about the value of collecting baseline data for randomized controlled trials. Of course, there are settings where the baseline outcome is degenerate by design, however when that is not the case, this paper points to the value of collecting baseline outcome data.