Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
47,921 characters · 9 sections · 58 citation commands
Revisiting the Analysis of Matched-Pair and Stratified Experiments in the Presence of Attrition
\and Meng Hsuan Hsieh\\ Ross School of Business\\ University of Michigan\\ [email removed]} \and Jizhou Liu\\ Booth School of Business\\ University of Chicago\\ [email removed]} \and Max Tabord-Meehan\\ Department of Economics\\ University of Chicago\\ [email removed]} }
KEYWORDS: Randomized controlled trial, attrition, matched pairs, stratified randomization, fixed effects
JEL classification codes: C12, C14
\thispagestyle{empty} \setcounter{page}{1}
In this paper we revisit some common recommendations regarding the analysis of matched-pair and stratified experimental designs in the presence of attrition. Here, we define attrition to mean that we do not observe outcomes for some subset of the experimental units. This situation may arise, for instance, if subjects refuse to participate in the experiment’s endline survey or if researchers lose track of subjects prior to observing their experimental outcomes.
Our main objective is to clarify a number of well-known claims about the practice of dropping pairs with an attrited unit in matched-pair designs. Specifically, when one unit in a pair is lost, several contradictory suggestions have been made in the literature about whether or not experimenters should drop the remaining unit in their analyses.\footnote{Appendix (ref) contains relevant excerpts from the referenced sources.} For instance, king2007politically and bruhn2009pursuit assert that a key advantage of matched-pair designs is that dropping pairs with an attrited unit may protect against attrition bias when attrition is a function of the matching variables. In contrast, glennerster2013running claim that dropping pairs may increase attrition bias, and point out that the widespread practice of including pair fixed effects in a regression of outcomes on treatment is equivalent to computing the difference-in-means estimator after dropping pairs. Accordingly, they go on to suggest that experimenters should instead stratify the units into larger groups if there is risk of attrition. donner2000design assert that dropping pairs with an attrited unit is a requirement in analyses of matched-pair designs with attrition, and characterize this as a weakness of matched-pair designs. As a result, they also recommend stratifying units into larger groups.
To address these claims, we first derive the estimands obtained from the difference-in-means estimator in a matched-pair design both when the observations from pairs with an attrited unit are retained and when they are dropped. We find that the estimand produced when retaining the units is simply the difference in the mean outcomes conditional on not attriting. In contrast, the estimand produced when dropping the units is a complicated function of the mean outcomes and attrition probabilities conditional on the matching variables. Using this result, we show that dropping pairs does not recover the average treatment effect when attrition is a function of the matching variables, and instead recovers a convex weighted average\footnote{Here and throughout the paper we define a convex weighted average to be a weighted average whose coefficients are non-negative and sum to one.} of conditional average treatment effects. Moreover, we argue that natural conditions under which this convex weighted average further collapses to the average treatment effect are in fact stronger than the condition that attrition is independent of experimental outcomes. From these results we conclude that, although dropping pairs may potentially help in recovering a convex weighted average of conditional average treatment effects, we find limited evidence to support the claims that dropping pairs in a matched-pair design helps protect against attrition bias more generally.
Next, to address the claims that the issues surrounding whether or not to drop pairs with an attrited unit can be resolved by instead stratifying the experiment into larger groups, we repeat the above exercise in the context of a stratified randomized experiment where the strata are made up of a large number of observations. To mirror the analysis carried out for matched pairs, we study the estimands obtained from a regression of outcomes on treatment with and without strata fixed effects. We find analogous results: the estimand produced when omitting strata fixed effects is once again the difference in mean outcomes conditional on not attriting, and the estimand produced when including strata fixed effects is a function of the mean outcomes and attrition probabilities conditional on the strata labels with very similar properties to what was obtained for matched pairs. From these results we conclude that we do not find compelling evidence to support the idea that stratifying into larger groups resolves the issues surrounding attrition that we explore in this paper.
Including pair fixed effects when conducting inference via linear regression is a widely adopted practice bruhn2009pursuit, and is numerically equivalent to dropping pairs with an attrited unit. As a consequence, inference considerations sometimes drive the discussion of whether or not to drop pairs glennerster2013running. However, in our view this should not play a primary role when deciding whether or not to drop pairs for three reasons. First, as we show in this paper, including vs. excluding pair fixed effects produces estimands with distinct interpretations in the presence of attrition. Second, as argued in bai2021inference and bugni2018inference (in settings without attrition), including pair/strata fixed effects is not a requirement for conducting valid inference on the ATE in matched-pair/stratified experiments, and there is no clear benefit obtained from doing so in general. Third, there are no formal results which justify the use of conventional robust standard errors in the presence of attrition (with or without fixed effects), and we conjecture that alternative inference procedures should be developed in this case (see Remark (ref) for a preliminary discussion). For these reasons, in this paper our primary focus is on studying the interpretation of the resulting estimands.
Finally, we explore the empirical relevance of our results using experimental data collected in groh2016macroinsurance as well as data collected from a systematic survey of all papers published in the American Economic Review (AER) and American Economic Journal: Applied Economics (AEJ: Applied) from 2020-2022 which conduct matched-pair or stratified experiments in the presence of attrition. Using these datasets we find that there can be noticeable differences between the point estimates obtained from dropping or retaining pairs with an attrited unit (or including/omitting stratum fixed effects), even when attrition is comparatively low. For instance, using the data in groh2016macroinsurance we find an average absolute percentage difference of $13.82\%$ in point estimates across a collection of outcomes even with an average attrition rate of only $1.4\%$.
Our paper is related to a large literature on the analysis of randomized experiments with attrition. Most of this literature focuses on developing methods to recover the average treatment effect, often by either modeling the missing data process heckman1979sample,rubin2004multiple, inverse probability weighting wooldridge2002inverse,little2019statistical, bounding manski2000analysis,lee2009training,behaghel2015please, or testing for the presence of attrition bias ghanem2021testing. Instead, the focus of our paper is on studying the behavior of commonly used estimators in the analysis of matched-pair and stratified experiments. To our knowledge, the paper most similar to ours is fukumoto2022nonignorable, who conducts finite population and super-population analyses of the bias and variance of the difference-in-means estimator in matched-pair designs with and without dropping pairs. However, his super-population analysis maintains a sampling framework where the observations are drawn together as pairs, whereas we consider a sampling framework where observations are drawn as individuals and then subsequently paired according to their covariates. As a consequence, his results and ours are not directly comparable fukumoto2022nonignorable. Moreover, fukumoto2022nonignorable exclusively focuses on the setting of matched-pair designs and thus does not derive results for stratified randomized experiments.
The rest of the paper is structured as follows. In Section (ref) we describe our setup and introduce the main assumptions we consider on the attrition process. Section (ref) presents the main results. In Section (ref) we present an empirical illustration. Finally, we conclude in Section (ref) with some recommendations for empirical practice.
Let $Y^*_i$ denote the realized outcome of interest for the $i$th unit in the absence of attrition, $D_i \in \{0,1\}$ denote treatment status for the $i$th unit and $X_i$ denote the observed, baseline covariates for the $i$th unit. Further denote by $Y_i(1)$ the potential outcome of the $i$th unit if treated and by $Y_i(0)$ the potential outcome if not treated. As usual, the realized outcome is related to the potential outcomes and treatment status by the relationship
We consider a framework which allows for the possibility that units collected in the baseline survey may drop out (attrit) after treatment is assigned. In particular, let $R_i \in \{0, 1\}$ be an indicator where $R_i = 1$ indicates the $i$th unit is present in the endline survey (i.e. has not attrited) and $R_i = 0$ indicates otherwise. Let $R_i(1)$ denote the potential attrition decision of the $i$th unit if treated, and $R_i(0)$ denote the potential attrition decision of the $i$th unit if not treated. As was the case for the realized outcome, the realized attrition decision is related to the potential attrition decisions and treatment status by the relationship
With these definitions in hand, we define the observed outcome to be
We note that the observed outcome is undefined if individual $i$ is not observed in the endline survey, and so we set it arbitrarily to zero in equation (ref).
We assume that we observe a sample $\{(Y_i, R_i, D_i, X_i): 1 \le i \le n\}$, obtained from i.i.d random variables $\{W_i : 1 \le i \le n\}$ where $W_i = (Y_i(1), Y_i(0), R_i(1), R_i(0), X_i)$. As a result, the distribution of the observed data is determined by ((ref)), ((ref)), ((ref)), $\{W_i : 1 \le i \le n\}$, and the mechanism for determining treatment assignment (which we specify in Sections (ref) and (ref)). We maintain the following assumption on $\{W_i: 1 \le i \le n\}$ throughout the entirety of the paper:
Assumption (ref)(a) imposes mild restrictions on the moments of the potential outcomes. Assumption (ref)(b) rules out situations where the probability of attrition is one for either treatment status.
Our parameter of interest is the average treatment effect, denoted as
Without further assumptions on the nature of attrition, $\theta$ is not point-identified from the observed data. As a consequence, in this paper we first study the estimands produced by commonly used estimators in the analysis of matched-pair and stratified randomized experiments, and then document if and when these estimands collapse to $\theta$ under well-known, albeit strong, assumptions on the attrition process; see Remark (ref) for further discussion. The first assumption we consider is that attrition is independent of the potential outcomes:
Under Assumption (ref), the average treatment effect $\theta$ is point-identified in a classical randomized experiment by simply comparing the mean outcomes under treatment and control for the non-attritors gerber2012field. The next assumption we consider is that attrition is independent of potential outcomes conditional on some set of observable characteristics:
Although Assumption (ref) does not necessarily imply Assumption (ref) or vice versa, it is often argued that Assumption (ref) may be easier to defend in practice moffit1999sample,hirano2001combining,gerber2012field,little2019statistical. Under Assumption (ref), $\theta$ is point-identified in a classical randomized experiment by first identifying the average treatment effect conditional on each value $C = c$ and then averaging these conditional treatment effects across $C$. Note that Assumption (ref) generalizes the assumption discussed in the introduction that attrition is a function of observable characteristics. The final assumption we consider is that attrition is independent of observable characteristics:
A useful observation for the discussion which follows is that, although Assumptions (ref) and (ref) are not nested, Assumptions (ref) and (ref) do in fact imply Assumption (ref). To see this, consider the following derivation:
where the first equality follows from the law of iterated expectations, the second equality from Assumption (ref), the third from Assumption (ref), and the fourth from the law of iterated expectations once again.
In this section we study the estimands produced by the difference-in-means estimator in a matched-pair design when the observations from pairs with an attrited unit are retained and when they are dropped. Before defining the estimators we provide a formal description of the treatment assignment mechanism. To simplify the exposition, we assume that $n$ is even for the remainder of Section (ref). For any random variable indexed by $i$, for example $D_i$, we denote by $D^{(n)}$ the random vector $(D_1, D_2, \ldots, D_n)$. Let $\pi = \pi_n(X^{(n)})$ be a permutation of $\{1, \ldots, n\}$, potentially dependent on $X^{(n)}$. The $n/2$ matched pairs are then represented by the sets \[ \left\{\{\pi(2j - 1), \pi(2j)\}: 1 \leq j \leq \frac{n}{2}\right\}~. \] In other words, pairs are formed by arranging observations in the order $\{\pi(1), \pi(2), \ldots, \pi(n)\}$ according to the permutation $\pi$, and then forming pairs from the adjacent units as $\{\pi(1), \pi(2)\}$, $\{\pi(3), \pi(4)\}$, etc. Next, given such a $\pi$, we assume treatment status is assigned as follows:
To summarize, the assignment mechanism first forms pairs of units (according to $\pi$) and then assigns both treatments exactly once in each pair at random. The first estimator we consider is the standard difference-in-means estimator computed on non-attritors:
Note that $\hat{\theta}_n$ may be obtained as the estimator of the coefficient on $D_i$ in an ordinary least squares regression of $Y_i$ on a constant and $D_i$, computed on the non-attritors. The second estimator we consider is the difference-in-means estimator computed by first dropping any observations belonging to a pair with an attritor:
Note that $\hat{\theta}_n^{\rm drop}$ corresponds to the estimator recommended in bruhn2009pursuit and king2007politically. We emphasize that, in the absence of attrition, $\hat{\theta}_n$ and $\hat{\theta}^{\rm drop}_n$ are numerically equivalent.
As a consequence of the Frisch-Waugh-Lovell theorem, $\hat{\theta}_n^{\rm drop}$ can equivalently be obtained as the ordinary least squares estimator of the coefficient on $D_i$ in the linear regression of $Y_i$ on $D_i$ and pair fixed effects computed on the non-attritors (i.e. individuals with $R_i = 1$):\footnote{See Appendix (ref) for a derivation of this fact.}
Similar regression specifications are extremely common in the analysis of matched-pair experiments. See, for example, ashraf2006deposit, angrist2009effects, crepon2015estimating, bruhn2016impact, and fryer2018pupil.
We impose the following assumption in addition to Assumption (ref):
Assumptions (ref)(a)--(b) are smoothness requirements that ensure that units that are “close” in terms of their baseline covariates are also “close” in terms of their potential attrition indicators and potential outcomes on average. Similar smoothness requirements are also imposed in bai2021inference and bai2022optimality.
Finally, we require that the matched-pair design is such that the units in each pair are “close” in terms of their baseline covariates in the following sense:
See bai2021inference for sufficient conditions for Assumption (ref). In particular, if $\mathrm{dim}(X_i) = 1$, then Assumption (ref) is satisfied if $E[|X_i|] < \infty$ and we construct pairs by simply ordering the units from smallest to largest according to $X_i$ and then pairing adjacent units. For the case $\mathrm{dim}(X_i) > 1$, bai2021inference provide sufficient conditions under which Assumption (ref) is satisfied when using the popular R package {\tt nbpMatching}. Using appropriate laws of large numbers developed in bai2021inference, we now establish the following result:
Theorem (ref) shows that the estimand produced by the difference-in-means estimator, $\theta^{\rm obs}$, is simply the difference in the mean outcomes conditional on not attriting (under the additional assumption that $R_i(1) = R_i(0)$ this could be interpreted as the average treatment effect for units who do not attrit: see Remark (ref) for details). It follows immediately that, under Assumption (ref), $\theta^{\rm obs} = \theta$ and thus under this assumption we recover the average treatment effect.
On the other hand, the estimand produced by first dropping units belonging to a pair with an attritor, $\theta^{\rm drop}$, is a complicated function of the mean outcomes and attrition probabilities conditional on the matching variables. First, note that unlike $\theta^{\rm obs}$, $\theta^{\rm drop}$ does not collapse to $\theta$ under Assumption (ref). Moreover, $\theta^{\rm drop}$ does not collapse to $\theta$ under Assumption (ref) with $C_i = X_i$ either. Instead, under Assumption (ref) with $C_i = X_i$, $\tau^{\rm obs}(x) = \tau(x)$ where $\tau(x) = E[Y_i(1) - Y_i(0)|X_i=x]$, so that \[\theta^{\rm drop} = E[\tau(X_i)\rho(X_i)]~,\] i.e. $\theta^{\rm drop}$ may be written as a convex weighted average of the conditional average treatment effects $\tau(x)$. In some special cases this convex-weighted average has a simple and transparent interpretation: consider for example a setting where $X_i$ is a binary variable, and suppose that attrition is such that units with $X_i = 1$ always appear in the endline survey, so that $R_i(1) = R_i(0) = 1$ if $X_i = 1$, but units with $X_i = 0$ appear only if they are treated, so that $R_i(1) = 1$ and $R_i(0) = 0$ if $X_i = 0$. Then, \[ \rho(1) = \frac{1}{P \{X_i = 1\}}~, \] and $\rho(0) = 0$. We thus have that in this case \[ \theta^{\rm drop} = E[Y_i(1) - Y_i(0) | X_i = 1]~, \] which is the average treatment effect for those units with $X_i = 1$. In contrast, $\theta^{\rm obs}$ does not lend itself to a straightforward causal interpretation in this example (however in Remark (ref) we provide a favorable interpretation of $\theta^{\rm obs}$ under Assumption (ref) and the additional assumption that $R_i(1) = R_i(0)$).
In general, straightforward algebra shows that $\rho(x) = 1$ if and only if
In words, $\rho(x) = 1$ if and only if the conditional probability of attrition under treatment is inversely proportional to the conditional probability of attrition under control. A natural assumption which guarantees ((ref)) for all $x$ is Assumption (ref) with $C_i = X_i$, so that attrition is independent of the matching variables $X_i$. Finally, we note that under Assumption (ref) with $C_i = X_i$, it follows that $\theta^{\rm drop} = \theta^{\rm obs}$. As a result, $\theta^{\rm drop} = \theta$ under Assumptions (ref) and (ref). We summarize the above discussion in the following corollary:
We conclude this section by noting that, as explained in the derivation following the statement of Assumption (ref), Assumptions (ref) and (ref) imply Assumption (ref). In other words, we see that the sufficient conditions provided in Corollary (ref) under which $\theta^{\rm drop} = \theta$ are in fact stronger than the conditions required for $\theta^{\rm obs} = \theta$. We thus find limited evidence to support the claims that dropping pairs in a matched-pair design helps in reducing attrition bias. However, we emphasize that dropping pairs may potentially help in recovering a convex weighted average of conditional average treatment effects.
In this section we repeat the exercise presented in Section (ref) but in the context of stratified designs. Before describing the estimators, we provide a description of the class of treatment assignment mechanisms we consider. In words, our results accommodate any treatment assignment mechanism which first partitions the covariate space into a finite number of “large” strata, and then performs treatment assignment independently across strata so as to achieve “balance” within each stratum. Formally, let $S:\text{supp}(X_i) \rightarrow \mathcal{S}$ be a function which maps the support of the covariates into a finite set $\mathcal{S}$ of strata labels. For $1 \le i \le n$, let $S_i = S(X_i)$ denote the strata label of individual $i$. For $s \in \mathcal{S}$, let \[D_n(s) = \sum_{1 \le i \le n}(D_i - \nu)I\{S_i = s\}~,\] where $\nu \in (0, 1)$ denotes the “target” proportion of units to assign to treatment in each stratum. Intuitively, $D_n(s)$ measures the amount of imbalance in stratum $s$ relative to the target proportion $\nu$. Our requirements on the treatment assignment mechanism can then be summarized as follows:
Assumption (ref)(a) simply requires that treatment assignment be exogenous conditional on the strata labels. Assumption (ref)(b) formalizes the requirement that the assignment mechanism performs treatment assignment so as to achieve “balance” within strata. Assumption (ref)(b) is a relatively mild assumption which is satisfied by most stratified randomization procedures employed in field experiments: see bugni2018inference for examples.
As before, the first estimator we consider is the standard difference-in-means estimator computed on non-attritors $\hat{\theta}_n$. The second estimator we consider, denoted $\hat{\theta}_n^{\rm sfe}$, is the estimator obtained as the estimator of the coefficient on $D_i$ in an ordinary least squares regression of $Y_i$ on $D_i$ and strata fixed effects computed on the non-attritors: \[Y_i = \theta^{\rm sfe}D_i + \sum_{s \in \mathcal{S}}\delta_sI\{S_i = s\} + \epsilon_i \hspace{3mm} \text{(for individuals with $R_i = 1$)}~.\] Similar regression specifications are extremely common in the analysis of stratified randomized experiments. See, for example, bruhn2009pursuit, duflo2015education, glennerster2013running, de_mel2019labor, and callen2020data. Using appropriate laws of large numbers developed in bugni2018inference, we now establish the following result:
The conclusions we draw from Theorem (ref) closely mirror those of Theorem (ref). In this case, under Assumption (ref) with $C_i = S_i$, \[\theta^{\rm sfe} = E\left[\tau(S_i)\lambda(S_i)\right]~,\] where $\tau(s) = E[Y_i(1) - Y_i(0)|S_i = s]$ and \[\lambda(s) = \left ( E \left [ \frac{E[R_i(1) | S_i] E[R_i(0) | S_i]}{\nu E[R_i(1) | S_i] + (1 - \nu) E[R_i(0) | S_i]} \right ] \right )^{-1}\times\frac{E[R_i(1)| S_i = s] E[R_i(0) | S_i = s]}{\nu E[R_i(1) | S_i=s] + (1 - \nu) E[R_i(0) | S_i=s]}~,\] so that $\theta^{\rm sfe}$ is also a convex weighted average of the strata-level treatment effects $\tau(s)$, although the weights $\lambda(s)$ are arguably more complicated to interpret than the weights $\rho(x)$ defined in Section (ref). Straightforward algebra shows that $\lambda(s) = 1$ if and only if
where $\Lambda = E \left [ \frac{E[R_i(1) | S_i] E[R_i(0) | S_i]}{\nu E[R_i(1) | S_i] + (1 - \nu) E[R_i(0) | S_i]} \right ]$. Conditions under which this holds seem difficult to articulate in words, but once again a natural assumption which guarantees ((ref)) for every $s \in \mathcal{S}$ is that Assumption (ref) is satisfied with $C_i = S_i$. We summarize these observations in the following corollary:
We conclude this section by stating that, given how closely the results presented in Section (ref) mirror those in Section (ref), we do not find compelling evidence to support the idea that stratifying into larger groups resolves the issues surrounding attrition that we explore in this paper.
In this section we illustrate the potential empirical relevance of deciding whether or not to drop pairs with an attrited unit using the experimental data collected in groh2016macroinsurance, which implemented a matched-pair design in the presence of attrition. The regression specifications in the paper contain pair fixed effects, which, as explained in Section (ref), is mechanically equivalent to dropping pairs with an attrited unit when regressing outcomes on a constant and treatment.
groh2016macroinsurance study the effect of insuring microenterprises (clients) against macroeconomic instability and political uncertainty in post-revolution Egypt. A baseline survey was completed for 2961 clients, who were then randomly assigned to treatment (1481 individuals) and control (1480 individuals) using a matched-pair design\footnote{Per the authors, they “created matched pairs [...] to minimize the Mahalanobis distance between the values of 13 variables that [they] hypothesized may determine loan take-up and investment decisions”. The final assignment contained one stratum with 16 individuals, each belonging to a different branch office. We follow the authors' methodology in keeping this stratum when we conduct our analysis in Table (ref). We drop these when we perform additional analyses in Table (ref).}. In Table (ref) we reproduce the intention-to-treat estimates from Table 7 of their paper, which presents estimated treatment effects on profits, revenues, employees and household consumption. “Original” corresponds to the estimates obtained from running the regression specifications in the original paper which include pair fixed effects, and $\hat{\theta}_n$ corresponds to estimates obtained from running an identical regression specification without pair fixed effects (we note that we were able successfully reproduce all of the reported estimates from the paper). We find an average absolute percentage difference of $13.82\%$\footnote{Here the absolute percentage difference is computed as $\left(\frac{|\text{Original} - \hat{\theta}_n|}{|\text{Original}|}\right)\times 100$.} for the point estimates of these effects, with the largest differences appearing for profits and revenue.
One caveat to the findings in Table (ref) is that the setting does not map exactly into our theoretical results: first, both regressions control for baseline covariates and second, the final assignment contained one stratum with 16 individuals, each belonging to a different branch office. Given this, in Table (ref) we report the intention-to-treat estimates without baseline covariates and without this additional stratum. In this case we find an average absolute percentage difference of $15.61\%$ for the point estimates of the effects. We emphasize that we consider these difference particularly salient given that attrition is quite low (on average $1.4\%$ across the outcomes), and that in the absence of attrition these estimates would be numerically identical, as illustrated from the estimates of the effect of treatment for monthly consumption.
Next, we perform a similar exercise using the data from a systematic survey of all papers published in the American Economic Review (AER) and the American Econonomic Journal: Applied Economics (AEJ: Applied) from 2020-2022 which conducted matched-pair or stratified randomized experiments in the presence of attrition. Our survey identified seven such papers: abebe2021selection, attanasio2020estimating, carter2021subsidies, casaburi2021using, dhar2022reshaping, hjort2021research, and romero2020outsourcing. For each paper, we collected a set of “relevant" regression specifications,\footnote{We note that in some papers such as attanasio2020estimating and casaburi2021using the primary results were not necessarily the output of a linear regression, and so in these cases we selected a collection of preliminary regression analyses. In other papers such as hjort2021research, the primary results were LATE estimates obtained via IV regression, and so in these cases we report the intention to treat analyses. Specific selection details for each paper are outlined in Appendix (ref).} and reproduced these regressions with and without pair/stratum fixed effects (we note that we were able to successfully reproduce all of the reported estimates from each paper). In Figure (ref) we report the average absolute percentage change (computed as $\left(\frac{|\text{Alternative} - \text{Original}|}{|\text{Original}|}\right)\times 100$, where “Original" corresponds to the point estimate computed in the paper, and “Alternative" corresponds to the estimate computed from the alternative specification with or without fixed effects) across all specifications for each paper. Similar to our findings for groh2016macroinsurance, we find that there can be noticeable differences in the point estimates with and without fixed effects (although we emphasize that we do not claim that these differences are necessarily statistically significant).
We conclude with some recommendations for empirical practice based on our theoretical results. Our main takeaway is that choosing whether or not to include pair/strata fixed effects when attrition is a concern can make a substantive difference to empirical findings and to the interpretation of the resulting estimand. In our view, unless practitioners are interested in recovering the convex-weighted averages produced by $\theta^{\rm drop}$ and $\theta^{\rm sfe}$ under a conditional independence assumption (Assumption (ref)), primary analyses should be based on regressions without pair/strata fixed effects: the resulting estimand $\theta^{\rm obs}$ has a simple interpretation in the absence of any assumptions, and collapses to the average treatment effect under arguably weaker assumptions than $\theta^{\rm drop}$ and $\theta^{\rm sfe}$. A secondary benefit of $\theta^{\rm obs}$ is that, under the additional assumption that $R_i(1) = R_i(0)$, $\theta^{\rm obs}$ also enjoys an interpretation as a convex-weighted average under Assumption (ref), with weights which may be more desirable than those appearing in $\theta^{\rm drop}$ or $\theta^{\rm sfe}$ in that they do not “double-up" on attrition: see Remark (ref) for details.