The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
47,731 characters
Revisiting the Analysis of Matched-Pair and Stratified Experiments in the Presence of Attrition
\author{
Yuehao Bai\\
Department of Economics\\
University of Southern California\\
[email removed]}
\and
Meng Hsuan Hsieh\\
Ross School of Business\\
University of Michigan\\
[email removed]}
\and
Jizhou Liu\\
Booth School of Business\\
University of Chicago\\
[email removed]}
\and
Max Tabord-Meehan\\
Department of Economics\\
University of Chicago\\
[email removed]}
}
\title{Revisiting the Analysis of Matched-Pair and Stratified Experiments in the Presence of Attrition \thanks{We thank Rachel Glennerster, Hongchang Guo, David McKenzie, Azeem Shaikh, Alex Torgovitsky, Ed Vytlacil, and three anonymous refereees for helpful comments. We also thank Lorenzo Casaburi and Tristan Reed for helpful comments and for sharing their data for one of our empirical applications. The fourth author acknowledges support from NSF grant SES-2149408.}}
\maketitle
\vspace{-0.3in}
\begin{spacing}{1.2}
\begin{abstract}
In this paper we revisit some common recommendations regarding the analysis of matched-pair and stratified experimental designs in the presence of attrition. Our main objective is to clarify a number of well-known claims about the practice of dropping pairs with an attrited unit when analyzing matched-pair designs. Contradictory advice appears in the literature about whether or not dropping pairs is beneficial or harmful, and stratifying into larger groups has been recommended as a resolution to the issue. To address these claims, we derive the estimands obtained from the difference-in-means estimator in a matched-pair design both when the observations from pairs with an attrited unit are retained and when they are dropped. We find limited evidence to support the claims that dropping pairs helps recover the average treatment effect, but we find that it may potentially help in recovering a convex weighted average of conditional average treatment effects. We report similar findings for stratified designs when studying the estimands obtained from a regression of outcomes on treatment with and without strata fixed effects.
\end{abstract}
\end{spacing}
\noindent \textsc{KEYWORDS}: Randomized controlled trial, attrition, matched pairs, stratified randomization, fixed effects
\noindent JEL classification codes: C12, C14
\thispagestyle{empty}
\newpage
\setcounter{page}{1}
\section{Introduction} \label{sec:intro}
In this paper we revisit some common recommendations regarding the analysis of matched-pair and stratified experimental designs in the presence of attrition. Here, we define attrition to mean that we do not observe outcomes for some subset of the experimental units. This situation may arise, for instance, if subjects refuse to participate in the experiment’s endline survey or if researchers lose track of subjects prior to observing their experimental outcomes.
Our main objective is to clarify a number of well-known claims about the practice of dropping pairs with an attrited unit in matched-pair designs. Specifically, when one unit in a pair is lost, several contradictory suggestions have been made in the literature about whether or not experimenters should drop the remaining unit in their analyses.\footnote{Appendix \ref{sec:quotes} contains relevant excerpts from the referenced sources.} For instance, \cite{king2007politically} and \cite{bruhn2009pursuit} assert that a key advantage of matched-pair designs is that dropping pairs with an attrited unit may protect against attrition bias when attrition is a function of the matching variables. In contrast, \cite{glennerster2013running} claim that dropping pairs may \emph{increase} attrition bias, and point out that the widespread practice of including pair fixed effects in a regression of outcomes on treatment is equivalent to computing the difference-in-means estimator after dropping pairs. Accordingly, they go on to suggest that experimenters should instead stratify the units into larger groups if there is risk of attrition. \cite{donner2000design} assert that dropping pairs with an attrited unit is a \emph{requirement} in analyses of matched-pair designs with attrition, and characterize this as a weakness of matched-pair designs. As a result, they also recommend stratifying units into larger groups.
To address these claims, we first derive the estimands obtained from the difference-in-means estimator in a matched-pair design both when the observations from pairs with an attrited unit are retained and when they are dropped. We find that the estimand produced when retaining the units is simply the difference in the mean outcomes conditional on not attriting. In contrast, the estimand produced when dropping the units is a complicated function of the mean outcomes and attrition probabilities conditional on the matching variables. Using this result, we show that dropping pairs does not recover the average treatment effect when attrition is a function of the matching variables, and instead recovers a convex weighted average\footnote{Here and throughout the paper we define a convex weighted average to be a weighted average whose coefficients are non-negative and sum to one.} of conditional average treatment effects. Moreover, we argue that natural conditions under which this convex weighted average further collapses to the average treatment effect are in fact stronger than the condition that attrition is independent of experimental outcomes. From these results we conclude that, although dropping pairs may potentially help in recovering a convex weighted average of conditional average treatment effects, we find limited evidence to support the claims that dropping pairs in a matched-pair design helps protect against attrition bias more generally.
Next, to address the claims that the issues surrounding whether or not to drop pairs with an attrited unit can be resolved by instead stratifying the experiment into larger groups, we repeat the above exercise in the context of a stratified randomized experiment where the strata are made up of a large number of observations. To mirror the analysis carried out for matched pairs, we study the estimands obtained from a regression of outcomes on treatment with and without strata fixed effects. We find analogous results: the estimand produced when omitting strata fixed effects is once again the difference in mean outcomes conditional on not attriting, and the estimand produced when including strata fixed effects is a function of the mean outcomes and attrition probabilities conditional on the strata labels with very similar properties to what was obtained for matched pairs. From these results we conclude that we do not find compelling evidence to support the idea that stratifying into larger groups resolves the issues surrounding attrition that we explore in this paper.
Including pair fixed effects when conducting inference via linear regression is a widely adopted practice \citep[see for instance the recommendations in][]{bruhn2009pursuit}, and is numerically equivalent to dropping pairs with an attrited unit. As a consequence, inference considerations sometimes drive the discussion of whether or not to drop pairs \citep[see for example Chapter 4, footnote 32 in][]{glennerster2013running}. However, in our view this should not play a primary role when deciding whether or not to drop pairs for three reasons. First, as we show in this paper, including vs. excluding pair fixed effects produces estimands with distinct interpretations in the presence of attrition. Second, as argued in \cite{bai2021inference} and \cite{bugni2018inference} (in settings without attrition), including pair/strata fixed effects is not a requirement for conducting valid inference on the ATE in matched-pair/stratified experiments, and there is no clear benefit obtained from doing so in general. Third, there are no formal results which justify the use of conventional robust standard errors in the presence of attrition (with or without fixed effects), and we conjecture that alternative inference procedures should be developed in this case (see Remark \ref{rem:inference} for a preliminary discussion). For these reasons, in this paper our primary focus is on studying the interpretation of the resulting estimands.
Finally, we explore the empirical relevance of our results using experimental data collected in \cite{groh2016macroinsurance} as well as data collected from a systematic survey of all papers published in the American Economic Review (AER) and American Economic Journal: Applied Economics (AEJ: Applied) from 2020-2022 which conduct matched-pair or stratified experiments in the presence of attrition. Using these datasets we find that there can be noticeable differences between the point estimates obtained from dropping or retaining pairs with an attrited unit (or including/omitting stratum fixed effects), even when attrition is comparatively low. For instance, using the data in \cite{groh2016macroinsurance} we find an average absolute percentage difference of $13.82\%$ in point estimates across a collection of outcomes even with an average attrition rate of only $1.4\%$.
Our paper is related to a large literature on the analysis of randomized experiments with attrition. Most of this literature focuses on developing methods to recover the average treatment effect, often by either modeling the missing data process \citep{heckman1979sample,rubin2004multiple}, inverse probability weighting \citep{wooldridge2002inverse,little2019statistical}, bounding \citep{manski2000analysis,lee2009training,behaghel2015please}, or testing for the presence of attrition bias \citep{ghanem2021testing}. Instead, the focus of our paper is on studying the behavior of commonly used estimators in the analysis of matched-pair and stratified experiments. To our knowledge, the paper most similar to ours is \cite{fukumoto2022nonignorable}, who conducts finite population and super-population analyses of the bias and variance of the difference-in-means estimator in matched-pair designs with and without dropping pairs. However, his super-population analysis maintains a sampling framework where the observations are drawn together as pairs, whereas we consider a sampling framework where observations are drawn as individuals and then subsequently paired according to their covariates. As a consequence, his results and ours are not directly comparable \citep[we note that every empirical application we consider in Section \ref{sec:application} describes a specific procedure by which they stratified their sample using available covariates, and thus does not feature a sample constructed from pre-formed strata as modelled in][]{fukumoto2022nonignorable}. Moreover, \cite{fukumoto2022nonignorable} exclusively focuses on the setting of matched-pair designs and thus does not derive results for stratified randomized experiments.
The rest of the paper is structured as follows. In Section \ref{sec:setup} we describe our setup and introduce the main assumptions we consider on the attrition process. Section \ref{sec:results} presents the main results. In Section \ref{sec:application} we present an empirical illustration. Finally, we conclude in Section \ref{sec:recs} with some recommendations for empirical practice.
\section{Setup and Notation}\label{sec:setup}
Let $Y^*_i$ denote the realized outcome of interest for the $i$th unit in the absence of attrition, $D_i \in \{0,1\}$ denote treatment status for the $i$th unit and $X_i$ denote the observed, baseline covariates for the $i$th unit. Further denote by $Y_i(1)$ the potential outcome of the $i$th unit if treated and by $Y_i(0)$ the potential outcome if not treated. As usual, the realized outcome is related to the potential outcomes and treatment status by the relationship
\begin{equation}\label{eq:PO}
Y^*_i = Y_i(1) D_i + Y_i(0) (1 - D_i)~.
\end{equation}
We consider a framework which allows for the possibility that units collected in the baseline survey may drop out (attrit) after treatment is assigned. In particular, let $R_i \in \{0, 1\}$ be an indicator where $R_i = 1$ indicates the $i$th unit is present in the endline survey (i.e. has \emph{not} attrited) and $R_i = 0$ indicates otherwise. Let $R_i(1)$ denote the potential attrition decision of the $i$th unit if treated, and $R_i(0)$ denote the potential attrition decision of the $i$th unit if not treated. As was the case for the realized outcome, the realized attrition decision is related to the potential attrition decisions and treatment status by the relationship
\begin{equation}\label{eq:PA}
R_i = R_i(1) D_i + R_i(0) (1 - D_i)~.
\end{equation}
With these definitions in hand, we define the observed outcome to be
\begin{equation}\label{eq:OO}
Y_i = Y^*_iR_i = Y_i(1)R_i(1)D_i + Y_i(0)R_i(0)(1 - D_i)~.
\end{equation}
We note that the observed outcome is undefined if individual $i$ is not observed in the endline survey, and so we set it arbitrarily to zero in equation \eqref{eq:OO}.
We assume that we observe a sample $\{(Y_i, R_i, D_i, X_i): 1 \le i \le n\}$, obtained from i.i.d random variables $\{W_i : 1 \le i \le n\}$ where $W_i = (Y_i(1), Y_i(0), R_i(1), R_i(0), X_i)$. As a result, the distribution of the observed data is determined by (\ref{eq:PO}), (\ref{eq:PA}), (\ref{eq:OO}), $\{W_i : 1 \le i \le n\}$, and the mechanism for determining treatment assignment (which we specify in Sections \ref{sec:pairs} and \ref{sec:sfe}). We maintain the following assumption on $\{W_i: 1 \le i \le n\}$ throughout the entirety of the paper:
\begin{assumption} \label{as:Q}
\hfill
\begin{enumerate}[(a)]
\item $E[|Y_i(d)|] < \infty$ for $d \in \{0, 1\}$.
\item $E[R_i(d)] > 0$ for $d \in \{0, 1\}$.
\end{enumerate}
\end{assumption}
Assumption \ref{as:Q}(a) imposes mild restrictions on the moments of the potential outcomes. Assumption \ref{as:Q}(b) rules out situations where the probability of attrition is one for either treatment status.
Our parameter of interest is the average treatment effect, denoted as
\begin{equation} \label{eq:ate-na}
\theta = E[Y_i(1) - Y_i(0)]~.
\end{equation}
Without further assumptions on the nature of attrition, $\theta$ is not point-identified from the observed data. As a consequence, in this paper we first study the estimands produced by commonly used estimators in the analysis of matched-pair and stratified randomized experiments, and then document if and when these estimands collapse to $\theta$ under well-known, albeit strong, assumptions on the attrition process; see Remark \ref{rem:bias} for further discussion. The first assumption we consider is that attrition is independent of the potential outcomes:
\begin{assumption}\label{as:MIPO}
\[(Y_i(1), Y_i(0)) \perp \!\!\! \perp (R_i(1), R_i(0))~.\]
\end{assumption}
Under Assumption \ref{as:MIPO}, the average treatment effect $\theta$ is point-identified in a classical randomized experiment by simply comparing the mean outcomes under treatment and control for the non-attritors \citep[see for instance][] {gerber2012field}. The next assumption we consider is that attrition is independent of potential outcomes conditional on some set of observable characteristics:
\begin{assumption}\label{as:MIPO|T}
For some set of observable characteristics $C_i$,
\[(Y_i(1), Y_i(0)) \perp \!\!\! \perp (R_i(1), R_i(0)) \hspace{1mm} | \hspace{1mm} C_i~.\]
\end{assumption}
Although Assumption \ref{as:MIPO} does not necessarily imply Assumption \ref{as:MIPO|T} or vice versa, it is often argued that Assumption \ref{as:MIPO|T} may be easier to defend in practice \citep{moffit1999sample,hirano2001combining,gerber2012field,little2019statistical}. Under Assumption \ref{as:MIPO|T}, $\theta$ is point-identified in a classical randomized experiment by first identifying the average treatment effect conditional on each value $C = c$ and then averaging these conditional treatment effects across $C$. Note that Assumption \ref{as:MIPO|T} generalizes the assumption discussed in the introduction that attrition is a function of observable characteristics. The final assumption we consider is that attrition is independent of observable characteristics:
\begin{assumption}\label{as:indep_attrition}
For some set of observable characteristics $C_i$,
\[C_i \perp \!\!\! \perp (R_i(1), R_i(0))~.\]
\end{assumption}
A useful observation for the discussion which follows is that, although Assumptions \ref{as:MIPO} and \ref{as:MIPO|T} are not nested, Assumptions \ref{as:MIPO|T} and \ref{as:indep_attrition} do in fact imply Assumption \ref{as:MIPO}. To see this, consider the following derivation:
\begin{align*}
& P\{(Y_i(1), Y_i(0)) \in U_1, (R_i(1), R_i(0)) \in U_2\} \\ &=
E\left[E[I\{(Y_i(1), Y_i(0)) \in U_1\}I\{(R_i(1), R_i(0)) \in U_2\}|C_i]\right]\\
& = E\left[E[I\{(Y_i(1), Y_i(0)) \in U_1 | C_i]E[I\{(R_i(1), R_i(0)) \in U_2\}|C_i]\right] \\
& = E\left[E[I\{(Y_i(1), Y_i(0)) \in U_1 | C_i]E[I\{(R_i(1), R_i(0)) \in U_2\}]\right] \\
& = E\left[I\{(Y_i(1), Y_i(0)) \in U_1\}\right]E[I\{(R_i(1), R_i(0)) \in U_2\}] \\
& = P\{(Y_i(1), Y_i(0)) \in U_1\}P\{(R_i(1), R_i(0)) \in U_2\}~,
\end{align*}
where the first equality follows from the law of iterated expectations, the second equality from Assumption \ref{as:MIPO|T}, the third from Assumption \ref{as:indep_attrition}, and the fourth from the law of iterated expectations once again.
\section{Main Results}\label{sec:results}
\subsection{Matched-Pair Designs with Attrition}\label{sec:pairs}
In this section we study the estimands produced by the difference-in-means estimator in a matched-pair design when the observations from pairs with an attrited unit are retained and when they are dropped. Before defining the estimators we provide a formal description of the treatment assignment mechanism. To simplify the exposition, we assume that $n$ is even for the remainder of Section \ref{sec:pairs}. For any random variable indexed by $i$, for example $D_i$, we denote by $D^{(n)}$ the random vector $(D_1, D_2, \ldots, D_n)$. Let $\pi = \pi_n(X^{(n)})$ be a permutation of $\{1, \ldots, n\}$, potentially dependent on $X^{(n)}$. The $n/2$ matched pairs are then represented by the sets
\[ \left\{\{\pi(2j - 1), \pi(2j)\}: 1 \leq j \leq \frac{n}{2}\right\}~. \]
In other words, pairs are formed by arranging observations in the order $\{\pi(1), \pi(2), \ldots, \pi(n)\}$ according to the permutation $\pi$, and then forming pairs from the adjacent units as $\{\pi(1), \pi(2)\}$, $\{\pi(3), \pi(4)\}$, etc. Next, given such a $\pi$, we assume treatment status is assigned as follows:
\begin{assumption} \label{as:mp}
Treatment status is assigned so that
\[ (Y^{(n)}(1), Y^{(n)}(0), R^{(n)}(1), R^{(n)}(0)) \perp \!\!\! \perp D^{(n)} | X^{(n)} \]
and, conditional on $X^{(n)}$, $(D_{\pi(2j - 1)}, D_{\pi(2j)}), 1 \leq j \leq n/2$ are i.i.d.\ and each uniformly distributed over $\{(0, 1), (1, 0)\}$.
\end{assumption}
To summarize, the assignment mechanism first forms pairs of units (according to $\pi$) and then assigns both treatments exactly once in each pair at random. The first estimator we consider is the standard difference-in-means estimator computed on non-attritors:
\begin{equation} \label{eq:est}
\hat \theta_n = \frac{\sum_{1 \leq i \leq n} Y_i R_i D_i}{\sum_{1 \leq i \leq n} R_i D_i} - \frac{\sum_{1 \leq i \leq n} Y_i R_i (1 - D_i)}{\sum_{1 \leq i \leq n} R_i (1 - D_i)}~.
\end{equation}
Note that $\hat{\theta}_n$ may be obtained as the estimator of the coefficient on $D_i$ in an ordinary least squares regression of $Y_i$ on a constant and $D_i$, computed on the non-attritors. The second estimator we consider is the difference-in-means estimator computed by first dropping any observations belonging to a pair with an attritor:
\begin{multline*}
\hat \theta_n^{\rm drop} = \left ( \sum_{1 \leq j \leq n/2} R_{\pi(2j - 1)} R_{\pi(2j)} \right )^{-1} \\
\times \left( \sum_{1 \leq j \leq n/2} R_{\pi(2j - 1)} R_{\pi(2j)} (Y_{\pi(2j - 1)} - Y_{\pi(2j)}) (D_{\pi(2j - 1)} - D_{\pi(2j)}) \right)~.
\end{multline*}
Note that $\hat{\theta}_n^{\rm drop}$ corresponds to the estimator recommended in \cite{bruhn2009pursuit} and \cite{king2007politically}. We emphasize that, in the absence of attrition, $\hat{\theta}_n$ and $\hat{\theta}^{\rm drop}_n$ are numerically equivalent.
As a consequence of the Frisch-Waugh-Lovell theorem, $\hat{\theta}_n^{\rm drop}$ can equivalently be obtained as the ordinary least squares estimator of the coefficient on $D_i$ in the linear regression of $Y_i$ on $D_i$ and pair fixed effects computed on the non-attritors (i.e. individuals with $R_i = 1$):\footnote{See Appendix \ref{sec:FWL} for a derivation of this fact.}
\begin{equation} \label{eq:pfe}
Y_i = \theta^{\rm drop} D_i + \sum_{1 \leq j \leq n/2} \delta_j I \{i \in \{\pi(2j - 1), \pi(2j)\}\} + \epsilon_i \hspace{3mm} \text{(for individuals with $R_i = 1$)}~.
\end{equation}
Similar regression specifications are extremely common in the analysis of matched-pair experiments. See, for example, \cite{ashraf2006deposit}, \cite{angrist2009effects}, \cite{crepon2015estimating}, \cite{bruhn2016impact}, and \cite{fryer2018pupil}.
We impose the following assumption in addition to Assumption \ref{as:Q}:
\begin{assumption} \label{as:Q-lip}
\hfill
\begin{enumerate}[(a)]
\item $E[R_i(d) | X_i = x]$ is Lipschitz in $x$ for $d \in \{0, 1\}$.
\item $E[Y_i(d) R_i(d) | X_i = x]$ is Lipschitz in $x$ for $d \in \mathcal \{0, 1\}$.
\end{enumerate}
\end{assumption}
\noindent Assumptions \ref{as:Q-lip}(a)--(b) are smoothness requirements that ensure that units that are ``close'' in terms of their baseline covariates are also ``close'' in terms of their potential attrition indicators and potential outcomes on average. Similar smoothness requirements are also imposed in \cite{bai2021inference} and \cite{bai2022optimality}.
Finally, we require that the matched-pair design is such that the units in each pair are ``close'' in terms of their baseline covariates in the following sense:
\begin{assumption} \label{as:close}
The pairs used in determining treatment status satisfy
\[ \frac{1}{n} \sum_{1 \leq j \leq n} \|X_{\pi(2j -1)} - X_{\pi(2j)}\| \stackrel{P}{\to} 0~. \]
\end{assumption}
See \cite{bai2021inference} for sufficient conditions for Assumption \ref{as:close}. In particular, if $\mathrm{dim}(X_i) = 1$, then Assumption \ref{as:close} is satisfied if $E[|X_i|] < \infty$ and we construct pairs by simply ordering the units from smallest to largest according to $X_i$ and then pairing adjacent units. For the case $\mathrm{dim}(X_i) > 1$, \cite{bai2021inference} provide sufficient conditions under which Assumption \ref{as:close} is satisfied when using the popular R package {\tt nbpMatching}. Using appropriate laws of large numbers developed in \cite{bai2021inference}, we now establish the following result:
\begin{theorem} \label{thm:pair}
Suppose the data satisfy Assumptions \ref{as:Q} and \ref{as:Q-lip} and the treatment assignment mechanism satisfies Assumptions \ref{as:mp} and \ref{as:close}. Then, as $n \to \infty$, $\hat \theta_n \stackrel{P}{\to} \theta^{\rm obs}$, where
\[\theta^{\rm obs} = \frac{E[R_i(1) Y_i(1)]}{E[R_i(1)]} - \frac{E[R_i(0) Y_i(0)]}{E[R_i(0)]} = E[Y_i(1)|R_i(1) = 1] - E[Y_i(0)|R_i(0) = 1]~, \]
and $\hat \theta_n^{\rm drop} \stackrel{P}{\to} \theta^{\rm drop}$, where
\[\theta^{\rm drop} = E[\tau^{\rm obs}(X_i) \rho(X_i)]~,\] with
\begin{align*}
\tau^{\rm obs}(x) & = E[Y_i(1) | R_i(1) = 1, X_i = x] - E[Y_i(0) | R_i(0) = 1, X_i = x] \\
\rho(x) & = \frac{E[R_i(0)|X_i=x]E[R_i(1)|X_i=x]}{E[E[R_i(0)|X_i]E[R_i(1)|X_i]]}~.
\end{align*}
\end{theorem}
Theorem \ref{thm:pair} shows that the estimand produced by the difference-in-means estimator, $\theta^{\rm obs}$, is simply the difference in the mean outcomes conditional on not attriting (under the additional assumption that $R_i(1) = R_i(0)$ this could be interpreted as the average treatment effect for units who do not attrit: see Remark \ref{rem:attrit_units} for details). It follows immediately that, under Assumption \ref{as:MIPO}, $\theta^{\rm obs} = \theta$ and thus under this assumption we recover the average treatment effect.
On the other hand, the estimand produced by first dropping units belonging to a pair with an attritor, $\theta^{\rm drop}$, is a complicated function of the mean outcomes and attrition probabilities conditional on the matching variables. First, note that unlike $\theta^{\rm obs}$, $\theta^{\rm drop}$ does not collapse to $\theta$ under Assumption \ref{as:MIPO}. Moreover, $\theta^{\rm drop}$ does not collapse to $\theta$ under Assumption \ref{as:MIPO|T} with $C_i = X_i$ either. Instead, under Assumption \ref{as:MIPO|T} with $C_i = X_i$, $\tau^{\rm obs}(x) = \tau(x)$ where $\tau(x) = E[Y_i(1) - Y_i(0)|X_i=x]$, so that
\[\theta^{\rm drop} = E[\tau(X_i)\rho(X_i)]~,\]
i.e. $\theta^{\rm drop}$ may be written as a convex weighted average of the conditional average treatment effects $\tau(x)$. In some special cases this convex-weighted average has a simple and transparent interpretation: consider for example a setting where $X_i$ is a binary variable, and suppose that attrition is such that units with $X_i = 1$ always appear in the endline survey, so that $R_i(1) = R_i(0) = 1$ if $X_i = 1$, but units with $X_i = 0$ appear only if they are treated, so that $R_i(1) = 1$ and $R_i(0) = 0$ if $X_i = 0$. Then,
\[ \rho(1) = \frac{1}{P \{X_i = 1\}}~, \]
and $\rho(0) = 0$. We thus have that in this case
\[ \theta^{\rm drop} = E[Y_i(1) - Y_i(0) | X_i = 1]~, \]
which is the average treatment effect for those units with $X_i = 1$. In contrast, $\theta^{\rm obs}$ does not lend itself to a straightforward causal interpretation in this example (however in Remark \ref{rem:attrit_units} we provide a favorable interpretation of $\theta^{\rm obs}$ under Assumption \ref{as:MIPO|T} and the additional assumption that $R_i(1) = R_i(0)$).
In general, straightforward algebra shows that $\rho(x) = 1$ if and only if
\begin{equation}\label{eq:proportion}
E[R_i(0)|X_i = x] = \frac{E[E[R_i(0)|X_i]E[R_i(1)|X_i]]}{E[R_i(1)|X_i = x]}~.
\end{equation}
In words, $\rho(x) = 1$ if and only if the conditional probability of attrition under treatment is inversely proportional to the conditional probability of attrition under control. A natural assumption which guarantees (\ref{eq:proportion}) for all $x$ is Assumption \ref{as:indep_attrition} with $C_i = X_i$, so that attrition is independent of the matching variables $X_i$. Finally, we note that under Assumption \ref{as:indep_attrition} with $C_i = X_i$, it follows that $\theta^{\rm drop} = \theta^{\rm obs}$. As a result, $\theta^{\rm drop} = \theta$ under Assumptions \ref{as:MIPO} and \ref{as:indep_attrition}. We summarize the above discussion in the following corollary:
\begin{corollary}\label{cor:drop_recover}
\hfill
\begin{enumerate}[(a)]
\item Under Assumption \ref{as:MIPO}, $\theta^{\rm obs} = \theta$.
\item Under Assumption \ref{as:MIPO|T} with $C_i = X_i$, $\theta^{\rm drop} = E[\tau(X_i)\rho(X_i)]$.
\item Under Assumption \ref{as:indep_attrition} with $C_i = X_i$, $\theta^{\rm drop} = \theta^{\rm obs}$.
\item Under Assumption \ref{as:indep_attrition} with $C_i = X_i$ and either Assumption \ref{as:MIPO} or Assumption \ref{as:MIPO|T} with $C_i = X_i$, $\theta^{\rm drop} = \theta$.
\end{enumerate}
\end{corollary}
We conclude this section by noting that, as explained in the derivation following the statement of Assumption \ref{as:indep_attrition}, Assumptions \ref{as:MIPO|T} and \ref{as:indep_attrition} \emph{imply} Assumption \ref{as:MIPO}. In other words, we see that the sufficient conditions provided in Corollary \ref{cor:drop_recover} under which $\theta^{\rm drop} = \theta$ are in fact \emph{stronger} than the conditions required for $\theta^{\rm obs} = \theta$. We thus find limited evidence to support the claims that dropping pairs in a matched-pair design helps in reducing attrition bias. However, we emphasize that dropping pairs may potentially help in recovering a convex weighted average of conditional average treatment effects.
\begin{remark}\label{rem:inference}
In Appendix \ref{sec:norm_theta} we develop the requisite distributional results to use $\hat{\theta}_n$ for inference about $\theta^{\rm obs}$. In contrast, the large sample distribution of $\hat{\theta}^{\rm drop}_n$ seems non-trivial to characterize and may in fact feature an asymptotic bias in general. For this reason we leave an in-depth study of the limiting distribution of $\hat{\theta}^{\rm drop}_n$ to future work.
\end{remark}
\begin{remark}\label{rem:bias}
We note that, in the absence of additional assumptions like Assumptions \ref{as:MIPO}--\ref{as:indep_attrition}, we are not able to conclude that either $\theta^{\rm obs}$ or $\theta^{\rm drop}$ is less biased for $\theta$ relative to the other, and in fact it is possible to construct data generating processes where either estimand is closer to the true average treatment effect. We present a concrete construction of such a set of DGPs in Appendix \ref{sec:numerical}.
\end{remark}
\begin{remark}\label{rem:attrit_units}
From Theorem \ref{thm:pair} we also observe that, in the absence of additional assumptions like Assumptions \ref{as:MIPO}--\ref{as:indep_attrition}, neither $\theta^{\rm obs}$ nor $\theta^{\rm drop}$ can be interpreted as a ``treatment effect parameter" (in the sense that neither parameter can be interpreted as an average treatment effect for some subset of individuals or more generally as a weighted average of treatment effects). This is because the subgroup of units who attrit under treatment ($R_i(1) = 0$) may not correspond to the subgroup of units who attrit under control ($R_i(0) = 0$). Under the additional assumption that $R_i(1) = R_i(0)$, so that these subgroups coincide, we obtain
\[\theta^{\rm obs} = E[Y_i(1) - Y_i(0) | R_i = 1]~,\]
and then $\theta^{\rm obs}$ could be understood as the average treatment effect for units who do not attrit. Imposing the same assumption for $\theta^{\rm drop}$ we obtain that
\[\theta^{\rm drop} = \frac{E[\left(Y_i(1) - Y_i(0)\right) R_i E[R_i | X_i]]}{E[R_iE[R_i | X_i]]}~,\]
and then $\theta^{\rm drop}$ could be understood as a ``probability of attrition"-weighted average of individual-level treatment effects for the non-attritors. If we additionally impose Assumption \ref{as:MIPO|T}, we alternatively obtain
\begin{equation*}
\theta^{\rm obs} = \frac{E[\tau(X_i)E[R_i | X_i]]}{P(R_i=1)}, \quad \theta^{\rm drop} = \frac{E[\tau(X_i)E[R_i | X_i]^2]}{E[E[R_i | X_i]^2]}~.
\end{equation*}
In this case, \emph{both} parameters can be interpreted as convex weighted averages of conditional average treatment effects, with the main difference being that $\theta^{\rm obs}$ is weighted using the the conditional attrition rate $E[R_i | X_i]$, whereas $\theta^{\rm drop}$ ``doubles down" by weighting using the squared conditional attrition rate $E[R_i | X_i]^2$.
\end{remark}
\begin{remark} \label{rem:bfe}
Regardless of whether or not a practitioner finds the interpretation of $\theta^{\rm drop}$ more or less attractive than the interpretation of $\theta^{\rm obs}$, it is crucial to note that, \emph{even in the absence of attrition}, inferences produced using robust standard errors obtained from a regression with pair fixed effects are generally conservative, but in some cases may in fact be \emph{invalid}, in the sense that the limiting rejection probability could be strictly larger than the nominal level. See \cite{bai2022inference} and \cite{de2020level} for details.
\end{remark}
\subsection{Stratified Designs with Attrition}\label{sec:sfe}
In this section we repeat the exercise presented in Section \ref{sec:pairs} but in the context of stratified designs. Before describing the estimators, we provide a description of the class of treatment assignment mechanisms we consider. In words, our results accommodate any treatment assignment mechanism which first partitions the covariate space into a finite number of ``large'' strata, and then performs treatment assignment independently across strata so as to achieve ``balance'' within each stratum. Formally, let $S:\text{supp}(X_i) \rightarrow \mathcal{S}$ be a function which maps the support of the covariates into a finite set $\mathcal{S}$ of strata labels. For $1 \le i \le n$, let $S_i = S(X_i)$ denote the strata label of individual $i$. For $s \in \mathcal{S}$, let
\[D_n(s) = \sum_{1 \le i \le n}(D_i - \nu)I\{S_i = s\}~,\]
where $\nu \in (0, 1)$ denotes the ``target'' proportion of units to assign to treatment in each stratum. Intuitively, $D_n(s)$ measures the amount of imbalance in stratum $s$ relative to the target proportion $\nu$. Our requirements on the treatment assignment mechanism can then be summarized as follows:
\begin{assumption} \label{as:strat}
The treatment assignment mechanism is such that
\begin{enumerate}[(a)]
\item $W^{(n)} \perp \!\!\! \perp D^{(n)} | S^{(n)}$.
\item $\frac{D_n(s)}{n} \stackrel{P}{\to} 0$ for every $s \in \mathcal{S}$.
\end{enumerate}
\end{assumption}
Assumption \ref{as:strat}(a) simply requires that treatment assignment be exogenous conditional on the strata labels. Assumption \ref{as:strat}(b) formalizes the requirement that the assignment mechanism performs treatment assignment so as to achieve ``balance'' within strata. Assumption \ref{as:strat}(b) is a relatively mild assumption which is satisfied by most stratified randomization procedures employed in field experiments: see \cite{bugni2018inference} for examples.
As before, the first estimator we consider is the standard difference-in-means estimator computed on non-attritors $\hat{\theta}_n$. The second estimator we consider, denoted $\hat{\theta}_n^{\rm sfe}$, is the estimator obtained as the estimator of the coefficient on $D_i$ in an ordinary least squares regression of $Y_i$ on $D_i$ and strata fixed effects computed on the non-attritors:
\[Y_i = \theta^{\rm sfe}D_i + \sum_{s \in \mathcal{S}}\delta_sI\{S_i = s\} + \epsilon_i \hspace{3mm} \text{(for individuals with $R_i = 1$)}~.\]
Similar regression specifications are extremely common in the analysis of stratified randomized experiments. See, for example, \cite{bruhn2009pursuit}, \cite{duflo2015education}, \cite{glennerster2013running}, \cite{de_mel2019labor}, and \cite{callen2020data}. Using appropriate laws of large numbers developed in \cite{bugni2018inference}, we now establish the following result:
\begin{theorem}\label{thm:covariate-adaptive}
Suppose the data satisfy Assumptions \ref{as:Q} and the treatment assignment mechanism satisfies Assumption \ref{as:strat}, Then as $n \to \infty$, $\hat{\theta}_n \stackrel{P}{\to} \theta^{\rm obs}$, where
\[\theta^{\rm obs} = \frac{E[R_i(1) Y_i(1)]}{E[R_i(1)]} - \frac{E[R_i(0) Y_i(0)]}{E[R_i(0)]} = E[Y_i(1)|R_i(1) = 1] - E[Y_i(0)|R_i(0) = 1]~,\]
and $\hat{\theta}^{\rm sfe}_n \stackrel{P}{\to} \theta^{\rm sfe}$, where
\begin{multline*}
\theta^{\rm sfe} = \left ( E \left [ \frac{E[R_i(1) | S_i] E[R_i(0) | S_i]}{\nu E[R_i(1) | S_i] + (1 - \nu) E[R_i(0) | S_i]} \right ] \right )^{-1} \\
\times
E \left [ \frac{E[R_i(1) Y_i(1) | S_i] E[R_i(0) | S_i] - E[R_i(0) Y_i(0) | S_i] E[R_i(1) | S_i]}{\nu E[R_i(1) | S_i] + (1 - \nu) E[R_i(0) | S_i]} \right ]~.
\end{multline*}
\end{theorem}
The conclusions we draw from Theorem \ref{thm:covariate-adaptive} closely mirror those of Theorem \ref{thm:pair}. In this case, under Assumption \ref{as:MIPO|T} with $C_i = S_i$,
\[\theta^{\rm sfe} = E\left[\tau(S_i)\lambda(S_i)\right]~,\]
where $\tau(s) = E[Y_i(1) - Y_i(0)|S_i = s]$ and
\[\lambda(s) = \left ( E \left [ \frac{E[R_i(1) | S_i] E[R_i(0) | S_i]}{\nu E[R_i(1) | S_i] + (1 - \nu) E[R_i(0) | S_i]} \right ] \right )^{-1}\times\frac{E[R_i(1)| S_i = s] E[R_i(0) | S_i = s]}{\nu E[R_i(1) | S_i=s] + (1 - \nu) E[R_i(0) | S_i=s]}~,\]
so that $\theta^{\rm sfe}$ is also a convex weighted average of the strata-level treatment effects $\tau(s)$, although the weights $\lambda(s)$ are arguably more complicated to interpret than the weights $\rho(x)$ defined in Section \ref{sec:pairs}. Straightforward algebra shows that $\lambda(s) = 1$ if and only if
\begin{equation}\label{eq:strat_proportion}
E[R_i(1)|S_i = s] = \frac{E[R_i(0)|S_i = s](1 - \nu)\Lambda}{E[R_i(0)|S_i = s] - \Lambda\nu}~,
\end{equation}
where $\Lambda = E \left [ \frac{E[R_i(1) | S_i] E[R_i(0) | S_i]}{\nu E[R_i(1) | S_i] + (1 - \nu) E[R_i(0) | S_i]} \right ]$. Conditions under which this holds seem difficult to articulate in words, but once again a natural assumption which guarantees (\ref{eq:strat_proportion}) for every $s \in \mathcal{S}$ is that Assumption \ref{as:indep_attrition} is satisfied with $C_i = S_i$. We summarize these observations in the following corollary:
\begin{corollary}\label{cor:strat_recover}
\hfill
\begin{enumerate}[(a)]
\item Under Assumption \ref{as:MIPO}, $\theta^{\rm obs} = \theta$.
\item Under Assumption \ref{as:MIPO|T} with $C_i = S_i$, $\theta^{\rm sfe} = E[\tau(S_i)\lambda(S_i)]$.
\item Under Assumption \ref{as:indep_attrition} with $C_i = S_i$, $\theta^{\rm sfe} = \theta^{\rm obs}$.
\item Under Assumption \ref{as:indep_attrition} with $C_i = S_i$ and either Assumption \ref{as:MIPO} or Assumption \ref{as:MIPO|T} with $C_i = S_i$, we obtain $\theta^{\rm sfe} = \theta$.
\end{enumerate}
\end{corollary}
We conclude this section by stating that, given how closely the results presented in Section \ref{sec:sfe} mirror those in Section \ref{sec:pairs}, we do not find compelling evidence to support the idea that stratifying into larger groups resolves the issues surrounding attrition that we explore in this paper.
\section{Empirical Illustrations}\label{sec:application}
\subsection{Re-analysis of \cite{groh2016macroinsurance}}
In this section we illustrate the potential empirical relevance of deciding whether or not to drop pairs with an attrited unit using the experimental data collected in \cite{groh2016macroinsurance}, which implemented a matched-pair design in the presence of attrition. The regression specifications in the paper contain pair fixed effects, which, as explained in Section \ref{sec:pairs}, is mechanically equivalent to dropping pairs with an attrited unit when regressing outcomes on a constant and treatment.
\cite{groh2016macroinsurance} study the effect of insuring microenterprises (clients) against macroeconomic instability and political uncertainty in post-revolution Egypt. A baseline survey was completed for 2961 clients, who were then randomly assigned to treatment (1481 individuals) and control (1480 individuals) using a matched-pair design\footnote{Per the authors, they ``created matched pairs [...] to minimize the Mahalanobis distance between the values of 13 variables that [they] hypothesized may determine loan take-up and investment decisions''. The final assignment contained \textit{one} stratum with 16 individuals, each belonging to a different branch office. We follow the authors' methodology in \textit{keeping} this stratum when we conduct our analysis in Table \ref{table:est-main-GM}. We drop these when we perform additional analyses in Table \ref{table:est-supp-GM}.}.
In Table \ref{table:est-main-GM} we reproduce the intention-to-treat estimates from Table 7 of their paper, which presents estimated treatment effects on profits, revenues, employees and household consumption. ``Original'' corresponds to the estimates obtained from running the regression specifications in the original paper which include pair fixed effects, and $\hat{\theta}_n$ corresponds to estimates obtained from running an identical regression specification without pair fixed effects (we note that we were able successfully reproduce all of the reported estimates from the paper). We find an average absolute percentage difference of $13.82\%$\footnote{Here the absolute percentage difference is computed as $\left(\frac{|\text{Original} - \hat{\theta}_n|}{|\text{Original}|}\right)\times 100$.} for the point estimates of these effects, with the largest differences appearing for profits and revenue.
\begin{table}[htbp]
\centering
\caption{Summary of Estimates Obtained from Empirical Application: \cite{groh2016macroinsurance}}
\begin{adjustbox}{max width=\linewidth,center}
\begin{tabular}{ccccccccc}
\toprule
& & High & & High & Number & Any & Owner's & Monthly \\
& Profits & Profit & Revenue & Revenue & Employees & Worker & Hours & Consumption \\
\midrule
Original & -59.702 & -0.009 & -737.199 & -0.020 & -0.024 & 0.008 & -0.655 & -7.551 \\
\addlinespace
$\widehat{\theta}_n$ & -38.642 & -0.007 & -692.818 & -0.020 & -0.023 & 0.007 & -0.773 & -7.337 \\
\addlinespace
Attrition (\%) & 2.086 & 2.086 & 2.153 & 2.153 & 1.783 & 1.783 & 1.480 & 0.000 \\
\bottomrule
\end{tabular}
\end{adjustbox}
\begin{tablenotes} \footnotesize
\item Note: For each outcome listed in Table 7 of \cite{groh2016macroinsurance}, we report (a) the original estimates obtained in paper (``Original''), (b) the estimate on treatment status without pair fixed effects ($\widehat{\theta}_n$), and (c) the attrition rate in \% by outcome, defined as [number of individuals with missing outcome / total number of individuals]. The regression specifications here include baseline covariates; see Table \ref{table:est-supp-GM} for analogous results without baseline covariates included.
\end{tablenotes}
\label{table:est-main-GM}\end{table}
One caveat to the findings in Table \ref{table:est-main-GM} is that the setting does not map exactly into our theoretical results: first, both regressions control for baseline covariates and second, the final assignment contained \emph{one} stratum with 16 individuals, each belonging to a different branch office. Given this, in Table \ref{table:est-supp-GM} we report the intention-to-treat estimates without baseline covariates and without this additional stratum. In this case we find an average absolute percentage difference of $15.61\%$ for the point estimates of the effects. We emphasize that we consider these difference particularly salient given that attrition is quite low (on average $1.4\%$ across the outcomes), and that in the absence of attrition these estimates would be \emph{numerically identical}, as illustrated from the estimates of the effect of treatment for monthly consumption.
\begin{table}[ht!]
\centering
\caption{Summary of Additional Estimates Obtained from Empirical Application: \cite{groh2016macroinsurance}}
\begin{adjustbox}{max width=\linewidth,center}
\begin{tabular}{ccccccccc}
\toprule
& & High & & High & Number & Any & Owner's & Monthly \\
& Profits & Profit & Revenue & Revenue & Employees & Worker & Hours & Consumption \\
\midrule
Original & -91.197 & -0.011 & -967.967 & -0.024 & -0.032 & 0.004 & -0.561 & -3.600 \\
\addlinespace
$\hat{\theta}_n$ & -80.058 & -0.009 & -888.608 & -0.023 & -0.026 & 0.005 & -0.481 & -3.600 \\
\addlinespace
Attrition (\%) & 1.755 & 1.755 & 1.824 & 1.824 & 1.411 & 1.411 & 1.514 & 0.000 \\
\bottomrule
\end{tabular}
\end{adjustbox}
\begin{tablenotes} \footnotesize
\item Note: For each outcome regression specification listed in Table 7 of \cite{groh2016macroinsurance}, we report (a) the original estimates obtained in paper (``Original''), (b) the estimate on treatment status without pair fixed effects ($\hat{\theta}_n$), and (c) the attrition rate in \% by outcome, defined as [number of individuals with missing outcome / total number of individuals]. The regression specifications here exclude baseline covariates from the authors' original work.
\end{tablenotes}
\label{table:est-supp-GM}\end{table}
\subsection{Re-analysis of Recent Publications in AER \& AEJ: Applied} \label{sec:application-aer-aej}
Next, we perform a similar exercise using the data from a systematic survey of all papers published in the American Economic Review (AER) and the American Econonomic Journal: Applied Economics (AEJ: Applied) from 2020-2022 which conducted matched-pair or stratified randomized experiments in the presence of attrition. Our survey identified seven such papers: \cite{abebe2021selection}, \cite{attanasio2020estimating}, \cite{carter2021subsidies}, \cite{casaburi2021using}, \cite{dhar2022reshaping}, \cite{hjort2021research}, and \cite{romero2020outsourcing}. For each paper, we collected a set of ``relevant" regression specifications,\footnote{We note that in some papers such as \cite{attanasio2020estimating} and \cite{casaburi2021using} the primary results were not necessarily the output of a linear regression, and so in these cases we selected a collection of preliminary regression analyses. In other papers such as \cite{hjort2021research}, the primary results were LATE estimates obtained via IV regression, and so in these cases we report the intention to treat analyses. Specific selection details for each paper are outlined in Appendix \ref{sec:addl-info-empirical}.} and reproduced these regressions with and without pair/stratum fixed effects (we note that we were able to successfully reproduce all of the reported estimates from each paper). In Figure \ref{fig:comparison} we report the average absolute percentage change (computed as $\left(\frac{|\text{Alternative} - \text{Original}|}{|\text{Original}|}\right)\times 100$, where ``Original" corresponds to the point estimate computed in the paper, and ``Alternative" corresponds to the estimate computed from the alternative specification with or without fixed effects) across all specifications for each paper. Similar to our findings for \cite{groh2016macroinsurance}, we find that there can be noticeable differences in the point estimates with and without fixed effects (although we emphasize that we do not claim that these differences are necessarily statistically significant).
\begin{figure}[ht!]
\centering
\includegraphics[width=\textwidth]{comparison-studies.jpg}
\caption{Average absolute percentage difference for ``Original" vs ``Alternative" point estimates. Average attrition rate, defined as [number of individuals with missing outcome / total number of individuals] is reported in parentheses below each author label.}
\label{fig:comparison}
\end{figure}
\section{Recommendations for Empirical Practice}\label{sec:recs}
We conclude with some recommendations for empirical practice based on our theoretical results. Our main takeaway is that choosing whether or not to include pair/strata fixed effects when attrition is a concern can make a substantive difference to empirical findings and to the interpretation of the resulting estimand. In our view, unless practitioners are interested in recovering the convex-weighted averages produced by $\theta^{\rm drop}$ and $\theta^{\rm sfe}$ under a conditional independence assumption (Assumption \ref{as:MIPO|T}), primary analyses should be based on regressions \emph{without} pair/strata fixed effects: the resulting estimand $\theta^{\rm obs}$ has a simple interpretation in the absence of any assumptions, and collapses to the average treatment effect under arguably weaker assumptions than $\theta^{\rm drop}$ and $\theta^{\rm sfe}$. A secondary benefit of $\theta^{\rm obs}$ is that, under the additional assumption that $R_i(1) = R_i(0)$, $\theta^{\rm obs}$ \emph{also} enjoys an interpretation as a convex-weighted average under Assumption \ref{as:MIPO|T}, with weights which may be more desirable than those appearing in $\theta^{\rm drop}$ or $\theta^{\rm sfe}$ in that they do not ``double-up" on attrition: see Remark \ref{rem:attrit_units} for details.
\clearpage
\bibliography{attrition}
\pagebreak