Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
171,647 characters · 30 sections · 65 citation commands
Randomization Inference with Sample Attrition
\onehalfspacing
{2pt} {2pt} {1pt} {1pt}
\allowdisplaybreaks
Randomization inference, also known as permutation inference, is widely used across the sciences, including fields such as economics, political science, medicine, and genomics gerber2012field, imbens2015causal, young2019channeling, bates2020causal, zhang2023randomization, chen2024role. Under a sharp null hypothesis of, say, no treatment effect for any unit, the Fisher randomization test (FRT) provides exact finite-sample inference that relies only on the known treatment assignment mechanism and does not need any distributional assumptions or asymptotic approximations fisher1935design, lehmann2006testing. However, its implementation requires complete outcome data. In practice, many experiments suffer from sample attrition—the failure to observe outcomes for some experimental units duflo2007using, gerber2012field, Zhaoetal2024. For example, ghanem2023testing find that more than 80% of experimental studies in leading economics journals report missing outcomes. Moreover, simply discarding units with missing outcomes and applying standard randomization tests to the remaining observed units may not guarantee validity, particularly when dropouts differ systematically from those who remain. This is a well-known issue in the missing data literature little2019statistical.
The prevalence of sample attrition in experiments and the widespread use of randomization inference raise several questions. What assumptions, explicit or implicit, do researchers make about missingness mechanisms? How can the size of a randomization test be controlled when testing a null hypothesis on treatment effects in the presence of attrition? How can the test be tailored to accommodate different assumptions on missingness mechanisms? This paper's contribution is to answer these questions, providing theoretical foundations, methodological tools, and empirical guidance for using randomization inference in experiments with sample attrition.
With sample attrition, the usual sharp null hypothesis---such as the assumption of no treatment effect for every unit---is no longer “sharp”, because the potential outcomes for each unit cannot be fully recovered from the observed data and the null hypothesis. As a result, the classical Fisher randomization test becomes infeasible. To address this, we propose using the supremum of the p-value from the FRT over all possible configurations of the unknown potential outcomes resulting from sample attrition. While this p-value is valid, it is often computationally intractable due to its complex dependence on potential outcomes that are unobserved and not imputable. In particular, both the test statistic and its null distribution may depend on the unknown potential outcomes, and the null distribution often lacks a simple form and requires Monte Carlo approximation. We overcome this challenge by leveraging distribution-free, rank-based test statistics caughey2021randomization, which allow us to transform the p-value computation into a worst-case optimization problem over solely the test statistics---one that admits a closed-form solution. This solution naturally aligns with the sharp bounds in manski1990nonparametric and horowitz2000nonparametric for, e.g., the average treatment effect among all units, where missing outcomes for treated and control units are assumed to lie at the boundary of the outcome support. Moreover, the use of ranks enhances robustness, enabling our method to yield informative inference even when the potential outcomes are unbounded.
We then consider a series of missingness mechanisms, ranging from general missingness to monotone missingness, and ultimately to sharp missingness. The FRT, when applied with test statistics based solely on potential outcomes, is unable to incorporate information from the missingness mechanism. To address this, we propose the use of a composite outcome variable that combines both the original potential outcomes and the potential missingness indicators. Importantly, this formulation incorporates the implications of the missingness assumptions into the test statistics. For example, under the monotone missingness assumption zhang2003estimation, lee2009training, treated units with missing outcomes would also be missing if assigned to control. We show that the FRT based on these composite outcomes can yield more powerful, yet still valid, tests by leveraging information in the specific missingness mechanism.
We further improve the power of the test based on the composite outcomes by exploiting the fact that the distribution of missingness types can be (partially) inferred from the observed data. Specifically, we propose a two-step procedure berger1994p. In the first step, we construct confidence or prediction sets for the missingness types. In the second step, we impute worst-case composite potential outcomes subject to the constraints implied by the first-step confidence or prediction sets and the missingness assumptions. By combining information from the observed missingness patterns and the assumed missingness mechanisms, this two-step procedure can further enhance the power of the test. Moreover, the resulting tests align with the sharp bounds in zhang2003estimation and lee2009training for the average treatment effect among units who would be observed under both treatment and control, under general and monotone missingness assumptions.
Our paper contributes to several strands of literature. First, our work contributes to the literature on statistical inference, particularly randomization-based inference, in the presence of sample attrition. Existing approaches often invoke assumptions such as missing at random or zero treatment effect on missingness Kennes2012Choice, ivanova2022randomization, li2021randomization, heng2023design, heussen2024randomization, Zhaoetal2024. In contrast, our method accommodates informative missingness, in which missingness can depend on potential outcomes, and provides valid randomization tests under various missingness mechanisms.
Second, this paper contributes to the growing literature on randomization tests beyond sharp null hypotheses. While classical randomization tests are exact under sharp nulls, recent research has extended them to test for weak nulls, such as those involving average treatment effects chung2013exact, li2017general, bugni2018inference, wu2021randomization, bai2022inference or quantiles of individual treatment effects caughey2021randomization, su2022treatment, chen2024enhanced. We extend this literature by developing finite-sample valid randomization tests under sample attrition, which renders the classical sharp null hypothesis non-sharp.
Third, it contributes to the literature on the partial identification of treatment effects under sample attrition, where missingness may be informative and may depend on unobserved outcomes. manski1990nonparametric and horowitz2000nonparametric provide sharp bounds on treatment effects under minimal assumptions, although these bounds can be wide. Subsequent studies have sought to tighten these bounds by introducing additional assumptions. For instance, zhang2003estimation and lee2009training impose a monotonicity condition on missingness, and derive bounds on treatment effects among only those units that would be observed under both treatment and control. In contrast, our paper addresses a different but related question: how to conduct randomization inference under sample attrition using a range of missingness assumptions, without imposing any assumptions on the outcome variable? Notably, the worst-case conservative inference procedure developed here directly connects to the bounds in horowitz2000nonparametric, zhang2003estimation, and lee2009training. To our knowledge, this is the first paper to formally and explicitly link robust randomization inference with the framework of partial identification.
Consider a randomized experiment conducted on $n$ units with two treatment arms. We assume that the stable unit treatment value assumption (SUTVA) holds rubin1980discussion, and use the potential outcome framework neyman1923application, rubin1974estimating. Let $Y_{i}(0)$ and $Y_{i}(1)$ denote the potential outcomes under control and treatment, respectively, for unit $i$ in the full sample. We use $\boldsymbol{Y}(0) = (Y_{1}(0), Y_{2}(0), \ldots, Y_{n}(0))^{\intercal}$ and $\boldsymbol{Y}(1) = (Y_{1}(1), Y_{2}(1), \ldots, Y_{n}(1))^{\intercal}$ to denote the control and treatment potential outcome vectors for all units in the full sample. Let $\tau_{i} = Y_{i}(1) - Y_{i}(0)$ be the individual treatment effect for unit $i$, and let $\boldsymbol{\tau} = (\tau_{1}, \tau_{2}, \ldots, \tau_{n})^{\intercal}$ be the vector of individual treatment effects for the full sample. Let $Z_{i}$ be the treatment assignment for unit $i$, where $Z_{i}$ equals 1 if the unit is assigned to treatment and $0$ otherwise. We use $\boldsymbol{Z} = (Z_{1}, Z_{2}, \ldots, Z_{n})^{\intercal}$ to denote the treatment assignment vector for all units. The full sample's outcome $Y_{i}^{\star}$ for each unit $i$ is realized through $Y_{i}^{\star} = (1 - Z_{i}) Y_{i}(0) + Z_{i} Y_{i}(1) $. We use $\boldsymbol{Y}^{\star} = (Y_{1}^{\star}, Y_{2}^{\star}, \ldots, Y_{n}^{\star})^{\intercal}$ to denote the outcome vector for all units in the full sample. Equivalently, the outcome vector is realized through $\boldsymbol{Y}^{\star} = (\boldsymbol{1} - \boldsymbol{Z}) \circ \boldsymbol{Y}(0) + \boldsymbol{Z} \circ \boldsymbol{Y}(1)$, where $\circ$ denotes element-wise multiplication.
We allow sample attrition in the experiment duflo2007using, gerber2012field. Let $M_{i}$ be the missingness indicator for unit $i$. Let $\boldsymbol{M} = (M_{1}, M_{2}, \ldots, M_{n})^{\intercal}$ be the vector of missingness indicators. We define the potential missingness indicators $M_i(0)$ and $M_i(1)$ for each units $i$. We use $\boldsymbol{M}(0) = (M_{1}(0), M_{2}(0), \ldots, M_{n}(0))^{\intercal}$ and $\boldsymbol{M}(1) = (M_{1}(1), M_{2}(1), \ldots, M_{n}(1))^{\intercal}$ to denote the vectors of potential missingness indicators under control and treatment, respectively, for all units. In other words, we assume SUTVA also holds for potential missingness. The observed missingness indicator is realized through $M_{i} = (1 - Z_{i}) M_{i}(0) + Z_{i} M_{i}(1)$, or in vector form, $\boldsymbol{M} = (\boldsymbol{1} - \boldsymbol{Z}) \circ \boldsymbol{M}(0) + \boldsymbol{Z} \circ \boldsymbol{M}(1)$. Note that some of our results hold even without assuming SUTVA for potential missingness; see Remark (ref).
Let $Y_{i}$ denote the observed outcome for unit $i$. Specifically, when $M_{i} = 1$, the realized outcome $Y_{i}^{\star}$ is observed, and we set $Y_{i} = Y_{i}^{\star}$. When $M_{i}=0$, the realized outcome $Y_{i}^{\star}$ is not observed, and we set $Y_i = \texttt{NA}$, where $\texttt{NA}$ indicates “not available”. Define $0 \cdot \texttt{NA} = 0$ and $0 + \texttt{NA} = \texttt{NA}$. Then the observed outcome $Y_{i}$ has the following equivalent forms:
Therefore, for units with missing outcomes (i.e., units with $M_{i} = 0$), none of their potential outcomes are observed. For units without missing outcomes (i.e., units with $M_{i} = 1$), one of their potential outcomes is observed. For descriptive convenience, we further introduce
It should be emphasized that sample attrition in this context can encompass a wide range of scenarios. Missing outcomes may result from self-selection heckman1974shadow, lee2009training, survey nonresponse rubin1987multiple, Little01071988, duflo2007using, or other structural reasons zhang2003estimation, frangakis_principal_2004.
In this paper, we conduct design-based inference (also known as randomization-based or finite population inference), treating all potential outcomes (and potential missingness) as fixed constants\footnote{ In the main paper, we treat all potential missingness indicators as fixed constants. We also consider a missing completely at random mechanism, under which these indicators are random; to avoid confusion, we defer this case to the supplementary material.}—or equivalently, as conditioned on—so that randomness solely arises from the experimental design neyman1923application, fisher1935design, li2017general, abadie2020sampling, abadie2023should, roth2023efficient, zhang2023randomization.
In design-based inference, the distribution of the treatment assignment $\boldsymbol{Z}$, also called the treatment assignment mechanism, is what governs the statistical inference. Throughout the paper, we will focus on the completely randomized experiment (CRE), under which
where $n_1$ and $n_0$ are two fixed positive integers denoting the numbers of treated and control units, respectively, under the CRE. Importantly, the design in (ref) requires that the treatment assignment is independent of both the potential outcomes and the potential missingness indicators.
The inference in this paper applies to randomized experiments with exchangeable treatment assignments caughey2021randomization, including CREs and Bernoulli randomized experiments (BREs). This is because, once we condition on the numbers of treated and control units, the assignment mechanism is equivalent to a CRE, allowing such experiments to be analyzed in the same way.
The missingness mechanism is central to the inference of treatment effects. Random assignment and SUTVA can be ensured by design. However, missingness often arises from self-selection that is neither controlled nor observed by researchers. This may induce dependence between potential missingness and potential outcomes, which we call informative missingness. A canonical case is threshold missingness, where outcomes are observed only if they exceed a given threshold. We give two examples of threshold missingness below.
Because researchers may have limited information about the missingness mechanism, we develop randomization inference procedures for sharp and bounded nulls that remain valid, though potentially conservative, under a general missingness mechanism allowing arbitrary dependence between potential outcomes and missingness. The mechanism is formally defined below.
We also consider additional assumptions on the missingness mechanism. Importantly, these additional assumptions can improve the power of our inference for treatment effects. This and the following subsections present the assumptions to be studied, along with motivating examples.
The next two assumptions require the treatment effect on the missingness indicator to be monotone, either non-positive or non-negative. We state them formally below.
zhang2003estimation and lee2009training assume Assumption (ref). Consider the job training experiment in Example (ref). The monotone missingness assumption holds when the training program weakly increases each individual's propensity to participate in the labor market. Consider the education experiment in Example (ref). The monotone missingness holds when the treatment weakly decreases each student’s propensity to drop out of school.
Note that, under monotone missingness, both Assumptions (ref) and (ref) can be satisfied through label switching and change of the outcome sign, under which the treatment effects remain unchanged. For example, if the original sample satisfies Assumption (ref), then, once we switch the treatment labels and change the outcome signs, Assumption (ref) will hold and the corresponding treatment effects will be the same as those in the original sample. Therefore, when analyzing data, we can proceed with either Assumption (ref) or (ref). However, since switching the treatment labels leads to different test statistics that emphasize slightly different aspects of the treatment effects, we still consider and distinguish between these two monotone missingness mechanisms; see Remark (ref) for related discussion.
An extreme case of Assumptions (ref) and (ref) occurs when $M_i(1) = M_i(0)$ for all $i$. When this condition holds, both assumptions are satisfied, implying that Fisher’s sharp null of no treatment effect holds for the missingness indicator fisher1935design. We refer to this as the sharp missingness assumption.
heng2023design show that Assumption (ref) can still hold even when missingness depends on potential outcomes, provided additionally that (i) Fisher’s null of no treatment effect holds for the outcome and (ii) treatment influences missingness only through the outcome.
A sharp null hypothesis specifies all individual treatment effects fisher1935design. It has the following general form:
where $\boldsymbol{\delta}\in \mathbb{R}^n$ is a prespecified constant vector representing the hypothesized treatment effects for all units. Under the null hypothesis $H_{\boldsymbol{\delta}}$ and in the absence of missing data, we are able to impute the potential outcomes for all individuals from the observed data:
Note that the equalities in (ref) hold if and only if the sharp null hypothesis $H_{\boldsymbol{\delta}}$ is true.
To test the null in (ref), consider a general test statistic of the form $t(\cdot, \cdot): \mathcal{Z} \times \mathbb{R}^{n} \xrightarrow[]{} \mathbb{R}$, which is a generic function with two arguments: a treatment assignment vector $\boldsymbol{z} \in \mathcal{Z}$ and an outcome vector $\boldsymbol{y} \in \mathbb{R}^{n}$. Following rosenbaum2002design, we use test statistics $t(\boldsymbol{Z}, \boldsymbol{Y}(0))$ that depend on the treatment assignment $\boldsymbol{Z}$ and the control potential outcomes in (ref) imputed under the sharp null (ref). The tail probability of the randomization distribution of the test statistic $t(\boldsymbol{Z}, \boldsymbol{Y}(0))$ under the sharp null is
where $\boldsymbol{A}$ is a generic random treatment assignment vector that follows the same distribution as the observed assignment vector $\boldsymbol{Z}$. The corresponding randomization p-value under the sharp null is
which is a valid p-value in the sense that $\mathbb{P}(p_{\boldsymbol{Z}} \le \alpha) \le \alpha$ for $\alpha \in (0,1)$.
When sample attrition occurs, the p-value in (ref) cannot be computed. In particular, missing outcomes imply that the control potential outcomes of the missing units cannot be imputed under the sharp null. As a result, both the realized test statistic $t(\boldsymbol{Z}, \boldsymbol{Y}(0))$ and its true randomization distribution (ref) are unknown.
Consequently, valid randomization-based inference in the presence of attrition requires procedures that maintain size control across all possible configurations of missing outcomes. This can be challenging, as attrition affects both the test statistic and its randomization distribution, the latter having no simple form and often requiring Monte Carlo approximation. We therefore employ the distribution-free, rank-based test statistics proposed by caughey2021randomization, specifically the two classes described below.
The first is a class of rank-sum statistics of the following form:
where $\boldsymbol{z}\in \{0,1\}^n$ is the treatment assignment vector, $\boldsymbol{y} = (y_1, \ldots, y_n)^\intercal \in \mathbb{R}^n$ is the outcome vector, $\textup{rank}_i(\boldsymbol{y})$ denotes the rank of the $i$-th coordinate of $\boldsymbol{y}$, $\phi(\cdot)$ is a prespecified nondecreasing rank transformation, and $\psi_{i,j}(\cdot, \cdot)$ is an indicator function for pairwise comparison, which will be defined shortly. Setting $\phi(r) = r$ reduces (ref) to the Wilcoxon rank-sum statistic wilcoxon1945individual. Alternatively, with $\phi(r) = \binom{r-1}{s-1} \mathbbm{1}(r \ge s)$ for a fixed integer $s > 1$, it becomes the Stephenson rank-sum statistic robert1985two. To break ties, we assume units are randomly ordered and use their indices as a tie-breaker. For example, $\textup{rank}_i(\boldsymbol{y}) < \textup{rank}_j(\boldsymbol{y})$ if and only if (i) $y_i < y_j$ or (ii) $y_i = y_j$ and $i<j$. Under this rule, the indicator function $\psi_{i,j}(\cdot, \cdot)$ in (ref) is defined for $1\le i, j\le n$ and $y, y'\in \mathbb{R}$ as:
Thus, $\textup{rank}_i(\boldsymbol{y}) = \sum_{j=1}^n \psi_{i,j}(y_i, y_j)$ for all $i$, justifying the second equality in (ref).
The second is a class of generalized Mann-Whitney-type U-statistics:
and
where $\boldsymbol{z}$, $\boldsymbol{y}$, $\phi(\cdot)$, and $\psi_{i, j}(\cdot, \cdot)$ are defined analogously as in (ref) and (ref). We consider choosing $\phi$ as $\phi(r) = r^{s-1}$, where $s$ is a parameter governing the weights assigned to responses with different ranks. When $s = 2$, both (ref) and (ref) reduce to the usual Mann-Whitney U-statistic. When $s > 2$, larger responses among the treated (or control) group receive more weight, making the statistic more sensitive to extreme values; this is analogous to rank transformations for Stephenson rank-sum statistics of the form (ref).
We briefly compare the rank-based statistics in (ref), (ref) and (ref). When $\phi$ is the identity, they are equivalent up to a constant shift, reflecting the well-known equivalence between the Wilcoxon rank-sum and Mann–Whitney U statistics. Both are distribution-free and therefore suitable for our proposed randomization inference procedures. Due to these similarities, we use the same notation for the two, specifying the form as needed throughout the paper. However, for non-identity $\phi$, the statistics differ: (ref) is based on full-sample ranks, whereas (ref) ranks treated units relative to controls and (ref) ranks control units relative to treated units. This distinction can be practically important: statistics of the form (ref) or (ref) may be easier to optimize when computing valid p-values, especially under monotone missingness.
Importantly, under the CRE and with random tie-breaking, the distribution of $t_{\textup{R}, \phi}(\boldsymbol{Z}, \boldsymbol{y})$ in either (ref), (ref) or (ref) does not depend on the value of $ \boldsymbol{y}$. Specifically, $t_{\textup{R}, \phi}(\boldsymbol{Z}, \boldsymbol{y}) \sim t_{\textup{R}, \phi}(\boldsymbol{Z}, \boldsymbol{y}')$ for any constant vectors $ \boldsymbol{y}, \boldsymbol{y}' \in \mathbb{R}^n$, where $\sim$ denotes equal in distribution caughey2021randomization. Thus, we can define
In other words, the randomization distribution $G_{\textup{R}, \phi}(c)$ is determined by the rank transformation $\phi(\cdot)$ and the numbers of treated and control units, $n_1$ and $n_0$.
Throughout the paper, we first focus on the sharp null hypothesis in (ref), and later extend to bounded null hypotheses in Section (ref). We now study randomization tests for the sharp null under the general missingness mechanism of Assumption (ref), which allows arbitrary dependence between potential missingness and potential outcomes (i.e., arbitrarily informative missingness).
Under sample attrition, the standard randomization p-value for the sharp null in (ref) is generally not computable because both the test statistic and its randomization distribution are unknown (see Section (ref)). To address this challenge, we use the rank statistic $t_{\textup{R}, \phi}(\boldsymbol{Z}, \boldsymbol{Y}(0))$ introduced in Section (ref), which depends only on control potential outcomes. Due to the distribution-free property of rank statistics, the randomization distribution in (ref) does not depend on $\boldsymbol{Y}(0)$ and is therefore known exactly. Consequently, the problem of obtaining a valid p-value that maintains size control over all possible configurations of missing outcomes reduces to minimizing the realized test statistic $t_{\textup{R}, \phi}(\boldsymbol{Z}, \boldsymbol{Y}(0))$ subject to the constraints imposed by the observed data and the sharp null hypothesis. In other words, this valid p-value is a worst-case p-value.
This strategy will be used throughout the paper. In later sections, we carefully choose test statistics that incorporate information about potential missingness. This allows us to exploit constraints imposed by the assumptions on the missingness mechanisms and by the distribution of missingness types implied by the observed data, thereby increasing the power of the resulting tests.
Panel A in Table (ref) summarizes the available information on potential outcomes and missingness indicators, based on the observed data and the sharp null hypothesis in (ref). From the table, we do not know the value of $Y_{i}(0)$ for units whose outcomes are missing. Intuitively, $t_{\textup{R}, \phi}(\boldsymbol{Z}, \boldsymbol{Y}(0))$ achieves its minimum value when the potential outcomes $Y_i(0)$ for treated units are as small as possible, while those for control units are as large as possible. This motivates us to consider the following control potential outcome vector: $\boldsymbol{Y}_{\boldsymbol{Z}, \boldsymbol{\delta}}^{\texttt{g}}(0) = (Y_{\boldsymbol{Z}, \boldsymbol{\delta}, 1}^{\texttt{g}}(0), \ldots, Y_{\boldsymbol{Z}, \boldsymbol{\delta}, n}^{\texttt{g}}(0))^\intercal,$ where the superscript $\texttt{g}$ indicates the general missingness mechanism, the subscript $\boldsymbol{\delta}$ represents the sharp null of interest in (ref), and the subscript $\boldsymbol{Z}$ emphasizes its dependence on the observed treatment assignment, and
Obviously, $\boldsymbol{Y}_{\boldsymbol{Z}, \boldsymbol{\delta}}^{\texttt{g}}(0)$ is a “feasible” imputed value of $\boldsymbol{Y}(0)$ based on the observed data and the sharp null hypothesis. The theorem below shows that $\boldsymbol{Y}_{\boldsymbol{Z}, \boldsymbol{\delta}}^{\texttt{g}}(0)$ indeed yields the “worst-case” configuration of the control potential outcomes, leading to a valid p-value for the sharp null hypothesis $H_{\boldsymbol{\delta}}$.
Our worst-case inferential procedure in Theorem (ref) has a similar flavor with the worst-case identification approach in the partial identification literature robins1989probability, robins1989estimability, manski1990nonparametric, horowitz2000nonparametric. This approach assumes that (i) the missing values in the treated group are at the lower bound of the support of the outcome and (ii) the missing values in the control group are at the upper bound of the support of the outcome. This can then give us a partially identified set for the parameter of interest, such as the average treatment effect. Although we are studying a conceptually different problem, namely, inference for sharp null hypothesis, the worst-case scenario for the p-value nevertheless coincides with the worst-case identification problem. Moreover, our approach has the advantage that it can still give informative inference results even when the support of the outcome is unbounded.
Note that imposing assumptions on potential missingness (such as Assumptions (ref), (ref) or (ref)) or incorporating information about its distribution, which can be learned as least partially from the observed data, does not change the worst-case imputation of the control potential outcomes. Consequently, they do not change the worst-case test statistic $t_{\textup{R}, \phi}(\boldsymbol{Z}, \boldsymbol{Y}_{\boldsymbol{Z}, \boldsymbol{\delta}}^{\texttt{g}}(0))$ and the resulting worst-case p-value. In the next section, we will carefully choose a test statistic that leverages these assumptions to refine worst-case imputation and improve statistical power.
To use the missingness assumptions in the randomization inference procedure, we modify the test statistic by incorporating potential missingness. Specifically, we introduce the composite control potential outcome for each unit $i$: $$ \tilde{Y}_{\boldsymbol{b},i}(0) \coloneqq Y_i(0) M_i(0) M_i(1) + b_{00} (1 - M_i(0)) (1 - M_i(1)) + b_{01} (1 - M_i(0)) M_i(1) + b_{10} M_i(0) (1 - M(1)),$$ where $\boldsymbol{b} = (b_{00}, b_{01}, b_{10})$ and $b_{00}, b_{01}, b_{10} \in\mathbb{R}$. Equivalently, in vector form, $\tilde{\boldsymbol{Y}}_{\boldsymbol{b}}(0) \coloneqq \boldsymbol{Y}(0) \circ \boldsymbol{M}(0) \circ \boldsymbol{M}(1) + b_{00} \cdot (\boldsymbol{1}-\boldsymbol{M}(0)) \circ (\boldsymbol{1}-\boldsymbol{M}(1)) + b_{01} \cdot (\boldsymbol{1}-\boldsymbol{M}(0)) \circ \boldsymbol{M}(1) + b_{10} \cdot \boldsymbol{M}(0) \circ (\boldsymbol{1}-\boldsymbol{M}(1))$. We use $t(\boldsymbol{Z}, \tilde{\boldsymbol{Y}}_{\boldsymbol{b}}(0))$, which we refer to as the composite test statistic, for randomization inference.
To see the benefit of incorporating potential missingness on inference, we compare the test statistics $t(\boldsymbol{Z}, \boldsymbol{Y}(0))$ and $t(\boldsymbol{Z}, \tilde{\boldsymbol{Y}}_{\boldsymbol{b}}(0))$. Both statistics contrast the treated and control groups, using either the original or the composite control potential outcomes. The composite potential outcome coincides with the original potential outcome only for units that would be observed under both treatment and control, and is set to some constant otherwise. Consequently, the statistic $t(\boldsymbol{Z}, \tilde{\boldsymbol{Y}}_{\boldsymbol{b}}(0))$ essentially focuses solely on units with $M_i(1) = M_i(0) = 1$; such a subgroup of units is known as a principal stratum frangakis_principal_2004. This focus is natural. For units outside this principal stratum, at least one potential outcome is unobservable regardless of treatment assignment and may take any value on the real line; such units therefore cannot provide evidence against the null hypothesis. The composite test statistic restricts attention to the principal stratum $\{i: M_i(1) = M_i(0) = 1\}$ rather than the full sample, and, as shown later, can yield sharper and more powerful tests of the null hypothesis.
We now consider testing the sharp null $H_{\boldsymbol{\delta}}$ in (ref) using the test statistic $t(\boldsymbol{Z}, \tilde{\boldsymbol{Y}}_{\boldsymbol{b}}(0))$ under the general missingness mechanism. Similar to Section (ref), we use a distribution-free statistic in (ref), (ref) or (ref), and we seek to minimize the realized value of the test statistic $t_{\textup{R}, \phi}(\boldsymbol{Z}, \tilde{\boldsymbol{Y}}_{\boldsymbol{b}}(0))$.
Panel A in Table (ref) summarizes the composite potential outcomes under the sharp null in (ref), which depend on treatment assignment, the realized outcome, and counterfactual missingness. Specifically, for treated units with observed outcomes, the composite potential outcome equals $Y_i - \delta$ if the unit would remain observed under control and $b_{01}$ if it would instead be missing under control; for treated units with missing outcomes, it equals $b_{00}$ if the unit would remain missing under control and $b_{10}$ if it would instead be observed under control. The corresponding expressions for control units’ composite potential outcomes can be obtained by replacing $Y_i - \delta$ with $Y_i$ and evaluating missingness under treatment.
As in Theorem (ref), we want $\tilde{Y}_{\boldsymbol{b},i}(0)$ to be as small as possible for treated units and as large as possible for control units in order to obtain a valid randomization p-value under the general missingness mechanism. This motivates us to consider the following “feasible” values for the composite potential outcome vector: $\tilde{\boldsymbol{Y}}_{\boldsymbol{Z},\boldsymbol{\delta}, \boldsymbol{b}}^{\texttt{g}}(0) = (\tilde{Y}_{\boldsymbol{Z},\boldsymbol{\delta}, \boldsymbol{b}, 1}^{\texttt{g}}(0), \ldots, \tilde{Y}_{\boldsymbol{Z},\boldsymbol{\delta}, \boldsymbol{b}, n}^{\texttt{g}}(0))^\intercal$ with
The theorem below shows that $ \tilde{\boldsymbol{Y}}_{\boldsymbol{Z},\boldsymbol{\delta}, \boldsymbol{b}}^{\texttt{g}}(0)$ indeed yields the worst-case configuration of the composite potential outcomes, leading to a valid p-value for testing the sharp null $H_{\boldsymbol{\delta}}$ in (ref).
The p-value in Theorem (ref) is valid for any pre-specified constants $b_{00}$, $b_{01}$, and $b_{10}$ used to construct the composite potential outcomes. We now discuss how the choice of these constants affects the power of the test in Theorem (ref). Given the form of the rank-based test statistic, the p-value in (ref) depends crucially on the relative magnitude of the worst-case configuration $\tilde{\boldsymbol{Y}}_{\boldsymbol{Z},\boldsymbol{\delta}, \boldsymbol{b}}^{\texttt{g}}(0)$ for treated and control units. Regardless of the values of $b_{00}$, $b_{01}$, and $b_{10}$, the composite potential outcomes for the observed treated units are always less than or equal to those for the missing control units, while those for the missing treated units are always less than or equal to those for all the control units. As $b_{01}$ increases and $b_{10}$ decreases, the composite potential outcomes for observed treated units are more likely to exceed those for observed control units. Therefore, to “maximize” the test statistic under the worst-case configuration, or equivalently to “minimize” the p-value, we suggest choosing $b_{01}$ as large as possible, i.e., $b_{01} = \infty$, and choosing $b_{10}$ as small as possible, i.e., $b_{10} = -\infty$. For descriptive convenience, we define $0 \cdot \infty = 0$ and $0 \cdot -\infty = 0$.
Panel A of Table (ref) summarizes the worst-case configurations of both the control potential outcomes and the composite potential outcomes for any fixed $b_{00}, b_{01}, b_{10} \in \mathbb{R}$, including the suggested choices $b_{01} = \infty$ and $b_{10} = -\infty$. When using these suggested values of $b_{01}$ and $b_{10}$, the p-value in (ref) coincides with that in (ref), because $\tilde{\boldsymbol{Y}}_{\boldsymbol{Z},\boldsymbol{\delta}, \boldsymbol{b}}^{\texttt{g}}(0)$ reduces to $\boldsymbol{Y}_{\boldsymbol{Z}, \boldsymbol{\delta}}^{\texttt{g}}(0)$. Accordingly, Theorem (ref) is a special case of Theorem (ref) with these suggested values of $b_{01}$ and $b_{10}$, and both yield the same valid p-value for testing the sharp null hypothesis. Although the test statistic $t(\boldsymbol{Z}, \tilde{\boldsymbol{Y}}_{\boldsymbol{b}}(0))$ offers no advantage over $t(\boldsymbol{Z}, \boldsymbol{Y}(0))$ under general missingness, it can be beneficial under specific missingness mechanisms, as discussed in the following subsection. Moreover, it can also offer advantages under general missingness when combined with a two-step approach described in Section (ref).
We test the sharp null in (ref) under missingness mechanisms in which treatment has a monotone effect on the missingness indicator. We consider two cases corresponding to non-negative and non-positive treatment effects, as formalized in Assumptions (ref) and (ref), respectively. As discussed at the end of Section (ref), under either Assumption (ref) or (ref), the worst-case configuration of $\boldsymbol{Y}(0)$ remains the same as in (ref), and the resulting p-value is the same as that in Theorem (ref). Thus, the monotone missingness assumptions do not improve the power of tests based on $t_{\textup{R}, \phi}(\boldsymbol{Z}, \boldsymbol{Y}(0))$.
We now focus on the test statistic $t_{\textup{R}, \phi}(\boldsymbol{Z}, \tilde{\boldsymbol{Y}}_{\boldsymbol{b}}(0))$. Panels B and C of Table (ref) summarize the available information on the composite potential outcomes implied by the observed data, the sharp null $H_{\boldsymbol{\delta}}$, and Assumption (ref) or (ref). Compared with Panel A under Assumption (ref), the monotone missingness assumptions allow us to impute additional $M_i(0)$ and $M_i(1)$ values. Specifically, Assumption (ref) implies $M_i(0)=0$ for missing treated units because $M_i(0)\le M_i(1)=M_i=0$. It also implies $M_i(1)=1$ for observed control units because $M_i(1)\ge M_i(0)=M_i=1$. Similarly, under Assumption (ref), $M_i(0)=1$ for observed treated units and $M_i(1)=0$ for missing control units. The full composite potential outcomes remain unknown, so we consider their worst-case configuration to ensure the validity of the randomization test. As in Section (ref), we minimize treated units' composite outcomes and maximize those of control units.
We therefore consider the following “feasible” configurations of the composite potential outcomes under Assumptions (ref) and (ref), as summarized in Panels B and C of Table (ref). Under Assumption (ref), we define $\tilde{\boldsymbol{Y}}^{\texttt{mp}}_{\boldsymbol{Z}, \boldsymbol{\delta}, \boldsymbol{b}}(0) = (\tilde{Y}^{\texttt{mp}}_{\boldsymbol{Z}, \boldsymbol{\delta}, \boldsymbol{b}, 1}(0), \ldots, \tilde{Y}^{\texttt{mp}}_{\boldsymbol{Z}, \boldsymbol{\delta}, \boldsymbol{b}, n}(0))^\intercal$, with the superscript indicating that the treatment has a monotone positive (more precisely, nonnegative) effect on the missingness indicator:
Under Assumption (ref), we define $\tilde{\boldsymbol{Y}}^{\texttt{mn}}_{\boldsymbol{Z}, \boldsymbol{\delta},\boldsymbol{b}}(0) = (\tilde{Y}^{\texttt{mn}}_{\boldsymbol{Z}, \boldsymbol{\delta},\boldsymbol{b}, 1}(0), \tilde{Y}^{\texttt{mn}}_{\boldsymbol{Z}, \boldsymbol{\delta},\boldsymbol{b}, 2}(0), \ldots, \tilde{Y}^{\texttt{mn}}_{\boldsymbol{Z}, \boldsymbol{\delta},\boldsymbol{b}, n}(0))^\intercal$, with the superscript indicating that the treatment has a monotone negative (more precisely, nonpositive) effect on the missingness indicator:
The following theorem shows that (ref) and (ref) indeed give the worst-case configurations under Assumptions (ref) and (ref), respectively, and thereby lead to valid p-values for testing the sharp null $H_{\boldsymbol{\delta}}$.
We now discuss the choice of $b_{00}$, $b_{01}$, and $b_{10}$, which are important for the power of the test. Consider first the worst-case configuration of the composite potential outcomes in (ref) under Assumption (ref). In this case, the composite outcomes of control units with observed outcomes do not depend on $b_{00}$, $b_{01}$, or $b_{10}$. Moreover, regardless of the values of $b_{00}$, $b_{01}$, and $b_{10}$, the composite potential outcomes of treated units are always less than or equal to those of control units with missing outcomes (which are $\max\{b_{00}, b_{01}\}$). As $b_{01}$ increases, observed treated units are more likely to exceed observed control units. Similarly, as $b_{00}$ increases, missing treated units are also more likely to exceed observed control units. Therefore, we recommend choosing $b_{00}$ and $b_{01}$ as large as possible for Theorem (ref) with $\square = \text{p}$: that is, using the p-value $\tilde{p}_{\boldsymbol{Z}, \boldsymbol{\delta},\boldsymbol{b}}^{\texttt{mp}}$ in (ref) with $b_{00} = b_{01} = \infty$.
We next consider Assumption (ref), under which the worst-case configuration of the composite potential outcomes is given in (ref). In this case, the composite potential outcomes of missing treated units are always bounded above by those of the control units. As $b_{00}$ and $b_{10}$ decrease, the composite potential outcomes of observed treated units are more likely to exceed those of the control units. Therefore, we recommend choosing $b_{00}$ and $b_{10}$ as small as possible in Theorem (ref) with $\square = \text{n}$; that is, using the p-value $\tilde{p}_{\boldsymbol{Z}, \boldsymbol{\delta},\boldsymbol{b}}^{\texttt{mn}}$ in (ref) with $b_{00} = b_{10} = -\infty$.
With the suggested $\boldsymbol{b}$ values, the worst-case configuration under Assumption (ref) or (ref) matches that under general missingness, except that $-\infty$ values imputed for missing treated units become $\infty$, or $\infty$ values imputed for missing control units become $-\infty$. Because of the effect increasing property of the rank-sum statistics in (ref), (ref) and (ref), the p-value in (ref) from Theorem (ref) under either Assumption (ref) or (ref) is smaller than or equal to $\tilde{p}_{\boldsymbol{Z}, \boldsymbol{\delta}, \boldsymbol{b}}^\texttt{g}$ under the general missing mechanism assumption. This holds more generally for any prespecified values of $(b_{00}, b_{01}, b_{10})$, and can be seen from the last two columns of Table (ref). In other words, by using the monotonicity assumptions in the missingness mechanism, Theorem (ref), which is based on the composite potential outcomes, yields more powerful tests than Theorem (ref).
Testing (ref) under the sharp missingness mechanism proceeds as in Sections (ref), minimizing the original or composite potential outcomes of treated units and maximizing those of control units. For the test statistic $t_{\textup{R}, \phi}(\boldsymbol{Z}, \boldsymbol{Y}(0))$, the worst-case configuration of $Y_i(0)$ again coincides with Panel A of Table (ref), yielding the same p-value $p_{\boldsymbol{Z}, \boldsymbol{\delta}}^\texttt{g}$ as in (ref) under general missingness mechanisms. In contrast, for $t_{\textup{R}, \phi}(\boldsymbol{Z}, \tilde{\boldsymbol{Y}}_{\boldsymbol{b}}(0))$, the composite potential outcomes are fully known: for all $1\le i \le n$, $\tilde{Y}_{\boldsymbol{b},i}(0) = Y_i(0) M_i(0) M_{i}(1) + b_{00} (1-M_i(0)) (1-M_i(1)) = (Y_i -\delta_i Z_i) M_i + b_{00} (1-M_i)$. The corresponding valid p-value is given in the following proposition.
Due to the effect increasing property of the test statistics in (ref), (ref) and (ref), the p-value in Proposition (ref) is smaller than or equal to that in Theorems (ref) and (ref) under the monotone missingness. Thus, using composite potential outcomes under sharp missingness yields more powerful tests.
We now discuss the choice of $b_{00}$ in Proposition (ref), which sets the composite potential outcomes for units with missing outcomes. Under the CRE and Assumption (ref), the numbers of missing outcomes are balanced on average between treatment and control groups, so $b_{00}$ may have little effect on test performance. Below we describe a strategy that avoids specifying $b_{00}$ and is often preferable in practice. Specifically, we propose excluding units with missing outcomes when conducting randomization tests under the sharp missingness mechanism. Specifically, we restrict attention to the subset $\mathcal{S}=\{ i:M_i=1, 1 \leq i \leq n \}$ of units with observed outcomes and view the resulting experiment as a CRE that assigns $n_{\mathcal{S}1}=\sum_{i\in\mathcal{S}} Z_i$ units to treatment and $n_{\mathcal{S}0}=|\mathcal{S}|-n_{\mathcal{S}1}$ to control. Standard randomization tests for sharp null hypotheses then apply, because no outcomes are missing. This procedure is valid because (i) under Assumption (ref), $\mathcal{S}=\{i:M_i(1)=M_i(0)=1, 1 \leq i \leq n\}$ is nonrandom, and (ii) the test is conditional on the treatment assignments of units in $\mathcal{S}^{\complement}=\{i:M_i(1)=M_i(0)=0 \}$. Consequently, concerns about missing outcomes and the choice of $b_{00}$ are avoided. We summarize the results in the following theorem.
Theorem (ref) essentially tests a “modified” sharp null hypothesis that is restricted to the subset of units whose outcomes are always observed under either treatment arm:
Moreover, because outcomes are fully observed within $\mathcal{S}$, the procedure in Theorem (ref) applies to generic test statistics, such as the difference-in-means statistic.
The power of the tests in Section (ref) can be further improved by exploiting information about potential missingness. The key intuition is that we can use the observed missingness for treated units to infer the counterfactual missingness for control units, and vice versa. Below we present a two-step procedure to improve the power of the tests in Section (ref). As detailed below, we first construct prediction sets for the distribution of potential missingness in the treated or control groups, and then use that to improve the lower bound of the worst-case test statistic values. The resulting comparisons between treated and control units in the improved p-values can mirror the corresponding sharp bounds in the partial identification literature, which in some sense implies the sharpness of our tests. In the following, we consider first the monotone missingness under Assumption (ref) or (ref), and then the general missingness under Assumption (ref).
Under Assumption (ref) and with the suggested $b_{00}=b_{01} = \infty$ from Section (ref), the composite potential outcome for each unit $i$ reduces to $\tilde{Y}_{\boldsymbol{b},i}(0) = Y_i(0) M_i(0) + \infty (1 - M_i(0))$. Under the sharp null in (ref), with information from the observed data and as summarized in Table (ref) Panel B, the composite potential outcomes are known except for observed treated units. Specifically, for an observed treated unit $i$, $\tilde{Y}_{\boldsymbol{b},i}(0)$ equals $Y_i - \delta_i$ if the unit would have been observed under control (i.e., $M_i(0)=1$), and equals $\infty$ if the unit would have been missing under control (i.e., $M_i(0)=0$). In the worst-case consideration as in (ref) and as summarized in the last column of Panel B in table (ref), we essentially pretend that all observed treated units would also be observed under control. This, however, may be overly conservative, and thus compromises the power of the test in Theorem (ref).
Indeed, the observed missingness in the control group can inform the number of observed treated units who would have been missing under control. Specifically, the number of observed control units, $\sum_{i=1}^n (1-Z_i) M_i = \sum_{i=1}^n (1-Z_i) M_i(0)$, follows a Hypergeometric distribution $\text{HG}(N, \sum_{i=1}^n M_i(0), n_0)$\footnote{$\text{HG}(N, m, n)$ denotes the distribution of the number of successes in a sample of size $n$, drawn without replacement from a finite population of size $N$ that contains exactly $m$ successes.} This can lead to an upper confidence bound for $\sum_{i=1}^n M_i(0)$ and consequently a lower prediction bound for the number of observed treated units who would have been missing under control. Importantly, this implies that at least a certain number of composite potential outcomes for the observed treated units must be $\infty$, rather than all being the imputed control potential outcomes, as in the last column of Panel B in Table (ref). Intuitively, this would increase the worst-case value of the test statistic and decrease the p-value, although we still need to take into account the error that may arise in the first step when inferring $\sum_{i=1}^n M_i(0)$.
Accordingly, we propose a two-step testing procedure. First, we construct an upper confidence bound for $\sum_{i=1}^n M_i(0)$ using, for example, the optimal method in Wang:2015, which also implies a lower prediction limit for the number of observed treated units with $M_i(0)=0$. Second, we impose this limit as a constraint when imputing the worst-case configuration of composite control potential outcomes for the observed treated units. We apply a Bonferroni-type correction to control the errors that may occur in both steps. Moreover, to simplify the constrained optimization in the second step, we consider only the Mann--Whitney-type U-statistic in (ref), which sums the relative ranks of each treated unit compared with all control units; in particular, it admits a closed-form solution. We summarize this two-step procedure in Algorithm (ref). Recall the definition in (ref).
The worst-case imputation in (ref) under the two-step procedure is closely connected to the Zhang--Rubin--Lee bounds zhang2003estimation, lee2009training. Note that when the sample size is large, due to randomization, the proportion of units with $M_i(0) = 1$ would be about the same in the treated and control groups, implying that $(n_{11} - \underline{m})/n_1$ is approximately $n_{01}/n_0$, where $n_{11}$ and $n_{01}$ are defined in (ref). This implies that the proportion of $\infty$ among treated units in (ref) is roughly $1 - (n_{11} - \underline{m})/n_1 \approx 1 - n_{01}/n_0$, the same as the proportion of $\infty$ among control units. Thus, in the randomization test and ignoring ties, we are essentially comparing the $(n_{11} - \underline{m})/n_{11} \approx (n_{01} / n_0) / (n_{11} / n_1)$ fraction of the observed treated units with the smallest imputed control potential outcomes to all the observed control units. Such a comparison exactly mirrors that used in the Zhang--Rubin--Lee lower bound on the average treatment effect within the principal stratum of units whose outcomes would be observed under both treatment and control; see zhang2003estimation and lee2009training.
Under Assumption (ref) and with the suggested $b_{00}=b_{10} = -\infty$ from Section (ref), the composite potential outcome for each unit $i$ reduces to $\tilde{Y}_{\boldsymbol{b},i}(0) = Y_i(0) M_i(1) - \infty (1 - M_i(1))$. Under the sharp null in (ref), with information from the observed data and as summarized in Table (ref) Panel C, the composite potential outcomes are known except for observed control units, for which they are either the observed outcome $Y_i$ or $-\infty$ depending on their counterfactual missingness had they been assigned to treatment. The worst-case imputation in Theorem (ref) with the suggested $\boldsymbol{b}$ value, as summarized in the last column in Table (ref) Panel C, essentially pretend that all the observed control units would also be observed had they been assigned to treatment, which may be overly conservative. Therefore, by the same logic as Section (ref), we can use a two-step method to further improve the power of the test in Theorem (ref). Moreover, to facilitate optimization for the worst-case composite potential outcomes, we consider only test statistics of form (ref), which admit closed-form solutions. We summarize it in the following algorithm.\footnote{We slightly abuse notation: the same symbols may represent different quantities across algorithms.}
In Algorithm (ref), we can construct the confidence limit $\hat{M}$ using the method in Wang:2015 and the fact that $\sum_{i=1}^n Z_i M_i = \sum_{i=1}^n Z_i M_i(1)$ follows $\text{HG}(N, \sum_{i=1}^n M_i(1), n_1)$.
Analogous to the discussion after Theorem (ref), in the worst-case imputation in (ref) and ignoring ties, we are essentially comparing all the observed treated units to approximately $(n_{11}/n_1)/(n_{01}/n_0)$ fraction of observed control units with the largest observed outcomes. This again mirrors the comparison in the Zhang–Rubin–Lee lower bound on the average treatment effect within the principal stratum of units who can be observed under both treatment and control.
We now consider the general missingness. In the worst-case imputation in (ref) with the suggested $b_{01} = \infty$ and $b_{10} = -\infty$ and as summarized in the last column of Panel A in Table (ref), we pretend that all treated units would have been observed under control, and all control units would have been observed under treatment. Similar to the monotone missingness studied in the previous two subsections, we can also use the observed missingness to infer counterfactual missingness, and then use a two-step approach improve the worst-case consideration in the randomization test in Theorems (ref) and (ref). However, both steps become more challenging under the general missingness. First, without the monotonicity in Assumption (ref) or (ref), the distribution of $(M_i(1), M_i(0))$ is generally not identifiable from the observed data, implying that we need a more refined design of the first step. Second, due to the non-identifiability of the joint distribution of the potential missingness, we need to consider a larger search space for the worst-case configuration, implying a more challenging optimization in the second step.
To address these challenges, in the first step, we construct a two-sided prediction interval for the number of treated units who would have been missing under control, together with an upper prediction bound on the difference between the treated and control groups in the proportions of units with $M_i(1)=M_i(0)=1$ (i.e., units that would be observed under both treatment and control). As illustrated shortly, these prediction bounds are, in a certain sense, sufficient to ensure the sharpness of our test; in particular, as discussed later, the resulting comparison under the worst-case scenario mirrors that in the sharp identification bound for the average treatment effect within the principle stratum of units that would be observed under both treatment and control.
In the second step, to ease the optimization, we set $b_{00} = \infty$ and use the test statistic of form (ref).\footnote{By a similar logic, we can also set $b_{00} = -\infty$ and use test statistics of form (ref); we omit the details for conciseness.} Moreover, to facilitate the optimization, we derive a closed-form but conservative solution for the minimum value of the test statistic under the constraints imposed by the observed data, the null hypothesis of interest, the prediction bounds on the distribution of potential missingness from the first step, and the number of treated units with $M_i(1)=M_i(0)=1$, with the last quantity requiring further enumeration. The conservativeness primarily arises from ties at $-\infty$. In particular, the solution is exact under certain orderings used to break ties, for example, when all missing treated units have smaller indices than all observed control units.
We now present the algorithm for the two-step testing procedure under the general missingness. The construction of prediction bounds in the first step is discussed immediately after the algorithm. Let $m^{\texttt{t}}_{11} = |\{i: Z_i=1, M_i(0)=M_i(1) = 1\}|$, $m^{\texttt{t}}_{10} = |\{i: Z_i=1, M_i(0) = 1, M_i(1)=0\}|$, and $m^{\texttt{c}}_{11} = |\{i: Z_i=0, M_i(0) = M_i(1) = 1\}|$.
One way to construct the prediction bounds $(\underline{m}, \overline{m}, \overline{d})$ in step (i) of Algorithm (ref) is as follows. Choose $\beta_1, \beta_2 \ge 0$ such that $\beta_1 + \beta_2 = \beta$. Let $[\hat{M}_1, \hat{M}_2]$ be a $1-\beta_1$ two-sided confidence interval for $\sum_{i=1}^n M_i(0)$ using the fact that $n_{01} \sim \text{HG}(N, \sum_{i=1}^n M_i(0), n_0)$ and the method in Wang:2015. Then set $\underline{m} = \hat{M}_1 - n_{01}$ and $\overline{m} = \hat{M}_2 - n_{01}$. Let $q_{\text{HG}}(p; n, m_{11}, n_0)$ denote the $p$ quantile of the Hypergeometric distribution with parameters $(n, m_{11}, n_0)$, and set $ \overline{d} = \max_{0\le m_{11} \le n_{11}+n_{01}} \{ n/(n_1n_0) \cdot q_{\text{HG}}(1-\beta_2; n, m_{11}, n_0) - m_{11}/n_1\}. $
Below we intuitively explain the connection between the randomization test in Theorem (ref) and the sharp bound in zhang2003estimation. When the sample size is large, due to randomization, both $\underline{m}$ and $\overline{m}$ are approximately $n_1 \cdot n_{01}/n_0$, while $\overline{d}$ is approximately $0$. Consequently, $\underline{K} \approx \max\{0, n_1 n_{01}/n_0 - n_{10}\}$ and $\overline{K} \approx \min\{n_{11}, n_1 n_{01}/n_0\}$. In addition, for each $K\in [\underline{K}, \overline{K}]$, when computing $T_K$ in step (iv) of Algorithm (ref), we have $J \approx K \cdot n_0/n_1$ and $L \approx n_1 n_{01}/n_0 - K$.
We now explain the test statistic $T_K$ in step (iv) of Algorithm (ref), where $K$ can be interpreted as a candidate value of $m_{11}^\texttt{t}$. Ignoring ties, $T_K$ corresponds to a comparison between treated and control groups under the following worst-case imputation of the composite potential outcomes. For the treated group, the imputed sample consists of $K$ observed treated units with the smallest imputed control potential outcome, $L \approx n_1 n_{01}/n_0 - K$ negative infinities, and approximately $n_1 - n_1 n_{01}/n_0$ positive infinities. For the control group, the imputed sample consists of $J \approx K \cdot n_0/n_1$ observed control units with the largest observed outcomes, $n_{01}-J \approx n_{01} - K \cdot n_0/n_1$ negative infinities, and $n_{00} = n_0 - n_{01}$ positive infinities. We can verify that the proportions of positive and minus infinities are approximately the same between the treated and control group. Therefore, in the test statistic $T_K$, we are essentially comparing the $K/n_{11}$ fraction of the observed treated units with the smallest imputed outcomes to the $Kn_0/(n_1 n_{01})$ fraction of the observed control units with the largest observed outcomes. In step (vi) of Algorithm (ref), we further take a worst case over $K\in [\underline{K}, \overline{K}]$.
To better connect to the bounds in zhang2003estimation, we introduce $\pi \coloneqq {n_{01}}/{n_0} - {K}/{n_1}$, which can be interpreted as a candidate value for the proportion of units with $M_i(0)=1$ and $M_i(1) = 0$. Then the above fractions of observed treated and control units being compared in $T_K$ are equivalently $n_{01}/n_0/(n_{11}/n_1) - \pi/(n_{11}/n_1)$ and $1 - {\pi}/{(n_{01}/n_0)}$, with a further worst-case consideration over $\pi \in [ \max\{0, n_{01}/n_0 - n_{11}/n_1\}, \ \min\{n_{01}/n_0, 1 - n_{11}/n_1\} ]$. This exactly mirrors the comparison underlying the sharp lower bound of the average treatment effect within the principle stratum of units who would be observed under both treatment and control.
We now consider the following bounded null hypothesis:\footnote{The bounded null hypothesis of form $H_{\succcurlyeq \boldsymbol{\delta}}: \boldsymbol{\tau} \succcurlyeq \boldsymbol{\delta}$ can be analogously tested by switching the treatment label or changing the outcome sign.}
where $\boldsymbol{\delta} \in \mathbb{R}^{n}$ is a fixed vector. The null hypothesis in (ref) states that each individual treatment effect is bounded above by the corresponding component of $\boldsymbol{ \delta}$. In the special case where $\boldsymbol{\delta} = \boldsymbol{0}$, this reduces to the null hypothesis that all individual treatment effects are non-positive.
We give a brief remark on the validity of the randomization p-values developed in Theorems (ref)–(ref) and Proposition (ref) for testing the bounded null in (ref). Because the test statistics in (ref), (ref) and (ref) are all distribution-free and effect increasing,\footnote{ A statistic $t(\boldsymbol{z}, \boldsymbol{y})$ is said to be effect increasing if its value weakly increases when the outcomes of treated units (i.e., $y_i$ for $z_i = 1$) are increased and the outcomes of control units (i.e., $y_i$ for $z_i = 0$) are decreased.} we can verify that these $p$-values are nondecreasing in the hypothesized effects $\delta_i$s. Therefore, under the bounded null in (ref), they are no less than their corresponding values using the true treatment effects for imputing potential outcomes. This explains their validity for testing the bounded null hypothesis.
We first conduct a simulation study to evaluate the finite-sample validity of the proposed randomization test under various missingness mechanisms. Specifically, we compare three inferential procedures: (1) the worst-case inferential procedure using the outcome-only test statistic in Section (ref); (2) the worst-case inferential procedure using the outcome-missingness composite test statistic in Section (ref) with different choices of $\boldsymbol{b}$; and (3) the naive procedure that drops units with missing outcomes and applies the standard randomization test to the observed sample.
We generate potential outcomes as $Y_i(0) = Y_i(1) \overset{\text{i.i.d.}}{\sim} \mathcal{N}(0, 1)$ with $N = 500$, and assign treatment via a CRE allocating half the units to treatment and half to control. We test Fisher's sharp null hypothesis of no treatment effect for any unit, formally stated as $H_0: \boldsymbol{\tau}=\boldsymbol{0}$, at $10\%$ nominal level. We consider four types of missingness mechanisms, each evaluated under two scenarios in which approximately 5% and 10% of units having missing outcomes:
Table (ref) reports the empirical Type I error rates of several tests under the missingness mechanisms described above. Column 5 corresponds to the proposed tests based on control potential outcomes. Columns 6--9 correspond to the proposed tests using composite potential outcomes with different values of $\boldsymbol{b}$, where we vary one element of $\boldsymbol{b}$ under each missingness mechanism. Column 10 reports the error rates of the “naive” method that excludes units with missing outcomes. All tests are based on the Wilcoxon rank-sum statistic.
Table (ref) reveals several key findings. First, the proposed worst-case inference procedures control the Type I error rate at or below the nominal level across all missingness mechanisms, confirming their validity. By contrast, the “naive” method exhibits substantial Type I error inflation under the threshold and monotone missingness mechanisms, with rejection rates exceeding $75\%$ in several cases, even when the overall proportion of missing outcomes is small. This result highlights an important concern for empirical practice: inference procedures that ignore the missingness mechanism may lead to severe size distortions and misleading conclusions. Under the sharp missingness mechanism, however, the “naive” method controls the Type I error rate, consistent with our theoretical results.
Second, the choice of $\boldsymbol{b}$ affects the conservativeness of the proposed tests based on composite potential outcomes under the threshold and monotone missingness mechanisms. The optimal choice varies with the missingness mechanism. Specifically, we recommend $b_{01} = \infty, b_{10} = -\infty$ under general missingness (which includes threshold missingness as a special case), $b_{00} = b_{01} = \infty$ under monotone positive missingness, and $b_{00} = b_{10} = -\infty$ under monotone negative missingness. The simulation results align with these recommendations: the empirical rejection probability is closest to the nominal level when the recommended $\boldsymbol{b}$ is used. Under the sharp missingness mechanism, the rejection probability remains close to the nominal level for all choices of $b_{00}$, consistent with the discussion in Section (ref).
We next conduct a simulation study to compare the power of the two-step methods with that of the one-step methods under various missingness mechanisms. Potential outcomes are generated according to $Y_i(0) \overset{\text{i.i.d.}}{\sim} \mathcal{N}(0,1)$ and $Y_i(1) = Y_i(0) + \delta$, with sample size $N = 500$ and $\delta$ ranging from 0 to 0.8. We again consider the CRE, which assigns half of the units to treatment and the remaining half to control. For each design, we test the sharp null hypothesis $H_0:\boldsymbol{\tau}=\boldsymbol{0}$ at the $10\%$ nominal level. Missingness indicators are generated as in Section (ref), with approximately $20\%$ of units having missing outcomes:
The results in Figure (ref) reveal two main findings. First, the two-step method achieves substantially higher power than the one-step method under both threshold and monotone missingness patterns, while maintaining type I error control at the nominal level. For example, under monotone positive missingness with a treatment effect of 0.5, power increases from 18% for the one-step method to 49% for the two-step method with $\beta = 0.1\alpha$. Second, the test is relatively robust to a range of positive $\beta$ values. Based on the simulation, we suggest $\beta = 0.1\alpha$ in practice.
We use data from the National Job Corps Study, a large randomized evaluation funded by the U.S. Department of Labor to assess the Job Corps program’s impact on labor market outcomes, especially wages lee2009training. As one of the largest federally funded job training programs, Job Corps serves economically disadvantaged youth aged 16–24, offering residential services, healthcare, and educational and vocational training. Participants typically spend about eight months in the program, receiving roughly 1,100 hours of instruction—equivalent to one academic year of high school.
Even with randomized assignment, estimating the impact of the training program on wages is empirically challenging due to sample selection. Wages are only observed for the employed, but employment itself may be influenced by the program heckman1974shadow. As a result, treatment and control groups may not be comparable among those employed. lee2009training argues that the missingness mechanism at weeks 90, 135, and 180 is monotone positive.
We begin by testing the sharp null hypothesis that the treatment effect is zero for all individuals. Table (ref) presents randomization p-values under various missingness assumptions for both one-step and two-step method. We fail to reject the null hypothesis at conventional significance levels under general missingness assumptions and monotone positive/negative missingness assumptions. However, under sharper assumptions—such as sharp missingness or missing at random\footnote{We discuss randomization inference under the missing-at-random mechanism in the supplementary material.}—the sharp null can be rejected at the 5% level. Note that the two-step procedure yields no improvement under either general or monotone missingness: the former is due to high overall missingness, and the latter is due to similar missing proportions between treated and control groups.
Table (ref) reports the 95% confidence intervals derived via Lehmann-style test inversion under the assumption of a constant treatment effect. At week 90 after the treatment, the Job Corps sample includes 5546 treated and 3599 control individuals. Under the general missingness assumption, the 95% confidence interval for the treatment effect is uninformative: $(-\infty, \infty)$. However, assuming monotone positive missingness, the interval narrows substantially to $[-0.042, 0.087]$, despite 54% of wage observations being missing. By the end of the 180-week follow-up period, the interval remains informative, ruling out effects more negative than -8% and more positive than 12%. Confidence intervals under stronger assumptions (sharp or random missingness) are even tighter, reflecting the additional statistical power provided by these assumptions.
Randomization inference has been widely used across scientific disciplines. However, naive application in experiments with sample attrition can lead to severe size distortions. We address this problem by developing computationally efficient methods for randomization inference that remain valid under a broad class of potentially informative missingness mechanisms. These proposed methods are valid for testing both sharp and bounded null hypotheses.
We expect our methods to be practically useful for applied researchers. When missingness mechanisms are largely unknown, our methods under general missingness should be used. If the experiment satisfies monotone or sharp missingness, the corresponding methods should be applied. In addition, as discussed in the supplementary material, when missingness is random or believed to be random, randomization inference can be applied directly to the observed sample. An R package implementing the proposed methods is publicly available at \url{https://github.com/peizansheng/riattrition}.
Finally, kwon2024testing note that their test of the sharp null of full mediation can also be used to test the sharp null of no treatment effect in our setting. This follows because the sharp null of no treatment effect implies no treatment effect for always-in-sample and never-in-sample units (i.e., $M_i(1)=M_i(0)=1$ and $M_i(1)=M_i(0)=0$, respectively). Their testing procedure differs from ours, and comparing the two approaches is an important direction for future work.
\setcounter{equation}{0} \setcounter{section}{0} \setcounter{figure}{0} \setcounter{example}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{theorem}{0} \setcounter{lemma}{0} \setcounter{table}{0} \setcounter{condition}{0} \setcounter{page}{1}
In this section, we assume non-informative missingness, under which we can infer treatment effects most precisely. Importantly, unlike Assumptions (ref)--(ref), the potential missingness indicators are treated as random variables and are not conditioned upon. Intuitively, we assume that the potential missingness indicators are independent of the potential outcomes little2019statistical, bai2024revisiting. Since the potential outcomes are already conditioned on in our randomization-based inference and treated as fixed (arbitrary) constants, this assumption effectively requires the potential missingness indicators to be i.i.d. across all units. We formally state this missing at random assumption below, explicitly including the conditioning on potential outcomes to emphasize the non-informativeness of the missingness.
Assumption (ref) does not restrict the dependence between $M(1)$ and $M(0)$, so the information on potential missingness and outcomes remains the same as under general missingness mechanisms (see Panel A of Table (ref)). However, Assumption (ref) implies that, after appropriate conditioning, the experiment can be viewed as a CRE restricted to units with observed outcomes, as in ghanem2023testing. Consequently, standard randomization tests can be applied without concern for missing data under this assumption. Specifically, we can show that
where $\mathcal{S} = \{i: M_i = 1, 1\le i \le n\}$, $\boldsymbol{Z}_{\mathcal{S}}$ and $\boldsymbol{Z}_{\mathcal{S}^\complement}$ denote the subvectors of $\boldsymbol{Z}$ corresponding to units in $\mathcal{S}$ and $\mathcal{S}^\complement$, respectively, $n_{\mathcal{S}1} = \sum_{i\in \mathcal{S}} Z_i$, and $n_{\mathcal{S}0} = |\mathcal{S}|-n_{\mathcal{S}1}$. That is, after proper conditioning, the treatment assignment for units in $\mathcal{S}$ follows the same distribution as that of a CRE, which randomly assigns $n_{\mathcal{S}1}$ units from $\mathcal{S}$ to treatment and the remaining $n_{\mathcal{S}0}$ to control. The following theorem establishes the validity of randomization tests conditional on units with observed outcomes.
Theorem (ref) is essentially testing the following “modified” sharp null hypothesis for units in $\mathcal{S}$ whose outcomes are observed:
The sharp null $H_{\boldsymbol{\delta}}$ implies $H_{\boldsymbol{\delta}, \mathcal{S}}$, so testing $H_{\boldsymbol{\delta}}$ reduces to conducting a randomization test for $H_{\boldsymbol{\delta}, \mathcal{S}}$ using only units with observed outcomes. Note that $\mathcal{S}$ is a random subset of units that generally depends on the treatment assignments. Moreover, the procedure in Theorem (ref) also works when using a generic test statistic, such as the difference-in-means test statistic, since there will be no missing outcomes once we focus only on units in $\mathcal{S}$.
Although the testing procedures in Theorems (ref) and (ref) are the same, the key difference between the two theorems is that $\mathcal{S}$ is random in Theorem (ref), whereas it is fixed under the sharp missingness mechanism in Theorem (ref).
Another remark regarding Theorem (ref) concerns the necessity direction. In general, random treatment assignment (such as in the CRE) and $\boldsymbol{Z} \perp \!\!\! \perp (\boldsymbol{Y}(0), \boldsymbol{Y}(1)) \mid \boldsymbol{M}$ do not imply $(\boldsymbol{M}(0), \boldsymbol{M}(1)) \perp \!\!\! \perp (\boldsymbol{Y}(0), \boldsymbol{Y}(1))$. A counterexample is the case where $\boldsymbol{M}(1) = \boldsymbol{M}(0) = \boldsymbol{Y}^{\star}(1) = \boldsymbol{Y}^{\star}(0)$, and $\boldsymbol{Z}$ is drawn from a CRE.