EconBase
← Back to paper

Don't Drop the Singletons: Efficient Inference for Pairwise Experiments with Independent Attrition

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

69,738 characters

Don't Drop the Singletons: Efficient Inference for Pairwise Experiments with Independent Attrition


\maketitle
\begin{abstract}
Pairwise randomization can yield substantial efficiency gains in experiments. Yet methodological guidance cautions against pairwise randomization, especially in settings with attrition, partly because common practices for estimation (i.e., pair fixed effects) imply discarding data from incomplete pairs thus exacerbating data loss from attrition. This practice of dropping incomplete pairs reduces statistical power of tests as well as precision of estimates, in paired experiments, compared to designs with less finely stratified treatment assignment. We argue that this concern is misplaced if attrition is independent of treatment status and potential outcomes, and that these issues follow from an inefficient use of the data that remains post-attrition. First, we show how, by using a specific permutation test, it is possible to use all observed units for inference (complete pairs and incomplete pairs where one unit attrits) while still exploiting the pairwise randomization design structure. The test procedure we suggest provides exact size control under the sharp null. Second, we study an optimally weighted estimator that efficiently combines within-pair and across-pair comparisons. Finally, we show that combining these two insights yields a test procedure that dominates the two commonly used inference methods (a paired \(t\)-test and the two-sample \(t\)-test) in power, for any level of attrition. Usefully for applied researchers, we show that the efficient procedure can be implemented via a weighted fixed effects regression, straightforward in standard software. In sum, our results provide researchers with practical tools for conducting experiments with pairwise randomization without sacrificing observations or statistical power when facing independent attrition.
\end{abstract}

\section{Introduction}\label{sec:introduction}

Experiments with pairwise randomization, i.e., randomly assigning treatment to one unit in blocks of two units each, can offer substantial efficiency gains over less finely stratified randomization, particularly when sample sizes are limited.
As surveyed in Bai, Shaikh, and Tabord-Meehan (forthcoming), pairwise randomization designs have unique advantages for estimation precision and test power. Imai, King, and Nall (2009) demonstrate for cluster-randomized experiments that pair matching can dramatically increase statistical power, arguing that ``from the perspective of bias, efficiency, power, robustness or research costs, and in large or small samples, pairing should be used in cluster-randomized experiments whenever feasible.'' Bruhn and McKenzie (2009) provide simulation evidence supporting this claim, finding that pairwise randomization can outperform alternative randomization strategies in achieving balance in potential outcomes when covariates with good predictive power for outcomes exist (and the researchers have a single outcome to focus on).
More recently, Bai (2022) establishes formal optimality results, showing that pairwise randomization designs emerge as optimal among a large set of randomization schemes.
These theoretical and empirical results suggest pairwise designs should be widely adopted where efficiency matters.

Despite these advantages, methodological guidance has cautioned against pairwise randomization designs when attrition is anticipated, for a number of reasons.
The widely-used evaluation handbook by Glennerster and Takavarasha (2013, 159) advises researchers facing attrition risk to ``use strata that have at least four units rather than pairwise randomization.''\footnote{Not to be confused with Athey and Imbens (2017)'s argument that groups larger than pairs can provide within-stratum variance estimators, that may improve variance estimation and thus inference, even without attrition. This is a different argument than the one about attrition, and we do not focus on it in this paper.} Similarly, Donner and Klar (2000) characterize the need to drop singletons (incomplete pairs) as a fundamental weakness of the approach (Bai et al. 2024 cover this debate in more detail). Likewise, in a blogpost, McKenzie (2022) identifies attrition as ``main reason for being cautious'' regarding pairwise randomization, also noting that when units drop out, the pairwise randomization designs are usually analyzed in a way that discards not only the unit that attrits but also its paired counterpart, the orphaned singletons, thus increasing effective attrition rates. Our paper specifically addresses this concern by suggesting and discussing a test method that does not discard these singletons.

The core of our paper is based on the observation that if attrition occurs independently of treatment assignment, the canonical test procedures for pairwise designs (such as the paired \(t\)-test) unnecessarily discard useful information.\footnote{Where attrition depends on treatment and potential outcomes, the problem is a different one and identification itself is at stake (Bai et al. 2024)}
We argue that this can be overcome; i.e., that it is possible to conduct valid and powerful inference while fully exploiting the pairwise design structure and using all observed units, including the orphaned singletons in pairs with attrition. Specifically, we show that pairwise designs analyzed with our methods dominate standard inference approaches. Whether they also dominate the alternatives of designs with larger strata is a related but different question that remains open and that we discuss in the conclusion and through a short simulation study.

Conveniently, our efficient procedure can be estimated via a single weighted fixed-effects regression. So researchers who anticipate exogenous attrition can retain the efficiency benefits of pairing at essentially no implementation cost, rather than falling back to a study design with coarser stratification.

We build on a literature on the econometrics of finely stratified experiments that has seen some new progress in the past years.
Bai, Romano, and Shaikh (2022) study inference under pairwise randomization \emph{without} attrition: they show that both the two-sample and the paired \(t\)-tests are asymptotically conservative, propose an adjusted variance estimator that restores asymptotic exactness, and show that the within-pair permutation test is finite-sample valid under the sharp null and, when studentized by the adjusted variance, asymptotically exact under the weak null.
Our point of departure is an observation in Bai (2022, Online Appendix C.3) :
the difference-in-means, computed without dropping singletons, remains consistent for the ATE under two conditions: attrition is independent of treatment assignment (conditional on baseline covariates if included), and
attrition is independent of the individual treatment effect.
No variance, test, or distinction between complete pairs and singletons is developed there.
We build on this observation and develop it into an inference procedure that separates the complete-pair and singleton components, derives their variance-minimizing combination, and provides exact randomization-based tests.
Bai et al. (2024) study pairwise randomization \emph{with} possibly endogenous attrition, deriving the estimands recovered by the difference-in-means estimator both when singletons are dropped and when they are retained. They find limited support for the practice of dropping singletons, showing that under heterogeneous treatment effects, the difference in means when singletons are dropped generally identifies a weighted average of conditional treatment effects rather than the ATE.
Their focus is identification, whereas we hold the estimand fixed by assuming independent attrition and optimize the inference procedure: we show how retained singletons should be combined with complete pairs, quantify the resulting power gain, and provide exact tests.
Relatedly, Fukumoto (2022) investigates implications of the two common ways of dealing with outcome-correlated attrition: either dropping or retaining singletons in the presence of attrition. These two testing approaches (below we refer to those as the paired \(t\)-test, often implemented via a regression with pair fixed effects, and two-sample \(t\)-test, often implemented via a regression without fixed effects) have been widely used in practice when analyzing pairwise randomized experiments with attrition.

Our main contribution is to show that the choice between the two standard approaches, paired \(t\)-test and two-sample \(t\)-test, is a false dichotomy, and to provide a combined alternative that produces valid inference, has optimal power in a clearly defined setting, and is easy to implement. Both standard approaches waste information in different ways: the paired \(t\)-test discards all singletons, thus losing information from singletons, while the two-sample \(t\)-test ignores the pairwise design structure, thus wasting information about the design. By combining randomization inference with an optimally weighted estimator that efficiently combines information from complete pairs and singletons, we can make full use of all information for the estimand, and by using randomization inference, we can conduct valid inference that respects the pairwise design structure, providing exact size control under the sharp null for any attrition rate, while dominating standard approaches in power.\footnote{A related literature in biostatistics on incomplete paired data dates back to Ekbohm (1976). Closest to us is Amro and Pauly (2017), who construct a permutation test also combining a complete-pairs test statistic with an unpaired-singletons statistic.
  Our approach differs in four respects:
  First, our framework allows for optimal combination. We combine two unbiased estimators with weights derived to minimize the variance of the combined estimator, which maximizes power of the resulting test for our setting. Second, we combine estimators directly rather than pre-standardized test statistics. Third, in our setting exactness follows from the treatment-assignment mechanism in an RCT, without distributional assumptions. Fourth, and building on the first three points, we derive closed-form power and a dominance result.}

In detail, our paper's contribution is threefold:
First, we propose a permutation-based randomization inference procedure that respects the pairwise design structure while using all observed units after attrition: complete pairs and singletons. This procedure provides exact size control under the sharp null hypothesis for any attrition rate when attrition is independent of treatment assignment and potential outcomes.
Second, we introduce an optimally weighted estimator that combines information from within-pair differences (from complete pairs) and between-unit comparisons (from singletons). This estimator is efficient in a clearly defined setting and, under constant effects, yields a test that dominates both the paired \(t\)-test (which discards observations) and the two-sample \(t\)-test (which ignores pairing) in power.
Third, we show that the optimally weighted estimator can be computed via a weighted regression with appropriate observation weights, enabling straightforward implementation in standard statistical software. This works via weighted OLS, by regressing the outcome on treatment status and \emph{pair fixed effects only for complete pairs}, while including singletons with a lower per-observation weight that reflects their contribution to the variance of the estimator.

When facing attrition, researchers often present results across multiple methods with varying identifying assumptions, from strongest to weakest. Our optimally weighted randomization inference procedure assumes attrition is independent of treatment assignment and, under this assumption, uses all observed units optimally.
Analyses based on our procedure should thus be presented next to more conservative approaches that relax independence, such as the endogenous-attrition identification analysis of Bai et al. (2024) or bounding methods for non-random attrition, with our estimate serving as the point obtained under the strongest assumption about attrition against which those weaker-assumption analyses can be compared.

The remainder of the paper is organized as follows. \Cref{sec:setting} presents our formal setting, describes the four inference procedures we compare, and introduces our optimally weighted estimator. This section also derives the asymptotic power of each procedure as a function of attrition rate and matching quality.
\Cref{sec:results} presents our main results: we illustrate the power functions across different scenarios, establish exact size control under the sharp null, demonstrate the dominance of our optimally weighted procedure, and extend our approach to weak-null testing via studentized randomization inference.
\Cref{sec:conclusion} concludes with a discussion of practical implementation and directions for future research. All proofs are collected in the appendix.

\section{Setting}\label{sec:setting}

\subsection{Model and Data Generating Process}\label{model-and-data-generating-process}

We consider experiments with \(N = 2n\) units organized into \(n\) pairs. Each pair is drawn independently from a common population of pairs, i.e., we study this question through the lens of a superpopulation model with sampling uncertainty, where pairs are sampled.
Our sampling model follows the pairs-as-primitive superpopulation perspective of Fukumoto (2022), who, following Imai (2008) and Imbens and Rubin (2015) (ch.~10), draws the sampled pairs from a superpopulation of pairs; we differ in drawing pairs i.i.d. from a bivariate distribution \(F\) rather than from a large finite population of pairs, and in restricting attention to attrition that is independent of assignment.\footnote{Our framework also differs from the superpopulation model for pairs of Bai, Romano, and Shaikh (2022), in which units are sampled i.i.d. and pairs are formed on observed covariates, so that within-pair dependence is endogenous and heterogeneous across pairs; we instead sample pairs i.i.d. from a fixed joint distribution \(F\) with a common within-pair correlation, trading covariate-adaptive generality for closed-form power expressions and finite-sample statements.
  Two remarks on this choice.
  First, it is made for ease of exposition of our main argument. With pairs as i.i.d. draws, match quality reduces to the single parameter \(r\), the optimal weights and all power functions below have closed forms, and the comparison across inference methods can be read directly off the resulting expressions.
  Second, the choice is not what drives our results and recommendation for applied researchers. The exact size control of the randomization tests (Proposition 6) is a finite-sample statement conditional on potential outcomes and the attrition pattern; it uses only the within-pair randomization and holds verbatim whether pairs are sampled from \(F\), units are sampled i.i.d. and matched on covariates as in Bai, Romano, and Shaikh (2022), or the sample is a fixed population as in Chaisemartin and Ramirez-Cuellar (2024). Likewise, the dominance of the optimally weighted procedure rests on inverse-variance weighting of two unbiased, uncorrelated component estimators, of which the mean within-pair difference weights \((1,0)\) and the difference-in-means weights \((1-q,\,q)\) are special cases; this logic is not specific to our sampling model. What is framework-specific is the closed-form characterization of weights and power in terms of \((q, r)\), and the validity of unstudentized weak-null inference, which we delimit in \Cref{weak-null-discussion}.
  So, a reader who prefers a different framework can retain the substantive conclusion: under attrition that is independent of treatment assignment, the default choice between discarding singletons and ignoring the pairing is dominated.} We index pairs by \(p=1, ..., n\) and units within pairs by \(g\in\{1,2\}\). Specifically, for a pair, the untreated potential outcomes of its two units, \((Y_{p,1}(0), Y_{p,2}(0))\), are an i.i.d. draw from a bivariate distribution \(F\). The two coordinates of \(F\) are exchangeable, i.e.~swapping which unit is labeled 1 vs.~2 leaves the distribution unchanged, so both members share a common marginal distribution. Every ``for large \(n\)'' statement below refers to \(n \to \infty\) pairs drawn from \(F\).\footnote{Asymptotic normality follows from the Lindeberg--Feller central limit theorem for triangular arrays (Vaart 1998, Proposition 2.27).} This framework is chosen to resemble a setting in which population pairs pre-exist (e.g., neighbors) or are formed through some matching procedure (e.g., pairing on Mahalanobis distance of baseline covariates), which is encoded in the dependence structure of \(F\).

We denote the marginal variance of the untreated potential outcome by \(\sigma^2_{Y_0} = \text{Var}(Y_{p,g}(0))\), and require it to be finite. The within-pair correlation
\[r = \text{Corr}(Y_{p,1}(0), Y_{p,2}(0)) \in [0,1]\]
is a functional of \(F\) and hence a population parameter. It is typically unobserved and can be understood as a measure of match quality.
Higher values of \(r\) indicate better matching quality. In the limit case, \(r = 0\), pairing provides no efficiency gain over complete randomization; conversely as \(r \to 1\), paired units have identical control potential outcomes.\footnote{If treatment effects are heterogeneous rather than a constant shift, \(r\) remains well-defined as a property of the untreated-outcome but its estimation from post-treatment data becomes challenging.}

For tractability, we first consider constant treatment effects. We define \(Y_{p,g}(1) := Y_{p,g}(0) + \tau\), so the treated potential outcome is derived from the drawn untreated outcome rather than sampled separately. Our power calculations use this constant-shift specification. Our results extend straightforwardly to a location-shift family, \(Y_{p,g}(1) \overset{d}{=} Y_{p,g}(0) + \tau\), which permits heterogeneous effects while fixing the average treatment effect at \(\tau\); we return to discussing unconstrained heterogeneity in the second half of the paper.

Two further layers of randomness exist on top of the pair sampling: treatment assignment and attrition.
First, treatment assignment: within every pair, one unit is assigned to treatment and the other to control by an independent coin flip, decided independently of the drawn outcomes and applied to all pairs before any attrition.
Treatment status is denoted by the indicator \(D_{p,g}\in\{0,1\}\); for notational convenience we write \(T\) (treated, \(D=1\)) and \(C\) (control, \(D=0\)) as subscripts on group means and variances.
Second, attrition: exactly a fraction \(q\) of outcomes is removed deterministically, yielding exactly \(m = n(1-q)^2\) complete pairs, \(k_1 = nq(1-q)\) observed treated singletons, and \(k_0 = nq(1-q)\) observed control singletons.\footnote{This deterministic specification keeps the observed counts fixed and preserves the feature that matters, balance of attrition by treatment status; under random unit-level attrition the counts become random but converge to the same values and the power formulas are unchanged in the limit. Further, we ignore integer-rounding of \(n(1-q)^2\), \(nq(1-q)\), and \(nq^2\) throughout, which affects the results only at \(o(1)\).}
We call pair \(p\) \emph{complete} if both units are observed, and a \emph{singleton} if only one unit is observed.

All results below are based on this framework, with its three independent sources of randomness: the draw of pairs from \(F\), the within-pair assignment, and attrition. The asymptotic power results (Propositions 1--5) are driven by the sampling of pairs from \(F\). Finite-sample exact size control (Proposition 6) follows by conditioning on the realized draw and attrition pattern and using only the within-pair assignment randomization; since this holds for every draw, it also holds unconditionally over the draw.

\subsection{Considered Inference Methods}\label{considered-inference-methods}

We consider four inference procedures, all based on two-sided tests at level \(\alpha\).
The first two methods are standard approaches commonly used in practice, each with known limitations under attrition.
The latter two methods are randomization inference procedures that respect the pairwise design structure while using all observed units.

\textbf{Paired \(t\)-test.} Using only the \(m = n(1-q)^2\) complete pairs, the estimator is the mean within-pair difference
\[\hat{\tau}_{\text{pair}} = \frac{1}{m}\sum_{p:\,\text{complete}} (Y_{p,T} - Y_{p,C}),\]
with sampling variance
\[\text{SE}_{\text{pair}}^2 = \frac{2\sigma^2_{Y_0}(1-r)}{n(1-q)^2},\]
(derived in the proof of Proposition 1). Inference uses the \(t\)-statistic \(\hat{\tau}_{\text{pair}}/\widehat{\text{SE}}_{\text{pair}}\) against a \(t\) reference with \(m-1\) degrees of freedom. Under attrition this discards all singleton observations. Empirically, economists typically implement this test in a fixed-effects regression: \(\hat{\tau}_{\text{pair}}\) is the coefficient on \(D_{p,g}\) in the OLS regression \(Y_{p,g} = \beta D_{p,g} + \alpha_p + u_{p,g}\) with pair fixed effects \(\alpha_p\), where the pair fixed effects drop the singletons.

\textbf{Two-sample \(t\)-test.} Using all observed units regardless of pair status, the estimator used for this test is the difference in means \[\hat{\tau} = \bar{Y}_T - \bar{Y}_C.\] The two-sample \(t\)-test computes its standard error under the (incorrect) assumption of complete randomization,
\[\text{SE}_{\text{assumed}}^2 = \frac{2\sigma^2_{Y_0}}{n(1-q)},\]
and refers \(\hat{\tau}/\widehat{\text{SE}}_{\text{assumed}}\) to a \(t\) (large-\(n\): normal) distribution. Empirically, economists typically implement this test in a plain OLS regression: \(\hat{\tau}\) is the coefficient on \(D_{p,g}\) in the OLS regression of \(Y_{p,g}\) on \(D_{p,g}\) with a constant.

The variance used for inference in this test, \(\text{SE}_{\text{assumed}}^2\), is not the true sampling variance of \(\hat{\tau}\). Under the pairwise design, the true sampling variance of \(\hat{\tau}\) over the draw from \(F\), the within-pair assignment, and attrition is
\begin{equation}
\text{SE}_\tau^2 = \frac{2\sigma^2_{Y_0}[1 - r + qr]}{n(1-q)}, \label{eq:se-tau}
\end{equation}
derived in the proof of Proposition 3. Since \(\text{SE}_\tau^2 \le \text{SE}_{\text{assumed}}^2\), the two-sample \(t\)-test is typically conservative: it ignores the design-based efficiency gain from pairing. The two-sample \(t\)-test and the difference-in-means randomization test below share the estimator \(\hat{\tau}\) and the true sampling variance \(\text{SE}_\tau^2\); they differ only in the reference distribution used for inference. The former uses a plug-in normal with the mis-specified \(\text{SE}_{\text{assumed}}\), the latter uses the correct permutation distribution.\footnote{A third variant is possible: refer \(\hat{\tau}\) to a normal using a consistent estimate of the correct variance \(\text{SE}_\tau^2\) (via \(\hat{\sigma}^2_{Y_0}\) and \(\hat{r}\)) rather than \(\text{SE}_{\text{assumed}}\). In the setting without attrition, inference based on a consistent estimate of the correct variance is the main proposal of Bai, Romano, and Shaikh (2022); our variant is in the same spirit, though we do not establish a formal correspondence between the two.}

\textbf{Randomization inference (RI) using the difference in means as test statistic.}
Using all observed units, we compute the observed difference in means \(\hat\tau = \bar{Y}_T - \bar{Y}_C\).
We obtain the null distribution by repeatedly permuting\footnote{For complete pairs this amounts to randomly exchanging the treatment status between the observations.
  For singleton observations, since attrition is by assumption balanced by treatment status (exactly \(nq(1-q)\) observed treated and the same number of control units from singleton pairs), each permutation randomly reassigns their treatment labels while maintaining this balance, exactly like complete randomization with fixed group sizes.} treatment assignment within each pair and recalculating the test statistic for each permutation.
The two-sided \(p\)-value is the proportion of permutations, \(\pi\), yielding \(|\tau^{(\pi)}| \geq |\hat\tau|\).

This procedure respects the pairwise design structure through the permutation scheme while using all observed units, thus avoiding the limitations of the paired \(t\)-test and two-sample \(t\)-test under attrition. The permutation scheme is design-aware: complete-pair members are swapped within pair, singletons are permuted unrestricted (or more precisely: within pair but with an unobserved counterpart). The test statistic, however, weights every observation equally in \(\bar{Y}_T - \bar{Y}_C\), and so does not combine the complete-pair and singleton information according to their differing precision. The optimally weighted procedure below also addresses this remaining inefficiency through weighting.

\textbf{Randomization inference using an optimally weighted estimator as test statistic.}
The above RI procedure uses a simple difference in means as the test statistic, treating all observed units identically.
However, under the pairwise design with attrition, we can partition our observed sample into two distinct subsets with different precision:
complete pairs (for which we observe both treatment and control outcomes and can thus difference out pair-level heterogeneity) and
singletons (for which we observe only one unit and must estimate treatment effects from between-pair comparisons).
This partitioning suggests the existence of an \emph{optimally weighted estimator} that combines both according to their relative precision, which we derive below. Its sampling variance \(\text{SE}_{\text{oracle}}^2\) is given in Proposition 4.
For inference, we use this optimally weighted estimator as the test statistic in the randomization inference procedure described above.
The next subsection derives this optimally weighted estimator in detail.

\subsubsection{Optimally weighted estimator}\label{optimally-weighted-estimator}

The optimally weighted estimator is a linear combination of the two unbiased estimators for the average treatment effect \(\tau\), so, as algebraic building blocks, we define separate estimators for the two data partitions: complete pairs and singletons.

For complete pairs,
\[\hat{\tau}_{\text{pair}} = \frac{1}{m}\sum_{p:\, \text{complete}} (Y_{p,T} - Y_{p,C})\]
where \(Y_{p,T}\) and \(Y_{p,C}\) denote the observed outcomes of the treated and control member of pair \(p\), respectively, \(m= n(1-q)^2\) is the number of complete pairs, and the within-pair difference eliminates pair-specific heterogeneity.

For singletons, \[\hat{\tau}_{\text{single}} = \bar{Y}_{T,\text{single}} - \bar{Y}_{C,\text{single}}\] uses between-unit comparisons (retaining full \(\sigma^2_{Y_0}\) variance).

Both estimators are unbiased and efficient for the average treatment effect \(\tau\) within their respective data partitions.
The optimally weighted treatment effect estimator is then:
\[\hat\tau_{\text{oracle}} = w^* \cdot \hat{\tau}_{\text{pair}} + (1-w^*) \cdot \hat{\tau}_{\text{single}}\]
where \(w^*\) is chosen to minimize the variance of \(\hat\tau_{\text{oracle}}\) via standard inverse-variance weighting:
\[w^* = \frac{V_{\text{single}}}{V_{\text{pair}} + V_{\text{single}}}\]
where \(V_{\text{pair}} = 2\sigma^2_{Y_0}(1-r)/[n(1-q)^2]\) and \(V_{\text{single}} = 2\sigma^2_{Y_0}/[nq(1-q)]\) are the variances of the two component estimators.

We call this the \emph{oracle} estimator because \(w^*\) depends on population parameters (useful for power calculations); the \emph{feasible} estimator \(\hat\tau_{\text{feasible}}\) replaces \(w^*\) with an estimate \(\hat{w}^*\) and is defined below.

\paragraph{Oracle optimally weighted estimator}\label{oracle-optimally-weighted-estimator}

The oracle optimal weight simplifies to
\[w^* = \frac{(1-q)}{(1-q) + q(1-r)}.\]
By efficiently combining two efficient estimators, the oracle estimator achieves the minimum possible variance among all unbiased linear combinations of \(\hat{\tau}_{\text{pair}}\) and \(\hat{\tau}_{\text{single}}\), and therefore provides an efficiency bound for this setting.\footnote{We conjecture this is the efficiency bound in a stronger sense, i.e., that no regular estimator using the observed units achieves smaller asymptotic variance, since the two components together exhaust the observed-data information about \(\tau\). We do not pursue this here.}

\paragraph{Feasible optimally weighted estimator}\label{feasible-optimally-weighted-estimator}

In practice, the population variances and correlation are unknown and must be estimated from the data. The optimal weight \(w^*\) depends on the attrition rate \(q\) and the within-pair correlation \(r\). The attrition rate can be directly observed as \(\hat{q} = 1 - n_{\text{observed}}/N\), where \(n_{\text{observed}} = 2m + k_1 + k_0\) is the number of non-attrited units and \(N = 2n\). Estimating the within-pair correlation \(\hat{r}\) requires more care, and we discuss two approaches.

When baseline outcomes \(Y_{p,g}^{\text{baseline}}\) are available, the within-pair correlation can be estimated directly from pre-treatment data. Compute the within-pair variance \(\hat{\sigma}_{WP,\text{baseline}}^2 = \frac{1}{n-1} \sum_{p=1}^n (Y_{p,1}^{\text{baseline}} - Y_{p,2}^{\text{baseline}})^2\) and the marginal variance \(\hat{\sigma}_{Y,\text{baseline}}^2 = \frac{1}{2n-1} \sum_{p=1}^{n}\sum_{g \in \{1,2\}} (Y_{p,g}^{\text{baseline}} - \bar{Y}^{\text{baseline}})^2\), then estimate \(\hat{r}^{\text{baseline}} = 1 - \hat{\sigma}_{WP,\text{baseline}}^2/(2\hat{\sigma}_{Y,\text{baseline}}^2)\). This approach directly measures the matching quality achieved by the pairing procedure and avoids contamination from post-treatment heterogeneity. It assumes that baseline correlation predicts the correlation in control potential outcomes, which is plausible when outcomes are relatively stable and matching was based on baseline predictors.

We conjecture that even absent baseline data, a pre-specified educated guess for \(r\) (for instance, based on pilot studies or similar experiments) is likely to improve power over standard practice of using a paired \(t\)-test. The paired \(t\)-test, which discards all singleton observations, implicitly corresponds to setting \(\hat{r} = 1\) (optimal only under perfect matching), an assumption that is almost never justified in practice. Our optimally weighted approach nests the paired \(t\)-test as the limiting case when \(r \to 1\), but allows for the more realistic scenario of imperfect matching, thereby recovering information from singletons that would otherwise be discarded.

Without baseline data or educated guesses, the within-pair correlation could be estimated from post-treatment outcomes:

\begin{itemize}

  \setlength{\itemsep}{0pt}\setlength{\parskip}{0pt}
\item
  Within-pair variance: \(\hat{\sigma}_{WP}^2 = \frac{1}{m-1} \sum_{p:\, \text{complete}} (\Delta_p - \hat{\tau}_{\text{pair}})^2\), where \(\Delta_p = Y_{p,T} - Y_{p,C}\) is the treated-minus-control within-pair difference and \(\hat{\tau}_{\text{pair}} = \frac{1}{m}\sum_{p:\,\text{complete}}\Delta_p\) their mean.
\item
  Marginal variance: \(\hat{\sigma}_{Y_0}^2 = \frac{1}{n_0-1} \sum_{(p,g): D_{p,g}=0} (Y_{p,g} - \bar{Y}_0)^2\) where \(n_0 = m + k_0\) is the number of observed control units (control members of complete pairs plus control singletons), with \(n_0 = n(1-q)\).
\item
  Estimated within-pair correlation: \(\hat{r} = 1 - \frac{\hat{\sigma}_{WP}^2}{2\hat{\sigma}_{Y_0}^2}\)
\item
  Observed attrition rate: \(\hat{q} = 1 - n_{\text{observed}}/N\)
\end{itemize}

It is worth noting that for the purpose of sharp hypothesis testing via randomization inference, this approach requires no assumptions beyond those already maintained for the sharp null hypothesis \(H_0: Y_{p,g}(1) = Y_{p,g}(0)\), since under this null the within-pair variance estimates the within-pair variance in control outcomes exactly, with no contribution from treatment effect heterogeneity. The baseline data approach, when available, provides more stable estimates using all pairs and avoids potential complications from heterogeneous treatment effects when the estimator is used for purposes beyond testing (such as estimation or power calculations).

When using the feasible estimator with sample-estimated components as a test statistic in randomization inference, all variance estimates and weights should be recomputed for each permutation. Specifically, for each permuted treatment assignment \(D^{(\pi)}\), compute \(\hat{\sigma}_{Y_0}^2(\pi)\) using observations with \(D^{(\pi)} = 0\), derive \(\hat{r}(\pi)\) and \(\hat{w}^*(\pi)\) accordingly, and use \(\hat{w}^*(\pi)\) to compute the permuted test statistic.
The within-pair variance is recomputed as \(\hat{\sigma}_{WP}^2(\pi) = \frac{1}{m-1}\sum_{p:\,\text{complete}}(\Delta_p^{(\pi)} - \hat{\tau}_{\text{pair}}(\pi))^2\) for each permutation, since \(\Delta_p^{(\pi)} = Y_{p,T(\pi)} - Y_{p,C(\pi)}\) and its mean both change sign for pairs whose assignment is flipped. Under the sharp null this recompute is immaterial (it coincides with the fixed member-label version at \(\tau = 0\)) while under the alternative it is what keeps the estimate unbiased.
This procedure ensures that the permutation distribution properly reflects uncertainty in both the treatment effect and the nuisance parameters, preserving exact size control under the sharp null.

\subsubsection{Estimation}\label{estimation}

This paper focuses on inference, but the optimally weighted estimator can also be used for estimation and is efficient. Conveniently, the estimator can also be computed via a weighted regression with appropriately chosen weights for singletons vs complete pairs.

\textbf{Lemma (Optimally Weighted Estimator via Weighted Regression).} The oracle optimally weighted estimator can be obtained as the treatment coefficient from the weighted OLS regression:
\[Y_{p,g} = \beta D_{p,g} + \alpha_0\mathbbm{1}\{p \text{ is singleton}\} +  \alpha_p\mathbbm{1}\{p \text{ is complete pair}\} + \varepsilon_{p,g}\]
with observation-level weights proportional to the precision of each design component:
\[\tilde{w}_{\text{pair}} \propto \frac{1}{1-r}, \quad \tilde{w}_{\text{single}} \propto 1\]

In plain language, this regression specification uses all observed units, but includes pair fixed effects only for complete pairs, while singletons are absorbed into a separate intercept and complete pairs are weighted more heavily, reflecting their higher precision in estimating the treatment effect.
The resulting coefficient on \(D_{p,g}\) is equivalent to the optimally weighted estimator \(\hat\tau_{\text{oracle}}\).

\emph{Proof.} See appendix \cref{proof:weighted-regression}.

The observation weights implement inverse-variance weighting based on design-specific information content. Paired observations identify \(\tau\) through within-pair differences with variance proportional to \((1-r)\), while singleton observations identify \(\tau\) through unpaired comparisons with variance proportional to \(1\). The weights optimally balance the efficiency gain from pairing: when \(r\) is high, paired observations are heavily upweighted; when \(r \to 0\), weights converge to equal weighting.

\subsection{Asymptotic power as a function of attrition and matching quality}\label{asymptotic-power-as-a-function-of-attrition-and-matching-quality}

Below, we derive the asymptotic power of each of the four inference procedures under constant treatment effects \(\tau\) as a function of attrition rate \(q\) and matching quality \(r\).\footnote{All power statements are along local alternatives \(\tau_n = h/\sqrt{n}\) for fixed \(h\): the formulas give the limiting rejection probability with \(\tau\) replaced by \(\tau_n\). This is needed because at a fixed \(\tau \neq 0\) the permutation reference distribution is inflated by terms of order \(\tau^2/\sigma^2_{Y_0}\) (e.g., for pairs \(E[\Delta_p^2] = 2\sigma^2_{Y_0}(1-r) + \tau^2\)), so the formulas evaluated at fixed \(\tau\) overstate power by a relative error of that order; along \(\tau_n = h/\sqrt{n}\) this term is \(o(1)\). Propositions 1 and 2 use the same convention for comparability. For the calibrations in \Cref{sec:results}, \(\tau^2/\sigma^2_{Y_0} \approx 1\%\), so the fixed-\(\tau\) reading is accurate to that order.}
First, the pairwise design with independent attrition efficiently uses information in complete pairs, but loses information from attrited units.
Second, the two-sample \(t\)-test uses all observed units but has two shortcomings.
The difference in means used in this test fails to exploit the pairwise design structure, by treating observations from complete pairs the same as singletons, thus inefficiently combining available information. The standard error estimate used in this test also fails to account for the pairwise design structure, leading to incorrect and typically conservative standard errors.
Third, the randomization inference procedure using the difference in means as test statistic remedies the second shortcoming by respecting the pairwise design structure through the permutation scheme, while still using all observed units. However, like the two-sample \(t\)-test, it treats all observed units identically in the construction of the test statistic and thus does not remedy the first shortcoming of inefficiently combining information from pairs and singletons.

Finally, the randomization inference procedure using the optimally weighted estimator as test statistic remedies both shortcomings, efficiently combining information from complete pairs and singletons according to their relative precision, while respecting the pairwise design structure through the permutation scheme.

\subsubsection{\texorpdfstring{Paired \(t\)-test}{Paired t-test}}\label{paired-t-test}

\textbf{Proposition 1.} Under constant treatment effects along local alternatives \(\tau_n = h/\sqrt{n}\), the asymptotic power of the paired \(t\)-test at significance level \(\alpha\) is approximately:
\[\text{Power}_{\text{Paired}} = \Phi\left(\frac{\tau\sqrt{n(1-q)^2}}{\sqrt{2\sigma^2_{Y_0}(1-r)}} - z_{\alpha/2}\right) + \Phi\left(-\frac{\tau\sqrt{n(1-q)^2}}{\sqrt{2\sigma^2_{Y_0}(1-r)}} - z_{\alpha/2}\right)\]
where \(\Phi\) is the standard normal CDF and \(z_{\alpha/2} = \Phi^{-1}(1-\alpha/2)\).

\emph{Proof.} See appendix \cref{proof:power-paired-t-test}.

\subsubsection{\texorpdfstring{Two-sample \(t\)-test (difference in means)}{Two-sample t-test (difference in means)}}\label{two-sample-t-test-difference-in-means}

\textbf{Proposition 2.} Under constant treatment effects, \(\tau\), along local alternatives and pairwise randomization, the asymptotic power of the two-sample \(t\)-test (which ignores pairing) at significance level \(\alpha\) is approximately:
\[\text{Power}_{\text{Mean Diff}} = \Phi\left(\lambda(\delta - z_{\alpha/2})\right) + \Phi\left(\lambda(-\delta - z_{\alpha/2})\right)\]
where
\[\delta = \frac{\tau\sqrt{n(1-q)}}{\sqrt{2\sigma^2_{Y_0}}}\]
is the non-centrality parameter under the \(t\)-test's assumption of complete randomization, and
\[\lambda = \sqrt{\frac{\text{SE}_{\text{assumed}}^2}{\text{SE}_{\tau}^2}}\]
is the variance inflation factor accounting for the mismatch between the \(t\)-test's assumption (complete randomization) and the actual design (pairwise randomization). Here:
\[\text{SE}_{\text{assumed}}^2 = \frac{2\sigma^2_{Y_0}}{n(1-q)}\]
\[\text{SE}_{\tau}^2 = \frac{2\sigma^2_{Y_0}[1 - r + qr]}{n(1-q)} \text{(as given by \eqref{eq:se-tau})}\]

\emph{Proof.} See appendix \cref{proof:power-2-sample-t-test}.

\subsubsection{Randomization Inference (RI) with Difference in Means as Test Statistic}\label{randomization-inference-ri-with-difference-in-means-as-test-statistic}

\textbf{Proposition 3.} Under constant treatment effects along local alternatives \(\tau_n = h/\sqrt{n}\), the asymptotic power of the randomization inference test at significance level \(\alpha\) is approximately:
\[\text{Power}_{\text{RI, Mean Diff}} = \Phi\left(\frac{\tau}{\text{SE}_{\tau}} - z_{\alpha/2}\right) + \Phi\left(-\frac{\tau}{\text{SE}_{\tau}} - z_{\alpha/2}\right)\]
where \(\text{SE}_{\tau}^2\) is the same as in Proposition 2.

\emph{Proof.} See appendix \cref{proof:power-ri-diff-means}.

\subsubsection{Randomization Inference with an Optimally Weighted Estimator as Test Statistic}\label{randomization-inference-with-an-optimally-weighted-estimator-as-test-statistic}

\paragraph{Oracle optimally weighted RI}\label{oracle-optimally-weighted-ri}

\textbf{Proposition 4.} Under constant treatment effects along local alternatives \(\tau_n = h/\sqrt{n}\) and known population parameters (variances and \(r\)), the asymptotic power of the oracle optimally weighted RI test at significance level \(\alpha\) is approximately:
\[\text{Power}_{\text{oracle}} = \Phi\left(\frac{\tau}{\text{SE}_{\text{oracle}}} - z_{\alpha/2}\right) + \Phi\left(-\frac{\tau}{\text{SE}_{\text{oracle}}} - z_{\alpha/2}\right)\]
where

\[\text{SE}_{\text{oracle}}^2 = \frac{2\sigma^2_{Y_0}(1-r)}{n(1-q)(1 - rq)}\]

\emph{Proof.} See appendix \cref{proof:power-ri-optimal-oracle}.

\textbf{Analytic (oracle) test.} Proposition 4 yields an analytic test alongside the randomization test: refer \(\hat{\tau}_{\text{oracle}}/\text{SE}_{\text{oracle}}\) to the standard normal. It shares the estimator and sampling variance of the optimally weighted randomization test, and by the CLT established in the proof of Proposition 4 has the same asymptotic power \(\text{Power}_{\text{oracle}}\); the two differ only in reference distribution. The randomization test is preferred because it retains exact size under the sharp null (Proposition 6), which the normal approximation does not. Replacing \(\text{SE}_{\text{oracle}}\) with a consistent estimate inflates the variance by \(O(n^{-3/2})\) (Proposition 5), so the feasible analytic test is asymptotically equivalent. The power curves in \Cref{sec:results} are computed from this closed-form expression.

\textbf{Proposition 5 (Asymptotic equivalence of feasible and oracle procedures).} Under constant treatment effects along local alternatives \(\tau_n = h/\sqrt{n}\) and assuming \(F\) has finite fourth moments, the power of the feasible optimally weighted RI test approaches that of the oracle procedure as \(n \to \infty\):
\[\text{Power}_{\text{feasible}} \to \text{Power}_{\text{oracle}} = \Phi\left(\frac{\tau}{\text{SE}_{\text{oracle}}} - z_{\alpha/2}\right) + \Phi\left(-\frac{\tau}{\text{SE}_{\text{oracle}}} - z_{\alpha/2}\right)\]

More precisely, the feasible estimator satisfies:
\[\text{SE}_{\text{feasible}}^2 = \text{SE}_{\text{oracle}}^2 + O(n^{-3/2})\]

\emph{Proof.} See appendix \cref{proof:power-ri-optimal-feasible}.

\section{Results and Discussion}\label{sec:results}

\subsection{Illustration}\label{illustration}

In figure \ref{fig:power-plots}, we illustrate the theoretical power results for the four inference methods across varying attrition rates and matching quality. As a useful benchmark case, we chose a parametrization where, in the absence of attrition (\(q=0\)), the paired \(t\)-test achieves 90\% power. This corresponds to an effect size of \(\tau = 0.146\) and \(n = 1000\) pairs. To enable direct comparison across matching quality levels, we hold the baseline noise variance constant while allowing the marginal variance \(\sigma^2_{Y_0}\) to vary with \(r\) (higher matching quality implies higher correlation in potential outcomes, thus higher marginal variance for fixed noise). For the chosen parametrization, we consider three scenarios of matching quality:

\textbf{Weak matching (\(r = 0.003\)):} All methods start off with similar power at \(q=0\), as matching provides minimal efficiency gain, and pairwise matching effectively resembles complete randomization. As attrition increases, the paired \(t\)-test loses power quickly due to attrition of complete pairs, while the two-sample \(t\)-test maintains power by using all observed units. Both RI methods perform similarly to the two-sample \(t\)-test.

\textbf{Moderate matching (\(r = 0.25\)):} The paired \(t\)-test starts off at par with the two RI methods at \(q=0\), but loses power quickly as attrition increases and the paired \(t\)-test suffers from the high attrition of complete pairs. At high attrition rates, the two-sample \(t\)-test outperforms the paired \(t\)-test, because it uses all observed units. Both RI methods dominate the two standard approaches across all attrition rates.

\textbf{Strong matching (\(r = 0.80\)):} The paired \(t\)-test starts off with substantially higher power at \(q=0\), leveraging the strong matching. However, as attrition increases, the paired \(t\)-test again loses power due to attrition of complete pairs. The two-sample \(t\)-test's power is below 50\%. RI with difference in means also performs poorly as attrition increases, as it fails to efficiently combine information from pairs and singletons. RI with the optimally weighted estimator dominates all other methods across all attrition rates, leveraging both the strong matching and using all observed units.

It is worth noting that in practice, we may not know the exact matching quality in advance. However, the optimally weighted RI procedure is robust to this uncertainty, as it estimates the optimal weights from the data, thus adapting to the actual matching quality.

\begin{figure}
\subfloat[Weak matching benefit\label{fig:power-plots-1}]{\includegraphics[width=0.33\linewidth]{paper_files/figure-latex/power-plots-1} }\subfloat[Moderate matching benefit\label{fig:power-plots-2}]{\includegraphics[width=0.33\linewidth]{paper_files/figure-latex/power-plots-2} }\subfloat[Strong matching benefit\label{fig:power-plots-3}]{\includegraphics[width=0.33\linewidth]{paper_files/figure-latex/power-plots-3} }\caption{Power as a function of attrition}\label{fig:power-plots}
\end{figure}

\subsection{RI is well sized also with attrition}\label{ri-is-well-sized-also-with-attrition}

As randomization inference directly leverages the random assignment mechanism, it provides exact finite-sample size control under the sharp null hypothesis of no treatment effect, even in the presence of attrition, and without a need for large-sample approximations or distributional assumptions.

\textbf{Proposition 6 (Exact size control).} Under the sharp null hypothesis \(H_0: Y_{p,g}(1) = Y_{p,g}(0)\) for all \((p,g)\) and the assumption that attrition is independent of treatment assignment
conditional on the pair's full vector of potential outcomes \(\mathbf{Y}_p := (Y_{p,1}(0), Y_{p,1}(1), Y_{p,2}(0), Y_{p,2}(1))\), i.e.~\(A_{p,g} \perp D_{p,g} \mid \mathbf{Y}_p\), both RI procedures (standard and optimally weighted) reject with probability at most \(\alpha\) (exactly \(\alpha\) under a randomized tie-break) for any finite sample size \(n\) and attrition rate \(q\).

\emph{Proof.} See appendix \cref{proof:exact-size-control}.

\textbf{Remark (holds also with larger strata).} This argument does not rely on the pair structure and holds for any finely-stratified design with assignment-independent attrition. The estimator, weights, and power expressions above, by contrast, are specific to pairs.

\textbf{Remark (Implementation with feasible weights).} In implementing the feasible weighted RI procedure, the weights \(\hat{w}^*(\pi)\) must be recomputed for each permutation \(\pi\) using the variance estimates corresponding to that permutation's treatment assignment. Specifically, for each permutation, compute \(\hat{\sigma}_{Y_0}^2(\pi)\) using observations labeled as control under that permutation, derive \(\hat{r}(\pi)\) and \(\hat{w}^*(\pi)\) accordingly, and compute the test statistic using these permutation-specific weights. This recomputation ensures exact size control under the sharp null, as the permutation distribution properly accounts for the joint randomization distribution of both the treatment effect estimator and the nuisance parameter estimates. Holding weights fixed across permutations (computed only from the observed treatment assignment) would break exactness because the control group variance estimate \(\hat{\sigma}_{Y_0}^2\) is not permutation-invariant.

\textbf{Remark on asymptotic vs exact results:} Unlike the power results (Propositions 1-5), which rely on asymptotic normality via the Lindeberg--Feller CLT, Proposition 6 provides exact finite-sample size control under the sharp null without any distributional assumptions beyond the independence of attrition from treatment assignment.

\subsection{Weak-null inference}\label{weak-null-discussion}

The derivations above are exact under the sharp null \(H_0: Y_{p,g}(1) = Y_{p,g}(0)\) for all \((p,g)\).
We now consider the weak null \(H_0: E[Y_{p,g}(1) - Y_{p,g}(0)] = 0\), which permits arbitrary effect heterogeneity.
Under the weak null, permutation-based inference is not automatically valid: the permutation distribution of a test statistic need not match its true sampling distribution once treatment effects vary across units (Chung and Romano 2013; Ding and Dasgupta 2018; Wu and Ding 2021). The standard remedy is studentization, which restores asymptotic validity under the weak null while typically preserving exact size under the sharp null. We show below that in our setting this remedy is not always needed.

\subsubsection{Without studentization}\label{without-studentization}

The reason is that both test statistics (the within-pair difference as well as differences between singletons), as well as their {[}weighted{]} combination, are balanced in treatment. Consider the singleton component first. Treated and control singletons are equal in number, so the difference in means is a two-sample comparison with equal group sizes. For such a comparison the permutation variance and the true sampling variance coincide when the two groups are equal in size, even if their outcome distributions differ (Chung and Romano 2013, Example 2.1). The weak null is therefore tested at the correct level without studentization. The pair component of the test statistic is a within-pair sign flip of the differences \(\Delta_p = Y_{p,T} - Y_{p,C}\). Its permutation variance is \(E[\Delta_p^2]\) and its true variance is \(E[\Delta_p^2] - (E\Delta_p)^2\); these agree exactly when \(E\Delta_p = 0\), which is the weak null. The sign flip is balanced by construction, so no tuning is needed. The full estimator combines the two components, which use disjoint partitions of the data with independent assignments. The two permutation pieces are therefore independent and each matches its sampling counterpart under the null, so their weighted sum does as well.

Balance is what these arguments rely on.
For singletons, we assumed exactly equal numbers of treated and control singletons for simplicity. But also more generally, when attrition is independent of treatment, the two counts are equal in expectation and their ratio tends to one, so the equal-size condition holds in the limit and the singleton component is asymptotically valid regardless of the realized counts. For pairs, symmetry holds by construction.\footnote{This does not contradict the studentization requirement of Bai, Romano, and Shaikh (2022). In their framework, units are sampled i.i.d. and then paired on covariates, and under the weak null the within-pair sign-flip permutation distribution of the mean pair difference has limiting variance exceeding the estimator's sampling variance by \(\tfrac{1}{2}E[(E[Y_i(1)-Y_i(0)\mid X_i])^2]\): matching makes the two members of a pair share the same conditional average effect, which the estimator averages over all \(2n\) sampled units while the permutation treats each pair's difference as a single draw. In our model the pair itself is the i.i.d. draw, so the permutation variance of the pair component, \(E[\Delta_p^2]\), coincides with its sampling variance \(\mathrm{Var}(\Delta_p)\) whenever \(E[\Delta_p]=0\), whatever the within-pair dependence of the effects. The cost is that our model cannot express effect heterogeneity systematically related to the pairing index across pairs; where that is a concern, the studentized procedures below apply.}

\subsubsection{With studentization}\label{with-studentization}

If one prefers not to rely on this balance, expects attrition to differ across arms while remaining independent, or operates in a different sampling framework,
the estimator can instead be studentized by a consistent standard error, which restores weak-null validity for any group sizes while keeping exact size under the sharp null.\footnote{In general, consistency for the sampling variance is not sufficient for permutation tests. The variance estimator used for studentization must be consistent for the variance of the randomization distribution, which may be different (Bai, Romano, and Shaikh 2022, Remark 3.16); in the framework that we used here, the two limits coincide under the weak null, so this distinction has no bite.} This requires a variance estimator appropriate to the pairwise design. Consistent alternatives are discussed for different cases in the literature: Bai, Romano, and Shaikh (2022), adapted to attrition by Bai et al. (2024), construct one by grouping adjacent pairs, and Chaisemartin and Ramirez-Cuellar (2024) discusses when clustering at the strata level is appropriate. A reader who wants full robustness to heterogeneous effects and unbalanced attrition can adopt either estimator, studentize each component, and test the weak null in the same permutation framework.

\subsection{Optimally weighted RI outperforms the other methods for all parameter values}\label{optimally-weighted-ri-outperforms-the-other-methods-for-all-parameter-values}

\textbf{Proposition 7 (Dominance of optimally weighted RI).} For any \(n, q, \sigma^2_{Y_0}, r\) with \(q \in (0,1)\) and \(r \in [0,1)\), the oracle optimally weighted RI procedure has power greater than or equal to both the paired \(t\)-test and the two-sample \(t\)-test.

\emph{Proof.} See appendix \cref{proof:dominance-optimal-ri}.

Figure \ref{fig:advantage-plots} shows the advantage of the optimally weighted RI procedure over the two standard methods across varying attrition rates and matching quality. The advantage is measured as the percentage reduction in required sample size to achieve the same power as the optimally weighted RI procedure, when using the paired \(t\)-test or the two-sample \(t\)-test instead.

\begin{figure}
\subfloat[RI with optimal weights vs. Paired t-test\label{fig:advantage-plots-1}]{\includegraphics[width=0.33\linewidth]{paper_files/figure-latex/advantage-plots-1} }\subfloat[RI with optimal weights vs. Two-sample t-test\label{fig:advantage-plots-2}]{\includegraphics[width=0.33\linewidth]{paper_files/figure-latex/advantage-plots-2} }\subfloat[RI versus best of both\label{fig:advantage-plots-3}]{\includegraphics[width=0.33\linewidth]{paper_files/figure-latex/advantage-plots-3} }\caption{Advantage of RI with optimal weights (\%-reduction in required sample size to achieve target power) over standard methods}\label{fig:advantage-plots}
\end{figure}

\section{Simulation to compare pair designs with our method to minimally larger strata}\label{simulation-to-compare-pair-designs-with-our-method-to-minimally-larger-strata}

\subsection{Simulation: Paired vs.~Larger Strata Designs}\label{simulation-paired-vs.-larger-strata-designs}

The methodological guidance discussed above commonly recommends using larger strata (e.g., strata of size four) rather than pairs when substantial attrition is anticipated.
Our theoretical results establish that randomization inference with optimal weights dominates standard approaches within pairwise designs, but leave open whether pairwise designs analyzed with our methods can compete with larger strata designs analyzed with standard methods.
We investigate this through simulation.

Strata of size four offer advantages in a specific regime: attrition rates high enough that many pairs become singletons, but not so high that four-unit strata lose all within-stratum variation.
However, this advantage depends critically on the trade-off between matching quality and effective sample size.
When tight matching gains are substantial, pairwise designs may still dominate even with higher effective attrition.
When gains from matching are modest, larger strata may be better---but in those cases, the benefits of efficiently including singletons are also higher, thus improving the performance of our optimally weighted inference.

\subsection{Simulation design}\label{simulation-design}

We simulate data from the theoretical model of \Cref{sec:setting}, with additional structure to enable comparison across designs. We generate \(N = 2n\) units with a baseline covariate \(X_{p,g} \sim N(0, \sigma_x^2)\) and idiosyncratic noise \(\varepsilon_{p,g} \sim N(0, \sigma_\varepsilon^2)\), yielding control potential outcomes \(Y_{p,g}(0) = X_{p,g} + \varepsilon_{p,g}\). Treatment effects are constant: \(Y_{p,g}(1) = Y_{p,g}(0) + \tau\).

To ensure fair comparison, both designs are constructed from the same realized data by sorting units on \(X_{p,g}\). The pairwise design assigns consecutive units to pairs; the strata-of-4 design assigns consecutive groups of four to strata.
This is to mimic the decision a researcher would make when forming matched pairs or strata based on a single pre-treatment covariate that is a decent predictor of outcomes.
Treatment is randomized independently within each pair or stratum.
Attrition occurs independently at rate \(q\) per unit, independent of the employed design.
This construction isolates the pure design comparison: pairs achieve tighter matching (higher within-pair correlation in \(Y_{p,g}(0)\)), while pairs have a higher probability of losing all within-stratum variation due to attrition.

This setup imposes stronger assumptions than our theoretical analysis requires, e.g., normal distributions, a specific matching procedure, and serves only to demonstrate existence of parameter configurations where our proposed method's efficiency gains offset the attrition disadvantage of pairwise designs.

The chosen setup is also such that stratifying at a higher level has virtually no disadvantages absent attrition. We chose the parametrization such that the within strata variance of potential outcomes is almost identical irrespective of whether strata of 2 or 4 are formed. This is reflected by the fact that in the simulations below, all designs achieve similar power, when there is no attrition. We do this to show that the advantage of our method does not only exist when larger strata are disadvantaged anyways.

\subsection{Results}\label{results}

\begin{table}[!h]
\centering
\caption{\label{tab:sim_results}Simulated power comparison: Paired design with optimal RI vs. strata-of-4 design.}
\centering
\begin{tabular}[t]{rrrr}
\toprule
\multicolumn{1}{c}{ } & \multicolumn{3}{c}{Power} \\
\cmidrule(l{3pt}r{3pt}){2-4}
\multicolumn{1}{c}{ } & \multicolumn{2}{c}{Pairs} & \multicolumn{1}{c}{Strata-of-4} \\
\cmidrule(l{3pt}r{3pt}){2-3} \cmidrule(l{3pt}r{3pt}){4-4}
Attrition Rate & Pair FE & RI Optimal & Strata FE\\
\midrule
0.00 & 0.892 & 0.895 & 0.891\\
0.05 & 0.864 & 0.874 & 0.853\\
0.10 & 0.813 & 0.839 & 0.834\\
0.15 & 0.759 & 0.802 & 0.789\\
0.20 & 0.711 & 0.779 & 0.759\\
\addlinespace
0.25 & 0.662 & 0.735 & 0.716\\
0.30 & 0.607 & 0.713 & 0.689\\
0.35 & 0.560 & 0.671 & 0.645\\
0.40 & 0.503 & 0.628 & 0.570\\
\bottomrule
\end{tabular}
\end{table}


Table \ref{tab:sim_results} presents simulated power at \(\alpha = 0.05\) for three methods
(pairwise randomization design with paired \(t\)-test,
strata-of-4 design with strata-fixed effects,
pairwise design with optimally weighted randomization inference) across attrition rates. Parameters are \(n = 50\) pairs, \(\tau = 0.146\), \(\sigma_x = 0.130\), \(\sigma_\varepsilon = 0.224\) (chosen to yield approximately 90\% power for the pairwise design without attrition). With no attrition, all methods achieve almost exactly the nominal power of 90\%.
As attrition increases to 15\%, pair fixed effects drops to 76\% power while the strata-of-4 design maintains 79\%.
However, optimally weighted RI on pairs achieves 80\%, surpassing both.
At high levels of 30\% attrition, the paired \(t\)-test drops to 61\%, reflecting that in expectation only \((1-0.3)^2 = 49\%\) of pairs remain complete.
The strata-of-4 design with fixed effects maintains 69\%, still marginally below the optimally weighted RI on pairs at 71\%.

These results demonstrate that efficiency gains from tight matching in pairs, when properly exploited through optimal weighting of complete pairs and singletons, can still dominate the effective sample size advantage of larger strata in terms of power.
Researchers facing anticipated attrition need not abandon pairwise designs if they adopt the inference methods we propose. And researchers who are constrained to pairwise designs can still achieve substantial power gains by employing our optimally weighted randomization inference approach.

\section{Conclusion and discussion}\label{sec:conclusion}

Some practitioner advice cautions against pairwise randomization when attrition is expected. One argument in that direction is that loss of complete pairs when inference is based on a paired t-test will reduce power. We argue that the case against pairwise randomization rests, under ``ignorable'' attrition, on the choice of a suboptimal inference method rather than on a defect of the design itself. When attrition is independent of treatment, i.e., ``ignorable'', tests that use every observed unit (complete pairs as well as orphaned singletons) while still exploiting the paired structure, can be constructed and retain exact size under the sharp null at any attrition rate. Weighting complete pairs and singletons by their precision yields a test that dominates both the paired \(t\)-test and the two-sample \(t\)-test in power under standard assumptions. The procedure we propose is straightforwardly implemented as a weighted fixed-effects regression, in standard software.

Our analysis assumes homoscedasticity in the sense that pairs are exchangeable draws from a single distribution \(F\), so that the within-pair correlation \(r\) and marginal variance \(\sigma^2_{Y_0}\) are common across pairs. Applied researchers may not want to make this assumption. Relaxing exchangeability, by allowing pair-specific correlations \(r_p\) and variances \(\sigma^2_{Y_0,p}\), as would arise when match quality varies across the sample, would make the optimal weight pair-specific. We conjecture that a generalization of the dominance result (Proposition 7) holds under such heterogeneity, but leave this extension to future work.

A fruitful extension could be to generalize our approach to more experimental designs, such as stratified randomization with larger strata, and compare the efficiency gains across designs \emph{and} methods (e.g., inference with the corrected variance estimator of Bugni, Canay, and Shaikh (2018)).
We believe the result of these comparisons is not obvious. Without attrition, block size has no first-order effect on asymptotic efficiency of the unadjusted estimator for RCTs with a fixed treated fraction (Bai et al. 2025); attrition is precisely what makes the design comparison interesting.
When attrition becomes an issue, designs with larger strata can continue to exploit within-stratum variation from partially complete strata. So they might outperform pairs if the gains from using within strata variation outweigh the gain coming from tighter matching in pairs. This is context-specific. Formalizing this trade-off into generalizable rules could be valuable at the design stage of an RCT, but we leave it for future research.

\section*{References}\label{references}
\addcontentsline{toc}{section}{References}

\protect\phantomsection\label{refs}
\begin{CSLReferences}{1}{0}
\bibitem[#2]{ref-amropauly2017}
Amro, Lubna, and Markus Pauly. 2017. {``Permuting Incomplete Paired Data: A Novel Exact and Asymptotic Correct Randomization Test.''} \emph{Journal of Statistical Computation and Simulation} 87 (6): 1148--59. \url{https://doi.org/10.1080/00949655.2016.1249871}.

\bibitem[#2]{ref-athey2017econometrics}
Athey, Susan, and Guido W Imbens. 2017. {``The Econometrics of Randomized Experiments.''} \emph{Handbook of Economic Field Experiments} 1: 73--140.

\bibitem[#2]{ref-bai2022optimality}
Bai, Yuehao. 2022. {``Optimality of Matched-Pair Designs in Randomized Controlled Trials.''} \emph{American Economic Review} 112 (12): 3911--40.

\bibitem[#2]{ref-bai2024revisiting}
Bai, Yuehao, Meng Hsuan Hsieh, Jizhou Liu, and Max Tabord-Meehan. 2024. {``Revisiting the Analysis of Matched-Pair and Stratified Experiments in the Presence of Attrition.''} \emph{Journal of Applied Econometrics} 39 (2): 256--68.

\bibitem[#2]{ref-bailiushaikhtabordmeehan2023}
Bai, Yuehao, Jizhou Liu, Azeem M. Shaikh, and Max Tabord-Meehan. 2025. {``On the Efficiency of Highly Stratified Experiments.''} arXiv preprint.

\bibitem[#2]{ref-bai2022inference}
Bai, Yuehao, Joseph P Romano, and Azeem M Shaikh. 2022. {``Inference in Experiments with Matched Pairs.''} \emph{Journal of the American Statistical Association} 117 (540): 1726--37.

\bibitem[#2]{ref-baishaikhtabordmeehan2024primer}
Bai, Yuehao, Azeem M. Shaikh, and Max Tabord-Meehan. forthcoming. {``A Primer on the Analysis of Randomized Experiments and a Survey of Some Recent Advances.''} \emph{Journal of Political Economy Microeconomics}.

\bibitem[#2]{ref-bruhn2009}
Bruhn, Miriam, and David McKenzie. 2009. {``In Pursuit of Balance: Randomization in Practice in Development Field Experiments.''} \emph{American Economic Journal: Applied Economics} 1 (4): 200--232.

\bibitem[#2]{ref-bugnicanayshaikh2018}
Bugni, Federico A., Ivan A. Canay, and Azeem M. Shaikh. 2018. {``Inference Under Covariate-Adaptive Randomization.''} \emph{Journal of the American Statistical Association} 113 (524): 1784--96. \url{https://doi.org/10.1080/01621459.2017.1375934}.

\bibitem[#2]{ref-dechaisemartin2024clustering}
Chaisemartin, Clément de, and Jaime Ramirez-Cuellar. 2024. {``At What Level Should One Cluster Standard Errors in Paired and Small-Strata Experiments?''} \emph{American Economic Journal: Applied Economics} 16 (1): 193--212.

\bibitem[#2]{ref-chung2013exact}
Chung, EunYi, and Joseph P Romano. 2013. {``Exact and Asymptotically Robust Permutation Tests.''} \emph{The Annals of Statistics}, 484--507.

\bibitem[#2]{ref-dingdasgupta2018}
Ding, Peng, and Tirthankar Dasgupta. 2018. {``A Randomization-Based Perspective on Analysis of Variance: A Test Statistic Robust to Treatment Effect Heterogeneity.''} \emph{Biometrika} 105 (1): 45--56. \url{https://doi.org/10.1093/biomet/asx059}.

\bibitem[#2]{ref-donner2000design}
Donner, Allan, and Neil Klar. 2000. \emph{Design and Analysis of Cluster Randomization Trials in Health Research}. Arnold London.

\bibitem[#2]{ref-ekbohm1976}
Ekbohm, Gunnar. 1976. {``Comparing Means in the Paired Case with Missing Data on One Response.''} \emph{Biometrika} 63 (1): 169--72. \url{https://doi.org/10.1093/biomet/63.1.169}.

\bibitem[#2]{ref-fukumoto2022nonignorable}
Fukumoto, Kentaro. 2022. {``Nonignorable Attrition in Pairwise Randomized Experiments.''} \emph{Political Analysis} 30 (1): 132--41.

\bibitem[#2]{ref-glennerster2013running}
Glennerster, Rachel, and Kudzai Takavarasha. 2013. \emph{Running Randomized Evaluations: A Practical Guide}. Princeton University Press.

\bibitem[#2]{ref-imai2008}
Imai, Kosuke. 2008. {``Variance Identification and Efficiency Analysis in Randomized Experiments Under the Matched-Pair Design.''} \emph{Statistics in Medicine} 27 (24): 4857--73. \url{https://doi.org/10.1002/sim.3337}.

\bibitem[#2]{ref-imai2009}
Imai, Kosuke, Gary King, and Clayton Nall. 2009. {``The Essential Role of Pair Matching in Cluster-Randomized Experiments, with Application to the Mexican Universal Health Insurance Evaluation.''} \emph{Statistical Science} 24: 29--53.

\bibitem[#2]{ref-imbensrubin2015}
Imbens, Guido W., and Donald B. Rubin. 2015. \emph{Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction}. New York: Cambridge University Press. \url{https://doi.org/10.1017/CBO9781139025751}.

\bibitem[#2]{ref-mckenzie2022_matched_pairs}
McKenzie, David. 2022. {``Why i Am Now More Cautious about Using or Recommending Matched Pair Randomization and Like Matched Quadruplets Instead.''} April 19, 2022. \url{https://blogs.worldbank.org/en/impactevaluations/why-i-am-now-more-cautious-about-using-or-recommending-matched-pair-randomization}.

\bibitem[#2]{ref-vandervaart1998}
Vaart, A. W. van der. 1998. \emph{Asymptotic Statistics}. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge: Cambridge University Press.

\bibitem[#2]{ref-wuding2021}
Wu, Jason, and Peng Ding. 2021. {``Randomization Tests for Weak Null Hypotheses in Randomized Experiments.''} \emph{Journal of the American Statistical Association} 116 (536): 1898--1913. \url{https://doi.org/10.1080/01621459.2020.1750415}.

\end{CSLReferences}

\clearpage