Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
33,020 characters · 9 sections · 15 citation commands
Optimal Screening in Experiments with Partial Compliance
\makeatletter \patchcmd{\@maketitle}{ \@title}{\fontsize{12}{12}\selectfont\@title} \makeatother
\font\myfont=cmr12 at 16pt
\setcounter{footnote}{0}
Keywords: Instrumental variables, Compliance, Experimental Design, Median Bias, Statistical Power \thispagestyle{empty}
In experiments with partial compliance, low take-up rates can make instrument variables (`IV') estimates of the average treatment effect on the treated both biased and imprecise. In response, recent econometrics research has proposed alternative methods aimed at improving precision Spiess2021, Hazard2025. Such efforts are based on the insight that placing greater weight on complier populations can enhance efficiency. For example, Spiess2021 propose an alternative estimator that weighs each observation according to its estimated compliance; and Hazard2025 exploits covariate information to restrict estimation to a subpopulation with non-zero compliance.
This note studies a related, but distinct, problem. It considers optimal experimental design under partial compliance when experimenters can screen participants prior to randomization. This differs from previous studies, which propose alternative estimators that can be applied to existing IV studies Spiess2021, Hazard2025. By contrast, this note addresses experimental design.
Our main aim is to establish the conditions under which a screened 2SLS estimator dominates the standard 2SLS estimator in terms of both bias and power. We study this question in the standard setup where both the instrument and treatment are binary. To begin, we consider an idealized setting where the experimenter observes participant types (i.e. compliers, always-takers, and never-takers). In later sections, we discuss how types can be elicited in practice.
A first-order concern is that screening can change the target parameter. This can render comparisons between screened and unscreened estimators ambiguous. Thus, to begin, we establish a sufficient condition which guarantees that the screened 2SLS estimator identifies the same Local Average Treatment Effect (`LATE') as the unscreened 2SLS estimator. Specifically, the sufficient condition states that any screening mechanism which retains all compliers identifies the same LATE. The intuition is simple: because the LATE only identifies the treatment effect for compliers, any screening mechanism retaining all compliers must preserve the target parameter.
Under this sufficient condition, we then show that the optimal screening mechanism is to retain all compliers and screen out all non-compliers (i.e. always-takers and never-takers). This screening mechanism is `optimal' in the sense that it minimizes bias and maximizes statistical power among the set of screening mechanisms that retain all compliers. Given that this set also includes no screening, the optimal screened 2SLS estimator therefore dominates the standard, unscreened 2SLS estimator in terms of both bias and power. Intuitively, the 2SLS estimator only identifies the average treatment effect among compliers, so that non-compliers only contribute noise to the estimator. In effect, removing non-compliers therefore `strengthens' the instrument; and this more than offsets decreases in precision due to a smaller sample size. It is noteworthy that the optimal screened estimator does not exhibit a bias-variance trade-off, which can be can be present in alternative approaches Spiess2021, Hazard2025, as both bias and efficiency improve.
The theoretical results rest on the assumption that the experimenter observes participants' types, which is not true in practice. Thus, to practically implement the screening estimator, the experimenter must be able to accurately elicit types. To do this, we suggest that experimenters simply ask participants two questions, with the proviso that their answers will have no bearing on whether they will be offered treatment: (i) “Would you take treatment $(D_i=1)$ if offered $(Z_i=1)$?” and (ii) “Would you take treatment $(D_i=1)$ if not offered $(Z_i=0)$”?\footnote{In an RCT with one-sided non-compliance, it is only necessary to ask participants question (i) because treatment is not available to those not offered.}
If reported compliance matches actual compliance, then types are accurately identified and the optimal screening estimator is feasible. To test whether this is the case, we propose an approach based on results in Hull2024, which allow one to estimate the probability that stated complier status equals one conditional on actually being a complier.\footnote{More formally, let $D_i$ be binary treatment, $CM_i$ be a binary variable for true complier type, and $\widetilde{CM}_i$ be a binary variable that equals one for reported compliers. Results in Hull2024 show that an IV regression (with the full sample) of $\widetilde{CM}_i \times D_i$ on $D_i$ estimates $\mathbbm{P}[ \widetilde{CM}_i=1 | CM_i=1]$.}
This provides a test of the identification assumption that all compliers are retained under screening, and has implications for which estimator should be reported. For example, if the null hypothesis that $\mathbbm{P}[ \widetilde{CM}_i=1 | CM_i=1]=1$ is rejected, then researchers should use the standard, unscreened estimator because the screened estimator identifies a different LATE. Alternatively, failing to reject the null provides support for the use of the screened estimator, which will identify the same LATE and have better statistical properties as compared with the unscreened estimator in terms of both power and bias.
Finally, it is important to add that we recommend that experimenters actually implement `pseudo-screening', where they ask the screening questions, but retain the full sample for the experiment. This allows them to calculate both the screened and unscreened estimator, estimate the power gain from the screened design, and also perform stronger tests of the identification assumption, as discussed in more detail in the main text.
This note proceeds as follows. Section (ref) sets up the problem and presents the main theoretical results. Section (ref) shows how to feasibly implement the screened estimator and test for the reliability of stated compliance. Finally, Section (ref) concludes with a brief discussion of future work.
In this section, we describe the set up and present the main theoretical results. Proofs are in Appendix (ref).
The main aim is to compare the standard unscreened IV estimator to an alternative approach where experimenters screen out certain participants and then run the experiment on the remaining subset. We call this the screened IV estimator.
The theoretical results are proved under an idealized setting in which participant types are observed by the experimenter (i.e. complier, always-taker, never-taker). We derive an `optimal' screened estimator under this assumption. Since types are not observed in reality, we discuss how to implement this estimator in practice in Section (ref).
Consider the standard linear IV model with heterogeneous treatment effects: $$ Y_i = \alpha + \beta_i D_i + u_i $$ $$ D_i = \phi + \pi Z_i + \eta_i $$ where $\mathbbm{E}[D_i u_i] \neq 0$ and we assume, for simplicity, that errors are homoscedastic. Let $S_i$ be a binary variable that equals one if $i$ is screened into the sample, and zero otherwise. We impose the standard assumptions under the LATE framework Imbens1994:
Compared to the standard LATE assumptions, we include one additional assumption, namely, that the instrument is assumed to be independent of screening. This is immediately satisfied in the scenario we consider, where screening occurs prior to randomization.
The main aim is to compare median bias and power of the screened IV estimator, $\hat{\beta}_r$, and the unscreened IV estimator, $\hat{\beta}_1$, where the subscript denotes the fraction of the original sample that is retained. To make meaningful comparisons, we assume throughout that screening is neither trivial nor degenerate, in that at least one unit is screened out of the experiment and not all are screened out: $r \in (0,1)$.
A fundamental concern for alternative IV estimators is whether or not they change the target parameter. For example, Spiess2021 and Hazard2025 propose alternative estimators that can provide gains in precision, but, in general, target a different LATE to standard IV estimation. This makes direct comparisons challenging. To avoid this issue, we establish a sufficient condition under which screening does not change the target parameter:
Under this sufficient condition, the target parameter is unchanged:
In words, this result says that any screening procedure which retains all compliers will target the same LATE as the unscreened estimator (i.e. the average treatment effect on compliers). We maintain this assumption throughout the paper and propose a simple approach to test it in Section (ref).
This subsection considers the optimal screening mechanism conditional on retaining all compliers (Assumption (ref)). The mechanism we propose is optimal in two senses: first, that it maximizes statistical power; and second, that it minimizes median bias. This differs from alternative IV estimators in the literature, which typically pose a bias-variance trade-off Spiess2021, Hazard2025.
First, consider statistical power, for which we make the following assumption:
The variance of the 2SLS estimator depends on the average residual variance in the estimation sample; under homoscedasticity, this average is unaffected by screening, ensuring that any reduction in the standard error comes solely from the strengthened first stage rather than from changes in error variance.\footnote{This can be weakened to allow for heteroscedasticity across units, provided that the average variance is equal across screened and unscreened samples. An even weaker requirement is that the screened sample have weakly lower average residual variance, $\sigma_{u,r}^{2} \le \sigma_{u,1}^{2}$, which also guarantees improved precision.} With this, we can state the main result:
This result shows that any (non-trivial) screening mechanism that retains all compliers lowers the standard error, which corresponds to a lower MDE and higher statistical power.\footnote{Statistical power is increasing in the ratio of the true treatment effect and the standard error. Since the true treatment effect is unchanged under Assumption (ref), a lower standard error immediately implies higher statistical power.} Moreover, the power gain is inversely related to the fraction of the full sample that is retained, $r$. This implies that the power-maximizing screening mechanism minimizes $r$, which is achieved by screening out all never-takers and always-takers and retaining only compliers.
The intuition for Proposition (ref) is that LATE identifies the effect for compliers only, so removing non-compliers reduces noise and thus increases power by `strengthening the instrument'. This general insight is well understood in the existing IV literature Spiess2021, Hull2025, Hazard2025. The primary contribution of Proposition (ref) relative to existing work is therefore conceptual, in that it shows that this insight can be used to optimize experimental design; this differs from the complementary but distinct aim of improving precision in existing IV studies.
To gauge the potential practical importance of this insight, consider tarozzi2015, which studies the impact of microcredit in Ethiopia using an RCT with partial compliance. The first-stage shows that assignment to access to micro-credit increased borrowing rates by 25 percentage points; note that this also provides an estimate of the share of compliers in the population.\footnote{Formally, Assumption (ref) implies: $\pi = \mathbbm{E}[D_i | Z_i = 1] - \mathbbm{E}[D_i | Z_i = 0] = \mathbbm{E}[D_i(1) - D_i(0)] = \mathbbm{P}[CM_i=1]$.}
Now suppose experimenters effectively screen out all non-compliers from the estimation sample. Then Lemma (ref) implies that the target parameter is unchanged. Moreover, under the assumptions of Proposition (ref), the standard error for the screened estimator will be halved relative to the standard, unscreened estimator ($\sqrt{0.25} = 0.5$). Thus, confidence intervals and the MDE would also be halved, which represents an extremely large gain in precision for zero cost.
Next, we turn to bias, where we make the following assumptions:
Assumption (ref) imposes normality of errors for analytical convenience. Following Angrist2024, Assumption (ref) assumes that experimenters screen the estimated first-stage to have the same sign as the population first-stage. In practice, this is a relatively mild assumption.
To illustrate, consider again the example of a microcredit intervention. The population first-stage is generally thought to be positive ($\pi>0$), since those who are offered microcredit ($Z_i=1$) should have (weakly) higher take-up rates than those not offered ($Z_i=0$). Screening the estimated first-stage in this case means that if the estimated take-up rate were lower in the group offered microcredit, then the experimenter would discard this quantitative results from the experiment. This is likely to hold in practice because such an outcome would likely invalidate the study since it suggests an extremely weak instrument.
With this, we can state the result:
This results state that any screening mechanism which retains all compliers lowers median bias -- and also that median bias will be lowest when all never-takers and always-takers are excluded. The intuition is similar to Proposition (ref): since we only ever identify the impact on compliers, removing always-takers and never-takers effectively increases the strength of the instrument, which in turn lowers bias.
To summarize, this section makes a number of complementary claims. First, that a sufficient condition for maintaining the same target parameter in the screened and unscreened designs is to retain all compliers. Second, that conditional on satisfying this sufficient condition, screening out all always-takers and never-takers will minimize bias and maximize statistical power. Next, we turn to feasible implementation of this screening design.
In the preceding section, the theoretical results were derived assuming the experimenter can directly observe participants' types (e.g. complier, always-taker, never-taker) and screen accordingly. In reality, types are unobservable. This section discusses how to implement the optimal screening estimator in practice, and proposes an approach to partially test whether screening effectively identifies types.
We recommend experimenters screen for compliers by simply asking them their intended compliance. Thus, for RCTs with two-sided noncompliance, we propose adding two questions to the questionnaire prior to randomisation:
Responses to these two questions identify participants' stated type. For example, if a participant respond no to (i) and yes to (ii), then their stated type is a complier. The effectiveness of screening depends on whether participants' stated type match their true type, which we discuss further in the next subsection.
Before that, it is useful to introduce some notation. Let $\widetilde{CM}_i=1$ if $i$ is a stated complier based on the two questions above, and zero otherwise; and let $CM_i=1$ if $i$ is a true complier and zero otherwise. Additionally, we say that Type-I screening error occurs when an individual is a stated complier ($\widetilde{CM}_i=1$) but not a true complier ($CM_i=0$). Alternatively, Type-II screening error occurs when an individual is a stated non-compliers ($\widetilde{CM}_i=0$) but is in fact a true complier ($CM_i=1$).
Identification requires that all compliers are retained; in other words, that there is no Type-II error. That, under screening for stated compliers ($S_i=1 \iff \widetilde{CM}_i=1$), Assumption (ref) becomes $\mathbbm{P}[ \widetilde{CM}_i=1 | CM_i=1]=1$, which we discuss how to estimate in the next subsection. Note that Type-I error does not, by contrast, harm identification, although it does reduce the power gain because non-compliers included in the sample will increase the variance of the estimator.
Finally, note that in RCTs with one-sided noncompliance only question (1) needs to be added to the questionnaire; this is because the answer to (2) is `no' by design for all participants.
We recommend that experimenters implement what we call pseudo-screening. This is where screening questions are asked to elicit stated types, but where the experiment nevertheless is conducted on the full sample. Pseudo-screening means that experimenters can calculate both the IV estimate and standard error for both the screened sample and the full sample, and allows experimenters to measure the power gains from the screened design.
More importantly, however, is that pseudo-screening provides experimenters with greater variation for testing the identification assumption, namely, that screening successfully retains all compliers, which ensure the same LATE as the unscreened estimator (Assumption (ref)). To see this, consider the case experimenters do not pseudo-screen but actually screened out stated non-compliers ($\widetilde{CM}_i=0$). Suppose further that a subset of stated non-compliers are in fact compliers ($CM_i=1$); this represents Type-II error, which violates the identification assumption.
However, under true screening, this type of violation is undetectable because non-compliers are excluded from the experiment, and their behaviour is therefore never observed. By contrast, under pseudo-screening, a subset of stated non-compliers will be offered treatment, and their choices can be observed to test whether stated non-compliance matches non-compliance in reality.
To see how this can be formally tested, consider first the more general result in Proposition 1 in Hull2024. This result can be used to show that for any characteristic $X_i$, the 2SLS coefficient from an IV regression of $X_i \times D_i$ on $D_i$ will estimate the average of that characteristic in the complier subpopulation i.e. $\mathbbm{E}[X_i | CM_i=1]$.
Returning to our setting, let $\widetilde{CM}_i$ be a binary variable that equals one if $i$ reports the intention to comply, and zero otherwise. From the general result above, it follows that an IV regression of $\widetilde{CM}_i \times D_i$ on $D_i$ will estimate the probability that stated compliance equals one conditional on actually being a complier i.e. $\mathbbm{P}[ \widetilde{CM}_i=1 | CM_i=1]$. This corresponds exactly the probability in Assumption (ref) when the screening mechanism is to retain all stated compliers.
Thus, we propose a one-sided test of the null hypothesis that $\mathbbm{P}[ \widetilde{CM}_i=1 | CM_i=1]=1$, where the alternative hypothesis is $\mathbbm{P}[ \widetilde{CM}_i=1 | CM_i=1]<1$.\footnote{An alternative test of the identification assumption would be to estimate whether treatment take-up is zero for the `pseudo-screened-out' group (i.e. stated non-compliers). We opt for the test in the main text because it utilizes the full sample and is therefore more powerful.} Rejecting the null hypothesis suggests that stated compliance is not reliable. This might justify using of the standard, unscreened estimator, which is, of course, calculable under pseudo-screening. On the other hand, failing to reject the null hypothesis would favor using the screened estimator, as this is guaranteed to have better statistical power (Proposition (ref)) and lower bias (Proposition (ref)).
Validating that $\mathbbm{P}[ \widetilde{CM}_i=1 | CM_i=1]=1$ tests for Type-II screening error; and hence it provides a direct test on whether the LATE is preserved and that the screen estimator has better bias and power properties than the unscreened estimator. However, it does not imply that there is no Type-I screening error, and hence whether the optimal screening design has been achieved. That is, some participants may state being compliers ($\widetilde{CM}_i=1$) but may not in fact be compliers $CM_i=0$. This outcome might arise, for example, due to Hawthorne effects. More formally, it is possible that $\mathbbm{P}[ \widetilde{CM}_i=0 | CM_i =0] \neq 1$, even if $\mathbbm{P}[ \widetilde{CM}_i=1 | CM_i =1]=1$.
Thus, to test whether the optimal screening design was obtained, the experimenter can additionally test the null hypothesis that $\mathbbm{P}[ \widetilde{CM}_i=0 | CM_i=0] = 1$. To calculate this quantity, note that the law of total probability implies the following: \[ \mathbb{P}\!\left[\widetilde{CM}_i = 0 \mid CM_i = 0\right] = \frac{ \mathbb{P}\!\left[\widetilde{CM}_i = 0\right] - \mathbb{P}(CM_i = 1)\, \mathbb{P}\!\left[\widetilde{CM}_i = 0 \mid CM_i = 1\right] }{ 1 - \mathbb{P}[CM_i = 1] } \]
where all quantities on the right-hand side are directly estimatable.
Hence, if both nulls are not rejected, then this provides evidence that stated responses are reliable and that the optimal screening estimator is achieved.
This note examines optimal experimental design under partial compliance when experimenters can screen participants prior to randomization. It shows that screening for compliers maintains the same LATE as no screening but maximizes statistical power and minimizes bias. Back-of-the-envelope calculations based on actual RCTs suggest that power gains could be quite large e.g. halving confidence intervals.
However, the feasibility of the screened estimator depends on the reliability of eliciting compliance behaviour. Thus, we are pursuing ongoing work to test the feasibility and potential advantages of the optimal screening design over a standard compliance design.