Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
62,891 characters · 9 sections · 8 citation commands
Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect
\singlespacing
We begin with a stylized example that demonstrates how a design-based model of an experiment can reveal evidence beyond the average effect. In Section (ref), we formalize the model and present a likelihood function for the joint distribution of potential outcomes in the sample, which we use to construct maximum likelihood estimates. In Section (ref), we use our design-based likelihood in two real-world health applications. Section (ref) discusses our contributions to the literature. Section (ref) presents implications for econometric and applied research.
You are a doctor at a rural health clinic, and you have six patients who are smokers. You would like them to quit smoking, so you consider offering them a payment contingent on quitting. You are optimistic that the payments would induce some patients to quit. However, your psychologist friend warns you that payments can crowd out intrinsic motivations \citep*{gneezy2000}. Your economist friend tells you that traditional economic models assume away the possibility \citep*{bjorklund1987, imbens1994, heckman1999, vytlacil2002}. But you have taken an oath to “do no harm.” Even if the payment works on average, that is little comfort if some of your patients would actually be harmed. The more patients who would quit otherwise but would not because of your payment, the worse you would feel. For patients who would quit regardless, your concern is more modest---you would rather not waste the payments. Because the payment is contingent on quitting, you would not waste payments on any patients who would remain smoking regardless of the payment, but as their doctor, you would still feel bad that your intervention cannot reach them. You would really like to know the count of patients of each of the four types.
You know that experiments are a tool to reveal causal evidence on the average effect. You also know that with traditional inference, you can obtain Fr\'{e}chet bounds \citep*{boole1854, hoeffding1940, frechet1957} on the distribution of effects in the population from which the smokers in your clinic were drawn. However, since your responsibilities are to your own patients, you do not focus on the population. You would like to estimate the full distribution of effects in your fixed sample of patients.
You plan an experiment, and you are deliberate about the design of your randomization. You choose a “completely randomized’’ design in which exactly three patients will be assigned to the payment intervention. You implement this design by using a computer to effectively pick three names out of a hat, without replacement.
You run the experiment, and the results are promising in terms of the average effect. Two of the three patients in intervention quit smoking; only one of the three patients in control quits. The intervention mean is 2/3, and the control mean is 1/3, so the average effect of the payment intervention is an increase in the quit rate of 1/3.
Why did you observe the data you observed? What would you have seen under the other possible randomized assignments? Can the experiment reveal any evidence about effect heterogeneity? In more technical terms, what is the joint distribution of “potential outcomes” \citep*{neyman1923,rubin1974, rubin1977} in the sample? Counterfactual questions like these have attracted recent interest in the study of causal inference \citep*{gelman2013, pearl2018, imbens2020, dawid2022}.
We assume, as in design-based inference, that the potential outcomes of each person are fixed, and therefore the joint distribution of potential outcomes in the sample is fixed, before randomization. Following \citet*{angrist1996}, we classify patients into types based on their potential outcomes in intervention and control: “always takers” take up the desired health behavior---they quit smoking---in intervention and control (effect=0); “compliers” take up in intervention but not control (effect=1); “defiers” do not take up in intervention but take up in control (effect=-1); and “never takers” do not take up in intervention or control (effect=0). We specify the joint distribution of potential outcomes with the numbers of always takers, compliers, defiers, and never takers in the sample. The joint distribution of potential outcomes provides more information than the average effect; it provides the full distribution of effects and disambiguates the zero effects of always takers from those of never takers.
We demonstrate that a design-based model of your experiment can reveal evidence on the numbers of always takers, compliers, defiers, and never takers in your clinic, beyond the evidence revealed by the estimated average effect and the resulting Fr\'{e}chet bounds. We develop a novel visualization of potential outcomes to illustrate how in Figure (ref). We represent each of the six patients in the sample with a circle. The left half of each circle represents the potential outcome in the payment intervention, and the right half represents the potential outcome in the no payment control. Orange represents “takeup” of the desired health behavior---quitting smoking; white represents “no takeup,” and grey represents “unobserved.”
Rectangles with lines down the middle represent experiments with intervention on the left and control on the right. Inside each experiment, assignment to intervention reveals the potential outcome in intervention---the left of the rectangle reveals the left of the circle, which can be orange for takeup or white for no takeup---but the potential outcome in control is unobserved, so it is grey. Similarly, assignment to control reveals the potential outcome in control---the right of the rectangle reveals the right of the circle, which can be orange for takeup or white for no takeup---but the potential outcome in intervention is unobserved, so it is grey.
The data from your experiment, depicted in the box at the top of Figure (ref), show takeup among two of three patients in intervention and one of three patients in control. Therefore, the estimated marginal distribution of the potential outcome in intervention is 2/3 takeup, and the estimated marginal distribution of the potential outcome in control is 1/3 takeup. These two marginal distributions define the “estimated Fr\'{e}chet set.’’ We refer to all joint distributions of potential outcomes with the same marginal distributions as members of the same “Fr\'{e}chet set.”
The three rows of Figure (ref) give the three joint distributions of potential outcomes with the same marginal distributions as the estimated Fr\'{e}chet set: half circles in which the potential outcomes are unobserved in the data are filled in so that they have the same marginal distributions as the data. Four of the six circles have orange on the left (takeup occurs among 2/3 of patients in intervention), and two of the six circles have orange on the right (takeup occurs among 1/3 of patients in control). As shown in the three rows, there are three ways to combine the colored half circles into colored full circles, each yielding a different number of defiers (circles with white on the left and orange on the right). This depiction demonstrates how the estimated Fr\'{e}chet set determines the estimated Fr\'{e}chet bounds on defiers, which extend here from a lower bound of 0 to an upper bound of 2. Is there any variation in the design-based likelihood across the joint distributions of potential outcomes in the three rows?
The joint distribution of potential outcomes in the first row is consistent with the \citet*{imbens1994} assumption of “monotonicity” in the sample, whereby there can be compliers or defiers, but not both. Since the estimated effect of 1/3 is positive, you assume away defiers. You then deduce that the person who takes up in control must be an always taker and the person who does not take up in intervention must be a never taker. Assuming that the count of each type is the same in each arm, the sample includes two always takers and two never takers. The remaining two patients must be compliers, so the joint distribution of potential outcomes consistent with the monotonicity assumption includes two always takers, two compliers, and two never takers. However, you are uncomfortable assuming monotonicity because you are concerned that there could be defiers, and you would prefer to rely on evidence.
The joint distributions of potential outcomes in the next two rows include defiers. The one in the second row includes one always taker, three compliers, one defier, and one never taker. The one in the third row includes four compliers and two defiers. There is a rationale in the literature for assuming the distribution in the second row---it has one defier, the midpoint \citep*{li2019} of the Fr\'{e}chet bounds on defiers. It can therefore seem surprising that evidence yields a rationale for assuming the distribution in the third row with two defiers.
Each row in Figure (ref) depicts evidence from a “design-based model’’ that combines the design of the randomization with a model of potential outcomes to specify a data generating process and use it to generate all possible data. As depicted in the top of each row, there are 20 possible randomized assignments of six patients that result in three in intervention ($6 \text{ choose } 3 = 20$). By the design of your experiment, each of these possible randomized assignments has the same probability, so you can make the assumption that they each have the same probability very compelling through careful implementation.
Within each main row, each of the six patients in the experiment has a separate row. The person’s type and randomized assignment to intervention or control determines the data the experiment reveals, which can be counterfactual to what you observed. The assignments that generate the data you observed---2/3 takeup in intervention and 1/3 takeup in control---are in lighter shading.
The number of randomized assignments that generate the data you observed gives the numerator of the design-based likelihood of the joint distribution of potential outcomes, reported in the third column. Paraphrasing the baby board book “Statistical Physics for Babies” \citep*{ferrie2017}, which provides style inspiration for our visualization, physicists refer to the number of ways that you could have seen what you have seen as “entropy.” The design-based likelihood is equal to the entropy---the number of assignments that generate the data (shown in lighter shading)---divided by the number of possible assignments. The first row, with zero defiers, generates the data in eight of 20 possible assignments, so the likelihood is $8/20=40\%$. The second row, with one defier, generates the data in six of 20 possible assignments, so the likelihood is $6/20=30\%$. The last row, with two defiers, generates the data in 12 of 20 possible assignments, so the likelihood is $12/20=60\%$.
Why do some joint distributions of potential outcomes generate the data in more randomized assignments than others and therefore have higher likelihoods than others? The bottom half of each row sorts the six patients by type and separately shows how the always takers, compliers, defiers, and never takers must be randomized into intervention and control to generate the data. In the first row, there are three types---always takers, compliers, and never takers---and there must be a balance between intervention and control within each type to generate the data. In the second row, there are four types---always takers, compliers, defiers, and never takers---and there must be an imbalance between intervention and control within each type to generate the data. In the third row, there are only two types---compliers and defiers---and there must be a balance between intervention and control within each type to generate the data.
Likelihoods are higher when the patients who generate the data are as similar as possible in two senses: they belong to fewer types, and they have a counterpart of the same type in the opposite arm such that those of the same type are balanced between intervention and control. The patients who generate the data are most similar in the third row: they belong to the fewest number of types (only two), and those of the same type are balanced between intervention and control. Comparison of the first and third rows illustrates how having fewer types can increase the likelihood. In the first row, there are two ways to randomize one of two always takers into intervention ($2 \text{ choose } 1=2$), two ways to randomize one of two compliers into intervention ($2 \text{ choose } 1=2$), and two ways to randomize one of two never takers into intervention ($2 \text{ choose } 1=2$), so there are eight assignments that generate the data ($2 \times 2 \times 2=8$). In the third row, there are six ways to randomize two of four compliers into intervention ($4 \text{ choose } 2=6$) and two ways to randomize one of two defiers into intervention ($2 \text{ choose } 1=2$), so there are 12 assignments that generate the data ($6 \times 2=12$). Joint distributions of potential outcomes other than those in Figure (ref) can generate the data; the true joint distribution of potential outcomes need not preserve the estimated marginal distributions. For example, the likelihood that the true average effect is zero and your sample contains only always and never takers is $9/20=45\%$. However, by exhaustive grid search, you determine that the global maximizer of the likelihood is the one depicted in the third row, which contains four compliers and two defiers.
As a doctor considering the distribution with the maximum likelihood, you are thrilled that it implies that the payment intervention would induce four of your six patients to quit smoking, which is even more promising than what you would conclude based on the average effect and a monotonicity assumption. You are also encouraged that it implies that you do not waste any payments on patients who would have quit regardless and that there are no patients who would continue smoking regardless. However, you are concerned that it implies that the payment would induce two patients who would otherwise have quit to continue to smoke.
Here, we formalize the design-based model of an experiment presented in our stylized example. A sample consists of $n$ subjects, each of whom has a binary potential outcome in intervention $Y_I \in \{0, 1\}$ and a binary potential outcome in control $Y_C \in \{0, 1\}$, where $1$ represents takeup and $0$ represents no takeup. A subject's observed outcome depends only on their own potential outcomes and their inclusion in the intervention or control arm, ruling out network-type effects through a “no interference’’ \citep*{cox1958} or “stable unit treatment value’’ \citep*{rubin1980} assumption.
Subjects receive a random assignment $Z$ to intervention ($Z=I$) or control ($Z=C$), which is independent of a subject’s potential outcomes $(Y_I, Y_C)$. The observed outcome $Y$ is:
where $\mathbf{1}_{\{\cdot\}}$ is the indicator function.
Subjects in the sample belong to one of four types or “principal strata,’’ defined by combinations of potential outcomes \citep*{frangakis2002}. Let $\theta_{11}$ represent the number of always takers $(Y_I = 1, Y_C = 1)$, $\theta_{10}$ the number of compliers $(Y_I = 1, Y_C = 0)$, $\theta_{01}$ the number of defiers $(Y_I = 0, Y_C = 1)$, and $\theta_{00}$ the number of never takers $(Y_I = 0, Y_C = 0)$. The collection of these four counts ${ \boldsymbol{\theta} } = (\theta_{11}, \theta_{10}, \theta_{01}, \theta_{00})$ constitutes the joint distribution of potential outcomes in the sample, which we write as counts rather than shares to ease the following likelihood notation. The experimental data ${ \boldsymbol{X} } = (X_{I1}, X_{I0}, X_{C1}, X_{C0})$ consist of the counts of subjects who take up in intervention $X_{I1}$, who do not take up in intervention $X_{I0}$, who take up in control $X_{C1}$, and who do not take up in control $X_{C0}$.
We adopt a “design-based” approach, in which the joint distribution of potential outcomes ${ \boldsymbol{\theta} }$ is fixed, but unknown, and all randomness in the experimental data ${ \boldsymbol{X} }$ comes from the random assignment of subjects to either intervention or control. As shown in (ref), randomization implies a distribution for the experimental data ${ \boldsymbol{X} }$ given the joint distribution of potential outcomes ${ \boldsymbol{\theta} }$, yielding a likelihood expression. For an experiment with Bernoulli randomization (also known as simple randomization or independent and identically distributed randomization) or a completely randomized experiment, the likelihood is proportional to:
where the set $\mathcal{I}({ \boldsymbol{x} }, { \boldsymbol{\theta} })$ restricts $j$ such that the binomial coefficients remain well defined. This expression appears in \citet*{copas1973}. The likelihoods we derive in the stylized example of Section (ref) are also proportional to ((ref)). For the first distribution with zero defiers and the third distribution with two defiers in Figure (ref), there is a single term in the likelihood summation. For the second distribution with one defier, there are two.
In the previous section, we present a likelihood for the joint distribution of potential outcomes in the sample. Each joint distribution of potential outcomes implies marginal distributions of potential outcomes in the intervention and control arms. The marginal distribution of the potential outcome in intervention is the number of subjects for whom $Y=1$ in intervention:
The marginal distribution of the potential outcome in control is the number of subjects for whom $Y=1$ in control:
Given a pair of marginal distributions of potential outcomes in intervention and control, there are numerous joint distributions consistent with them. boole1854, hoeffding1940, and frechet1957 derive bounds on the possible copulas connecting marginal distributions of random variables into joint distributions. By applying these bounds to a pair of marginal distributions of potential outcomes, we can determine the set of all joint distributions of potential outcomes consistent with the pair of marginal distributions. We refer to the set of joint distributions of potential outcomes consistent with a given pair of marginal distributions as the “Fr\'{e}chet set’’ $\mathcal{F}(\theta_{1 \bullet}, \theta_{\bullet 1};n)$ of those marginal distributions:
In Figure (ref), the three rows comprise the Fr\'{e}chet set for $\theta_{1 \bullet} = 4$ and $\theta_{\bullet 1} = 2$. We index the distributions in a Fr\'{e}chet set by their number of defiers, $\theta_{01}$, which can take values in the range
These inequalities provide the so-called “Fr\'{e}chet bounds’’ on the number of defiers.
We can form estimates of the marginal distributions of potential outcomes in intervention and control using the observed takeup rates in intervention and control:
In our design-based model of an experiment, these are unbiased estimators of the sample marginal distributions of potential outcomes for a completely randomized experiment. In a standard sampling-based model of an experiment, dividing by $n$ yields consistent estimates of the population marginal distributions of potential outcomes.
It is well-known that, in the sampling-based model of an experiment, the takeup counts in intervention and control contain no information about the joint distribution of potential outcomes in the population beyond its marginal distributions. In (ref), we recreate this result by showing that the likelihood of the population-level joint distribution of potential outcomes is constant, conditional on the population-level marginal distributions. In other words, the likelihood function is flat within every population-level Fr\'{e}chet set.
In our design-based model of an experiment, however, the data can contain information about the sample joint distribution of potential outcomes beyond the sample marginal distributions. That is, the likelihood function can vary with the number of defiers within sample-level Fr\'{e}chet sets. Figure (ref) plots the likelihood values of each joint distribution of potential outcomes in the estimated Fr\'{e}chet set formed by the data $(x_{I1}, x_{I0}, x_{C1}, x_{C0})= (154, 152, 111, 195)$, a hypothetical version of the first empirical example we discuss in Section (ref). We index the joint distributions in this Fr\'{e}chet set by their number of defiers along the horizontal axis. The height of the bar is the likelihood of the joint distribution in the Fr\'{e}chet set with the specified number of defiers. Within this Fr\'{e}chet set, the likelihood is highest at the distribution with 222 defiers, the upper Fr\'{e}chet bound on defiers. As the number of defiers increases, the likelihood initially decreases, and then increases. If we normalize likelihoods in this Fr\'{e}chet set to sum to one, the smallest collection of likelihoods to contain 95% of the mass excludes 105 to 116 defiers, indicated by the lighter shading. Notably, it excludes 111 defiers, the midpoint \citep*{li2019} of the Fr\'{e}chet set.
One way to exploit the curvature in the design-based likelihood is to estimate the joint distribution of potential outcomes with maximum likelihood:
There are finitely many vectors of four integers that sum to the number of participants in the experiment $n$, so at least one value ${ \boldsymbol{\theta} }$ maximizes the likelihood. We solve this optimization problem through an exhaustive grid search over the finitely many joint distributions of potential outcomes consisting of $n$ subjects.
While maximum likelihood estimation is common in practice, we also provide three justifications in (ref), which begins in (ref) by characterizing our estimates as the output of a statistical decision rule using the framework of statistical decision theory. First, we demonstrate that our estimates vary systematically with the data, as we show in Figure (ref) and discuss in (ref). Second, in (ref), we further justify our estimates as Bayes optimal under the appropriate conditions, including a uniform prior. A Bayesian perspective also allows us to quantify the strength of evidence by constructing credible sets. Third, in (ref), we discuss how the maximum likelihood estimate is optimal based on the principle of maximum entropy \citep*{jaynes1957a, jaynes1957b}, which “assumes the least" \citep*{jaynes1968} about the joint distribution of potential outcomes conditional on the observed data.
We consider two published experiments with positive, statistically significant average effects on takeup of desired health behaviors and plausible defiers. Figure (ref) summarizes the results from all possible empirical applications in experiments with given sizes, illustrating empirical conditions under which our maximum likelihood estimates yield defiers. It is predictable from the data for each application and the patterns in Figure (ref) that our first application yields no defiers and our second application yields defiers, which we confirm with exhaustive grid search. Our focus here is to provide examples to illustrate how applied researchers can present and interpret quantitative results and to provide substantive insights. The top of Table (ref) summarizes the context and design for both applications. The statistics in the table can be calculated using our Python- and Stata-compatible dbmle package, documented at \href{https://pypi.org/project/dbmle/}{https://pypi.org/project/dbmle/} \citep*{christy2025dbmle}.
As in our stylized example, in tappin2015, the definition of “takeup’’ is to quit smoking, the intervention involves payment, and the design is completely randomized with half assigned to intervention. The sample includes 612 pregnant smokers. Monotonicity assumptions have a long history in the context of interventions that induce pregnant women to quit smoking. In a study prominently discussed by \citet*{angrist1996}, \citet*{permutt1989} propose an early monotonicity assumption requiring that women who would quit smoking in control would also quit smoking in an arm that received a multifaceted intervention. However, the intervention in tappin2015 involves payment, and as the psychologist friend warned in the introduction, literature since the monotonicity assumption was proposed expresses concern that payments can backfire in many contexts, including on-time pickup of children from day care \citep*{gneezy2000}, blood donation \citep*{mellstrom2008, lacetera2010}, and vaccination \citep*{schneider2023}. Defiers who would quit smoking in control but not in intervention are plausible in tappin2015 because a payment could backfire by crowding out the intrinsic motivations of pregnant women to quit smoking based on moral ideals and health goals.
The average effect of the payment relative to usual care is a statistically significant 14 percentage point increase in the smoking quit rate on a base of 8% in control. Our maximum likelihood estimates reveal evidence beyond the average effect: $52/612=8\%$ always takers (effect=0), $86/612=14\%$ compliers (effect=1), $0/612=0\%$ defiers (effect=-1), and $474/612=77\%$ never takers (effect=0). The weighted average effect across the four types ($8 \times 0+14 \times 1+0 \times (-1)+77 \times 0$) is equal to the estimated average effect of 14 percentage points, which is also equal to the share of compliers. We would obtain the same results if we assumed monotonicity in the sample, so our estimates provide evidence in favor of monotonicity without assuming it.
Returning to the concerns of the doctor in the introduction, the maximum likelihood estimates are reassuring. No patients would be harmed by payment---the doctor's greatest concern. The 8% who are always takers represent a modest waste, although the doctor is not too disappointed about giving payments to pregnant women who quit smoking. The 77% of patients who are never takers cannot be reached by the intervention, which is disappointing but not costly given that payment is contingent on quitting. The 14% who are compliers are exactly the patients the doctor hoped to help.
Our design-based perspective affects the interpretation of our results. The doctor in the introduction cares about the patients in the sample, and our design-based likelihood is informative because it conveys information about the sample at hand. However, if the doctor were interested in the effects of a payment policy in the population from which the sample was drawn, the sampling-based likelihood we derive in (ref) would reveal no information about defiers in the population beyond the Fr\'{e}chet bounds---at least, not without a departure from standard sampling assumptions, which we explore in (ref).
We emphasize that even though our maximum likelihood estimates provide evidence that there are no defiers in the sample, considerable uncertainty remains. As discussed in (ref), we compute a 95% smallest credible set on the joint distribution of potential outcomes, and we report the ranges for the numbers of always takers, compliers, defiers, and never takers in that set in Table (ref). The 95% smallest credible set includes distributions with 0 to 71 defiers, representing 0% to 71/612=12% of the sample. The magnitudes of the values in this 95% smallest credible set help us to highlight that evidence on defiers is weak when the estimated average effect is positive.
Building on the evidence here, researchers could go on to make the stronger assumption of monotonicity and proceed as usual. With additional data on covariates, they could calculate average characteristics of always takers, compliers, and never takers \citep*{imbens1997, katz2001, abadie2002, abadie2003}. With additional data on a second stage outcome like birth weight, they could interpret the resulting instrumental variable estimate as a local average treatment effect on compliers \citep*{imbens1994}.
Note that the joint distribution of potential outcomes can change dramatically even when the average effect does not. If the control takeup rate in tappin2015 were $111/306=36\%$ instead of 8%, with the same average effect of 14 percentage points (takeup among an additional 43 people in intervention), the maximum likelihood estimates would include 36% defiers rather than zero. Note that these are the same data used for Figure (ref) at the end of Section (ref). Our second application illustrates the estimation of a positive count of defiers with real data.
Our second application, \citet*{johnson2003}, is a high-profile experiment in which a simple change from an opt-in default to an opt-out default increases organ donation almost twofold from 43% in control to 82% in intervention. Results from \citet*{johnson2003} have inspired state policy changes to encourage life-saving organ donations to combat chronic organ shortages \citep*{kessler2025}. However, in a recent experiment by \citet*{kessler2025}, in which takeup represents actually becoming an organ donor instead of hypothetically becoming an organ donor, the average effect is close to zero. Especially given the hypothetical outcome in \citet*{johnson2003}, both compliers and defiers are plausible. Compliers may think of the default as an endorsement by the experimenter and thus follow it regardless of what it is. Defiers may think of a default as a threat to their autonomy and thus go against it regardless of what it is. For this application, the population from which the sample is drawn has clear policy relevance, so we begin by engaging with standard sampling-based results.
Given the plausibility of defiers in this context, a monotonicity assumption might not be reasonable, but there is no obvious alternative, so standard practice is to engage with bounds on defiers, which can be large. As shown at the bottom of Table (ref), in which we report auxiliary statistics, the smallest possible number of defiers in the sample is 0 because it is always possible for all people who take up to be always takers and all people who do not take up to be never takers, as in the null of Fisher’s exact test. The largest possible number of defiers in the sample occurs if all $11(=61-50)$ people who do not take up in intervention and all 23 people who take up in control are defiers, yielding a maximum defier share of $30\% ( = (23+11)/115$). The estimated Fr\'{e}chet bounds are tighter because they incorporate the estimated marginal distributions of potential outcomes in each arm using the observed takeup rates. To calculate the upper Fr\'{e}chet bound from ((ref)), we consider the bounds on the share of defiers implied by the takeup rates in intervention and control. The intervention takeup rate implies that the share of defiers in intervention and thus in the sample is at most $18\%(=100\%-82\%)$, and the control takeup rate implies that the share of defiers in control and thus in the sample is at most 43%, so the share in intervention binds, and the upper Fr\'{e}chet bound is 18%.
The Fr\'{e}chet bounds are estimated with error, so there is even greater uncertainty in the number of defiers. The number of defiers in our 95% smallest credible set ranges from 0 to 34, which represents 0% to 30% of the sample of 115. The 95% confidence interval on the population fraction of defiers from \citet*{imbens2004}, incorporating the adjustment recommended by \citet*{jun2023} with pretesting at the 0.01 level, also accounts for estimation of the Fr\'{e}chet bounds, so we report it for comparison in Table (ref). It is slightly narrower, extending from 0% to 27%, but still includes the estimated Fr\'{e}chet bounds. Accounting for estimation of the Fr\'{e}chet bounds is empirically important. As demonstrated in Table (ref), if we do not account for estimation of the Fr\'{e}chet bounds and instead construct a 95% smallest credible set within the estimated Fr\'{e}chet set, variation in the likelihood allows us to exclude 9 defiers---a count in the middle of the Fr\'{e}chet bounds translated into whole numbers from 0 to 21.
The maximum likelihood estimates indicate that the sample includes $21/115=18\%$ defiers, which is equal to the estimated upper Fr\'{e}chet bound. The maximum likelihood estimates also include $28/115=24\%$ always takers, $66/115=57\%$ compliers, and no never takers. The maximum likelihood estimates are the same as the estimate that would be obtained under an assumption of no never takers. Under an assumption of no defiers, the estimated average effect of 39% is entirely due to compliers, who represent 39% of the sample. The maximum likelihood estimates preserve the estimated average effect ($39\%=57\% \times (1)+18\% \times (-1)$), but they tell a very different story. Compliers representing 57% of the sample are affected in one direction and defiers representing 18% of the sample are affected in the other. The total share of people affected of $75\%=(57\%+18\%)$ is almost two times as large as the number of people affected under monotonicity.
In this application, evidence from the design-based model suggests a weaker alternative assumption to monotonicity---an assumption of no never takers---that can subsequently be imposed. Monotonicity can, of course, still be imposed with acknowledgment that it is a stronger assumption. While the assumption of no never takers is still strong, by the principle of maximum entropy, it “assumes the least’’ \citep*{jaynes1968} about the joint distribution of potential outcomes. From the design-based likelihood in ((ref)), the assumption of no never takers can generate the data in at most $(28 \text{ choose } 50-(66-(54-23))) \times (66 \text{ choose } 66-(54-23)) \times (21 \text{ choose } (61-50))=8.5 \times 10^{31}$ randomized assignments representing 3.4% of the total randomized assignments ($(115 \text{ choose } 61)=2.5 \times 10^{33}$). The assumption of no defiers can generate the data in at most $(49 \text{ choose } 26) \times (45 \text{ choose } 24) \times (21 \text{ choose } 11)=7.8 \times 10^{31}$ randomized assignments, representing 3.1% of the total randomized assignments. An ex post justification for the assumption of no never takers is that though experimental subjects might feel comfortable complying with or defying defaults imposed by the experimenter, they might feel uncomfortable revealing to the experimenter that they would not donate their organs under either default, especially since they do not know when they are asked a hypothetical question if a later question will ask them for a response under the other default. If an assumption of no never takers is not palatable, evidence for defiers could be used to motivate methods for examining effect heterogeneity, like the collection of additional data or the use of machine learning methods.
If a second stage outcome such as actual organ donation is available, under an assumption of no never takers, the instrumental variable estimate is no longer interpretable as an average treatment effect on compliers without further assumptions. As discussed by \citet*{angrist1996}, under an additional assumption that average treatment effects are the same for compliers and defiers, the instrumental variable estimate is again interpretable as a LATE. \citet*{dechaisemartin2017} and \citet*{tchetgen2024} propose alternative assumptions to interpret the instrumental variable estimand in the presence of compliers and defiers.
Although the design-based maximum likelihood estimates target the distribution of potential outcomes in the sample rather than the population, an evidence-based assumption of no never takers in the sample can facilitate learning about defiers in the population if data are available on covariates. Under the assumption of no never takers, it is possible to label some specific people in the sample as compliers and others as defiers: all people in intervention who do not take up must be defiers, and all people who do not take up in control must be compliers. This reasoning allows us to make claims that the intervention had a causal effect on some specific people in the sample. Using terminology from \citet*{pearl1999}, under an assumption of no never takers, the intervention was a “necessary’’ cause of no takeup for all 11 people assigned intervention who do not take up, and it would have been a “sufficient’’ cause of takeup for all 31 people assigned control who do not take up. Under a monotonicity assumption, it is not possible to directly label any specific people as compliers or defiers because people who take up in intervention can be compliers or always takers, and people who do not take up in control can be compliers or never takers. An assumption that allows us to label some specific people as compliers and others as defiers could be an important first step to learn their characteristics, improve targeting in future samples, and obtain a larger average effect.
Our primary insight is that a randomized experiment can reveal evidence on an estimand more primitive than the average effect: the joint distribution of potential outcomes in the sample. The evidence can support a monotonicity assumption or inform a specific alternative assumption. We obtain evidence using the structure of the implemented randomization process to construct a design-based causal model and its resulting likelihood function. Our work contributes to the literature on design-based econometrics, statistical decision theory, and causal inference without monotonicity.
In his 1935 book on the design of experiments, Fisher notes that the design of Charles Darwin’s experiments was “greatly superior” to the “statistical methods available at the time” \citep*{fisher1935}. In their 2017 handbook chapter on the econometrics of experiments, Susan Athey and Guido Imbens write that they “recommend using statistical methods that are directly justified by randomization, in contrast to the more traditional sampling-based approach that is commonly used in econometrics” \citep*{athey2017}. The sampling-based approach assumes that people in an experiment are randomly sampled from a population, and it attempts to learn about the population. The design-based approach, which we adopt, attempts to learn about the sample of people in the experiment, which is useful since participation in experiments and clinical trials can be far from random \citep*{alsan2024} and some samples do not come from a meaningful population \citep*{abadie2020}.
Our use of a design-based model rather than a sampling-based model enables our contributions. Our ability to learn about defiers with a design-based likelihood may seem paradoxical because it is well known that the traditional sampling-based likelihood does not vary with the number of defiers within the Fr\'{e}chet bounds determined by the estimated marginal distributions of potential outcomes in intervention and control, as we review in (ref). Furthermore, it is straightforward to show that the Fr\'{e}chet bounds on defiers always include zero if the estimated average effect shows compliers on average. A large literature has focused on specifying what we can learn from the nonparametric Fr\'{e}chet bounds in the traditional sampling-based model (\citealt*{balke1997, heckman1997, manski1997mixing, tian2000, zhang2003, imbens2004, fan2010, mullahy2018, li2019, bai2024, semenova2024}).
Our use of the likelihood implied by the design-based model is central to our contributions. The design-based model defines counterfactuals using causal models of potential outcomes from \citet*{neyman1923}, \citet*{welch1937}, \citet*{kempthorne1952}, \citet*{copas1973}, \citet*{rubin1974, rubin1977}, \citet*{greenland1986}, \citet*{holland1986}, and others, and then uses the randomization structure to derive the design-based likelihood. While \citet*{copas1973} derives the design-based likelihood and uses it to test hypotheses about the average effect and the Fisher null hypothesis, our advance is to use the likelihood to learn about a more primitive object: the joint distribution of potential outcomes in the sample, including the number of defiers. We revisit Copas' design-based model in light of modern causal models, modern assumptions about the absence of defiers, and modern computational advances. Likelihood-based models can provide exact distributions of experimental data, test statistics, and confidence intervals, but they remain uncommon relative to simulation methods in randomization inference. One exception is \citet*{ding2019}, who use the design-based likelihood to examine the sensitivity of inference on aggregate parameters to assumptions about the joint distribution of potential outcomes. \citet*{rigdon2015} and \citet*{li2016} construct exact confidence intervals without the design-based likelihood. We use the design-based likelihood for estimation as well as inference.
While the design-based model provides evidence about defiers beyond the Fr\'{e}chet bounds in the sample, it is important to consider what evidence, if any, it can provide about the population from which the sample was drawn. Sometimes, evidence about the population could be more relevant for policy than evidence about the sample. Under standard sampling assumptions, the experiment does not provide evidence about defiers beyond the Fr\'{e}chet bounds in the population. While standard sampling models assume that the data consist of independent and identically distributed random draws from a population, this assumption is often far from accurate---especially for small, finite populations sampled without replacement. When the population is finite and the sample is drawn without replacement, departure from standard sampling assumptions can propagate information about defiers in the sample to the population. Suppose that the 6 people in our stylized example were randomly drawn without replacement from a finite population of 24 people to participate in a pilot study. As we show in (ref), it is possible to derive an alternative sampling-based likelihood of the joint distribution of potential outcomes in the finite population under sampling without replacement, and it can reveal evidence within the Fr\'{e}chet bounds. In our stylized example, the maximum likelihood estimate for the finite population of 24 based on an experimental sample that was one fourth of the population includes 16 compliers and 8 defiers, exactly four times the estimate in the sample. In terms of likelihood ratios, the evidence revealed within the finite sample is weak, and the evidence revealed within the finite population is even weaker. Furthermore, as the finite population grows, the evidence gets weaker until the finite population becomes infinite, and there is no evidence within the Fr\'{e}chet bounds. Nevertheless, weak evidence can improve decision making over no evidence at all.
Our engagement with statistical decision theory in the style of \citet*{wald1949}, rather than classical hypothesis tests, facilitates our progress. Classical hypothesis tests specify a null hypothesis and place the burden of evidence on the alternative hypothesis. This asymmetric treatment of Type I and Type II errors can imply an unrealistically conservative decision maker \citep*{tetenov2012} and exclude statistical evidence that, while weak, is still informative. To make use of weak evidence from our design-based likelihood, we focus on the statistical decision problem of choosing the correct joint distribution of potential outcomes in the sample, rather than on testing hypotheses about a prespecified distribution under prespecified size and power. We do not claim to provide a “test’’ of a monotonicity assumption. We provide a “decision’’ about the joint distribution of potential outcomes that can provide evidence for or against a monotonicity assumption.
Our use of design-based statistical decision theory contributes to the integration of statistical decision theory into econometrics \citep*{manski2004, dehejia2005, manski2007, hirano2008, hirano2009, stoye2012, kitagawa2018, manski2018, manski2019, manski2021, fernandez2024}. Econometrics is a natural discipline to embrace statistical decision theory because statistical decision theory has such a tight relationship with economic theory. Statistical decision theory is natural to apply in our design-based setting because it does not require large sample assumptions or asymptotic approximations. Our use of design-based statistical decision theory departs from “asymptotic analysis of statistical decision rules in econometrics” as reviewed in \citet*{hirano2020} in favor of the literature that focuses on statistical decision theory in finite sample settings \citep*{canner1970, manski2007admissible, schlag2007, stoye2007, stoye2009, tetenov2012}.
We justify our maximum likelihood estimates with justifications commonly employed in statistical decision theory, starting with the justification that our estimates vary systematically with the data, as we show in Figure (ref). One challenge with statistical decision criteria such as minimax and minimax regret is that researchers have shown instances in which they yield rules that do not vary with the data \citep*{schlag2003, manski2004, hirano2009, stoye2009}. In his influential introduction to statistical decision theory, \citet*{ferguson1967} recognizes that because maximum likelihood rules vary with the data, they are “reasonable” because they are better than “just guessing.” “However,” he continues, “decision theory, as developed here, is devoted to the problem of finding optimal rules, so that we do not refer to maximum likelihood estimates [...] unless they turn out naturally to be optimal in some sense.”
In (ref), we further justify our maximum likelihood estimates as Bayes optimal under the appropriate conditions, which advances Bayesian decision theory \citep*{chamberlain2011} by incorporating the design-based likelihood. Bayes optimality implies that our rule is admissible---that is, there is no alternative rule that performs at least as well for all possible joint distributions of potential outcomes, and strictly better for some \citep*{ferguson1967}. Furthermore, it ensures that our rule cannot be bested in a betting framework \citep*{freedman1969}. The Bayesian interpretation of our estimates facilitates our auxiliary construction of credible sets on the joint distribution of potential outcomes in Table (ref), providing an exact Bayesian alternative to classical asymptotic inference on Fr\'{e}chet bounds \citep*{manski1992, horowitz2000, tamer2004, jun2023} and the number of defiers within the Fr\'{e}chet bounds \citep*{imbens2004}.
In addition to standard decision theoretic justifications, we also offer a perspective that our design-based maximum likelihood estimates are optimal through the principle of maximum entropy \citep*{jaynes1957a, jaynes1957b}, which does not require specification of a prior. \citet*{golan2002} reviews a literature in econometrics that uses the principle of maximum entropy to abstract away from functional form assumptions on the likelihood function. In our setting, the design of the randomization, which is under the control of the experimenter, determines the functional form of the likelihood function, which also determines the functional form of the entropy.
Our results provide evidence that can support a monotonicity assumption or a specific alternative, which can be useful because monotonicity assumptions have become widespread, but they are not always realistic. Influential work by \citet*{imbens1994} proposes a monotonicity assumption that assumes away compliers or defiers in the first stage of an instrumental variable model, and \citet*{manski1997monotone} proposes an analogous assumption on treatment response. Monotonicity assumptions are ubiquitous in the analysis of first stages, and they are gaining traction elsewhere \citep*{alsan2025}. First stage monotonicity assumptions are useful for the interpretation of instrumental variable estimates as local average treatment effects on compliers, and we view them as reasonable in many contexts. However, they can be difficult to defend in some contexts, particularly in health care settings that motivate our work because interventions may be helpful for some but harmful for others. We interpret our binary outcome as “takeup” in our stylized example and applications so that we can use the familiar terminology of always and never “takers” developed for the first stage \citep*{angrist1996}, but the design-based likelihood is valid for other binary outcomes, like survival. Given concerns about side effects in medicine, we see potential for useful applications that have a single binary outcome that represents survival. In those applications, evidence of defiers could motivate researchers to reduce the dosage of a potentially toxic intervention.
Previous approaches have revealed evidence that violates monotonicity assumptions, but only with additional data beyond our binary intervention and outcome. Machine learning approaches of \citet*{wager2018} and \citet*{semenova2024} require data on covariates to reveal heterogeneous effects. Analysis of side effects in medicine requires data on secondary outcomes (e.g. bernard2001). Specification tests of instrumental variable model assumptions can reveal evidence against monotonicity in the first stage \citep*{imbens1997, richardson2010, huber2012, huber2015testing, kitagawa2015, mourifie2017, machado2019}, and marginal treatment effect methods can reveal evidence against monotonicity in the second stage \citep*{bjorklund1987, heckman1999, kowalski2023behavior, kowalski2023reconciling}, but such approaches require the additional structure of a two-stage model as well as data on both stages. Evidence against first stage monotonicity by \citet*{chan2022} requires a multi-valued instrument and the ability to observe a second stage outcome for one value of the first stage outcome. We contribute to the literature on violations of monotonicity by demonstrating that, even without additional data, our design-based model of an experiment can reveal evidence for or against monotonicity.
We view the development of our visualizations as a secondary contribution. Figure (ref) allows us to depict how the same data can arise from different distributions of potential outcomes within a Fr\'{e}chet set. Our visualization of how the MLE varies with the data in Figure (ref) allows us to depict all possible results from an experiment of a given size, demonstrating that we have not engaged in selective reporting. The visualization allows us to characterize patterns that would be difficult to infer from isolated empirical applications and demonstrates the value of statistical decision theory over testing.
We also offer substantive insights in real empirical applications. In one, a payment intervention to induce pregnant women to quit smoking, we find no defiers. In another, a nudge intervention to get more people to hypothetically donate their organs, we find both compliers and the upper Fr\'{e}chet bounds on defiers. In both applications, the weak evidence that we find can motivate stronger assumptions that researchers can use to learn more.
We see our demonstration that a design-based model can reveal evidence beyond the average effect with very minimal assumptions and data as a proof-of-concept that could pave the way for the production of stronger evidence that will achieve the ultimate goal of personalizing decisions. Even if the average effect is the decision-relevant parameter in some contexts, learning about heterogeneity that underlies the average effect is a first step toward targeting that can ultimately yield a larger average effect. In our second empirical application, under the maximum likelihood joint distribution of potential outcomes, which includes no never takers, we can deduce that the people who are untreated in intervention must be defiers, and the people who are untreated in control must be compliers. Under monotonicity, it is never possible to deduce that a given individual is a complier, and defiers are assumed away. Our hope is that the ability to look directly at some compliers and defiers under alternative assumptions motivated by our estimates will pave the way to characterize them, target interventions toward compliers and away from noncompliers, and improve the average effect of future interventions.
Our findings have already inspired new research from statisticians and econometricians. In a previous version of our paper \citep*{christy2024a}, we engage directly with utility functions that depend on the joint distribution of potential outcomes, motivated by our work on mammograms, which can help some women but harm others through overdiagnosis \citep*{kowalski2023behavior, kowalski2023reconciling}. Our work contributes to the literature on “asymmetric counterfactual utilities’’ \citep*{li2019,mueller2023,benmichael2024}. Alongside the literature, it has prompted engagement from \citet*{koch2025} and inspired \citet*{gelman2025}. Our work has also inspired \citet*{chen2026}, which shows that the joint distribution of potential outcomes is formally identified using the design-based likelihood. Despite identification, like us, they emphasize that the potential for learning is limited. They find that there exist Bayesian priors that do not update about monotonicity under any realization of the data. In (ref), we demonstrate that while such priors exist, they are rare, and those they construct assign zero prior weight to many distributions, violating Cromwell's rule \citep*{lindley1985making}. A challenge for future research is to explore how best to use weak evidence on defiers in the sample and how to make it stronger.
Our work suggests several areas for further methodological research. Design-based likelihoods can be derived for more complex experimental settings, including those with two stages, multi-valued variables, and alternative randomization designs, such as matched pair, stratified, randomized permuted block, and designs that allow for rerandomization. A promising direction is applying design-based likelihoods to parameters traditionally studied under the lens of partial identification, such as the average treatment effect in experiments with imperfect compliance, missing data, attrition, spillovers, and measurement error. Another is using design-based likelihoods and entropy to motivate alternative parametric assumptions in auctions \citep*{jun2024} and other designed markets \citep*{abdulkadiroğlu2017}. Approximation methods and advances in algorithmic efficiency could facilitate calculations in a wide array of samples. Optimal estimators under alternative priors and utility functions could be implemented with straightforward modifications. Estimators could also be developed for alternative estimands, like asymmetric counterfactual utilities, measures of discrimination from audit studies \citep*{kline2021}, and potential outcome quantiles \citep*{cui2023,guggenberger2024}. Finally, while previous results on optimal experimental design focus on estimating the average effect \citep*{bai2022}, other designs, especially those that incorporate additional data, could reveal stronger evidence on the joint distribution of potential outcomes.
Design-based methods have been around for a long time, and we have demonstrated that it is feasible to implement our design-based decision rule, but a practical impediment to widespread implementation is that applied researchers often do not report enough information about their experiments to replicate them with design-based methods \citep*{young2019}. We use large language models to pull (OpenAI’s GPT-4o-mini) and categorize (OpenAI’s GPT-o3-mini) the randomization processes from 2,080 papers associated with randomized controlled trials from the Abdul Latif Jameel Poverty Action Lab (J-PAL). Only 61% have enough description for the large language model to confidently categorize the randomization method from the paper or associated AEA RCT Registry entry. However, of those, 78% use designs captured by design-based likelihoods that we consider. In the full set of papers, over 60% describe some mechanism through which there could be defiers. Another practical impediment to design-based methods is that instead of reporting absolute numbers, several researchers only report rates or regression results. Mandatory reporting of experimental design and standard sample counts by J-PAL, medical journals, and other authorities could facilitate the application of our design-based likelihood, allowing researchers to learn more from their experiments.
\setcounter{figure}{0}\setcounter{table}{0}
\titleformat{\section} {\thesection}{1em}
\titleformat{\subsection} {Appendix \Alph{subsection}}{1em}{\normalfont}
\titleformat{\subsubsection} {Appendix \Alph{subsection} .\arabic{subsubsection}}{1em}{\normalfont}