EconBase
← Back to paper

Finite Population Identification and Design-Based Sensitivity Analysis

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

134,571 characters · 34 sections · 157 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Finite Population Identification and Design-Based Sensitivity Analysis

abstractWe develop a new approach for quantifying uncertainty in finite populations, by using design distributions to calibrate sensitivity parameters in finite population identified sets. This yields uncertainty intervals that can be interpreted as identified sets, robust Bayesian credible sets, or uniform frequentist design-based confidence sets. We focus on quantifying uncertainty about the average treatment effect, where our approach (1) yields design-based confidence intervals which allow for heterogeneous treatment effects without using asymptotics, (2) provides a new motivation for examining covariate balance, and (3) gives a new formal analysis of the role of randomization. We illustrate our approach in three empirical applications.

JEL classification: C18; C21; C25; C51

Keywords: Treatment Effects, Partial Identification, Randomization, Uncertainty Quantification

\onehalfspacing

Introduction

Many empirical papers use variation across a small number of units to learn about causal effects. For example, in many settings this variation is across units like states or industries. How does identification of causal effects work in such settings, and how should empirical researchers quantify uncertainty? Our paper gives new answers to these questions, focusing on the context of randomized experiments with a finite number of units.

Traditionally, the $N$ units are assumed to be sampled from a hypothetical infinite “super” population, but recently there has been a resurgence in the “design-based inference” literature (reviewed in section (ref)) which studies explicitly finite populations.\footnote{The “design-based” versus “super-population” distinction and jargon come from the survey sampling literature (e.g., SarndalEtAl1978), which is how we use them here. In economics, some researchers now use the term “design-based” in a related but not equivalent sense, to refer to specific identification strategies (e.g. Card2022). We focus on randomized experiments in this paper and hence both meanings of “design-based” apply.} With very few exceptions (e.g., ManskiPepper2018, BorusyakHullJaravel2024, and RambachanRoth2025), most papers that study finite populations do not discuss identification and instead focus solely on proving frequentist properties like consistency or confidence interval validity. In contrast, there is a vast and rich literature on identification in infinite populations\footnote{See ChristensenConnault2023, Obradovic2024, Spini2024, Deaner2025, Sloczynski2025, BlandholEtAl2025, and CaetanoEtAl2026 for a few recent examples, among many others; for comprehensive surveys see Matzkin2007, Manski2009, BontempsMagnac2017, Lewbel2019, ChesherRosen2020, Molinari2020, and KlineTamer2023.}; no corresponding literature on identification in finite populations exists. Our first main contribution, therefore, is to develop a formal theory of identification in finite populations.

The literature on identification in finite populations is so sparse that we are not aware of any source which formally defines identification in an explicitly finite population setting (though some references come close). Nonetheless, it is quite common and “conventional” in the prior literature (following the term used by HeckmanVytlacil2007, footnote 6, page 4786) to informally state that a parameter is point identified if there is a known mapping from the distribution of the data induced by randomization (of treatment assignment, for example) into the parameter. We state this prior definition formally, and then argue that it has two major problems.

First, we show that this prior conventional definition relies on hypothetical data that is impossible to observe, since it involves averaging over counterfactual worlds that are incompatible with each other. As a consequence, unlike in infinite populations, we prove that this prior conventional definition in finite populations implies the values of every potential outcome for every unit are point identified in any setting where the probability of each treatment value is positive (a type of overlap restriction), including non-experimental settings where treatment assignment depends almost arbitrarily on potential outcomes (this finding builds on ideas in Imbens2018 and ChenRothSpiess2026, as we discuss in section (ref)). For example, according to the prior conventional definition, assumptions like unconfoundedness or instrument exogeneity are unnecessary to obtain point identification of treatment effects, so long as an overlap restriction holds. Therefore, we argue that the prior conventional definition is not adequate for assessing how informative the actual data one has are about treatment effects.

Second, the prior conventional definition of identification, as well as the broader literature on finite population design-based inference, pre-supposes that there is some kind of known a priori randomization. For example, that randomization may involve treatment assignment, an instrument assignment, or sampling, depending on the setting. These may be plausible in randomized experiments or designed surveys, but in many observational settings it is not clear that the variable of interest (like treatment assignment) is randomly chosen, let alone that the distribution it was drawn from is known to the researcher (a point discussed in ManskiPepper2018, pages 234--235). Consequently, the prior conventional definition of identification in finite populations does not apply to such cases.

To address these two problems, we develop a different, formal definition of identification for a parameter in a finite population setting. Our definition is analogous to the usual definition of identification in infinite population settings (e.g., Manski2003, Lewbel2019, Molinari2020): The finite population identified set for a parameter (such as the average treatment effect, ATE) is the set of parameter values consistent with (a) the finite matrix of data the researcher actually has and (b) any assumptions that they make.

Our definition directly addresses the two problems with the prior conventional definition described above: First, our definition does not require any knowledge about how treatment is assigned. Hence it can be applied in a wider variety of settings. Second, it uses actually observed data, rather than hypothetical data that is impossible to observe. Relatedly, it does not require any kind of asymptotics or hypothetical infinite populations. Consequently, any parameter which cannot be written as a known function of the actually observed data will typically be partially identified. The size of the identified set, however, will generally vary with the assumptions and data available, as in the usual infinite population definition of identified sets.

As a leading example of our identification analysis, we show that the average treatment effect in a randomized experiment is generally partially identified in finite populations.\footnote{For a recent survey of infinite population approaches to randomized experiments, see BaiShaikhTabordMeehan2025.} This conclusion arises from a subtle distinction between the finite and infinite population settings: Randomization guarantees that the treatment and control groups are exactly balanced in an infinite population, but not in finite populations. Since randomization cannot guarantee any level of balance between the treatment and control groups in a finite population, it does not have any identifying power.

Nonetheless, we show how to use our finite population identified sets combined with assumptions on treatment assignment to obtain a new way to quantify uncertainty in finite populations, which we call design-based sensitivity analysis. This is our second main contribution.

Overall, our method yields uncertainty intervals that can be interpreted as (1) identified sets, (2) robust Bayesian credible sets, or (3) uniform frequentist design-based confidence sets. This approach is non-asymptotic, which allows it to be reliably applied in small populations. And it allows researchers to flexibly impose additional assumptions on unobserved values, transparently quantifying the impact of these assumptions on conclusions. We summarize our approach in section (ref).

In particular, focusing on their frequentist interpretation, our uncertainty intervals avoid several limitations of the confidence intervals that are currently available in the literature: (i) They do not rely on large-$N$ asymptotics, which can perform poorly if $N$ is small (e.g., $N=10$) and (ii) they allow for heterogeneous treatment effects. They are also valid uniformly over the set of all possible matrices of potential outcomes consistent with the assumptions (an important feature in partially identified settings; see CanayShaikh2017). Furthermore, unlike the confidence intervals in the literature, our uncertainty intervals have both Bayesian and identification-based interpretations. This allows empirical researchers to choose their most preferred interpretation. This interpretational flexibility is one benefit obtained from a quantification of uncertainty that begins with identification. In particular, our uncertainty intervals are not based on hypothesis test-inversion, but are instead based on ideas from partial identification.

As mentioned above, we develop this approach in the context of randomized experiments with a binary treatment, where the goal is to quantify the uncertainty in the average treatment effect due to missing potential outcomes. We extend the analysis in three main directions: In section (ref) we give a new motivation for examining covariate balance. Specifically, by making an assumption that explicitly links covariates to potential outcomes, we show that observed imbalances in covariates have identifying power for the unobserved realized imbalance in potential outcomes, which allows researchers to obtain tighter bounds on ATE. In section (ref) we study instrumental variable based analysis under one-sided noncompliance. And in section (ref) we discuss uncertainty due to non-sampled units.

Our Approach to Quantifying Uncertainty in Experiments

\nocite{Manski1990}

Here we sketch our approach in the context of causal inference. See appendix (ref) for a parallel exposition in the context of survey sampling. Consider the population of six units in table (ref) (left). Here $Y_i(1)$ and $Y_i(0)$ denote potential outcomes and $X_i \in \{0,1\}$ denotes realized treatment. The true ATE is 0.55, the difference between the average of $Y_i(1)$ over $i=1,\ldots,6$ and the average of $Y_i(0)$ over $i=1,\ldots,6$. The data (table (ref) right) only reveals half of the potential outcomes, however. Suppose we know that all outcomes must be between 0 and 1. Then without further assumptions, all we can say for sure about ATE is that it is in the set

align[align omitted — 274 chars of source]

This is an example of what we call the finite population identified set for ATE; in this case, it is the set of ATE values consistent with the observed data and known bounds on outcomes, but no other assumptions are imposed (see section (ref) for a formal definition). It is obtained by filling in the unobserved values in table (ref) (right) with values in $[0,1]$ to either maximize or minimize the corresponding ATE value (analogous to Manski's 1990 derivations).

table[table omitted — 2,434 chars of source]

Suppose we also knew that treatment was randomly assigned. In infinite populations, this assumption is routinely interpreted to imply $\ensuremath{\mathbb{E}}[Y(x) \mid X=1] = \ensuremath{\mathbb{E}}[Y(x) \mid X=0]$ for $x \in \{0,1\}$, a post-randomization exact balance condition which leads to point identification of ATE (see appendix (ref) for a discussion of this notation in the context of finite populations). Unfortunately, as is well known (e.g., Altman1985, GreenlandRobins1986, and Greenland1990), in finite populations randomization does not guarantee any kind of post-randomization balance in potential outcomes, even if it makes such balance “likely”. Consequently, ATE is not point identified in finite populations, even when treatment is randomly assigned; its identified set continues to be equation (ref). We formalize this in theorem (ref).

What explains the difference in the identification status of ATE under randomly assigned treatment between the infinite and finite population settings? We first emphasize what is not different about the two settings: In both settings, sampling uncertainty has been assumed away---data on all units in the population are observed. Moreover, in both settings we only observe data from a single realization of random assignment of treatment for each unit in the population. In particular, classical infinite population identification analyses (e.g., HeckmanVytlacil2007) always assume that we only observe one potential outcome per unit, because of the fundamental problem of causal inference---they can either get treatment or not. Analogously, our finite population identification analysis is based on data like table (ref) right. We discuss this similarity in more detail at the end of section (ref).

This discussion shows that the only difference in the identification analysis between the two settings is population size: Exact balance obtains for infinitely many units but not finitely many. We discuss that distinction further in section (ref).

Although random assignment of treatment does not guarantee exact balance in finite populations, it is “likely” to deliver “some” balance. We formalize this idea in two steps, starting with a non-probabilistic analysis and then moving to a probabilistic analysis. Specifically, suppose we were willing to assume that (a) the average of the unobserved values in the dashed box is not farther than $K$ from the observed mean $(0.1+0+0.2)/3 = 0.1$ in the solid box and (b) the average of the unobserved values in the dashed oval is not farther than $K$ from the observed mean $(0.6+0.9+0.8)/3 = 0.77$ in the solid oval (a finite population version of an assumption proposed in Manski2003). We call this the $K$-approximate mean balance assumption and derive the identified set for ATE under this assumption in theorem (ref), which we denote by $\Theta_I(K)$. For sufficiently small $K$, we can conclude that ATE is in a strictly smaller set than equation (ref), where the size of this set depends on $K$, the maximal magnitude of imbalance between the treatment and control groups. Note, in particular, that ATE is point identified if $K=0$ (exact balance), as in infinite population analyses, although exact balance is implausible for small finite $N$, as discussed earlier.

figure[figure omitted — 751 chars of source]

The top plot in figure (ref) shows an example of these identified sets $\Theta_I(K)$ as a function of $K$ (using real data; see section (ref)). By examining how these sets change with $K$, we can examine the sensitivity of conclusions about ATE to assumptions about the magnitude of the realized differences in average potential outcomes across the treatment and control group. These sets provide a non-probabilistic approach to quantifying uncertainty about ATE.

Researchers may be unsure how to select plausible values of $K$, however. In order for $\Theta_I(K)$ to contain the true ATE, $K$ must be weakly larger than the true, but unknown magnitude of imbalance between the treatment and control groups, which we denote by $K^\text{true,max}$ (e.g., it is $\max \{0.13, 0.1\}$ in table (ref)). Consequently, we next use random assignment of treatment to probabilistically quantify uncertainty in $K^\text{true,max}$. We use this to calibrate values of $K$, which will lead to a probabilistic quantification of uncertainty in ATE.

Specifically, suppose momentarily that the true values of all potential outcomes were known. Then, since we know the treatment assignment distribution, we can compute the ex ante probability that the difference in potential outcome means across the treatment and control groups is at most $K$. This can be done for any $K$. This yields a distribution of possible values of the ex post imbalance $K^\text{true,max}$. Since we do not actually know all potential outcomes, however, this distribution itself is partially identified. Nonetheless, we can find the smallest possible value of this probability, across all logically possible completions of table (ref) (right). We do this using modern computational methods, described in section (ref). Denote this smallest probability by $\underline{p}(K)$. We can now plot this function $\underline{p}(\cdot)$, which converts $K$-values into ex ante design-probabilities. The bottom plot in figure (ref) shows an example.

Finally, leading to our main empirical recommendation, we combine the two plots in figure (ref) as follows: For any desired ex ante probability of balance $1-\alpha \in (0,1)$, use the bottom plot to find the smallest magnitude of imbalance that occurs with at least $100(1-\alpha)$% ex ante probability; denote this value by $K(\alpha)$. For example, if we set $1-\alpha = 0.9$ then the figure shows us that $K=18.9$ satisfies $\underline{p}(K) = 1-\alpha$ (see the dotted lines). Then use the top plot to read off bounds on ATE for this value of $K$, as shown in the dotted lines. In this case the bounds are about $[-6, 22]$.

Our main construction is a set of uncertainty intervals, such as those shown in figure (ref), which plots the bounds $\Theta_I(K(\alpha))$ for all values of $1-\alpha$.\footnote{Note: For $1-\alpha$ close to zero but strictly positive, the bounds in figure (ref) are non-singleton. This arises since exact balance is impossible in the worst case completions of table (ref); this is why the plot of $\underline{p}$ in figure (ref) is flat at zero up to a strictly positive value of $K$. We discuss this point further in section (ref).} We show that these uncertainty intervals have three interpretations:

enumerate• First, $\Theta_I(K(\alpha))$ is the finite population identified set for ATE, under the deterministic assumption that the maximum magnitude of ex post imbalance is at most $K(\alpha)$. • Second, $\Theta_I(K(\alpha))$ is a robust empirical Bayesian $1-\alpha$ credible set for ATE. That is, $\Theta_I(K(\alpha))$ contains the true ATE with posterior probability at least $1-\alpha$ for any prior over the vector of all potential outcomes (such as the 12 potential outcomes in table (ref) left) that satisfies our bounded outcome assumption. Essentially, we show that even the “most pessimistic” Bayesian has a posterior distribution over the true, but unknown, magnitude of imbalance in potential outcomes $K^\text{true,max}$ that is no more pessimistic than the distribution $\underline{p}$ that we constructed above (and we show that $\underline{p}$ is in fact a valid cdf). • Third, $\Theta_I(K(\alpha))$ is a uniformly valid design-based $1-\alpha$ confidence interval for ATE. That is, across repeated re-assignments of treatment, these bounds will contain the true ATE at least $100(1-\alpha)$% of the time. Uniformity here means that this property holds for all possible vectors of potential outcomes that satisfy our bounded outcome assumption (such as the 12 potential outcomes in table (ref) left).

We discuss each of these interpretations further in section (ref). Overall, we recommend that researchers report plots like figure (ref), which can then be interpreted any of these three ways.

Besides reporting interval quantifications of uncertainty, researchers also often report summary statistics that are meant to measure the “strength of evidence” in favor of a hypothesis. Our results also allow for such analyses. For example, we may ask: How much evidence is there in the data for the conclusion that ATE is nonnegative? The horizontal intercept in figure (ref), called a breakdown point, provides a simple answer. In figure (ref), we first compute the horizontal intercept in the top plot, which is the largest magnitude of imbalance between mean potential outcomes that can be allowed while still allowing us to conclude that ATE is nonnegative. We then plug it into $\underline{p}$ to convert it to design probabilities (see the dashed lines). Here we obtain $\underline{p}(12.871) = 0.65$. This tells us that, given the data, and regardless of what the true unknown potential outcomes actually are, there was an ex ante probability of at least 65% that the treatment and control groups would be sufficiently balanced to ensure that the identified set for ATE only contains non-negative numbers. Or, using the Bayesian interpretation of our analysis, this number says there is at least a 65% posterior probability that ATE is non-negative.

We also show how to obtain tighter bounds by making further assumptions about unobserved potential outcomes, which illustrates the flexibility of our approach. For example, in many experiments researchers observe covariates for each unit. It is often reasonable to expect that these covariates are predictive of potential outcomes. If researchers are willing to place a lower bound on this predictive power, as measured by the R-squared from the regression of potential outcomes onto covariates, then they can obtain tighter uncertainty intervals for ATE. We illustrate this analysis in figure (ref) in section (ref).

Beyond this basic setup, we discuss a variety of extensions and complementary results throughout the paper, including (1) instrumental variables and noncompliance, (2) random sampling of units, in addition to randomization, (3) other restrictions on the unobserved potential outcomes, like having an a priori bounded variance, (4) identification of parameters besides the average treatment effect, and (5) using other measures of balance beyond means.

Related Literature

Our paper relates to a variety of literatures, but for brevity we only discuss the most closely related work, rather than attempt a comprehensive survey. ManskiPepper2018 derive identified sets in an explicitly finite population setting. Like them, we also consider the identifying power of bounded variation type assumptions. A main difference is that they focus on a setting with observational data whereas we primarily consider randomized experiments. Our focus on experiments allows us to use randomization itself to calibrate the magnitude of the bounded-variation sensitivity parameter $K$. Like us, GreenlandRobins1986 also describe the problem of drawing conclusions in finite populations as an identification problem which must be solved by making assumptions about unobserved values of variables. They formally show how exact balance assumptions yield point identification. They then informally explain that “if we randomize...when both samples are large, random differences will in probability be small” (page 415). Many other papers make similar informal remarks, like Lindley1980 and RoyallPfeffermann1982. Relatedly, the distinction between ex ante and ex post balance is well known; Greenland1990 gives a particularly clear discussion, where he considers a randomized experiment with $N=2$. GreenlandRobins2009 more recently discuss this distinction, and conclude that “randomization (or more generally, ignorability) does not impose “no [ex post] confounding”...rather, it provides...a randomization-based (“objective”) derivation of a prior...that applies after allocation as well as before, and becomes more narrowly centered around zero as the sample size increases. This is a key post-allocation benefit of randomization” (page 5, emphasis in the original). That paper as well as the rest of this prior literature, however, is largely verbal discussion. One of our main contributions is to build on those observations to provide a new, detailed, and formal approach to identification and quantifying uncertainty in finite populations.

The standard approach to uncertainty quantification in finite populations is to construct frequentist design-based confidences sets. For example, chapter 5 of ImbensRubin2015 discusses confidence sets which are valid for finite $N$ but are based on homogeneous treatment effect assumptions, while chapter 6 discusses confidence sets which allow for heterogeneous treatment effects but which are based on asymptotics. Our uncertainty intervals also have a confidence set interpretation, and so fit within this standard approach. Importantly, however, our intervals also have two additional non-frequentist interpretations, as either identified sets or Bayesian credible sets. We are not aware of any other intervals in the literature which have these three distinct interpretations. Moreover, viewed purely as confidence sets for ATE, our intervals have valid coverage for any fixed $N$, allow for heterogeneous treatment effects, and do not require outcomes to be binary. We are only aware of one feasible alternative design-based confidence interval with these properties, which we call the Hoeffding CI, as discussed in AronowEtAl2025 and Ding2025 (who build on the binary outcome analysis of RobinsRitov1997). In our empirical applications this interval is always at least 3.5 times wider than ours, and can be up to 20 times wider, depending on the nominal coverage probability (see appendix (ref)). Since both intervals have valid coverage, this suggests that the Hoeffding CI is substantially less powerful than our interval. We discuss these issues as well as further related literature in appendix (ref).

Finally, many recent papers on causal inference study explicitly finite populations, including LiDing2017, LiDingMiratrix2017, AronowSamii2017, Ding2017, AtheyEcklesImbens2018, KangPeckKeele2018, AbadieAtheyImbensWooldridge2020, AbadieAtheyImbensWooldridge2023, HongLeungLi2020, WuDing2021, EcklesEtAl2020, ImbensMenzel2021, ZhaoDing2021, Savje2021, BojinovRambachanShephard2021, Xu2021, XuWooldridge2022, AtheyImbens2022, RothSantAnna2023, Pollmann2023, Wooldridge2023, StartzSteigerwald2023, Sancibrian2024, BorusyakHullJaravel2024, RambachanRoth2025, and BaiEtAl2025, among many others. Also see ImbensRubin2015 and Ding2024 for book level surveys. Empirical researchers have also recently argued in favor of these methods, leading to renewed applied interest as well; see Young2019. The theoretical literature largely follows a template laid out by Neyman (1923) that jumps straight to deriving various frequentist properties, without doing any identification analysis. Two exceptions are BorusyakHullJaravel2024 and RambachanRoth2025 which both discuss identification, based on the conventional definition, and do asymptotics-based inference (section 4.3 of BorusyakHullJaravel2024 also discuss exact inference under a constant effects assumption). Our finding that all potential outcomes are point identified under the conventional definition builds on ideas in Imbens2018 and ChenRothSpiess2026, as we discuss in section (ref). Rather than further studying identification, ChenRothSpiess2026 derive various frequentist testing and Bayesian updating implications of the conventional definition of identification, concluding that this conventional definition “translates to little, if any, practical learning” (page 14).

\nocite{Neyman1923}

Finite Population Identification

We begin by describing our approach to identification: We define the finite population identified set in section (ref). We then use this definition to derive identified sets for finite population average treatment effects under various assumptions in section (ref), and we discuss the role of randomization for identification in section (ref). We explain our concerns with the prior definition of identification in section (ref).

The Finite Population Identified Set

Consider the standard potential outcomes model with a binary treatment. There are four components required to conduct an identification analysis:

1. The population: Suppose there are $N$ units in our population. Let $\mathcal{I} \coloneqq \{1,\ldots,N \}$ be the set of indices for these units. Each unit $i \in \mathcal{I}$ is associated with the vector of numbers $(Y_i(1), Y_i(0), X_i, W_i)$, where $Y_i(1)$ and $Y_i(0)$ are potential outcomes, $X_i \in \{0,1\}$ is a realized binary treatment, and $W_i$ is a $d_W$-vector of covariates. The population is the $N \times (3+d_W)$ matrix of all of these numbers. Let $\ensuremath{\mathbf{P}}^\text{true} \coloneqq (\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0),\ensuremath{\mathbf{X}},\ensuremath{\mathbf{W}})$ denote this population matrix. We use $\ensuremath{\mathbf{P}}$ to denote alternative possible values of this population matrix. Table (ref) shows an example population.

2. Assumptions: We do not observe all elements of $\ensuremath{\mathbf{P}}^\text{true}$ because of the fundamental problem of causal inference. We typically make assumptions about its missing elements, the missing potential outcomes. Formalize these assumptions as the restriction that $\ensuremath{\mathbf{P}}^\text{true} \in \mathcal{P}$ where $\mathcal{P}$ is a known set of $N \times (3+d_W)$ matrices. We give examples of $\mathcal{P}$ in section (ref).

3. Parameters: Parameters of interest are functionals of the population matrix. For example, the average treatment effect is defined as $\text{ATE} \coloneqq \frac{1}{N} \sum_{i=1}^N \big( Y_i(1) - Y_i(0) \big)$. Similarly, the average effect of treatment on the treated is $\text{ATT} \coloneqq \Big( \sum_{i=1}^N \big( Y_i(1) - Y_i(0) \big) \allowbreak \ensuremath{\mathbbm{1}}(X_i = 1) \Big) / \sum_{i=1}^N \ensuremath{\mathbbm{1}}(X_i = 1)$, assuming the denominator is nonzero. In general, let $\theta(\ensuremath{\mathbf{P}})$ be a functional defined on the set of logically possible values of the population matrix. Let $\Theta$ denote the set of logically possible values of this parameter. Let $\theta^\text{true} \coloneqq \theta(\ensuremath{\mathbf{P}}^\text{true})$ denote its true value.

4. The data: Finally, we must describe the data that is observed to the econometrician. To allow for sampling, let $\ensuremath{\mathbf{S}} = (S_1,\ldots,S_N)$ be a vector of indicators where $S_i = 1$ if unit $i$ appears in the dataset, and zero otherwise. Let $\ensuremath{\mathbf{Y}} = (Y_1,\ldots,Y_N)$ where $Y_i \coloneqq Y_i(1) X_i + Y_i(0) (1-X_i)$ for all $i \in \mathcal{I}$. Then $\ensuremath{\mathbf{P}}^\text{data}$ is the matrix $(\ensuremath{\mathbf{Y}}, \ensuremath{\mathbf{X}}, \ensuremath{\mathbf{W}})$ with row $i$ empty if $S_i = 0$. Formalize this construction via the function $\text{MakeData}(\ensuremath{\mathbf{P}},\ensuremath{\mathbf{S}})$, so that $\ensuremath{\mathbf{P}}^\text{data} = \text{MakeData}(\ensuremath{\mathbf{P}}^\text{true}, \ensuremath{\mathbf{S}})$.\footnote{We define the population at the time period after which treatment has been assigned but before units are sampled. This is standard in the identification literature (e.g., Manski2003, ch.\ 7). One could modify the notation to further distinguish population values of potential outcomes and covariates from realized treatment assignment, as none of our concepts or procedures depend on this particular notational choice, but at the cost of additional notation throughout, especially since we mainly focus on the case where all units are sampled.}

We are now ready to define the identified set.

definitionThe identified set for $\theta$ is \[ \Theta_I \coloneqq \{ \theta \in \Theta : \theta = \theta(\ensuremath{\mathbf{P}}) \text{ for some } \ensuremath{\mathbf{P}} \in \mathcal{P} \text{ such that } \text{MakeData}(\ensuremath{\mathbf{P}},\ensuremath{\mathbf{S}}) = \ensuremath{\mathbf{P}}^\text{data} \}. \]

\nocite{Koopmans1949}

This definition follows the usual description of the identified set as “the set of parameters...that are consistent with the model and the data” (Tamer2010, page 184), but with close attention to what is meant by “data”: In a finite population, the “data” is defined by a finite dimensional matrix rather than a joint distribution of random variables (as in def. 3.1 on page 5324 of Matzkin2007, for example; also see our appendix (ref) for a related discussion). Our definition also allows for data on some units to be missing (i.e., not sampled, so that $S_i=0$); we discuss this further in section (ref).

$\Theta_I$ has two key features: (1) It can always be computed with the data the analyst actually has at hand, and (2) like the usual definition of the identified set in an infinite population, it is guaranteed to contain the true parameter of interest, so long as the model is not false. We will revisit these properties below.

Identified Sets for ATE

To simplify the exposition, we generally assume that all units in the population are observed ($S_i=1$ for all $i \in \mathcal{I}$). We show that the extension to incorporate sampling is straightforward in section (ref). We also focus on the identification of the average treatment effect for brevity. We briefly discuss other parameters in section (ref).

It is well known that bounds on mean based parameters are usually infinite without some kind of restriction on the values they can take. Hence we maintain the following assumption throughout the paper, which is standard in the literature on partial identification.

Aassumption[Bounded outcomes] There are known values $-\infty < y_\text{min} < y_\text{max} < \infty$ such that $Y_i(x) \in [y_\text{min},y_\text{max}]$ for all $i \in \mathcal{I}$, for each $x \in \{0,1\}$.

In some applications the values of $y_\text{min}$ and $y_\text{max}$ can be set to their logical values, such as 0 to 100 for test scores. In other settings, like when outcomes are wages, these values are sensitivity parameters that reflect our beliefs about the smallest and largest possible values of potential outcomes in the population under consideration. We discuss several variations on this assumption in section (ref), including an assumption that does not require selecting a priori bounds on outcomes.

(ref) alone has identifying power for ATE, via a finite population version of the Manski Manski1990 bounds. Those bounds are a special case of theorem (ref) below, which uses the following assumption.

Aassumption[$K$-approximate mean balance] There is a known $K \geq 0$ such that $K^\text{true}(x) \leq K$ for $x \in \{0,1\}$, where \begin{equation} K^true(x) \coloneqq \left| \frac{1}{N_1} \sum_{i=1}^N Y_i(x) \ensuremath{\mathbbm{1}}(X_i = 1) - \frac{1}{N_0} \sum_{i=1}^N Y_i(x) \ensuremath{\mathbbm{1}}(X_i = 0) \right| \end{equation} and $N_x \coloneqq \sum_{i=1}^N \ensuremath{\mathbbm{1}}(X_i = x)$ is the number of units who receive treatment $x$, with $N_1, N_0 > 0$.

This assumption was proposed by Manski2003, who called it approximate mean independence. He noted that if $K = 0$ then it is equivalent to mean independence of potential outcomes from realized treatment. ManskiPepper2018 call this a bounded variation type assumption. (ref) could be generalized to allow a different $K$ for each potential outcome; we use a common $K$ for simplicity. Let $\overline{\ensuremath{\mathbf{Y}}(x)} \coloneqq \frac{1}{N} \sum_{i=1}^N Y_i(x)$ and $\overline{Y}_x \coloneqq \frac{1}{N_x} \sum_{i=1}^N Y_i \ensuremath{\mathbbm{1}}(X_i = x)$. With a minor variation on Manski's Manski1990 derivations, we obtain the following characterization of the finite population identified set for ATE under the $K$-approximate mean balance assumption.

theoremSuppose (ref) and (ref) hold, and $\ensuremath{\mathbf{P}}^\text{data}$ is known. Then, for each $x \in \{0,1\}$, the identified set for $\overline{\ensuremath{\mathbf{Y}}(x)}$ is $[\text{LB}_K(x), \text{UB}_K(x)]$ where \begin{align*} LB_K(x) &\coloneqq \overline{Y}_x \frac{N_x}{N} + \max \{ y_min, \overline{Y}_x - K \} \frac{N_{1-x}}{N}, and \\ UB_K(x) &\coloneqq \overline{Y}_x \frac{N_x}{N} + \min \{ y_max, \overline{Y}_x + K \} \frac{N_{1-x}}{N}. \end{align*} Moreover, the identified set for ATE is $\Theta_I(K) \coloneqq [\text{LB}_K(1) - \text{UB}_K(0), \; \text{UB}_K(1) - \text{LB}_K(0)]$.

The magnitude of uncertainty about ATE depends on the choice of $K$. At one extreme, for sufficiently large values ($K \geq \overline{K} \coloneqq \max \{ \overline{K}_1, \overline{K}_0 \}$ where $\overline{K}_x \coloneqq \max \{ y_\text{max} - \overline{Y}_x, \; \overline{Y}_x - y_\text{min} \}$), $\Theta_I(K)$ gives what Manski2003 calls a “domain of consensus”, a set of ATE values that are consistent with the observed data and (ref), but otherwise no restriction at all on the relationship between realized treatment and potential outcomes. Denote this set by $\Theta_I(\infty)$. At the other extreme, $K=0$, mean independence holds. In this case, ATE is point identified and equals the observed difference in means, $\overline{Y}_1 - \overline{Y}_0$. In general, smaller $K$'s lead to smaller identified sets. So how should researchers assess the credibility of a choice of $K$?

Randomization and Identification in Finite Populations

In practice, researchers often motivate assumptions that two groups are in some sense comparable by appealing to random assignment. So consider the following common formalization of random assignment in finite populations.

Aassumption[Random assignment] The size of the treatment and control groups, $N_1$ and $N_0$, are fixed a priori, with $0 < N_1 < N$. $\ensuremath{\mathbf{X}} = (X_1,\ldots,X_N)$ is a single realization from the known probability distribution $\ensuremath{\mathbb{P}}_\text{design}$ on $\{ 0,1\}^N$ defined by \[ \ensuremath{\mathbb{P}}_\text{design}(X_1^\text{new} = x_1,\ldots,X_N^\text{new} = x_N) = 1 \Big/ {N \choose N_1} \] for all $(x_1,\ldots,x_N) \in \{0,1\}^N$ with $\sum_{i=1}^N x_i = N_1$, and equal to 0 otherwise.

In this assumption we use the notation $\ensuremath{\mathbf{X}}^\text{new} = (X_1^\text{new},\ldots,X_N^\text{new})$ to denote the random vector with distribution $\ensuremath{\mathbb{P}}_\text{design}$. $\ensuremath{\mathbf{X}} = (X_1,\ldots,X_N)$ is a single, non-random realization of this random variable. Call $\ensuremath{\mathbb{P}}_\text{design}$ the design distribution of treatment. This particular choice of design distribution is often called uniform randomization. It is a standard formalization of randomization in the design-based inference literature; for example, see ImbensRubin2015. We conjecture that most of our results extend to other design distributions commonly used in randomized experiments (as surveyed in RosenbergerLachin2015, for example), but we focus on uniform randomization for brevity.

As has long been recognized, randomization in finite populations does not guarantee any level of ex post balance (e.g., Greenland1990). The identified set, by definition, must contain the true parameter value whenever the model assumptions are true. Consequently, for any finite population, randomization has no identifying power. The following theorem formalizes this result.

theoremSuppose (ref) and (ref) hold and $\ensuremath{\mathbf{P}}^\text{data}$ is known. Then the identified set for ATE is $\Theta_I(\infty)$.

Put differently, after treatment has been assigned, we cannot rule out the possibility that potential outcomes are substantially imbalanced across the treatment and control groups, even if it was unlikely a priori. Thus the assumption that treatment was randomly assigned does not rule out any values of the unknown potential outcomes. Hence it does not shrink the domain-of-consensus bounds in finite populations. Nonetheless, in section (ref) we will reinterpret randomization as a procedure that affects our beliefs about ex post balance. Specifically, we will use randomization to assess the plausibility of a specific choice of the sensitivity parameter $K$. This will let us use the identification result in theorem (ref) to perform a sensitivity analysis motivated by random assignment of treatment. That analysis will yield uncertainty intervals that have interpretations as identified sets, robust Bayesian credible sets, or uniform frequentist confidence sets.

Exact Balance in Infinite versus Finite Populations

In infinite population identification analysis, random assignment is typically formalized as the assumption that realized treatment is statistically independent from potential outcomes, $X \mathbin{ \mathpalette{\@indep}{} } (Y(1),Y(0))$. This assumption implies mean independence, which in a finite population is equivalent to $K=0$. But in a finite population, $K=0$ is an exact balance assumption. It requires that $\frac{1}{N_1} \sum_{i=1}^N Y_i(x) \ensuremath{\mathbbm{1}}(X_i = 1) = \frac{1}{N_0} \sum_{i=1}^N Y_i(x) \ensuremath{\mathbbm{1}}(X_i = 0)$ for $x \in \{0,1\}$. In fact, for many values of the vectors $\ensuremath{\mathbf{Y}}(x) = (Y_1(x),\ldots, Y_N(x))$, it is impossible for exact balance to hold regardless of the values of realized treatment $\ensuremath{\mathbf{X}} = (X_1,\ldots,X_N)$, a point noted by GreenlandRobins1986. Hence statistical independence is not an appropriate formalization of random assignment in finite populations.

This qualitative distinction between the finite and infinite setting explains why theorem (ref) comes to a different conclusion about identification of ATE from traditional infinite population identification analyses: In the infinite population setting, exact balance is assumed to hold and hence ATE is point identified. But in the finite population setting, exact balance is not guaranteed even under random assignment, and is sometimes even logically impossible, and hence ATE is partially identified. That said, because any level of approximate balance is more likely when $N$ is large, there is a connection between these two results, which we study in section (ref). Also see section (ref), where we discuss various measures of balance.

The Prior Definition of Finite Population Identification

Now that we have explained our approach to identification, we turn to examine the prior definition. First, what is the prior definition? We are not aware of any books or papers that explicitly and formally define identification for finite populations. For example, BorusyakHullJaravel2024 discuss identification but do not provide a formal definition, and as representative of the rest of the literature, their first formal result is a consistency proof, not an identification theorem. The book by Tille2020 is an extensive and detailed survey of sampling from finite populations, but does not discuss identification. The textbook by Ding2024 discusses causal inference in finite populations in detail, starting from chapter 2, but a definition of identification does not appear until chapter 10, only after shifting focus from randomized experiments to observational data. That definition (10.1 on page 128) states that “a parameter $\theta$ is nonparametrically identifiable if it can be written as a function of the distribution of the observed data without any parametric model assumptions”. Similar definitions appear in other references, going back to the earliest sources (such as Koopmans1949, KoopmansReiersol1950, Hurwicz1950, Rothenberg1971). Verbal definitions like these are ambiguous about the meaning of “the distribution of the observed data” (Ding2024) or the “joint distribution of the observations” (Koopmans1949, page 125). There are different ways to formalize such statements and they lead to substantially different definitions of identification. As we argue below, this issue is particularly important in the finite population setting.

Our definition (ref) is one such formalization, where by “data” we mean the matrix $\ensuremath{\mathbf{P}}^\text{data}$ that the researcher actually has at hand. In this case the “distribution” of this data can be formalized as a specific discrete distribution (see appendix (ref) for details). However, based on the informal discussions in the literature, and our personal communication with researchers in this field, “data” is conventionally interpreted as the design distribution of the observed data, rather than the dataset that is actually observed. This interpretation leads to the following definition of point identification for finite populations.

definition[Conventional] Let $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0),\ensuremath{\mathbf{W}})$ be the $N \times (2+ d_W)$ matrix of potential outcomes and covariate values for all units in the population $\mathcal{I} = \{1,\ldots,N\}$. Let $\theta(\cdot)$ be a functional defined on the set of logically possible values of this matrix. Let $\theta^\text{true} \coloneqq \theta((\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0),\ensuremath{\mathbf{W}}))$ denote its true value. For each $i \in \mathcal{I}$, let $X_i^\text{new} \in \{0,1\}$ and $S_i^\text{new} \in \{0,1\}$ be random variables. Let $\ensuremath{\mathbf{X}}^\text{new} \coloneqq (X_1^\text{new},\ldots,X_N^\text{new})$ and $\ensuremath{\mathbf{S}}^\text{new} \coloneqq (S_1^\text{new}, \ldots, S_N^\text{new})$. Let $\ensuremath{\mathbb{P}}_\text{design}$ denote the joint distribution of $(\ensuremath{\mathbf{X}}^\text{new},\ensuremath{\mathbf{S}}^\text{new})$; it is called the design distribution. Let $Y_i^\text{new} \coloneqq Y_i(X_i^\text{new})$. Let $\mathcal{S} = \{ i \in \mathcal{I} : S_i^\text{new} = 1\}$ denote the random set of sampled indices. Let $\mathcal{D} \coloneqq \{ (i, Y_i^\text{new}, X_i^\text{new}, W_i) \in \mathcal{I} \times \ensuremath{\mathbb{R}}^{2+d_W} : i \in \mathcal{S} \}$. Say $\theta^\text{true}$ is point identified if there is a known mapping from the probability distribution of the random set $\mathcal{D}$ to $\theta^\text{true}$.

We believe definition (ref) accurately formalizes the conventional definition of finite population identification in the prior literature. For example, HeckmanVytlacil2007 say “identification in small samples requires establishing the sampling distribution of estimators, and adopting bias as the criterion for identifiability” (see our proposition (ref) below) and that “this approach is conventional in classical statistics,” although they do not give a formal definition which we could compare against. This definition also resembles the definition of “sampling identification” in FlorensSimoni2021, although that paper is not about design-based inference or finite populations.

Definition (ref) has the following seemingly innocuous and familiar implication.

propositionConsider the setting of definition (ref). Suppose there is a known function $m$ such that $\theta^\text{true}$ is the unique solution to $\ensuremath{\mathbb{E}}_\text{design}[ m(\mathcal{D}, \theta)] = 0$. Then $\theta^\text{true}$ is point identified.

This result says that a parameter is point identified if it is the unique solution to a set of design-based moment conditions. A special case arises when there is an $m$ is such that we can write $\theta^\text{true} = \ensuremath{\mathbb{E}}_\text{design}[m(\mathcal{D})]$. In this case the function $m(\mathcal{D})$ is usually called a design-unbiased estimator of $\theta^\text{true}$. Moment-based reasoning like in proposition (ref) is used in the identification analysis of BorusyakHullJaravel2024 and RambachanRoth2025. Proposition (ref), however, implies the following result.

propositionConsider the setting of definition (ref). Suppose $\ensuremath{\mathbb{P}}_\text{design}(S_i^\text{new}=1) = 1$ for all $i \in \mathcal{I}$. Then, according to definition (ref), the following hold. \begin{enumerate} • $p_i \coloneqq \ensuremath{\mathbb{P}}_\text{design}(X_i^\text{new} = 1)$ is point identified for all $i \in \mathcal{I}$. • $Y_i(1)$ is point identified for each $i \in \mathcal{I}$ with $p_i > 0$. $Y_i(0)$ is point identified for each $i \in \mathcal{I}$ with $p_i < 1$. \end{enumerate}
proof[Proof of proposition (ref)] Part 1 follows from $p_i = \ensuremath{\mathbb{E}}_\text{design}(X_i^\text{new})$ and proposition (ref). For part 2, define \[ \widehat{Y_i(1)} \coloneqq \frac{X_i^\text{new} Y_i^\text{new}}{p_i} \qquad \text{and} \qquad \widehat{Y_i(0)} \coloneqq \frac{(1-X_i^\text{new}) Y_i^\text{new}}{1-p_i}. \] Then \[ \ensuremath{\mathbb{E}}_\text{design}[\widehat{Y_i(1)}] = \frac{1}{p_i} \ensuremath{\mathbb{E}}_\text{design}[ X_i^\text{new} Y_i(X_i^\text{new}) ] \\ = \frac{1}{p_i} \ensuremath{\mathbb{E}}_\text{design}[ X_i^\text{new} ] Y_i(1) \\ = Y_i(1). \] Similarly, $\ensuremath{\mathbb{E}}_\text{design}[\widehat{Y_i(0)}] = Y_i(0)$. Apply proposition (ref).\footnote{An alternative proof defines $\theta \coloneqq Y_i(1)$, sets $m(\mathcal{D},\theta) \coloneqq Y_i^\text{new} X_i^\text{new} - \theta X_i^\text{new}$, and then notes that $\theta$ solves $\ensuremath{\mathbb{E}}_\text{design}[m(\mathcal{D},\theta)] = 0$.}

We assumed all units are sampled in proposition (ref) for simplicity, but the result can be generalized to allow for sampling. Proposition (ref) combines ideas from Imbens2018 and ChenRothSpiess2026. Specifically, using a different proof strategy and restricting attention to the binary outcome case, ChenRothSpiess2026 showed that the proportion of units who benefit from treatment is point identified, according to the conventional definition (ref) of identification in finite populations. Imbens2018 considered a randomized experiment with $N=1$ and $\ensuremath{\mathbb{P}}_\text{design}(X_1^\text{new}=1) = 0.5$ and showed that there exists a design-unbiased estimator of that units' unit-level treatment effect, $Y_1(1) - Y_1(0)$. Our proposition (ref) builds on these ideas to show that all potential outcomes are point identified, according to definition (ref), as long as the probability of treatment is strictly between zero and one for all units.\footnote{Following the original design-based literature in survey sampling (e.g., page 19 of CasselSarndalWretman1977 or page 10 of Thompson1997) and more recent design-based literature in economics such as AbadieAtheyImbensWooldridge2020 (who work with data matrices where the row position encodes the unit identifier), in proposition (ref) we assume that unit identifiers are observed in the data set. For example, if the population consists of all U.S. states, then the data consists of each state's name $i$ and its corresponding realized treatment $X_i$ and outcome $Y_i$. Chen et al's ChenRothSpiess2026 proposition 3.1 shows that the proportion of units who benefit from treatment is point identified according to the conventional definition (ref) even if only anonymized data is observed, meaning that the data is a list of $(Y_i,X_i)$ values, but it is not known which units these values belong to. Our proposition (ref) does not apply to this kind of anonymized data. However, suppose that the anonymized data includes covariates $W_i$. Then as long as the covariates are rich enough to uniquely identify units---which will typically be the case as long as there is a single “continuous” covariate, for example---the conclusion of proposition (ref) will continue to hold, even with anonymized data. This follows because the covariates effectively become unit identifiers.}

Proposition (ref) applies to randomized experiments, which typically guarantee $p_i \in (0,1)$ by design (e.g., (ref) holds), but it also applies to many observational settings as well, because this result does not assume that $p_i$ is known for all $i \in \mathcal{I}$. For example, unconfoundedness restricts the treatment assignment probabilities $p_i$ to be functions of observed covariates (in which case it is called the propensity score). Proposition (ref) implies that such a priori restrictions on treatment probabilities are unnecessary to point identify potential outcomes, according to definition (ref); it could even be that the unit-level treatment assignment probabilities $p_i = f(Y_i(1),Y_i(0))$ are functions of potential outcomes themselves; all that matters is that the ex ante probability of treatment is not degenerate on 0 or 1. Put differently, unconfoundedness is unnecessary according to the conventional definition of identification; overlap ($p_i \in (0,1)$ for all $i \in \mathcal{I}$) is sufficient for point identification of all unit-level treatment effects. Similarly, the conventional definition implies that all unit-level treatment effects are point identified for compliers in an instrumental variables setting, even if the instrument is endogenous (see appendix (ref) for details).

Essentially all of these conclusions implied by proposition (ref) differ from the conclusions of classical infinite population identification theory. The fact that a parameter can be partially identified under infinite population theory but point identified according the conventional definition was first pointed out by ChenRothSpiess2026, who focused on the proportion of units who benefit from treatment with a binary outcome.\footnote{They use this finding to motivate a careful study of the implications of the conventional definition of identification for Bayesian updating and various frequentist concepts, rather than exploring different definitions of identification as we do.} In a randomized experiment, that parameter is generally partially identified in infinite population theory (Makarov1982, Manski1997, FanPark2010), and yet the conventional definition (ref) implies that it is point identified. As they note, in the binary treatment and binary instrument setting of ImbensAngrist1994, this implies that the proportion of compliers is point identified even without a no-defiers assumption, another inconsistency with infinite population theory. Our result goes further to show that the compliance type (i.e., complier, defier, always taker, or never taker) is point identified for every unit, according to the conventional definition of identification (see appendix (ref)), a further inconsistency with infinite population theory.

Discussion: Averaging over Units or over Counterfactual Worlds

The different conclusions implied by different definitions of identification are a direct consequence of how “data” is defined. The conventional definition (ref) uses data that is not even in principle available to real-world researchers. That follows since means like $\ensuremath{\mathbb{E}}_\text{design}$ average over hypothetical, counterfactual worlds that are mutually incompatible with each other---such as a world where unit $i$ is treated and a world where unit $i$ is not treated. By allowing identification to depend on information from both such worlds, this definition allows one to conclude that both $Y_i(1)$ and $Y_i(0)$ are point identified. This conclusion suggests that there is, in fact, no fundamental problem of causal inference (Holland1986).

In stark contrast, both our definition (ref) and the standard definition of infinite population identification do not assume one can observe “data” from such mutually incompatible worlds. This holds for our definition (ref) since the identified set only depends on $\ensuremath{\mathbf{P}}^\text{data}$, the matrix of data that is actually observed by real empirical researchers. The standard infinite population definition is also not based on data from incompatible worlds, and hence does not lead to results like proposition (ref). Specifically, consider a joint distribution of random variables $(Y,X)$, representing the population “data” of realized outcomes and treatments. As explained in HeckmanVytlacil2007, for example, each realization $(Y(\omega), X(\omega))$ from this distribution represents a single unit $\omega$ in some set of units $\Omega$ (they assume $\Omega = [0,1]$) who is either treated ($X(\omega)=1)$ or not ($X(\omega)=0$). Consequently, even though the population is infinite, each unit is only treated once. The “data” available to researchers comes from a single assignment of each of the infinitely many units to be treated or not. Concretely, infinite population expectations like $\ensuremath{\mathbb{E}} \left[ \frac{Y X}{\ensuremath{\mathbb{P}}(X=1 \mid W)} \right]$, as used in a selection-on-observables analysis (e.g., ImbensRubin2015, section 12.4), do not involve averaging worlds where some units are treated with other worlds where those same units are not treated. Rather, such averages are over different units in the infinite population instead of over different counterfactual worlds involving different treatment assignments, as in $\ensuremath{\mathbb{E}}_\text{design}$.

In summary, these issues with the conventional definition motivate our development of a different theory of identification, based on definition (ref).

Design-Based Sensitivity Analysis

Thus far we have defined and studied identification in finite populations. We also derived the identified set for ATE, $\Theta_I(K)$, under the $K$-approximate mean balance assumption. That set provides a non-stochastic quantification of uncertainty about ATE, a mapping from a class of deterministic assumptions about balance into values of ATE consistent with the observed data and the assumptions. Researchers may be unsure about which specific values of $K$ are plausible, however. In this section, we show how to use a known design distribution of treatment to probabilistically quantify uncertainty about $K$, which can then be combined with our identified sets to probabilistically quantify uncertainty about ATE.

A Design-Based Approach to Calibrating $K$

The identified set $\Theta_I(K)$ as a function of the sensitivity parameter $K$, as in the top plot of figure (ref), shows how our conclusions can vary from point identification of ATE (under exact balance, $K=0$) to partial identification under approximate balance ($K > 0$). Like any sensitivity analysis, however, there is an important question: Which values of $K$ are most plausible? In this section, we provide an objective approach to calibrating this sensitivity parameter, based on an assumption that treatment was randomly assigned. Specifically, define

multline*[multline* omitted — 356 chars of source]

This is the design-probability that the $K$-approximate mean balance assumption holds, when the true potential outcomes are $\ensuremath{\mathbf{Y}}(1)$ and $\ensuremath{\mathbf{Y}}(0)$. Although $\ensuremath{\mathbb{P}}_\text{design}$ is known, $p(K,\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$ is unknown since it depends on unknown values of the potential outcomes.

We can, however, think of $p(K,\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$ itself as a parameter---it is a functional of the population matrix of potential outcomes, holding $K$ fixed. We will show that this parameter is itself partially identified, and we can derive its identified set, in the same sense as we developed in section (ref). Specifically, we obtain bounds on this probability by using two constraints: (1) Half of all potential outcomes are known (namely, those associated with the observed treatments) and (2) Outcomes are bounded ((ref)).

We focus on the lower bound on $p(K,\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$, since this tells us the smallest design probability such that $K$-approximate mean balance is guaranteed. So let

equation[equation omitted — 419 chars of source]

where $\times$ denotes component-wise multiplication and $\bold{1}$ is an $N$-vector of ones. $\underline{p}$ is a known function. It depends on the realized data $(\ensuremath{\mathbf{Y}},\ensuremath{\mathbf{X}})$ and the design distribution $\ensuremath{\mathbb{P}}_\text{design}$. We discuss how to feasibly compute this function in section (ref). For now we focus on its interpretation and use. Note that, by construction, $\underline{p}(K) \leq p(K,\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$ for any $\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0)$ satisfying the constraints in (ref); in particular, this holds for the true population values of potential outcomes.

We use $\underline{p}(K)$ to calibrate plausible values of the sensitivity parameter $K$ in the identification results of section (ref). Specifically, we suggest performing what we call a design-based sensitivity analysis:

enumerate• First, plot $\Theta_I(K)$, the identified set for ATE, as a function of $K$. Below this, plot the function $\underline{p}(K)$. Figure (ref) gives an example of this paired plot. The top graph shows the sequence of finite population identified sets alone. The second plot then converts its horizontal axis values of $K$ into design-probabilities. • Focus on particular values of $K$. There are several reasonable ways to do this: \begin{enumerate} • Define the breakdown point $K^\text{bp} \coloneqq \sup \{ K \geq 0 : 0 \notin \Theta_I(K) \}$, the largest relaxation of exact balance such that zero is not in the identified set. Researchers can compute $\underline{p}(K^\text{bp})$ to interpret this value. In figure (ref) it is 0.65. Thus there was at least a 65% ex ante probability that potential outcomes would be sufficiently balanced that we can conclude that ATE is positive. • Next, notice that $\underline{p}(K)$ decreases as $K$ decreases, because smaller $K$ implies a stronger balance assumption that is therefore less likely to hold. Let $\alpha \in (0,1)$. Define $K(\alpha) \coloneqq \inf \{ K \geq 0 : \underline{p}(K) \geq 1-\alpha \}$, the closest we can get to exact balance while still ensuring that approximate balance holds with design-probability at least $1-\alpha$. Researchers can then present the set $\Theta_I \big( K(\alpha) \big)$ as a function of $1-\alpha$. Figure (ref) gives an example. \end{enumerate}

Interpreting $\Theta_I(K(\alpha))$

Our main recommendation is that researchers report the set $\Theta_I(K(\alpha))$ as a function of $1-\alpha$, as in figure (ref). This set can then be interpreted three different ways. The first interpretation is non-probabilistic while the second two are probabilistic.

First, $\Theta_I(K(\alpha))$ can be interpreted as the finite population identified set for ATE under different deterministic assumptions about the true, but unknown, magnitude of ex post imbalance, $K^\text{true,max} \coloneqq \max \{ K^\text{true}(1), \allowbreak K^\text{true}(0) \}$ (recall that $K^\text{true}(x)$ is defined in eq.\ (ref)); namely, under the assumption that $K^\text{true,max}$ is at most $K(\alpha)$. This follows immediately from the identification analysis of section (ref). However, since researchers may not know how large $K^\text{true,max}$ is likely to be, our second and third interpretations view $[0, K(\alpha)]$ as an uncertainty interval for $K^\text{true,max}$, using either a robust Bayesian or uniform frequentist interpretation. Using our identification analysis in section (ref), these uncertainty intervals over $K^\text{true,max}$ then translate into uncertainty intervals over ATE, yielding $\Theta_I(K(\alpha))$, as we discuss next.

Second, viewed as a function of $\alpha$, $\Theta_I(K(\alpha))$ can be interpreted as a robust empirical Bayesian $1-\alpha$ credible set for ATE. To see this, consider the classical, “fully” Bayesian approach (e.g., Rubin1978 or section 8.4 of ImbensRubin2015). This approach starts with a prior on the entire matrix $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$. The random assignment of treatment assumption (ref) then implies corresponding beliefs over the magnitude of imbalance between the treatment and control groups. This is sufficient to obtain a distribution over $K^\text{true,max}$. Knowing that distribution alone would allow us construct a Bayesian credible set for ATE, by using our identified sets $\Theta_I(K)$ from theorem (ref) and treating $K$ as the uncertain quantity.

However, as usual with Bayesian analyses, it is unclear what specific prior over $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$ should be used. To address this, we first show that $\underline{p}$ is a valid cdf on $\ensuremath{\mathbb{R}}_+$:

propositionSuppose (ref) and (ref) hold. Then for any realization $\ensuremath{\mathbf{X}}$, $\underline{p}$ is monotonic, $\underline{p}$ is right continuous, $\lim_{K \rightarrow 0} \; \underline{p}(K) = 0$, and $\lim_{K \rightarrow \infty} \; \underline{p}(K) = 1$.

From this result and the definition of $\underline{p}$, we can show that $\underline{p}$ is a “worst case” or “robust” distribution over $K^\text{true,max}$, in the sense that every prior on $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$ yields a distribution over $K^\text{true,max}$ that is “more optimistic” about balance than the distribution $\underline{p}$. In other words, $[0,K(\alpha)]$ is a $1-\alpha$ credible set for $K^\text{true,max}$, based on the distribution $\underline{p}$, and this set is also a $1-\alpha$ credible set based on any prior distribution over $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$ satisfying (ref).

This distribution $\underline{p}$ over $K^\text{true,max}$ induces a distribution over the sets $\Theta_I(K)$. Since $\Theta_I(K)$ contains the true ATE so long as $K \leq K^\text{true,max}$ and the model is not falsified, this implies that $\Theta_I(K(\alpha))$ is a robust empirical Bayesian $1-\alpha$ credible set for ATE: It is a $1-\alpha$ credible set for ATE for any prior on $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$ that is consistent with our baseline assumption on potential outcomes, (ref). We call it an “empirical” Bayes credible set because the distribution $\underline{p}$ depends on the observed data $(\ensuremath{\mathbf{Y}},\ensuremath{\mathbf{X}})$. We give additional details and discussions of these results in appendix (ref).

This Bayesian interpretation is based on obtaining a distribution over the sets $\Theta_I(K)$, rather than obtaining a posterior distribution for ATE itself. This distinction is well known from the previous literature on Bayesian inference in partially identified settings. One consequence is that, like that previous literature (e.g., poirier1998revising, moon2012bayesian, kline2016bayesian, giacomini2021robust), we cannot make probabilistic statements about where ATE is within the set $\Theta_I(K(\alpha))$. This analysis nonetheless allows us to make some probabilistic statements about ATE directly. For example, in figure (ref), we can say that there is at least a 90% posterior probability that ATE is in $\Theta_I(K(0.1)) = [-6,22]$, and that there is at least a 65% posterior probability that ATE is non-negative, regardless of what one's prior on $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$ is.

Third, $\Theta_I(K(\alpha))$ has a frequentist interpretation as a uniform design-based confidence interval. To make the frequentist thought experiment explicit, write $\mathcal{C}(\ensuremath{\mathbf{Y}}(1) \times \ensuremath{\mathbf{X}} + \ensuremath{\mathbf{Y}}(0) \times (\bold{1}-\ensuremath{\mathbf{X}}),\ensuremath{\mathbf{X}}) = \Theta_I(K(\alpha))$ to emphasize that both the identified set $\Theta_I(K)$ and $K(\alpha)$ depend on the vector of realized treatments $\ensuremath{\mathbf{X}}$ and the potential outcome vectors $\ensuremath{\mathbf{Y}}(1)$ and $\ensuremath{\mathbf{Y}}(0)$. Let $\theta\big( \ensuremath{\mathbf{Y}}(1), \ensuremath{\mathbf{Y}}(0) \big) \coloneqq \frac{1}{N} \sum_{i=1}^N Y_i(1) - Y_i(0)$ denote the average treatment effect functional.

theoremSuppose (ref) and (ref) hold. Then \[ \resizebox{0.98\linewidth}{!}{$ \displaystyle \inf_{\ensuremath{\mathbf{Y}}(1), \ensuremath{\mathbf{Y}}(0) \in [y_\text{min},y_\text{max}]^N} \; \ensuremath{\mathbb{P}}_\text{design}\Big( \mathcal{C}(\ensuremath{\mathbf{Y}}(1) \times \ensuremath{\mathbf{X}}^\text{new} + \ensuremath{\mathbf{Y}}(0) \times (\bold{1}-\ensuremath{\mathbf{X}}^\text{new}),\ensuremath{\mathbf{X}}^\text{new}) \ni \theta \big( \ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0) \big) \Big) \geq 1-\alpha. $} \]

As usual in frequentist design-based finite population analyses, $\ensuremath{\mathbf{X}}^\text{new}$ is the only random quantity under consideration. Theorem (ref) shows that $\Theta_I(K(\alpha))$ is a $100 (1-\alpha)$% design-based confidence interval. That is, across repeated random assignments of treatment, $\Theta_I(K(\alpha))$ will contain the true parameter value with design-probability at least $1-\alpha$.

Viewed as a confidence interval, $\Theta_I(K(\alpha))$ has several key features that distinguish it from alternatives in the literature: First, it is valid for any fixed $N$; that is, it does not rely on an assumption that a large-$N$ asymptotic approximation holds. Second, it does not require strong assumptions like treatment effect homogeneity. And third, it is uniformly valid over any $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0))$ matrix satisfying (ref). Uniform validity is well known to be an important property in partially identified settings (e.g, CanayShaikh2017). In appendix (ref) we discuss all of these features in more detail, give further discussion of the literature on design-based confidence intervals for ATE, and use simulations to study coverage probabilities.

These features of $\Theta_I(K(\alpha))$ all follow from its construction as a finite population identified set for ATE with sensitivity parameter $K = K(\alpha)$ calibrated to be large enough that it will be larger than $K^\text{true,max}$ often enough, across repeated random assignments of treatment. That is, our approach of first studying identification of ATE under non-probabilistic assumptions helps focus attention on $K^\text{true,max}$ as a key unknown parameter, which led us to probabilistically quantify uncertainty on ATE by first probabilistically quantifying uncertainty in $K^\text{true,max}$. This approach helps us overcome the downsides of the traditional methods for constructing design-based confidence intervals based on inverting hypothesis tests for sharp nulls or using large-$N$ asymptotics.

The Value of Randomization

In section (ref) we showed that, in any finite population, random assignment of treatment does not have any identifying power because it does not guarantee any particular level of ex-post balance. However, random assignment of treatment does impact our beliefs about balance. This led to our construction of robust Bayesian credible sets above. Our next result shows that our set $\Theta_I(K(\alpha))$ yields tight conclusions about ATE in large populations.

theoremConsider a sequence of finite populations, $\{ (Y_i(0)_N, Y_i(1)_N) : i =1,\ldots,N \}$. Assume that for each $x \in \{0,1\}$ there is a constant $\mu(x) \in \ensuremath{\mathbb{R}}$ such that $\frac{1}{N} \sum_{i=1}^N Y_i(x)_N \rightarrow \mu(x)$ as $N \rightarrow \infty$. Suppose (ref) holds and that the bounds $y_\text{min}$ and $y_\text{max}$ do not depend on $N$. Assume $N_1 / N \rightarrow \rho$ for some constant $\rho \in (0,1)$. Suppose that for each $N$, a vector of treatment assignments $\ensuremath{\mathbf{X}}_N$ is drawn from the random vector $\ensuremath{\mathbf{X}}_N^\text{new}$ that satisfies (ref). Then: \begin{enumerate} • If $\mu(1) - \mu(0) \neq 0$, $\underline{p}(K^\text{bp}) \xrightarrow{p} 1$ as $N \rightarrow \infty$. • Let $\text{ATE}_N$ denote the finite population ATE. For any $\alpha \in (0,1)$, \[ \sup_{\theta \in \Theta_I(K(\alpha))} | \theta - \text{ATE}_N | \xrightarrow{p} 0 \qquad \text{as $N \rightarrow \infty$.} \] \end{enumerate}

Theorem (ref) is one way to formally show that randomly assigning treatment helps with learning about causal effects, in an explicitly finite population setting. Essentially, unobserved potential outcomes are “more likely” to be “more balanced” across the treatment and control groups in larger populations than small ones.

The first part of theorem (ref) concerns identification of the sign of ATE. It shows that, in large enough finite populations, the assumption of $K^\text{bp}$-approximate mean balance will indeed hold with probability approaching $1$. By definition of the breakdown point, this means that the “naive” conclusion about the sign of ATE (based on the exact balance assumption of $K = 0$) will be robust, in the sense that it is almost guaranteed that the amount of imbalance required to overturn that conclusion will not occur.

The second part of theorem (ref) concerns the behavior of the credible sets and confidences sets constructed above, $\Theta_I(K(\alpha))$. It is a consistency result: It shows that this set converges to the singleton true value of ATE in larger finite populations. This implies, for example, that if you invert $\Theta_I(K(\alpha))$ to construct a design-based test of the hypothesis $H_0: \text{ATE}_N = \theta_0$ then this test is consistent; it will reject false nulls with high design-probability when $N$ is large enough.

Measuring the Strength of Evidence: $\underline{p}(K^\text{bp})$ as an Alternative to $p$-values

$p$-values are commonly interpreted as quantitative measures of the evidence against the null hypothesis, with small values interpreted as “strong evidence” against the null and larger values interpreted as weaker evidence against it. This is a controversial interpretation (e.g., see the entire 2019 special issue of The American Statistician on “a world beyond $p < 0.05$”). In light of this debate about $p$-values, we suggest that our measure $\underline{p}(K^\text{bp})$ can be viewed as an alternative summary statistic that has a well justified interpretation as a quantitative measure of how much evidence the data provides for a specific conclusion. In particular, because our analysis has a Bayesian interpretation, it allows us to make certain probabilistic statements about the true value of the parameter. For example, when the naive difference in means estimate is positive, $\underline{p}(K^\text{bp})$ is a lower bound on the posterior probability that the true ATE is nonnegative.

Computing Bounds on the Design-Probability of $K$-Approximate Balance

Our approach requires computing the function $\underline{p}(K)$, which involves solving an optimization problem over $N$ variables. In appendix (ref), we show that $\underline{p}(K)$ is the optimized value in a mixed integer linear programming (MILP) problem. Consequently, standard software for solving these problems can be used. However, as we discuss in section (ref), the MILP solver can often be very slow. So we also recommend that users try alternative solvers to compute $\underline{p}$. We have found that genetic algorithms (KochenderferWheeler2019, pages 148--156) work exceptionally well in this setting, delivering results that are very close to the solution from MILP but in a small fraction of the time. See section (ref) for a discussion of the numerical evidence.

A second issue concerns computation of the objective function, which involves a summation over the set of all possible treatment assignments. For small values like $N=10$ this is feasible but it can become infeasible for moderate values like $N=100$. This issue is well known in design-based inference; e.g., sec.\ 5.8 in ImbensRubin2015. Like them, we address this issue by sampling from the set of all treatment assignments and using this to approximate the objective function. In section (ref) we show that our results are quite insensitive to the choice of sample size.

Numerical Illustration

Next we use simulated data to illustrate the analysis of section (ref). We generated a nested sequence of five populations, with $N \in \{ 10, 20, 40, 100, 400 \}$. These populations have $[y_\text{min}, y_\text{max}]=[0,1]$ and heterogeneous treatment effects such that $\text{ATE}_N$ converges to 0.25 as $N$ gets large. We assigned treatment via (ref) such that $N_1 = N/2$ for all $N$ while ensuring the populations are nested. In this illustration the assignment happens only once. Appendix (ref) gives further details on how we generated the data.

First we demonstrate the convergence results of theorem (ref). Its first part shows that $\underline{p}(K^\text{bp}) \xrightarrow{p} 1$ as $N \rightarrow \infty$. In the proof, we showed that $\underline{p}(\cdot)$ converges to a step function at zero. Figure (ref) demonstrates this convergence, which implies convergence of $\underline{p}(K^\text{bp})$ to 1 so long as the limiting breakdown point is nonzero. The second part of theorem (ref) shows that, for any $\alpha \in (0,1)$, the distance between the set $\Theta_I(K(\alpha))$ and ATE converges to zero as $N \rightarrow \infty$. Figure (ref) plots $\Theta_I(K(\alpha))$ as a function of $1-\alpha$, for the five different values of $N$. The bounds are centered at the naive difference-in-means estimand $\overline{Y}_1 - \overline{Y}_0$, which varies with $N$ but converges to 0.25 by construction. Again, consistent with the theory, we see that for any $\alpha \in (0,1)$ our bounds shrink as $N$ gets larger. This plot also shows the first part of theorem (ref), since $\underline{p}(K^\text{bp})$ is the point at which the bounds intersect the horizontal axis at zero. We see that this point converges to 1 as $N$ gets large. Indeed, for $N=100$ our bounds are strictly positive for almost all values of $1-\alpha$.

figure[figure omitted — 853 chars of source]

These results used a genetic algorithm (GA) to solve the optimization problem (ref). Next we verify that this algorithm is obtaining accurate results by comparing its output with output from the mixed-integer linear programming (MILP) approach that we discussed in section (ref). For $N=10$, the left plot in figure (ref) shows the function $\underline{p}$ obtained using the genetic algorithm as a solid line, and the same function obtained using mixed-integer linear programming as a dashed line. The two lines are very similar, showing that the genetic algorithm is able to closely match the output of the MILP approach, despite being substantially faster. Specifically, for this plot the genetic algorithm took about 4 minutes, whereas MILP took about 34 hours. Similarly, the right plot in figure (ref) shows $\Theta_I(K(\alpha))$ as a function of $1-\alpha$, as obtained by both algorithms. Again, the genetic algorithm is able to closely match the output from MILP. Finally, appendix figure (ref) shows that our results are robust to the choice of the number of treatment assignments used to approximate the objective function.

figure[figure omitted — 544 chars of source]

The Role of Balance in Covariates

We next illustrate the flexibility of our approach to finite population causal inference by extending it to incorporate covariates (sometimes called “attributes”) in this section and noncompliance in section (ref).

For each unit $i \in \mathcal{I}$, let $W_i$ denote a vector of covariate values. Let $\ensuremath{\mathbf{W}} = (W_1,\ldots,W_N)$ denote the collection of covariate values for all units in the population. In practice, researchers commonly examine the observed magnitude of balance in covariates across the treatment and control groups. In this section we give a new perspective on this kind of covariate balance analysis: Observed imbalances in covariates can be used to help identify the unobserved imbalance in potential outcomes. Informally, this additional information arises when the covariates are predictive of potential outcomes. For covariates to have identifying power, however, we must make an explicit assumption on this relationship. There are many different formal assumptions one could consider. For brevity we focus on a particularly simple one here, but it would be useful to explore variations in future work. We illustrate this analysis empirically in section (ref).

Let $\ensuremath{\mathbf{Q}} = (q(W_1)', \ldots, q(W_N)')'$ be a $N \times \text{dim}(q(W_i))$ matrix of known transformations of the observed covariates. Assume $\ensuremath{\mathbf{Q}}' \ensuremath{\mathbf{Q}}$ is invertible. Let $\beta(x) \coloneqq (\ensuremath{\mathbf{Q}}' \ensuremath{\mathbf{Q}})^{-1} \ensuremath{\mathbf{Q}}' \ensuremath{\mathbf{Y}}(x)$ denote the population regression coefficient from OLS of potential outcomes onto the transformed covariates. Let $U_i \coloneqq Y_i(x) - q(W_i)' \beta(x)$ denote the regression residual for unit $i$, with $\ensuremath{\mathbf{U}} \coloneqq (U_1,\ldots, U_N)$. Let $R_{Y(x) \sim q(W)}^2 \coloneqq 1 - \operatorname*{var}(\ensuremath{\mathbf{U}}) / \operatorname*{var}[\ensuremath{\mathbf{Y}}(x)]$, where for any vector $\ensuremath{\mathbf{A}} \coloneqq (A_1,\ldots,A_N)$ we let $\overline{\ensuremath{\mathbf{A}}} \coloneqq \frac{1}{N} \sum_{i=1}^N A_i$ denote its mean and $\operatorname*{var}(\ensuremath{\mathbf{A}}) \coloneqq \frac{1}{N} \sum_{i=1}^N (A_i - \overline{\ensuremath{\mathbf{A}}})^2$ denote its variance. Assume $q(W_i)$ includes a constant, so that $\overline{\ensuremath{\mathbf{U}}} = 0$ by construction, since these are OLS residuals. Not all potential outcomes are observed. Hence $\beta(x)$, $(U_1,\ldots,U_N)$, and $R_{Y(x) \sim q(W)}^2$ are not point identified. Instead, we make assumptions about them, as follows.

Aassumption[Predictive covariates] For each $x \in \{0,1\}$, $R_{Y(x) \sim q(W)}^2 \geq \lambda$ for a known $\lambda \in [0,1]$.

$\lambda$ is a sensitivity parameter that controls the minimal predictive power of the covariates relative to unobserved variables that determine variation in potential outcomes. Larger values of $\lambda$ require a closer connection between potential outcomes and covariates. Since some potential outcomes are known, not all values of the sensitivity parameter $\lambda$ are consistent with the data. The largest value of $\lambda$ that is consistent with the data and all of the maintained assumptions is called the falsification point (MastenPoirier2021). Any $\lambda$ values below this point will lead to a non-falsified model, and therefore a non-empty identified set for ATE. Note that we could use different $\lambda$ values for each potential outcome; we use a common value for simplicity.

(ref) has two implications: First, it has direct identifying power for potential outcomes $\ensuremath{\mathbf{Y}}(1)$ and $\ensuremath{\mathbf{Y}}(0)$, since it imposes a constraint the unobserved values of potential outcomes. Second, it affects $\underline{p}(K)$ because that constraint also shrinks the constraint set for the optimization in equation (ref). We consider each of these implications next.

The Identifying Power of Covariates

In section (ref) we derived an explicit expression for the finite population identified set for ATE. In this section we provide a numerical procedure for computing the identified set for ATE under the additional assumption of predictive covariates ((ref)). Specifically, the upper bound on ATE solves \[ \max_{\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0) \in [y_\text{min},y_\text{max}]^{2N}} \; \frac{1}{N} \sum_{i=1}^N \big( Y_i(1) - Y_i(0) \big) \] subject to (i) the data constraints that $Y_i(X_i) = Y_i$ for all $i=1,\ldots,N$, (ii) the $K$-approximate means balance constraint \[ -K \leq \frac{1}{N_1} \sum_{i=1}^N Y_i(x) \ensuremath{\mathbbm{1}}(X_i = 1) - \frac{1}{N_0} \sum_{i=1}^N Y_i(x) \ensuremath{\mathbbm{1}}(X_i = 0) \leq K, \] and (iii) the predictive covariates constraint $R_{Y(x) \sim q(W)}^2 \geq \lambda$, which is equivalent to

equation[equation omitted — 176 chars of source]

where $\beta(x) = (\ensuremath{\mathbf{Q}}' \ensuremath{\mathbf{Q}})^{-1} \ensuremath{\mathbf{Q}}' \ensuremath{\mathbf{Y}}(x)$. The lower bound can be obtained by computing the minimum rather than the maximum. The constraint (ref) is smooth in the unknowns and the objective function and other constraints are linear, so any gradient-based nonlinear solver can be used to solve this program.

The Impact of Covariates on Design Probabilities of Balance

Because the predictive covariates assumption restricts the feasible values of unobserved potential outcomes, it also affects the design probability that the $K$-approximate mean balance assumption (ref) is guaranteed to hold. Specifically, we modify the definition of $\underline{p}$ to also impose the constraint in equation (ref):

equation[equation omitted — 491 chars of source]

This additional constraint does not meaningfully affect the computational time. Because adding constraints weakly reduces the set of feasible values, $\underline{p}^\text{mod}(K,\lambda)$ is weakly increasing in $\lambda$ for any fixed $K$.

Interpretation and Discussion

To build intuition for the identifying power of covariates, consider the case with a binary covariate $W_i$ and focus on the treated potential outcomes $(Y_1(1),\ldots,Y_N(1))$. Let $q(W_i) = (1,W_i)'$, and denote the corresponding components of $\beta(x)$ by $\beta_0(x)$ and $\beta_1(x)$. For any vector $(A_1,\ldots,A_N)$ let $\ensuremath{\mathbb{E}}(A \mid X=x) \coloneqq \left( \sum_{i=1}^N A_i \ensuremath{\mathbbm{1}}(X_i=x) \right) / \sum_{i=1}^N \ensuremath{\mathbbm{1}}(X_i=x)$. By the definition of the residuals $U_i(1)$, $\ensuremath{\mathbb{E}}[Y(1) \mid X=x] = \beta_0(1) + \beta_1(1) \ensuremath{\mathbb{E}}(W \mid X=x) + \ensuremath{\mathbb{E}}[U(1) \mid X=x]$ for $x \in \{0,1\}$ and hence

align*[align* omitted — 294 chars of source]

This equation shows that the magnitude of imbalance in potential outcomes depends directly on the magnitude of imbalance in covariates. It also depends on the magnitude of imbalance in the residuals, which is controlled by $\lambda$. In particular, for $\delta \coloneqq \sqrt{N (1-\lambda) \operatorname*{var}(Y)}$, $\ensuremath{\mathbb{E}}[U(1) \mid X=x] \in [-\delta,\delta]$ for each $x \in \{0,1\}$. Hence \[ \big| \ensuremath{\mathbb{E}}[U(1) \mid X=0] - \ensuremath{\mathbb{E}}[U(1) \mid X=1] \big| \leq 2 \delta. \] This bound combined with the observed imbalance in covariates directly constrain the imbalance in potential outcomes, which is the source of identifying power of covariates. A similar analysis applies to balance in $Y(0)$. Note that this discussion is for intuition only; section (ref) describes how we obtain identified sets in general.

Random assignment of treatment does not imply anything about the value of $R_{Y(x) \sim q(W)}^2$; it is a population level parameter that does not depend on how treatment is assigned. Consequently, we cannot use random assignment of treatment to calibrate $\lambda$. However, assumptions like (ref) are not necessary for drawing relatively tight conclusions about ATE in large finite populations, as in the second part of theorem (ref). Rather, additional assumptions like (ref) can be useful to provide tighter bounds in smaller finite populations.

We conclude this section with several remarks about the literature. First, covariate balance is sometimes examined to test whether treatment was actually randomized. Here we simply assume treatment was in fact randomized. Second, the traditional design-based inference framework primarily uses covariates to motivate different choices of test statistics, with the goal of increasing the power of hypothesis tests. For example, see ImbensRubin2015 or ZhaoDing2021. This approach, like ours, uses covariates to derive stronger conclusions about the parameter of interest. Mathematically, however, our approach uses covariates for identification and does not rely on hypothesis testing theory. Finally, covariates are also used to define the parameter of interest (e.g., AbadieAtheyImbensWooldridge2020). In principle this can be done in our approach too, but we leave this to future work (also see section (ref)).

Noncompliance

We have focused on the case where all units comply with their treatment assignment. In this section we extend the analysis to a finite population version of the ImbensAngrist1994 model (as in section 3 of HongLeungLi2020), allowing us to study settings with noncompliance. When certain exact balance conditions hold, we show that the Wald estimand point identifies a realized local average effect of treatment on the treated (LATT) parameter. We then study finite population identification under $K$-approximate mean balance type assumptions. In the special case of one-sided noncompliance, we show that the finite population identified set for the realized LATT has a simple form that is analogous to classical infinite population results. We then show how to use this result to do design-based sensitivity analysis. In all of this analysis we allow for heterogeneous treatment effects while still not relying on asymptotics (in contrast to some prior work on design-based instrumental variable analysis, such as Rosenbaum1996).\footnote{See ImbensRosenbaum2005, BaiocchiSmallLorchRosenbaum2010, KeeleSmallGrieve2017, KangPeckKeele2018, and RambachanRoth2025 for additional prior work on design-based instrumental variable analysis.}

Setup

As before, there is a finite population of units $i=1,\ldots,N$. For each unit $i$: Let $Y_i(1)$ and $Y_i(0)$ denote potential outcomes, $X_i(1)$ and $X_i(0)$ the binary potential treatments, $Z_i$ the realized value of a binary instrument (assigned treatment in the noncompliance setting), $X_i = X_i(Z_i)$ the realized treatment, and $Y_i = Y_i(X_i)$ the realized outcome. Note that we impose the exclusion restriction implicitly here and maintain it throughout this section. Without loss of generality, write $Y_i(x) = \beta_i \cdot x + U_i$ for $x \in \{0,1\}$, where we defined $U_i \coloneqq Y_i(0)$ and $\beta_i \coloneqq Y_i(1) - Y_i(0)$. Let $N_1 = \sum_{i=1}^N \ensuremath{\mathbbm{1}}(Z_i = 1)$ denote the number of units whose instrument value equals 1. We'll call this the instrument-on group. Let $N_0 = N - N_1$. We'll call the set of units with $Z_i = 0$ the instrument-off group. Let $\overline{Y}_{z=1} \coloneqq \frac{1}{N_1} \sum_{i : Z_i = 1} Y_i$ denote the average outcome in the instrument-on group. Define $\overline{Y}_{z=0}$, $\overline{X}_{z=1}$, $\overline{X}_{z=0}$, $\overline{U}_{z=1}$, and $\overline{U}_{z=0}$ similarly.

Identification under Exact Balance

Define the compliance type variable \[ T_i =

casesc &if $X_i(1)=1, X_i(0) = 0$ \\ a &if $X_i(1)=1, X_i(0)=1$ \\ n &if $X_i(1)=0, X_i(0)=0$ \\ d &if $X_i(1)=0, X_i(0) =1$.

\] We maintain the following finite population version of the monotonicity / no defiers assumption in ImbensAngrist1994.

Bassumption[No defiers] $T_i \neq d$ for all $i=1,\ldots,N$.

Similarly, we assume the following finite population version of relevance holds for the specific realization of treatment assignment that is observed.

Bassumption[Relevance] $\overline{X}_{z=1} \neq \overline{X}_{z=0}$.

Let $\overline{T}_1(a) \coloneqq \frac{\sum_{i=1}^N \ensuremath{\mathbbm{1}}(Z_i=1) \ensuremath{\mathbbm{1}}(T_i=a)}{\sum_{i=1}^N \ensuremath{\mathbbm{1}}(Z_i=1)}$ denote the proportion of always takers in the instrument-on group. Likewise, let $\overline{T}_0(a)$ denote the proportion of always takers in the instrument-off group, and $\overline{T}_1(c)$ the proportion of compliers in the instrument-on group. Let $\overline{\beta}_1(a) = \left( \sum_{i : Z_i=1, T_i=a} \beta_i \right) / \allowbreak \sum_{i=1}^N \ensuremath{\mathbbm{1}}(Z_i=1) \ensuremath{\mathbbm{1}}(T_i=a)$ denote the average treatment effect among the instrument-on always takers. Define $\overline{\beta}_0(a)$ and $\overline{\beta}_1(c)$ similarly.

lemmaSuppose (ref) (no defiers) and (ref) (relevance) hold. Then \begin{equation} \resizebox{0.94\linewidth}{!}{$ \frac{\overline{Y}_{z=1} - \overline{Y}_{z=0}}{\overline{X}_{z=1} - \overline{X}_{z=0}} = \frac{\overline{U}_{z=1} - \overline{U}_{z=0}}{\overline{X}_{z=1} - \overline{X}_{z=0} } + \frac{\overline{T}_1(a) \overline{\beta}_1(a) - \overline{T}_0(a) \overline{\beta}_0(a)}{\overline{X}_{z=1} - \overline{X}_{z=0} } + \frac{\overline{T}_1(c)}{\big( \overline{T}_1(a) - \overline{T}_0(a) \big) + \overline{T}_1(c)} \; \text{\footnotesize $\overline{\beta}_1(c)$}. $} \end{equation}

The left hand side of equation (ref) is the finite population version of the Wald estimand. So this lemma decomposes the Wald estimand into three pieces that each depend on the magnitude of various realized imbalances. Consider the following exact mean balance assumption.

Bassumption[Exact balance] (i) $\overline{U}_{z=1} = \overline{U}_{z=0}$, (ii) $\overline{T}_1(a) = \overline{T}_0(a)$, and (iii) $\overline{\beta}_1(a) = \overline{\beta}_0(a)$.

Part (i) says that $Y_i(0)$ is balanced across the instrument-on and instrument-off groups. This is an instrument exogeneity assumption, with respect to the potential outcomes. It is analogous to $Z \mathbin{ \mathpalette{\@indep}{} } Y(0)$ in an infinite population analysis. Part (ii) says that the proportion of always takers is the same in the instrument-on and instrument-off groups. It is also an instrument exogeneity assumption. It is often called “unconfounded types”, because it is about the relationship between the instrument and the potential treatment variables. It is analogous to $Z \mathbin{ \mathpalette{\@indep}{} } (X(1), X(0))$ in an infinite population analysis. Finally, part (iii) says that the average treatment effect for always takers is the same in the instrument-on and instrument-off groups. In an infinite population analysis, this kind of mean balance condition would hold if $Z \mathbin{ \mathpalette{\@indep}{} } (Y(0), Y(1), X(0), X(1))$.

propositionSuppose (ref) (no defiers), (ref) (relevance), and (ref) (exact balance) hold. Then $\frac{\overline{Y}_{z=1} - \overline{Y}_{z=0}}{\overline{X}_{z=1} - \overline{X}_{z=0}} = \overline{\beta}_1(c)$.

Proposition (ref) shows that $\overline{\beta}_1(c)$ is point identified in finite populations under exact balance. In particular, it equals the finite population Wald estimand. This result is a finite population version of the ImbensAngrist1994 result, with one slight difference: The point identified parameter $\overline{\beta}_1(c) \coloneqq \left( \sum_{i=1}^N \beta_i \cdot \ensuremath{\mathbbm{1}}(T_i=c) \ensuremath{\mathbbm{1}}(Z_i=1) \right) / \sum_{i=1}^N \ensuremath{\mathbbm{1}}(T_i=c) \ensuremath{\mathbbm{1}}(Z_i=1)$ is a realized local average effect of treatment on the treated (LATT)---it is the average unit-level causal effect among treated compliers (recall that $Z_i=X_i$ for compliers). In contrast, the population LATE is $\overline{\beta}(c) \coloneqq \left( \sum_{i=1}^N \beta_i \ensuremath{\mathbbm{1}}(T_i=c) \right) / \sum_{i=1}^N \ensuremath{\mathbbm{1}}(T_i=c)$. The LATT is an ex ante random parameter, since the set of units who will be treated, and thus who appear in the parameter's definition, depends on the realization of treatment assignment. This is analogous to Rosenbaum's Rosenbaum2001 finite population analysis of ATT, where he noted that the ATT parameter is also ex ante random (he called the ATT the “attributable effect”). Another slight difference from the usual infinite population analysis is that when there is one-sided noncompliance (so there are no always takers), the ATT and LATT are the same, but they do not equal LATE because there can be non-treated compliers. Finally, note that we can decompose LATE as $\overline{\beta}(c) = p_1(c) \overline{\beta}_1(c) + p_0(c) \overline{\beta}_0(c)$ where $p_1(c) \coloneqq \left( \sum_{i=1}^N \ensuremath{\mathbbm{1}}(T_i=c) \ensuremath{\mathbbm{1}}(Z_i=1) \right) / \sum_{i=1}^N \ensuremath{\mathbbm{1}}(T_i=c)$ is the proportion of units who are treated, among all compliers, and likewise for $p_0(c)$. So if we further assume that the average treatment effect for compliers in the instrument-on group is the same as for compliers in the instrument-off group---$\overline{\beta}_1(c) = \overline{\beta}_0(c)$---then the finite population Wald estimand equals LATE. This additional condition is analogous to part (iii) of (ref).

Design-Based Sensitivity Analysis

Proposition (ref) shows that the Wald estimand equals LATT under an exact balance assumption. However, as discussed in section (ref), exact balance rarely holds in finite populations. Instead, we can apply the same ideas from earlier to perform a design-based sensitivity analysis: We can derive identified sets for LATT under approximate balance assumptions and then use randomization to calibrate the sensitivity parameters. Here we briefly sketch the analysis in the one-sided noncompliance case.

Bassumption[One-sided noncompliance] $T_i \neq a$ for all $i=1,\ldots,N$.

Without always takers, the second term in equation (ref) disappears, and the third term becomes $\overline{\beta}_1(c)$. Hence the only remaining term is a difference in non-treated potential outcomes, which we bound via the following assumption.

Bassumption[$K$-approximate mean balance for $Y(0)$] There is a known $K \geq 0$ such that $| \overline{U}_{z=1} - \overline{U}_{z=0} | \leq K$.

The next result bounds the realized LATT as a function of $K$. Here we let $\pi \coloneqq \overline{X}_{z=1} - \overline{X}_{z=0}$ denote the first stage difference in means.

theoremSuppose (ref) (no defiers), (ref) (relevance), (ref) (one-sided noncompliance), and (ref) ($K$-approximate mean balance for $Y(0)$) hold. Then the finite population identified set for $\overline{\beta}_1(c)$ is \[ \left[ \frac{\overline{Y}_{z=1} - \overline{Y}_{z=0}}{\overline{X}_{z=1} - \overline{X}_{z=0}} - \frac{K}{\pi}, \; \frac{\overline{Y}_{z=1} - \overline{Y}_{z=0}}{\overline{X}_{z=1} - \overline{X}_{z=0}} + \frac{K}{\pi} \right]. \]

The identified set in theorem (ref) is analogous to infinite population identified sets where instrument exogeneity is relaxed at the population level. For example, see BoundJaegerBaker1995 or ConleyHansenRossi2012. The identified set in theorem (ref) can be further adjusted to impose the bounded outcome assumption (ref), like in theorem (ref). Then, assuming the instrument is randomly assigned according to a known distribution, we can construct a function similar to $\underline{p}$ in equation (ref) and use this to calibrate the value of $K$. We omit the details for brevity. The general two-sided noncompliance case is more complicated, since it involves more than just a single balance condition (i.e., relaxations of the three conditions in (ref)). We conjecture that our analysis extends to this case but leave a full exploration to future work.

Further Extensions

Sampling

For most of this paper we assumed for simplicity that there was no sampling---all units in the finite population are observed, and the only uncertainty is about the unknown potential outcomes. Here we briefly discuss two extensions: An analysis of sampling by itself, and an analysis that combines both sampling and random assignment (e.g., as in AbadieAtheyImbensWooldridge2020). In both cases we reframe the classical question of inferring population quantities from sample data as an identification problem, which allows us to provide a single approach to quantify uncertainty from missing units and from missing potential outcomes. Keep in mind that this extension will generally make bounds wider, since it accounts for an additional source of uncertainty.

First consider the sampling setup from section (ref). The population are the numbers $\ensuremath{\mathbf{Y}} = (Y_1,\ldots,Y_N)$. Now suppose we only observe $n < N$ of these units, and the goal is to learn about the population mean $\overline{Y} \coloneqq \frac{1}{N} \sum_{i=1}^N Y_i$ (rather than a treatment effect parameter). As before, $S_i \in \{0,1\}$ denotes whether unit $i$ is sampled or not. From an identification perspective, sampling is a missing data problem---we observe $Y_i$ when $S_i=1$ but have no data whatsoever on unit $i$ when $S_i=0$. Consequently, if all we know are that all outcomes $Y_i$ lie in known bounds (an assumption similar to (ref)), then all we can say for sure about the population mean is that it lies in domain-of-consensus bounds analogous to $\Theta_I(\infty)$; this is the motivation behind including sample indicators in definition (ref). However, we can shrink the identified set by making assumptions like \[ \left| \frac{1}{n} \sum_{i=1}^N S_i Y_i - \frac{1}{N-n} \sum_{i=1}^N (1-S_i) Y_i \right| \leq K, \] which is analogous to the $K$-approximate mean balance assumption (ref). Under this assumption, we can derive identified sets for $\overline{Y}$ as a function of the sensitivity parameter $K$. Finally, suppose we know the sample was obtained via simple random sampling (SRS), for example. This is the sampling analog of uniform randomization ((ref))---all possible samples of size $n$ from $N$ have equal probability of being selected. Then, as in section (ref), we can compute worst case design probabilities of imbalance, which we can use to calibrate the sensitivity parameter $K$. This allows us to perform a design-based sensitivity analysis for sampling.

Next consider the randomized experiment setting we focused on in this paper. Suppose we only observe a sample of $n < N$ units. In this case we need to address the identification problem that arises from missing potential outcomes as well as missing data on some units altogether (see Manski1996 and KlineTamer2018 for related identification analyses that combine experiments and sampling, but in the infinite population setting). This setting does not require any new conceptual ideas, and so we only discuss it briefly. Consider the average treated potential outcome. Let $n_1 < n$ denote the number of observed treated units. Assume both $n$ and $n_1$ are fixed a priori, and both sampling and randomization are performed independently according to simple random sampling and uniform randomization. We observe the average treated potential outcome among sampled units who are treated, $\frac{1}{n_1} \sum_{i : S_i=1, X_i=1} Y_i(1)$. We do not know the average $\frac{1}{N-n_1} \sum_{i : S_i=0 \text{ or } X_i=0} Y_i(1)$. However, we can make a $K$-approximate mean balance assumption that says this unobserved mean is not too far from the observed one. Then we can use our sampling and randomization assumptions to calibrate the value $K$. The same analysis can be done for the non-treated potential outcome, and they can be combined to do a design-based sensitivity analysis for the population ATE.

The Distinction Between Being Identified Versus Identifiable

Our definition (ref) treats sampling and randomization symmetrically---both lead to missing data which can be thought of as a finite population identification problem, and therefore analyzed using the tools of partial identification. This is perhaps unusual, since uncertainty due to sampling is traditionally separated from other types of uncertainty, as in Koopmans' (1949, page 132) original definition of identification. Nonetheless, as sketched above, definition (ref) provides a foundation to quantify uncertainty from both missing units and missing potential outcomes. That said, here we briefly discuss a distinction between the two, which helps reconcile the difference between our definition (ref) and traditional definitions of identification in infinite populations like Koopmans' Koopmans1949 which assume away sampling uncertainty.

Consider the set of all designs, which are mappings from $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0),\ensuremath{\mathbf{W}})$ into joint probabilities over sampling and treatment assignments. Say a parameter is point identifiable with respect to a set of assumptions $\mathcal{P}$ if there exists a design such that, for any $(\ensuremath{\mathbf{Y}}(1),\ensuremath{\mathbf{Y}}(0),\ensuremath{\mathbf{W}})$ satisfying the assumptions $\mathcal{P}$, and for any realization $(\ensuremath{\mathbf{S}},\ensuremath{\mathbf{X}})$ from that design, the subsequent identified set $\Theta_I$ is a singleton. Then the population mean $\overline{Y} \coloneqq \frac{1}{N} \sum_{i=1}^N Y_i(X_i)$ is point identifiable with respect to (ref), since we can in principle sample all units with probability 1. But in any fixed dataset with $n < N$ it will only be partially identified. In contrast, ATE is not point identifiable with respect to (ref) since we can never treat everyone and treat no-one. Thus, even though our definition of the identified set treats the ex post uncertainty due to missing units and missing potential outcomes symmetrically, our definition still allows these two kinds of uncertainty to be ex ante different, since one can in principle sample all units whereas we cannot even in principle treat all units.

Variations on the Bounded Outcomes Assumption

Obtaining nontrivial bounds on ATE usually requires some kind of assumption that restricts the range of possible potential outcomes. Throughout this paper we used the familiar uniform bound $[y_\text{min}, y_\text{max}]$ on potential outcomes, assumption (ref). It is important to recognize that this is a substantive identifying assumption, not a regularity condition---if these bounds are large then the researcher is explicitly allowing for a wide range of possible outcomes, which allows for large magnitudes of imbalance.

The specific form of the assumption we used can be replaced or augmented with a variety of similar assumptions, and all of our methods will continue to apply. Here we give just a few examples: (i) A simple extension is to allow unit specific bounds, $Y_i(x) \in [y_{\text{min},i}, y_{\text{max},i}]$. We use this version in one of our empirical applications, where the units are groups and the outcome depends on group size. (ii) One could impose a bounded unit-level treatment effect assumption: $| Y_i(1) - Y_i(0) | \leq M$ for all $i \in \mathcal{I}$, where $M$ is a known sensitivity parameter. This assumption implies unit-specific bounds on the unobserved potential outcomes: $Y_i(x) \in [Y_i - M, Y_i + M]$. (iii) Alternatively, one could assume the sum of unit-level treatment effect magnitudes is bounded: $\sum_{i=1}^N | Y_i(1) - Y_i(0) | \leq M$ for a known $M$. This would allow for some units to have very large treatment effects, so long as not too many do. This condition implies that ATE can be no larger than $M / N$. (iv) Or one could restrict the population variance of potential outcomes to be no larger than a known $M$. This assumption implies that $Y_i(x)$ is in the interval $\frac{1}{N_x} \sum_{i=1}^N Y_i \ensuremath{\mathbbm{1}}(X_i = x) \pm 2 \sqrt{N_x M}$. Hence it does not require users to specify a priori bounds $[y_\text{min},y_\text{max}]$; they only must choose $M$. (v) Finally, the covariate restrictions in section (ref) can also be viewed as one way of restricting the range of possible potential outcomes.

Distributional Balance and Parameters Beyond ATE

Our analysis has focused on balance defined via differences in means. This corresponds to our focus on ATE, where balance in means is the “fundamental identification condition” (page 263, HeckmanIchimuraTodd1998). However, there are many other ways to quantify balance across the treatment and control groups; see chapter 14 of ImbensRubin2015. For example, letting $\ensuremath{\mathbb{P}}^\text{true}(Y(x) \leq y \mid X=x) \coloneqq \frac{1}{N_x} \sum_{i : X_i=x} \ensuremath{\mathbbm{1}}[Y_i(x) \leq y]$ (see appendix (ref)), we could consider the assumption

equation[equation omitted — 215 chars of source]

which bounds the sup-norm distance between the population distributions of potential outcomes in the treatment and control groups. We could then derive identified sets for the parameter of interest under this assumption, for a fixed $K$. Given a randomization design, we could then compute the worst case design-probability that (ref) holds, which would lead to a design-based sensitivity analysis.

Different forms of balance might lead to different bounds than those based on mean balance. Thus the choice of balance metric in our analysis could be loosely thought of as analogous to the choice of the test statistic in classical randomization tests. Exactly how much the bounds vary with the choice of balance metric will likely depend on the parameter of interest. Alternative balance metrics may be more appropriate for studying parameters beyond ATE. For example, the sup-norm distance in equation (ref) will likely work well for identification of quantile treatment effects (QTEs) since these are defined using inverses of the unconditional population cdfs $\ensuremath{\mathbb{P}}^\text{true}(Y(x) \leq y) \coloneqq \frac{1}{N} \sum_{i=1}^N \ensuremath{\mathbbm{1}}[Y_i(x) \leq y]$. We leave a full exploration of these questions to future work.

Empirical Applications

In this section we illustrate our approach in two empirical applications with small, finite populations ($N=10$ and $N=17$). While our methods apply to any population size, and are feasible for larger population sizes (see appendix (ref) for a third application with $N=722$), these two applications show that it is still possible to do meaningful inference in small datasets. In appendix (ref) we show three alternative frequentist confidence intervals for comparison.

Getting Parents to Pick Their Kids Up on Time

Our first application uses data from the well known paper GneezyRustichini2000, which has about 3600 Google Scholar citations as of March 2026. They study day care centers, where administrators were frustrated with parents showing up late to pick up their kids. They asked: Would a monetary fine incentivize parents to be on time? The population is 10 centers. The treatment is the introduction of a center-wide late fee. 6 centers were treated and 4 were not. The outcome variable is the number of late parents in a week. The logical lower bound is zero and the logical upper bound is five times the number of kids in that center (assuming every parent is late every day of the week). We assume that regardless of treatment, on average, each child has a late parent at most once per week. That is, we set $y_{\text{max},i}$ equal to the number of children in center $i$ (see (i) in section (ref)). In the data there are between 28 and 37 kids per center. We let $y_{\text{min},i} = 0$ for all units.

The authors gathered baseline data on outcomes for 4 weeks. The fee was introduced at treated centers at the beginning of week 5. It was removed at the beginning of week 17. The authors gathered another 4 weeks of data after removal of the fee, for a total of 20 weeks of data. Their table 1 provides the full dataset. For simplicity we only use data from one post-treatment week, week 19. It would be interesting to study how to extend our results to use the panel dimension of this dataset, but we leave this to future work.

The ATE point estimate is $13$ late parents, suggesting that the addition of a fee increased the number of late parents per week by 13. To quantify the uncertainty around this estimate, we conduct a design-based sensitivity analysis. Figure (ref) in the introduction shows the results, which we discussed already. First consider the breakdown point, $\underline{p}(K^\text{bp}) = 65$%. So there is at least a 65% chance that the ATE is non-negative, according to our robust empirical Bayesian interpretation. Although this does not attain conventional levels of “significance”, like 95%, it is nonetheless a non-trivial inference given that this dataset only contains 10 units. This conclusion can also be seen in figure (ref), which plots $\Theta_I(K(\alpha))$ as a function of $1-\alpha$. All sets with probabilities smaller than 64% contain only positive values. Moreover, even for large probabilities, the sets $\Theta_I(K(\alpha))$ are still mostly in the positive region. For example, the 90% set is about $[-6, 22]$. There is at least a 90% chance that the true ATE is in this set. Overall, these findings suggest that there is reasonably strong evidence that the ATE is positive.

The Long Run Adoption of Management Practices

\nocite{BloomData}

Our second application uses data from BloomEtAl2013 and BloomEtAl2020. These are influential papers with about 2650 total Google scholar citations as of March 2026. These papers asked whether large observed differences in productivity across firms are driven by variation in firms' management practices. To answer this, they ran a randomized experiment in a population of 17 woven cotton fabric firms. These were large and old firms, with an average of 270 employees per firm and an average age of 20 years old at baseline. Their control group received a one-month diagnostic about their management practices. The treatment group received the diagnostic plus four months of support for implementing the management changes. They consider many different outcomes of interest, but we will focus on the long run outcome from their 2020 paper. Specifically, the outcome variable is the proportion of 38 management practices adopted in 2017, which was about 8 to 9 years after they received treatment. Since this outcome is a proportion, we set $[y_\text{min}, y_\text{max}] = [0,1]$.

figure[figure omitted — 842 chars of source]

Some of the 17 firms in the population have multiple plants. Treatment was administered at and varies at the plant-level. This implies that there are non-treated plants at firms with a different treated plant. The authors use this data to study within-firm spillovers. For simplicity we ignore all data from non-treated plants at treated-firms. We use the authors' “experimental” dataset (as described in figure 1 of BloomEtAl2020), where each unit is a single plant. There are 11 treated plants and 6 control plants.\footnote{See pages 16--17 of BloomEtAl2013 for more details on the experimental design.}

The ATE point estimate is 0.13, suggesting that providing 4 months of support for changing management practices leads to a 13 percentage point increase in the proportion of practices that are still in place about one decade later. The authors emphasize that it is important to also measure the uncertainty associated with this point estimate, however:

quote“The major challenge of our experiment was its small cross-sectional sample size. We have data on only 28 plants across 17 firms. To address concerns over statistical inference in small samples, we implemented permutation tests whose properties are independent of sample size.” (page 4, BloomEtAl2013)

That is, to deal with the small population size they performed exact tests of the sharp null of no unit-level treatment effects, assuming uniform randomization. To complement their results, we implement our design-based sensitivity analysis. Figure (ref) shows the main results. First consider the breakdown point, $\underline{p}(K^\text{bp}) = 20$%. So there is at least a 20% chance that the ATE is non-negative, according to our robust empirical Bayesian interpretation. So with this dataset it is unlikely that potential outcomes would be balanced enough to ensure that $K$ is small enough that we can rule out negative ATE values. This is also shown in the outer-most bounds of figure (ref), which plots $\Theta_I(K(\alpha))$ as a function of $1-\alpha$. Here we see that the sets all contain negative numbers for probabilities larger than 20%. For example, the 90% interval is $[-0.28, 0.49]$. So there is at least a 90% chance that ATE is inside this interval.

We can obtain tighter bounds by adding assumptions about covariates, as in section (ref). Figure (ref) also shows bounds on ATE that impose the predictive covariates restriction (ref) for $\lambda = 0$ (no covariate restrictions), $0.2$, and $0.4$, going from the darkest, outer-most bands to the lightest, inner-most bands. We stop at $\lambda = 0.4$ since larger values are generally falsified (and hence lead to empty identified sets). Following the authors' analysis (table 2 in BloomEtAl2020), we only use a single covariate, baseline management score in 2008, which is essentially a lagged outcome, and specify a simple linear model $q(W_i) = (1,W_i)'$. For this choice, the predictive covariates assumption tightens the bounds somewhat, bringing the breakdown point up to about 45%, compared to its unconstrained value of 20%. The bounds could potentially be further tightened by using additional covariates (e.g., those in their table 1), or by making additional assumptions about the unobserved potential outcomes, as in section (ref).

Conclusion

In this paper we studied identification in finite populations. We showed that the prior conventional definition of identification implies that every potential outcome for every unit is point identified under an overlap condition alone, even in observational settings where treatment depends on potential outcomes. Hence we argued that it is not adequate for understanding what can be learned about treatment effects in real datasets.

As our first main contribution, we developed a definition of identification that is not subject to this problem. Moreover, unlike the traditional design-based approaches to inference in finite populations, our definition of identification does not pre-suppose a known distribution from which treatment, an instrument, or a “shock,” was drawn. This allows it to be used in a wide variety of settings, including both experimental and observational studies.

Building on our definition of identification, in our second main contribution we developed a new approach to quantifying uncertainty in finite populations, called a design-based sensitivity analysis. This approach combines design distributions and our finite population identified sets to construct new uncertainty intervals that have (1) identification, (2) robust Bayesian, and (3) uniform frequentist interpretations.