Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,442 characters · 23 sections · 46 citation commands
On the relationship between prediction intervals, tests of sharp nulls and inference on realized treatment effects in settings with few treated units
\onehalfspacing
Inference on treatment effects in settings with few treated units presents significant challenges alvarez2025inferencetreatedunits. When the number of treated units is small --- with the extreme case being a single treated unit --- there is limited information on potential outcomes under treatment, making uncertainty quantification a difficult task. As a consequence, inference methods that remain valid in the presence of few treated units often rely on strong assumptions, particularly regarding treatment effect heterogeneity. In this paper, we focus on methods derived in “model-based” settings in which treatment assignment is viewed as fixed (or conditioned on) and where uncertainty stems from stochasticity in potential outcomes.\footnote{See alvarez2025inferencetreatedunits for a discussion on the appropriateness of such modeling in settings with few treated units} In this scenario, many of the existing approaches rely on treatment effect homogeneity assumptions, imposing that treatment effects are non-stochastic. Approaches that allow for stochastic treatment effects typically shift the inferential focus: instead of constructing confidence intervals for the average treatment effect on the treated (ATT), they focus on (i) testing sharp null hypotheses, (ii) conducting inference on the realized treatment effect, or (iii) constructing prediction intervals.
This first alternative, rather than testing hypotheses about the ATT, tests sharp null hypotheses, which assess whether the treatment had no effect on any treated unit.\footnote{More generally, these tests can also be used to test a null that the treatment effect takes a specific value for each treated unit, with probability one.} While commonly used in design-based approaches Imbens_matching,young_QJE, this strategy has also been applied in model-based settings Chung2021,Bugni2018,LeeShaikh2014,Heckman2024. This second alternative seeks to provide inferential statements concerning the realization of treatment effects in the sample at hand, instead of averaging over possible realizations of treatment effects. This inferential target is particularly relevant when the goal is to understand the specific context in which treatment was delivered rather than to generalize to other settings. There are many papers that consider treatment effects as deterministic but potentially heterogeneous across individuals, which can be implicitly seen as a setting in which we condition on the stochastic treatment effects conley20211inference,carvalho2018arco,ferman2019inference,synthetic_did,alvarez2023extensions,alvarez2023inference,chernozhukov2024ttest. Finally, this third alternative constructs set-valued functions of the data --- prediction intervals --- that aim to cover a random variable with a given confidence over repeated samples. These intervals have a long tradition in the forecasting literature (e.g. Phillips1979,Brockwell1991). More recently, this type of construction has been considered in causal inference settings with the explicit goal of constructing bands that cover the in-sample treatment effects on the treated (viewed as random variables) with a pre-specified probability Candes,kivaranovic2020conformal,chernozhukov2021exact,CWZ2021Pnas,cattaneo2021prediction,cattaneo2023uncertainty.
This note provides a unifying perspective on these different inferential approaches for settings with few treated units. As a leading example, we consider first a difference-in-means estimator and then extend our results to more general estimators. We show that all these inferential approaches are deeply interconnected: they are either equivalent under the same set of assumptions required for their individual validity or become equivalent under additional assumptions. Some of these equivalences are straightforward, while others require novel results that we derive in this paper.
First, we show that inference methods for the ATT that remain valid with few treated units under homogeneous treatment effects are also valid for testing sharp null hypotheses, and that, conversely, any test of sharp nulls provides a valid test of the ATT under treatment effect homogeneity. This follows directly from the fact that imposing a sharp null is equivalent to assuming, under the null, that treatment effects are non-stochastic and homogeneous.
Next, we show that a broad class of inference methods originally developed under the assumption of non-stochastic and homogeneous treatment effects can also be applied in settings with stochastic treatment effects for inference on the realized treatment effect, under strong assumptions --- a sufficient one being independence between the (stochastic) treatment effects and the potential outcomes under no treatment. To the best of our knowledge, this discussion of the assumptions required for valid inference on the realized treatment effect is novel.
We also derive new results providing conditions under which inference methods valid for the realized treatment effect, when treatment effects are independent of untreated potential outcomes, can also be used to construct prediction intervals, even when treatment effects and untreated potential outcomes exhibit arbitrary dependence. This is the case when we have a consistent estimator for the distribution of the untreated potential outcomes of the treated (without conditioning on the realized treatment effect).
Finally, we show that prediction interval construction and sharp null hypothesis testing are intrinsically linked: under a broad class of inference methods used in small-sample settings, each can be obtained by “inverting” the other, remaining valid under the same set of assumptions. While such equivalence has been noted in specific settings (e.g. Lei2013 and chernozhukov2021exact), we believe that, by laying it out in a general setting, we contribute to clarifying the interpretation of prediction intervals. In particular, we show that the usual practice of checking whether prediction intervals contain zero can always be seen as a valid test of the sharp null that treatment effects are homogeneous and equal to zero in the treated population, when uncertainty stems from sampling from a well-defined population. The latter connection is not specific to a few treated setting, being true in any situation where prediction intervals for treatment effects are reported.
Overall, by establishing the connections among these approaches, this paper provides new theoretical justifications for inference methods originally developed under the assumption of homogeneous treatment effects (for example, conley20211inference and ferman2019inference), demonstrating their validity for alternative inferential targets in settings where treatment effects are stochastic. Specifically, we show that, in settings with stochastic treatment effects, inference methods developed under treatment effect homogeneity (i) remain valid for testing sharp null hypotheses; (ii) remain valid for inference on the realized treatment effect under strong assumptions on the dependence between treatment effects and untreated potential outcomes; and (iii) can be used to construct prediction intervals under arbitrary dependence between treatment effects and untreated potential outcomes, subject to additional assumptions.
The aim of this section is to illustrate our main results by means of a simple example.
Consider a setting where a researcher has access to a sample with $N$ units from a population of interest. There is an outcome of interest $Y$ and a policy intervention affecting a subset of the units in the population. We denote by $D$ the indicator for the units exposed to the intervention.
We define potential outcomes $Y(1)$ and $Y(0)$, corresponding to the outcomes that would have been observed for a unit in the population were the unit assigned, respectively, treatment and non-treatment. Observed outcomes are thus given by $Y = Y(0) + D(Y(1) - Y(0))$. We consider a model-based approach, in which we condition on treatment assignment and focus on uncertainty coming from potentially unobservable shocks that determine the potential outcomes (see alvarez2025inferencetreatedunits for further discussion on the use of model-based designs in settings with few treated units).
Consider the difference-in-means estimator that is computed with a sample of $N$ units, with $N_1$ being treated and $N_0$ not. This estimator is given by:
where $\alpha_i = Y_{i}(1)-Y_{i}(0)$ is the individual treatment effect.
Under the mean independence assumption, $\mathbb{E}[Y_{i}(0)|D_1,\ldots, D_N] =\mu_0$, for every $i=1,\ldots, N$, $\hat{\beta}_{\text{DM}}$ is unbiased for the expected sample average treatment effect on the treated, i.e. $$\mathbb{E}[\hat\beta_{\text{DM}}|D_1,\ldots, D_N] = \frac{1}{N_1}\sum_{i=1}^N D_i \mathbb{E}[\alpha_i|D_1,D_2,\ldots D_N] \eqqcolon \beta. $$
The parameter $\beta$ captures the average expected effect of the intervention on the treated units in the sample, where expectations are taken with respect to the distribution of economic uncertainty determining potential outcomes.
Even though $\hat \beta_{\text{DM}}$ is unbiased for $\beta$, it is difficult to quantify uncertainty regarding it when $N_1$ is small. To see this, observe that:
Uncertainty regarding $\hat \beta_{\text{DM}}$ can be decomposed into uncertainty concerning treatment effect heterogeneity on the treated units, $D_i(\alpha_i - \beta)$, uncertainty regarding untreated potential outcomes in the treatment group, $D_iY_{i}(0)$, and uncertainty regarding untreated potential outcomes in the control group, $(1-D_i)Y_{i}(0)$.
In a setting with a fixed number of treated units but many controls, we can estimate the distribution of controls $(1-D_i)Y_{i}(0)$ when $N_0\rightarrow \infty$ under weak dependence assumptions conley20211inference,alvarez2023inference. Moreover, if we assume that the distribution of $D_iY_{i}(0)$ for the treated is the same as the distribution of $(1-D_i)Y_{i}(0)$ for the controls, then we can extrapolate the information from the controls to quantify uncertainty on the untreated potential outcomes of the treated, even when we have only a single treated observation. For example, we can strengthen the mean independence assumption to the following assumption.
In this case, the distribution of potential outcomes in the control group may be used to infer the distribution in the treatment group conley20211inference. This assumption can be relaxed to accommodate heteroskedasticity depending on a set of observed covariates, as shown in ferman2019inference. See alvarez2025inferencetreatedunits for other alternatives, including panel data settings in which such extrapolations may come from pre-treatment periods.
Quantifying this type of uncertainty is more complicated, because it depends on potential outcomes under treatment, for which, in settings with few treated units, we observe limited information. In the extreme case in which there is only a single treated unit, we observe only a single observation of $Y_i(1)$.
We describe here common alternatives that have been considered in the literature to circumvent this problem in settings with few treated units (including the case with $N_1=1$).
In this case, we assume that there exists a constant $\alpha$ such that $\alpha_i=\alpha$ for every treated unit $i$. This means that the distribution of the $Y_i(1)$ for treated units can be obtained from a (common) location shift from the distribution of $Y_i(0)$. In this case, for each unit $i$, the treatment effect would be the same across all realizations of the uncertainty that determine the potential outcomes, and homogeneous across units.
Under this (arguably strong) assumption, $\beta = \alpha$, and variability due treatment effect heterogeneity disappears from the distribution of $\hat \beta_{\text{DM}}$. Therefore, if we account for the distribution of $Y_i(0)$ for both treated and control units using one of the methods discussed in Section (ref), then we can conduct valid inference.
For example, consider the case in which only $i=1$ is treated. Under treatment effect homogeneity and Assumption (ref), conley20211inference note that, when $N_0 \rightarrow \infty$, $(\hat \beta_{\text{DM}} - \beta) \overset{d}{\rightarrow} (Y_1(0) - \mathbb{E}[Y_1(0)])$, where the distribution of $(Y_1(0) - \mathbb{E}[Y_1(0)])$ can be estimated using the control residuals to construct p-values that are asymptotically valid when $N_1$ is fixed and $N_0 \rightarrow \infty$. More specifically, if we set $\hat p_{CT} = \frac{1}{N} \sum_{i=1}^N \mathbf{1}\left\{ |\hat \beta_{\text{DM}}-c| \leq |\hat u_i | \right\}$, where $\hat u_i$ are the residuals, this would provide an asymptotically valid p-value (when $N_0 \rightarrow \infty$) for the null $H_0: \beta = c$.
\begin{example_contd}[(ref)] In this example, treatment effect homogeneity would mean that the treatment would have the same causal effect on agricultural outcomes irrespectively of whether we had a positive or negative weather shock, which may be an unreasonable assumption in many settings. We can imagine settings in which the treatment effect would be stronger when farmers were hit by a negative shock, which would invalidate the assumption of treatment effect homogeneity. \end{example_contd}
Another possibility consists of considering a different null hypothesis. Instead of testing a hypothesis regarding the ATT, $\beta$, we could test a sharp null Chung2021,Bugni2018. In this case, the researcher is interested in testing, for some $c \in \mathbb{R}$:
This null implies that treatment effects on treated units are constant over repeated realizations of sampling uncertainty, and equal to some value $c$. More commonly, by setting $c=0$ we have a null that treatment effect has no effect whatsoever, meaning that for (almost) every realization of uncertainty on potential outcomes, the treatment has no effect on any treated unit.
Under the null (ref), Equation (ref) subsumes to:
$$\hat \beta - c = \frac{1}{N_1}\sum_{i=1}^N D_i Y_i(0) - \frac{1}{N_0}\sum_{i=1}^N (1-D_i)Y_i(0) \, .$$
Consequently, we can construct valid tests for the sharp null by leveraging methods that quantify the uncertainty in untreated potential outcomes (Section (ref)). For example, under Assumption (ref), the statistic
where $$\hat{G}(x) = \frac{1}{|\Pi|}\sum_{\pi \in \Pi} \mathbf{1}\left\{\left| \frac{1}{N_1}\sum_{i=1}^N D_{i} (Y_{\pi(i)} - c D_{\pi(i)} )- \frac{1}{N_0}\sum_{i=1}^N (1-D_i)(Y_{\pi(i)} - c D_{\pi(i)} ) \right|\leq x\right\}\, ,$$ and $\Pi$ is the set of permutations on $\{1,\ldots, N\}$, is a valid p-value for testing the sharp null. This property follows from standard results on randomization tests (see Appendix A of alvarez2025inferencetreatedunits for details). When Assumption (ref) does not hold, it is not generally possible to construct exact tests. However, when other restrictions enabling extrapolation from the control group are assumed it can be possible to construct tests that are asymptotically valid as $N_0\rightarrow \infty$ ferman2019inference.
\begin{example_contd}[(ref)] In our example, by setting $c=0$, we would be testing a null hypothesis that the technical assistance has no effect whatsoever on agricultural outcomes of the treated units, regardless of the realizations of the random variables that determine the potential outcomes. \end{example_contd}
A third alternative for inference in settings with few treated units consists in targeting the realized sample average treatment effect, i.e. one is concerned in constructing a test statistic for the null
that controls size conditionally on the realized treatment effects $\mathcal{A} = \{\alpha_i: D_i=1\}$, i.e. one considers a decision rule $\tilde{\phi}_c$ with the property that $\mathbb{E}[\tilde{\phi}_c|D_1,\ldots, D_n, \mathcal{A}] \leq \gamma$ if (ref) holds. We can interpret settings that consider $\alpha_i$ as non-stochastic but potentially different across $i$ conley20211inference,carvalho2018arco,ferman2019inference, synthetic_did,alvarez2023inference,alvarez2023extensions,chernozhukov2024ttest as implicitly considering a setting with stochastic treatment effects but conditioning on the realized effects. In this case, constructing a test that controls size can be seen as equivalent, in a setting where effects are stochastic, to construct a test that controls size conditionally on $\mathcal{A}$.
This inferential focus is especially pertinent when the objective is to learn about the treatment effect in the particular context where the intervention occurred, rather than to extrapolate to other environments. Pragmatically, in the extreme case of a single treated unit, ATT inference becomes infeasible if treatment effects are stochastic, since we only observe one realization of the treated potential outcome. In this case, focusing on uncertainty quantification for the realized treatment effects observed in the sample may be more feasible.
For inference on realized effects to be valid, one requires that the assumptions discussed in Section (ref) are valid even when we condition on $\mathcal{A}$. As we discuss in more detail in Section (ref), in many settings this would require relatively strong assumptions.
\begin{example_contd}[(ref)] In our example, we may consider settings in which the technical assistance is only helpful when there is a negative weather shock, so the effect could be $\tau$ if there is a negative weather shock in farm $i$ (which happens with probability $\pi$) and zero otherwise. In this case, the ATT would be $\tau \times \pi$, but we would be drawing inference on either $\tau$ or $0$, depending on whether or not we had a negative weather shock. In this setting, inference on the realized treatment effect could be interesting if the goal is understanding the effects of that particular implementation of the program. In contrast, if the goal is generalizing to the expected effect across different scenarios, then the ATT would be a more interesting target parameter. In settings with only a single treated unit, note that we only observe a realization of the treated outcome either when there was or when there was not a negative weather shock. Therefore, there would not be much hope in drawing inferences on the ATT, but we could have some hope in drawing inferences on the realized treatment effect. \end{example_contd}
\begin{example_contd}[(ref)] As another illustration, we can consider settings in which the quality of the technical assistance is stochastic, and that generates variability of the causal effect of such treatment. In this case, we would be considering inference on the treatment effect conditional on the quality of the technical assistance. \end{example_contd}
Another alternative is to construct prediction sets, rather than confidence intervals. Prediction sets are set-valued statistics ${C}$ with the property that, with a given confidence $1- \gamma \in (0,1)$:
Prediction sets are popular in the conformal prediction literature Lei2013, and have been considered in causal settings with few treated units by cattaneo2021prediction,cattaneo2023uncertainty and chernozhukov2021exact. It has been argued that such construction is especially useful in settings with substantial treatment effect heterogeneity -- and where the goal is to personalize treatment allocation --, in which case assessing uncertainty regarding individual treatment effects may be more relevant than focusing on average treatment effects from the sampled population Candes,kivaranovic2020conformal.
A set-valued function of the data satisfying (ref) has the property of containing the sample average treatment effect on the treated with probability at least $1-\gamma$, over repeated realizations of sampling uncertainty. If one had knowledge of the quantile function $Q_{\bar{Y}(0)|D_1,\ldots,D_N}$ of the conditional distribution of $\frac{1}{N_1}\sum_{i=1}^{N_1} D_i Y_i(0)$ given $D_1,\ldots, D_N$, then it would be possible to take $\mathcal{C}$ as the interval:
$$\left[\frac{1}{N_1}\sum_{i=1}^N D_i Y_i -Q_{\bar{Y}(0)|D_1,\ldots,D_N}(1-\gamma/2), \frac{1}{N_1}\sum_{i=1}^N D_i Y_i -Q_{\bar{Y}(0)|D_1,\ldots,D_N}(\gamma/2)\right]\, .$$
Since the quantile function $Q_{\bar{Y}(0)|D_1,\ldots,D_N}$ is typically unknown, one approach to constructing prediction intervals is to focus on estimators of these quantities with good statistical properties cattaneo2021prediction. We further discuss this type of construction in Section (ref).
In this paper, we establish several connections between the four alternative inference approaches described earlier. In this section, we present these connections for the simpler case of a difference-in-means estimator. Some of the connections rely on new results that we derive in a more general setting in Section (ref). Others are not new, but we believe their interpretation has been somewhat underappreciated.
It is easy to see that the sharp null (ref) is equivalent to the conditions $\beta=c$ and $\mathbb{P}[\cap_{i: D_i=1}\{\alpha_i = \beta\}|D_1,\ldots, D_N]=1$, where the first term states that the ATT is equal to $c$, while the second term means that treatment effects are non-stochastic and homogeneous. Consequently, under the sharp null (ref), we have that the assumptions considered in the methods discussed in Section (ref) for inference on the ATT are valid.
In other words, if we have a test that is valid for inference on the ATT under the assumption that there is no treatment effect heterogeneity, then such test would also be valid for testing a sharp null. Conversely, if we have a test that is valid for inference on a sharp null, then this test would also be valid for inference on the ATT under the assumption of no treatment effect heterogeneity.
Consider a test of the null $H_0 : \beta = c$ that is valid under the assumption that there is no treatment effect heterogeneity at the $\gamma$ significance level. Let the test statistic be $\hat \phi = g\left(\left\{Y_i,D_i\right\}_{i=1}^N\right)$, where $g$ is a map that is not a function of the data.\footnote{The map $g$ may depend on other random variables that are independent from the data, such as in the case of randomized decision rules.} Assume that this statistic satisfies the following condition:
In other words, Condition (ref) restricts consideration to tests whose conclusions depend on the actual treatment effect heterogeneity only through the average of effects on the treated $\frac{1}{N_1}\sum_{i=1}^{N} D_i \alpha_i$.
We show that a test satisfying Condition (ref) that is valid under treatment effect homogeneity at the $\gamma$ significance level would also be valid for inference on the realized treatment effect under assumptions that limit dependence between treatment effects and untreated potential outcomes. Specifically, under the assumption:
the test is valid for testing the null $H_0: \frac{1}{N_1}\sum_{i=1}^{N_1}D_i \alpha_i=c$ conditionally on $\mathcal{A}$, since, on the event that the null holds,
The intuition for the above result is that, conditional on $\mathcal{A}$, we have a model with homogeneous treatment effects where the distribution of $\{Y_i(0)\}_{i=1}^n |D_1,\ldots, D_N$ satisfies the properties required for valid inference when there is no treatment effect heterogeneity.
Notice that Assumption (ref) is potentially restrictive: in particular this assumption is not necessarily valid even if we assume that $\{Y_i(0)\}_{i=1}^N$ is iid and treatment is randomly assigned. In contrast, Assumption (ref) would be satisfied under these conditions.
Assumption (ref) is satisfied if we consider a model in which (stochastic) treatment effects $\mathcal{A} = \Gamma(U)$ are a function of unobserved variables $U$, while potential outcomes when untreated $\{Y_i(0)\}_{i=1}^n = \Lambda(V)$ is a function of unobserved variable $V$. In this case, Assumption (ref) means that the sources of variability that determine $\mathcal{A}$ are independent from the sources of variability that determine $\{Y_i(0)\}_{i=1}^n = \Lambda(V)$, that is, $U \perp V | D_1,...,D_N$.
Another alternative is that there is dependence between $\mathcal{A}$ and $\{Y_i(0)\}_{i=1}^n$ (conditionally on $D_1,\ldots, D_N$), but we have that the distribution $\{Y_i(0)\}_{i=1}^n |\mathcal{A},D_1,\ldots, D_N$ satisfy the assumptions necessary so that we can learn about $Y_i(0)$ of the treated using the outcomes from the controls, e.g. if we replace Assumption (ref) with $Y_i(0) \overset{iid}{\sim} F_{Y(0)|\mathcal{A},D_1,\ldots,D_N}$.
\begin{example_contd}[(ref)] Considering again our leading example, we can think that the independence between treatment effects and potential outcomes when untreated is reasonable when the heterogeneity in treatment effects come from variability in the quality of the institution implementing the treatment. However, this assumption would be less reasonable if the heterogeneity in treatment effects comes from variables that may be related to potential outcomes when untreated, such as weather shocks. The key point in this case is whether we have different weather shocks for each farm, or we have a single weather shock that affects all farms in a similar way. In the first case, even if Assumption (ref) is valid unconditionally, it would be hard to justify that it would also be valid conditional on $\mathcal{A}$. In the second case, it should be more reasonable to assume that Assumption (ref) is reasonable conditional on $\mathcal{A}$, even if we have that the treatment effect heterogeneity is potentially related to the potential outcomes when untreated. \end{example_contd}
\begin{example_contd}[(ref)] Consider again the case treatment effect depends only observable whether shocks, which we denote by $Z_i$. In this case, it might be reasonable to assume that Assumption (ref) is valid conditional on the treatment effect, if we also condition on $Z_i$. In this case, one would conduct inference on the realized treatment effect separately for each subsample defined by the values of $Z_i$. \end{example_contd}
Tests of sharp nulls and prediction sets are intimately connected. As an example, for a given significance level $\gamma \in (0,1)$, it is possible to invert the procedure given by the decision rule $\phi_c = \mathbf{1}\{\hat{p}_c \leq \gamma\}$, where $\hat{p}_c$ denotes the p-value in (ref), to construct a set $\mathcal{I} = \{c \in \mathbb{R}: \phi_c = 0\}$. Doing so will result in a {prediction set} that satisfies (ref) when Assumption (ref) is satisfied.
More generally, in the next section, we show that, any set-valued function satisfying (ref) defines a valid test of sharp null (ref); and, conversely, for the types of test statistics typically considered in settings with few treated units, inverting a test of the sharp null (ref) will result in a prediction set satisfying (ref). While unsurprising, this result seems to be unappreciated in the literature. Specifically, it provides a rationale for reporting prediction sets of treatment effects in settings with few treated units, as doing so is equivalent to reporting those values for which a test of a sharp null is not rejected by the data. Finally, we note that the class of tests whose inversion leads to valid prediction intervals is quite large. Indeed, consider a valid family of tests of sharp nulls $\{\hat{\phi}_c\}_{c \in \mathbb{R}}$ that satisfy the following condition:
In this case, our results show that inversion of this family of tests will result in valid prediction sets. Condition (ref) requires that the conclusion of the tests $\hat{\phi}_c$ depend on the in-sample average treatment on the treated $\frac{1}{N_1}\sum_{i=1}^{N_1}D_i \alpha_i$ and the null value $c$ only through the difference $\frac{1}{N_1}\sum_{i=1}^{N_1}D_i \alpha_i-c$. Notice that, if we have a family of tests that satisfy Condition (ref), then each individual test has the structure in Condition (ref).
If we have a test that is valid for inference on the realized treatment effects, then it is clear that it would also be valid to construct prediction intervals. The intuition is that a confidence interval for the realized treatment effect would have valid coverage conditional on the treatment effect, so when we integrate over the distribution of treatment effects it would lead to a valid prediction interval.
The other direction, however, is more tricky. It might be that we have a valid prediction interval, but it does not lead to valid confidence intervals for the realized treatment effects (that is, once we condition on the treatment effects). The reason is that we may have some realizations of the treatment effects in which the prediction interval has conditional (on treatment effects) undercoverage, which is compensated by other realizations of the treatment effects in which we have over-coverage. In Appendix (ref) we provide a simple example in which this may happen (see also CWZ2023_comment for a related discussion).
Therefore, we do not have a direct equivalence between prediction intervals and inference on the realized treatment effects. Still, we show that the following result is valid: for the types of test statistics typically considered in the few treated literature,\footnote{Specifically, we consider inference methods based on statistics of the form $\frac{1}{N_1}\sum_{i=1}^{N_1} D_i (Y_i-\hat M_i)$, where $\hat{M}_i$ is a consistent estimator of a proxy $\hat M_i$ of the untreated potential outcome $Y_i(0)$. These cover several methods in the few treated literature. See Section (ref) for further details.} if we have a test that is valid for the realized treatment effect under the assumption that the treatment effects are independent of untreated potential outcomes (Assumption (ref)), then inversion of this test will also produce a valid prediction interval, even when we allow for arbitrary dependence between the stochastic treatment effects and the potential outcomes when untreated.
Put another way, our results provide conditions ensuring that, if one has a consistent estimator of the distribution of $\frac{1}{N_1}\sum_{i=1}^N D_i Y_i(0)$ (conditionally on assignments), then it is possible to construct an interval that is simultaneously an (asymptotically) valid prediction interval {with no restriction on the dependence between treatment effects and untreated potential outcomes}, as well as an (asymptotically) valid confidence interval for the realized treatment effects under the additional assumption of independence between treatment effects and untreated potential outcomes. For example, we would be able to construct a consistent estimator of the distribution of $\frac{1}{N_1}\sum_{i=1}^N D_i Y_i(0)$ (conditionally on assignments) if $Y_i(0)$ has the same marginal distribution across $i=1,...,N$ and $\{Y_i(0)\}_{i:D_=0}$ is weakly dependent alvarez2023inference,conley20211inference.
\begin{example_contd}[(ref)] In our empirical illustration, imagine that we have weather shocks that may affect agricultural outputs. If we are in a setting in which all farms are in the same geographical region, then we should expect weather shocks that affect all units {in a similar way}, so the assumption for the connection between prediction intervals and inference on realized treatment effects would not be valid, because in this case it would be difficult to find a consistent estimator for the conditional-on-assignment distribution of $\frac{1}{N_1}\sum_{i=1}^N D_i Y_i(0)$. In contrast, this assumption would be more reasonable if the farms are geographically more distant, so they may experience different weather shocks. \end{example_contd}
Since we show that sharp nulls and prediction intervals are intimately connected, this means that we also have a connection between tests of sharp nulls and inference on the realized treatment effects. {Indeed, note that inference on the realized treatment effects under Assumptions (ref) and (ref) will result in exactly the same p-value $\hat{p}_c$ given in (ref) as the test of sharp nulls under Assumption (ref). Therefore, the decision rule to reject the null if $\hat{p}_c$ is below some significance value produces a test of the sharp null (ref) under Assumption (ref), and a test on realized average affects (ref) under Assumptions (ref) and (ref). As we show in the next section, this is a more general feature of tests designed for settings with few treated units: under some assumptions, a test that is asymptotically valid for testing the sharp null (ref), is also valid for testing (ref) under the Assumption (ref) that restricts the relation between treatment effects and untreated potential outcomes.}
Fix a probability space $(\Omega, \Sigma, P)$. We consider a setting with $N_1$ units of interest, indexed by $i=1,\ldots N_1$. For each of these units, we define potential outcomes $Y_i(0)$ and $Y_i(1)$ corresponding to the thought experiment of manipulating a binary treatment over an outcome $Y$ of interest. In line with the model-based setting of earlier sections, we view $(Y_i(0),Y_i(1))$ as random variables defined on $(\Omega, \Sigma, P)$. Uncertainty on potential outcomes may reflect a sampling process from a population or, in those settings where such interpretation is not warranted, economic uncertainty that determines potential outcomes.
{We remain agnostic about the interpretation of the unit index $i$: it may reflect different individuals (or counties, states etc.), or the same individual (or counties, states etc.) at different points in time, or reflecting both different individuals over different time periods}. This allows us to establish connection between inference approaches in both settings where the target are averages of effects across individuals and/or across the time. The causal effects of the policy on each unit are denoted by the random variables $\alpha_i=Y_i(1)-Y_i(0)$, $i=1$.
We assume that the econometrician observes a sample $\{Y_i,X_i\}_{i=1}^{N_1} \cup \{\xi\}$ of random variables defined on $(\Omega, \Sigma, P)$, where $Y_i = Y_i(1)$, $i=1,\ldots N_1$. In other words, the econometrician observes the outcomes of the $N_1$ units under the treatment regime. The random variables $\{X_i\}_{i=1}^{N_1}$ denote a set of unit-level covariates used in the analysis, while $\xi$ denotes auxiliary data used in the imputation of the missing potential outcomes $Y_i(0)$ that is always required, either implicitly or explicitly, by any causal inference approach athey2021matrix. The auxiliary data $\xi$ can include past values of the outcome $Y$ for the $N_1$ units -- when the assumptions underlying the causal inference approach enable extrapolation of the missing potential outcome from the past --, the values of $Y$ for a group of controls -- as in methods exploiting cross-sectional variation in assignment -- or a combination of both -- as in panel data methods such as differences-in-differences or the synthetic control method.
In what follows, we denote the push-through probability measure of $(Y_i(0),Y_i(1),X_i)$ by $\mathbb{P}_i$. Expectations with respect to a probability law $Q$ is denoted by $E_Q$.
We now formalize, in the setup described in Section (ref), the alternative inferential targets introduced in Section (ref). We begin with a definition of a test of a sharp null. In this section, we focus on definitions that assume finite-sample validity of the different inferential approaches, though it is immediate to extend these to asymptotic sequences where validity is only approximate. Asymptotic definitions are discussed in Section (ref).
In our setting, a sharp null is a hypothesis on both the level of the treatment effects on the treated units, as well as on the absence of treatment effect heterogeneity across these units. Notice that, when uncertainty stems from a sampling process from a well defined population, the sharp null implies that treatment effects are constant and homogeneous in the (treated) population from which the sample is drawn.
A prediction set is a function of the sample that, over repeated samples, contains the (possibly stochastic) sample average treatment effect in at least $100(1-\gamma) \%$ of the cases.
The third type of inference concerns inference on the realized sample average effect on the treated.
Clearly, by iterated expectations, any valid confidence set for the realized effect is a valid prediction set for the sample average effect. We note, however, that, given its conditonal nature, construction of confidence sets for realized effects will typically require assumptions on the relation between treatment effects $\alpha$ and untreated outcomes $Y(0)$ in the treated population that may be hard to justify in practice. In the next sections, we show that inference procedures devised for inference on the realized effect may be reinterpreted as confidence sets for the sample average effect under a different set of assumptions that may be more palatable to applied researchers.
In this section, we lay out results that connect tests of sharp nulls and prediction sets. These results hold more generally, and are not restricted to few-treated-asymptotics. We discuss connections under a few-treated asymptotic sequence in the next section.
The first result shows that prediction intervals always yield tests of sharp nulls.
The previous result is straightforward, though its interpretation may be novel. The result shows that the procedure of checking whether a value is within a prediction set for the sample average treatment effect may be seen as a test of a sharp null.
In what follows, we provide a partial converse to the above result. For that, we consider an analog of Condition (ref) for this more general setting.
For this class of decision rules, we show that the inversion of a testing procedure of sharp nulls generates a valid prediction interval.
The previous lemma shows that, for a rather general class of tests of sharp nulls, the inversion of these tests produces a valid prediction region for the sample average treatment effect on the treated $\frac{1}{N_1}\sum_{i=1}^{N_1} \alpha_i$. The assumption, spelled as Condition (ref), that the decision rule depends on $\frac{1}{N_1}\sum_{i=1}^{N_1} \alpha_i$ and the hypothesized value $c$ only through the difference $\frac{1}{N_1}\sum_{i=1}^{N_1} \alpha_i -c$ cannot be generally dispensed with. Indeed, in Appendix (ref), we provide a simple example of a valid family of tests of sharp nulls that does not possess this structure and whose inversion does not produce a valid prediction interval for the average effect. Following the same discussion as in Remark (ref), we also note that Condition (ref) would not be satisfied if we consider an unequal variance t-test.
In the next section, we show that a broad class of tests employed in the few treated literature enjoys the structure in the statement of Lemma (ref). As a consequence, these tests can be used both as (asymptotically) valid tests of sharp nulls, and may as well be inverted to construct (asymptotically) valid prediction sets for the sample average effect on the treated.
In this section, we specialize the structure of our problem to the one typically adopted in few treated settings. Specifically, we follow alvarez2025inferencetreatedunits and decompose the unobserved potential outcome of treated observations as:
$$Y_{i}(0) = M_i + \epsilon_i, \quad i=1,\ldots N_1 \, ,$$ where $M_i$ is a proxy for the untreated potential outcome for the treated observations. Let $\hat{M}_i$ be an estimator for this proxy. The average effect may then be estimated as:
$$\widehat{\boldsymbol{\alpha}} = \frac{1}{N_1} \sum_{i=1}^{N_1} (Y_i - \hat{M}_i)\, .$$
In what follows, we shall assume that the estimators $\hat{M}_i$ adopted by the researcher are consistent in an asymptotic sequence where the number of treated units is fixed. Importantly, we shall remain agnostic on whether the estimator relies on a large time series or on a large number of controls (or both). This allows us to cover a wide range of methods.\footnote{See chernozhukov2021exact and alvarez2025inferencetreatedunits for examples of proxies that rely on time series or cross-sectional variation or both.}
To formally state asymptotic inference results, we embed our setting in an asymptotic framework indexed by $s \in \mathbb{N}$, and consider the behavior of the inference procedures as $s \to \infty$. We allow both the sample law, now indexed by $s$ and denoted by $P_s$, as well as the dimension of the vector of auxiliary random variables, denoted by $\xi_s$, to vary with $s$. Doing so allows us to capture settings where consistent estimation of the proxies relies on a large number of pre-treatment variables or on a large number of controls (or both).\footnote{Moreover, by allowing the sample law to vary with $s$, it is possible, if one considers arbitrary sequences of laws $P_s \in \mathcal{P}_s$ in a sequence of spaces $\mathcal{P}_s$, $s \in \mathbb{N}$, to obtain uniform-in-law asymptotic coverage results Canay2017,Andrews2020.} In keeping with the few treated literature, we consider an asymptotic sequence where $N_1$ is fixed.
The assumption of consistent estimation of the proxies can be stated in our asymptotic framework as follows:
Notice that, under an asymptotic sequence that satisfies Assumption (ref), we have that:
$$\widehat{\boldsymbol{\alpha}} - \frac{1}{N_1}\sum_{i=1}^{N_1}\alpha_i = \frac{1}{N_1}\sum_{i=1}^{N_1}\epsilon_i + o_{P_s}{(1)} \, .$$
Consequently, for asymptotically valid tests of sharp nulls, it suffices to approximate the distribution of the $\epsilon_i$.\footnote{Following cattaneo2021prediction, we can also consider improvements to asymptotically valid inference methods that explicitly take into account the $o_{P_s}(1)$ estimation error of the proxy. Given that these constructions offer asymptotically vanishing improvements, they do not alter the main connections established in this section.} For example, let $u \mapsto \hat{Q}_s(u)$ be a pointwise consistent estimator of the quantile function $Q_s$ of the distribution of $\frac{1}{N_1}\sum_{i=1}^{N_1}\epsilon_i$, meaning that, for each $u \in (0,1)$: $$|\hat{Q}_s(u)-Q_s(u)|\overset{P_s}{\to}0\, .$$
For testing the sharp null $P_s[\cap_{i=1}^{N_1}\{\alpha_i = c\}] = 1$ at significance level $\gamma$, the researcher could consider the test function:
This decision rule produces an asymptotic size $\gamma$-test of the sharp null if the distribution of $\frac{1}{N_1}\sum_{i=1}^{N_1}\epsilon_i$ is continuous, by which we mean that:
$$\lim_{s \to \infty} P_s[\phi_c] = \gamma\, ,$$ for a sequence of laws $(P_s)_s$ where the null holds (see the proof of the lemma below).
The next result shows that, for tests with the structure (ref), test inversion produces an asymptotically valid prediction set under the same set of assumptions that justify the test.
The previous result, when combined with Lemma (ref), shows that, for the usual approaches undertaken in the few treated literature, tests of sharp nulls and prediction sets are intricately connected, in the sense that prediction sets always produce (asymptotic) tests of sharp nulls, and tests of sharp nulls with the structure (ref) can always be inverted to construct an asymptotically valid prediction set.
Finally, we consider the construction of asymptotically valid confidence intervals for the realised effect, by which we mean set-valued functions of the data $\mathcal{C}$ with the property that:
A natural question to be asked is whether, in our triangular array setup where the distribution of the data varies with $s$, Assumption (ref) is equivalent to consistency, conditional on $(\alpha_1,\ldots \alpha_{N_1})$, of the proxies “in probability” -- the notion of consistency that one requires to construct sets satisfying (ref). The next lemma shows that this is in fact true.
At its essence, the previous result reveals that the way we treat treatment effect heterogeneity -- either as stochastic or as “realized and conditioned-on” -- is “irrelevant” for assessing the correct notion of consistency of the proxies required by different inference methods. Indeed, our result shows that verifying if the proxies are consistent -- which is required for asymptotic validity of prediction intervals -- or conditionally consistent in probability -- which is required for validity of methods for inference on the realized effect -- always produces the same conclusion, given that both notions are equivalent.
Given the result in Lemma (ref), it follows that, in constructing confidence sets for the realized effect, one could adopt a similar formulation to (ref), but replacing the estimator of the quantiles of the distribution of $\frac{1}{N_1}\sum_{i=1}^{N_1}\epsilon_i$ with a consistent estimator of the quantiles of the conditional on $(\alpha_i, \ldots, \alpha_{N_1})$ distribution of $\frac{1}{N_1}\sum_{i=1}^{N_1}\epsilon_i$. However, as argued in earlier sections, one difficulty of this approach is that consistency of the estimator of conditional quantiles requires assumptions on the relation between $\alpha_i$ and $\epsilon_i$ that may be hard to justify in practice.
In light of this difficulty, our final result provides a novel connection between confidence sets for the realized treatment effects and prediction sets. Specifically, we show that methods for conducting inference on treatment effects that rely on a consistent estimator of the distribution $e \mapsto P_s\left[\frac{1}{N_1}\sum_{i=1}^{N_1}\epsilon_i \leq e \right]$ can be either seen as valid prediction intervals, or as a valid method for inference on the realized effect, under the additional assumption that treatment effects are independent of the $\epsilon_i$.