Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
66,607 characters · 23 sections · 35 citation commands
Individual Treatment Effect: Prediction Intervals and Sharp Bounds
\doublespacing
The traditional causal inference literature has been focused on population level treatment effect parameters such as the average treatment effect (ATE) and conditional average treatment effect (CATE). Although the individual treatment effect (ITE) is often regarded as the ideal parameter of interest for personalized decision making, it is not, in general identifiable even if we know the outcome for an individual and have data from a large randomized experiment. Recent works by lei2021conformal, jin2023sensitivity, chernozhukov2023toward,wang2025conformal discuss conformal inference methods to estimate ITE and prediction intervals for the ITE. Another recent debate from mueller2022personalized and dawid2023personalised also explores the possibility of using ITE and bounds on ITE to help personalized decision making.
Bounds and relationships on the probability of the counterfactual treatment effect has been studied in terms of probability of causation in robins1989probability and tian2000probabilities. Some more recent work from mueller2021causes,sani2023bounding,kawakami2025mediation extends these ideas to learn individual responses and bounds from causal diagrams and mediation analysis. Inference and examples of probability of causation have been discussed in dawid2016statistical. fan2010sharp study sharp bounds on the cdf of the individual treatment effect using copulas based on the previous result from frank1987best and williamson1990probabilistic. mullahy2018individual applies the bounds in fan2010sharp in health economics applications. Bounds and optimal policy for binary treatment and outcome have been studied in kallus2022treatment, kallus2022s, and dawid2023personalised. Individual treatment effects on ordinal outcomes have been studied in lu2018treatment. Numerical approaches for solving causal inference problems in discrete settings have been explored in duarte2024automated. Exact inference on individual treatment effects based on permutation has been studied in blaker2000confidence,rigdon2015randomization,chiba2015exact. The construction of confidence intervals for causal parameters has been discussed in robins1988confidence, imbens2018causal and brennan2024causal.
In this work, we try to understand the prediction intervals and bounds for the ITE given that we have observations from well-conducted large randomized control trials (RCT). In Section (ref), we start from the binary treatment and outcome model and provide a complete characterization of ITE prediction intervals under this simple model setting. We characterize when a degenerate interval consisting of a single value is a valid prediction interval, when is the valid prediction interval for ITE non-negative/non-positive, and when the only valid prediction interval is a trivial interval. Additionally, we give conditions on the observed data for there to exist a consistent joint distribution under which a given non-trivial prediction interval would be valid. In Section (ref), we extend our insights to continuous and ordinal outcomes. Our approach leverages cdf bounds for the difference of two random variables from fan2010sharp and zhang2024bounds. In Section (ref), we provide sharp bounds on the cumulative distribution function (cdf) and probability mass function (pmf) of the ITE under binary treatment and outcome model. We explore the relationships between the cdf/pmf bounds on ITE and the Fr\'{e}chet-Hoeffding bounds on two random variables. This motivates the general sharp bounds on the pmf of the ITE for ordinal outcome in section (ref). Lastly, we compare ITE prediction intervals and ATE confidence intervals and provide a synthetic data example to discuss the implications of them in section (ref).
Throughout the paper, we consider a binary treatment setting with treatment $D=0,1$. Let $Y_1$ be the potential outcome when receiving the treatment and $Y_0$ be the potential outcome when not receiving the treatment. We define the individual treatment effect as:
Let $Y$ be the observed variable. By consistency, $Y=Y_0$ when $D=0$ and $Y=Y_1$ when $D=1$. In this section, we will examine what is possible to learn in the limit of a large sample size.
A $(1-\alpha)$ {\em prediction interval} for an individual treatment effect is an interval such that \[ P\left((Y_1-Y_0) \in [L,R]\right) \geq 1-\alpha. \] Throughout the paper, we assume $\alpha$ is sufficiently bounded away from $0.5$.
To further narrow the discussion, we first consider the case in which the treatment and response are both binary. Following copas1973randomization who characterizes individual patients in the binary treatment and outcome model, we have the following four types and individual treatment effects:
Notice that in this setting the individual treatment effects take three possible values: $-1,0,1$.
In the simple setting that we consider, since there are only $3$ possible values taken by the ITE, there are $6$ possible prediction intervals:
Note that prediction intervals $[-1,0],[0,1],[-1, 1]$ here are sets correspond to $\{-1,0\}, \{0,1\}, \{-1,0,1\}$. Singleton sets can be viewed as degenerate intervals where the starting point equals to the end point. In this simple setting, there is only one prediction set that does not correspond to an interval, namely $\{-1,1\}$. We discussed it in Remark (ref).
Under the conditions stated in Corollary (ref), the set $\{0\}$ fails to provide sufficient coverage; yet each of $[-1,0]$ and $[0,1]$ captures a sufficiently large portion of the probability mass (at least $1-\alpha$). Both of these intervals therefore qualify as valid $(1-\alpha)\%$ prediction intervals for $\mathrm{ITE}$, and each achieves the same minimal length (one unit). Consequently, there is no single “shortest” interval that strictly dominates the other. In general, even given the joint distribution of $Y_0, Y_1$, there could be multiple minimal-length valid prediction intervals.
Under randomization, the relationship between the counterfactual distribution $P(Y_0, Y_1)$ and the observed distributions $\left\{P(Y\mid D=0), P(Y\mid D=1)\right\}$ is given by this table:
Here $P(Y\!=\!i\mid D\!=\!j) = P( Y_j\!=\!i)$ due to randomization.
Equivalently we may write this in terms of types:
Since under randomization the joint distribution must add up to satisfy the observed distributions in the two treatment arms, we can parametrize the distribution of types using $P(AR)=t$. We then have the following solution set:
where, by Proposition (ref), we have
There are circumstances in which the only valid prediction interval for the individual treatment effect is the trivial interval: $[-1,1]$!
The interval will be trivial when the set of people of type Hurt can be larger than $\alpha$, {\em and} the set of people of type Helped can also be larger than $\alpha$.
Notice that the proportion Helped and Hurt, {\em both } achieve their maximum value when the proportion Always Recover ${\textcolor{red}{ t}}$ achieves its minimum value. Hence the only valid prediction interval will be trivial whenever we have both:
this may be equivalently expressed as:
Consequently, provided we have:
and
then the only valid prediction interval will be trivial. In concrete terms, if $\alpha=0.05$ and the proportions recovering under treatment ($D=1$) and control ($D=0$) both lie between $5\%$ and $95\%$ then, given that we know the conditional distributions $P(Y|D)$, the only valid 95% prediction interval for the individual treatment effect will be $[-1,1]$.
These results show that for randomized experiments in which the proportion of recovery in both arms lies between $\alpha$ and $(1-\alpha)$, the observed data is entirely uninformative regarding the ITE.
Notwithstanding the results in the previous section, perhaps surprisingly, there are situations in which a valid $(1-\alpha)\%$ prediction interval is a singleton. We now characterize when this occurs. Detailed derivations are provided in the Appendix (ref). There are three cases to consider:
Following the previous result, one can rule out negative (or positive) treatment effects if the proportion of HU (or HE) can never exceed $\alpha$. This occurs when either $P(Y=1 \mid D=0)$ or $1 - P(Y=1 \mid D=1)$ is below $\alpha$ (and similarly for ruling out positive effects). See Appendix (ref).
Based on the result in (ref), (ref), and (ref), under randomization and in the limit of a large sample size where we observe the true marginals $P(Y=1\mid D=0), P(Y=1\mid D=1)$ and their counterparts, we can characterize the corresponding ITE prediction intervals. These intervals are valid no matter the joint distribution over $Y_0$ and $Y_1$ (provided it is compatible with $P(Y\mid D)$. These intervals are also “the best we can do" without additional assumptions or information, e.g. from a cross-over study.
Notice that we have two overlapping triangles each with area $\frac{1}{2}\alpha^2$ in Figure (ref) where both $[0,1]$ and $[-1,0]$ can serve as valid $\alpha-$level prediction interval for individual treatment effects. Assume that for the same length of intervals, intervals with higher coverage are "better". Then we can further decompose these triangular areas to obtain the "best" prediction intervals; see Figure (ref).
Suppose we are given a prediction interval and an observed distribution from an RCT, we can consider when does there exist a joint distribution over potential outcomes that is compatible with the observed distributions and for the prediction interval to be valid and of minimal length. In Appendix (ref), we derive the precise constraints on the marginal distributions \(\,P(Y=1 \mid D=0)\) and \(P(Y=1 \mid D=1)\) that guarantee each of the intervals can serve as a valid (or “best”) prediction interval for the individual treatment effect for some joint distribution over potential outcome. The necessary conditions take the form of simple inequalities relating the two response probabilities; they characterize scenarios in which each the true treatment effect with probability at least \(1-\alpha\) lies within the given set for som population compatible with $P(Y\mid D)$. Full details, derivations, and proofs appear in Appendix (ref).
Figure (ref) and (ref) provide visualizations of the necessary conditions on $P(Y=1\mid D=0)$ and $P(Y=1\mid D=1)$ for a given prediction interval to be valid/best. Note that if for a given point in the unit square, we consider the intervals (from Figure (ref) and (ref)) containing $(P(Y=1\mid D=1), P(Y=1\mid D=0))$ and then select the longest, we recover Figure (ref).
We now turn to the case where the outcome is continuous, rather than binary. Our goal is to address questions analogous to those explored in the binary setting. In particular, we begin by investigating how to construct valid prediction intervals based solely on the marginal distributions, without imposing assumptions on the joint distribution of the potential outcomes. We then examine the conditions under which these intervals can be bounded away from zero. Further, we solve the problem: given a prediction interval (or set), what conditions on the observed data ensure the existence of a consistent joint distribution under which the interval is valid?
The marginal distributions of $Y_1$ and $Y_0$ are identified from randomized experiments. Let $[L_0, R_0]$ be an interval such that $P\left(Y_0\in [L_0,R_0]\right) \geq 1-\alpha/2.$ Let $[L_1, R_1]$ be an interval such that $P\left(Y_1\in [L_1,R_1]\right) \geq 1-\alpha/2.$ Then
This result is similar to the “naive" conformal inference prediction interval for the ITE given in lei2021conformal Section 4.1. Though we further assume an infinite sample size and do not use covariates.
One might argue that the bounds in ((ref)) are too conservative. We will take a different perspective to see what are the points that must be included in the interval. Note that to be valid, an ITE prediction interval must be valid for all joint distributions consistent with the observed data, and hence will in general be wider than that resulting from knowledge of this joint distribution. Let $[L_0', R_0']$ be an interval such that \[L_0' := \min \Bigl\{\ell \in \mathbb{R} : P\bigl(Y_0 < \ell\bigr) > \alpha\Bigr\}\] and \[R_0' := \max \Bigl\{\ell \in \mathbb{R} : P\bigl(Y_0 > \ell\bigr) > \alpha\Bigr\}.\] In other words, $L'_0$ and $R'_0$ are the $\alpha$-quantile and $(1-\alpha)$-quantile of $Y_i(0)$. Similarly we can define \[L_1' := \min \Bigl\{\ell \in \mathbb{R} : P\bigl(Y_1 < \ell\bigr) > \alpha\Bigr\}\] and \[R_1' := \max \Bigl\{\ell \in \mathbb{R} : P\bigl(Y_1 > \ell\bigr) > \alpha\Bigr\}.\] Then a valid $(1-\alpha)$ prediction interval for the ITE must include these points:
This follows because otherwise there exists a joint distribution of $Y_1, Y_0$ such that there is more than $\alpha$ mass outside of the prediction interval. This is because the only constraint imposed on the joint distribution by the marginals is given by the Fréchet inequalities, \[P(Y_1> R_1', Y_0<L_0')\leq \min\{P(Y_1> R_1'), P(Y_0<L'_0)\}\] By construction the minimum is greater than $\alpha$, meaning there exists a joint distribution that $P(Y_1> R_1', Y_0<L_0')>\alpha$. Hence under this joint distribution, $P(Y_1-Y_0>R_1'-L_0')>\alpha$, therefore, any interval that does not include $R_1'-L_0'$ will not be a valid $\alpha$ level prediction interval.
Even small tails in marginal distributions of $Y_1, Y_0$ can force the prediction interval to expand considerably at both extremes. We illustrate this in Figure (ref).
In the previous section, we gave conditions under which a point is included in every valid ITE prediction interval. Here we now give necessary and sufficient conditions for a valid interval to exclude the half line $(0,\infty)$ (or $(-\infty,0)$). Though the arguments extend to any constant, not just zero, we now address the question: can a valid prediction interval for the individual treatment effect (ITE) be bounded away from zero? Suppose there exists a joint distribution of the potential outcomes such that $P(Y_1 - Y_0 \leq 0) > \alpha$. Then, without further assumptions on the joint distribution, the left endpoint of any valid $(1-\alpha)$ prediction interval must be less than or equal to zero. This is closely related to Kolmogorov's problem and the results in fan2010sharp, zhang2024bounds, which we review in Appendix (ref). In particular, consider the following upper bound on $P(Y_1 - Y_0 \leq 0)$:
If this upper bound exceeds $\alpha$, then zero must lie within the prediction interval, i.e., the left endpoint cannot be strictly greater than zero.
Similarly, if $P(Y_1 - Y_0 \geq 0) > \alpha$, then the right endpoint of any valid prediction interval must be at least zero. This probability can be written as \[ P(Y_1 - Y_0 \geq 0) = 1 - P(Y_1 - Y_0 < 0), \] so we may use a lower bound on $P(Y_1 - Y_0 < 0)$ to assess whether the right endpoint must include zero. The lower bound is given by:
If this lower bound is less than $1 - \alpha$, then the right endpoint must be at least zero.
Therefore, if both the left endpoint must be less than or equal to zero and the right endpoint must be greater than or equal to zero, the prediction interval necessarily includes zero and cannot be bounded away from it.
In fact, if the distribution of $Y_1-Y_0$ is absolutely continuous with respect to Lebesgue measure so $P(Y_1-Y_0=0)=0$, then we can combine the conditions in ((ref)) and ((ref)) to state that for a valid $(1-\alpha)$ prediction interval for the ITE to exclude zero, either the lower bound on $P(Y_1 - Y_0 \leq 0)$ to exceed $1 - \alpha$ or the upper bound to be less than $\alpha$—conditions that are only met in highly extreme cases.
In Appendix (ref), we briefly discuss the cases of bounding ITE prediction intervals in the ordinal outcome setting, using a new result that we will prove in Section (ref).
In the previous section, we used bounds on the cumulative distribution function (cdf) of the individual treatment effect (ITE) under known marginal distributions for potential outcomes. We now turn to the corresponding problem for the probability mass function (pmf). As before, we begin with the binary treatment and outcome setting. We first consider the binary treatment and binary outcome model where we know the marginals for $Y_0$ is $p, 1-p$ and the marginals for $Y_1$ is $q, 1-q$. We can parameterize the joint density using $P(Y_0=1, Y_1=0)=t$ and Table $\ref{table_binary}$.
Note: Here we make the distinction that Fr\'{e}chet-Hoeffding bounds are bounding the joint cdf of $Y_1, Y_0$ while the Fr\'{e}chet inequality bounds are bounding the joint pmf of $Y_1, Y_0$.
In the binary treatment and outcome model, a perfect positive dependence of $Y_1, Y_0$ is achieved by $t=\min\{p,q\}$ and a perfect negative dependence of $Y_1,Y_0$ is achieved by $t=\max\{0,p+q-1\}$.
From Table (ref), we have $P(\text{ITE}=-1)=q-t$, $P(\text{ITE}=0)=1-p-q+2t$, $P(\text{ITE}=1)=p-t$. The upper and lower bounds for each ITE value are reached at the Fr\'{e}chet-Hoeffding bound on the joint distribution of $Y_1, Y_0$ (i.e. when $t$ takes minimum or maximum value).
We further observe that for $P(\text{ITE}=1)$, $P(\text{ITE}=-1)$, we can obtain pmf bounds by applying Fr\'{e}chet inequality bounds on each cell directly. And for $P(\text{ITE}=0)$, we can first obtain Fr\'{e}chet inequality bounds on the two cells $t\in[\max\{0,p+q-1\},\min\{p, q\}]$ and $1-q-p+t\in[\max\{0,1-p-q\},\min\{1-p, 1-q\}]$ and then add up the lower and upper bounds to obtain $1-q-p+2t\in [\max\{0,p+q-1\}+\max\{0,1-p-q\},\min\{p, q\}+\min\{1-p, 1-q\}]$. This simplifies to the bounds given by $1-p-q+2t, t\in [\max\{0,p+q-1\},\min\{p, q\}]$ as before. We will generalize this observation in section (ref).
Fr\'{e}chet-Hoeffding bound on the joint distribution of $Y_1, Y_0$ implies that $t\in[\max\{0,p+q-1\},\min\{p, q\}]$. Since we can parameterize the joint probability with one parameter $t$, we have $P(\text{ITE}=-1)=q-t$, $P(\text{ITE}=0)=1-p-q+2t$. In the binary case, the cdf for the ITE is characterized by the values $F(-1)=P(\text{ITE}=-1)=q-t$, $F(0)=P(\text{ITE}\leq 0)=P(\text{ITE}=-1)+P(\text{ITE}=0)=1-p+t$, $F(1)=1$.
As noted in Section $2.1$ of fan2010sharp, sharp bounds on the cdf of the individual treatment effect are not achieved at the Fr\'{e}chet-Hoeffding lower and upper bounds (perfectly positive/ perfectly negative dependence) for the distribution of $Y_1, Y_0$. As a special case, the sharp bounds on cdf of ITE are achieved at the Fr\'{e}chet-Hoeffding bounds in binary treatment and outcome model.
Using the cdf bounds for $Y_1-Y_0$ in Proposition (ref), we have $$P(\mathrm{ITE}=1)=F(1)-F(0),$$ with $F(1)=1, F(0)=1-p+t\in[1-p+\max\{0,p+q-1\},1-p+\min\{p,q\}]$. Threrefore, $$P(\mathrm{ITE}=1)\in [p-\min\{p,q\}, p-\max\{0,p+q-1\}].$$ Similarly, $P(\mathrm{ITE}=0)=F(0)-F(-1)$ for $F(-1)\in [q-\min\{p,q\}, q-\max\{0,p+q-1\}]$. Thus, $$P(\mathrm{ITE}=0)\in[1-p-q+2\max\{0,p+q-1\},1-p-q+2\min\{p,q\}].$$ And $$P(\mathrm{ITE}=-1)=F(-1)=q-t, t\in[\max\{0,p+q-1\},\min\{p, q\}].$$ These bounds agree with the sharp bounds derived in Proposition (ref). However, in general, the pmf bounds derived from cdf bounds using $(\ref{eq:CDF_to_PMF})$ are not sharp. Therefore, we will explore sharp pmf bounds in Section (ref).
A natural question to ask here is: given the marginals of $Y_1, Y_0$, can we derive sharp bounds on the probability mass function (pmf) of ITE, for example, what are the bounds for $P(\text{ITE}=0)?$
This is given by adding up the Fr\'{e}chet inequality bounds for each $P(Y_1=i, Y_0=i-\delta)$.
First, we show that the bound in ((ref)) is a valid bound. Notice
And for all $i$, Fr\'{e}chet inequality bounds states that
where
So the sum must be within the bound in $(\ref{PMF_bound})$.\\
Now, we will show that the bounds in $(\ref{PMF_bound})$ are tight in a sense that there exist joint distributions compatible with the marginal conditions that achieve the lower/upper bounds. First, we will state a very useful result in koperberg2024couplings.
Let $A$ and $B$ be sets and $R \subseteq A \times B$ a relation. Then for each $\mathrm{U} \subseteq A$ the set of neighbours of $\mathrm{U}$ in $R$ is denoted by $$ \mathcal{N}_{\mathrm{R}}(\mathrm{U})=\{\mathrm{b} \in \mathrm{B}:(\mathrm{U} \times\{\mathrm{b}\}) \cap \mathrm{R} \neq \varnothing\}. $$
Now we can begin the proof with a proposition.
We will now show that there exist a joint distribution of $Y_1, Y_0$ such that $P(Y_1=i, Y_0=i-\delta)=\max\{P(Y_1=i)+P( Y_0=i-\delta)-1,0\}=L_i$ for all $i$. By Proposition (ref), $P(Y_1=i, Y_0=i-\delta)\neq 0$ for at most one $i$. There are thus two cases to consider here: when $L_i=0$ for all $i$ and when $L_i=0$ for all $i$ except one.\\
Case $1$: when $P(Y_1=i, Y_0=i-\delta)=0$ for all $i$. Case $1$ implies that
We apply Theorem (ref). Let $A,B$ be the support of the specified margins for $Y_1, Y_0$ respectively. Naturally, we have $P(Y_1), P(Y_0)$ as two probability measures on $A,B$. Let $R=\{(i,j):i\in A, j\in B, i-j\neq \delta\}$. Strassen's theorem states that there exists a coupling $\hat{P}$ of $P(Y_1), P(Y_0)$ with $\hat{P}(R)=1$ if and only if
where $$ \mathcal{N}_{R}(U)=\{b \in B:(U \times\{b\}) \cap R \neq \varnothing\}. $$ By construction of $R$, for any $i\in A$,
Therefore for any $U\subseteq A$ with more than one element, we have $\mathcal{N}_{R}(U)=B$ and:
Thus the only non trivial constraints are for singleton sets $U$ given by:
The constraints in $(\ref{lower_bound_zero_constraint_strassen})$ are already satisfied by the assumption in $(\ref{lower_bound_zero_constraint})$. Therefore, by Strassen's theorem, there must exist a coupling $P^\prime$ such that $P^\prime(R)=1$. Thus, the lower bounds $L_i=0$ for all $i$ are achievable in case $1$.\\
Case $2$: Consider the case that $L_i=0$ for all $i$ except for $i=j$. We will explicitly construct a joint distribution for $Y_1, Y_0$ that satisfies:
Consider a distribution satisfies by ((ref)) and ((ref)). Further, let $P(Y_1=i, Y_0=k)=0$ for any $i\neq j$ or $k\neq j-\delta$. Let $P(Y_1=j, Y_0=k)=P(Y_0=k), k\neq j-\delta$ and $P(Y_1=i, Y_0=j-\delta)=P(Y_1=i), i\neq j$. For all $i\neq j$, we have
For all $k\neq j-\delta$, we have
Also, we have
where the second equality follows definition of $j$. Similarly,
All the $P(Y_1=i, Y_0=j)$ are non-negative. This is a valid joint distribution that satisfies the given marginals. Thus, the lower bounds are achievable. An example of the construction of the joint probability matrix in case 2 is given in Figure (ref). \\
Next, we show that there exists a joint distribution of $Y_1,Y_0$ such that $P(Y_1=i, Y_0=i-\delta)=\min\{P(Y_1=i), P( Y_0=i-\delta)\}=U_i$ for all $i$. We will construct a joint distribution of $Y_1, Y_0$ using the joint probability table explicitly. We first permute the joint probability table based on whether the upper bound is achieved at $P(Y_1=i)$ or $P(Y_0=i-\delta)$. Let $J_1$ be the set of $i$ such that $P(Y_1=i, Y_0=i-\delta)=P(Y_1=i)$ for $i\in J_1$ and $J_2$ be the set of $j$ such that $P(Y_1=j+\delta, Y_0=j)=P(Y_0=j)$ for $j\in J_2$. We can rearrange the joint probability table based on the partition of $J_1, J_2$. Let $|J_1|=n_1, |J_2|=n_2$. Let $N_1$ be the number of elements in the support of $Y_1$. Let $N_2$ be the number of elements in the support of $Y_0$. We permute the joint probability matrix of $Y_1, Y_0$ such that the first $n_1$ rows denotes the probability of elements in $J_1$ and the last $n_2$ rows denotes the probability of elements in $J_2$. Figure $\ref{Probability Table Permutation}$ shows an example of such permutation.
By construction, the first $n_1$ rows and the last $n_2$ columns are filled with $0$'s and $U_i$'s. We have an empty $(N_1-n_1)\times(N_2\times n_2)$ sub-matrix to fill. We also obtain a new set of margin constraints on the sub-matrix by subtracting the marginals with what we have already filled. Notice that the row/column sum of these new constraints for the $(N_1-n_1)\times(N_2\times n_2)$ matrix is given by
We can fill each entry of the sub-matrix by
Thus, we construct a valid joint distribution of $Y_1, Y_0$ with $P(Y_1=i, Y_0=i-\delta)=\min\{P(Y_1=i), P( Y_0=i-\delta)\}=U_i$ for all $i$ and satisfies the given marginals. This concludes our proof of theorem (ref). We note that the same proof technique can be used to obtain the pmf bounds on the sum of two discrete random variables given the marginals.
In this section, we discuss the relationship between the average treatment effect (ATE) and the individual treatment effect (ITE). We focus on the implications for prediction intervals and hypothesis testing. The ATE is a well-defined parameter and is identifiable under standard assumptions. Consequently, as the sample size increases, confidence intervals for the ATE will eventually shrink to a point. In contrast, the individual treatment effect (ITE) is not a parameter in the classical sense: it is a random quantity defined at the unit level. Although prediction intervals for the ITE may become narrower with larger sample sizes, they do not, in general, converge to zero width. This reflects the inherent uncertainty in predicting individual level responses, even when the population-level effect is precisely estimated.
The Neyman null hypothesis posits a zero effect on average, while the Fisher null hypothesis posits no effect for every individual. By construction, a true Fisher null implies the Neyman null; if all individual effects are zero, their average must be zero. A true Fisher null also implies that $\{0\}$ will be a valid prediction interval for any level $\alpha$. A true Neyman null implies that confidence interval for ATE will contain zero with probability $(1-\alpha)\%$.
When performing hypothesis testing, $\{0\}$ being a valid prediction interval can be considered as evidence in support of the Fisher null while confidence interval does not contain $0$ can be considered as evidence against Neyman null. However, in finite samples, a failure to reject the Fisher null does not imply the ATE is in fact zero, and rejecting the Neyman null (i.e. finding a nonzero estimated ATE) does not automatically imply that $\{0\}$ cannot be a valid prediction interval.
It is possible to construct a valid prediction interval for ITE as $\{0\}$, even when the data yield a confidence interval for the ATE that excludes zero. Formally, if $\alpha = 0.05$, it is possible that $95\%$ of the individuals have a zero treatment effect, while the ATE analysis, being an aggregate statement, detects a statistically meaningful difference overall from the rest of the $5\%$ population. This implies, perhaps paradoxically, we can have evidence against Neyman null but also evidence in support of Fisher null.
Non-individualized decision policies can consequently outperform individualized decision policies in some partial identification settings; see cui2021individualized. Even with an ATE significantly different from zero, substantial uncertainty at the individual level (reflected in wide ITE intervals) can yield a scenario where assigning the same treatment to everyone is more reliable than any personalized rule that attempts to exploit covariate information. This outcome highlights how partial identification and weakly informative data can obscure the individual-level signal, despite showing clear evidence of an overall effect.
Consider a randomized trial for $n=50000$ participants with binary outcomes, divided into treatment and control groups using a coin flip. Suppose that treatment cures $3\%$ of the population, hurts $1\%$ of the population, while having no effect on $96\%$ of the populations. We observe the following marginals in Table (ref).
We estimate the average treatment effect to be $0.0198$ with $95\%$ confidence interval $[0.0174 , 0.0223]$. This yields a confidence interval for the average treatment effect excluding zero. However, based on the result in Section (ref), we will conclude a $95\%$ prediction interval for ITE as a singleton $\{0\}$.
Although this situation might seem paradoxical, it simply reflects the fact that population-level inferences about the mean effect can diverge from inferences about individual-level effects under uncertainty. On the one hand, one can accumulate enough information to claim that the average effect is nonzero, while on the other hand, one lacks the precision required to identify which individuals truly benefit.
Even with unconfoundedness (or randomization), the difficulty of estimating ITE lies in the unknown structure of potential outcomes $Y_1, Y_0$. If $Y_1$ and $Y_0$ are independent, then the fundamental problem of causal inference does not exist. For example, the introduction section of yin2022conformal can be confusing. Similarly, the section on numerical experiments in lei2021conformal also assumes the independence of $Y_1$ and $Y_0$.
A complementary strategy under partial identification is to condition on covariates $X$, seeking subpopulations where treatment effects are more homogeneous. Our results can be extended to this scenario when the treatment is conditionally independent of the potential outcomes. For example, if there exist covariates for which $P(Y=1 \mid D=1,X)$ or $P(Y=1 \mid D=0,X)$ are close to $0$ or $1$, the uncertainty about individual-level effects in that subgroup may be greatly reduced, enabling tighter bounds on $P\bigl(Y_{1}-Y_{0} \in [L,R] \mid X\bigr)$. In Appendix (ref), we consider the ITE prediction intervals in the setting of randomized experiments where additional covariates are observed for each individual and provide an example of how the conditional ITE prediction interval can be even wider compared to the normal ITE prediction interval. We also need to be cautious with such conditional individual treatment effect prediction intervals as they could disagree with the conditional average treatment effect (CATE) estimation results.
We thank James M. Robins, Carlos Cinelli, Yanqin Fan, Ting Ye and Gary Chan for valuable input and discussion.
\singlespacing
\doublespacing