EconBase
← Back to paper

Individual Treatment Effect: Prediction Intervals and Sharp Bounds

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

66,607 characters · 23 sections · 35 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Individual Treatment Effect: Prediction Intervals and Sharp Bounds

\doublespacing

abstractIndividual treatment effect (ITE) is often regarded as the ideal target of inference in causal analyses and has been the focus of several recent studies. In this paper, we describe the intrinsic limits regarding what can be learned concerning ITEs given data from large randomized experiments. We consider when a valid prediction interval for the ITE is informative and when it can be bounded away from zero. The joint distribution over potential outcomes is only partially identified from a randomized trial. Consequently, to be valid, an ITE prediction interval must be valid for all joint distribution consistent with the observed data and hence will in general be wider than that resulting from knowledge of this joint distribution. We characterize prediction intervals in the binary treatment and outcome setting, and extend these insights to models with continuous and ordinal outcomes. We derive sharp bounds on the probability mass function (pmf) of the individual treatment effect (ITE). Finally, we contrast prediction intervals for the ITE and confidence intervals for the average treatment effect (ATE). This also leads to the consideration of Fisher versus Neyman null hypotheses. While confidence intervals for the ATE shrink with increasing sample size due to its status as a population parameter, prediction intervals for the ITE generally do not vanish, leading to scenarios where one may reject the Neyman null yet still find evidence consistent with the Fisher null, highlighting the challenges of individualized decision-making under partial identification.

Introduction

The traditional causal inference literature has been focused on population level treatment effect parameters such as the average treatment effect (ATE) and conditional average treatment effect (CATE). Although the individual treatment effect (ITE) is often regarded as the ideal parameter of interest for personalized decision making, it is not, in general identifiable even if we know the outcome for an individual and have data from a large randomized experiment. Recent works by lei2021conformal, jin2023sensitivity, chernozhukov2023toward,wang2025conformal discuss conformal inference methods to estimate ITE and prediction intervals for the ITE. Another recent debate from mueller2022personalized and dawid2023personalised also explores the possibility of using ITE and bounds on ITE to help personalized decision making.

Bounds and relationships on the probability of the counterfactual treatment effect has been studied in terms of probability of causation in robins1989probability and tian2000probabilities. Some more recent work from mueller2021causes,sani2023bounding,kawakami2025mediation extends these ideas to learn individual responses and bounds from causal diagrams and mediation analysis. Inference and examples of probability of causation have been discussed in dawid2016statistical. fan2010sharp study sharp bounds on the cdf of the individual treatment effect using copulas based on the previous result from frank1987best and williamson1990probabilistic. mullahy2018individual applies the bounds in fan2010sharp in health economics applications. Bounds and optimal policy for binary treatment and outcome have been studied in kallus2022treatment, kallus2022s, and dawid2023personalised. Individual treatment effects on ordinal outcomes have been studied in lu2018treatment. Numerical approaches for solving causal inference problems in discrete settings have been explored in duarte2024automated. Exact inference on individual treatment effects based on permutation has been studied in blaker2000confidence,rigdon2015randomization,chiba2015exact. The construction of confidence intervals for causal parameters has been discussed in robins1988confidence, imbens2018causal and brennan2024causal.

In this work, we try to understand the prediction intervals and bounds for the ITE given that we have observations from well-conducted large randomized control trials (RCT). In Section (ref), we start from the binary treatment and outcome model and provide a complete characterization of ITE prediction intervals under this simple model setting. We characterize when a degenerate interval consisting of a single value is a valid prediction interval, when is the valid prediction interval for ITE non-negative/non-positive, and when the only valid prediction interval is a trivial interval. Additionally, we give conditions on the observed data for there to exist a consistent joint distribution under which a given non-trivial prediction interval would be valid. In Section (ref), we extend our insights to continuous and ordinal outcomes. Our approach leverages cdf bounds for the difference of two random variables from fan2010sharp and zhang2024bounds. In Section (ref), we provide sharp bounds on the cumulative distribution function (cdf) and probability mass function (pmf) of the ITE under binary treatment and outcome model. We explore the relationships between the cdf/pmf bounds on ITE and the Fr\'{e}chet-Hoeffding bounds on two random variables. This motivates the general sharp bounds on the pmf of the ITE for ordinal outcome in section (ref). Lastly, we compare ITE prediction intervals and ATE confidence intervals and provide a synthetic data example to discuss the implications of them in section (ref).

The limits to inference for individual treatment effects in binary treatment and outcome model

Throughout the paper, we consider a binary treatment setting with treatment $D=0,1$. Let $Y_1$ be the potential outcome when receiving the treatment and $Y_0$ be the potential outcome when not receiving the treatment. We define the individual treatment effect as:

align*[align* omitted — 35 chars of source]

Let $Y$ be the observed variable. By consistency, $Y=Y_0$ when $D=0$ and $Y=Y_1$ when $D=1$. In this section, we will examine what is possible to learn in the limit of a large sample size.

Definition of a prediction interval for the individual treatment effect (ITE)

A $(1-\alpha)$ {\em prediction interval} for an individual treatment effect is an interval such that \[ P\left((Y_1-Y_0) \in [L,R]\right) \geq 1-\alpha. \] Throughout the paper, we assume $\alpha$ is sufficiently bounded away from $0.5$.

propositionSuppose an interval $\mathbb{I}$ is a valid $(1-\alpha)\%$ prediction interval. If $P(\mathrm{ITE}\in A)>\alpha$ for some set $A$, then $\mathbb{I}\cap A \neq \emptyset$. If $P(\mathrm{ITE}\in A)\leq \alpha$ for some set $A$, then there exists a $(1-\alpha)$ valid prediction set that does not intersect with $A$. In particular, if $\mathbb{R}\setminus A$ is an interval, then it will be a valid $(1-\alpha)$ prediction interval.
proofIf $P(\mathrm{ITE}\in A)> \alpha$ and $\mathbb{I}\cap A = \emptyset$, then $P(\mathrm{ITE}\in\mathbb{I})\leq 1-P(\mathrm{ITE}\in A)<1-\alpha$. The interval $\mathbb{I}$ does not have $1-\alpha$ coverage. If $P(\mathrm{ITE}\in A)\leq \alpha$, then $P(\mathrm{ITE}\in (\mathbb{R}\setminus A))= 1- P(\mathrm{ITE}\in A)\geq 1-\alpha$. $\mathbb{R}\setminus A$ is a valid $(1-\alpha)$ prediction set.
corollaryIn discrete settings, whenever $P(ITE=i)>\alpha$, we must have $i\in \mathbb{I}$.

Binary Treatment and Outcome Model

To further narrow the discussion, we first consider the case in which the treatment and response are both binary. Following copas1973randomization who characterizes individual patients in the binary treatment and outcome model, we have the following four types and individual treatment effects:

center[center omitted — 204 chars of source]

Notice that in this setting the individual treatment effects take three possible values: $-1,0,1$.

In the simple setting that we consider, since there are only $3$ possible values taken by the ITE, there are $6$ possible prediction intervals:

equation*[equation* omitted — 81 chars of source]

Note that prediction intervals $[-1,0],[0,1],[-1, 1]$ here are sets correspond to $\{-1,0\}, \{0,1\}, \{-1,0,1\}$. Singleton sets can be viewed as degenerate intervals where the starting point equals to the end point. In this simple setting, there is only one prediction set that does not correspond to an interval, namely $\{-1,1\}$. We discussed it in Remark (ref).

When will the valid prediction interval of minimal length for a given joint distribution not be unique?

corollaryIf $P(HE\cup HU)>\alpha$ and $\max\{P(HE),P(HU)\}\leq \alpha$, then $\{0\}$ is not a valid $(1-\alpha)\%$ prediction interval but both $[-1,0]$ and $[0,1]$ are valid and minimal length $(1-\alpha)\%$ prediction intervals for the ITE.

Under the conditions stated in Corollary (ref), the set $\{0\}$ fails to provide sufficient coverage; yet each of $[-1,0]$ and $[0,1]$ captures a sufficiently large portion of the probability mass (at least $1-\alpha$). Both of these intervals therefore qualify as valid $(1-\alpha)\%$ prediction intervals for $\mathrm{ITE}$, and each achieves the same minimal length (one unit). Consequently, there is no single “shortest” interval that strictly dominates the other. In general, even given the joint distribution of $Y_0, Y_1$, there could be multiple minimal-length valid prediction intervals.

Partial Identification of the Joint Distribution over Types under Randomization

Under randomization, the relationship between the counterfactual distribution $P(Y_0, Y_1)$ and the observed distributions $\left\{P(Y\mid D=0), P(Y\mid D=1)\right\}$ is given by this table:

center[center omitted — 313 chars of source]

Here $P(Y\!=\!i\mid D\!=\!j) = P( Y_j\!=\!i)$ due to randomization.

Equivalently we may write this in terms of types:

center[center omitted — 246 chars of source]
proposition[Fr\'{e}chet inequality bounds] For two real valued random variables $Y_0, Y_1$ and any $(y_0,y_1)\in \mathbb{R}^2$, suppose that we know $P(Y_0=y_0)=a$ and $P(Y_1=y_1)=b$, then \begin{align*} \max\{0,a+b-1\}\leq P(Y_0=y_0,Y_1=y_1)\leq \min\{a,b\} \end{align*} This is also known as the Boole-Fr\'{e}chet inequality.

Since under randomization the joint distribution must add up to satisfy the observed distributions in the two treatment arms, we can parametrize the distribution of types using $P(AR)=t$. We then have the following solution set:

equation*[equation* omitted — 358 chars of source]

where, by Proposition (ref), we have

align*[align* omitted — 213 chars of source]

When is the only valid prediction interval trivial?

There are circumstances in which the only valid prediction interval for the individual treatment effect is the trivial interval: $[-1,1]$!

The interval will be trivial when the set of people of type Hurt can be larger than $\alpha$, {\em and} the set of people of type Helped can also be larger than $\alpha$.

Notice that the proportion Helped and Hurt, {\em both } achieve their maximum value when the proportion Always Recover ${\textcolor{red}{ t}}$ achieves its minimum value. Hence the only valid prediction interval will be trivial whenever we have both:

align*[align* omitted — 235 chars of source]

this may be equivalently expressed as:

align*[align* omitted — 177 chars of source]

Consequently, provided we have:

align*[align* omitted — 79 chars of source]

and

align*[align* omitted — 82 chars of source]

then the only valid prediction interval will be trivial. In concrete terms, if $\alpha=0.05$ and the proportions recovering under treatment ($D=1$) and control ($D=0$) both lie between $5\%$ and $95\%$ then, given that we know the conditional distributions $P(Y|D)$, the only valid 95% prediction interval for the individual treatment effect will be $[-1,1]$.

These results show that for randomized experiments in which the proportion of recovery in both arms lies between $\alpha$ and $(1-\alpha)$, the observed data is entirely uninformative regarding the ITE.

When is the valid prediction interval for the individual treatment effect a singleton?

Notwithstanding the results in the previous section, perhaps surprisingly, there are situations in which a valid $(1-\alpha)\%$ prediction interval is a singleton. We now characterize when this occurs. Detailed derivations are provided in the Appendix (ref). There are three cases to consider:

itemize$\{0\}$ is valid if the sum of AR and NR types is at least $1-\alpha$, i.e., both arms have nearly $\alpha\%$ or nearly $1-\alpha\%$ response such that the sum of the proportions of individuals across the two arms with the less common outcome is less than $\alpha$. Concretely, either \begin{align*} P(Y\!=\!1 \mid D\!=\!0)+ P(Y\!=\!1 \mid D\!=\!1)\leq \alpha \end{align*} or \begin{align*} P(Y\!=\!0 \mid D\!=\!0)+ P(Y\!=\!0 \mid D\!=\!1)\leq \alpha. \end{align*} Note however, that in such a case the average treatment effect, though less than $\alpha$, may be non-zero if $P(Y=0\mid D=0)\neq P(Y=0\mid D=1)$. Thus, given sufficiently large sample sizes, confidence intervals for the Average Treatment Effect will not include zero. We note that the singleton ITE prediction interval $\{0\}$ is equivalent to establishing the Fisherian sharp null hypothesis fisher1936design holds for at least $(1-\alpha)$ of the population. We will further discuss this in section (ref). • $\{1\}$ is valid if the HE type alone can exceed $1-\alpha$. This requires \begin{align*} P(Y=1\mid D=1) - P(Y=1\mid D=0) \geq (1-\alpha). \end{align*} meaning the average treatment effect is at least $1-\alpha$. • $\{-1\}$ is valid if the HU type alone can exceed $1-\alpha$. This requires \begin{align*} P(Y=1\mid D=0) - P(Y=1\mid D=1) \geq (1-\alpha). \end{align*} i.e., the average treatment effect is less than $-(1-\alpha)$.

When is the valid prediction interval for the individual treatment effect non-negative/non-positive?

Following the previous result, one can rule out negative (or positive) treatment effects if the proportion of HU (or HE) can never exceed $\alpha$. This occurs when either $P(Y=1 \mid D=0)$ or $1 - P(Y=1 \mid D=1)$ is below $\alpha$ (and similarly for ruling out positive effects). See Appendix (ref).

remark[] It is not possible to conclude $\{-1,1\}$ as a best prediction set given the observed marginals $P(Y=1\mid D=1), P(Y=1\mid D=0)$. For $\{-1,1\}$ to be a valid prediction set, the proportion of people of type Helped plus the proportion of people of type Hurt should always be greater than or equal to $1-\alpha$. From the lower bounds on proportion of people of type Helped and the proportion of people of type Hurt, we need \begin{equation*} \begin{aligned} &P(Y=1\mid D=1)+P(Y=1\mid D=0) \&-2\min\{P(Y=1\mid D=1),P(Y=1\mid D=0)\}\geq 1-\alpha, \end{aligned} \end{equation*} equivalently, either \begin{align*} P(Y=1\mid D=1)-P(Y=1\mid D=0)\geq 1-\alpha \end{align*} or \begin{align*} P(Y=1\mid D=0)-P(Y=1\mid D=1)\geq 1-\alpha. \end{align*} However, based on the result in Section (ref), in either setting, we would just conclude the singleton $\{-1\}$ or $\{1\}$ respectively as the prediction set instead of $\{-1,1\}$. Therefore, it is not possible to conclude $\{-1,1\}$ as a best prediction set from the observed marginals $P(Y=1\mid D=1), P(Y=1\mid D=0)$. Note that $\{-1,1\}$ is a theoretically possible prediction set which can be optimal for certain disributions over potential outcomes. It is just we will not be able to conclude it based on observed distribution from randomization.

Visual summary of results on valid ITE prediction intervals

Based on the result in (ref), (ref), and (ref), under randomization and in the limit of a large sample size where we observe the true marginals $P(Y=1\mid D=0), P(Y=1\mid D=1)$ and their counterparts, we can characterize the corresponding ITE prediction intervals. These intervals are valid no matter the joint distribution over $Y_0$ and $Y_1$ (provided it is compatible with $P(Y\mid D)$. These intervals are also “the best we can do" without additional assumptions or information, e.g. from a cross-over study.

figure[figure omitted — 189 chars of source]

Notice that we have two overlapping triangles each with area $\frac{1}{2}\alpha^2$ in Figure (ref) where both $[0,1]$ and $[-1,0]$ can serve as valid $\alpha-$level prediction interval for individual treatment effects. Assume that for the same length of intervals, intervals with higher coverage are "better". Then we can further decompose these triangular areas to obtain the "best" prediction intervals; see Figure (ref).

figure[figure omitted — 2,606 chars of source]

Necessary conditions for a given prediction interval to be valid and of minimal length

Suppose we are given a prediction interval and an observed distribution from an RCT, we can consider when does there exist a joint distribution over potential outcomes that is compatible with the observed distributions and for the prediction interval to be valid and of minimal length. In Appendix (ref), we derive the precise constraints on the marginal distributions \(\,P(Y=1 \mid D=0)\) and \(P(Y=1 \mid D=1)\) that guarantee each of the intervals can serve as a valid (or “best”) prediction interval for the individual treatment effect for some joint distribution over potential outcome. The necessary conditions take the form of simple inequalities relating the two response probabilities; they characterize scenarios in which each the true treatment effect with probability at least \(1-\alpha\) lies within the given set for som population compatible with $P(Y\mid D)$. Full details, derivations, and proofs appear in Appendix (ref).

Visualization of Necessary Conditions given an ITE prediction interval to be the best

Figure (ref) and (ref) provide visualizations of the necessary conditions on $P(Y=1\mid D=0)$ and $P(Y=1\mid D=1)$ for a given prediction interval to be valid/best. Note that if for a given point in the unit square, we consider the intervals (from Figure (ref) and (ref)) containing $(P(Y=1\mid D=1), P(Y=1\mid D=0))$ and then select the longest, we recover Figure (ref).

figure[figure omitted — 2,290 chars of source]
figure[figure omitted — 1,875 chars of source]

Beyond binary outcomes

We now turn to the case where the outcome is continuous, rather than binary. Our goal is to address questions analogous to those explored in the binary setting. In particular, we begin by investigating how to construct valid prediction intervals based solely on the marginal distributions, without imposing assumptions on the joint distribution of the potential outcomes. We then examine the conditions under which these intervals can be bounded away from zero. Further, we solve the problem: given a prediction interval (or set), what conditions on the observed data ensure the existence of a consistent joint distribution under which the interval is valid?

A conservative interval for continuous outcomes

The marginal distributions of $Y_1$ and $Y_0$ are identified from randomized experiments. Let $[L_0, R_0]$ be an interval such that $P\left(Y_0\in [L_0,R_0]\right) \geq 1-\alpha/2.$ Let $[L_1, R_1]$ be an interval such that $P\left(Y_1\in [L_1,R_1]\right) \geq 1-\alpha/2.$ Then

align[align omitted — 105 chars of source]
proofNote that if $Y_1-Y_0$ is not in $[L_1-R_0, R_1-L_0]$, then either $Y_1$ is not in $[L_1, R_1]$ or $Y_0$ is not in $[L_0,R_0]$. Put into set notation, \[ \{Y_1 - Y_0 \notin [L_1 - R_0,\; R_1 - L_0] \} \subseteq \{Y_1 \notin [L_1,R_1]\} \cup \{Y_0 \notin [L_0,R_0]\}. \] Applying the union bound to that set containment gives \[ P\Bigl(Y_1 - Y_0 \notin [L_1 - R_0, R_1 - L_0]\Bigr) \leq P\bigl(Y_1 \notin [L_1,R_1]\bigr) + P\bigl(Y_0\notin [L_0,R_0]\bigr). \] By hypothesis, $P\left(Y_0\notin [L_0,R_0]\right) < \alpha/2$ and $P\left(Y_1\notin [L_1,R_1]\right) < \alpha/2$, therefore, \[P\bigl(Y_1-Y_0 \notin [L_1,R_1]\bigr) < \alpha.\] Thus, $$P\left((Y_1-Y_0) \in [L_1-R_0, R_1-L_0]\right) \geq 1-\alpha.$$

This result is similar to the “naive" conformal inference prediction interval for the ITE given in lei2021conformal Section 4.1. Though we further assume an infinite sample size and do not use covariates.

Points that must be included in every valid $(1-\alpha)$ prediction interval

One might argue that the bounds in ((ref)) are too conservative. We will take a different perspective to see what are the points that must be included in the interval. Note that to be valid, an ITE prediction interval must be valid for all joint distributions consistent with the observed data, and hence will in general be wider than that resulting from knowledge of this joint distribution. Let $[L_0', R_0']$ be an interval such that \[L_0' := \min \Bigl\{\ell \in \mathbb{R} : P\bigl(Y_0 < \ell\bigr) > \alpha\Bigr\}\] and \[R_0' := \max \Bigl\{\ell \in \mathbb{R} : P\bigl(Y_0 > \ell\bigr) > \alpha\Bigr\}.\] In other words, $L'_0$ and $R'_0$ are the $\alpha$-quantile and $(1-\alpha)$-quantile of $Y_i(0)$. Similarly we can define \[L_1' := \min \Bigl\{\ell \in \mathbb{R} : P\bigl(Y_1 < \ell\bigr) > \alpha\Bigr\}\] and \[R_1' := \max \Bigl\{\ell \in \mathbb{R} : P\bigl(Y_1 > \ell\bigr) > \alpha\Bigr\}.\] Then a valid $(1-\alpha)$ prediction interval for the ITE must include these points:

itemize$R_1'-L_0'$, • $L_1'-R_0'$.

This follows because otherwise there exists a joint distribution of $Y_1, Y_0$ such that there is more than $\alpha$ mass outside of the prediction interval. This is because the only constraint imposed on the joint distribution by the marginals is given by the Fréchet inequalities, \[P(Y_1> R_1', Y_0<L_0')\leq \min\{P(Y_1> R_1'), P(Y_0<L'_0)\}\] By construction the minimum is greater than $\alpha$, meaning there exists a joint distribution that $P(Y_1> R_1', Y_0<L_0')>\alpha$. Hence under this joint distribution, $P(Y_1-Y_0>R_1'-L_0')>\alpha$, therefore, any interval that does not include $R_1'-L_0'$ will not be a valid $\alpha$ level prediction interval.

Even small tails in marginal distributions of $Y_1, Y_0$ can force the prediction interval to expand considerably at both extremes. We illustrate this in Figure (ref).

figure[figure omitted — 3,745 chars of source]

Can we obtain a prediction interval bounded away from zero?

In the previous section, we gave conditions under which a point is included in every valid ITE prediction interval. Here we now give necessary and sufficient conditions for a valid interval to exclude the half line $(0,\infty)$ (or $(-\infty,0)$). Though the arguments extend to any constant, not just zero, we now address the question: can a valid prediction interval for the individual treatment effect (ITE) be bounded away from zero? Suppose there exists a joint distribution of the potential outcomes such that $P(Y_1 - Y_0 \leq 0) > \alpha$. Then, without further assumptions on the joint distribution, the left endpoint of any valid $(1-\alpha)$ prediction interval must be less than or equal to zero. This is closely related to Kolmogorov's problem and the results in fan2010sharp, zhang2024bounds, which we review in Appendix (ref). In particular, consider the following upper bound on $P(Y_1 - Y_0 \leq 0)$:

align[align omitted — 195 chars of source]

If this upper bound exceeds $\alpha$, then zero must lie within the prediction interval, i.e., the left endpoint cannot be strictly greater than zero.

Similarly, if $P(Y_1 - Y_0 \geq 0) > \alpha$, then the right endpoint of any valid prediction interval must be at least zero. This probability can be written as \[ P(Y_1 - Y_0 \geq 0) = 1 - P(Y_1 - Y_0 < 0), \] so we may use a lower bound on $P(Y_1 - Y_0 < 0)$ to assess whether the right endpoint must include zero. The lower bound is given by:

align[align omitted — 78 chars of source]

If this lower bound is less than $1 - \alpha$, then the right endpoint must be at least zero.

Therefore, if both the left endpoint must be less than or equal to zero and the right endpoint must be greater than or equal to zero, the prediction interval necessarily includes zero and cannot be bounded away from it.

In fact, if the distribution of $Y_1-Y_0$ is absolutely continuous with respect to Lebesgue measure so $P(Y_1-Y_0=0)=0$, then we can combine the conditions in ((ref)) and ((ref)) to state that for a valid $(1-\alpha)$ prediction interval for the ITE to exclude zero, either the lower bound on $P(Y_1 - Y_0 \leq 0)$ to exceed $1 - \alpha$ or the upper bound to be less than $\alpha$—conditions that are only met in highly extreme cases.

In Appendix (ref), we briefly discuss the cases of bounding ITE prediction intervals in the ordinal outcome setting, using a new result that we will prove in Section (ref).

Fr\'{e}chet-Hoeffding bound on the pmf and cdf of ITE under binary treatment and outcome model

In the previous section, we used bounds on the cumulative distribution function (cdf) of the individual treatment effect (ITE) under known marginal distributions for potential outcomes. We now turn to the corresponding problem for the probability mass function (pmf). As before, we begin with the binary treatment and outcome setting. We first consider the binary treatment and binary outcome model where we know the marginals for $Y_0$ is $p, 1-p$ and the marginals for $Y_1$ is $q, 1-q$. We can parameterize the joint density using $P(Y_0=1, Y_1=0)=t$ and Table $\ref{table_binary}$.

center[center omitted — 404 chars of source]
definition[Fr\'{e}chet-Hoeffding bounds] For two real valued random variables $Y_1,Y_0$ and any $y_1,y_0\in \mathbb{R}$, suppose that we know $P(Y_0\leq y_0)=a$ and $P(Y_1\leq y_1)=b$, then \begin{align*} \max\{0,a+b-1\}\leq P(Y_0\leq y_0, Y_1\leq y_1)\leq \min\{a,b\}. \end{align*}

Note: Here we make the distinction that Fr\'{e}chet-Hoeffding bounds are bounding the joint cdf of $Y_1, Y_0$ while the Fr\'{e}chet inequality bounds are bounding the joint pmf of $Y_1, Y_0$.

definition[Comonotonicity] The bivariate random vector $Y=(Y_0,Y_1)\in \mathbb{R}^2$ is called comonotonic if \begin{align*} P(Y_0\leq y_0, Y_1\leq y_1)=\min\{P(Y_0\leq y_0),P(Y_1\leq y_1)\} \end{align*} for all $y_0,y_1\in \mathbb{R}$. In this case, the bivariate distribution functions $F(y_1,y_0)=P(Y_0\leq y_0, Y_1\leq y_1)$ achieves the Fr\'{e}chet-Hoeffding upper bound and we call the two random variables $Y_0, Y_1$ perfectly positively dependent.
definition[Countermonotonicity] The bivariate random vector $Y=(Y_0,Y_1)\in \mathbb{R}^2$ is called countermonotonic if \begin{align*} P(Y_0\leq y_0, Y_1\leq y_1)=\max\{0,P(Y_0\leq y_0)+P(Y_1\leq y_1)-1\} \end{align*} for all $y_0,y_1\in \mathbb{R}$. In this case, the bivariate distribution functions $F(y_1,y_0)=P(Y_0\leq y_0, Y_1\leq y_1)$ achieves the Fr\'{e}chet-Hoeffding lower bound and we call the two random variables $Y_0, Y_1$ perfectly negatively dependent.

In the binary treatment and outcome model, a perfect positive dependence of $Y_1, Y_0$ is achieved by $t=\min\{p,q\}$ and a perfect negative dependence of $Y_1,Y_0$ is achieved by $t=\max\{0,p+q-1\}$.

propositionWhen the treatment and outcome are both binary, the sharp pmf bounds on $Y_1-Y_0$ are achieved at the Fr\'{e}chet-Hoeffding bounds.

From Table (ref), we have $P(\text{ITE}=-1)=q-t$, $P(\text{ITE}=0)=1-p-q+2t$, $P(\text{ITE}=1)=p-t$. The upper and lower bounds for each ITE value are reached at the Fr\'{e}chet-Hoeffding bound on the joint distribution of $Y_1, Y_0$ (i.e. when $t$ takes minimum or maximum value).

We further observe that for $P(\text{ITE}=1)$, $P(\text{ITE}=-1)$, we can obtain pmf bounds by applying Fr\'{e}chet inequality bounds on each cell directly. And for $P(\text{ITE}=0)$, we can first obtain Fr\'{e}chet inequality bounds on the two cells $t\in[\max\{0,p+q-1\},\min\{p, q\}]$ and $1-q-p+t\in[\max\{0,1-p-q\},\min\{1-p, 1-q\}]$ and then add up the lower and upper bounds to obtain $1-q-p+2t\in [\max\{0,p+q-1\}+\max\{0,1-p-q\},\min\{p, q\}+\min\{1-p, 1-q\}]$. This simplifies to the bounds given by $1-p-q+2t, t\in [\max\{0,p+q-1\},\min\{p, q\}]$ as before. We will generalize this observation in section (ref).

propositionWhen the treatment and outcome are both binary, the sharp bounds on cdf of $Y_1-Y_0$ are achieved at the Fr\'{e}chet-Hoeffding bounds.

Fr\'{e}chet-Hoeffding bound on the joint distribution of $Y_1, Y_0$ implies that $t\in[\max\{0,p+q-1\},\min\{p, q\}]$. Since we can parameterize the joint probability with one parameter $t$, we have $P(\text{ITE}=-1)=q-t$, $P(\text{ITE}=0)=1-p-q+2t$. In the binary case, the cdf for the ITE is characterized by the values $F(-1)=P(\text{ITE}=-1)=q-t$, $F(0)=P(\text{ITE}\leq 0)=P(\text{ITE}=-1)+P(\text{ITE}=0)=1-p+t$, $F(1)=1$.

As noted in Section $2.1$ of fan2010sharp, sharp bounds on the cdf of the individual treatment effect are not achieved at the Fr\'{e}chet-Hoeffding lower and upper bounds (perfectly positive/ perfectly negative dependence) for the distribution of $Y_1, Y_0$. As a special case, the sharp bounds on cdf of ITE are achieved at the Fr\'{e}chet-Hoeffding bounds in binary treatment and outcome model.

propositionIn general, we can obtain valid pmf bounds from the cdf bounds. For example, if the ITE takes integer values, then $P(\mathrm{ITE}=i)= P(\mathrm{ITE}\leq i)-P(\mathrm{ITE}< i)=F(i)-F(i-1)$. If we know $F(i)\in [a,b]$ and $F(i-1)\in[c,d]$, we can obtain that \begin{align} P(\mathrm{ITE}=i)\in [a-d, b-c]. \end{align} When the treatment and outcome are both binary, the pmf bounds obtained from the sharp bounds on the cdf of $Y_1-Y_0$ using the above method are sharp.

Using the cdf bounds for $Y_1-Y_0$ in Proposition (ref), we have $$P(\mathrm{ITE}=1)=F(1)-F(0),$$ with $F(1)=1, F(0)=1-p+t\in[1-p+\max\{0,p+q-1\},1-p+\min\{p,q\}]$. Threrefore, $$P(\mathrm{ITE}=1)\in [p-\min\{p,q\}, p-\max\{0,p+q-1\}].$$ Similarly, $P(\mathrm{ITE}=0)=F(0)-F(-1)$ for $F(-1)\in [q-\min\{p,q\}, q-\max\{0,p+q-1\}]$. Thus, $$P(\mathrm{ITE}=0)\in[1-p-q+2\max\{0,p+q-1\},1-p-q+2\min\{p,q\}].$$ And $$P(\mathrm{ITE}=-1)=F(-1)=q-t, t\in[\max\{0,p+q-1\},\min\{p, q\}].$$ These bounds agree with the sharp bounds derived in Proposition (ref). However, in general, the pmf bounds derived from cdf bounds using $(\ref{eq:CDF_to_PMF})$ are not sharp. Therefore, we will explore sharp pmf bounds in Section (ref).

Sharp bounds on the pmf of Individual Treatment Effects

A natural question to ask here is: given the marginals of $Y_1, Y_0$, can we derive sharp bounds on the probability mass function (pmf) of ITE, for example, what are the bounds for $P(\text{ITE}=0)?$

This is given by adding up the Fr\'{e}chet inequality bounds for each $P(Y_1=i, Y_0=i-\delta)$.

theoremFor discrete random variables $Y_1, Y_0$ with given marginals and a given $\delta$, let $[L_i, U_i]=[\max\{P(Y_1=i)+P( Y_0=i-\delta)-1,0\},\min\{P(Y_1=i), P( Y_0=i-\delta)\}]$ be the Fr\'{e}chet inequality bounds for the joint probability $P(Y_1=i, Y_0=i-\delta)$. We claim that \begin{align} P(Y_1-Y_0=\delta)\in[\sum_i L_i, \sum_i U_i]. \end{align} Furthermore, the bounds in ((ref)) are sharp as there exists joint distributions of $Y_1,Y_0$ that satisfies the given marginals and achieves the upper or lower bound on $P(Y_1-Y_0=\delta)$.

Proof of Theorem (ref)

First, we show that the bound in ((ref)) is a valid bound. Notice

align*[align* omitted — 61 chars of source]

And for all $i$, Fr\'{e}chet inequality bounds states that

align*[align* omitted — 51 chars of source]

where

align*[align* omitted — 101 chars of source]

So the sum must be within the bound in $(\ref{PMF_bound})$.\\

Now, we will show that the bounds in $(\ref{PMF_bound})$ are tight in a sense that there exist joint distributions compatible with the marginal conditions that achieve the lower/upper bounds. First, we will state a very useful result in koperberg2024couplings.

Let $A$ and $B$ be sets and $R \subseteq A \times B$ a relation. Then for each $\mathrm{U} \subseteq A$ the set of neighbours of $\mathrm{U}$ in $R$ is denoted by $$ \mathcal{N}_{\mathrm{R}}(\mathrm{U})=\{\mathrm{b} \in \mathrm{B}:(\mathrm{U} \times\{\mathrm{b}\}) \cap \mathrm{R} \neq \varnothing\}. $$

theorem[Strassen's theorem for finite sets] Let $\mathrm{A}$ and $\mathrm{B}$ be finite sets, $\mathrm{P}$ and $\mathrm{P}^{\prime}$ probability measures on $\mathrm{A}$ and $\mathrm{B}$ respectively and $\mathrm{R} \subseteq A \times \mathrm{B}$ a relation. Then there exists a coupling $\hat{\mathrm{P}}$ of $\mathrm{P}$ and $\mathrm{P}^{\prime}$ that satisfies $\hat{\mathrm{P}}(\mathrm{R})=1$ if and only if $$ \mathrm{P}(\mathrm{U}) \leqslant \mathrm{P}^{\prime}\left(\mathcal{N}_{\mathrm{R}}(\mathrm{U})\right) \text {, for all } \mathrm{U} \subseteq A \text {. } $$

Now we can begin the proof with a proposition.

propositionIf \begin{align*} &\max\{P(Y_1=i)+P( Y_0=j)-1,0\}>0\\ and \quad &\max\{P(Y_1=k)+P( Y_0=l)-1,0\}>0 \end{align*} then we have either $i=k$ or $j=l$. In other words, at most one Fr\'{e}chet inequality lower bound can be non-zero if $i\neq k$ and $j\neq l$.
proofSuppose, for a contradiction that $i\neq j$, $k\neq l$ and \begin{align} P(Y_1=i)+P( Y_0=j)-1&>0 ;\\ P(Y_1=k)+P( Y_0=l)-1&>0. \end{align} Summing up $(\ref{s_p_1})$ and $(\ref{s_p_2})$, we get \begin{align*} P(Y_1=i)+P( Y_0=j)+ P(Y_1=k)+P( Y_0=l)-2>0. \end{align*} However, we also have \begin{align*} P(Y_1=i)+&P( Y_0=j)+P(Y_1=k)+P( Y_0=l)\&\leq \sum_iP(Y_1=i)+\sum_jP(Y_0=j)=2. \end{align*} Therefore, we got a contradiction.

We will now show that there exist a joint distribution of $Y_1, Y_0$ such that $P(Y_1=i, Y_0=i-\delta)=\max\{P(Y_1=i)+P( Y_0=i-\delta)-1,0\}=L_i$ for all $i$. By Proposition (ref), $P(Y_1=i, Y_0=i-\delta)\neq 0$ for at most one $i$. There are thus two cases to consider here: when $L_i=0$ for all $i$ and when $L_i=0$ for all $i$ except one.\\

Case $1$: when $P(Y_1=i, Y_0=i-\delta)=0$ for all $i$. Case $1$ implies that

align[align omitted — 177 chars of source]

We apply Theorem (ref). Let $A,B$ be the support of the specified margins for $Y_1, Y_0$ respectively. Naturally, we have $P(Y_1), P(Y_0)$ as two probability measures on $A,B$. Let $R=\{(i,j):i\in A, j\in B, i-j\neq \delta\}$. Strassen's theorem states that there exists a coupling $\hat{P}$ of $P(Y_1), P(Y_0)$ with $\hat{P}(R)=1$ if and only if

align*[align* omitted — 109 chars of source]

where $$ \mathcal{N}_{R}(U)=\{b \in B:(U \times\{b\}) \cap R \neq \varnothing\}. $$ By construction of $R$, for any $i\in A$,

align*[align* omitted — 59 chars of source]

Therefore for any $U\subseteq A$ with more than one element, we have $\mathcal{N}_{R}(U)=B$ and:

align*[align* omitted — 114 chars of source]

Thus the only non trivial constraints are for singleton sets $U$ given by:

align[align omitted — 119 chars of source]

The constraints in $(\ref{lower_bound_zero_constraint_strassen})$ are already satisfied by the assumption in $(\ref{lower_bound_zero_constraint})$. Therefore, by Strassen's theorem, there must exist a coupling $P^\prime$ such that $P^\prime(R)=1$. Thus, the lower bounds $L_i=0$ for all $i$ are achievable in case $1$.\\

Case $2$: Consider the case that $L_i=0$ for all $i$ except for $i=j$. We will explicitly construct a joint distribution for $Y_1, Y_0$ that satisfies:

align[align omitted — 196 chars of source]

Consider a distribution satisfies by ((ref)) and ((ref)). Further, let $P(Y_1=i, Y_0=k)=0$ for any $i\neq j$ or $k\neq j-\delta$. Let $P(Y_1=j, Y_0=k)=P(Y_0=k), k\neq j-\delta$ and $P(Y_1=i, Y_0=j-\delta)=P(Y_1=i), i\neq j$. For all $i\neq j$, we have

align*[align* omitted — 72 chars of source]

For all $k\neq j-\delta$, we have

align*[align* omitted — 64 chars of source]

Also, we have

align*[align* omitted — 237 chars of source]

where the second equality follows definition of $j$. Similarly,

align*[align* omitted — 237 chars of source]

All the $P(Y_1=i, Y_0=j)$ are non-negative. This is a valid joint distribution that satisfies the given marginals. Thus, the lower bounds are achievable. An example of the construction of the joint probability matrix in case 2 is given in Figure (ref). \\

figure[figure omitted — 2,060 chars of source]

Next, we show that there exists a joint distribution of $Y_1,Y_0$ such that $P(Y_1=i, Y_0=i-\delta)=\min\{P(Y_1=i), P( Y_0=i-\delta)\}=U_i$ for all $i$. We will construct a joint distribution of $Y_1, Y_0$ using the joint probability table explicitly. We first permute the joint probability table based on whether the upper bound is achieved at $P(Y_1=i)$ or $P(Y_0=i-\delta)$. Let $J_1$ be the set of $i$ such that $P(Y_1=i, Y_0=i-\delta)=P(Y_1=i)$ for $i\in J_1$ and $J_2$ be the set of $j$ such that $P(Y_1=j+\delta, Y_0=j)=P(Y_0=j)$ for $j\in J_2$. We can rearrange the joint probability table based on the partition of $J_1, J_2$. Let $|J_1|=n_1, |J_2|=n_2$. Let $N_1$ be the number of elements in the support of $Y_1$. Let $N_2$ be the number of elements in the support of $Y_0$. We permute the joint probability matrix of $Y_1, Y_0$ such that the first $n_1$ rows denotes the probability of elements in $J_1$ and the last $n_2$ rows denotes the probability of elements in $J_2$. Figure $\ref{Probability Table Permutation}$ shows an example of such permutation.

figure[figure omitted — 2,623 chars of source]

By construction, the first $n_1$ rows and the last $n_2$ columns are filled with $0$'s and $U_i$'s. We have an empty $(N_1-n_1)\times(N_2\times n_2)$ sub-matrix to fill. We also obtain a new set of margin constraints on the sub-matrix by subtracting the marginals with what we have already filled. Notice that the row/column sum of these new constraints for the $(N_1-n_1)\times(N_2\times n_2)$ matrix is given by

align*[align* omitted — 66 chars of source]

We can fill each entry of the sub-matrix by

align*[align* omitted — 136 chars of source]

Thus, we construct a valid joint distribution of $Y_1, Y_0$ with $P(Y_1=i, Y_0=i-\delta)=\min\{P(Y_1=i), P( Y_0=i-\delta)\}=U_i$ for all $i$ and satisfies the given marginals. This concludes our proof of theorem (ref). We note that the same proof technique can be used to obtain the pmf bounds on the sum of two discrete random variables given the marginals.

corollaryIn the case where the distribution of one potential outcome is degenerate, the individual treatment effect is fully identified. For example, if we assume $P(Y_0=0)=1$, then the upper and lower bounds in Theorem (ref) coincide and we can identify $P(Y_1-Y_0=\delta)$.
proofThis follows from the fact that $[L_i, U_i]=[\max\{P(Y_1=i)+P( Y_0=i-\delta)-1,0\},\min\{P(Y_1=i), P( Y_0=i-\delta)\}]=[0,0]$ when $i-\delta\neq 0$ and $[L_i, U_i]=[\max\{P(Y_1=i)+P( Y_0=i-\delta)-1,0\},\min\{P(Y_1=i), P( Y_0=i-\delta)\}]=[P(Y_1=i),P(Y_1=i)]$ when $i-\delta =0$. Therefore, $\sum_iL_i=\sum_iU_i=P(Y_1=\delta)$. This means that in the case where $P(Y_0=0)=1$, the bounds in Theorem (ref) tell us that $P(Y_1-Y_0=\delta)=P(Y_1=\delta)$.

Discussion of ATE versus ITE

In this section, we discuss the relationship between the average treatment effect (ATE) and the individual treatment effect (ITE). We focus on the implications for prediction intervals and hypothesis testing. The ATE is a well-defined parameter and is identifiable under standard assumptions. Consequently, as the sample size increases, confidence intervals for the ATE will eventually shrink to a point. In contrast, the individual treatment effect (ITE) is not a parameter in the classical sense: it is a random quantity defined at the unit level. Although prediction intervals for the ITE may become narrower with larger sample sizes, they do not, in general, converge to zero width. This reflects the inherent uncertainty in predicting individual level responses, even when the population-level effect is precisely estimated.

The Neyman null hypothesis posits a zero effect on average, while the Fisher null hypothesis posits no effect for every individual. By construction, a true Fisher null implies the Neyman null; if all individual effects are zero, their average must be zero. A true Fisher null also implies that $\{0\}$ will be a valid prediction interval for any level $\alpha$. A true Neyman null implies that confidence interval for ATE will contain zero with probability $(1-\alpha)\%$.

When performing hypothesis testing, $\{0\}$ being a valid prediction interval can be considered as evidence in support of the Fisher null while confidence interval does not contain $0$ can be considered as evidence against Neyman null. However, in finite samples, a failure to reject the Fisher null does not imply the ATE is in fact zero, and rejecting the Neyman null (i.e. finding a nonzero estimated ATE) does not automatically imply that $\{0\}$ cannot be a valid prediction interval.

It is possible to construct a valid prediction interval for ITE as $\{0\}$, even when the data yield a confidence interval for the ATE that excludes zero. Formally, if $\alpha = 0.05$, it is possible that $95\%$ of the individuals have a zero treatment effect, while the ATE analysis, being an aggregate statement, detects a statistically meaningful difference overall from the rest of the $5\%$ population. This implies, perhaps paradoxically, we can have evidence against Neyman null but also evidence in support of Fisher null.

Non-individualized decision policies can consequently outperform individualized decision policies in some partial identification settings; see cui2021individualized. Even with an ATE significantly different from zero, substantial uncertainty at the individual level (reflected in wide ITE intervals) can yield a scenario where assigning the same treatment to everyone is more reliable than any personalized rule that attempts to exploit covariate information. This outcome highlights how partial identification and weakly informative data can obscure the individual-level signal, despite showing clear evidence of an overall effect.

Synthetic data example

Consider a randomized trial for $n=50000$ participants with binary outcomes, divided into treatment and control groups using a coin flip. Suppose that treatment cures $3\%$ of the population, hurts $1\%$ of the population, while having no effect on $96\%$ of the populations. We observe the following marginals in Table (ref).

table[table omitted — 456 chars of source]

We estimate the average treatment effect to be $0.0198$ with $95\%$ confidence interval $[0.0174 , 0.0223]$. This yields a confidence interval for the average treatment effect excluding zero. However, based on the result in Section (ref), we will conclude a $95\%$ prediction interval for ITE as a singleton $\{0\}$.

Although this situation might seem paradoxical, it simply reflects the fact that population-level inferences about the mean effect can diverge from inferences about individual-level effects under uncertainty. On the one hand, one can accumulate enough information to claim that the average effect is nonzero, while on the other hand, one lacks the precision required to identify which individuals truly benefit.

Further discussions

Even with unconfoundedness (or randomization), the difficulty of estimating ITE lies in the unknown structure of potential outcomes $Y_1, Y_0$. If $Y_1$ and $Y_0$ are independent, then the fundamental problem of causal inference does not exist. For example, the introduction section of yin2022conformal can be confusing. Similarly, the section on numerical experiments in lei2021conformal also assumes the independence of $Y_1$ and $Y_0$.

A complementary strategy under partial identification is to condition on covariates $X$, seeking subpopulations where treatment effects are more homogeneous. Our results can be extended to this scenario when the treatment is conditionally independent of the potential outcomes. For example, if there exist covariates for which $P(Y=1 \mid D=1,X)$ or $P(Y=1 \mid D=0,X)$ are close to $0$ or $1$, the uncertainty about individual-level effects in that subgroup may be greatly reduced, enabling tighter bounds on $P\bigl(Y_{1}-Y_{0} \in [L,R] \mid X\bigr)$. In Appendix (ref), we consider the ITE prediction intervals in the setting of randomized experiments where additional covariates are observed for each individual and provide an example of how the conditional ITE prediction interval can be even wider compared to the normal ITE prediction interval. We also need to be cautious with such conditional individual treatment effect prediction intervals as they could disagree with the conditional average treatment effect (CATE) estimation results.

Acknowledgments

We thank James M. Robins, Carlos Cinelli, Yanqin Fan, Ting Ye and Gary Chan for valuable input and discussion.

\singlespacing

\doublespacing