Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
163,296 characters · 22 sections · 109 citation commands
What Makes Treatment Effects Identifiable? Characterizations and Estimators Beyond Unconfoundedness
{ \linespread{1} }
Understanding cause and effect is a central goal in science and decision-making. Across disciplines, we ask: What is the effect of a new drug on disease rates? How does a policy impact growth? Is technology driving economic growth? Causal inference tackles such questions by disentangling correlation from causation. Unlike statistical learning, which predicts outcomes from data, causal inference estimates the effects of interventions that alter the data-generating process.
A fundamental challenge in causal inference is that we can never observe both potential outcomes for the same individual. For example, if a patient takes a medication and recovers, we do not know whether the patient would have recovered without it. This fundamental problem of causal inference implies that causal effects must be inferred under certain assumptions holland1986statistics.
To formalize this challenge, we consider the widely-used potential outcomes model introduced by neyman1990applications (originally published in 1923) and later formalized by rubin1974estimating; see also \citet*{hernan2023causal,rosenbaum2002observational,chernozhukov2024appliedcausalinferencepowered}. Here, for a unit with covariates $X \in \mathbb{R}^d$, $Y(1)$ and $Y(0)$ denote potential outcomes under treatment and control, respectively. Since only the outcome $Y(T)$ corresponding to the assigned treatment $T$ is observed, certain assumptions are needed to estimate the average treatment effect (ATE), defined as $\tau \coloneqq \operatornamewithlimits{\mathbb{E}}[Y(1) - Y(0)]$, where $(T, Y(1), Y(0))$ are random variables whose distribution may depend on $X$. This framework underpins many modern causal inference methods -- both practical and theoretical -- and can capture many treatment effects, apart from $\tau,$ such as the average treatment effect on the treated (ATT), defined as $\gamma\coloneqq \operatornamewithlimits{\mathbb{E}}[Y(1)-Y(0)\mid T{=}1]$. Two fundamental questions under this framework, studied since cochran1965observationalStudies, rubin1974estimating, rubin1978randomization, heckman1979SelectionBias, are as follows:
Due to the missingness in data (explained above), even the identification problem is unsolvable without making structural assumptions on the distribution of $(X, T, Y(T))$, which is a censored version of the (complete) data distribution $(X,T,Y(1),Y(0))$. The earliest and most widely used such assumptions are unconfoundedness and overlap.
Unconfoundedness (a.k.a.,\ ignorability, conditional exogeneity, conditional independence, selection on observables) and overlap (a.k.a.,\ positivity and common support) are essential for unbiased estimation of the average treatment effect and are widely studied across {Statistics (e.g., rosenbaum2002observational,hernan2023causal,rubin1974estimating,rubin1977regressionDiscontinuity,rubin1978randomization,rosenbaum1983central) and many other disciplines, including} Medicine (e.g., rosenbaum1983central), Economics (e.g., athey2017CausalityReview,dehejia1998causal,dehejia2002propensity,abadie2006large,abadie2016matching), Political Science (e.g., brunell2004turnout,sekhon2004quality,ho2007matching), Sociology (e.g., morgan2006matching,lee2009estimation,oakes2017methods), and other fields (e.g., austin2008critical). Despite their wide use across different disciplines, there are fundamental instances where unconfoundedness or overlap are easily violated.
Unconfoundedness is often violated in observational studies, where treatments or exposures are not assigned by the researcher but observed in a natural setting. In a prospective cohort study, for example, individuals are followed over time to assess how exposures influence outcomes. A common violation arises when key confounders are unmeasured. For instance, in studying smoking’s impact on health, omitting socioeconomic status (SES), which affects both smoking habits and health, can bias results, as lower SES correlates with higher smoking rates and poorer health, independent of smoking.
Overlap is violated when certain covariate values make treatment assignments nearly deterministic. In a marketing study estimating the effect of personalized advertisements on purchases, covariates like demographics, browsing history, and preferences define a high-dimensional feature space. As this space grows, many user profiles either always or never receive the ad, leading to lack of overlap damour2021highDimensional. Without comparable treated and untreated units, causal inference methods struggle to estimate counterfactual outcomes, yielding unreliable effect estimates.
We refer the reader to (ref) for an in-depth discussion of scenarios demonstrating the fragility of unconfoundedness and overlap. Further, while Randomized Controlled Trials (RCTs) can eliminate hidden factors that lead to violation of unconfoundedness or overlap, they are often very expensive and, even unethical, for treatments that can harm individuals. Moreover, even RCTs can violate unconfoundedness due to participant non-compliance; see (ref).
These examples lead us to the following question, which we answer.
This question is not new and can be traced back to at least the work of rubin1977regressionDiscontinuity, who recognized that, without substantial overlap between treatment and control groups, identification of treatment effects necessarily requires additional prior assumptions. To the best of our knowledge, the present work provides the first formal characterization of the precise assumptions required to identify treatment effects in scenarios lacking substantial overlap, unconfoundedness, or both.
The main conceptual contribution of this work is a learning-theoretic approach that enables a characterization of when identification and estimation of {treatment effects} are possible. Before presenting this approach, it is instructive to reconsider how {unconfoundedness and overlap} enable identification {of the simplest and most widely used treatment effect -- the average treatment effect}: Given the observational study $\euscr{D}$, which is a distribution over $(X,T,Y(0),Y(1))$, unconfoundedness and overlap put a strong constraint on $\euscr{D}$: they require that $Y(t) \perp T~|~X{=}x$ for each $t \in \{0,1\}$ and $x \in \mathbb{R}^d$ and that the propensity scores $e(x) = \Pr[T {=} 1 | X{=}x]$ are bounded away from 0 and 1. Under these assumptions, identification and estimation of ATE $\tau = \tau_{\euscr{D}}$ is possible given censored samples $(X,T,Y(T))$ due to the following decomposition of $\tau_\euscr{D}$ for a fixed $x \in \mathbb{R}^d$ (we integrate over the $x$-marginal to get $\tau_\euscr{D}$): \[ \operatornamewithlimits{\mathbb{E}}_{Y(0), Y(1)} \left[Y(1)-Y(0) \mid X{=}x\right] = \operatornamewithlimits{\mathbb{E}}_{Y(0), Y(1),T}\left[ \frac{Y(1)\cdot T}{e(X)} - \frac{Y(0)\cdot (1-T)}{1-e(X)} \;\middle|\; X{=}x\right] \,, \addtocounter{equation}{1}\tag{\theequation}\label{eq:decomposition:unconfoundedness} \] {where we use overlap to divide with $e(X), 1-e(X)$ and unconfoundedness to obtain the equation $\operatornamewithlimits{\mathbb{E}}[Y(1)\cdot T\mid X] =\operatornamewithlimits{\mathbb{E}}[Y(1)\mid X]\cdot \Pr[T{=}1\mid X] $ and analogously for $Y(0)$.} Note that all the quantities appearing in the RHS are identifiable and estimable\footnote{We remark that the problem of estimating the propensity scores $e(x) = \Pr[T{=}1|X{=}x]$ is identical to the classical problem of learning probabilistic concepts kearns1994pconcept. We refer the reader to (ref) for details.} from the censored distribution $\euscr{C}_\euscr{D}$ rubin1978randomization, which is defined over $(X,T,Y(T))$.
{When} no {constraints are} put on $\euscr{D}$, identification of ATE is impossible in general imbens2015causal. Without unconfoundedness, propensity scores $\Pr[T{=}1\mid X{=}x]$ are not sufficient to identify the distribution of $T$, which can also depend on the outcomes $Y(0)$ and $Y(1)$ (conditioned on $X{=}x$). Instead, we can decompose the expression of $\tau_\euscr{D}$ for a fixed $x \in \mathbb{R}^d$ as follows: \[ \operatornamewithlimits{\mathbb{E}}_{\substack{Y(0),Y(1)}}\left[Y(1)-Y(0)\mid X{=}x\right] = \operatornamewithlimits{\mathbb{E}}_{\substack{Y(0), Y(1), T}}\left[ \frac{Y(1)\cdot T}{\Pr[T{=}1|X, Y(1)]} - \frac{Y(0)\cdot (1-T)}{\Pr[T{=}0|X, Y(0)]} \;\middle|\; X{=}x \right]\,. \] If unconfoundedness holds, then we could recover (ref) since then $T$ would not depend on $Y(1), Y(0)$ given $X.$ {However, unlike} the previous decomposition of (ref), the above equation always holds and crucially utilizes the generalized propensity scores $p_t(x,y) = \Pr[T{=}t\mid X{=}x, Y(t) {=}y]$ with $t \in \{0,1\}$.\footnote{Observe that we need some overlap condition to divide by $p_0(\cdot)$ and $p_1(\cdot)$ in the above equation. In our main results, however, we do not follow this decomposition and will not need such overlap conditions.} Unfortunately, these generalized propensity scores, in contrast to the standard propensity scores, are not always identifiable from data. To understand when these are identifiable, we need to consider the joint distribution of covariates and outcomes $\euscr{D}_{X,Y(t)}$ for each $t \in \{0,1\}$.
To this end, we adopt an approach inspired by statistical learning theory valiant1984theory,vapnik1999overview,blumer1989learnability,hastie2013elements,AnthonyBartlett1999NNLearning,alon1997scale,lugosi2002pattern,massartNoise2006,vapnik2006estimation,bousquet2003introduction,bousquet2003new. We introduce concept classes for the two key quantities derived by the above discussion $p_t$ and $\euscr{D}_{X,Y(t)}$ (for each $t \in \{0,1\}$) that will place some restrictions on the observational study $\euscr{D}$ {towards understanding which conditions enable} identification and estimation. In the remainder of the paper, we assume that all distributions are continuous and have a density. (All results also extend to discrete domains by replacing densities by probability mass functions.)
We are interested in the structure of two concept classes: the class of generalized propensity scores $\mathbbmss{P} \subseteq \{p \colon \mathbb{R}^d \times \mathbb{R} \to [0,1]\}$ and the class of {covariate-outcome} distributions $\mathbbmss{D} \subseteq \Delta(\mathbb{R}^d \times \mathbb{R})$. As in classical {statistical} learning theory, having fixed the concept classes, our next step is to restrict the underlying distribution $\euscr{D}$ to be realizable with respect to the {pair of concept classes $(\mathbbmss{P}, \mathbbmss{D})$}. An observational study is said to be realizable with respect to the concept class pair $(\mathbbmss{P}, \mathbbmss{D})$ if the generalized propensity scores {$p_0(\cdot),p_1(\cdot)$} induced by $\euscr{D}$ belong to $\mathbbmss{P}$ and $\euscr{D}_{X,Y(t)} \in \mathbbmss{D}$ for {each} $t \in \{0,1\}$. This learning-theoretic framework is quite expressive. For instance, it can capture unconfoundedness and overlap\footnote{ We refer to overlap as $c$-overlap: for some absolute constant $c \in (0,\nicefrac{1}{2})$, $c < {p_0(x,y),p_1(x,y)} < 1-c$.} by letting $\mathbbmss{D}$ be the set of all distributions over $\mathbb{R}^d\times \mathbb{R}$, {denoted by $\mathbbmss{D}_{\rm all}$}, and restricting $\mathbbmss{P}$ to be the following class \[ \phantom{.}\mathbbmss{P}_{\rm OU}(c)\coloneqq \left\{p\colon\mathbb{R}^d\times \mathbb{R}\to [0,1]\;\middle|\; p(x,y)=p(x,z) \text{ and } c<p(x,y)<1-c \text{ for each $(x,y,z)$}\right\}. \hspace{-2mm} \addtocounter{equation}{1}\tag{\theequation}\label{eq:PinScenarioI} \] That is, $\euscr{D}$ satisfies unconfoundedness and $c$-overlap if and only if it is realizable with respect to the pair of classes $(\mathbbmss{P}_{\rm OU}(c), \mathbbmss{D}_{\rm all})$.
Before proceeding to our results, we introduce some further terminology. Given classes $(\mathbbmss{P}, \mathbbmss{D})$, we are {particularly} interested in the generalized propensity scores in $\mathbbmss{P}$ and covariate-outcome distributions in $\mathbbmss{D}$ that induce a valid observational study. To this end, we say that {a} tuple $(p,\euscr{P}) \in \mathbbmss{P} \times \mathbbmss{D}$ is compatible with classes $(\mathbbmss{P},\mathbbmss{D})$ (henceforth, just compatible) if there exists another tuple $(\wh p, \wh \euscr{P})\in \mathbbmss{P}\times \mathbbmss{D}$ such that setting $(p_0,p_1,\euscr{D}_{X,Y(0)},\euscr{D}_{X,Y(1)})=(p,\widehat{p},\euscr{P},\widehat{P})$ (or equivalently the re-ordered {assignment} $(p_0,p_1,\euscr{D}_{X,Y(0)},\euscr{D}_{X,Y(1)})=(\widehat{p},p,\widehat{P},\euscr{P})$) defines a valid observational study, i.e., a valid distribution over $(X,T,Y(0),Y(1))$.
We say that a certain treatment effect ${\eta}_\euscr{D}$ is identifiable from the censored distribution $\euscr{C}_\euscr{D}$ when $(\mathbbmss{P}, \mathbbmss{D})$ satisfy {some} Condition C, if there is a mapping $f$ such that $f(\euscr{C}_\euscr{D}) = {\eta}_\euscr{D}$ for any observational study $\euscr{D}$ realizable with respect to $(\mathbbmss{P}, \mathbbmss{D})$ that satisfy C; in other words, if $\eta_{\euscr{D}_1} \neq \eta_{\euscr{D}_2}$ then it should be $\euscr{C}_{\euscr{D}_1} \neq \euscr{C}_{\euscr{D}_2}$ (see also (ref) for a formal definition). Having set the stage, we now ask our first main question:
As a first contribution, we identify a condition on the classes $(\mathbbmss{P}, \mathbbmss{D})$ that will be crucial for the results on the identification of ATE and ATT that proceed.
Observe that in (ref) we only focus on tuples that are compatible. This is due to the fact that incompatible tuples of $\mathbbmss{P} \times \mathbbmss{D}$ cannot be part of a realization of any valid observational study and hence properties of incompatible tuples are not relevant to our characterizations.
To gain some intuition for (ref), consider two observational studies $\euscr{D}_1$ and $\euscr{D}_2$ which correspond to the pairs $(p, \euscr{P})$ and $(q,\euscr{Q})$ respectively, where $\euscr{P}$ and $\euscr{Q}$ are distributions of $(X,Y(1)).$ Assume that the true observational study $\euscr{D}$ is either $\euscr{D}_1$ or $\euscr{D}_2$. Given the censored distribution $\euscr{C}_\euscr{D}$, we want to identify $\operatornamewithlimits{\mathbb{E}}_{\euscr{D}}[Y(1)].$ First, suppose that the tuples $(p, \euscr{P}), (q, \euscr{Q})$ satisfy Requirement 1 in (ref). Then we are done since we only care about the expected outcomes $\operatornamewithlimits{\mathbb{E}}_{\euscr{D}}[Y(1)] = \operatornamewithlimits{\mathbb{E}}_{(x,y) \sim \euscr{P}}[y] = \operatornamewithlimits{\mathbb{E}}_{(x,y) \sim \euscr{Q}}[y]$ which are the same under both distributions. Next, let us assume that Requirement 1 is violated and, hence, the expected treatment outcome is different between the null and the alternative hypothesis. In this case, if Requirement 2 is satisfied, then we can distinguish $\euscr{P}$ and $\euscr{Q}$ from $\euscr{C}_{\euscr{D}}$ (by comparing $\euscr{P}_X$ and $\euscr{Q}_X$ to the covariate marginal of $\euscr{C}_{\euscr{D}}$) and, hence, distinguish between $\euscr{D}_1$ and $\euscr{D}_2$. Finally, if both Requirements 1 and 2 fail but Requirement 3 holds, then $p(x,y) \euscr{P}(x,y)$ is proportional to the density of $(X,T,Y(1))$ in the censored distribution on each point $(x,y)$. Using this, we can again distinguish between the null and the alternative hypothesis. (Notice that, {in both the second and third steps,} we can distinguish between distributions that differ on a measure-zero set {since we allow the identification algorithms to be a function of the whole density. If one does not allow this, then one needs to consider the “almost everywhere” analogue of (ref).})
Our first result states that (ref) fully characterizes the ATE {in} any observational study $\euscr{D}$ realizable with respect to $(\mathbbmss{P}, \mathbbmss{D})$.
Interestingly, we show that (ref) also characterizes the identifiability of the average treatment effect on the treated (ATT), i.e., $\gamma_\euscr{D}\coloneqq \operatornamewithlimits{\mathbb{E}}[Y(1)-Y(0)|T{=}1]$. ({When talking about ATT, to simplify the results and exposition, we assume that $\Pr[T{=}1] > 0$}.)
Thus, the above two results, together, imply that {identifiability of} ATT and ATE is characterized by the same (ref); up to the mild assumption that $\Pr[T{=}1]>0$ which holds in any practical scenario where ATT identification is meaningful.
\noindentDiscussion. The above collection of results adds to classical {identifiability} conditions in Statistics (e.g., everitt2013finite,teicher1963identifiability), Statistical Learning Theory (e.g., angluin1980inductive,angluin1988identifying\footnote{The characterizing condition in language identification concerns pairs of languages angluin1980inductive. This is also the case in our setting (see (ref)). Intuitively, this is expected since identification in both problems requires being able to distinguish between pairs of task instances that have distinct ”identities.”}), and Econometrics (e.g., manski1990nonparametric,athey2002identification). To the best of our knowledge, these are the first tight characterizations of when ATE and ATT identification is possible in observational studies. For an overview of the proofs, see the technical overview in (ref). While we focus on the average treatment effect and the average treatment effect on the treated, the proposed concept class-based framework is flexible and allows us to characterize when other types of treatment effects are identifiable; see (ref) for an application to the heterogeneous treatment effect.
For (ref) to be useful {given the} other existing conditions (such as unconfoundedness and overlap), it needs to capture interesting examples not captured by existing conditions. In what follows, we revisit several well-studied scenarios in causal inference or their generalizations and, for each scenario, provide identification results based on (ref) -- in the process -- obtaining several novel identification results. Finally, we also give finite sample complexity guarantees for each of these scenarios.
\noindentScenario I: Unconfoundedness and Overlap. At the end of (ref), we mentioned that our framework can capture unconfoundedness and overlap. Identification in this scenario is standard and can also be deduced using (ref); see (ref). Estimation in this setting is also standard imbens2015causal and we discuss how our framework captures it in (ref).
\noindentScenario II: Overlap without Unconfoundedness. Next, we consider observational studies $\euscr{D}$ which satisfy $c$-overlap for some $c\in (0,\nicefrac{1}{2})$ but may not satisfy unconfoundedness. We are going to use our framework to characterize the subset of these studies $\euscr{D}$ for which ATE is identifiable. Since overlap holds with some parameter $c \in (0,\nicefrac{1}{2})$, it restricts the concept class $\mathbbmss{P}$ to be $\mathbbmss{P}_{\rm O}(c)$ where $c < p(x,y) < 1-c$ for any $(x,y)$ and $p \in \mathbbmss{P}_{\rm O}(c)$. This case generalizes several models studied in the causal inference literature tan2006distributional,rosenbaum2002observational,rosenbaum1987sensitivity,kallus2021minimax; see the discussion after (ref). Under this scenario, we can ask: which conditions should the covariate-outcome distributions $\mathbbmss{D}$ satisfy for $\tau$ to be identifiable, i.e., for which observational studies realizable by $(\mathbbmss{P}_{\rm O}(c), \mathbbmss{D})$ is the ATE identifiable? Our result is the following.
The above condition for identification is quite similar to (ref) and is satisfied by setting the outcomes marginal of $\euscr{P} \in \mathbbmss{D}$ to be, e.g., Gaussian, Pareto, or Laplace, and letting the $x$-marginal $\euscr{P}_X$ be unrestricted. This captures important practical models where the outcomes are modeled as a generalized linear model with Gaussian noise rosenbaum2002observational,chernozhukov2024appliedcausalinferencepowered.
\noindentConnections to Prior Work. Since we do not require unconfoundedness in any form, the requirements {on the generalized propensity score class $\mathbbmss{P}_{\rm O}{}(c)$,} in this scenario, are very mild and are already satisfied by most existing frameworks that relax unconfoundedness while retaining overlap. {The restriction on the propensity score class $\mathbbmss{P}_{\rm O}{}(c)$} relaxes tan2006distributional's model and rosenbaum2002observational's odds-ratio model, which are widely used in the literature on sensitivity analysis; see kallus2021minimax,rosenbaum2002observational and the references therein. Both of these models roughly speaking restrict the range of generalized propensity scores $p_0(x,y),p_1(x,y)$ for the same covariate $x$, while already assuming overlap; see (ref) for a detailed discussion. The range of the propensity scores in Tan's and Rosenbaum's models is parameterized by certain constants $\Lambda,\Gamma\geq 1$ respectively, where $\Lambda=\Gamma=1$ corresponds to unconfoundedness, and the extent of violation of unconfoundedness increases with $\Lambda$ and $\Gamma$. The parameter $c$ relates to $\Lambda$ and $\Gamma$ as $\Lambda,\Gamma=O\left(\nicefrac{(1-c)^2}{c^2}\right)>1$. As tan2006distributional,rosenbaum2002observational note, when $\Lambda,\Gamma>1$, without distributional assumptions, $\tau$ can only be identified up to $O(\Lambda)$ and $O(\Gamma)$ factors respectively. Hence, from earlier results, it is not clear which distribution classes $\mathbbmss{D}$ enable the identification of $\tau$; this is answered by (ref).
\noindentFinite-Sample Complexity. Given the above characterization of when the identification of ATE is possible when only overlap holds, one can ask for finite sample estimation. We complement the above result with the following sample complexity guarantee.
The sample complexity depends on the fat-shattering dimension alon1997scale,talagrand2003vc of the class $\mathbbmss{P} = \mathbbmss{P}_{\rm O}{(c)}$ and the covering number $\log N_\varepsilon$ of the class of distributions $\mathbbmss{D}$. Moreover, the mass function $M(\cdot)$ appearing in the sample complexity depends on the class of distributions studied (for illustrations, we refer to (ref)). To the best of our knowledge, this result is the first sample complexity result for such a general setting. For further details, we refer to (ref).
\noindentScenario III: Unconfoundedness without Overlap. We now consider the setting where overlap may fail but unconfoundedness holds. Without additional assumptions, this allows for degenerate cases in which everyone (or no one) receives the treatment, making identification of the ATE impossible. To rule out such extremes, one can assume that some nontrivial subset of covariates satisfies overlap. Concretely, there is a set $S\subseteq\mathbb{R}^d$ with Lebesgue measure $\textrm{\rm vol}(S)\ge c$ such that for each $(x,y)\in S\times\mathbb{R}$, we have $c < p_0(x,y),\,p_1(x,y) < 1 - c$.\footnote{In general, we do not require the lower bound on $\textrm{\rm vol}(S)$ to match the lower bound on the propensity scores. We assume them to be the same in the exposition for notational convenience. Our approach still applies when the two lower bounds differ, and (ref) continue to hold with appropriate modifications.} This is already significantly weaker than the usual $c$-overlap assumption, which demands the previous inequalities pointwise for every $(x,y)\in\mathbb{R}^d\times\mathbb{R}$. We relax it further into the notion of $c$-weak-overlap by removing the upper bound on $p_0(x, y)$ and $p_1(x, y)$ (defined formally in (ref)). Then, we define the class $\mathbbmss{P} = \mathbbmss{P}_{\rm U}(c)$ which captures both unconfoundedness and $c$-weak-overlap; see (ref).
Scenarios with unconfoundedness but without full overlap frequently arise in practice. Classic examples include regression discontinuity designs imbens2008regressionDiscontinuity,lee2010regressionDiscontinuity,angrist2009mostlyHarmless {(see also cook2008waitingforLife)} and observational studies with extreme propensity scores crump2009dealing,li2018overlapWeights,khan2024trimming,kalavasis2024cipw; see further discussion after (ref). As before, we ask which conditions on $\mathbbmss{D}$ enable identification of ATE, i.e., for which observational studies realizable with respect to $(\mathbbmss{P}_{\rm U}{(c)}, \mathbbmss{D})$, can one identify the ATE?
We refer the reader to (ref) for a formal discussion on this condition and result. We would like to stress that the above characterization has a novel conceptual connection with an important field of statistics, called truncated statistics Galton1897,cohen1991truncated,woodroofe1985truncated,cohen1950truncated,laiYing1991truncation. The main task in truncated statistics concerns extrapolation: given a true density $\euscr{D}$ over some domain $X$ and a set $S \subseteq X$, the question is whether the structure of $\euscr{D}$ can be identified from truncated samples, i.e., samples from the conditional density of $\euscr{D}$ on $S$. The condition of the above result requires the pairs $\euscr{P},\euscr{Q}$ to be distinguishable on any set of the form $S\times \mathbb{R}$ (where $S$ has sufficient volume). In other words, any $\euscr{P}$ and $\euscr{Q}$ (with $\euscr{P}_X=\euscr{Q}_X$) whose truncations to the set $S\times \mathbb{R}$ are identical must also have the same untruncated means. Roughly speaking, this condition holds for any family $\mathbbmss{D}$ whose elements $\euscr{P}$ can be extrapolated given samples from their truncations to full-dimensional sets, a problem which is well-studied and provides us with multiple applications Kontonis2019EfficientTS,daskalakis2021statistical,lee2024efficient (see (ref)). We refer to (ref) for a more extensive discussion.
\noindentConnections to Prior Work. This scenario captures two important and practical settings. First, as mentioned before, it captures regression discontinuity (RD) designs where propensity scores violate the overlap assumption for a large fraction of individuals but unconfoundedness holds. These designs were introduced by thistlethwaite1960regressionDiscontinuity, {were independently discovered in many fields cook2008waitingforLife,} and have found applications in various contexts from Education thistlethwaite1960regressionDiscontinuity,angrist1999classSizeRD,klaauw2002regressionDiscontinuityEnrollment,black1999regressionDiscontinuity, to Public Health moscoe2015rdPublicHealth, to Labor Economics lee2010regressionDiscontinuity. Formally, in an RD design, the treatment is a known deterministic function of the covariates: there is some known set $S$ and $T=1$ if and only if $x\in S$.
To the best of our knowledge in RD designs, ATE is only known to be identifiable under strong linearity assumptions on the expected outcomes hahn2001regressionDiscontinuity. Due to that, recent work focuses on identifying certain local treatment effects, which, roughly speaking, measure the effect of the treatment for individuals close to the “decision boundary” imbens2008regressionDiscontinuity. In contrast, (ref) enables us to achieve identification under much weaker restrictions, e.g., it allows the expected outcomes to be any polynomial functions of the covariates (see (ref)).
Apart from RD designs, the above scenario also captures observational studies where certain individuals have extreme propensity scores -- close to 0 or 1. This is a challenging case for the de facto inverse propensity weighted (IPW) estimators of $\tau$, whose error scales with $\sup_x \nicefrac{1}{\left(e(x)(1-e(x))\right)}$ li2018overlapWeights,crump2009dealing,imbens2015causal, and, hence, can be arbitrarily large even when overlap is violated for a single covariate $x$ kalavasis2024cipw. In contrast to such estimators, (ref) enables us to identify ATE even when propensity scores are violated for a large fraction of the covariates.
\noindentFinite-Sample Complexity. As before, we complement the identification result with a finite sample complexity guarantee under a robust version of the above {identifiability} condition.
As in the previous estimation result, the sample complexity depends on the fat-shattering dimension of $\mathbbmss{P} = \mathbbmss{P}_{\rm U}{(c)}$ and the covering number of $\mathbbmss{D}$. An interesting technical observation is that the estimation of (generalized) propensity scores corresponds to a well-known problem in learning theory, that of probabilistic-concept learning of kearns1994pconcept. This connection allows us to get estimation algorithms for classes of bounded fat-shattering dimension.
\noindentScenario IV: Neither Unconfoundedness nor Overlap. {A natural extension of Scenarios II and III arises when both unconfoundedness and overlap fail simultaneously. In this setting, neither the overlap‐based arguments from Scenario II nor the unconfoundedness‐based arguments from Scenario III apply, making identification particularly challenging. Nevertheless, there are some special cases under this scenario where (ref) holds and, hence, ATE is identifiable. We illustrate one such example below, but we do not explore this scenario further because, to our knowledge, the resulting identifiable instances do not directly connect with existing causal inference literature.
}
Our work is related to and connects several lines of work in causal inference and learning theory. We believe that an important contribution of our work is bridging these previously disconnected areas, possibly opening up new paths for applying learning-theoretic insights to causal inference problems. {We discuss the relevant lines of work below.}
We begin with related work from the Causal Inference literature. Here, our work is related to the literature on sensitivity analysis -- which explores the sensitivity of results on deviations from unconfoundedness and is related to results in Scenario II (e.g., cochran1965observationalStudies,rosenbaum1991sensitivity,tan2006distributional), the works on RD designs (e.g., hahn2001regressionDiscontinuity,imbens2008regressionDiscontinuity,cook2008waitingforLife) -- which are a special case of Scenario III -- and to works on handling extreme propensity scores (close to 0 or 1) which arise when overlap is violated and is considered in Scenario III (e.g., crump2009dealing,li2018overlapWeights,khan2024trimming,kalavasis2024cipw).
\noindentExtreme Propensity Scores. Extreme propensity scores (those close to 0 or 1) are a common problem in observational studies. They pose an important challenge since the variance of most standard estimators of, e.g., the average treatment effect, rapidly increases as the propensity scores approach 0 or 1 -- leading to poor estimates. {A} large body of work {designs estimators with lower variance} crump2009dealing,li2018overlapWeights,khan2024trimming,kalavasis2024cipw. While these estimators are widely used, they introduce bias in the estimation of ATE, hence, they do not lead to point identification or consistent estimation, which is the focus of our work. We refer the reader to petersen2012diagnosing for an extensive overview of violations of unconfoundedness and to \citet*{leger2022causal,li2018overlapWeights} for an empirical evaluation of the robustness of existing estimators in the absence of overlap.
\noindentSensitivity Analysis. Sensitivity analysis methods in causal inference assess how unmeasured confounding can bias estimated treatment effects. {The idea dates back to \citet*{cornfield1958smoking}, who studied the causal effect of smoking on developing lung cancer and showed that an unmeasured confounder needed to be nine times more prevalent in smokers than non-smokers to nullify the causal link between smoking and lung cancer -- since this was unlikely, it strengthened the belief that smoking had harmful effects on health. rosenbaum1983sensitivity, subsequently, proposed a sensitivity model for categorical variables. Since then, many works have extended the analysis of Rosenbaum's sensitivity model and introduced alternative parameterizations of the extent of confounding (e.g., rosenbaum2002observational,tan2006distributional,carnegie2016assessing,oster2019unobservable). A notable line of work refines these models to obtain tight intervals in which the ATE lies with the desired confidence level zhao2019sensitivity,dorn2024doublyvalidSharpAnalysis,jin2022sensitivityanalysisfsensitivitymodels,dorn2023sharpSensitivityAnalysis,chernozhukov2023personalizedITE. } {While these works construct valid uncertainty intervals that are valid without distributional assumptions, they do not achieve point identification. Finding the distributional assumptions necessary for point identification is the focus of our work.}
\noindentAdversarial Errors in Propensity Scores. Even with unconfoundedness, propensity scores have to be learned from data {(e.g., mccaffrey2004propensity,athey2019generalized,WESTREICH2010826)}, and errors in the estimation of propensity scores propagate to the estimate of ATE. While under overlap, works from sensitivity analysis (discussed above) provide intervals containing the ATE, these intervals become vacuous even if overlap is violated for a single covariate. {kalavasis2024cipw estimate ATE despite of adversarial errors and outliers, under specific assumptions, by merging outliers with nearby inliers to form “coarse” covariates.} {Our work is orthogonal to theirs in terms of both assumptions and objectives. They obtain interval estimates of treatment effects that are robust to adversarial errors, provided unconfoundedness holds. In contrast, we characterize settings where treatment effects can be point identified without adversarial errors, even when unconfoundedness or overlap fail.}
\noindentRegression Discontinuity Designs. {Regression discontinuity designs were introduced by thistlethwaite1960regressionDiscontinuity in 1960, and have since been independently re-discovered\footnote{{Though there is some debate around this; see cook2008waitingforLife.}} and studied in several disciplines, including Statistics (e.g., rubin1977regressionDiscontinuity,sacks1978regressionDiscontinuity) and Economics (e.g., goldberger1972selection). See cook2008waitingforLife for a detailed overview.} Today, there are two main types of Regression discontinuity (RD) designs: sharp RD designs, where treatment is deterministically assigned based on whether an observed covariate crosses a fixed cutoff,\footnote{We note that typically RD designs consider one-dimensional covariates and where the set $S$ (from (ref)) is an interval of the form $(\alpha,\infty)$ for some constant $\alpha$. In this work, we allow for high-dimensional covariates and any measurable set $S$ satisfying some mild assumptions on its volume.} and fuzzy RD designs, in which treatment assignment is probabilistic near the cutoff (e.g., lee2010regressionDiscontinuity, imbens2008regressionDiscontinuity,hahn2001regressionDiscontinuity). In this work, we considered sharp RD designs, although our framework can also be applied to some fuzzy RD settings and exploring this further is a promising direction for future research. {Recent works in regression discontinuity designs use} local linear regression to estimate the treatment effect at the cutoff (e.g., fan1996local,porter2003estimation,calonico2014robust). These approaches yield only a local average treatment effect and often require linearity or other strong parametric assumptions to “extrapolate” to a global average treatment effect (ATE); see hahn2001regressionDiscontinuity,cattaneo2019practical,chernozhukov2024appliedcausalinferencepowered. In contrast, our work facilitates point identification of the ATE in more general settings, by utilizing recent developments in truncated statistics (see (ref)). Finding interesting classes (apart from the ones mentioned in this work) that can be extrapolated is an interesting open question in truncated statistics, and any progress on it will also enable applications of our framework to these classes.
Next, we discuss relevant work in Learning Theory. Here, we draw on foundational results on probabilistic-concept learning kearns1994pconcept,alon1997scale to get sample complexity bounds. Moreover, to satisfy the extrapolation condition in Scenario III ((ref)), we leverage recent advances in truncated statistics daskalakis2021statistical.
\noindentProbabilistic Concepts. {Most} prior works {in causal inference} assume {access to} an oracle that {estimates the propensity scores $e(x) = \Pr[T{=}1 | X{=}x]$}. The propensity scores $e(\cdot)$ are $[0,1]$-valued, but the feedback provided to the learning algorithm is binary; it is the result of a coin toss where for each $x$, the probability of observing 1 is $e(x)$. Inference in this setting is well-studied in learning theory and corresponds to the problem of learning probabilistic concepts (or $p$-concepts), introduced by kearns1994pconcept. Learnability of a concept class of $p$-concepts is characterized by the finiteness of the fat-shattering dimension of the class (see \citet*{alon1997scale}). To the best of our knowledge, this connection was not reported in the area of causal inference prior to our work.
\noindentTruncated Statistics. Our work and in particular applications which violate overlap are closely related to the area of truncated statistics maddala1986limited,Galton1897,cohen1991truncated,woodroofe1985truncated,cohen1950truncated,laiYing1991truncation. Recently, there has been extensive work on truncated statistics regarding the design of efficient algorithms daskalakis2018efficient,plevrakis2021learning,fotakis2020efficient,lee2025learningpositiveimperfectunlabeled. However, all these works focus on computationally efficient learning of parametric families, while we focus on identification and estimation of treatment effects {and do not impose constraints on computational efficiency}.
An observational study involves units (e.g., patients) with covariates $X\in\mathbb{R}^d$ (e.g., medical history). Each unit receives a binary treatment $T\in\{0,1\}$ (e.g., medication) with a fixed but unknown probability, independent across units, and we observe a treatment-dependent outcome $Y(T)\in\mathbb{R}$ (e.g., symptom severity). The tuple $(X,Y(0),Y(1),T)$ follows an unknown joint distribution $\euscr{D}$, which defines the study. For each $t\in\{0,1\}$, $\euscr{D}_{X,Y(t)}$ denotes the marginal over $X$ and $Y(t)$ and $\euscr{D}_X$ the marginal over $X$. To simplify the exposition, we assume that $\euscr{D}_{X,Y(0)}$ and $\euscr{D}_{X,Y(1)}$ are continuous distributions with densities throughout.
\noindentTreatment Effects. An important goal in causal inference is to identify treatment effects. The Average Treatment Effect $\tau_\euscr{D}$ (ATE) and the Average Treatment Effect on the Treated $\gamma_\euscr{D}$ (ATT) imbens2015causal, hernan2023causal, rosenbaum2002observational are defined as \[ \tau_\euscr{D} \coloneqq \operatornamewithlimits{\mathbb{E}}\nolimits_\euscr{D}\left[Y(1)-Y(0)\right] \qquad\text{and}\qquad \gamma_\euscr{D} \coloneqq \operatornamewithlimits{\mathbb{E}}\nolimits_\euscr{D}\left[Y(1)-Y(0)|T{=}1\right] \,. \] Since instead of observing full samples $(X,Y(0),Y(1),T)$, we only see the censored version $(X,Y(T),T)$, $\tau_\euscr{D}$ and $\gamma_\euscr{D}$ are non-identifiable without further assumptions chernozhukov2024appliedcausalinferencepowered,rosenbaum2002observational.\footnote{In particular, $\operatornamewithlimits{\mathbb{E}}_\euscr{D}\left[Y(1)\right]$ is unobserved and may differ from $\operatornamewithlimits{\mathbb{E}}_\euscr{D}\left[Y(1)\mid T{=}1\right]$ by an arbitrary amount.} This brings us to our main tasks (presented in terms of ATE but also relevant for any treatment effect):
When the distribution $\euscr{D}$ is clear from context, we write $\tau$ and $\euscr{C}$ for $\tau_\euscr{D}$ and $\euscr{C}_\euscr{D}$, respectively. In general, $\tau_\euscr{D}$ cannot be identified from censored samples. This is because there exist $\euscr{D}^{(1)}$ and $\euscr{D}^{(2)}$ with $\abs{\tau_{\euscr{D}^{(1)}}-\tau_{\euscr{D}^{(2)}}}\gg1$ but $\euscr{C}_{\euscr{D}^{(1)}}=\euscr{C}_{\euscr{D}^{(2)}}$. Hence, one needs some assumptions on $\euscr{D}$ to have any hope of solving (ref). The above can be naturally adapted to ATT.
\noindentUnconfoundedness and Overlap. Unconfoundedness and overlap are common sufficient assumptions that enable the identification and estimation of ATE, and have been utilized in a number of important studies; see imbens2015causal,hernan2023causal,rosenbaum2002observational and (ref). The {observational study} $\euscr{D}$ is said to satisfy unconfoundedness if, for each $x\in \mathbb{R}^d$, it holds: $Y(0)~\bot~ T~\mid~ X{=}x$ and $Y(1)~\bot~ T~\mid~ X{=}x$. In other words, the potential outcomes are independent of the treatment $T$ given $X{=}x$. Next, we move to overlap, which ensures that treatment probabilities are bounded away from 0 and 1. The observational study $\euscr{D}$ is said to satisfy overlap if, for each $x\in \mathbb{R}^d$, $0<\Pr_\euscr{D}[T{=}1\mid X{=}x]<1$. Given a constant $c\in(0,\nicefrac{1}{2})$, if $\euscr{D}$ satisfies $c<\Pr_\euscr{D}[T{=}1\mid X{=}x]<1-c$ (for each $x\in \mathbb{R}^d$) then $\euscr{D}$ is said to satisfy the $c$-overlap condition. Although unconfoundedness and overlap suffice to estimate $\tau$ with enough samples, they are not necessary. Unconfoundedness and overlap are often violated (see (ref) for a discussion and examples). To derive necessary and sufficient conditions for identifying $\tau$, we now introduce certain conditional probabilities.
For the reader familiar with causal inference terminology, note that the generalized propensity scores differ from the “usual” propensity score $e(x)\coloneqq\Pr_\euscr{D}[T{=}1\mid X{=}x]$: while $e(\cdot)$ is always identifiable from the data, $p_0(\cdot)$ and $p_1(\cdot)$ in general are not.\footnote{There exist $\euscr{D}^{(1)}$ and $\euscr{D}^{(2)}$ with very different generalized propensity scores but identical censored distributions.} To succinctly state assumptions on generalized propensity scores and $\euscr{D}$, we adopt a statistical learning theory notion of realizability.
Realizability couples the observational study $\euscr{D}$ with the pair of concept classes $\left(\mathbbmss{P},\mathbbmss{D}\right)$.
If $\euscr{D}$ only satisfies $p_0(\cdot),p_1(\cdot)\in \mathbbmss{P}$ (respectively $\euscr{D}_{X,Y(0)},\euscr{D}_{X,Y(1)}\in \mathbbmss{D}$), then $\euscr{D}$ is said to be realizable with respect to $\mathbbmss{P}$ (respectively $\mathbbmss{D}$). We are interested in conditions on the pair $(\mathbbmss{P}, \mathbbmss{D}).$ Finally, we use the following terminology for {certain} elements of {interest in} $\mathbbmss{P} \times \mathbbmss{D}.$
{Note that the tuples $(p,\euscr{P})$ and $(\wh p, \wh \euscr{P})$ in the above definition are not necessarily distinct.}
In this section, we prove {characterizations of identifiability of} ATE and ATT ((ref)) and provide an overview of our estimation algorithms. The characterizations of (ref) consist of two parts. In (ref), we show that (ref) is sufficient for {identifying} ATE {and ATT}. Then, in (ref), we prove that {it is also necessary}. Finally, we provide an overview of our algorithms for estimating ATE in Scenarios I, II, and III in (ref).
(ref) is our main tool to obtain our identification characterizations for ATE and ATT ((ref)). In this section, we explain our technique for identifying these treatment effects from the censored distribution $\euscr{C}_\euscr{D}$ under (ref). We proceed with the following claim.
In this section, we complete the proofs of (ref) by further showing that (ref) is necessary for identification. To do that, we show that if the condition does not hold, then one can find two observational studies with different treatment effects but identical censored distributions. We prove the following claim, which deals with the case of ATT.
By combining (ref), we complete the proofs of (ref).
In this section, we overview our algorithms for estimating ATE in Scenarios I-III. We refer the reader to (ref) for formal statements of results and algorithms.
\noindentStandard Approach to Estimate ATE. We begin with the standard scenario (Scenario I) where unconfoundedness and $c$-overlap hold (and where methods to estimate ATE are already known). Recall that in this scenario, $\tau_\euscr{D}$ can be decomposed as in (ref), which leads to the following finite sample version: given estimates $\widehat{e}(\cdot)$ of the propensity scores $e(\cdot)$, \[ \widehat{\tau} = \sum_i \frac{y_i t_i}{\widehat{e}(x_i)} - \sum_i \frac{y_i(1-t_i)}{1-\widehat{e}(x_i)}\,. \addtocounter{equation}{1}\tag{\theequation}\label{eq:overview:decomposition} \] This decomposition has several useful properties. First, when the outcomes are bounded -- a standard setting (see, e.g., kallus2021minimax) -- each term in the decomposition (i.e., $\nicefrac{y_i t_i}{\widehat{e}(x)}$ and $\nicefrac{y_i (1-t_i)}{(1-\widehat{e}(x))}$) is a bounded random variable. Roughly speaking, under the assumption that $\nicefrac{1}{\widehat{e}(\cdot)}\approx \nicefrac{1}{e(\cdot)}$, this enables one to use the Central Limit Theorem to deduce that, given $n$ samples, $\abs{\tau-\widehat{\tau}}\leq O\left(\nicefrac{1}{\sqrt{n}}\right)$ with high probability. Second, because we assume $c$-overlap, one can show that if $\widehat{e}(\cdot)$ is close to $e(\cdot)$ (e.g., $\int\abs{e(x)-\widehat{e}(x)}\euscr{D}_X(x){\rm d} x\approx 0$), then their inverses $\nicefrac{1}{\widehat{e}(\cdot)}$ and $\nicefrac{1}{e(\cdot)}$ -- which show up in the above decomposition -- are also close to each other. The sample complexity of learning $e(\cdot)$ can be bounded by observing that the problem is equivalent to estimating probabilistic concepts, introduced by kearns1994pconcept and imposing the family of propensity scores to have a finite fat-shattering dimension. While the equivalence to probabilistic concept learning is straightforward, we have not been able to find a reference for it in the learning theory or causal inference literature (which usually assume an estimation oracle {with a small, e.g., $L_2$-error,} as a black-box{; see} kennedy2024agnostic,foster2023orthognalSL,jin2024structureagnosticoptimalitydoublyrobust). For completeness, we present details on obtaining the sample complexity in (ref).
\noindentHurdles in Using (ref) in General Scenarios. None of these ideas work in the more general Scenarios II and III.
Thus, different ideas are needed to estimate $\tau$ in general scenarios.
\noindentOur Approach. We take a completely different approach to estimation, based on (ref). Since (ref) is sufficient for identifying ATE in all scenarios, our approach is quite general -- we present two algorithms -- one for Scenario II and one for Scenario III -- that work for all of the interesting and widely studied special cases of these scenarios discussed in (ref). Having general estimators can be useful since, like unconfoundedness, distributional assumptions and, hence, (ref), are not testable from $\euscr{C}_\euscr{D}$.\footnote{{In particular, given censored samples from $\euscr{C}_\euscr{D}$ and concept classes $(\mathbbmss{P},\mathbbmss{D})$, it is impossible to verify whether $\euscr{D}$ is realizable with respect to $(\mathbbmss{P},\mathbbmss{D})$; then it could be the case that is either realizable with respect to $(\mathbbmss{P},\mathbbmss{D})$ or with respect to alternative classes $(\mathbbmss{P}',\mathbbmss{D}')$ by balancing the products accordingly.}} Thus one cannot pick the estimator based on whether specific assumptions hold.
\noindentEstimator for Scenario II. Our estimator is simple: it first uses the censored samples to find $(p,\euscr{P}),(q,\euscr{Q})\in \mathbbmss{P}\times \mathbbmss{D}$ such that $p\,\euscr{P}$ approximates $p_0\,\euscr{D}_{X,Y(0)}$ and $q\,\euscr{Q}$ approximates $p_1\,\euscr{D}_{X,Y(1)}$ in the following sense: for a sufficiently small $\varepsilon>0$,
Then, it outputs $\widehat{\tau} = \operatornamewithlimits{\mathbb{E}}_{\euscr{P}}[y] - \operatornamewithlimits{\mathbb{E}}_{\euscr{Q}}[y]$. Here, $\operatornamewithlimits{\mathbb{E}}_{\euscr{P}}[y]$ estimates $\operatornamewithlimits{\mathbb{E}}_\euscr{D}[Y(1)] $ and $\operatornamewithlimits{\mathbb{E}}_{\euscr{Q}}[y]$ estimates $\operatornamewithlimits{\mathbb{E}}_\euscr{D}[Y(0)] $.
{The correctness of the estimator follows because under (ref) (a robust version of the condition in (ref)), we show that \[ \sabs{ \operatornamewithlimits{\mathbb{E}}\nolimits_{\euscr{P}}[y] - \operatornamewithlimits{\mathbb{E}}\nolimits_\euscr{D}[Y(1)]}\leq f(\varepsilon) \quad\text{and}\quad \sabs{ \operatornamewithlimits{\mathbb{E}}\nolimits_{\euscr{Q}}[y] - \operatornamewithlimits{\mathbb{E}}\nolimits_\euscr{D}[Y(0)]}\leq f(\varepsilon)\,, \] where $f(\cdot)$ is a function determined by (ref) with the property that $\lim_{z\to 0^+} f(z) = 0$.} {To see that this procedure can be implemented,} note that each product $p_t\,\euscr{D}_{X,Y(t)}$ (for $t\in\{0,1\}$) is identified from the censored data. {To obtain} finite-sample guarantees, we use the following standard assumptions:
{Under these assumptions, we can construct finite covers $C_{P}$ of $\mathbbmss{P}$ and $C_{D}$ of $\mathbbmss{D}$, so that $C_{P}\times C_{D}$ is an $O(\varepsilon)$‐cover of $\mathbbmss{P}\times\mathbbmss{D}$.} {This, in particular, ensures that to find the pairs $(p,\euscr{P})$ and $(q,\euscr{Q})$ in (ref), it suffices to select the elements of the cover $C_P\times C_D$ that are closest to $(p_0,\euscr{D}_{X,Y(0)})$ and $(p_1,\euscr{D}_{X,Y(1)})$ respectively -- as estimated from a suitably large set of samples (see (ref)).} {Hence, the} estimation of $\widehat{\tau}$ {reduces} to finding $(\widehat{p},\widehat{\euscr{P}})$ of $C_P\times C_D$ that is closest to the {empirical distribution induced by the censored samples}. {We note that the size of the cover $C_P\times C_D$ is} exponential in $O(\log(\nicefrac{1}{\varepsilon})) \cdot \mathrm{fat}_{O(\varepsilon)}(\mathbbmss{P}) \cdot \log N_{O(\varepsilon)}(\mathbbmss{D})$ {and this is why we obtain the} sample complexity {claimed in} (ref).
The pseudo-code of the algorithm appears in Algorithm (ref).
\noindentEstimator for Scenario III. In this scenario, unconfoundedness holds, but overlap is very weak: there are sets $S_0,S_1\subseteq\mathbb{R}^d$ with $\textrm{\rm vol}(S_0),\textrm{\rm vol}(S_1)\geq c$ such that for each $(x,y)\in S_t\times \mathbb{R}$, $p_t(x,y)\geq c$ (for each $t\in \ensuremath{\left\{0, 1\right\}}$). If one has membership access to sets $S_0$ and $S_1$ and query access to the propensity scores $e(\cdot)$, then a slight modification of the Scenario II estimator would suffice: one can find $\left(p,\euscr{P}\right)$ such that the product $p\euscr{P}$ approximates the product $p_1\euscr{D}_{X,Y(1)}$ over $S_1$, and output $\operatornamewithlimits{\mathbb{E}}_{\euscr{P}}[y]$ as an estimate for $\operatornamewithlimits{\mathbb{E}}_\euscr{D}[Y(1)]$. (With an analogous algorithm to estimate $\operatornamewithlimits{\mathbb{E}}_\euscr{D}[Y(0)]$.) The correctness of this algorithm {follows} from a robust version of the condition in (ref) (see (ref)) -- which guarantees that if $\euscr{P}_{S_1}$ (the truncation of $\euscr{P}$ to ${S_1}$) is close in TV distance to $(\euscr{D}_{X,Y(1)})_{S_1}$ (the truncation of $\euscr{D}_{X,Y(1)}$ to ${S_1}$), then their means are also close. However, because we neither have access to $S_0,S_1$ nor to $e(\cdot)$, we must estimate both of them from samples and carefully handle the estimation errors.
Next, we describe our estimator for $\operatornamewithlimits{\mathbb{E}}_\euscr{D}[Y(1)]$, the estimator for $\operatornamewithlimits{\mathbb{E}}_\euscr{D}[Y(0)]$ is symmetric, {and} the estimator for ATE follows by subtracting the two. Our estimator proceeds in three phases:
Here, as in the algorithm in Scenario II, one might be tempted to use $\operatornamewithlimits{\mathbb{E}}_{(x,y)\sim\euscr{P}}[y]$ (instead of $\operatornamewithlimits{\mathbb{E}}_{(x,y)\sim\euscr{P}'}[y]$) as an estimate for $\operatornamewithlimits{\mathbb{E}}_\euscr{D}[Y(1)]$. However, this fails because $\euscr{P}$ does not approximate $\euscr{D}_{X,Y(1)}$ well in regions outside of $S_1$ -- where overlap is violated. This is also why Step 2 above is necessary: intuitively, in Step 2, we find a distribution $\euscr{P}'$ which approximates $\nicefrac{p(x,y)\euscr{P}(x,y)}{\widehat{e}(x,y)}$ over the set $\widehat{S}$ -- restricting the optimization to $\widehat{S}$ is important because over $\widehat{S}$, it holds that $\nicefrac{p(x,y)\euscr{P}(x,y)}{\widehat{e}(x,y)}\approx \euscr{P}(x,y)\approx \euscr{D}_{X,Y(1)}$. Now, the correctness follows due to a robust version of the condition in (ref) which, at a high level, ensures that $\mathbbmss{P}'$ extrapolates and is a good estimate of $\euscr{D}_{X,Y(1)}$ over the whole domain and not just $\widehat{S}.$ {We provide the pseudo-code of our algorithm in} Algorithm (ref).
As for the previous algorithm, all the quantities estimated by this algorithm are also identifiable from the censored samples. For finite sample guarantees, we use the same standard assumptions as for the previous scenario. As mentioned above, proving the correctness of this estimator is much more challenging than for the previous estimator and requires a careful analysis; see (ref).
In this section, we present several scenarios, including many novel ones, that satisfy (ref) and, hence, enable the identification of average treatment effect $\tau$. Later, in the upcoming (ref), we show that, under natural assumptions, $\tau$ can also be estimated from finite samples in all of these scenarios.
To gain some intuition about (ref), we begin with the classical scenario where unconfoundedness and overlap both hold. We verify that this scenario satisfies (ref). Before proceeding, we note that in this scenario $\tau$ is already known to be identifiable and, under mild additional assumptions, one also has finite sample estimators for it imbens2015causal,chernozhukov2024appliedcausalinferencepowered. To verify that (ref) is satisfied, we first need to put this scenario in the context of (ref) by identifying the structure of the concept classes $\mathbbmss{P}$ and $\mathbbmss{D}$. As mentioned in (ref), an observational study $\euscr{D}$ satisfies unconfoundedness and overlap if and only if it is realizable with respect to $\mathbbmss{P}_{\rm OU}(c)$ (see (ref)).\footnote{To see that if $\euscr{D}$ satisfies unconfoundedness and $c$-overlap it belongs to $\mathbbmss{P}_{\rm OU}(c)$ consider that $p_t(x,y_1) = \Pr[T{=}t \mid X{=}x, Y(t){=}y_1] = \Pr[T{=}t \mid X{=}x] = \Pr[T{=}t \mid X{=}x, Y(t){=}y_2] = p_t(x,y_2)$ whenever $T\bot Y(t) \mid X$ for $t\in\{0,1\}$. {Next, to see that if} $\euscr{D}$ belongs to $\mathbbmss{P}_{\rm OU}(c)$, {then} it satisfies unconfoundedness and $c$-overlap consider that for $t\in\{0,1\}$ $p_t(x,y) = \Pr[T{=}t\mid X{=}x]$ by the first property and, so $c$-overlap holds and $\Pr\left[T{=}t, Y(t)\in S\mid X{=}x\right] =\Pr[T{=}t\mid X{=}x]\cdot \int_{S}\euscr{D}_{Y(t)\mid X{=}x}(y) {\rm d} y =\Pr[T{=}t\mid X{=}x]\cdot \euscr{D}_{Y(t)\mid X{=}x}(S)$, i.e., $T\bot Y(1), Y(0) \mid X$. } Since unconfoundedness and overlap place no restrictions on the concept class $\mathbbmss{D}$, we let $\mathbbmss{D}$ be the set of all distributions over $\mathbb{R}^d\times \mathbb{R}$, which we denote by $\mathbbmss{D}_{\rm all}$. Now, we are ready to verify that unconfoundedness and overlap satisfy (ref).
Hence, if an observational study $\euscr{D}$ is realizable with respect to $\left(\mathbbmss{P}_{\rm OU}, \mathbbmss{D}_{\rm all}\right)$, then $\tau_\euscr{D}$ can be identified. The proof of (ref) appears in (ref).
Next, we consider the scenario where overlap holds but unconfoundedness may not. Concretely, in this scenario, the generalized propensity scores belong to the following concept class.
This is a very weak requirement on the generalized propensity scores. Since it makes no assumptions related to unconfoundedness, it already captures the many existing models for relaxing unconfoundedness in the literature tan2006distributional,rosenbaum2002observational,rosenbaum1987sensitivity,kallus2021minimax.
Under the scenario we consider we can ensure that $\Gamma=\Lambda=O(\nicefrac{(1-c)^2}{c^2})>1$. However, as noted by rosenbaum2002observational,tan2006distributional, if $\Gamma, \Lambda > 1$, then without distributional assumptions, $\tau$ cannot be identified up to factors better than $O(\Gamma)$ and $O(\Lambda)$ respectively. Hence, based on earlier results, it is not clear when $\tau$ can be identified. Our main result in this section is a characterization of the concept class $\mathbbmss{D}$ that enables identifiability in the above scenario -- where overlap holds but unconfoundedness may not. Its proof appears in (ref).
This condition is similar to (ref). Each tuple $(p,\euscr{P})$ corresponds to some propensity score $p_t(\cdot)$ and distribution $\euscr{D}_{X,Y(t)}$ for some $t\in \ensuremath{\left\{0, 1\right\}}$. The above condition ensures that any two tuples that lead to different guesses for $\tau$, are distinguishable from the available samples. This is because of two reasons. First, as before, the marginal $\euscr{D}_X$ can be identified from data and, hence, all distributions $\euscr{P}$ with $\euscr{P}_X\neq \euscr{D}_X$ can be eliminated. Now, all remaining distributions have the same marginal over $X$. Since any two propensity scores $p,q$, their ratio $\nicefrac{p(x,y)}{q(x,y)}\in \left(\nicefrac{c}{(1-c)}, \nicefrac{(1-c)}{c}\right)$ (due to $c$-overlap). The above condition ensures that $p(x,y)\cdot \euscr{P}(x,y)\neq q(x,y)\cdot \euscr{Q}(x,y)$ for some $x,y$ {enabling} us to distinguish $\left(p,\euscr{P}\right)$ and $\left(q,\euscr{Q}\right)$ as in (ref).
The above result is valuable because a number of common distribution families, including the Gaussian distributions, Pareto distributions, and Laplace distributions, can be shown to satisfy (ref) (for any $c>0$). Hence, the above characterization shows that overlap alone already enables identifiability for many distribution families. A specific, interesting, and practically relevant example captured by this condition is generalized linear models (GLMs): in this setting, for each $t\in \ensuremath{\left\{0, 1\right\}}$, $Y(t)=\mu_t(x)+\xi_t$ for some function $\mu_t(\cdot)$ and noise $\xi_t\sim \euscr{N}(0,1)$.
{Next, we consider the scenario where unconfoundedness holds but overlap may not. Without further assumptions, this includes the extreme cases where either no one receives the treatment or everyone receives the treatment, i.e., for any $t\in \ensuremath{\left\{0, 1\right\}}$, \[ \forall_{x\in \mathbb{R}^d}\,,~~ \forall_{y\in \mathbb{R}}\,,\quad p_t(x,y) = 0 \qquad\text{or}\qquad \forall_{x\in \mathbb{R}^d}\,,~~ \forall_{y\in \mathbb{R}}\,,\quad p_t(x,y) = 1\,. \] Clearly, in these cases, identifying ATE is impossible. To avoid these extreme cases, we assume that at least some non-trivial set of covariates satisfies overlap. A natural way to satisfy this is to require that there is some set $S$ of covariates with $\textrm{\rm vol}(S)\geq \Omega(1)$ such that for each $(x,y)\in S\times \mathbb{R}$ overlap holds, i.e., $c < p_0(x,y), p_1(x,y) < 1-c$. This requirement is already significantly weaker than $c$-overlap which requires $c<p_0(x,y),p_1(x,y)<1-c$ to hold pointwise for each $(x,y)\in \mathbb{R}^d\times \mathbb{R}$. We make an even weaker requirement, which we call $c$-weak-overlap that removes the lower bound on $p_0(\cdot)$ and $p_1(\cdot)$:
The following class encodes the resulting scenario.
Two remarks are in order. First, to simplify the notation, we use the same constant $c$ to denote the lower bound on $\textrm{\rm vol}(S)$ and the values of $p(\cdot)$. One can extend our results to use different constants $c_S,c_p>0$. Second, for the above guarantee to be meaningful, the set $S$ must be a subset of $\operatorname{supp}(\euscr{D}_X)$; otherwise, the propensity scores could be 0 for each $x\in \operatorname{supp}(\euscr{D}_X)$ or 1 for each $x\in \operatorname{supp}(\euscr{D}_X)$, returning us to the extreme cases described above where ATE is clearly not identifiable. To ensure that this is always the case, in this section, we make the simplifying assumption $\operatorname{supp}(\euscr{D}_X)=\mathbb{R}^d$ and, hence, also assume for each $\mathbbmss{P}\in \mathbbmss{D}$, $\operatorname{supp}(\euscr{P}_X)=\mathbb{R}^d$ (otherwise, we can remove $\euscr{P}$ from $\mathbbmss{D}$).}
The identification and estimation methods we develop in this scenario are relevant to many well-studied topics in causal inference.
Next, we present the class of conditional outcome distributions that, together with the propensity scores in (ref), characterize the identifiability of $\tau$.
As for the other conditions we discussed so far, this condition allows us to distinguish any pair of tuples $(p,\euscr{P})$ and $(q,\euscr{Q})$ that lead to a different prediction for $\tau$. The requirement for the marginal of $\euscr{P}$ and $\euscr{Q}$ over $X$ to match is the same as in (ref), so let us consider the second requirement. It requires the pairs $\euscr{P},\euscr{Q}$ to be distinguishable on any set of the form $S\times \mathbb{R}$ where $S$ is a full-dimensional set. In other words, any $\euscr{P}$ and $\euscr{Q}$ (with $\euscr{P}_X=\euscr{Q}_X$) whose truncations to the set $S\times \mathbb{R}$ are identical must also have the same untruncated means. Roughly speaking, this condition holds for any family $\mathbbmss{D}$ whose elements $\euscr{P}$ can be extrapolated given samples from their truncations to full-dimensional sets. While this might seem like a strong requirement at first, it is satisfied by many families of parametric densities: For instance, using Taylor's theorem, one can show that it holds for distributions of the form $\propto e^{f(x,y)}$ for any polynomial $f(x,y)$ (see (ref)). This already includes several exponential families, including the Gaussian family.
Now, we are ready to state the main result of this section: a characterization of when $\tau$ is identifiable under unconfoundedness when overlap may not hold. The proof of this result appears in (ref).
{The requirement that $\operatorname{supp}(\euscr{P}_X)=\mathbb{R}^d$ for each $\euscr{P}\in \mathbbmss{D}$, in particular, ensures that $\operatorname{supp}(\euscr{D}_X)=\mathbb{R}^d$, which is necessary to ensure that the definition of $c$-weak-overlap is meaningful. Recall that if it does not hold and one can select a set $S$ with $\textrm{\rm vol}(S)>c$ disjoint from $\operatorname{supp}(\euscr{D}_X)$, then one can satisfy $c$-weak-overlap even in cases where no one receives the treatment or everyone receives the treatment, where ATE is clearly non-identifiable. That said, we note that the above result can be generalized to require $\operatorname{supp}(\euscr{P}_X)=K$ for any full-dimensional set $K$.}
Our next result presents several examples of families of distributions that satisfy (ref).
These distribution families capture a broad range of parametric assumptions commonly used in causal inference. The polynomial log-density framework includes widely applied exponential families, such as Gaussian outcome models with arbitrary distributions over covariates $X$. The second family allows for polynomial conditional expectations, covering popular linear and polynomial regressions chernozhukov2024appliedcausalinferencepowered. Both families leave the marginal distribution of $X$ unrestricted, allowing for rich covariate distributions while ensuring identifiability under the present scenario. The proof of (ref) appears in (ref).
\noindentRegression Discontinuity Design. As a concrete application of Scenario III, we study regression discontinuity (RD) designs hahn2001regressionDiscontinuity,thistlethwaite1960regressionDiscontinuity,imbens2008regressionDiscontinuity,lee2010regressionDiscontinuity,rubin1977regressionDiscontinuity,sacks1978regressionDiscontinuity which were introduced by and studied in several disciplines rubin1977regressionDiscontinuity,sacks1978regressionDiscontinuity,goldberger1972selection (see cook2008waitingforLife for an overview), and have found applications in various contexts from Education thistlethwaite1960regressionDiscontinuity,angrist1999classSizeRD,klaauw2002regressionDiscontinuityEnrollment,black1999regressionDiscontinuity, to Public Health moscoe2015rdPublicHealth, to Labor Economics lee2010regressionDiscontinuity. In an RD design, the treatment assignment is a known deterministic function of the covariates. \defRD* Since the treatment assignment is only a function of the covariates and not the outcomes, unconfoundedness is immediately satisfied. However, overlap may fail since any covariate $x$ outside of the treatment set $S$ does not receive treatment, while individuals within $S$ always receive treatment. To avoid degenerate cases in which the entire population is treated (or untreated), we require the treatment set $S$ and its complement to have a positive volume. Under these conditions, RD designs become a special case of Scenario III, where the generalized propensity scores lie in $\mathbbmss{P}_{\rm U}(c)$. The following corollary of (ref) shows that ATE can be identified in any RD design.
To the best of our knowledge, all results for identifying ATE in RD assume linear outcome regressions, i.e., that $\operatornamewithlimits{\mathbb{E}}[Y(t)\mid X{=}x]$ is a linear function of $x$ (for each $t\in \ensuremath{\left\{0, 1\right\}}$). (ref) substantially broadens these assumptions and is applicable in very general and practical models where $\operatornamewithlimits{\mathbb{E}}[Y(t)\mid X{=}x]$ are polynomial functions of $x$ and the distribution of covariates is arbitrary; see (ref) for a proof.
In this section, we study the estimation of the average treatment effect $\tau$ from finite samples in the scenarios presented in (ref). We show that, under mild additional assumptions, the estimation of the ATE is possible in all of them.
We begin with estimating ATE under the classical assumptions of unconfoundedness and $c$-overlap. As mentioned before, given access to propensity scores, estimators for ATE are already known in this scenario imbens2015causal,chernozhukov2024appliedcausalinferencepowered. For completeness, we prove ATE's end-to-end estimability (the proof appears in (ref)).
The assumption on the range of the outcomes being bounded is standard in the causal inference literature when one aims to get sample complexity (e.g., kallus2021minimax), and the bound on the fat-shattering dimension is expected because of the reduction to probabilistic concepts from statistical learning theory alon1997scale.
Next, we estimate ATE in Scenario II where $c$-overlap holds but unconfoundedness does not. In (ref), we characterized the identifiability of ATE under this scenario: ATE was identifiable if and only if the class $\mathbbmss{D}$ satisfied (ref). To estimate ATE, we need the following quantitative version of (ref).
(ref) and (ref) differ in two key aspects: First, (ref) scales the bounds on the ratio between any pair of distributions $\euscr{P}, \euscr{Q} \in \mathbbmss{D}$ by a factor of $2$. This factor is arbitrary and can be replaced by any constant strictly greater than 1. The crucial aspect of (ref) is that the bound on the ratio of densities holds not just at a single point but on a set $S$ with non-trivial probability mass. Intuitively, this ensures that differences between distributions can be detected using finite samples, allowing us to correctly identify the underlying distribution. In the next result, we formalize this intuition, showing that the sample complexity naturally depends on the mass function $M(\cdot)$.
The proof of (ref) appears in (ref).
\noindentProof Sketch of (ref). The argument proceeds in two steps.
\noindentConstruction of estimator $\widehat{\tau}$. At a high level, the assumptions on $\mathbbmss{P}$ and $\mathbbmss{D}$ enable one to create a cover of $\mathbbmss{P}\times \mathbbmss{D}$ {with respect to the $L_1$-norm}. ({Where, we define the $L_1$-norm between $\alpha(x,y)$ and $\beta(x,y)$ as $\|\alpha - \beta\|_1 \coloneqq \iint \bigl|\alpha(x,y) - \beta(x,y)\bigr|\;{\rm d} x{\rm d} y.$)} This, in turn, is sufficient to get $(p,\euscr{P})$ and $(q,\euscr{Q})$ such that the products $p\euscr{P}$ and $q\euscr{Q}$ are good approximations for the products $p_1\euscr{D}_{X,Y(1)}$ and $p_0\euscr{D}_{X,Y(0)}$ respectively. Concretely, they satisfy the following guarantee \[ \ensuremath{\left\lVert p_1\euscr{D}_{X,Y(1)} - p\euscr{P} \right\rVert}_1 < M(O(\varepsilon)) \qquad\text{ and }\qquad \ensuremath{\left\lVert p_0\euscr{D}_{X,Y(0)} - q\euscr{Q} \right\rVert}_1 < M(O(\varepsilon)) \,, \addtocounter{equation}{1}\tag{\theequation}\label{sec:est:overlap:guarantee} \] where we define the $L_1$-norm between $\alpha(x,y)$ and $\beta(x,y)$ as $\|\alpha - \beta\|_1 \coloneqq \iint \bigl|\alpha(x,y) - \beta(x,y)\bigr|\;{\rm d} x{\rm d} y.$ We present the details of constructing the cover and finding the tuples $(p,\euscr{P})$ and $(q,\euscr{Q})$ using finite samples in (ref). We then define our estimator as \[ \widehat{\tau} \;\;=\;\; \abs{\operatornamewithlimits{\mathbb{E}}\nolimits_{(x,y)\sim\euscr{P}}[y] - \operatornamewithlimits{\mathbb{E}}\nolimits_{(x,y)\sim\euscr{Q}}[y]}\,. \]
\noindentProof of accuracy of $\widehat{\tau}$. Due to (ref) and the fact that all elements of $\mathbbmss{P}$ satisfy overlap, if $\operatornamewithlimits{\mathbb{E}}_{\euscr{D}_{X,Y(1)}}[y]$ is $\varepsilon$-far from $\operatornamewithlimits{\mathbb{E}}_{\euscr{P}}[y]$, then $\euscr{D}_{X,Y(1)}(x,y)/\euscr{P}(x,y)$ must be very large or very small (concretely, outside $\left(\nicefrac{c}{2(1-c)}, \nicefrac{2(1-c)}{c}\right)$) for each $(x,y)\in S$ where $S$ is a set with measure at least $M(\varepsilon)$ under $\euscr{P}$ and $\euscr{Q}$. Because $p_1,p\in \mathbbmss{P}_{\rm O}$, their ratios are bounded and always lie in $\left(\nicefrac{c}{(1-c)}, \nicefrac{(1-c)}{c}\right)$.
Our proof relies on the following observation: intuitively, (ref) forces any two distributions, say $\euscr{P}$ and $\euscr{D}_{X,Y(1)}$ in $\mathbbmss{D}$, with a large difference in mean-outcomes to have a large (multiplicative) difference in their densities over a set of measure at least $M(\varepsilon)$. Concretely, if $\sabs{\operatornamewithlimits{\mathbb{E}}_{\euscr{D}_{X,Y(1)}}[y]-\operatornamewithlimits{\mathbb{E}}_{\euscr{P}}[y]}\geq O(\varepsilon)$, then $\euscr{D}_{X,Y(1)}(x,y)/\euscr{P}(x,y)\not\in \left(\nicefrac{c}{2(1-c)}, \nicefrac{2(1-c)}{c}\right)$ on at least a set $S$ of mass $M(\varepsilon)$ under both $\euscr{P}$ and $\euscr{D}_{X,Y(1)}$. Further, the ratios of propensity scores $p(\cdot)$ and $p_1(\cdot)$ are bounded between $\left(\nicefrac{c}{(1-c)}, \nicefrac{(1-c)}{c}\right)$. The combination of these facts allows one to show that if $\sabs{\operatornamewithlimits{\mathbb{E}}_{\euscr{D}_{X, Y(1)}}[y]-\operatornamewithlimits{\mathbb{E}}_{\euscr{P}}[y]}\geq O(\varepsilon)$, then \[ \ensuremath{\left\lVert p_1\euscr{D}_{X,Y(1)} - p\euscr{P} \right\rVert} > M(O(\varepsilon))\,, \] which contradicts the guarantee in (ref). Thus, due to the contradiction, one can conclude that $\sabs{\operatornamewithlimits{\mathbb{E}}_{\euscr{D}_{X,Y(1)}}[y]-\operatornamewithlimits{\mathbb{E}}_{\euscr{P}}[y]}\leq O(\varepsilon)$. An analogous proof shows $\sabs{\operatornamewithlimits{\mathbb{E}}_{\euscr{D}_{X,Y(0)}}[y]-\operatornamewithlimits{\mathbb{E}}_{\euscr{Q}}[y]}\leq O(\varepsilon)$. Together, these are sufficient to conclude the proof.
Next, we study estimation under Scenario III, where unconfoundedness is guaranteed but overlap is not. Recall that this scenario is captured by the following class of propensity scores. \[ \mathbbmss{P}_{\rm U}(c) \coloneqq \left\{ p\colon \mathbb{R}^d\times \mathbb{R}\to [0,1]\;\middle|\;
\right\}\,. \] In (ref), we showed that, in this case, the identifiability of $\tau$ is characterized by (ref). In this section, we show that $\tau$ can be estimated from finite samples under the following quantitative version of (ref).
To gain some intuition, fix a set $S$. Now, the above condition holds if whenever the truncated distributions $\euscr{P}(S\times \mathbb{R})$ and $\euscr{Q}(S\times \mathbb{R})$ are close, then the means of the untruncated distributions $\euscr{P},\euscr{Q}$ are also close. (ref) requires this for any set $S$ of sufficient volume. At a high level, this holds whenever the truncated distribution can be “approximately extended” to the whole domain to recover the original distribution -- i.e., whenever extrapolation is possible. At the end of this section, in (ref), we show that -- under some mild assumptions -- a rich class of distributions can be extrapolated and, hence, satisfy (ref). Now, we are ready to state our estimation result.
We expect that the $\nicefrac{1}{\varepsilon^4}$ dependence on the sample complexity can be improved using boosting, but we did not try to optimize it. We refer the reader to (ref) for a sketch of the proof of (ref) and to (ref) for a formal proof. {Before showing that (ref) is satisfied by interesting distribution families, we pause to note that apart from the constraints on concept classes $\mathbbmss{P}$ and $\mathbbmss{D}$, we require the additional requirement that $c < \Pr_\euscr{D}[T{=}1] < 1-c$. First, observe that this is a mild requirement and is significantly weaker than overlap, which requires $c < \Pr_\euscr{D}[T{=}1|X={x}] < 1-c$ for each $x$. (To see why it is a mild requirement, observe that it allows the propensity scores $e(x)=$ to be 0 or 1 for all covariates as in regression discontinuity designs.) Second, our constraints on $\mathbbmss{P}$ and $\mathbbmss{D}$ already ensure that $\Pr[T{=}1]\in (0,1)$, which was sufficient for identification; however, they allow $\Pr[T{=}1]$ to approach 0 or 1, which makes estimation impossible. We require this constraint to avoid these extreme cases.}
Next, we show that a rich family of distributions satisfies (ref) (also see (ref)).
In particular, when $M,k=O(1)$ and $c=\Omega(1)$, the constant is $C=O(1)$. This result is a corollary of Lemma 4.5 in daskalakis2021statistical and relies on the anti-concentration properties of polynomials carbery2001distributional. Moreover, the conclusion can be generalized to the case where $K$ is any convex subset of $\mathbb{R}^d\times\mathbb{R}$. Specifically, if $[a,b]^{d+1}\subseteq K\subseteq [c,d]^{d+1}$, then the constant $C$ will scale linearly with the diameter of $K$ and with a function of the aspect ratio $\frac{d-c}{b-a}$. The proof of (ref) appears in (ref).
\noindentEstimation under Regression Discontinuity Design. Next, we consider the estimation of $\tau$ with regression discontinuity (RD) designs. As mentioned before, RD designs are a special case of Scenario III and, hence, we get the following result as an immediate corollary of (ref).
This work extends the identification and estimation regimes for treatment effects beyond the standard assumptions of unconfoundedness and overlap, which are often violated in observational studies. Inspired by classical learning theory, we introduce a new condition that is both sufficient and necessary for the identification of ATE, even in scenarios where treatment assignment is deterministic or hidden biases exist. This condition allows us to build a framework that unifies and extends prior identification results by characterizing the distributional assumptions required for identifying ATE without the standard assumptions of unconfoundedness and overlap tan2006distributional,rosenbaum2002observational,thistlethwaite1960regressionDiscontinuity. Beyond immediate theoretical contributions, our results establish a deeper connection between learning theory and causal inference, opening new directions for analyzing treatment effects in observational studies with complex treatment mechanisms.
This project is in part supported by NSF Award CCF-2342642. Alkis Kalavasis was supported by the Institute for Foundations of Data Science at Yale. We thank the anonymous COLT reviewers for helpful suggestions on presentation and for suggesting to include a fourth scenario.
\printbibliography