Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
112,216 characters · 27 sections · 71 citation commands
Quantifying the Internal Validity of Weighted Estimands
\singlespacing
\newsavebox{\tablebox} \newlength{\tableboxwidth}
\thispagestyle{empty}
\onehalfspacing
\setcounter{section}{1}
Estimating average treatment effects is an important goal in many areas of empirical research. Applied researchers usually believe that treatment effects are heterogeneous, which means that they vary across units. Yet, many researchers also favor using well-established estimation methods that were not originally designed with treatment effect heterogeneity in mind. These methods may be chosen because of their computational simplicity, comparability across studies, effectiveness at incorporating high-dimensional covariates, and other reasons. In turn, these methods often lead to estimands that can be represented as weighted averages of the underlying treatment effects of interest.
For example, consider a scenario where unconfoundedness holds given covariates $X$. Let treatment $D$ be binary, $(Y(1),Y(0))$ be potential outcomes, and let $\tau_0(X) = \mathbb{E}[Y(1) - Y(0)\mid X]$ be the conditional average treatment effect, or CATE, for covariate value $X$. Following Angrist1998, if we additionally assume that $\mathbb{P}(D=1 \mid X)$ is linear in $X$, the population regression of $Y$ on a constant, treatment $D$, and covariates $X$ yields a coefficient on $D$ that can be written as
a weighted average of CATEs with nonnegative weights that integrate to 1. This parameter will be equal to the average treatment effect, $\mathbb{E}[Y(1) - Y(0)]$, if and only if $\operatorname{var}(D\mid X)$ and $\tau_0(X)$ are uncorrelated.
In this paper we are concerned with a general class of weighted estimands that can be expressed as follows:
where $W_0 \in \{0,1\}$ is an indicator for a subpopulation, $\tau_0(X) = \mathbb{E}[Y(1) - Y(0)\mid W_0 = 1, X]$ are the CATEs given covariates $X$ in the same subpopulation $W_0$, $w_0(X) = \mathbb{P}(W_0 = 1\mid X)$ is the probability of being in this subpopulation given $X$, and $a(X)$ is an identified weight function. The regression estimand above belongs to this class, which can be seen by letting $W_0 = 1$ with probability 1, and letting the weight function $a(X)$ be the conditional variance of treatment given covariates. Under some assumptions, this class also includes the two-stage least squares (2SLS) and two-way fixed effects (TWFE) estimands in instrumental variables and difference-in-differences settings, as well as many other parameters. Here, the leading cases of $W_0$ are compliers in the case of 2SLS and treated units in the case of TWFE\@.
There are two main questions that this paper seeks to answer. The first is whether, and under what conditions, the estimand in (ref) corresponds to an average treatment effect of the form $\mathbb{E}[Y(1)-Y(0)\mid W^*=1]$, where $W^*\in \{0,1\}$ is an indicator for a (possibly latent) subpopulation of $W_0$. An affirmative answer to this question would endow a specific estimand with some degree of validity as a causal parameter, given that it would then measure the average effect of treatment for a subset of all units.
The second and primary aim of this paper is to quantify the degree of validity of $\mu(a,\tau_0)$ as a causal parameter. To do this, we characterize the size, and the size relative to $W_0$, of subpopulations $W^*$ associated with the estimand in (ref). More plainly, we ask how large $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1\mid W_0=1)$ can be in the representation $\mu(a,\tau_0) = \mathbb{E}[Y(1) - Y(0)\mid W^* = 1]$. If these probabilities can be large, the estimand corresponds to the average treatment effect for a (relatively) large subpopulation, and when they are small, it corresponds to the average effect for a (relatively) small number of units. If $\mathbb{E}[Y(1) - Y(0)\mid W_0 = 1]$ is the target parameter, we interpret a large value of $\mathbb{P}(W^*=1\mid W_0=1)$ as evidence of a high degree of internal validity of $\mu(a,\tau_0)$ with respect to the target.\footnote{We use the term “internal validity” because it is often associated with the question of whether the probability limit of an estimator is equal to the parameter of interest. If the researcher reports estimates of $\mu(a,\tau_0)$ when $\mathbb{E}[Y(1) - Y(0)\mid W_0 = 1]$ is their ultimate target parameter, then the estimator is internally valid for the target when $\mathbb{P}(W^*=1\mid W_0=1) = 1$. If $\mathbb{P}(W^*=1\mid W_0=1)<1$, then $\mu(a,\tau_0)$ will be more informative about the target as the value of $\mathbb{P}(W^*=1\mid W_0=1)$ increases.} If $\mathbb{P}(W^*=1)$, the corresponding marginal probability, is large, we say that $\mu(a,\tau_0)$ is highly representative of the underlying population.
The answer to our questions about subpopulation existence and size depends on the information we have about the CATE function, $\tau_0$. Specifically, in one case, we may want to know whether $\mu(a,\tau_0)$ can be written as $\mathbb{E}[Y(1) - Y(0)\mid W^* = 1]$ for any choice of $\tau_0$, or without any restrictions on this function. If this is the case, then we know that the interpretation of $\mu(a,\tau_0)$ as a causal parameter is robust to heterogeneous treatment effects of any form, including the most adversarial CATE functions. We can also answer the second question about the maximum values of $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1\mid W_0=1)$ without needing to estimate or know the structure of the CATEs. In a second case, we may want to know how representative $\mu(a,\tau_0)$ is given knowledge of the CATE function. While the resulting maximum values of $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1\mid W_0=1)$ are less useful as measures of robustness than in the first case---after all, if the researcher knows or estimates the entire CATE function, they can as well report any average of $\tau_0(X)$ that may be relevant---we consider this problem to be of independent theoretical interest. Additionally, if the researcher estimates and compares the maximum values of $\mathbb{P}(W^*=1)$ or $\mathbb{P}(W^*=1\mid W_0=1)$ in both cases, they can evaluate the importance of treatment effect heterogeneity for the interpretation of $\mu(a,\tau_0)$ in a given application.
In the first case, when the CATE function is unrestricted, we formally show that $\mu(a,\tau_0)$ can be written as the average treatment effect for a subpopulation of $W_0$ if and only if $a(X) \geq 0$ with probability 1 given $W_0 = 1$. The contrapositive of this statement is that the incidence of “negative weights,” that is, $\mathbb{P}(a(X) < 0 \mid W_0 = 1) > 0$, implies that $\mu(a,\tau_0)$ cannot be represented as an average treatment effect for some subpopulation uniformly in $\tau_0$. This result provides a novel justification for the common requirement that the weights underlying a suitable estimand must be nonnegative. In a related contribution, BlandholBonneyMogstadTorgovitsky2022 show that the lack of negative weights and level dependence is a sufficient and necessary condition for the weighted estimand to be “weakly causal,” that is, to guarantee that the sign of $\tau_0$ will be preserved whenever it is uniform across all units. We also provide simple expressions for the maxima of $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1\mid W_0=1)$. We show how knowledge of the estimand and these expressions can be used to construct simple bounds on the target parameter. We also propose an analog estimator for our measure of internal validity. We establish the nonstandard limiting distribution of this estimator and describe inference procedures for it in Appendix (ref).
In the second case, when the CATE function is assumed to be known, we show that $\mu(a,\tau_0)$ can be written as an average treatment effect whenever it lies in the convex hull of CATE values, a weaker criterion than having nonnegative weights. The maximum values of $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1\mid W_0=1)$ now depend on $\tau_0$, and can be obtained via linear programming when $X$ is discrete. We show the solution to this linear program also admits a closed-form expression even when the support of $X$ includes discrete, continuous, and mixed components. This expression can be used to derive plug-in estimators.
Besides theoretical interest, we argue that $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1\mid W_0=1)$ represent a valuable diagnostic for empirical research. First, as stated above, our initial results demonstrate that two popular criteria for weighted estimands---that they lack negative weights and that they lie in the convex hull of CATE values---are necessary and sufficient (under different assumptions) for the existence of their causal representation, that is, for $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1\mid W_0=1)$ to be strictly positive. This suggests that researchers invoking these criteria are (indirectly) interested in whether their estimands can be represented as an average treatment effect for some subpopulation. If this is the case, it makes sense to better understand this implicit subpopulation, similar to how it is standard practice in instrumental variables settings to study the subpopulation of compliers. Relatedly, even though our main results concern subpopulation size, we also show that the distribution of covariates in the implicit subpopulation is identified. Thus, practitioners can examine whether this subpopulation has similar characteristics as the entire population, and report the associated sample statistics.
Second, we argue that when $\mathbb{E}[Y(1) - Y(0)\mid W_0 = 1]$ is the parameter of interest, it is reassuring for $\mathbb{P}(W^*=1\mid W_0=1)$ to be large. For one, related claims have been made by other researchers. MT2024 argue that “[t]arget parameters that reflect larger subpopulations of the population of interest are more interesting than those that reflect smaller and more specific subpopulations.” In a setting with multiple instrumental variables, HLM2024 suggest that the largest subpopulation of compliers is generally more interesting than other complier subpopulations. However, we also formalize this claim and show how to construct bounds on $\mathbb{E}[Y(1) - Y(0)\mid W_0 = 1]$ that only depend on $\mathbb{P}(W^*=1\mid W_0=1)$, $\mu(a,\tau_0)$, and a support restriction. The bounds are easy to compute and collapse to a point as $\mathbb{P}(W^*=1\mid W_0=1)$ approaches 1. Indeed, large values of $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1\mid W_0=1)$, our primary measures of interest, guarantee that the weighted estimand is not “too different” from $\mathbb{E}[Y(1) - Y(0)]$ and $\mathbb{E}[Y(1) - Y(0)\mid W_0 = 1]$. This makes our measures a practical diagnostic tool to evaluate the robustness of weighted estimands to heterogeneous treatment effects of any form.
This paper is related to a large literature studying weighted average representations of common estimands, including ordinary least squares (OLS), 2SLS, and TWFE in additive linear models. Some of the contributions to this literature include Angrist1998, AronowSamii2016, Sloczynski2022, Chen2024, Goldsmith-PinkhamHullKolesar2024, and Humphreys2025 for OLS; ImbensAngrist1994, AngristImbens1995, Kolesar2013, Sloczynski2020, and BlandholBonneyMogstadTorgovitsky2022 for 2SLS; and ChaisemartinDHaultfoeuille2020, Goodman-Bacon2021, SunAbraham2021, AI2022, CaetanoCallaway2023, BJS2024, and CallawayGoodman-BaconSantAnna2024 for TWFE\@.
A common view in much of this literature, attributable to ImbensAngrist1994, is that causal interpretability of weighted estimands requires all weights to be nonnegative. BlandholBonneyMogstadTorgovitsky2022 show that the lack of negative weights and level dependence is necessary and sufficient for an estimand to be “weakly causal,” that is, to guarantee sign preservation when all treatment effects have the same sign. In this paper we focus on the related problem of whether a weighted estimand can be written as the average treatment effect over a subpopulation. While the lack of negative weights is essential for subpopulation existence when the CATE function is unrestricted, nonuniform weights, too, have a detrimental effect on subpopulation size. This point is related to the negative view of both negative and nonuniform weights in CallawayGoodman-BaconSantAnna2024.
Some papers focus on weighted averages of heterogeneous treatment effects as legitimate targets in their own right rather than as probability limits of existing estimators. HIR2003 introduce the class of weighted average treatment effects, which are a subclass of the more general class of estimands in (ref). LMZ2018 discuss the connection between weighted average treatment effects and implicit target subpopulations. However, the internal validity and representativeness of weighted estimands have received little attention to date.
One exception is Chaisemartin2012,Chaisemartin2017, who focuses on the interpretation of the instrumental variables (IV) estimand. First, Chaisemartin2012 studies the size of the largest subpopulation whose average treatment effect is equal to that of compliers. While this question is similar to ours, the corresponding subpopulation size is not point identified, unlike in our paper. Second, when the usual monotonicity assumption is violated, Chaisemartin2012,Chaisemartin2017 uses a specific restriction on the CATE function to reinterpret the IV estimand as the average treatment effect for a subset of compliers. In our framework, this result can be seen as an existence result in an intermediate case between the setting where $\tau_0$ is unrestricted and where it is fixed.
Another exception is AronowSamii2016, who focus on whether mean covariate values are similar in the entire sample and in the “effective sample” used by OLS\@. We focus on the size of the implicit subpopulation, which is different and complementary. We also extend the results on mean covariate values to the entire distribution of covariates and to other weighted estimands besides OLS\@.
Yet another exception is MSG2023, who focus on fixed effects estimands and argue that it is problematic if “switchers” are a small subset of the sample. We similarly argue that if a given weighted estimand corresponds to the average treatment effect for a small subpopulation, then it may not be an appropriate target parameter, unless that subpopulation is interesting in its own right.
We organize the paper as follows. In Section (ref), we briefly discuss our motivating example of the OLS estimand. In Section (ref), we develop our theoretical framework and examine the conditions under which the estimand in (ref) has a causal representation as an average treatment effect over a population. In Section (ref), we establish our main results on the size of subpopulations associated with the estimand in (ref). In Section (ref), we revisit our motivating example from Section (ref) and apply our theoretical results to additional examples of weighted estimands. In Section (ref), we briefly discuss estimation and inference for the proposed measures. In Section (ref), we provide an empirical application to the effects of unilateral divorce laws on female suicide, as in SW2006 and Goodman-Bacon2021. In Section (ref), we conclude. The appendices contain our proofs as well as several additional results and derivations.
\endinput
Here we provide further discussion of the OLS estimand. We postpone the discussion of the 2SLS and TWFE estimands to Section (ref). In the initial example, we have a binary treatment $D \in \{0,1\}$, potential outcomes $(Y(1),Y(0))$, covariate vector $X$, and realized outcome $Y = Y(D)$. We make the following assumption.
Following Angrist1998, we can establish that $\beta_\text{OLS}$, the coefficient on $D$ in the linear projection of $Y$ on $(1,D,X)$, satisfies the representation in (ref) under a restriction on the propensity score. The following proposition summarizes Angrist1998's Angrist1998 result.
The linearity assumption can be removed if we instead regress $Y$ on $(1,D,h(X))$ where $h(X)$ is a vector of functions of $X$ such that $p(X)$ is in their linear span. The overlap assumption can also be weakened since it is not required for $\beta_\text{OLS}$ to be defined.
Proposition (ref) implies that we can write $\beta_\text{OLS}$ as a weighted estimand satisfying the representation in (ref) where $a(X) = p(X)(1-p(X))$ and $\tau_0(X) = \mathbb{E}[Y(1) - Y(0)\mid X]$. Here we implicitly set $W_0 = 1$ with probability 1. Thus, the regression coefficient $\beta_\text{OLS}$ is a weighted average of CATEs whose weights are $p(X)(1-p(X))$. Note that $\beta_\text{OLS} = \text{ATE} \coloneqq \mathbb{E}[Y(1) - Y(0)]$ if and only if $a(X)$ and $\tau_0(X)$ are uncorrelated, which is the case, for example, when $p(X)$ or $\tau_0(X)$ is constant.
An alternative representation of this estimand can be obtained by focusing on the subpopulation of treated units, $D=1$. Let $W_0 = D$, $w_0(X) = p(X)$, $\tau_0(X) = \mathbb{E}[Y(1) - Y(0)\mid D=1, X] = \mathbb{E}[Y(1) - Y(0)\mid X]$, which follows from conditional independence, and let $\tilde{a}(X) = 1-p(X)$. Then, we can write
Yet another representation can be obtained when focusing on the subpopulation of untreated units by letting $W_0 = 1-D$. We omit details for brevity.
We will return to this example in Section (ref) after establishing conditions under which weighted estimands have a causal representation (Section (ref)) and identifying the size of subpopulations that are represented by these estimands (Section (ref)).
\endinput
In this section, we consider a general class of weighted estimands. We show necessary and sufficient conditions for an estimand in this class to have a causal representation as an average treatment effect over a subpopulation. We provide these conditions under various assumptions---including no assumptions---on treatment effect heterogeneity.
Recall the earlier setting where we let $D \in \{0,1\}$ denote a treatment variable, and let $(Y(0),Y(1))$ denote the corresponding potential outcomes. Let $X \in \text{supp}(X) \subseteq \mathbb{R}^{d_X}$ denote a $d_X$-vector of covariates, where $\text{supp}(\cdot)$ denotes the support. We suppose that $(Y(1),Y(0),D,X)$ are drawn from a common population distribution $F_{Y(1),Y(0),D,X}$.
Let $W_0 \in \{0,1\}$ be an indicator variable used to denote a subpopulation $\{W_0 = 1\}$ and let $\tau_0(X) = \mathbb{E}[Y(1) - Y(0)\mid W_0 = 1,X]$ denote the conditional average treatment effect given $X$ in that subpopulation. For example, this subpopulation can be the entire population by setting $W_0 = 1$ almost surely, in which case $\tau_0$ denotes the usual CATE function. It can also denote the subpopulation of treated units by setting $W_0 = D$. In the presence of a binary instrument $Z$, the complier subpopulation is defined by setting $W_0 = \mathbbm{1}(D(1) > D(0))$, where $D(1)$ and $D(0)$ are potential treatments. In this case, $\tau_0$ denotes the conditional local average treatment effect.
Note that $\tau_0$ is defined for all values of $X$ such that $w_0(X) = \mathbb{P}(W_0=1\mid X) > 0$.\footnote{While $\tau_0(X)$ is only defined when $w_0(X) > 0$, we set $\tau_0(X)w_0(X) = 0$ when $w_0(X) = 0$.} Throughout this paper, we assume that $\mathbb{P}(W_0 = 1) > 0$, so that this subpopulation has a positive mass, which avoids technical issues associated with conditioning on zero-probability events.
Also recall the weighted estimands of equation (ref):
The estimands we consider have the above representation and satisfy the following regularity condition.
The first two restrictions ensure the existence of the numerator of $\mu(a,\tau_0)$. We rule out $\mathbb{E}[a(X)\mid W_0=1] = 0$ since it implies the estimand does not exist. The estimand in (ref) is unchanged if the sign of $a(X)$ is reversed, so $\mathbb{E}[a(X)\mid W_0=1] > 0$ is a sign normalization.
The weighted estimands of equation (ref) can also be written as a weighted sum when $X$ is discrete, or an integral when $X$ is continuous. In the discrete case, let $\text{supp}(X) = \{x_1,\ldots,x_K\}$, let $p_k \coloneqq \mathbb{P}(X=x_k) > 0$ for $k = 1,\ldots,K$, and assume $W_0 = 1$ almost surely for simplicity. Then,
which are weights that sum to one. The representations in (ref) and (ref) are equivalent since we can obtain $a(x_k)$ (up to scale) as the ratio $\omega_k/p_k$, and $\omega_k$ is defined as a function of $\{(a(x_k),p_k)\}_{k=1}^K$ in equation (ref).
From equation (ref), we can see that $a(x_k)$ being constant ensures $\omega_k = p_k$, or that the estimand is the ATE\@. Moreover, $\frac{a(x_k)}{a(x_{k'})} = \frac{\omega_k}{\omega_{k'}} \big/ \frac{p_k}{p_{k'}}$, which is the ratio of the relative weights of covariate cells $\{X = x_k\}$ and $\{X = x_{k'}\}$ in the estimand $(\omega_k/\omega_{k'})$ and in the population $(p_k/p_{k'})$. The inequality $a(x_k) > a(x_{k'})$ indicates that covariate cell $\{X = x_k\}$ is overweighted by the estimand relative to $\{X = x_{k'}\}$, when compared to their relative weights in the population. Similar algebra can be used to write the estimand as an integral when $X$ is continuously distributed. We instead focus on the representation in equation (ref) since it seamlessly accommodates discrete, continuous, and mixed covariates.
The first question we address is whether an estimand defined by (ref) can be represented as $\mathbb{E}[Y(1) - Y(0)\mid W^* = 1]$, where $W^* \in \{0,1\}$ is binary and $\{W^* = 1\}$ characterizes a subpopulation of $\{W_0 = 1\}$. Formally, $\{W^*= 1\}$ forms a subpopulation if $\{W^*= 1\} \subseteq \{W_0 = 1\}$ or, equivalently, if $W^* \leq W_0$ almost surely.
We impose some structure on this problem by restricting how these subpopulations may be formed. We only consider what we call “regular subpopulations,” defined here.
For convenience, we will abbreviate this as “$W^*$ is a regular subpopulation of $W_0$”. We denote the set of regular subpopulations of $W_0$ as
These subpopulations have positive masses and are subsets of $\{W_0 = 1\}$. The other substantive requirement is that they do not depend on potential outcomes when conditioning on $X$ and the original population $W_0 = 1$. While this may seem restrictive, it allows for rich and natural classes of subpopulations. For example, consider the unconfoundedness restriction of Section (ref) and let $W_0$ be the entire population, i.e., $\mathbb{P}(W_0 =1) = 1$. In this case, regular subpopulations must satisfy $W^* \perp \! \! \! \perp (Y(1),Y(0))\mid X$, or be unconfounded. Regular subpopulations include the population of all treated (or untreated) individuals, i.e., $W^* = D$ (or $W^* = 1-D$), and any subpopulation characterized by a subset of $\text{supp}(X)$. More generally, they include any subpopulation that can be described through a combination of $(D,X,U)$ where $U$ is independent from $(Y(1),Y(0),X)$. For example, a subpopulation characterized by “fraction $a(x)$ of units with covariate $X=x$ for all $x \in \text{supp}(X)$” can be constructed as $W^* = \mathbbm{1}(U \leq a(X))$ where $U \sim \text{Unif}(0,1)$ is independent from $(Y(1),Y(0),X)$.
The conditional independence requirement rules out subpopulations that directly depend on the potential outcomes such as $W^* = \mathbbm{1}(Y(1) \geq Y(0))$, i.e., the subpopulation of those who benefit from treatment. Note that $\mathbb{P}(W^*=1\mid X) = \mathbb{P}(Y(1) \geq Y(0)\mid X)$ and $\mathbb{P}(W^*=1) = \mathbb{P}(Y(1) \geq Y(0))$ are not point-identified under unconfoundedness. Another way to view this requirement is that regular subpopulations are policy relevant in the sense that we could design a policy that targets a regular subpopulation. Indeed, a policy maker may observe $X$ and can use $U$ to randomly target a fraction of units with specific values of $X$, but cannot observe potential outcomes.
These particular subpopulations enjoy a number of useful properties. Two of them are characterized in the following proposition.
The first part of this proposition shows that average effects within the original population $W_0$ and regular subpopulation $W^*$ are the same when conditioning on $X$. For example, this holds under unconfoundedness for the subpopulation of treated units, $W^* = D$. The second property allows us to write the average treatment effect for $W^* = 1$ using the same functional $\mu(\cdot,\cdot)$ that was used to characterize the class of estimands we analyze. This property will be used when studying the mapping between weighted estimands and average treatment effects for regular subpopulations of $W_0$.
We now consider necessary and sufficient conditions for the weighted estimand $\mu(a,\tau_0)$ to be written as the average treatment effect within a regular subpopulation of $W_0$. As we will show, these conditions depend on what is assumed about the function $\tau_0 = \mathbb{E}[Y(1) - Y(0)\mid X = \cdot,W_0=1]$.
For example, if $\tau_0$ is constant in $X$, then any weighted estimand satisfying (ref) equals $\mathbb{E}[Y(1) - Y(0)\mid W_0 = 1]$, the average treatment effect within population $\{W_0 = 1\}$. This is the case even when the sign of weight function $a(X)$ varies with $X$. However, if $\tau_0$ is nonconstant, the existence of causal representations will depend on the weight function $a(X)$. Among other cases, we will consider the case where no restrictions are placed on function $\tau_0$. In this case, the existence of a causal representation of $\mu(a,\tau_0)$ will require the sign of $a(X)$ to be constant.
To formalize this, let $\mathcal{T}$ denote a class of functions such that $\tau_0 \in \mathcal{T}$ and define
This is the set of regular subpopulations of $W_0$ such that the estimand $\mu(a,\tau_0) = \mathbb{E}[Y(1) - Y(0)\mid W^* = 1]$ for all functions $\tau_0$ in the set $\mathcal{T}$. If the set $\mathcal{W}(a;W_0,\mathcal{T})$ is empty, then the estimand $\mu(a,\tau_0)$ cannot be written as an average treatment effect over a regular subpopulation of $W_0$ uniformly in $\tau_0 \in \mathcal{T}$. We use this set to formally define a notion of uniform causal representation.
Recall that $\mathbb{E}[Y(1) - Y(0)\mid W^* = 1] = \mu(\underline{w}^*,\tau_0)$ where $\underline{w}^*(X) = \mathbb{P}(W^* = 1\mid W_0=1,X)$, so $W^* \in \mathcal{W}(a;W_0,\mathcal{T})$ if $\mu(a,\tau_0) = \mu(\underline{w}^*,\tau_0)$ for all $\tau_0 \in \mathcal{T}$. We further examine several cases for the set $\mathcal{T}$.
We begin by considering the largest class of functions in which $\tau_0$ lies: the class of all functions, subject to the moment condition in Assumption (ref) that ensures the existence of $\mu(a,\tau_0)$. We denote this class by
In this function class, we show the existence of a causal representation is equivalent to the estimand's weights being nonnegative. In what follows, let $a_{\max} \coloneqq \sup(\text{supp}(a(X)\mid W_0=1))$ be the essential supremum of $a(X)$ given $W_0=1$.
A uniform (in $\mathcal{T}_\text{all}$) causal representation exists if and only if $a(X)$ is nonnegative when $W_0 = 1$. To give some intuition on why the sign of $a(X)$ must be nonnegative with probability 1, we present a contradiction that occurs when $a(X)$ can be negative. Suppose that $\mathbb{P}(a(X) < 0 \mid W_0 = 1) > 0$ and consider the “adversarial” CATE function $\tau^-(X) = \mathbbm{1}(a(X) < 0)$. This CATE function is nonnegative for all $X$, and implies a positive average effect only for units with negative weights. However, it yields a strictly negative weighted estimand, $\mu(a,\tau^-) = \mathbb{E}[a(X)\mathbbm{1}(a(X) < 0)\mid W_0=1]/\mathbb{E}[a(X)\mid W_0=1] < 0$. Clearly, $\mu(a,\tau^-) $ cannot be the average treatment effect for any subpopulation of $W_0$, because averaging a nonnegative CATE function over any subpopulation cannot yield a negative average.
Conversely, if $a(X) \geq 0$, our proof constructively defines a regular subpopulation $W^*$ for which the average treatment effect is equal to the weighted estimand $\mu(a,\tau_0)$ uniformly in $\tau_0 \in \mathcal{T}_\text{all}$. Let
where $U \sim \text{Unif}(0,1) \perp \! \! \! \perp (Y(1),Y(0),X,W_0)$. This is a regular subpopulation of $W_0$ for which the probability of inclusion, conditional on $X$ and $W_0=1$, is proportional to $a(X)$. We can also interpret $\mu(a,\tau_0)$ as the average effect of an intervention in which units with covariate value $X$ are treated with probability $a(X)/a_{\max}$ given $W_0 = 1$. From this construction, we can see that $\underline{w}^*(X) = \mathbb{P}(W^*=1\mid W_0=1,X)$ is proportional to $a(X)$, and therefore $\mathbb{E}[Y(1) - Y(0)\mid W^*=1] = \mu(\underline{w}^*,\tau_0) = \mu(a,\tau_0)$ uniformly in $\tau_0 \in \mathcal{T}_\text{all}$.
The condition $a_{\max} < \infty$ restricts our attention to subpopulations with positive mass. We note that $a_{\max}$ is bounded above in each of our theoretical examples in Sections (ref) and (ref), implying that this condition trivially holds in these cases.
As mentioned earlier, the condition that weights are nonnegative is well established. BlandholBonneyMogstadTorgovitsky2022 show that it is equivalent to an estimand being “weakly causal,” which means that it will match the sign of $\tau_0$ whenever that sign is the same across all units. Thus, in the class of weighted estimands we consider, estimands have a causal representation uniformly in $\mathcal{T}_\text{all}$ if and only if they are weakly causal. This connection is formally established in Appendix (ref).
We now provide an existence result that requires the causal representation to exist only for the given $\tau_0$, rather than uniformly for $\tau_0$ in the larger set $\mathcal{T}_\text{all}$. The following result depends on the CATE function $\tau_0$ in the population, whereas Theorem (ref)'s condition depended only on the weight function $a(X)$. Thus, the distribution of the potential outcomes will have an impact on the existence of a causal representation given $\tau_0$. Using the notation from Definition (ref), a causal representation exists if and only if $\mathcal{W}(a;W_0,\{\tau_0\}) \neq \emptyset$. The following theorem characterizes this existence.
The existence condition in Theorem (ref) is weaker than the one in Theorem (ref) since we no longer require this representation to be valid for any CATE function, but rather just for the one that is identified from the population. The necessary and sufficient condition in this theorem only requires that the estimand is in the convex hull of the support of the CATEs. This means $\mu(a,\tau_0)$ has a causal representation even with negative weights, as long as there are CATEs smaller and greater than $\mu(a,\tau_0)$. We can see this support condition holds for all $\tau_0 \in \mathcal{T}_\text{all}$ if and only if $\mu(a,\tau_0)$ is in the support of $\tau_0(X)$ for any $\tau_0 \in \mathcal{T}_\text{all}$. This is precisely the case when the weights $a(X)$ are nonnegative since it guarantees $\inf(\text{supp}(\tau_0(X)\mid W_0=1)) \leq \mu(a,\tau_0) \leq \sup(\text{supp}(\tau_0(X)\mid W_0=1))$ for any $\tau_0$.
Analyzing the causal representation of an estimand under no restrictions on $\tau_0$ could be viewed as unnecessarily conservative in some settings. At the other extreme, assuming knowledge of $\tau_0$ may be unrealistic, especially in scenarios where $X$ has many components which makes the estimation of $\tau_0$ more challenging. For example, some shape constraints may be known to hold for $\tau_0$. In some economic applications one may posit that $\tau_0$ is monotonic or convex in some components of $X$, or positive/negative over a subset of $\text{supp}(X\mid W_0=1)$. In these cases, the existence of a causal representation may occur under weaker conditions than those in Theorem (ref), but stronger than those in Theorem (ref). In particular, one may be able to relax the requirement that $a(X) \geq 0$ without requiring that $\tau_0$ be completely known to the researcher. The following proposition shows this is the case when $\tau_0(X)$ is assumed to be linear in $X$.
The above proposition shows that placing restrictions on $\mathcal{T}$ may remove the requirement that $a(X) \geq 0$ for the existence of a uniform causal representation for an estimand. In particular, the requirement here is that $\frac{\mathbb{E}[a(X)X\mid W_0=1]}{\mathbb{E}[a(X)\mid W_0=1]}$ lies in the convex hull of the support of $X$ given $W_0=1$. When $X$ is scalar, this consists of an interval. This condition does not require $a(X)$ be nonnegative. For example, if $\text{supp}(X) = \{0,1,2\}$ and $W_0 = 1$ almost surely, then any combination of values of $(a(0),a(1),a(2))$ such that $\mathbb{E}[a(X)X]/\mathbb{E}[a(X)] \in [0,2]= \text{conv}(\text{supp}(X))$ implies a causal representation. Let $\mathbb{P}(X = x) = 1/3$ for $x \in \{0,1,2\}$ and $(a(0),a(1),a(2)) = (1,-1,1)$. Here units with $X=1$ have a negative weight, but $\mathbb{E}[a(X)X]/\mathbb{E}[a(X)] = 1 \in [0,2]$, implying that the corresponding weighted estimand has a causal representation uniformly in $\tau_0 \in\mathcal{T}_\text{lin}$. The result is stated for discrete $X$, but an estimand with negative weights can have a causal representation even when $X$ has continuous components.
We consider another class of CATE functions that restricts their heterogeneity. For $K \geq 0$, let
This function class uniformly bounds differences of the CATE function. When $K=0$, the CATE function is constant, and thus equal to $\mathbb{E}[Y(1) - Y(0)\mid W_0=1]$. When $K > 0$, CATEs may differ in value, but the maximum discrepancy between two CATEs is bounded above by $K$. We show that restricting the CATEs to satisfy this bounded difference assumption does not remove the requirement that $a(X)$ be nonnegative, unless $K = 0$, in which case all $a(\cdot)$ functions yield a causal representation uniformly in $\mathcal{T}_\text{BD}(0)$. We formalize this in the next proposition.
To understand this proposition, consider the adversarial CATE function $\tau^-(X) = K \cdot \mathbbm{1}(a(X) < 0)$, a member of $\mathcal{T}_\text{BD}(K)$, and assume $\mathbb{P}(a(X) \geq 0) < 1$. Then we obtain the same contradiction we discussed after Theorem (ref), where the CATE is nonnegative for all covariate values but the estimand is negative.
These last two propositions show that the impact of restrictions on $\tau_0$ on the requirement that $a(X)$ be nonnegative critically depends on the nature of these restrictions. Generalizations to additional or empirically motivated function classes are left for future work.
Many estimands will admit causal representations, but their associated subpopulations $\{W^* = 1\}$ will generally differ. Also, a weighted estimand may not always correspond to the target estimand a researcher is interested in. For example, a researcher may be interested in setting $\mathbb{E}[Y(1) - Y(0)\mid W_0=1]$, the average effect in population $\{W_0 = 1\}$, as the target parameter. In general, this parameter differs from $\mu(a,\tau_0)$.
However, the set of subpopulations corresponding to a weighted estimand can be used to understand how representative the weighted estimand is of the target. For example, we may seek estimands for which $\mathbb{P}(W^*=1\mid W_0=1)$ attains values closest to 1, since they have a higher degree of internal validity with respect to the target $\mathbb{E}[Y(1) - Y(0)\mid W_0=1]$. At one extreme, an estimand for which $\mathbb{P}(W^*=1\mid W_0=1) =1$ would be deemed to have the highest degree of internal validity for this target parameter since it would equal $\mathbb{E}[Y(1) - Y(0)\mid W_0=1]$. We convert this interpretation in a formal measure of internal validity that we define here.
Formally, $\overline{P}(a,W_0;\mathcal{T})$ is the sharp upper bound on $\mathbb{P}(W^* = 1\mid W_0=1)$ for any regular subpopulation $W^*$ of $W_0$ such that the weighted estimand $\mu(a,\tau_0)$ has a causal representation as the average treatment effect over subpopulation $W^*$. Note that we set $\overline{P}(a,W_0;\mathcal{T}) = 0$ when $\mathcal{W}(a;W_0,\mathcal{T})$ is empty. This object depends on the chosen function class $\mathcal{T}$, as did Theorems (ref) and (ref) in the previous section. Given the above terminology and assuming that $\mathbb{E}[Y(1) - Y(0)\mid W_0=1]$ is the target, we call $\overline{P}(a,W_0;\mathcal{T})$ a measure of the internal validity of estimand $\mu(a,\tau_0)$, and we use this definition in the remainder of the paper.
We can also compute the maximum value of $\mathbb{P}(W^* = 1)$ across $W^* \in \mathcal{W}(a;W_0,\mathcal{T})$, which measures the largest share of the entire population for which the weighted estimand has a causal representation. We refer to this measure as a measure of representativeness. The measures of internal validity and representativeness are the same when $W_0 = 1$ almost surely.
Note that $\mathbb{P}(W^*=1) = \mathbb{P}(W^*=1\mid W_0=1) \cdot \mathbb{P}(W_0=1)$ since $W^*$ is a subpopulation of $W_0$. The maximum value of $\mathbb{P}(W^* = 1)$ gives the internal validity of the weighted estimand with respect to target estimand $\mathbb{E}[Y(1) - Y(0)]$, the average treatment effect in the population from which the sample is drawn. Our measures of internal validity and representativeness are closely linked and a subpopulation will maximize $\mathbb{P}(W^* = 1)$ if and only if it maximizes $\mathbb{P}(W^* = 1 \mid W_0 = 1)$. We will also show how to use these measures to obtain simple bounds on target estimands.
We now derive explicit expressions for $\overline{P}(a,W_0;\mathcal{T})$. We focus on two cases, the first being when $\tau_0$ is unrestricted.
Without imposing any restrictions on the CATE function, except for the existence of second moments, the maximum value that $\mathbb{P}(W^*=1\mid W_0=1)$ can achieve is given by the following theorem.
Here we see that the maximum size of a subpopulation characterizing the estimand $\mu(a,\tau_0)$ depends on $a(X)$ through two terms: its conditional mean in the numerator, and its supremum $a_{\max}$ in the denominator. This bound can be computed at what ImbensRubin2015 call the “design stage” of the study, that is, without any knowledge of the conditional distribution of the outcome.
To understand the supremum's role in this expression, let $\underline{w}^*(X) = \mathbb{P}(W^* = 1\mid X,W_0=1)$ and note that $\mu(a,\tau_0) = \mathbb{E}[Y(1) - Y(0)\mid W^* = 1]$ is equivalent to writing
for all $\tau_0 \in \mathcal{T}_{\text{all}}$. Equation (ref) holding for all $\tau_0$ requires $\underline{w}^*(X)$ to be exactly proportional to $a(X)$. While the range of $a(X)$ is unconstrained, $\underline{w}^*(X)$ must lie in $[0,1]$ to be a valid conditional probability. Since we seek to maximize $\mathbb{P}(W^*=1\mid W_0=1) = \mathbb{E}[\underline{w}^*(X)\mid W_0=1]$, we let $\underline{w}^*(X)$ be the largest multiple of $a(X)$ that lies in $[0,1]$ with probability 1, which is defined below:
Here, $U \sim \text{Unif}(0,1)$ and $U \perp \! \! \! \perp (Y(1),Y(0),X,W_0)$. This population places relatively more weight on units with larger values of $a(X)$. Specifically, the population $\{W^* = 1\}$ contains a random subset of $\{W_0 = 1\}$ where the probability of inclusion is proportional to $a(X)$. Thus, units with larger values of $a(X)$ are more likely to be included in $W^*$. All units in $\{W_0 = 1\}$ with $X$ such that $a(X) = a_{\max}$ are included in $W^*$, whereas no units where $a(X) = 0$ are included.
The construction of this subpopulation is illustrated in Figure (ref) for the case where $x$ is continuous and where we omit the conditioning on $W_0 = 1$ for simplicity. We seek to maximize $\mathbb{P}(W^* = 1) = \int \underline{w}^*(x) f_X(x) \, dx$ with the requirement that $\underline{w}^*(x) \leq 1$ (or, equivalently, $\underline{w}^*(x)f_X(x) \leq f_X(x)$) and that $\underline{w}^*(x)$ is a multiple of $a(x)$. In the figure, we see that $a_{\max} > 1$ and thus the largest multiple of $a(x)$ that is weakly smaller than 1 is illustrated by the gray curve. The area under this curve is precisely $\mathbb{P}(W^* = 1)$. Note that the area under $f_X(x)$ is one, so closer alignment of the gray curve and the density $f_X(x)$ corresponds to more representative estimands.
Several further comments about Theorem (ref) are in order.
We now consider a simple example to give further intuition for Theorem (ref).
Consider an estimand $\mu(a,\tau_0)$ where $W_0 = 1$ almost surely, $a(X) \geq 0$, and where $X$ is binary with support $\text{supp}(X) = \{1,2\}$. Let $p_x = \mathbb{P}(X=x)$ for $x \in \{1,2\}$. As in equation (ref), $\mu(a,\tau_0)$ can be written as a linear combination of the two CATEs:
Let $\text{ATE} = \mathbb{E}[Y(1) - Y(0)]$ be the target estimand, which can be written as
If $a(1) = a(2)$, the relative weights placed on $\{X=1\}$ and $\{X=2\}$ by the estimand are equal to $p_1/p_2$, the ratio of the weights placed by the ATE\@. Therefore, the estimand equals the ATE and thus clearly has the maximum degree of internal validity with respect to the ATE\@. Applying Theorem (ref), we can directly see that, when $a(1) = a(2)$, $\overline{P}(a,W_0;\mathcal{T}_\text{all}) = \mathbb{E}[a(X)]/\sup_{x \in \{1,2\}} a(x) = a(2)/a(2) = 1$.
However, when $a(1) \neq a(2)$, the estimand's weights differ from $(p_1,p_2)$, the population weights for the two covariate cells. For concreteness, let $(p_1,p_2) = (0.2,0.8)$ and $(a(1),a(2)) = (0.24,0.09)$, where the latter correspond, for example, to the OLS weights of Proposition (ref) when the propensity score is $(p(1),p(2)) = (0.4,0.1)$. In this case, $(\omega_1,\omega_2) = (0.4,0.6)$ and thus
Relative to the ATE, the weighted estimand overrepresents the population with $X=1$ and underrepresents the population with $X=2$\@. The largest subpopulation $\{W^* = 1\}$ that causally represents the estimand can be constructed by combining subsets of the subpopulations defined by $\{X=1\}$ and $\{X=2\}$. Specifically, let
where $U \sim \text{Unif}(0,1)$ is independent of $(Y(1),Y(0),X)$. This is a regular subpopulation that contains all units with $X=1$ and three eighths of units with $X=2$, selected uniformly at random. Therefore
which yields $\mathbb{P}(W^* = 1) = p_1 \underline{w}^*(1) + p_2 \underline{w}^*(2) = 0.5$. The same quantity can be obtained from Theorem (ref), which implies that $\overline{P}(a,W_0;\mathcal{T}_\text{all}) = \mathbb{E}[a(X)] / \left( \sup_{x \in \{1,2\}} a(x) \right) = \left( a(1) p_1 + a(2) p_2 \right) / a(1) = 0.5$. The average effect in this subpopulation is given by \begingroup \allowdisplaybreaks
\endgroup which equals $\mu(a,\tau_0)$ for any choice of $\tau_0$. Note that the relative weights placed on $\{X=1\}$ and $\{X=2\}$ in subpopulation $\{W^* = 1\}$ are given by $\frac{\mathbb{P}(X=1\mid W^*=1)}{\mathbb{P}(X=2\mid W^*=1)} = 0.4/0.6 = \omega_1/\omega_2$, matching the ratio of the weights on $\{X=1\}$ and $\{X=2\}$ assigned by the estimand. The subpopulation $\{W^* = 1\}$ cannot expand while preserving this ratio since it already includes all units with $X=1$. Therefore, $W^*$ is the largest subpopulation for which $\mu(a,\tau_0) = \mathbb{E}[Y(1) - Y(0)\mid W^*=1]$ for any $\tau_0$.
The subpopulation size in Theorem (ref) can be used to bound the target estimand $\mathbb{E}[Y(1) - Y(0)\mid W_0=1]$. Consider a scenario where only the weighted estimand $\mu(a,\tau_0)$, which we assume has a causal representation uniformly in $\mathcal{T}_\text{all}$, and its internal validity are known. For example, this could be the case if a researcher uses a weighted estimand (e.g., OLS) and reports the measure we propose in Definition (ref) to quantify its degree of internal validity for the ATE\@. To simplify notation, assume that $W_0 = 1$ almost surely and denote $\overline{P}(a,W_0;\mathcal{T}_\text{all})$ by $\overline{P}_\text{all}(a)$. Abstracting from sample uncertainty, we only assume knowledge of the weighted estimand and its internal validity. We can decompose the target estimand, here $\mathbb{E}[Y(1) - Y(0)]$, as \begingroup \allowdisplaybreaks
\endgroup for a $W^* \in \mathcal{W}(a;W_0,\mathcal{T}_\text{all})$. If we have knowledge of bounds for the treatment effect $\mathbb{E}[Y(1) - Y(0)\mid W^* = 0]$, e.g., from the support of the potential outcomes, we can obtain bounds on $\mathbb{E}[Y(1) - Y(0)]$. For example, if $\text{supp}(Y(1) - Y(0)) \subseteq [B_\ell,B_u]$, bounds for the target estimand are given by \begingroup \allowdisplaybreaks
\endgroup The width of these bounds is minimized when $\mathbb{P}(W^*=1)$ is maximized, or when it equals the measure of internal validity for $\mu(a,\tau_0)$, given by $\overline{P}_\text{all}(a)$. The resulting bounds for $\mathbb{E}[Y(1) - Y(0)]$ can be written as \begingroup \allowdisplaybreaks
\endgroup If $\overline{P}_\text{all}(a) = 1$, it is easy to see that the estimand equals the ATE and that the bounds in (ref) collapse to a point. However, the ATE is not uniquely determined from $(\mu(a,\tau_0),\overline{P}_\text{all}(a))$ when $\overline{P}_\text{all}(a) < 1$.
The width of these bounds is $(B_u - B_\ell) \cdot (1 - \overline{P}_\text{all}(a))$. Hence for fixed $(B_\ell,B_u)$, this width decreases linearly with $\overline{P}_\text{all}(a)$. Moreover, values of $\overline{P}_\text{all}(a)$ close to 1, or high degrees of internal validity, lead to narrow bounds. It is easy to obtain a sample analog of these bounds by combining estimators for $\mu(a,\tau_0)$ and our proposed estimator for $\overline{P}_\text{all}(a)$ from Section (ref) below.
We note that bounds on $\mathbb{E}[Y(1) - Y(0)]$ may be tightened by assuming knowledge of other aspects of the joint distribution of $(Y(1),Y(0),D,X)$. For example, if knowledge of $a(\cdot)$ is assumed, additional constraints on $\mathbb{E}[Y(1) - Y(0)]$ can help narrow the bounds given in (ref). We focus here on the case where we add a single piece of additional information to $\mu(a,\tau_0)$, namely its internal validity, and how simple bounds can be obtained from the estimand and our proposed measure. We leave refinements of such bounds under different information sets to future work.
Now consider a case where the weighted estimand $\mu(a,\tau_0)$ has weights that can be negative, i.e., $\mathbb{P}(a(X) < 0) > 0$. We continue to assume that $W_0 = 1$ almost surely. In this case, we know that $\mu(a,\tau_0)$ does not have a causal representation uniformly in $\tau_0 \in \mathcal{T}_\text{all}$ and thus we cannot write $\mu(a,\tau_0) = \mathbb{E}[Y(1) - Y(0) \mid W^* = 1]$ uniformly in $\tau_0$. However, simple algebra reveals that an estimand with negative weights can be written as a weighted difference of two nonnegatively weighted estimands: \begingroup \allowdisplaybreaks
\endgroup where $\omega^+ \coloneqq \mathbb{E}[\max\{a(X),0\}]/ \mathbb{E}[a(X)]$ and $\omega^- \coloneqq \mathbb{E}[-\min\{a(X),0\}]/ \mathbb{E}[a(X)]$ are both nonnegative, and $\omega^+ - \omega^- = 1$. We note that $(\omega^+,\omega^-) = (1,0)$ when the estimand's weights are nonnegative, so this decomposition can be obtained regardless of the sign of $a(\cdot)$. Thus, by Theorem (ref), we can write
where $W^+$ and $W^-$ characterize two disjoint, regular subpopulations. As above, suppose we want to bound $\mathbb{E}[Y(1) - Y(0)]$, the ATE\@. Using the law of iterated expectations, we can write $\mathbb{E}[Y(1) - Y(0)]$ as \begingroup \allowdisplaybreaks
\endgroup Substituting equation (ref) in (ref) and assuming that $\mathbb{E}[Y(1) - Y(0)\mid W^- = 1]$ and $\mathbb{E}[Y(1) - Y(0)\mid W^+ + W^- = 0]$ lie in $[B_\ell,B_u]$ yields \begingroup \allowdisplaybreaks
\endgroup as an upper bound for the ATE\@. A lower bound is obtained by replacing $B_u$ with $B_\ell$. Thus, the ATE lies in the interval \begingroup \allowdisplaybreaks
\endgroup This interval is similar to the interval in (ref), but the latter is only valid when weights are nonnegative. These intervals are identical when weights are nonnegative because $\omega^+ = 1$ and $\mathbb{P}(W^+ = 1) = \mathbb{P}(W^* = 1)$ in that case. In order to compute the interval in (ref) and minimize its length, one needs to maximize the value of $\mathbb{P}(W^+ = 1)$, which corresponds to the level of internal validity of the estimand $\mu(\max\{a,0\},\tau_0)$ where $\max\{a,0\} \geq 0$, and compute the value of $\omega^+$. The resulting interval depends only on the ratio of the two quantities, which can be written as \begingroup \allowdisplaybreaks
\endgroup This last expression equals the level of internal validity of the original estimand $\mu(a,\tau_0)$ when it is assumed (perhaps incorrectly) to have nonnegative weights. It follows that the bounds in (ref) are valid regardless of whether weights are nonnegative, because minimizing the length of the interval in (ref) yields precisely the bounds in (ref). As mentioned earlier, these bounds do not make use of the entire distribution of $(Y,D,X)$, but simply of the original estimand $\mu(a,\tau_0)$ and of $\mathbb{E}[a(X)]/a_{\max}$, the expression for the level of internal validity under nonnegative weights.
We can also ask how internally valid a weighted estimand can be, given knowledge of the CATE function. In this case, the object of interest is
where $\tau_0$ is a given CATE function. Since $\tau_0$ is known, the condition $W^* \in \mathcal{W}(a;W_0,\{\tau_0\})$ can be written as $\mu_0 = \mu(\underline{w}^*,\tau_0)$, where we let $\mu_0 \coloneqq \mu(a,\tau_0)$ to simplify the notation. This condition is equivalent to $\mathbb{E}[(\tau_0(X)-\mu_0) \underline{w}^*(X)\mid W_0=1] = 0$, a linear constraint on the conditional probability of being in subpopulation $W^*$. Additionally, the objective function $\mathbb{P}(W^*=1\mid W_0=1) = \mathbb{E}[\underline{w}^*(X)\mid W_0=1]$ is linear in $\underline{w}^*$. Thus, the optimization in (ref) can be cast as a linear program. To see this, consider as an example the case where $W_0 = 1$ almost surely and where $X$ is discrete with finite support, i.e., $\text{supp}(X) = \{x_1,\ldots,x_K\}$. Let $f_k \coloneqq \mathbb{P}(W^*=1, X = x_k)$ and note that $f_k \in [0,p_k]$ where $p_k = \mathbb{P}(X=x_k)$.
We can write the above optimization problem as
a finite-dimensional linear program. This program has a feasible solution if $\tau_0(x_k) - \mu_0$ is not strictly positive or strictly negative for all $k$, meaning that the weighted estimand lies in the convex hull of CATE values, which is precisely stated in the condition for Theorem (ref). While there exist many methods for solving linear programs, the value function can be obtained through an algorithm that is simple to describe analytically.
When $\mu_0$ exceeds $\mathbb{E}[Y(1) - Y(0)\mid W_0 = 1]$, this algorithm reduces the weights associated with smallest CATEs until $\mu_0$ equals $\mathbb{E}[Y(1) - Y(0)\mid W^*=1]$ for some subpopulation. When $\mu_0 < \mathbb{E}[Y(1) - Y(0)\mid W_0 = 1]$, the same procedure is instead applied to the largest CATEs. The support assumption of Theorem (ref) guarantees that this algorithm ends.
When $X$ is not discretely supported, the problem can still be cast as a linear program, but its dimension is infinite, which generates difficulties in implementation. However, we show this program has an analytical solution even when $X$'s components are allowed to be discrete, continuous, and mixed, as is often the case in practice.
The computation of these bounds can be done using a linear programming algorithm when $X$ is discrete, or through plug-in estimators of the terms in equation (ref) regardless of the nature of the support of $X$.
To illustrate this theorem, let $\tau_0(X)$ be continuously distributed with support $[\underline{\tau},\overline{\tau}]$, and suppose $\mu_0 \in [\underline{\tau},\overline{\tau}]$. Without loss of generality, assume $E_0 \geq \mu_0$. If $E_0 = \mu_0$, then the estimand is perfectly representative of the population since it equals the average treatment effect over it. In the case where $E_0 > \mu_0$, the estimand is not representative of the entire population. We are searching for the largest subpopulation $\{W^* = 1\}$ such that $\mathbb{E}[Y(1) - Y(0)\mid W^*=1] = \mu_0$. Initializing $W^*$ at $W_0$, removing the subpopulation with the largest values of $\tau_0(x)$ yields the steepest decrease in $\mathbb{E}[Y(1) - Y(0)\mid W^* = 1]$. Therefore, $ \overline{P}(a,W_0;\{\tau_0\})$ is obtained by removing a subpopulation of the kind $W^-(\alpha) = \mathbbm{1}(\tau_0(X) > \alpha) \cdot W_0$ for a given threshold $\alpha$. This threshold is determined by the constraint
Thus, the remaining subpopulation $W^*$ corresponds to $W^* = \mathbbm{1}(\tau_0(X) \leq \alpha^*) \cdot W_0$ where $\alpha^*$ is the unique solution to (ref). In this setting, the value $\overline{P}(a,W_0;\{\tau_0\})$ is larger when the truncated subpopulations are smaller. In particular, this is the case when there are a few units with extreme values of $\tau_0$ whose removal has a large impact on the estimand, but a small impact on the share of the population.
Theorem (ref) can also be illustrated visually. In Figure (ref), the probability density function of $\tau_0(X)$ is drawn. In this figure, it is assumed that $\tau_0(X)$ is continuously distributed, that $W_0 = 1$ almost surely, and that $\mu < \mathbb{E}[\tau_0(X)] = \text{ATE}$\@. The representative subpopulation is obtained by trimming away covariate values that correspond to $\tau_0(X) \geq \alpha^+$, where $\alpha^+$ is determined by the equation $\mathbb{E}[\tau_0(X)\mid \tau_0(X) \leq \alpha^+] = \mu$. The size of the shaded area is the measure of internal validity.
Here we consider three identification strategies where commonly used estimands follow the structure of equation (ref). We show how the results in Sections (ref) and (ref) apply in each of these cases. For simplicity, we assume that $a_{\max} = \sup(\text{supp}(a(X)\mid W_0=1)) = \sup_{x \in \text{supp}(X\mid W_0=1)} a(x)$ in this section. This condition is satisfied when $a(\cdot)$ is continuous or when $X$ has finite support, among other cases. We also note that our assumption $a_{\max} < \infty$ holds trivially in every case considered below.
In Section (ref), we provided the expression for the coefficient on $D$ in a population regression of $Y$ on $(1,D,X)$:
Suppose the target estimand is the average treatment effect, i.e., $W_0 = 1$ almost surely. By Theorem (ref), there exists a regular subpopulation $W^*$ such that $\beta_\text{OLS}$ equals the average treatment effect over $W^*$ since the weight function $a(X) = p(X)(1-p(X))$ is nonnegative. By Theorem (ref), the upper bound on the size of subpopulation $W^*$ is
A corresponding subpopulation $W^*$ can be written as
where $U \perp \! \! \! \perp (Y(1),Y(0),X)$ and $U \sim \text{Unif}(0,1)$. This is a subpopulation where units with a larger variation in treatment given their covariate values are more likely to be included. The size of this subpopulation is largest when $\operatorname{var}(D\mid X) = p(X)(1-p(X))$ is constant, in which case $\mathbb{P}(W^* = 1\mid X) = 1$. This is the case if and only if $p(X)$ has support contained in $\{b,1-b\}$ for some $b \in (0,1)$. Whenever $\operatorname{var}(p(X)(1-p(X))) > 0$, $\{W^* = 1\}$ will be a strict subpopulation.
The size of this subpopulation is the expectation of $\operatorname{var}(D\mid X)$ divided by its maximum value. There are a few ways this expression can be further simplified or bounded. Its numerator is bounded above by $\operatorname{var}(D) = \mathbb{P}(D=1) \cdot \mathbb{P}(D=0)$, which is particularly simple to estimate. As for the denominator, it is a nonsmooth functional of $p(\cdot)$. However, if $X$ is continuously distributed, it may be likely that $p(X)$ is continuously distributed and thus that $1/2 \in \text{supp}(p(X))$. If this is the case, $\sup_{x \in \text{supp}(X)} p(x)(1-p(x)) = 1/4$. Combining these two approximations yields
when the support of $p(X)$ includes 1/2. This bound is trivial when $\mathbb{P}(D=1) = 1/2$, but is informative when the unconditional treatment probability is close to 0 or 1. For example, if $\mathbb{P}(D=1) = 0.1$, the OLS estimand cannot causally represent more than 36% of the population. This is consistent with the result in Sloczynski2022 that the OLS estimand is more similar to the ATE when $\mathbb{P}(D=1)$ is close to 1/2.
When $1/2 \in \text{supp}(p(X))$, we can also compute bounds on the ATE derived from $\beta_\text{OLS}$, bounds on the support of $(Y(1),Y(0))$, and our measure of internal validity $\overline{P}(a,W_0;\mathcal{T}_\text{all})$. Following the expression in (ref), bounds on the ATE are given by
Estimating these bounds requires the estimation of one additional quantity beyond the OLS estimand, which is the expectation of $\operatorname{var}(D\mid X)$. The width of these bounds depends crucially on $B_u - B_\ell$, the width of the support for unit-level treatment effects.
Alternatively, we can assess the internal validity of $\beta_\text{OLS}$ with respect to an alternative estimand such as $\mathbb{E}[Y(1) - Y(0)\mid D=1]$, the average treatment effect on the treated. In this case, we consider an alternative representation of the estimand:
Applying Theorem (ref) yields that
is the largest value that $\mathbb{P}(W^*=1\mid D=1)$ can take. Once again, this bound depends only on the propensity score and the distribution of $X$. This subpopulation satisfies \[\mathbb{P}(W^*=1\mid X,D=1) = \frac{1-p(X)}{1 - \inf_{x \in \text{supp}(X\mid D=1)} p(X)},\] so units with smaller propensity scores are more likely to be included in $W^*$, given that they are treated. $\overline{P}(1-p,D;\mathcal{T}_\text{all})$ is maximized at 1 when $p(X)$ is constant, or if $D \perp \! \! \! \perp X$. In this case, $\mathbb{P}(W^* = 1\mid D=1) = 1$ and $\mathbb{P}(W^* = 1) = \mathbb{P}(D=1)$.
If $p(X)$ takes values close to 0, the bound satisfies
This suggests that the OLS estimand is more representative of the ATT when the fraction of untreated units is larger. This again echoes the results in Sloczynski2022 on the relationship between $\mathbb{P}(D=1)$ and the interpretation of the OLS estimand.
We can also assess the internal validity of $\beta_\text{OLS}$ given $\tau_0(X) = \mathbb{E}[Y(1) - Y(0)\mid X]$. For simplicity, assume that $\tau_0(X)$ has a continuous distribution and, without loss of generality, assume that $\text{ATE} > \beta_\text{OLS}$\@. Then, using Theorem (ref), we obtain
where $\alpha^*$ satisfies $\mathbb{E}[\tau_0(X)\mid \tau_0(X) \leq \alpha^*] = \beta_\text{OLS}$. The quantity $\overline{P}(a,W_0;\{\tau_0\})$ is largest when the least amount of trimming needs to be applied. This is the case when the trimmed values are largest, or when $\mathbb{E}[\tau_0(X)\mid \tau_0(X) \geq \alpha]$ is large for large $\alpha$.
Now consider a binary treatment $D \in \{0,1\}$ and a binary instrument $Z \in \{0,1\}$. Potential treatments, denoted by $(D(1),D(0))$, are linked to the realized treatment through $Z$, that is, $D = D(Z)$. Potential outcomes, $Y(d,z)$ for $d,z\in \{0,1\}$, may depend on both $D$ and $Z$ in the absence of an exclusion restriction. Let $Y = Y(D,Z)$ be the realized outcome. As before, let $X$ denote covariates. We make the following assumptions.
The first instrumental variables estimand we consider was originally studied by AngristImbens1995. In addition to Assumption (ref), suppose that the model for $X$ is saturated, with $K$ possible combinations of covariate values, i.e., let $\text{supp}(X) = \{x_1,\ldots,x_K\}$. Let $X_S = \left( 1, \mathbbm{1}(X = x_1), \ldots, \mathbbm{1}(X = x_{K-1}) \right)$ and $Z_S = \left( Z, Z\cdot\mathbbm{1}(X = x_1), \ldots, Z\cdot\mathbbm{1}(X = x_{K-1}) \right) = ZX_S$, where $Z_S$ is the constructed instrument vector. The estimand in AngristImbens1995 is given by
where $W_S = \left( D, X_S \right)$, $Q_S = \left( Z_S, X_S \right)$, and $\left[ \cdot \right] _k$ denotes the $k$th element of the corresponding vector. This estimand has been studied by AngristImbens1995, Kolesar2013, Sloczynski2020, and BlandholBonneyMogstadTorgovitsky2022, and the representation in Proposition (ref) follows from Sloczynski2020.
Thus, $\beta_\text{2SLS}$ satisfies the representation in (ref) with $a_\text{2SLS}(X) = |\operatorname{cov}(D,Z\mid X)|$, $\tau_0(X) = \mathbb{E}[Y(1) - Y(0)\mid D(1) > D(0),X]$, and $W_0 = \mathbbm{1}(D(1) > D(0))$. Note that $\beta_\text{2SLS} = \text{LATE} \coloneqq \mathbb{E}[Y(1) - Y(0) \mid D(1) > D(0)]$ if and only if $a_\text{2SLS}(X)$ is uncorrelated with $\tau_0(X)$ given $D(1) > D(0)$.
The practical limitation of focusing on $\beta_\text{2SLS}$ is that applied researchers rarely create multiple interacted instruments BlandholBonneyMogstadTorgovitsky2022, which is how $Z_S$ is defined and used to obtain $\beta_\text{2SLS}$ above. A more practically relevant estimand is the “noninteracted” IV estimand,
where $Q = \left( Z, X \right)$ and $W = \left( D, X \right)$. We also make the following “rich covariates” assumption on the instrument propensity score, which is implied by the saturated specification in Proposition (ref).
Under the instrument validity assumption and the rich covariates assumption, Sloczynski2020 obtains the following representation of the “noninteracted” IV estimand.
It follows that $\beta_\text{IV}$ is a weighted estimand satisfying (ref) with weights $a_\text{IV}(X) = \operatorname{var}(Z\mid X)$, CATE function $ \tau_0(X) = \mathbb{E}[Y(1) - Y(0)\mid D(1) > D(0),X]$, and where the average is again taken over the complier subpopulation, i.e., $W_0 = \mathbbm{1}(D(1) > D(0))$.
First consider the estimand $\beta_\text{2SLS}$, which can be characterized as $\mu(a_\text{2SLS},\tau_0)$. Since $a_\text{2SLS}(X) \geq 0$, there exists a subpopulation of $\{D(1) > D(0)\}$ such that $\beta_\text{2SLS}$ is an average treatment effect over that subpopulation. The maximum size of that subpopulation is given by
The maximum value of $\mathbb{P}(W^* = 1\mid W_0=1)$ is obtained when $|\operatorname{cov}(D,Z\mid X)|$ does not depend on $X$\@. This occurs, for example, when the instrument and the fraction of units for which $D(1) > D(0)$ are independent of $X$. In this case, we have that $\beta_\text{2SLS} = \mathbb{E}[Y(1) - Y(0)\mid D(1) > D(0)]$, the average treatment effect for compliers.
Under the representation in Proposition (ref), the IV estimand has the same $W_0$, but has $a_\text{IV}(X) = \operatorname{var}(Z\mid X)$ instead. Here, $a_\text{IV}(X) \geq 0$ and
The internal validity of the IV estimand is maximized when $\operatorname{var}(Z\mid X)$ is constant, which occurs when $Z$ is independent of $X$\@. In this case, $\beta_\text{IV}$ equals LATE\@. The quantities $\overline{P}(a_\text{IV},W_0;\mathcal{T}_\text{all})$ and $\overline{P}(a_\text{2SLS},W_0;\mathcal{T}_\text{all})$ are not ranked uniformly in the distributions of $(D(1),D(0),X,Z)$ as there are data-generating processes that make each of these two quantities larger than the other. For example, if $\operatorname{var}(Z \mid X)$ is constant but $\mathbb{P}(D(1) > D(0)\mid X)$ is not, then $\overline{P}(a_\text{2SLS},W_0;\mathcal{T}_\text{all}) < \overline{P}(a_\text{IV},W_0;\mathcal{T}_\text{all})$. This scenario is plausible if $Z$ is randomly assigned and $X$ is a vector of pre-assignment characteristics. This inequality is reversed if $a_\text{2SLS}(X) = |\operatorname{cov}(D,Z\mid X)|$ is constant but $\mathbb{P}(D(1) > D(0)\mid X)$ is not. They are equally representative when $\mathbb{P}(D(1) > D(0)\mid X)$ is constant. In this case, the estimands are equal, so this is not unexpected.
Now suppose units are observed for $T$ periods and, for $t \in \{1,\ldots,T\}$, denote binary treatment by $D_t \in \{0,1\}$, potential outcomes $(Y_t(1),Y_t(0))$, and realized outcome $Y_t = Y_t(D_t)$. We assume units are untreated prior to period $G \in \{2,3,\ldots,T\} \cup \{+\infty\}$, receive the treatment in period $G$, and remain treated thereafter. We assume no units are treated in the first time period. This may include a group that remains untreated throughout, for which $G = +\infty$. Thus, $D_t = \mathbbm{1}(G \leq t)$. The panel is balanced, that is, no group appears or disappears over time.
The two-way fixed effects estimand is often used in this setting, and consists of regressing the outcome on the treatment indicator, group indicators, and period indicators. By partitioned regression results, the coefficient on treatment indicator is
where $\ddot{D}_t = D_t - \frac{1}{T}\sum_{s=1}^T D_s - \mathbb{E}[D_t] + \frac{1}{T}\sum_{s=1}^T \mathbb{E}[D_s]$.
We assume a version of parallel trends most similar to the one in ChaisemartinDHaultfoeuille2020.
We use a proposition that is essentially a special case of Theorem 1 in ChaisemartinDHaultfoeuille2020 to obtain a representation of $\beta_\text{TWFE}$ as a weighted average.
We show the above representation satisfies equation (ref) by introducing an auxiliary variable $P$ that is uniformly distributed on $\{1,\ldots,T\}$ independently from $\{(Y_t(0),Y_t(1),G)\}_{t=1}^T$. This period variable denotes the time period and we use it to define $(Y(1),Y(0),Y,D) \coloneqq (Y_P(1),Y_P(0),Y_P,D_P)$, which are potential outcomes, the realized outcome, and treatment at random period $P$, respectively.
Letting $X = (G,P)$, this means we can write $\beta_\text{TWFE}$ as
where $\tau_0(X) = \mathbb{E}[Y(1) - Y(0) \mid D = 1, G, P]$, $a_\text{TWFE}(X)$ is not generally nonnegative, and the average is taken over the treated units, i.e., $W_0 = D$.\footnote{$\tau_0(X)$ is what CallawaySantAnna2021 call “the group-time average treatment effect.”} A nonnegative weight function can be obtained under the assumption that $\tau_0(X)$ is constant over time. This property was described in ChaisemartinDHaultfoeuille2020 and Goodman-Bacon2021, and the resulting representations of the TWFE estimand are given in their Theorem S2 and equation (16), respectively. The following proposition yields a simple expression for the weights in our setting.
As is the case of the representation in Proposition (ref), the two-way fixed effects estimand in Proposition (ref) satisfies the representation in (ref), with $X = G$, $W_0 = D$, $\tau_0(X) = \mathbb{E}[Y(1) - Y(0)\mid D=1,G]$, and the weight function $a_\text{TWFE,H}(G) \geq 0$. This weight function is derived in Appendix (ref), with additional comparisons with the weights in Goodman-Bacon2021 in Appendix (ref).
We now consider the weights obtained in Proposition (ref) under its assumptions. These weights are nonnegative and therefore Theorem (ref) guarantees the existence of a causal representation for $\beta_\text{TWFE}$ uniformly in $\tau_0 \in \mathcal{T}_\text{all}$.\footnote{In the context of Proposition (ref), $\tau_0$ is a function of $G$ only, thus $\mathcal{T}_\text{all}$ denotes the set of all functions of $G$ with finite second moments. Note that this is a strict subset of all “time-heterogeneous” conditional average treatment effects, $\mathbb{E}[Y(1) - Y(0) \mid D=1, G,P]$.} Using Theorem (ref), the internal validity of $\beta_\text{TWFE}$ relative to target parameter $\mathbb{E}[Y(1) - Y(0)\mid D=1]$ is given by \begingroup \allowdisplaybreaks
\endgroup Due to the absorbing nature of the treatment in our setting, all expressions involving the distribution of $D$ given $P$ or $G$ can be derived as a function of the marginal distribution of $G$. Therefore, $\overline{P}(a_\text{TWFE,H},D;\mathcal{T}_\text{all})$ depends only on $\{\mathbb{P}(G=g)\}_{g \in \{2,\ldots,T\}}$.
To give some intuition, consider the case where $T = 3$ and therefore $G \in \{2,3,+\infty\}$. In this case, calculations yield that
where $\omega = \frac{4 - 2\mathbb{P}(G=2) - 4\mathbb{P}(G=3)}{2 - 2\mathbb{P}(G=2) - \mathbb{P}(G=3)} = \frac{a_\text{TWFE,H}(2)}{a_\text{TWFE,H}(3)}$. Therefore, $\beta_\text{TWFE}$ is perfectly representative of the ATT if and only if $\omega = 1$, which occurs if and only if $\mathbb{P}(G = 3) = 2/3$. Thus, $\overline{P}(a_\text{TWFE,H},D;\mathcal{T}_\text{all}) = 1$ when $\mathbb{P}(G=3) = 2/3$ and the internal validity of $\beta_\text{TWFE}$ declines as $|\mathbb{P}(G=3) - 2/3|$ increases. This is due to the weight function $a_\text{TWFE,H}(g)$ being constant in $g$ if and only if $\mathbb{P}(G = 3) = 2/3$. As before, constant weights imply that the weighted estimand equals the average treatment effect over $\{W_0 = 1\}$.
\endinput
We now consider estimation and inference for our measures of internal validity and representativeness. We focus our attention on the case when $\mathcal{T} = \mathcal{T}_\text{all}$ and briefly discuss the case where $\mathcal{T} = \{\tau_0\}$ in Appendix (ref).
To measure internal validity, we seek to estimate
Suppose we observe a random sample of size $n$, $\{(W_i,X_i)\}_{i=1}^n$, where $W_i$ is a set of variables needed to estimate $a(\cdot)$ and $w_0(\cdot)$. For example, under unconfoundedness we can let $W_i = D_i$ since the distribution of $(D,X)$ is sufficient to identify $a(\cdot)$; the outcome's distribution does not affect $\overline{P}(a,W_0;\mathcal{T}_\text{all})$. In our instrumental variables examples, we let $W_i = (D_i,Z_i)$.
Assuming the existence of estimators for $a(\cdot)$ and $w_0(\cdot)$, we consider the following analog estimator of $\overline{P}(a,W_0;\mathcal{T}_\text{all})$:
Here $c_n$ is a tuning parameter that converges to 0 as $n$ diverges. We start by noting that we can estimate $\mathbb{E}[a(X)\mid W_0=1]$ via $\frac{\frac{1}{n}\sum_{i=1}^n \widehat{a}(X_i)\widehat{w}_0(X_i)}{\frac{1}{n}\sum_{i=1}^n \widehat{w}_0(X_i)}$, which will be consistent under standard conditions on $\widehat{a}(\cdot)$ and $\widehat{w}_0(\cdot)$. Estimating $a_{\max} = \sup(\text{supp}(a(X) \mid W_0 = 1))$ is more delicate. In some of our examples, this supremum is known or can be bounded above without using data. For example, the OLS estimand under unconfoundedness has weights $a(X) = \operatorname{var}(D \mid X)$ which are bounded above by $1/4$. If $X$ is continuously distributed, then 1/2 may lie in the support of $p(X)$, and thus we may avoid the estimation of $a_{\max}$. The IV estimand of Section (ref) has weights $a_\text{IV}(X) = \operatorname{var}(Z\mid X)$ which are similarly bounded above by $1/4$. Similarly, $a_\text{2SLS}(X) = |\operatorname{cov}(D,Z\mid X)| \leq 1/4$ by the Cauchy--Schwarz inequality. If knowledge of $a_{\max}$ is not assumed, but $\text{supp}(X\mid W_0=1)$ is known and $a(x)$ is continuous,\footnote{Note that $a(x)$ is trivially continuous on finite support.} then $\sup_{x \in \text{supp}(X\mid W_0=1)} \widehat{a}(x)$ will be consistent for $a_{\max}$ when $\widehat{a}(x)$ is consistent for $a(x)$ uniformly in $x \in \text{supp}(X\mid W_0=1)$. Many parametric and nonparametric estimators for $a(\cdot)$ satisfy this requirement.
In Appendix (ref), we prove the consistency of $\widehat{\overline{P}}$ and derive its limiting distribution. We also provide a step-by-step bootstrap algorithm that can be employed to conduct inference and prove its validity. This bootstrap approach is based on FangSantos2019 and is nonstandard, but yields valid inferences even when $a_{\max}$ is estimated, as opposed to standard bootstrap approaches, such as the empirical bootstrap.
\endinput
\defcitealias{CallawaySantAnna2021}{CS}
In this section, we implement the proposed tools in an application to the effects of unilateral divorce laws in the U.S. on female suicide, as in SW2006. Between 1969 and 1985, 37 states (including D.C.) reformed their law by enabling each spouse to seek divorce without the other spouse's consent. SW2006 argue that these “unilateral” or “no-fault” divorce laws reduced female suicide, domestic violence, and spousal homicide. The results on female suicide are also replicated by Goodman-Bacon2021, whose analysis we follow here.
Our sample consists of 41 states observed over the 1964--1996 period. The outcome of interest is the state- and year-specific female suicide rate, as computed by the National Center for Health Statistics. The treatment is whether the state allowed unilateral divorce in a given year. Following Goodman-Bacon2021, our sample omits Alaska and Hawaii. We also omit eight further states which had unilateral divorce laws preceding 1964 and are therefore always treated within our timeframe.
Panel A of Table (ref) reports our baseline estimates of the average effects of unilateral divorce laws on female suicide. After we drop the eight always-treated states, the TWFE estimate, --0.604, becomes much smaller in absolute value than the corresponding estimate in Goodman-Bacon2021, --3.080. Unlike that estimate, ours is also statistically insignificant, with $p$-value = 0.819.
The conclusion changes, however, when we explicitly target the average treatment effect on the treated (ATT), that is, the average effect for the largest subpopulation for which such an effect is identified under standard assumptions. Using the approach of CallawaySantAnna2021, we obtain an estimate of --10.220 with a $p$-value of 0.001. The approach of Wooldridge2025 produces an estimate of --5.530 and a $p$-value of 0.138. These estimates are more strongly suggestive of a causal effect of unilateral divorce laws than the TWFE estimate.
While the TWFE estimate and the two estimates of the ATT are quite different, this paper focuses on another implication of the nonuniformity of the TWFE weight function. We ask: How representative of the underlying population is the TWFE estimand? What is the internal validity of this estimand if we are interested in the treated subpopulation? Panel B of Table (ref) reports our estimates of $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1 \mid D=1)$, based on the representation of the TWFE estimand in ChaisemartinDHaultfoeuille2020, revisited in our Proposition (ref). First, because the weights on some group-time average treatment effects are negative, the TWFE estimand does not have a causal interpretation uniformly in $\tau_0$, $\widehat{\mathbb{P}}(W^*=1) = \widehat{\mathbb{P}}(W^*=1 \mid D=1) = 0$. Second, when we estimate the CATE function and use these estimates in constructing the bounds, we conclude that the TWFE estimand corresponds to the average treatment effect for at most 62.16% of the treated units or 38.73% of the entire population.
Panel C of Table (ref) revisits these questions on the basis of the representation of the TWFE estimand in Proposition (ref). Here, we assume that group-time average treatment effects are constant over time, which eliminates the problem of negative weights. Indeed, we now conclude that the TWFE estimand has a causal interpretation uniformly in $\tau_0$, even if it is still not particularly representative of the underlying population or the treated subpopulation. Our estimates of $\mathbb{P}(W^*=1)$ and $\mathbb{P}(W^*=1 \mid D=1)$ are equal to 14.00% and 22.46%, respectively. When we use the estimated CATE function in constructing the bounds, these estimates increase to 48.02% and 77.07%. This is obviously much more than our initial estimate of 0, but still substantially less than 1, guaranteed in the case of $\mathbb{P}(W^*=1 \mid D=1)$ when using the estimation methods in CallawaySantAnna2021, Wooldridge2025, and other recent papers, each of which explicitly targets the ATT\@.
\endinput
In this paper, we studied the representativeness and internal validity of a class of weighted estimands, which includes the popular OLS, 2SLS, and TWFE estimands in additive linear models. We examined the conditions under which such estimands can be written as the average treatment effect over a subpopulation. When a given estimand can be shown to correspond to the average treatment effect for a large subset of the population of interest, we say its internal validity is high. In our main results, we derived the sharp upper bound on the size of that subpopulation under different assumptions on treatment effect heterogeneity, which offers a practical tool to quantify the internal validity of weighted estimands.
\endinput
\setlength\bibsep{0pt}