Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
60,625 characters · 9 sections · 48 citation commands
Assessing External Validity Over Worst-case Subpopulations
\RUNTITLE{Assessing External Validity Over Worst-case Subpopulations}
\TITLE{Assessing External Validity Over Worst-case Subpopulations}
\ARTICLEAUTHORS{ \AUTHOR{Sookyo Jeong} \AFF{Lyft Inc., \EMAIL{[email removed]}}
\AUTHOR{Hongseok Namkoong} \AFF{Decision, Risk, and Operations Division, Columbia Business School, New York, NY 10027, \EMAIL{[email removed]}} , \URL{hsnamkoong.github.io} }
\ABSTRACT{Study populations are typically sampled from limited points in space and time, and marginalized groups are underrepresented. To assess the external validity of randomized and observational studies, we propose and evaluate the worst-case treatment effect (WTE) across all subpopulations of a given size, which guarantees positive findings remain valid over subpopulations. We develop a semiparametrically efficient estimator for the WTE that analyzes the external validity of the augmented inverse propensity weighted estimator for the average treatment effect. Our cross-fitting procedure leverages flexible nonparametric and machine learning-based estimates of nuisance parameters and is a regular root-$n$ estimator even when nuisance estimates converge more slowly. On real examples where external validity is of core concern, our proposed framework guards against brittle findings that are invalidated by unanticipated population shifts.
}
\KEYWORDS{external validity, distributional robustness, semiparametrics} \HISTORY{This paper was first submitted on Jan, 2022.}
\else
\makeatletter \long\def\@makecaption#1#2{ \vskip 0.8ex \setbox\@tempboxa\hbox{ {\bf #1:} #2} \parindent 1.5em \dimen0=\hsize \advance\dimen0 by -3em \ifdim \wd\@tempboxa >\dimen0 \hbox to \hsize{ \parindent 0em \hfil \parbox{\dimen0}{ {\bf #1.} #2 } \hfil} \else \hbox to \hsize{\hfil \box\@tempboxa \hfil} \fi } \makeatother
When the study population is different from those affected by the treatment, the external validity of a study's finding may be called into question. Study populations are often sampled from a particular set of points in space and time and may not represent future populations of interest CampbellSt63, Manski13, RosenzweigUd20, DehejiaPoSa21. Furthermore, study populations often lack diversity, and minority groups are underrepresented. For example, out of $10,000+$ cancer clinical trials funded by the National Cancer Institute, less than 5% of participants were non-white ChenLaDaPaKe14, ShenoyHa15. When the treatment effect is heterogeneous, both randomized and observational studies lose external validity outside the study population. While large-scale randomized trials offer a “gold standard” for internal validity, their external validity can be nevertheless called into question over spatiotemporal changes in the population Deaton10, BasuSuHa17.
Existing approaches for assessing external validity require the knowledge of the target population HotzImMo05, AngristFe10, ColeSt10, StuartCoBrLe11, Tipton13, LeskoBuWeEdHuCo17, AndrewsOs17, Meager19. However, target populations chosen at the time of analysis may not be sufficient to guarantee external validity over unforeseen shifts in the population that occur post-analysis. Heuristic approaches such as estimating treatment effects over fixed subgroups are similarly limited as effects typically vary over a combination of multiple characteristics like race, gender, age, and income, a phenomenon we refer to as intersectionality.
To illustrate these challenges, consider estimating the effect of Medicaid enrollment on doctors' office utilization based on the National Health Interview Survey (NHIS) in 2009, where we focus on whether the individual made any visits to doctors two weeks prior to the survey date as the main outcome. Observed treatment effects and resulting decisions in 2009 must remain valid over a priori unknown shifts in the population over time. We observe substantial intersectionality in the treatment effect (Figure (ref)a): for Hispanic U.S. citizens, the treatment effect changes signs depending on professional degree attainment. In the decade following 2009 (“future”), there is a major shift in the underlying population (Figure (ref)b), and observed effects in 2009 are no longer valid in the future, as we later illustrate in Section (ref).
To assess external validity over unanticipated population shifts in the population, we propose and study the worst-case treatment effect, $\mbox{WTE}_{\alpha}$, defined over all subpopulations that comprise at least $\alpha$-fraction of the study population (see Eq. (ref) to come). As the $\mbox{WTE}_{\alpha}$ bounds the average treatment effect (ATE) and reduces to it when $\alpha=100\%$, the WTE analyzes the sensitivity of an average-case finding under population shifts, guaranteeing that conservative findings using the WTE remain valid uniformly over subpopulations. For example, if low-income Hispanic U.S. citizens with professional degrees comprise at least $20\%$ of the study population, positive findings with respect to $\mbox{\rm WTE}_{.2}$ guarantee the treatment remains effective over this subgroup.
We develop a semiparametrically efficient estimator of $\mbox{\rm WTE}_{\alpha}$, analyzing the external validity of the augmented inverse propensity weighted (AIPW) estimator RobinsRoZh94, RobinsRo95 for the ATE. Our $K$-fold cross-fitting procedure leverages machine learning-based and nonparametric estimators of nuisance parameters and provides a worst-case bound on the cross-fitted AIPW for the ATE ChernozhukovChDeDuHaNeRo18; our estimator reduces to the AIPW when $\alpha = 100\%$. On real datasets where external validity is of core concern, our worst-case sensitivity approach identifies disadvantaged subpopulations based on a priori nontrivial demographic groupings and guards against brittle findings that are invalidated under population shifts (Section (ref)).
Specifically, we exploit the dual representation of the worst-case over subpopulations to derive our semiparametric estimator (Section (ref)). By virtue of satisfying an orthogonality property (similar to the AIPW for the ATE), under standard assumptions required for the identification and estimation of the ATE, our augmented estimator of $\mbox{WTE}_{\alpha}$ enjoys central limit rates even when estimates of the nuisance parameters converge at slower-than-parameteric rates (Section (ref)). Since $\mbox{WTE}_{\alpha}$ is nonlinear in the underlying probability measure, our main asymptotic result (Theorem (ref)) requires a novel theoretical analysis different from the estimating equations framework (method of moments) studied by ChernozhukovChDeDuHaNeRo18.
We prove that our augmented estimator for the $\mbox{WTE}_{\alpha}$ is semiparametrically efficient, in both observational and randomized studies (Section (ref)). Our semiparametric efficiency bound informs experimental design with external validity as a central concern. Power calculations based on our efficiency bound provide the minimal sample size required to detect a specified effect size for the $\mbox{WTE}_{\alpha}$. Our bounds quantify how testing external validity against smaller subpopulations ($\alpha$) requires a correspondingly larger sample sizes.
\paragraph*{Related work}
Assessing the external validity of randomized and observational studies is an active area of research in causal inference DahabrehHe19. When the target population is known but different from the study population, many authors have leveraged the relationship between the two populations to guarantee external validity. A prevalent approach is to view selection into the study as another “treatment”, adjusting estimates of the ATE based on some information about the target population ColeSt10, StuartCoBrLe11, Tipton13, KernStHiGr16, LeskoBuWeEdHuCo17, AndrewsOs17, Meager19. StuartCoBrLe11 and Tipton13 use the probability of being included in the study to adjust for population bias, assuming that sample selection decisions only depend on observed covariates. HotzImMo05 apply bias-corrected matching methods to predict the impact of a program by using observations collected from a different location. In the context of structural causal models, Bareinboim, Pearl and colleagues identify settings that allow external validity in a series of works BareinboimPe12, BareinboimPe16.
External validity is of particular concern when identification strategies only allow studying a local notion of treatment effect. In such scenarios, several authors aim to connect estimates for a local population to a (known) broader population by leveraging a postulated structure between the two populations. For instrumental variable strategies, there is a line of work (see, for example, AngristImRu96, Angrist04, AngristFe10) studying when the treatment effects for compliers (LATE) can inform effects for a broader population. Most recently, RosenzweigUd20 directly estimate external validity over time when exogenous aggregate shocks are observed. For regression discontinuity designs, DongLe15, AngristRo15, BertanhaIm20 analyze settings where local estimates are externally valid and develop corresponding statistical tests.
Compared to the above methods, our worst-case sensitivity approach does not assume knowledge of the target population. Our conservative approach is agnostic to the unknown shifts in the population and provides uniform guarantees over subpopulations comprising at least $\alpha$-fraction of the study population. This is conceptually related to recent works on distributionally robust optimization in operations research and supervised learning, where models are trained to optimize a worst-case loss over distribution shifts DuchiHaNa20. Relatedly, BoGa21 studied external validity as a notion of stability of the joint distribution between the observed outcome and treatment assignments.
Study populations must be designed to be as diverse as possible across demographics, space, and time. Our approach can guarantee meaningful external validity only if the study includes heterogeneous subpopulations. (When the study population is not representative, our approach can nevertheless raise alarms.) A design-based approach to external validity complements our sensitivity framework by promoting diversity in the study population as a central concern. Several works in development and labor economics aim to improve the external validity of a (quasi-) experiment by collecting data over multiple sites and temporal points CrucesGa07, BanerjeeKaZi15, GertlerShAlCaMaPa15, DupasKaRoUb18, RosenzweigUd20, DehejiaPoSa21. Tipton and colleagues develop methodologies for measuring the diversity of a study population alongside practical experimental design guidelines Tipton14, TiptonPe17, TiptonRo18.
Our worst-case approach is broadly related to previous works that estimate treatment effects beyond mean differences Rothe10, KimKiKe18. ChernozhukovFeLu18 study sorted effects, a collection of sorted quantiles of the conditional average treatment effect (CATE). They develop central limit results for this nonparametric estimand under Donsker conditions (i.e., functional CLT) on CATE estimates. The $\mbox{WTE}_{\alpha}$ we introduce in Section (ref) is a tail-average of sorted effects, and our semiparametric approach extends the AIPW under unanticipated shifts in the population. Theoretically, we prove central limit results for our estimator without requiring Donsker conditions on CATE estimates. Our approach is not to be confused with quantile treatment effects Firpo07, which measures the difference between quantiles of $Y(1)$ and $Y(0)$.
Using the potential outcomes notation to denote counterfactuals, we let $Y(1)$ and $Y(0)$ be outcomes corresponding to treatment and control and let $Z \in \{0, 1\}$ be the assigned treatment Rubin74. We study both randomized and observational studies, assuming the analyst has access to observed covariates $X \sim P_X$. A standard goal is to estimate the average treatment effect, $\mbox{ATE} \defeq \E[Y(1) - Y(0)] = \E_{X \sim P_X}\left[\mu\opt(X)\right]$, where $\mu\opt(X) \defeq \E[Y(1) - Y(0)|X]$ is the conditional average treatment effect (CATE). Throughout, we assume the distribution $Y(1), Y(0) \mid X$ remains unchanged over subpopulations.
As the study population $P_X$ may not be representative of those affected by the treatment, we are interested in measuring the sensitivity of a study's finding to shifts in the underlying population. We consider the set of all subpopulations (probabilities) $Q_X$ that comprise more than $\alpha \in (0, 1]$ fraction of the study population $P_X$
As a convention, we assume the desired sign of the treatment effect is negative (the positive case is symmetric). We propose and study the worst-case subpopulation treatment effect
When treatment effects are highly heterogeneous and external validity is of particular concern, the worst-case subpopulation treatment effect (ref) will be substantially different from the ATE. The worst-case bound $\mbox{WTE}_{\alpha}$ reduces to the ATE when $\alpha = 100\%$.
The modeling choice of $\alpha$ in the definition (ref) is important, and it should be informed by domain knowledge. When a study population is not representative, we recommend selecting a smaller value of $\alpha$. For example, the analyst may reason about the level of bias anticipated in the data collection process or use the size of proxy target groups in the study. In the latter case, positive findings with the chosen level of $\alpha$ guarantee uniformly valid treatment effects over all minority subpopulations of size $\alpha$, not just the proxy targets. The choice of $\alpha$ should also consider the data size: as $\alpha$ becomes small, inference becomes difficult, as our semiparametric efficiency bounds demonstrate in Section (ref). Even when the analyst does not commit to a single level of $\alpha$, evaluating the WTE over a range of $\alpha$'s can offer a practical diagnostic. The level of $\alpha$ at which $\mbox{WTE}_{\alpha}$ crosses a threshold (e.g., 0) is often of particular practical interest as it represents the smallest subpopulation size over which average-case findings remain valid.
To derive our augmented estimator, we begin by simplifying the primal problem (ref) over (infinite-dimensional) covariate distributions $Q_X \in \mc{Q}_{\alpha}$ to its dual representation over a one-dimensional threshold on the CATE $\mu\opt(X) = \E[Y(1) - Y(0) \mid X]$. The dual reformulation shows an equivalence between worst-case subpopulation performances and tail-averages. We rely on this relationship heavily to derive our augmented estimator and to prove its asymptotic properties. We make the dependence on the underlying probability explicit and write $\E_Q[X]$, except for when $Q = P$, the data-generating distribution. The following lemma is a consequence of ShapiroDeRu09.
The dual optimum is attained at $P_{1-\alpha}^{-1}(\mu\opt)$ giving the second equality. This tail-average is known as the conditional value-at-risk (CVaR), a common risk measure in portfolio optimization RockafellarUr00. In contrast to the $\mbox{WTE}_{\alpha}$ involving an unknown nuisance parameter $\mu\opt(X)$ that needs to be estimated, the CVaR is typically considered over an observable random variable---this gives rise to a salient semiparametric structure. The dual shows the worst-off subpopulation is given by those who get disproportionately and adversely affected by the treatment, measured by $X$ such that $\mu\opt(X) \ge P_{1-\alpha}^{-1}(\mu\opt)$. To illustrate how the WTE (ref) accounts for heterogeneity across subpopulations, consider $\mu\opt(X) \sim N(-.1, 1)$ so that there is substantial heterogeneity across covariates. Although the $\mbox{ATE} = -.1$ suggests a negative treatment effect, $\mbox{WTE}_{.9} = 0.184$, meaning the treatment effect goes in the reverse direction of the ATE for even $90\%$ of the study population.
To identify causal effects, we assume no unobserved confounding and overlap between the treated and control groups.
We also assume that units do not interact with each other (no interference), and that we observe i.i.d. units $D_i = (X_i, Y_i, Z_i)$ for $i = 1, \ldots, n$ (stable unit treatment value assumption Rubin80). Finally, we require the following standard condition that uniformly bounds the conditional variance of the residuals for $z \in \{0, 1\}$.
These standard assumptions are also required to identify and estimate the ATE ImbensRu15, ChernozhukovChDeDuHaNeRo18.
Recalling that $h\opt$ is a nuisance parameter determining the worst-off subpopulation in Lemma (ref), we consider the following key nuisance parameters
Letting $D = (X, Y, Z)$ be the tuple of observed data and $(\mu_0, \mu_1, e, h)$ be the tuple of nuisance parameters, we consider the augmentation term
Under the stated assumptions, we have $\E[\kappa(D; \mu_0\opt, \mu_1\opt, e\opt, h\opt)] = 0$. Instead of estimating $\mbox{WTE}_{\alpha}$, we estimate the augmented form $\mbox{WTE}_{\alpha} + \E[\kappa(D; \mu_0\opt, \mu_1\opt, e\opt, h\opt)]$. When $\alpha = 1$ so that $\mbox{WTE}_{1} = \mbox{ATE}$, our estimator reduces to the augmented inverse probability weighted (AIPW) estimator for the ATE. Thus, our estimator can be viewed as an extension of the AIPW estimator under shifts in the underlying population.
\ifdefined\usemsstyle
\else
\fi
We now formally define our cross-fitted augmented estimator $\what{\omega}_{\alpha}$ and an estimate $\what{\sigma}^2_{\alpha}$ of its asymptotic variance
As we show in Section (ref), these estimates give an asymptotically exact confidence interval $\P(\mbox{WTE}_{\alpha} \in [\what{\omega}_{\alpha} \pm z_{\delta} \what{\sigma}_{\alpha} / \sqrt{n}]) \to 1-\delta$ if we set $z_{\delta}$ to be the $(1 - \delta/2)$-quantile of a standard normal distribution. The asymptotic variance (ref) is the best attainable in the typical semiparametric sense as our semiparametric efficiency bounds in Section (ref) show.
We fit estimators of nuisance parameters on the auxiliary sample, and combine them via the augmented dual form (ref) to evaluate the treatment effect on the worst-off subpopulation. Our approach is agnostic to the nuisance estimation method, and in particular, allows flexible use of machine learning models and nonparametric techniques to estimate $\mu\opt_z$ and $e\opt$. To estimate the threshold function $h\opt(X)$ that determines the worst-case subpopulation, we first compute an estimator $\what{q}$ of $P_{1-\alpha}^{-1}(\mu\opt)$ based on the auxiliary data, and take $\what{h}(x) \defeq \frac{1}{\alpha} \indic{(\what{\mu}_1 - \what{\mu}_0)(x) \ge \what{q}}$. In some applications, large quantities of unlabeled covariate observations can be cheaply collected even when labeled observations $(X, Y, Z)$ are expensive. Then a particularly nice estimator $\what{q}$ of $P_{1-\alpha}^{-1}(\mu\opt)$ can be constructed by evaluating the $(1-\alpha)$-quantile of the CATE estimator $\what{\mu} = \what{\mu}_1 - \what{\mu}_0$ on unlabeled observations. With cheap unlabeled covariates, such an estimator can be made arbitrarily close to $P_{1-\alpha}^{-1}(\what{\mu})$, and hence close to $P_{1-\alpha}^{-1}(\mu\opt)$ if $\what{\mu}$ is sufficiently close to $\mu\opt$ as we show in Section (ref).
To utilize the entire sample, we take a cross-fitting approach, partitioning the data into $K$ folds and switching the roles of the main and auxiliary datasets on each fold. We adapt the original cross-fitting algorithm for estimating equations (due to ChernozhukovChDeDuHaNeRo18) to estimating the WTE. Denoting the $k$-th fold $I_k$ and its complement $I_{k}^c = [n] \backslash I_k$, we fit nuisance parameters $(\what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})$ on the $k$-th auxiliary data $\{D_i\}_{i \in I_{k}^c}$. Using $\what{P}_k$ to denote the empirical distribution on the $k$-th main data $\{D_i\}_{i \in I_k}$, we summarize our procedure in Algorithm (ref). Our estimator can be computed in both randomized control trials using the true propensity score or in observational studies where $\what{e}_{k}$ needs to be estimated using suitable statistical models.
We can also derive natural analogues of the direct method (DM) and the inverse probability weighted estimator (IPW) for estimating $\mbox{WTE}_{\alpha}$ {
} Again, the above estimators reduce to their counterparts for estimating the ATE when $\alpha = 1$. In our subsequent analysis and experiments, we focus on the augmented estimator presented in Algorithm (ref) since unlike the two approaches above (ref), the augmented version satisfies Neyman orthogonality and achieves the semiparametric efficiency bound.
We empirically demonstrate how our worst-case sensitivity approach guards against spurious findings that are invalidated under shifts in the underlying population. Our WTE estimator (Algorithm (ref)) automatically detects subgroups adversely affected by the treatment and guarantees validity of findings over subpopulations on both randomized and observational studies. We observe that our WTE estimator remains stable even when typical CATE estimators---based on undersmoothed machine learning models---vary significantly across estimation methods and sample sizes.
Throughout our empirical analysis, we consider $\mbox{WTE}_{\alpha}$ for $\alpha \in \{ 1., .8, .6, .4, .2\}$ (recall $\mbox{ATE} = \mbox{WTE}_{1}$), where without mention we replace the supremum with an infimum in the definition (ref) if the desired sign of the treatment effect is positive. We use $K = 3$ in our cross-fitting procedure and use random forests to estimate outcome models with two-fold cross-validation. Our estimator bounds the usual cross-fitted AIPW ChernozhukovChDeDuHaNeRo18 and reduces to the AIPW when $\alpha = 1$, which we use as the topline estimator for the ATE. \ifdefined\usemsstyle \else \footnote{The code for all experiments can be found in \url{https://github.com/sookyojeong/worst-ate}.} \fi
\ifdefined\useectastyle \fi
Since the passage of the Affordable Care Act, Medicaid has expanded to 38 states in the U.S., aiming to increase healthcare coverage for low-income individuals. We study the effectiveness of Medicaid in increasing healthcare access as measured by the post-enrollment change in doctors' office utilization. Taking the viewpoint of an analyst in 2009, we illustrate how findings in 2009 (“present”) may no longer hold in the subsequent decade (“future”) due to unforeseen population shifts. Our worst-case sensitivity approach accounts for latent intersectionality using “present” data alone (Figure (ref)) and calls into question the external validity of present-day findings.
Using ten years of data (2009-18) from the National Health Interview Survey (NHIS), we focus on the binary outcome that indicates whether the individual made any visits to doctors in the two-weeks prior to the survey date CurrieGr96, LiptonDe15, CurrieFa05, KahendeMaEnZhMoXuSeRo17. Although survey respondents have limited control on the outcome variable as the date of survey depends on the availability of field agents, Medicaid enrollments are nevertheless non-random. Those in poor health and higher intention of seeking treatment are more likely to enroll in Medicaid and thus more likely to visit the doctors, leading to an upward bias in confounded estimates. To address the potential bias, we control for a rich set of baseline covariates ($d=396$) including demographics, medical history, employment, earnings, health limitations, whether they require help with daily tasks, and enrollment status in insurance and government programs. We posit that there are no unobserved confounders (Assumption (ref)); as this is a strong assumption, we also study a randomized experiment in the following subsection. We restrict attention to the Census West region to ensure uniformity in the experiences of the control group across different states. Although Medicaid eligibility cannot be determined based on the NHIS survey data (number of family members is missing), we restrict attention to those with annual income $\le$\$65K so that respondents have a positive probability of eligibility.
The $\mbox{WTE}_{\alpha}$ (ref) allows analyzing external validity over shifts in the study population. Using only data from 2009 ($n=82,993$ observations), in Figure (ref) we plot the cross-fitted estimates of $\mbox{WTE}_{\alpha}$ across a range of subpopulation sizes $\alpha$. While the ATE is positive in 2009 (significant at 99%), observed effects fail to be significant even over subpopulations that comprise $\alpha = 80\%$ of the study population. The WTE identifies subgroups disparately affected by the treatment and shows the average-case findings are invalidated under small changes to the study population. Our sensitivity framework thus brings into question whether the observed effects in 2009 can endure the test of time.
As predicted from 2009 data alone, treatment effects fail to be statistically significant over time (Figure (ref)a). To heuristically investigate the temporal population shift in the ATE, in Figure (ref)a we analyzed covariates with the largest mean difference between 2009 and 2018. The share of Hispanics and those who have been in the U.S. for less than 10 years decreased, but the share of U.S. citizens and those with high educational attainment increased. In Figure (ref)b, we observe that groups whose share decreased have higher treatment effects (and vice versa), contributing to the overall decrease in the treatment effect over time.
\ifdefined\useectastyle \fi
A large group of Americans harbors antipathy towards programs labeled “welfare,” a phenomenon that has generated much interest. We are interested in measuring how seemingly insubstantial wording changes in the description of social welfare programs affect public support. Disdain towards “welfare” has been associated with racist stereotypes towards welfare recipients HenryReWe04, Federico04 and political ideology KluegelSm86. Previous works have observed substantial heterogeneity with respect to variables such as level of racism, education levels, and political leanings Jacoby00, Federico04, GreenKe12.
We study an experiment on welfare attitudes in the General Social Survey (GSS) from 1986 to 2010 GreenKe12. We focus on the binary outcome indicating whether the respondents state too much is being spent on either “welfare” (treatment) or “assistance to the poor” (control). Both questions about public spending are identical except for the wording change. Covariates include attitude towards Blacks, political views, party identification, educational attainment, and age; there are $n = 20,783$ data points, and $d = 22$ covariates. To illustrate the flexibility of our estimation approach, we estimate the propensity score using a logistic regression with elastic net regularization as if it was unknown. (We observe similar results when we use the true propensity score $e\opt(X) \equiv \half$.)
By virtue of our semiparametric approach, we observe that even when CATE estimates vary significantly across sample size and outcome model classes, estimates of $\mbox{WTE}_{\alpha}$ align around a single value, a (empirical) stability property shared with estimators of ATE CarvalhoFeMuWoYe19. To illustrate the stability of WTE estimates, we plot them over different sample sizes (5-20K) and outcome model classes. Even at small sample sizes, wording changes in the survey have a resoundingly strong effect on attitude towards government welfare programs (Figure (ref)a). This observed effect is uniformly significant over subpopulations as small as 20% of the collected data. Such robust evidence instills confidence in the external validity of the finding across spatiotemporal changes in the demographic composition of respondents. While estimates of $\mbox{WTE}_{\alpha}$ and ATE remain relatively stable across different sample sizes, CATE estimates vary considerably over different sample sizes, a typical behavior for undersmoothed ML models (Figure (ref)b, c).
The stability of WTE estimates persists over different nuisance estimation approaches. Figure (ref) explores several common model classes for the conditional outcome $\E[Y(z) \mid X]$: linear models with elastic net regularizers, random forests, and gradient boosted regression trees. All hyperparameters are chosen using 2-fold cross-validation as before. Figure (ref)a highlights the stability of WTE and ATE estimates along different models. In Figures (ref)b and c, we observe that CATE estimates vary up to 25 times, especially around the endpoint of bins, a common phenomenon often attributed to bias at the boundaries of the support of the feature space WagerAt18. We anticipate that WTE estimators will similarly suffer instability issues when $\alpha$ is exceedingly small due to limited sample size.
We now show that our cross-fitted augmented estimator (Algorithm (ref)) enjoys central limit behavior $\frac{\sqrt{n}}{\what{\sigma}_{\alpha}} (\what{\omega}_{\alpha} - \mbox{WTE}_{\alpha}) \cd N(0, 1)$ even when we can only estimate the nuisance parameters (ref) at slower-than-parametric rates. The Neyman orthogonality condition Neyman59 serves a central role in our analysis.
We use the augmented form (ref) as our statistical functional $T(Q; \mu_0, \mu_1, e, h)$
where we use $\mu(x) \defeq \mu_1(x) - \mu_0(x)$ as usual. The first term is the dual form of $\mbox{WTE}_{\alpha}$ (Lemma (ref)), and the second term is the augmentation term defined in Eq. (ref). Since $\E_{D \sim P}[\kappa(D; \mu_0\opt, \mu_1\opt, e\opt, h\opt)]=0$ under ignorability (Assumption (ref)), we have $T(P; \mu_0\opt, \mu_1\opt, e\opt, h\opt) = \mbox{WTE}_{\alpha}$. \ifdefined\useectastyle An application of the envelope theorem confirms Neyman orthogonality of the augmented functional (ref). \else
To build intuition, we first informally argue that this augmented functional satisfies Neyman orthogonality. We first compute the (Gateaux) derivative of the first dual infimization term in the functional (ref). From Danskin's theorem BonnansSh00, the derivative of the dual formulation is the derivative of the objective at the unique optimal solution. Under sufficient regularity conditions, the unique solution to the dual is given by the quantile $P_{1-\alpha}^{-1}(\mu\opt)$, and the derivative at $r = 0$ is
To compute the derivative of the second term, let $\gamma = (\mu_0, \mu_1, e, h)$ to ease notation. So long as we can interchange derivatives and expectations, it is straightforward to calculate
from ignorability (Assumption (ref)). This verifies Neyman orthogonality of the functional (ref).
\fi
Orthogonality allows us to show a central limit theorem for our augmented cross-fitting estimator $\what{\omega}_{\alpha}$ under the following weak rate requirements for the nuisance parameters. Recall the bound $c >0$ on the propensity score given in Assumption (ref).
In particular, Assumption (ref) does not require a $\sqrt{n}$-convergence rate (Donsker condition) on the estimators of nuisance parameters. The first two conditions are standard convergence rates ChernozhukovChDeDuHaNeRo18, also required for proving a central limit result for the $\mbox{ATE}$. They hold, in particular, when $\Ltwo{\what{e} - e\opt} = o_p(n^{-1/4})$ and $\Ltwo{\what{\mu}_z - \mu\opt_z} = o_p(n^{-1/4})$ for $z \in \{0, 1\}$. The third condition guarantees approximation of the threshold function $h\opt(x) = \alpha^{-1} \indic{\mu\opt(x) \ge P_{1-\alpha}^{-1}(\mu\opt)}$ at suitably fast rates. The requirement $\Linf{\what{\mu}_{k} - \mu\opt} \le \delta_n n^{-1/3}$ states that the CATE be estimated at a somewhat faster rate compared to the case for estimating the ATE. \ifdefined\useectastyle \else In Appendix (ref), we provide detailed examples of model classes and learning methods where these convergence rates hold. \fi As noted in Section (ref), when unlabeled covariates (those without corresponding outcomes nor treatments) are cheaply available, the second part of $(c)$ can be easily achieved.
As the WTE is a tail-average of the CATE above the quantile $P_{1-\alpha}^{-1}(\mu\opt)$ (recall Lemma (ref)), to estimate $\mbox{WTE}_{\alpha}$ we need to estimate the quantile $P_{1-\alpha}^{-1}(\mu\opt)$. Towards this goal, we require that a positive density exists at its $(1-\alpha)$-quantile, a standard condition required for quantile estimation VanDerVaartWe96. Let $\mc{U}$ be a subset of measurable functions $\mu: \mc{X} \to \R$ such that the following holds:
We require this holds for our estimators $\mu = \what{\mu}_{k}$ with high probability.
In particular, Assumption (ref) requires $\mu\opt(X)$ to have a positive density at $P_{1-\alpha}^{-1}(\mu\opt)$.
We are now ready to give our main technical result which shows that the augmented cross-fitting estimator $\what{\omega}_{\alpha}$ enjoys central limit rates with the influence function
Indeed, $\psi(D)$ is a valid influence function under Assumption (ref): we have $\E[\psi(D)] = 0$, and $\var(\psi(D)) = \sigma^2_{\alpha}$ where $\sigma_{\alpha}^2$ is the asymptotic variance defined in expression (ref).
Since the proof of Theorem (ref) is involved, we give its sketch below, emphasizing how we leverage Neyman orthogonality of the augmented functional (ref) to obtain our result. We provide rigorous details of the proof in Appendices (ref), (ref), (ref). ChernozhukovChDeDuHaNeRo18 showed that the solution of a Neyman orthogonal estimating equation could be estimated at the usual central limit rate, even when nuisance parameters converge at slower rates. Their results do not apply to the $\mbox{WTE}_{\alpha}$ as it is a nonlinear statistical functional of the underlying probability measure $P$. To account for such nonlinearity, our theoretical analysis considerably extends existing results. Using tools from empirical process theory VanDerVaartWe96, our main result shows that for uniformly Hadamard differentiable functionals, orthogonality still allows insensitivity to estimation error in nuisance parameters. Even when $\alpha = 100\%$ so that $\mbox{WTE}_{1} = \mbox{ATE}$ and our estimator reduces to the cross-fitted AIPW for the ATE, our argument provides a different proof to that given by ChernozhukovChDeDuHaNeRo18.
Our proof proceeds in three parts. Recalling $\what{P}_k$, the empirical distribution on the $k$-th fold, our cross-fitted estimator can be written succinctly as $\frac{1}{K} \sum_{k = 1}^{K} T(\what{P}_k; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})$. We emphasize (empirical) expectations over $Q$ in the definition (ref) are taken only over $D$, and not over the randomness in $(\what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})$.
The first two parts show that for each fold $k \in [K]$, \ifdefined\useectastyle {
} \else
\fi Since $T(P; \mu_0\opt, \mu_1\opt, e\opt, h\opt) = \mbox{WTE}_{\alpha}$ by ignorability (Assumption (ref)), this gives our first result. Towards this goal, decompose the left hand side of the equality (ref) into
In Part I of the proof (Appendix (ref)), we prove that the first term (ref) is asymptotically equal to $\frac{1}{\sqrt{|I_k|}} \sum_{i \in I_k} \psi(D_i)$. Leveraging tools from empirical process theory VanDerVaartWe96, we first show that the functional $Q \mapsto T(Q; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})$ satisfies a uniform variant of Hadamard differentiability so that the functional delta method applies. A careful application of a uniform version of the Lindeberg-Feller central limit theorem over functions gives the desired conclusion.
In Part II (Appendix (ref)), we use Neyman orthogonality of our augmented estimator to show that the second term (ref) vanishes asymptotically. Define $\mathfrak{R}_k: [0, 1] \to \R$ \ifdefined\useectastyle {
} \else
\fi so that $\mathfrak{R}_k(0) = 0$, and $\mathfrak{R}_k(1)$ is equal to the second term (ref). $\mathfrak{R}_k(r)$ is continuously differentiable under suitable conditions and mean value theorem gives
for some $r \in [0, 1]$. From Neyman orthogonality, we have $\mathfrak{R}_k'(0) = 0$. Building on this, we show that all values of $\mathfrak{R}_k'(r)$ are sufficiently small: under postulated convergence rates for the nuisance parameters in Assumption (ref), $\sup_{r \in [0, 1]} |\mathfrak{R}_k'(r)| = o_p(n^{-1/2})$.
In Part III (Appendix (ref)), we show consistency of our variance estimator: $\what{\sigma}^2_{\alpha} \cp \sigma^2_{\alpha}$. Combining this with the central limit result (ref), Slutsky's lemma gives our second result. $\diamond$
We now establish a semiparametric efficiency bound, showing that all (regular) estimators of $\mbox{WTE}_{\alpha}$ necessarily have asymptotic variance larger than $\sigma_{\alpha}^2$, both when the true propensity score $e\opt(\cdot)$ is known and unknown. In particular, this implies that our augmented estimator $\what{\omega}_{\alpha}$ achieves the optimal asymptotic variance and that its influence function (ref) is the efficient influence function for estimating $\mbox{WTE}_{\alpha}$. Our efficiency bound i) informs the design of experiments by providing the minimum number of study participants required to reach a conclusion that is valid across subpopulations no smaller than $\alpha$ and ii) quantifies how the required sample size grows with the level of desired external validity $\alpha$.
For parametric problems, Hajek-Le Cam theorems VanDerVaart98 give a lower bound on the asymptotic variance of regular estimators, more generally, a lower bound on the mean squared error for any estimator. Since these bounds coincide with the Cramer-Rao bound for unbiased estimators, they can be considered as an asymptotic Cramer-Rao bound. We consider parametric submodels of our semiparametric problem, i.e., finite-dimensional parameterizations that contain the truth. Since the asymptotic variance of any (smooth enough) semiparametric estimator is worse than the Hajek-Le Cam bound---equivalently, the Cramer-Rao bound---of any parametric submodel, the semiparametric efficiency bound is defined as the supremum of the Hajek-Le Cam bound over all parametric submodels. As these definitions are standard yet tedious Newey90, BickelKlRiWe93, we defer a formal treatment to Appendix (ref).
We now characterize the semiparametric efficiency bound for estimating $\mbox{WTE}_{\alpha}$.
See Appendix (ref) for the proof. Theorem (ref) and (ref) show that our cross-fitted augmented estimator $\what{\omega}_{\alpha}$ achieves the semiparametric efficiency bound $\mathbb{V}(\mbox{WTE}_{\alpha}; P)$. Similar to the efficiency bound for the ATE Hahn98, the knowledge of $e\opt(\cdot)$ does not affect the efficiency bound $\mathbb{V}(\mbox{WTE}_{\alpha}; P)$. We conclude that our estimator is semiparametrically efficient for both observational studies and randomized control trials.
Our semiparametric efficiency bound allows calculating the minimum required sample size for finding conclusions that are robust against all subpopulations of size at least $\alpha$. Consider the hypothesis test $H_0: \mbox{WTE}_{\alpha} \ge 0 ~\mbox{v.}~H_{\rm A}: \mbox{WTE}_{\alpha} < -\epsilon$, where $\epsilon > 0$ is the specified minimum detectable effect size (recall that the desired sign of the treatment effect is negative). As an illustration, consider a test with size ($\P(\mbox{Type I error})$) at most $5\%$ and power ($1 -\P(\mbox{Type II error})$) at least $80\%$. A standard power calculation shows $n \ge 6.2 \frac{\sigma_{\alpha}^2}{\epsilon^2}$ samples are needed to detect a worst-case subpopulation treatment effect of size $-\epsilon$. In particular, this number grows as we require a stronger level of external validity, or equivalently as the worst-case subpopulation size $\alpha$ becomes small.
Motivated by challenges in evaluating treatments under unanticipated population shifts, we proposed a sensitivity analysis framework based on the worst-case subpopulation treatment effect. We advocate for such conservatism in important policy decisions that need to benefit all subpopulations uniformly. While the WTE (ref) provides a strong notion of robustness, it may be overly conservative in scenarios where one is concerned with more structured covariate shifts. When covariate shift on only a small subset of $X$ is of interest, the definition (ref) can be modified over this subset. More broadly, studying structured shifts in the covariate distribution $P_X$ is an exciting direction of future work.
Our theoretical developments show that our cross-fitted augmented estimator inherits two advantageous inferential properties of the AIPW estimator for the ATE: orthogonality and efficiency. Elementary derivations show, however, that our estimator does not satisfy the doubly robust property due to the additional nuisance parameter $h\opt(x) = \alpha^{-1} \indic{\mu\opt(x) \ge P_{1-\alpha}^{-1}(\mu\opt)}$. Another limitation of our augmented estimator is that it is not necessarily increasing in $\alpha$. Though the extension of the direct method and the inverse probability weighted estimator are increasing in $\alpha$, it is unclear whether an orthogonal, efficient, and monotone estimator of the WTE exists. Finally, estimators of tail-averages often suffer higher variance and may be more sensitive to lack of overlap and unobserved confounding. Further investigation of these issues is a topic for future research.
Sometimes it is of interest to estimate the average treatment effect on the treated $\E[Y(1) - Y(0) \mid Z = 1]$. There are two natural ways to extend the worst-case definition (ref) to measure the effect on the treated, depending on whether the worst-case is still taken over $Q_X$, or over the conditional distribution $Q_{X \mid Z = 1}$. We leave a systematic study of the two definitions and corresponding inferential frameworks to future work.
\paragraph*{Acknowledgments} We thank Steve Yadlowsky for helpful comments.
{.7em}
\ifdefined\usemsstyle
\ECSwitch
\ECHead{Appendix}
\else