EconBase
← Back to paper

Assessing External Validity Over Worst-case Subpopulations

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

60,605 characters

Assessing External Validity Over Worst-case Subpopulations





\RUNTITLE{Assessing External Validity Over Worst-case Subpopulations}

\TITLE{Assessing External Validity Over Worst-case Subpopulations}

\ARTICLEAUTHORS{
  \AUTHOR{Sookyo Jeong}
\AFF{Lyft Inc., \EMAIL{[email removed]}}

\AUTHOR{Hongseok Namkoong}
\AFF{Decision, Risk, and Operations Division, Columbia Business School, New York, NY 10027, \EMAIL{[email removed]}} , \URL{hsnamkoong.github.io}
}

\ABSTRACT{Study populations are typically sampled from limited points in space and time,
and marginalized groups are underrepresented. To assess the external validity
of randomized and observational studies, we propose and evaluate the
worst-case treatment effect (WTE) across all subpopulations of a given size,
which guarantees positive findings remain valid over subpopulations. We
develop a semiparametrically efficient estimator for the WTE that analyzes the
external validity of the augmented inverse propensity weighted estimator for
the average treatment effect.  Our cross-fitting procedure leverages flexible
nonparametric and machine learning-based estimates of nuisance parameters and
is a regular root-$n$ estimator even when nuisance estimates converge more
slowly. On real examples where external validity is of core concern, our
proposed framework guards against brittle findings that are invalidated by
unanticipated population shifts.



}


\KEYWORDS{external validity, distributional robustness, semiparametrics}
\HISTORY{This paper was first submitted on Jan, 2022.}

\maketitle



\else


\documentclass[11pt]{article}
\usepackage[numbers]{natbib}
\usepackage{fullpage}
\usepackage{./statistics-macros}
\usepackage{setspace}

\usepackage{hyperref}
\usepackage{pgfplotstable}
\usepackage{graphicx}
\usepackage{subcaption}
\usepackage{float}

\usepackage{algorithm}
\usepackage{algpseudocode}
\usepackage{tabularx}

\usepackage{overpic}
\usepackage{tikz}
\usepackage{rotating}
\usepackage{psfrag}

\usepackage{url}


\makeatletter
\long\def\@makecaption#1#2{
  \vskip 0.8ex
  \setbox\@tempboxa\hbox{\small {\bf #1:} #2}
  \parindent 1.5em
  \dimen0=\hsize
  \advance\dimen0 by -3em
  \ifdim \wd\@tempboxa >\dimen0
  \hbox to \hsize{
    \parindent 0em
    \hfil
    \parbox{\dimen0}{\small
      {\bf #1.} #2
    }
    \hfil}
  \else \hbox to \hsize{\hfil \box\@tempboxa \hfil}
  \fi
}
\makeatother

\begin{document}
\abovedisplayskip=8pt plus0pt minus3pt
\belowdisplayskip=8pt plus0pt minus3pt


\begin{center}
  {\LARGE Assessing External Validity Over  Worst-case Subpopulations\footnote{An extended abstract for an earlier version of this work appeared at
  the Conference in Learning Theory 2020 entitled ``Robust Causal Inference
  Under Covariate Shift via Worst-Case Subpopulation Treatment Effects''.}} \\
  \vspace{.5cm}
  {\Large Sookyo Jeong$^{1}$ ~~~~  Hongseok Namkoong$^{2}$ } \\
  \vspace{.2cm} {\large
    $^1$Lyft Inc. \\
    \vspace{.1cm}   $^2$Decision, Risk, and Operations Division, Columbia Business School
    \\ \vspace{.1cm} }
  \vspace{.2cm}
  {\tt [email removed], [email removed]}
\end{center}




\begin{abstract}

\end{abstract}

\fi





\definecolor{innerboxcolor}{rgb}{.9,.95,1}
\definecolor{outerlinecolor}{rgb}{.6,0,.2}


















\section{Introduction}
\label{section:introduction}

When the study population is different from those affected by the treatment,
the external validity of a study's finding may be called into question. Study
populations are often sampled from a particular set of points in space and
time and may not represent future populations of interest~\citep{CampbellSt63,
  Manski13, RosenzweigUd20, DehejiaPoSa21}. Furthermore, study populations
often lack diversity, and minority groups are underrepresented. For example,
out of $10,000+$ cancer clinical trials funded by the National Cancer
Institute, less than 5\% of participants were non-white~\citep{ChenLaDaPaKe14,
  ShenoyHa15}.
When the treatment effect is heterogeneous, both randomized and observational
studies lose external validity outside the study population. While large-scale
randomized trials offer a ``gold standard'' for internal validity, their
external validity can be nevertheless called into question over spatiotemporal
changes in the population~\citep{Deaton10, BasuSuHa17}.



\begin{figure}[t]
  \centering
  \begin{tabular}{cc}
      \includegraphics[scale=.3]{./figures/wte_nhis_intersectionality.png}
      &
          \includegraphics[scale=.3]{./figures/wte_nhis_share.png}
    \\
    \small{(a) Intersectionality of treatment effects (2009)}

      & \small{(b) Change in shares from 2009 to 2013 \& 2018 }
  \end{tabular}
  \vspace{10pt}
  \caption{
    \label{fig:intro} Effect of Medicaid enrollment on doctors' visit for
    low-income adults (NHIS) } \end{figure}

Existing approaches for assessing external validity require the knowledge of
the target population~\citep{HotzImMo05, AngristFe10, ColeSt10,
  StuartCoBrLe11, Tipton13, LeskoBuWeEdHuCo17, AndrewsOs17,
  Meager19}. However, target populations chosen at the time of analysis may
not be sufficient to guarantee external validity over unforeseen shifts in the
population that occur post-analysis.  Heuristic approaches such as estimating
treatment effects over fixed subgroups are similarly limited as effects
typically vary over a combination of multiple characteristics like race,
gender, age, and income, a phenomenon we refer to as \emph{intersectionality}.

To illustrate these challenges, consider estimating the effect of Medicaid
enrollment on doctors' office utilization based on the National Health
Interview Survey (NHIS) in 2009, where we focus on whether the individual made
any visits to doctors two weeks prior to the survey date as the main
outcome. Observed treatment effects and resulting decisions in 2009 must
remain valid over \emph{a priori unknown shifts} in the population over time.
We observe substantial intersectionality in the treatment effect
(Figure~\ref{fig:intro}a): for Hispanic U.S.  citizens, the treatment effect
changes signs depending on professional degree attainment.  In the decade
following 2009 (``future''), there is a major shift in the underlying
population (Figure~\ref{fig:intro}b), and observed effects in 2009 are no
longer valid in the future, as we later illustrate in
Section~\ref{section:experiments}.




To assess external validity over unanticipated population shifts in the
population, we propose and study the worst-case treatment effect,
$\mbox{WTE}_{\alpha}$, defined over \emph{all subpopulations} that comprise at
least $\alpha$-fraction of the study population (see Eq.~\eqref{eqn:wte} to
come).  As the $\mbox{WTE}_{\alpha}$ bounds the average treatment effect (ATE)
and reduces to it when $\alpha=100\%$, the WTE analyzes the sensitivity of an
average-case finding under population shifts, guaranteeing that conservative
findings using the WTE remain valid uniformly over subpopulations. For
example, if low-income Hispanic U.S. citizens with professional degrees
comprise at least $20\%$ of the study population,  positive findings with
respect to $\mbox{\rm WTE}_{.2}$ guarantee the treatment remains effective
over this subgroup.








We develop a semiparametrically efficient estimator of
$\mbox{\rm WTE}_{\alpha}$, analyzing the external validity of the augmented
inverse propensity weighted (AIPW) estimator~\citep{RobinsRoZh94, RobinsRo95}
for the ATE. Our $K$-fold cross-fitting procedure leverages machine
learning-based and nonparametric estimators of nuisance parameters and
provides a worst-case bound on the cross-fitted AIPW for the
ATE~\citep{ChernozhukovChDeDuHaNeRo18}; our estimator reduces to the AIPW when
$\alpha = 100\%$. On real datasets where external validity is of core concern,
our worst-case sensitivity approach identifies disadvantaged subpopulations
based on a priori nontrivial demographic groupings and guards against brittle
findings that are invalidated under population shifts
(Section~\ref{section:experiments}).


Specifically, we exploit the dual representation of the worst-case over
subpopulations to derive our semiparametric estimator
(Section~\ref{section:approach}).  By virtue of satisfying an orthogonality
property (similar to the AIPW for the ATE), under standard assumptions
required for the identification and estimation of the ATE, our augmented
estimator of $\mbox{WTE}_{\alpha}$ enjoys central limit rates even when
estimates of the nuisance parameters converge at slower-than-parameteric rates
(Section~\ref{section:asymptotics}).  Since $\mbox{WTE}_{\alpha}$ is nonlinear
in the underlying probability measure, our main asymptotic result
(Theorem~\ref{theorem:clt}) requires a novel theoretical analysis different from
the estimating equations framework (method of moments) studied
by~\citet{ChernozhukovChDeDuHaNeRo18}.

We prove that our augmented estimator for the $\mbox{WTE}_{\alpha}$ is
semiparametrically efficient, in both observational and randomized studies
(Section~\ref{section:efficiency}). Our semiparametric efficiency bound
informs experimental design with external validity as a central
concern.
Power calculations based on our efficiency bound provide the minimal sample
size required to detect a specified effect size for the
$\mbox{WTE}_{\alpha}$. Our bounds quantify how testing external validity
against smaller subpopulations ($\alpha$) requires a correspondingly larger
sample sizes.




\paragraph*{Related work}

Assessing the external validity of randomized and observational studies is an
active area of research in causal inference~\citep{DahabrehHe19}.
When the target population is \emph{known} but different from the study
population, many authors have leveraged the relationship between the two
populations to guarantee external validity. A prevalent approach is to view
selection into the study as another ``treatment'', adjusting estimates of the
ATE based on some information about the target population~\citep{ColeSt10,
  StuartCoBrLe11, Tipton13, KernStHiGr16, LeskoBuWeEdHuCo17, AndrewsOs17,
  Meager19}.~\citet{StuartCoBrLe11} and~\citet{Tipton13} use the probability
of being included in the study to adjust for population bias, assuming that
sample selection decisions only depend on observed covariates.
\citet{HotzImMo05} apply bias-corrected matching methods to predict the impact
of a program by using observations collected from a different location. In the
context of structural causal models, Bareinboim, Pearl and colleagues identify
settings that allow external validity in a series of
works~\citep{BareinboimPe12, BareinboimPe16}.


External validity is of particular concern when identification strategies only
allow studying a local notion of treatment effect. In such scenarios, several
authors aim to connect estimates for a local population to a (known) broader
population by leveraging a postulated structure between the two
populations. For instrumental variable strategies, there is a line of work
(see, for example,~\citep{AngristImRu96, Angrist04, AngristFe10}) studying
when the treatment effects for compliers (LATE) can inform effects for a
broader population. Most recently,~\citet{RosenzweigUd20} directly estimate
external validity over time when exogenous aggregate shocks are observed. For
regression discontinuity designs,~\citet{DongLe15, AngristRo15, BertanhaIm20}
analyze settings where local estimates are externally valid and develop
corresponding statistical tests.






Compared to the above methods, our worst-case sensitivity approach does not
assume knowledge of the target population. Our conservative approach is
agnostic to the unknown shifts in the population and provides uniform
guarantees over subpopulations comprising at least $\alpha$-fraction of the
study population. This is conceptually related to recent works on
distributionally robust optimization in operations research and supervised
learning, where models are trained to optimize a worst-case loss over
distribution shifts~\citep{DuchiHaNa20}. Relatedly,~\citet{BoGa21} studied
external validity as a notion of stability of the joint distribution between
the observed outcome and treatment assignments.

Study populations must be designed to be as diverse as possible across
demographics, space, and time.  Our approach can guarantee meaningful external
validity only if the study includes heterogeneous subpopulations. (When the
study population is not representative, our approach can nevertheless raise
alarms.)  A \emph{design-based} approach to external validity complements our
sensitivity framework by promoting diversity in the study population as a
central concern. Several works in development and labor economics aim to
improve the external validity of a (quasi-) experiment by collecting data over
multiple sites and temporal points~\citep{CrucesGa07, BanerjeeKaZi15,
  GertlerShAlCaMaPa15, DupasKaRoUb18, RosenzweigUd20, DehejiaPoSa21}.  Tipton
and colleagues develop methodologies for measuring the diversity of a study
population alongside practical experimental design guidelines~\citep{Tipton14,
  TiptonPe17, TiptonRo18}.


Our worst-case approach is broadly related to previous works that estimate
treatment effects beyond mean differences~\citep{Rothe10,
  KimKiKe18}.~\citet{ChernozhukovFeLu18} study \emph{sorted effects}, a
collection of sorted quantiles of the conditional average treatment effect
(CATE). They develop central limit results for this nonparametric estimand
under Donsker conditions (i.e., functional CLT) on CATE estimates. The
$\mbox{WTE}_{\alpha}$ we introduce in Section~\ref{section:approach} is a
tail-average of sorted effects, and our semiparametric approach extends the
AIPW under unanticipated shifts in the population.  Theoretically, we prove
central limit results for our estimator \emph{without requiring Donsker
  conditions on CATE estimates}. Our approach is not to be confused with
quantile treatment effects~\citep{Firpo07}, which measures the difference
between quantiles of $Y(1)$ and $Y(0)$.























\section{Approach}
\label{section:approach}






Using the potential outcomes notation to denote counterfactuals, we let $Y(1)$
and $Y(0)$ be outcomes corresponding to treatment and control and let
$Z \in \{0, 1\}$ be the assigned treatment~\citep{Rubin74}.  We study both
randomized and observational studies, assuming the analyst has access to
observed covariates $X \sim P_X$. A standard goal is to estimate the
\emph{average treatment effect},
$\mbox{ATE} \defeq \E[Y(1) - Y(0)] = \E_{X \sim P_X}\left[\mu\opt(X)\right]$,
where $\mu\opt(X) \defeq \E[Y(1) - Y(0)|X]$ is the \emph{conditional average
  treatment effect (CATE)}.  Throughout, we assume the distribution
$Y(1), Y(0) \mid X$ remains unchanged over subpopulations.

As the study population $P_X$ may not be representative of those affected by
the treatment, we are interested in measuring the sensitivity of a study's
finding to shifts in the underlying population. We consider the set of all
subpopulations (probabilities) $Q_X$ that comprise more than
$\alpha \in (0, 1]$ fraction of the study population $P_X$
\begin{equation}
  \label{eqn:subpopulations}
  \mc{Q}_\alpha \defeq \left\{Q_X \mid
    P_X = a Q_X + (1-a) Q_X'~\mbox{for some}~a \ge \alpha,~\mbox{and subpopulation}~Q_X'
  \right\}.
\end{equation}
As a convention, we assume the desired sign of the treatment effect is
negative (the positive case is symmetric).  We propose and study the worst-case
subpopulation treatment effect
\begin{equation}
  \label{eqn:wte}
  \mbox{\rm WTE}_{\alpha} \defeq \sup_{Q_X \in \mc{Q}_{\alpha}} \E_{X \sim
    Q_X}\E[Y(1) - Y(0)| X].
\end{equation}
When treatment effects are highly heterogeneous and external validity is of
particular concern, the worst-case subpopulation treatment
effect~\eqref{eqn:wte} will be substantially different from the ATE. The
worst-case bound $\mbox{WTE}_{\alpha}$ reduces to the ATE when
$\alpha = 100\%$.





The modeling choice of $\alpha$ in the definition~\eqref{eqn:wte} is
important, and it should be informed by domain knowledge. When a study
population is not representative, we recommend selecting a smaller value of
$\alpha$. For example, the analyst may reason about the level of bias
anticipated in the data collection process or use the size of proxy target
groups in the study.  In the latter case, positive findings with the chosen
level of $\alpha$ guarantee uniformly valid treatment effects over all
minority subpopulations of size $\alpha$, not just the proxy targets.  The
choice of $\alpha$ should also consider the data size: as $\alpha$ becomes
small, inference becomes difficult, as our semiparametric efficiency bounds
demonstrate in Section~\ref{section:efficiency}. Even when the analyst does
not commit to a single level of $\alpha$, evaluating the WTE over a range of
$\alpha$'s can offer a practical diagnostic. The level of $\alpha$ at which
$\mbox{WTE}_{\alpha}$ crosses a threshold (e.g., 0) is often of particular
practical interest as it represents the smallest subpopulation size over which
average-case findings remain valid.


To derive our augmented estimator, we begin by simplifying the primal
problem~\eqref{eqn:wte} over (infinite-dimensional) covariate distributions
$Q_X \in \mc{Q}_{\alpha}$ to its dual representation over a one-dimensional
threshold on the CATE $\mu\opt(X) = \E[Y(1) - Y(0) \mid X]$. The dual
reformulation shows an equivalence between worst-case subpopulation
performances and tail-averages. We rely on this relationship heavily to derive
our augmented estimator and to prove its asymptotic properties.
We make the dependence on the underlying probability explicit and write
$\E_Q[X]$, except for when $Q = P$, the data-generating distribution.  The
following lemma is a consequence of~\citet[Example 6.19]{ShapiroDeRu09}.
\begin{lemma}
  \label{lemma:dual}
  Let $P_{1-\alpha}^{-1}(\mu\opt)$ be the $(1-\alpha)$-quantile of $\mu\opt(X)$, and denote
  $\hinge{\cdot} \defeq \max(\cdot, 0)$ and
  $h\opt(x) \defeq \frac{1}{\alpha} \indic{\mu\opt(x) \ge P_{1-\alpha}^{-1}(\mu\opt)}$. If
  $\E[\mu\opt(X)_+] < \infty$, then
  {\small
  \begin{align*}
    \mbox{WTE}_{\alpha}
    = \inf_{\eta \in \R} \left\{
    \frac{1}{\alpha} \E\hinge{\mu\opt(X) - \eta} + \eta
      \right\}
    = \E \left[ \mu\opt(X) \mid \mu\opt(X) \ge
    P_{1-\alpha}^{-1}(\mu\opt) \right] = \E[\mu\opt(X) h\opt(X)].
  \end{align*}
  }
\end{lemma}

The dual optimum is attained at $P_{1-\alpha}^{-1}(\mu\opt)$ giving the second equality. This
tail-average is known as the conditional value-at-risk (CVaR), a common risk
measure in portfolio optimization~\citep{RockafellarUr00}. In contrast to the
$\mbox{WTE}_{\alpha}$ involving an unknown nuisance parameter $\mu\opt(X)$ that
needs to be estimated, the CVaR is typically considered over an observable
random variable---this gives rise to a salient semiparametric structure. The
dual shows the worst-off subpopulation is given by those who get
disproportionately and adversely affected by the treatment, measured by $X$
such that $\mu\opt(X) \ge P_{1-\alpha}^{-1}(\mu\opt)$.  To illustrate how the WTE~\eqref{eqn:wte}
accounts for heterogeneity across subpopulations, consider
$\mu\opt(X) \sim N(-.1, 1)$ so that there is substantial heterogeneity across
covariates. Although the $\mbox{ATE} = -.1$ suggests a negative treatment
effect,
$\mbox{WTE}_{.9} = 0.184$, meaning the treatment effect goes in the reverse
direction of the ATE for even $90\%$ of the study population.

To identify causal effects, we assume no unobserved confounding and overlap
between the treated and control groups.
\begin{assumption}
  \label{assumption:ig}
  Ignorability~ $Y(0), Y(1) \indep ~Z \mid X$
\end{assumption}
\begin{assumption}
  \label{assumption:overlap}
  Overlap: There exists $c> 0$ such that  $\P(e\opt(X) \in [c, 1-c]) = 1$.
\end{assumption}
\noindent We also assume that units do not interact with each other (no
interference), and that we observe i.i.d. units $D_i = (X_i, Y_i, Z_i)$ for
$i = 1, \ldots, n$ (stable unit treatment value
assumption~\citep{Rubin80}). Finally, we require the following standard
condition that uniformly bounds the conditional variance of the residuals for
$z \in \{0, 1\}$.
\begin{assumption}
  \label{assumption:residuals}
  Bounded residuals
  {\small $\E[Y(z)^2] + \linfstatnorm{\E[ (Y(z) - \mu_z\opt(X))^2 \mid X]} < \infty$}
\end{assumption}
\noindent These standard assumptions are also required to identify and
estimate the ATE~\citep{ImbensRu15, ChernozhukovChDeDuHaNeRo18}.

Recalling that $h\opt$ is a nuisance parameter determining the worst-off
subpopulation in Lemma~\ref{lemma:dual}, we consider the following key nuisance parameters
\begin{equation}
  \label{eqn:nuisance-def}
  \begin{split}
  \mbox{outcome models}~& ~\mu\opt_z(x) \defeq \E[Y(0) \mid X =x, Z=z]~~\mbox{for}~~z \in \{0,1\},\\
  \mbox{propensity score}~& ~e\opt(x) \defeq \P(Z = 1 \mid X = x) \\
  \mbox{threshold function}~& ~h\opt(x) \defeq \frac{1}{\alpha} \indic{\mu\opt(x) \ge P_{1-\alpha}^{-1}(\mu\opt)}.
  \end{split}
  \end{equation}
Letting $D = (X, Y, Z)$ be the tuple of observed data and
$(\mu_0, \mu_1, e, h)$ be the tuple of nuisance parameters, we consider the
augmentation term
\begin{equation}
  \label{eqn:aug}
  \kappa(D; (\mu_0, \mu_1, e, h)) \defeq
  h(X) \left(
    \frac{Z}{e(X)} (Y - \mu_1(X)) - \frac{1-Z}{1-e(X)} (Y - \mu_0(X))
  \right).
\end{equation}
Under the stated assumptions, we have $\E[\kappa(D; \mu_0\opt, \mu_1\opt, e\opt, h\opt)] = 0$. Instead of
estimating $\mbox{WTE}_{\alpha}$, we estimate the augmented form
$\mbox{WTE}_{\alpha} + \E[\kappa(D; \mu_0\opt, \mu_1\opt, e\opt, h\opt)]$. When $\alpha = 1$ so that
$\mbox{WTE}_{1} = \mbox{ATE}$, our estimator reduces to the augmented inverse
probability weighted (AIPW) estimator for the ATE. Thus, our estimator can be
viewed as an extension of the AIPW estimator under shifts in the underlying
population.

\ifdefined\usemsstyle
\begin{algorithm}[t]
  \caption{\label{alg:cross-fitting} Cross-fitting for $\mbox{WTE}_{\alpha}$}
  \begin{algorithmic}[1]
    \STATE \textsc{Input: $K$-fold partition $\cup_{k=1}^K I_k = [n]$ of
     $\{(X_i, Y_i, Z_i)\}_{i=1}^n$ s.t. $|I_k| = \frac{n}{K}$}
   \STATE \textbf{\textsc{For}} $k \in [K]$
   \STATE \hspace{10pt} \textbf{Estimate nuisance parameters}
   Using the data $\{D_i\}_{i \in I_{k}^c}$, fit estimators
   \STATE \hspace{10pt} 1. $\what{\mu}_{z, k}(\cdot)$ of
      $\mu_z\opt(\cdot) = \E[Y(z) \mid X = \cdot, Z = z]$ for $z \in \{0, 1\}$
      \STATE \hspace{10pt} 2. $\what{e}_{k}(\cdot)$ of $e\opt(\cdot) = \P(Z = 1 \mid X =\cdot)$
      \STATE \hspace{10pt} 3. $\what{h}_{k}(x) \defeq \frac{1}{\alpha} \indic{\what{\mu}_{k}(x) \ge \what{q}_{k}}$, where
      $\what{q}_{k}$ is an estimator of $P_{1-\alpha}^{-1}(\mu\opt)$
    \STATE \hspace{10pt} \textbf{Compute augmented estimator} Using the data
    $\{D_i = (X_i, Y_i, Z_i) \}_{i \in I_k}$, compute
    \begin{align*}
      \what{\omega}_{\alpha, k}
      & \defeq \inf_{\eta}
        \left\{ \frac{1}{\alpha} \E_{X \sim \what{P}_k}
        \hinge{\what{\mu}_{k}(X) - \eta} + \eta  \right\}
        + \E_{D \sim \what{P}_k}\left[\kappa\left(D; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k}\right)\right] \\
      \what{\sigma}^2_{\alpha, k}
      & \defeq \frac{1}{\alpha^2} \var_{X \sim \what{P}_k} \hinge{\what{\mu}_{k}(X) - \what{q}_{k}}
        + \var_{D \sim \what{P}_k} \left( \kappa\left(D; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k}\right)\right)
    \end{align*}
    \STATE \textbf{\textsc{Return}} Estimator
    $\what{\omega}_{\alpha} = \frac{1}{K} \sum_{k \in [K]}
    \what{\omega}_{\alpha, k}$, and variance estimate
    $\what{\sigma}^2_{\alpha} = \frac{1}{K} \sum_{k \in [K]}
    \what{\sigma}^2_{\alpha, k}$
  \end{algorithmic}
\end{algorithm}
\else
\begin{algorithm}[t]
  \caption{\label{alg:cross-fitting} Cross-fitting for $\mbox{WTE}_{\alpha}$}
  \begin{algorithmic}[]
    \State \textsc{Input: $K$-fold partition $\cup_{k=1}^K I_k = [n]$ of
     $\{(X_i, Y_i, Z_i)\}_{i=1}^n$ s.t. $|I_k| = \frac{n}{K}$}
   \State \textbf{\textsc{For}} $k \in [K]$
   \State \hspace{10pt} \textbf{Estimate nuisance parameters}
   Using the data $\{D_i\}_{i \in I_{k}^c}$, fit estimators
   \State \hspace{10pt} 1. $\what{\mu}_{z, k}(\cdot)$ of
      $\mu_z\opt(\cdot) = \E[Y(z) \mid X = \cdot, Z = z]$ for $z \in \{0, 1\}$
      \State \hspace{10pt} 2. $\what{e}_{k}(\cdot)$ of $e\opt(\cdot) = \P(Z = 1 \mid X =\cdot)$
      \State \hspace{10pt} 3. $\what{h}_{k}(x) \defeq \frac{1}{\alpha} \indic{\what{\mu}_{k}(x) \ge \what{q}_{k}}$, where
      $\what{q}_{k}$ is an estimator of $P_{1-\alpha}^{-1}(\mu\opt)$
    \State \hspace{10pt} \textbf{Compute augmented estimator} Using the data
    $\{D_i = (X_i, Y_i, Z_i) \}_{i \in I_k}$, compute
    \begin{align*}
      \what{\omega}_{\alpha, k}
      & \defeq \inf_{\eta}
        \left\{ \frac{1}{\alpha} \E_{X \sim \what{P}_k}
        \hinge{\what{\mu}_{k}(X) - \eta} + \eta  \right\}
        + \E_{D \sim \what{P}_k}\left[\kappa\left(D; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k}\right)\right] \\
      \what{\sigma}^2_{\alpha, k}
      & \defeq \frac{1}{\alpha^2} \var_{X \sim \what{P}_k} \hinge{\what{\mu}_{k}(X) - \what{q}_{k}}
        + \var_{D \sim \what{P}_k} \left( \kappa\left(D; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k}\right)\right)
    \end{align*}
    \State \textbf{\textsc{Return}} Estimator
    $\what{\omega}_{\alpha} = \frac{1}{K} \sum_{k \in [K]}
    \what{\omega}_{\alpha, k}$, and variance estimate
    $\what{\sigma}^2_{\alpha} = \frac{1}{K} \sum_{k \in [K]}
    \what{\sigma}^2_{\alpha, k}$
  \end{algorithmic}
\end{algorithm}
\fi

We now formally define our cross-fitted augmented estimator
$\what{\omega}_{\alpha}$ and an estimate $\what{\sigma}^2_{\alpha}$ of its
asymptotic variance
\begin{align}
  \label{eqn:var}
  \sigma^2_{\alpha}
  & \defeq \frac{1}{\alpha^2} \var\left(\hinge{\mu\opt(X) - P_{1-\alpha}^{-1}(\mu\opt)} \right)
    + \var\left( \kappa\left(D; \mu_0\opt, \mu_1\opt, e\opt, h\opt\right) \right).
\end{align}
As we show in Section~\ref{section:asymptotics}, these estimates give an
asymptotically exact confidence interval
$\P(\mbox{WTE}_{\alpha} \in [\what{\omega}_{\alpha} \pm z_{\delta}
\what{\sigma}_{\alpha} / \sqrt{n}]) \to 1-\delta$ if we set $z_{\delta}$ to be
the $(1 - \delta/2)$-quantile of a standard normal distribution. The
asymptotic variance~\eqref{eqn:var} is the best attainable in the typical
semiparametric sense as our semiparametric efficiency bounds in
Section~\ref{section:efficiency} show.


We fit estimators of nuisance parameters on the auxiliary sample, and combine
them via the augmented dual form~\eqref{eqn:aug} to evaluate the treatment
effect on the worst-off subpopulation. Our approach is agnostic to the
nuisance estimation method, and in particular, allows flexible use of machine
learning models and nonparametric techniques to estimate $\mu\opt_z$ and
$e\opt$. To estimate the threshold function $h\opt(X)$ that determines the
worst-case subpopulation, we first compute an estimator $\what{q}$ of
$P_{1-\alpha}^{-1}(\mu\opt)$ based on the auxiliary data, and take
$\what{h}(x) \defeq \frac{1}{\alpha} \indic{(\what{\mu}_1 - \what{\mu}_0)(x)
  \ge \what{q}}$.  In some applications, large quantities of \emph{unlabeled}
covariate observations can be cheaply collected even when \emph{labeled}
observations $(X, Y, Z)$ are expensive. Then a particularly nice estimator
$\what{q}$ of $P_{1-\alpha}^{-1}(\mu\opt)$ can be constructed by evaluating the
$(1-\alpha)$-quantile of the CATE estimator
$\what{\mu} = \what{\mu}_1 - \what{\mu}_0$ on unlabeled observations. With
cheap unlabeled covariates, such an estimator can be made arbitrarily close to
$P_{1-\alpha}^{-1}(\what{\mu})$, and hence close to $P_{1-\alpha}^{-1}(\mu\opt)$ if $\what{\mu}$ is
sufficiently close to $\mu\opt$ as we show in Section~\ref{section:asymptotics}.



To utilize the entire sample, we take a cross-fitting approach, partitioning
the data into $K$ folds and switching the roles of the main and auxiliary
datasets on each fold. We adapt the original cross-fitting algorithm for
estimating equations (due to~\citet{ChernozhukovChDeDuHaNeRo18}) to estimating
the WTE.  Denoting the $k$-th fold $I_k$ and its complement
$I_{k}^c = [n] \backslash I_k$, we fit nuisance parameters $(\what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})$ on the
$k$-th auxiliary data $\{D_i\}_{i \in I_{k}^c}$.
Using $\what{P}_k$ to denote the empirical distribution on the $k$-th main
data $\{D_i\}_{i \in I_k}$, we summarize our procedure in
Algorithm~\ref{alg:cross-fitting}.  Our estimator can be computed in both
randomized control trials using the true propensity score or in observational
studies where $\what{e}_{k}$ needs to be estimated using suitable statistical
models.



We can also derive natural analogues of the direct method (DM) and the inverse
probability weighted estimator (IPW) for estimating $\mbox{WTE}_{\alpha}$
{\small
\begin{subequations}
  \label{eqn:dm-ipw}
\begin{align}
  \what{\mbox{DM}}_{\alpha}
   & \defeq \frac{1}{K} \sum_{k=1}^K \inf_{\eta} \left\{ \frac{1}{\alpha} \E_{\what{P}_k}
    \hinge{\what{\mu}_k(X) - \eta} + \eta \right\}, \\
  ~\what{\mbox{IPW}}_{\alpha, k}
   & \defeq \frac{1}{K} \sum_{k=1}^K  \E_{\what{P}_k} \left[\what{h}_{k}(X) Y \left(
    \frac{Z}{\what{e}_{k}(X)}  - \frac{1-Z}{1-\what{e}_{k}(X)} \right) \right].
\end{align}
\end{subequations}
} Again, the above estimators reduce to their counterparts for estimating the
ATE when $\alpha = 1$. In our subsequent analysis and experiments, we focus on
the augmented estimator presented in Algorithm~\ref{alg:cross-fitting} since
unlike the two approaches above~\eqref{eqn:dm-ipw}, the augmented version
satisfies Neyman orthogonality and achieves the semiparametric efficiency
bound.








\section{Empirical results}
\label{section:experiments}


We empirically demonstrate how our worst-case sensitivity approach guards
against spurious findings that are invalidated under shifts in the underlying
population. Our WTE estimator (Algorithm~\ref{alg:cross-fitting})
automatically detects subgroups adversely affected by the treatment and
guarantees validity of findings over subpopulations on both randomized and
observational studies. We observe that our WTE estimator remains stable even
when typical CATE estimators---based on undersmoothed machine learning
models---vary significantly across estimation methods and sample sizes.

Throughout our empirical analysis, we consider $\mbox{WTE}_{\alpha}$ for
$\alpha \in \{ 1., .8, .6, .4, .2\}$ (recall $\mbox{ATE} = \mbox{WTE}_{1}$),
where without mention we replace the supremum with an infimum in the
definition~\eqref{eqn:wte} if the desired sign of the treatment effect is
positive. We use $K = 3$ in our cross-fitting procedure and use random forests
to estimate outcome models with two-fold cross-validation. Our estimator
bounds the usual cross-fitted AIPW~\citep{ChernozhukovChDeDuHaNeRo18} and
reduces to the AIPW when $\alpha = 1$, which we use as the topline estimator
for the ATE.
\ifdefined\usemsstyle
\else
\footnote{The code for all experiments can be found in
  \url{https://github.com/sookyojeong/worst-ate}.}
\fi









\ifdefined\useectastyle
\vspace{-15pt}
\fi

\subsection{Effect of Medicaid on doctor visits over time}

Since the passage of the Affordable Care Act, Medicaid has expanded to 38
states in the U.S., aiming to increase healthcare coverage for low-income
individuals.  We study the effectiveness of Medicaid in increasing healthcare
access as measured by the post-enrollment change in doctors' office
utilization.  Taking the viewpoint of an analyst in 2009, we illustrate how
findings in 2009 (``present'') may no longer hold in the subsequent decade
(``future'') due to unforeseen population shifts. Our worst-case sensitivity
approach accounts for latent intersectionality using ``present'' data alone
(Figure~\ref{fig:intro}) and calls into question the external validity of
present-day findings.



\begin{figure}[t]
\centering
\includegraphics[scale=.4]{./figures/wte_nhis_alpha.png}
\caption{ \label{fig:negex_share} Effect of Medicaid enrollment on doctors
  visits in 2009}
\end{figure}




Using ten years of data (2009-18) from the National Health Interview Survey
(NHIS), we focus on the binary outcome that indicates whether the individual
made any visits to doctors in the two-weeks prior to the survey
date~\citep{CurrieGr96, LiptonDe15, CurrieFa05,
  KahendeMaEnZhMoXuSeRo17}. Although survey respondents have limited control
on the outcome variable as the date of survey depends on the availability of
field agents, Medicaid enrollments are nevertheless non-random. Those in poor
health and higher intention of seeking treatment are more likely to enroll in
Medicaid and thus more likely to visit the doctors, leading to an upward bias
in confounded estimates. To address the potential bias, we control for a rich
set of baseline covariates ($d=396$) including demographics, medical history,
employment, earnings, health limitations, whether they require help with daily
tasks, and enrollment status in insurance and government programs.  We posit
that there are no unobserved confounders (Assumption~\ref{assumption:ig}); as
this is a strong assumption, we also study a randomized experiment in the
following subsection. We restrict attention to the Census West region to
ensure uniformity in the experiences of the control group across different
states. Although Medicaid eligibility cannot be determined based on the NHIS
survey data (number of family members is missing), we restrict attention to
those with annual income $\le$\$65K so that respondents have a positive
probability of eligibility.


The $\mbox{WTE}_{\alpha}$~\eqref{eqn:wte} allows analyzing external validity
over shifts in the study population. Using only data from 2009 ($n=82,993$
observations), in Figure~\ref{fig:negex_share} we plot the cross-fitted
estimates of $\mbox{WTE}_{\alpha}$ across a range of subpopulation sizes
$\alpha$. While the ATE is positive in 2009 (significant at 99\%), observed
effects fail to be significant even over subpopulations that comprise
$\alpha = 80\%$ of the study population. The WTE identifies subgroups
disparately affected by the treatment and shows the average-case findings are
invalidated under small changes to the study population. Our sensitivity
framework thus brings into question whether the observed effects in 2009 can
endure the test of time.



\begin{figure}[t]
\centering
\begin{tabular}{cc}
  \includegraphics[scale=.27]{./figures/wte_nhis_ts.png}
  &
    \includegraphics[scale=.23]{./figures/wte_nhis_breakdown.png}
  \\
  \small{(a) ATE estimates in 2009-18 with 95\% CIs}
  & \small{(b) CATE by covariates in 2009}
\end{tabular}
\vspace{10pt}
\caption{ \label{fig:negex_ts} Temporal changes in the ATE due to population shift}
\end{figure}


As predicted from 2009 data alone, treatment effects fail to be statistically
significant over time (Figure~\ref{fig:negex_ts}a).  To heuristically
investigate the temporal population shift in the ATE, in
Figure~\ref{fig:intro}a we analyzed covariates with the largest mean
difference between 2009 and 2018. The share of Hispanics and those who have
been in the U.S. for less than 10 years decreased, but the share of
U.S. citizens and those with high educational attainment increased. In
Figure~\ref{fig:negex_ts}b, we observe that groups whose share decreased
have higher treatment effects (and vice versa), contributing to the overall
decrease in the  treatment effect over time.






\ifdefined\useectastyle
\vspace{-15pt}
\fi
\subsection{Welfare attitudes experiment}

A large group of Americans harbors antipathy towards programs labeled
``welfare,'' a phenomenon that has generated much interest.  We are interested
in measuring how seemingly insubstantial wording changes in the description of
social welfare programs affect public support. Disdain towards ``welfare'' has
been associated with racist stereotypes towards welfare
recipients~\citep{HenryReWe04, Federico04}
and political ideology~\citep{KluegelSm86}.
Previous works have observed substantial heterogeneity with
respect to variables such as level of racism, education levels, and political
leanings~\citep{Jacoby00, Federico04, GreenKe12}.

We study an experiment on welfare attitudes in the General Social Survey (GSS)
from 1986 to 2010~\citep{GreenKe12}. We focus on the binary outcome indicating
whether the respondents state too much is being spent on either ``welfare''
(treatment) or ``assistance to the poor'' (control). Both questions about
public spending are identical except for the wording change. Covariates
include attitude towards Blacks, political views, party identification,
educational attainment, and age; there are $n = 20,783$ data points, and
$d = 22$ covariates. To illustrate the flexibility of our estimation approach,
we estimate the propensity score using a logistic regression with elastic net
regularization as if it was unknown. (We observe similar results when we use
the true propensity score $e\opt(X) \equiv \half$.)



\begin{figure}[t]
\centering
\begin{tabular}{ccc}
\includegraphics[scale=.34]{./figures/varyN_wte.png}
&
\includegraphics[scale=.34]{./figures/varyN_educ.png}
&
\includegraphics[scale=.34]{./figures/varyN_age.png}
\\
\small{(a) ATE and WTE$_\alpha$} & \small{(b) CATE by years of education}& \small{(c) CATE by age}
\end{tabular}
\vspace{10pt}
\caption{ \label{fig:posex} Treatment effects by number of observations. }
\end{figure}


By virtue of our semiparametric approach, we observe that even when CATE
estimates vary significantly across sample size and outcome model classes,
estimates of $\mbox{WTE}_{\alpha}$ align around a single value, a (empirical)
stability property shared with estimators of
ATE~\citep{CarvalhoFeMuWoYe19}. To illustrate the stability of WTE estimates,
we plot them over different sample sizes (5-20K) and outcome model classes.
Even at small sample sizes, wording changes in the survey have a resoundingly
strong effect on attitude towards government welfare programs (Figure
\ref{fig:posex}a). This observed effect is uniformly significant over
subpopulations as small as 20\% of the collected data. Such robust evidence
instills confidence in the external validity of the finding across
spatiotemporal changes in the demographic composition of respondents.  While
estimates of $\mbox{WTE}_{\alpha}$ and ATE remain relatively stable across
different sample sizes, CATE estimates vary considerably over different sample
sizes, a typical behavior for undersmoothed ML models (Figure~\ref{fig:posex}b,
c).

The stability of WTE estimates persists over different nuisance estimation
approaches.  Figure~\ref{fig:posex2} explores several common model classes for
the conditional outcome $\E[Y(z) \mid X]$: linear models with elastic net
regularizers, random forests, and gradient boosted regression trees. All
hyperparameters are chosen using 2-fold cross-validation as
before. Figure~\ref{fig:posex2}a highlights the stability of WTE and ATE
estimates along different models. In Figures~\ref{fig:posex2}b and c, we
observe that CATE estimates vary up to 25 times, especially around the
endpoint of bins, a common phenomenon often attributed to bias at the
boundaries of the support of the feature space~\citep{WagerAt18}.
We anticipate that WTE estimators will similarly suffer instability issues
when $\alpha$ is exceedingly small due to limited sample size.


\begin{figure}[t]
\centering
\begin{tabular}{ccc}
\includegraphics[scale=.34]{./figures/varyModel_wte.png}
&
\includegraphics[scale=.34]{./figures/varyModel_educ.png}
&
\includegraphics[scale=.34]{./figures/varyModel_age.png}
\\
\small{(a) ATE and WTE$_\alpha$} & \small{(b) CATE by years of education}& \small{(c) CATE by age}
\end{tabular}
\vspace{10pt}
\caption{ \label{fig:posex2} Treatment effects by model class.}
\end{figure}






\section{Asymptotics}
\label{section:asymptotics}


We now show that our cross-fitted augmented estimator
(Algorithm~\ref{alg:cross-fitting}) enjoys central limit behavior
$\frac{\sqrt{n}}{\what{\sigma}_{\alpha}} (\what{\omega}_{\alpha} -
\mbox{WTE}_{\alpha}) \cd N(0, 1)$ even when we can only estimate the nuisance
parameters~\eqref{eqn:nuisance-def} at slower-than-parametric rates.  The
Neyman orthogonality condition~\citep{Neyman59} serves a central role in our
analysis.
\begin{definition}
  \label{def:neyman}
  Let $Q \mapsto T(Q; \gamma)$ be a statistical functional with nuisance
  parameter $\gamma \in \Gamma$, where we take $\Gamma$ to be a subset of a
  normed vector space containing the true nuisance parameter $\gamma\opt$. The
  functional $T$ is Neyman orthogonal at $P$ if for all $\gamma \in \Gamma$,
  the derivative $\frac{d}{dr} \E_P[T(P; \gamma\opt + r(\gamma-\gamma\opt))]$
  exists for $r \in [0, 1)$, and is zero at $r = 0$.
\end{definition}
We use the augmented form~\eqref{eqn:aug} as our statistical functional
$T(Q; \mu_0, \mu_1, e, h)$
\begin{align}
  \label{eqn:functional}
  T(Q; \mu_0, \mu_1, e, h)
  \defeq \inf_{\eta} \left\{ \frac{1}{\alpha}
  \E_{D \sim Q}\hinge{\mu(X) - \eta} + \eta
  \right\} + \E_{D \sim Q}[\kappa(D; \mu_0, \mu_1, e, h)],
\end{align}
where we use $\mu(x) \defeq \mu_1(x) - \mu_0(x)$ as usual. The first term is
the dual form of $\mbox{WTE}_{\alpha}$ (Lemma~\ref{lemma:dual}), and the
second term is the augmentation term defined in Eq.~\eqref{eqn:aug}. Since
$\E_{D \sim P}[\kappa(D; \mu_0\opt, \mu_1\opt, e\opt, h\opt)]=0$ under ignorability
(Assumption~\ref{assumption:ig}), we have $T(P; \mu_0\opt, \mu_1\opt, e\opt, h\opt) = \mbox{WTE}_{\alpha}$.
\ifdefined\useectastyle An application of the envelope theorem confirms Neyman
orthogonality of the augmented functional~\eqref{eqn:functional}.
\else

To build intuition, we first informally argue that this augmented functional
satisfies Neyman orthogonality. We first compute the (Gateaux) derivative of
the first dual infimization term in the
functional~\eqref{eqn:functional}. From Danskin's theorem~\citep[Theorem
4.13]{BonnansSh00}, the derivative of the dual formulation is the derivative
of the objective at the unique optimal solution. Under sufficient regularity
conditions, the unique solution to the dual is given by the quantile
$P_{1-\alpha}^{-1}(\mu\opt)$, and the derivative at $r = 0$ is
\begin{equation*}
  \frac{d}{dr} \inf_{\eta} \left\{ \frac{1}{\alpha}
  \E\hinge{(\mu\opt + r (\mu - \mu\opt))(X)  - \eta} + \eta
  \right\} \Bigg|_{r=0} = \E[h\opt(X) (\mu - \mu\opt)(X)].
\end{equation*}
To compute the derivative of the second term, let
$\gamma = (\mu_0, \mu_1, e, h)$ to ease notation. So long as we can
interchange derivatives and expectations, it is straightforward to calculate
\begin{align*}
  \frac{d}{dr} \E\left[\kappa\left(D; \gamma\opt + r (\what{\gamma}_{k} -
    \gamma)\opt\right)\right] |_{r = 0} = -\E[h\opt(X) (\mu - \mu\opt)(X)]
\end{align*}
from ignorability (Assumption~\ref{assumption:ig}). This verifies Neyman
orthogonality of the functional~\eqref{eqn:functional}.

\fi

Orthogonality allows us to show a central limit theorem for our augmented
cross-fitting estimator $\what{\omega}_{\alpha}$ under the following weak rate
requirements for the nuisance parameters. Recall the bound $c >0$ on the
propensity score given in Assumption~\ref{assumption:overlap}.
\begin{assumption}
  \label{assumption:neyman} Let $\Lone{\what{\mu}_{k} -\mu\opt} \cas 0$, and let there
  exist an envelope function $\bar{\mu}: \mc{X} \to \R$ satisfying
  $\E[\bar{\mu}(X)^2] < \infty$ and $\max(|\what{\mu}_{0, k}|, |\what{\mu}_{1, k}|) \le \bar{\mu}$.  There
  exists $\delta_n, \Delta_n \downarrow 0$, and $M_h > 0$ such that with probability at least
  $1-\Delta_n$, for all $k \in [K]$
  \begin{enumerate}[(a)]
  \item
    {\small $\Ltwo{\what{e}_{k} - e\opt} + \Ltwo{\what{\mu}_{0, k} - \mu_0\opt} + \Ltwo{\what{\mu}_{1, k} -
      \mu_1\opt} \le \delta_n$, $\what{e}_{k} \in [c, 1-c]$}, and $|\what{h}_{k}| \le M_h$
\item {\small $\Ltwo{\what{e}_{k} - e\opt}\left(\Ltwo{\what{\mu}_{0, k} - \mu_0\opt} + \Ltwo{\what{\mu}_{1, k} - \mu_1\opt}\right)
    \le \delta_n n^{-1/2}$}
  \item $\Linf{\what{\mu}_{k} - \mu\opt} \le \delta_n n^{-1/3}$, and
    $|\what{q}_{k} - P_{1-\alpha}^{-1}(\what{\mu}_{k})| \le \delta_n n^{-1/3}$
  \end{enumerate}
\end{assumption}

In particular, Assumption~\ref{assumption:neyman} does not require a
$\sqrt{n}$-convergence rate (Donsker condition) on the estimators of nuisance
parameters. The first two conditions are standard convergence
rates~\citep{ChernozhukovChDeDuHaNeRo18}, also required for proving a central
limit result for the $\mbox{ATE}$. They hold, in particular, when
$\Ltwo{\what{e} - e\opt} = o_p(n^{-1/4})$ and
$\Ltwo{\what{\mu}_z - \mu\opt_z} = o_p(n^{-1/4})$ for $z \in \{0, 1\}$. The third
condition guarantees approximation of the threshold function
$h\opt(x) = \alpha^{-1} \indic{\mu\opt(x) \ge P_{1-\alpha}^{-1}(\mu\opt)}$ at suitably fast rates.
The requirement $\Linf{\what{\mu}_{k} - \mu\opt} \le \delta_n n^{-1/3}$ states that the
CATE be estimated at a somewhat faster rate compared to the case for
estimating the ATE.  \ifdefined\useectastyle \else In
Appendix~\ref{section:example}, we provide detailed examples of model classes
and learning methods where these convergence rates hold.  \fi As noted in
Section~\ref{section:approach}, when \emph{unlabeled} covariates (those
without corresponding outcomes nor treatments) are cheaply available, the
second part of $(c)$ can be easily achieved.


As the WTE is a tail-average of the CATE above the quantile $P_{1-\alpha}^{-1}(\mu\opt)$
(recall Lemma~\ref{lemma:dual}), to estimate $\mbox{WTE}_{\alpha}$ we need to
estimate the quantile $P_{1-\alpha}^{-1}(\mu\opt)$. Towards this goal, we require that a
positive density exists at its $(1-\alpha)$-quantile, a standard condition
required for quantile estimation~\cite[Chapter 3.7]{VanDerVaartWe96}.
Let~$\mc{U}$ be a subset of measurable functions $\mu: \mc{X} \to \R$ such
that the following holds:
\begin{quote}
  $F_{r, \mu}$, the cumulative distribution of
  $(\mu\opt + r(\mu - \mu\opt))(X)$, is uniformly differentiable in $r \in [0, 1]$
  at $P_{1-\alpha}^{-1}(\mu\opt + r(\mu - \mu\opt))$, with a positive density. Formally, if we let
  $q_{r, \mu} \defeq P_{1-\alpha}^{-1}(\mu\opt + r(\mu - \mu\opt))$, then for each $r \in [0, 1]$,
  there is a positive density $f_{r, \mu}(q_{r, \mu}) > 0$ such that
\begin{equation}
  \label{eqn:unif-diff}
  \lim_{t \to 0} \sup_{r \in [0, 1]}
  \left| \frac{1}{t} \left(F_{r, \mu}(q_{r, \mu} + t)
      - F_{r, \mu}(q_{r, \mu})\right)
    - f_{r, \mu}(q_{r, \mu}) \right| = 0.
\end{equation}
\end{quote}
We require this holds for our estimators $\mu = \what{\mu}_{k}$ with high probability.
\begin{assumption}
  \label{assumption:regularity}
 $\exists \Delta_n' \downarrow 0$ s.t. with probability at least
  $1-\Delta_n'$, $\what{\mu}_{k} \in \mc{U}$ for all $k \in [K]$.
\end{assumption}
\noindent In particular, Assumption~\ref{assumption:regularity} requires
$\mu\opt(X)$ to have a positive density at $P_{1-\alpha}^{-1}(\mu\opt)$.

We are now ready to give our main technical result which shows that the
augmented cross-fitting estimator $\what{\omega}_{\alpha}$ enjoys central
limit rates with the influence function
\begin{align}
  \label{eqn:influence}
  \psi(D) \defeq
  \frac{1}{\alpha} \hinge{\mu\opt(X) - P_{1-\alpha}^{-1}(\mu\opt)} + P_{1-\alpha}^{-1}(\mu\opt)
  - \mbox{WTE}_{\alpha}
  + \kappa\left(D; \mu_0\opt, \mu_1\opt, e\opt, h\opt\right).
\end{align}
Indeed, $\psi(D)$ is a valid influence function under
Assumption~\ref{assumption:ig}: we have $\E[\psi(D)] = 0$, and
$\var(\psi(D)) = \sigma^2_{\alpha}$ where $\sigma_{\alpha}^2$ is the
asymptotic variance defined in expression~\eqref{eqn:var}.
\begin{theorem}
  \label{theorem:clt}
  Under
  Assumptions~\ref{assumption:ig}-\ref{assumption:regularity},
  $\sqrt{n} (\what{\omega}_{\alpha} - \mbox{WTE}_{\alpha}) =
  \frac{1}{\sqrt{n}} \sum_{i=1}^n \psi(D_i) + o_p(1)$.
  Further, we have $\what{\sigma}^2_{\alpha} \cp \sigma^2_{\alpha}$ and
  $\frac{\sqrt{n}}{\what{\sigma}_{\alpha}} (\what{\omega}_{\alpha} -
  \mbox{WTE}_{\alpha}) \cd N(0, 1)$.
\end{theorem}
\noindent Since the proof of Theorem~\ref{theorem:clt} is involved, we give
its sketch below, emphasizing how we leverage Neyman orthogonality of the
augmented functional~\eqref{eqn:functional} to obtain our result. We provide
rigorous details of the proof in
Appendices~\ref{section:proof-clt-classical},~\ref{section:proof-clt-neyman},~\ref{section:proof-clt-var}.
\citet{ChernozhukovChDeDuHaNeRo18} showed that the solution of a Neyman
orthogonal \emph{estimating equation} could be estimated at the usual central
limit rate, even when nuisance parameters converge at slower rates.  Their
results do not apply to the $\mbox{WTE}_{\alpha}$ as it is a nonlinear
statistical functional of the underlying probability measure $P$. To account
for such nonlinearity, our theoretical analysis considerably extends existing
results.  Using tools from empirical process theory~\citep{VanDerVaartWe96},
our main result shows that for uniformly Hadamard differentiable functionals,
orthogonality still allows insensitivity to estimation error in nuisance
parameters. Even when $\alpha = 100\%$ so that $\mbox{WTE}_{1} = \mbox{ATE}$
and our estimator reduces to the cross-fitted AIPW for the ATE, our argument
provides a different proof to that given
by~\citet{ChernozhukovChDeDuHaNeRo18}.

\subsection*{Sketch of proof for Theorem~\ref{theorem:clt}}
\label{section:proof-clt}

Our proof proceeds in three parts.  Recalling $\what{P}_k$, the empirical
distribution on the $k$-th fold, our cross-fitted estimator can be written
succinctly as $\frac{1}{K} \sum_{k = 1}^{K} T(\what{P}_k; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})$. We
emphasize (empirical) expectations over $Q$ in the
definition~\eqref{eqn:functional} are taken only over $D$, and not over the
randomness in $(\what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})$.

The first two parts show that for each fold $k \in [K]$,
\ifdefined\useectastyle
{\small
\begin{align}
  \label{eqn:fold-clt}
  \sqrt{|I_k|} \left( T\left(\what{P}_k; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k}\right)
  - T(P; \mu_0\opt, \mu_1\opt, e\opt, h\opt) \right) = \frac{1}{\sqrt{|I_k|}} \sum_{i \in I_k} \psi(D_i)
  + o_p(1).
\end{align}
}
\else
\begin{align}
  \label{eqn:fold-clt}
  \sqrt{|I_k|} \left( T\left(\what{P}_k; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k}\right)
  - T(P; \mu_0\opt, \mu_1\opt, e\opt, h\opt) \right) = \frac{1}{\sqrt{|I_k|}} \sum_{i \in I_k} \psi(D_i)
  + o_p(1).
\end{align}
\fi
Since $T(P; \mu_0\opt, \mu_1\opt, e\opt, h\opt) = \mbox{WTE}_{\alpha}$ by ignorability
(Assumption~\ref{assumption:ig}), this gives our first result. Towards this
goal, decompose the left hand side of the equality~\eqref{eqn:fold-clt} into
\begin{subequations}
  \label{eqn:two-parts}
  \begin{align}
    & \sqrt{|I_k|} \left( T\left(\what{P}_k; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k}\right) - T\left( P; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k}\right)
      \right) \label{eqn:part-classical} \\
    & + \sqrt{|I_k|} \left( T\left(P; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k}\right) - T(P; \mu_0\opt, \mu_1\opt, e\opt, h\opt) \right).
      \label{eqn:part-neyman}
  \end{align}
\end{subequations}
In Part I of the proof (Appendix~\ref{section:proof-clt-classical}), we prove
that the first term~\eqref{eqn:part-classical} is asymptotically equal to
$\frac{1}{\sqrt{|I_k|}} \sum_{i \in I_k} \psi(D_i)$. Leveraging tools from empirical
process theory~\citep{VanDerVaartWe96}, we first show that the functional
$Q \mapsto T(Q; \what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})$ satisfies a uniform variant of Hadamard
differentiability so that the functional delta method applies. A careful
application of a uniform version of the Lindeberg-Feller central limit theorem
over functions gives the desired conclusion.

In Part II (Appendix~\ref{section:proof-clt-neyman}), we use Neyman
orthogonality of our augmented estimator to show that the second
term~\eqref{eqn:part-neyman} vanishes asymptotically.  Define
$\mathfrak{R}_k: [0, 1] \to \R$ \ifdefined\useectastyle {\small
\begin{align}
  \label{eqn:remainder}
  \mathfrak{R}_k(r) \defeq
  T\left(P; (1-r) (\mu_0\opt, \mu_1\opt, e\opt, h\opt) + r (\what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})\right) - T(P; \mu_0\opt, \mu_1\opt, e\opt, h\opt),
\end{align}
}
\else
\begin{align}
  \label{eqn:remainder}
  \mathfrak{R}_k(r) \defeq
  T\left(P; (1-r) (\mu_0\opt, \mu_1\opt, e\opt, h\opt) + r (\what{\mu}_{0, k}, \what{\mu}_{1, k}, \what{e}_{k}, \what{h}_{k})\right) - T(P; \mu_0\opt, \mu_1\opt, e\opt, h\opt),
\end{align}
\fi
so that $\mathfrak{R}_k(0) = 0$, and $\mathfrak{R}_k(1)$ is equal to the second
term~\eqref{eqn:part-neyman}. $\mathfrak{R}_k(r)$ is continuously differentiable under
suitable conditions and mean value theorem gives
\begin{align*}
  \mathfrak{R}_k(1) = \mathfrak{R}_k(0) + \mathfrak{R}_k'(r) \cdot (1-0) = \mathfrak{R}_k'(r)
\end{align*}
for some $r \in [0, 1]$.  From Neyman orthogonality, we have $\mathfrak{R}_k'(0) = 0$.
Building on this, we show that all values of $\mathfrak{R}_k'(r)$ are sufficiently
small: under postulated convergence rates for the nuisance parameters in
Assumption~\ref{assumption:neyman},
$\sup_{r \in [0, 1]} |\mathfrak{R}_k'(r)| = o_p(n^{-1/2})$.

In Part III (Appendix~\ref{section:proof-clt-var}), we show consistency of our
variance estimator: $\what{\sigma}^2_{\alpha} \cp
\sigma^2_{\alpha}$. Combining this with the central limit
result~\eqref{eqn:fold-clt}, Slutsky's lemma gives our second
result. $\diamond$








\section{Semiparametric efficiency bound}
\label{section:efficiency}

We now establish a semiparametric efficiency bound, showing that all (regular)
estimators of $\mbox{WTE}_{\alpha}$ necessarily have asymptotic variance
larger than $\sigma_{\alpha}^2$, both when the true propensity score
$e\opt(\cdot)$ is known and unknown. In particular, this implies that our
augmented estimator $\what{\omega}_{\alpha}$ achieves the \emph{optimal}
asymptotic variance and that its influence function~\eqref{eqn:influence} is
the efficient influence function for estimating $\mbox{WTE}_{\alpha}$. Our
efficiency bound i) informs the design of experiments by providing the minimum
number of study participants required to reach a conclusion that is valid
across subpopulations no smaller than $\alpha$ and ii) quantifies how the
required sample size grows with the level of desired external validity
$\alpha$.

For parametric problems, Hajek-Le Cam theorems~\cite[Chapter 8]{VanDerVaart98}
give a lower bound on the asymptotic variance of regular estimators, more
generally, a lower bound on the mean squared error for any estimator. Since
these bounds coincide with the Cramer-Rao bound for unbiased estimators, they
can be considered as an asymptotic Cramer-Rao bound. We consider parametric
submodels of our semiparametric problem, i.e., finite-dimensional
parameterizations that contain the truth. Since the asymptotic variance of any
(smooth enough) semiparametric estimator is worse than the Hajek-Le Cam
bound---equivalently, the Cramer-Rao bound---of any parametric submodel, the
semiparametric efficiency bound is defined as the supremum of the Hajek-Le Cam
bound over all parametric submodels.  As these definitions are standard yet
tedious~\citep{Newey90, BickelKlRiWe93}, we defer a formal treatment to
Appendix~\ref{section:proof-efficiency}.

We now characterize the semiparametric efficiency bound for estimating
$\mbox{WTE}_{\alpha}$.
\begin{assumption}
  \label{assumption:efficiency}
  $\mu\opt(X)$ has a positive density, and {\small
    $\Linf{\E[|Y(z)| \mid X]} < \infty$, $z \in \{0, 1\}$}.
\end{assumption}
\begin{theorem}
  \label{theorem:efficiency}
  Let  Assumptions~\ref{assumption:ig},~\ref{assumption:overlap},~\ref{assumption:residuals},~\ref{assumption:efficiency}
  hold, and $\sigma^2_{\alpha} = \E[\psi(D)^2] > 0$, where $\psi$ was defined in
  Eq.~\eqref{eqn:influence}. Then, $\mbox{WTE}_{\alpha}$ is a differentiable
  parameter in the sense of Definition~\ref{def:diff}, and its semiparametric
  efficiency bound is given by $\E[\psi(D)^2]$. The efficiency bound is
  identical if the true propensity score is known.
\end{theorem}
\noindent See Appendix~\ref{section:proof-efficiency} for the
proof. Theorem~\ref{theorem:clt} and~\ref{theorem:efficiency} show that our
cross-fitted augmented estimator $\what{\omega}_{\alpha}$ achieves the
semiparametric efficiency bound $\mathbb{V}(\mbox{WTE}_{\alpha}; P)$. Similar
to the efficiency bound for the ATE~\citep{Hahn98}, the knowledge of
$e\opt(\cdot)$ does not affect the efficiency bound
$\mathbb{V}(\mbox{WTE}_{\alpha}; P)$. We conclude that our estimator is
semiparametrically efficient for \emph{both} observational studies and
randomized control trials.


Our semiparametric efficiency bound allows calculating the minimum required
sample size for finding conclusions that are robust against all subpopulations
of size at least $\alpha$. Consider the hypothesis test
$H_0: \mbox{WTE}_{\alpha} \ge 0 ~\mbox{v.}~H_{\rm A}: \mbox{WTE}_{\alpha} <
-\epsilon$, where $\epsilon > 0$ is the specified minimum detectable effect
size (recall that the desired sign of the treatment effect is negative). As an
illustration, consider a test with size ($\P(\mbox{Type I error})$) at most
$5\%$ and power ($1 -\P(\mbox{Type II error})$) at least $80\%$. A standard
power calculation shows $n \ge 6.2 \frac{\sigma_{\alpha}^2}{\epsilon^2}$
samples are needed to detect a worst-case subpopulation treatment effect of
size $-\epsilon$. In particular, this number grows as we require a stronger
level of external validity, or equivalently as the worst-case subpopulation
size $\alpha$ becomes small.



\section{Discussion}
\label{section:discussion}






Motivated by challenges in evaluating treatments under unanticipated
population shifts, we proposed a sensitivity analysis framework based on the
worst-case subpopulation treatment effect. We advocate for such conservatism
in important policy decisions that need to benefit all subpopulations
uniformly.
While the WTE~\eqref{eqn:wte} provides a strong notion of robustness, it may
be overly conservative in scenarios where one is concerned with more
structured covariate shifts. When covariate shift on only a small subset of
$X$ is of interest, the definition~\eqref{eqn:wte} can be modified over this
subset. More broadly, studying structured shifts in the covariate distribution
$P_X$ is an exciting direction of future work.

Our theoretical developments show that our cross-fitted augmented estimator
inherits two advantageous inferential properties of the AIPW estimator for the
ATE: orthogonality and efficiency. Elementary derivations show, however, that
our estimator does not satisfy the doubly robust property due to the
additional nuisance parameter
$h\opt(x) = \alpha^{-1} \indic{\mu\opt(x) \ge P_{1-\alpha}^{-1}(\mu\opt)}$. Another limitation of
our augmented estimator is that it is not necessarily increasing in $\alpha$.
Though the extension of the direct method and the inverse probability weighted
estimator are increasing in $\alpha$, it is unclear whether an orthogonal,
efficient, and monotone estimator of the WTE exists. Finally, estimators of
tail-averages often suffer higher variance and may be more sensitive to lack
of overlap and unobserved confounding. Further investigation of these issues
is a topic for future research.

Sometimes it is of interest to estimate the \emph{average treatment effect on
  the treated} $\E[Y(1) - Y(0) \mid Z = 1]$. There are two natural ways to
extend the worst-case definition~\eqref{eqn:wte} to measure the effect on the
treated, depending on whether the worst-case is still taken over $Q_X$, or
over the conditional distribution $Q_{X \mid Z = 1}$. We leave a systematic
study of the two definitions and corresponding inferential frameworks to
future work.





\paragraph*{Acknowledgments} We thank Steve Yadlowsky for helpful comments.




\bibliographystyle{abbrvnat}

\setlength{\smallskipamount}{.7em}
\bibliography{bib}


\ifdefined\usemsstyle

\ECSwitch



\ECHead{Appendix}

\else
\newpage