The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
125,979 characters
Efficient Semiparametric Estimation of Average Treatment Effects Under Covariate Adaptive Randomization
\maketitle
\begin{abstract}
Experiments that use covariate adaptive randomization (CAR) are commonplace in
applied economics and other fields. In such experiments, the experimenter first
stratifies the sample according to observed baseline covariates and then assigns
treatment randomly within these strata so as to achieve balance according to
pre-specified stratum-specific target assignment proportions. In this paper, we
compute the semiparametric efficiency bound for estimating the average treatment
effect (ATE) in such experiments with binary treatments allowing for the class
of CAR procedures considered in
\citet{2018bugniInferenceCovariateAdaptiveRandomization,
2019bugniInferenceCovariateAdaptiveRandomization}. This is a broad class of
procedures and is motivated by those used in practice. The stratum-specific
target proportions play the role of the propensity score conditional on all
baseline covariates (and not just the strata) in these experiments. Thus, the
efficiency bound is a special case of the bound in
\citet{1998hahnRolePropensityScore}, but conditional on all baseline
covariates. Additionally, this efficiency bound is shown to be achievable under
the same conditions as those used to derive the bound by using a cross-fitted
Nadaraya-Watson kernel estimator to form nonparametric regression adjustments.
\end{abstract}
\noindent
{\it Keywords:} Efficient semiparametric estimation, average treatment effect,
randomized experiments, covariate adaptive randomization
\hfill
\noindent
{\it JEL Classification:} C14, C90
\newpage
\section{Introduction}
Experiments that use covariate adaptive randomization (CAR) are commonplace in
applied economics and other fields. In such experiments, the experimenter first
groups sample units according to observed baseline covariates (stratification)
and then assigns treatment randomly to achieve \emph{balance} within these
groups (strata). The term balance here means that the experimenter additionally
specifies stratum-specific target treatment proportions and assigns treatment so
that the proportion of units assigned to treatment reaches the corresponding
target as the sample size grows across strata. A textbook treatment of covariate
adaptive and stratified randomization in clinical trials can be found in
\citet{2015rosenbergerRandomizationClinicalTrials}. For review articles on their
use in development economics, see \citet{2007dufloChapter61Using},
\citet{2009bruhnPursuitBalanceRandomization} and
\citet{2017atheyChapterEconometricsRandomized}. In experiments comparing
outcomes from binary treatments, the quantity of interest is often the average
treatment effect (ATE). In this paper, we are concerned with efficient
semiparametric estimation of the ATE while allowing for a broad class of CAR
procedures motivated by those used in practice. We have two main questions. The
first is: is there a well-defined semiparametric efficiency bound
(SPEB) for this class of procedures? The second is: does there exist a
semiparametric estimator that achieves this bound and under what conditions will
this happen? Our answers to both are affirmative. For the first, we show that
a version of the bound in \citet{1998hahnRolePropensityScore} is valid. For the
second, we show that under the same (weak) conditions used for derivation of the
bound, there is a feasible estimator that achieves the bound asymptotically.
In randomized experiments, correctly implemented randomization ensures that in
expectation, confounding factors are equally distributed across treatment arms
so that differences in outcomes are solely due to the differences in
treatment. Stratified randomization additionally aims to ensure that along
observable dimesions, this also remains true in practice. In CAR experiments,
both discrete and continuously distributed covariates are combined to form
strata. As shown in \citet[Section 7.2]{2017atheyChapterEconometricsRandomized},
the main statistical benefit of stratified randomization for ATE estimation is
improved precision. The standard recommendation for ATE estimation in stratified
experiments is to regress the observed outcomes on indicators for strata and
their interactions with treatment status in a linear regression equation (a
fully saturated linear regression model). The coefficients on the interaction
terms from this regression are then combined with sample stratum proportions to
construct the ATE estimate. The resulting estimator of the ATE is analogous to
the Horvitz-Thompson estimator
(\citet{1952horvitzGeneralizationSamplingReplacement}) and does not use
information beyond the strata. However, experimenters concerned with estimation
precision may want to use information in the baseline covariates not contained
in the strata.
An additional complication in the CAR context comes from the choice of treatment
assignment mechanism by the experimenter. Many popular treatment assignment
mechanisms used in practice aim for faster targeting of the target assignment
proportions than simple independent and identically distributed (i.i.d.)
assignment and in doing so, induce dependence in the observed outcomes through
dependence in the treatment assignments. One such example is stratified permuted
block randomization (SPBR,
\th\ref{eg--sbr}). \citet{2018bugniInferenceCovariateAdaptiveRandomization}
provide additional examples. The dependence induced by CAR designs can affect
the behavior of ATE estimators in surprising ways. Analyzing the case where
target assignment proportions are constant across strata,
\citet{2018bugniInferenceCovariateAdaptiveRandomization} show that the standard
difference in treatment and control group means can have a limit variance that
depends explicitly on the choice of treatment assignment mechanism. For example,
all else held equal, the same estimator exhibits a larger limit variance under
i.i.d. treatment assignments than when treatments are assigned according to
SPBR, even when both mechanisms have the same target
proportions. \citet{2018bugniInferenceCovariateAdaptiveRandomization} also show
that the same phenomenon holds true of the ``stratum fixed effects'' estimator
which is recommended by \citet{2009bruhnPursuitBalanceRandomization}.
\citet{2019bugniInferenceCovariateAdaptiveRandomization}
extends this work and consider both multiple treatments and target assignment
proportions that vary by strata. They show that the estimator of the ATE from a
fully saturated regression has a limit variance that depends on the target
proportions (among other things), but not on the particular choice of assignment
mechanism. They also show that for the stratum fixed effects estimator however,
the limit variance is still affected explicitly by the choice of assignment
mechanism. \citet{2018bugniInferenceCovariateAdaptiveRandomization,
2019bugniInferenceCovariateAdaptiveRandomization} do not consider efficiency
questions.
There is a literature that considers efficiency gains in experiments from using
baseline covariates via linear regression adjustments. However, whether there
are gains at all depends on the linear regression specification.
In the case without stratification,
\citet{2008freedmanRegressionAdjustmentsExperimental}
shows that linear regression adjustments (without interactions between treatment
status and baseline covariates) can hurt asymptotic precision of the ATE
estimates though the estimates remain consistent. However, work by
\citet{2001yangEfficiencyStudyEstimators} and
\citet{2013linAgnosticNotesRegression} show that appropriately formed regression
adjustments (i.e. with the correct interaction terms) cannot hurt (and indeed
can improve) asymptotic precision in ATE estimation. For CAR experiments,
\citet{2022maRegressionAnalysisCovariate} build on the results of
\citet{2018bugniInferenceCovariateAdaptiveRandomization,
2019bugniInferenceCovariateAdaptiveRandomization} and show that
correctly formed linear regression adjustments cannot hurt (and can improve)
asymptotic precision under CAR. These works do not treat semiparametrically
efficient estimation.
In this paper, we are concerned with the semiparametrically efficient estimation
of the ATE in the CAR framework under the minimal set of assumptions laid out by
\citet{2019bugniInferenceCovariateAdaptiveRandomization}. There is a large
statistics literature around semiparametric efficiency, starting with the
seminal work of \citet{1956steinEfficientNonparametricTesting}. Most of these
are developed assuming i.i.d. data.
\citet{1998bickelEfficientAdaptiveEstimation} offers a comprehensive textbook
treatment and \citet{1990neweySemiparametricEfficiencyBounds} provides an
approachable review. Notable examples with non-i.i.d. data can be found in
\citet{2001bickelInferenceSemiparametricModels},
\citet{2004greenwoodIntroductionEfficientEstimation} and
\citet{2010komunjerSemiparametricEfficiencyBound} as well in references
therein. For treatment effects, \citet{1998hahnRolePropensityScore} derives the
SPEB for the ATE in observational studies with i.i.d. data when treatment
assignment is ignorable conditional on observable covariates (ignorable as in
\citet[Section 1.3]{1983rosenbaumCentralRolePropensity}). The
\citet{1998hahnRolePropensityScore} bound does not apply immediately to our
setting since the treatment assignment rule is allowed to depend on the entire
profile of sample strata. For instance, it is not clear \emph{a priori} if the
choice of assignment mechanism will affect the efficiency bound since it can
clearly influence limit behavior of estimators as in
\citet{2018bugniInferenceCovariateAdaptiveRandomization,
2019bugniInferenceCovariateAdaptiveRandomization}. Furthermore, even when
treatment assignments are i.i.d. in a CAR context, the covariates being
conditioned on during assignment are the strata, so it is again not immediate
from the bound in \citet{1998hahnRolePropensityScore} what role the additional
baseline covariates can play in providing efficiency gains. This paper shows
that a version of the \citet{1998hahnRolePropensityScore} bound accounting for
both stratum-specific target proportions and all baseline covariates is the SPEB
across all CAR experimental procedures. The target proportions play the role of
the propensity score conditional on baseline covariates. Additionally, the
choice of assignment mechanism does not affect this bound, only the target
proportions do. We derive this by using the partial sums arguments of
\citet{2018bugniInferenceCovariateAdaptiveRandomization,
2019bugniInferenceCovariateAdaptiveRandomization} to show that the log
likelihood ratios under \(1 / \sqrt{n}\) local alternatives has the local
asymptotic normality (LAN) property of
\citet{1960lecamLocallyAsymptoticallyNormal}.
In concurrent work, \citet{2022armstrongAsymptoticEfficiencyBounds} proves the
LAN property under a larger class of experimental designs using martingale
methods. The experiments considered there include the CAR framework and
additionally allows for arbitrary dependence of the treatment rule on the
covariates as well as observations of past outcomes (to account for sequential
assignment). \citet{2022armstrongAsymptoticEfficiencyBounds}'s main goal is to
show that an optimized form of the \citet{1998hahnRolePropensityScore} bound is
a lower bound on limit variance of the ATE across all experimental designs in
that class. This is motivated by recent papers on optimal design of experiments
by using past waves to either allocate treatment sequentially
(e.g. \citet{2011hahnAdaptiveExperimentalDesign}) or to form optimal strata in a
main experiment (e.g. \citet{2022tabord-meehanStratificationTreesAdaptive},
\citet{2022baiOptimalityMatchedPairDesigns} and
\citet{2021cytrynbaumDesigningRepresentativeBalanced}). The optimized bound is
analogous to implementing a ``conditional on covariates'' Neyman allocation
(\citet{1934neymanTwoDifferentAspects}). We differ from
\citet{2022armstrongAsymptoticEfficiencyBounds} work on two counts. First, while
their lower bound holds across the designs the aforementioned class, they do not
provide efficiency bounds in specific instances within the class. We provide
efficiency bounds for a given fixed stratification scheme and we do
not consider the question of optimizing the bound. As a result our efficiency
bound as a variance lower bound is higher than the optimized one in
\citet{2022armstrongAsymptoticEfficiencyBounds} and hence sharper for a given
fixed stratification scheme. Second,
\citet{2022armstrongAsymptoticEfficiencyBounds} does not consider the question
of when the bounds are achievable. We do this for the CAR framework explicitly
as explained in the subsequent two paragraphs.
Once an efficiency bound is established, the question of its sharpness remains,
in the sense of whether or not it is achievable. A well known phenomenon in the
literature is that conditions under which a SPEB is achievable are typically
much stronger than those required to derive the
bound. \citet{1990ritovAchievingInformationBounds} provide counterexamples where
a finite and non-singular SPEB exists but in certain non-trivial submodels, even
consistent (let alone efficient) estimation of the parameter of interest is
impossible. One of their examples is the partially linear model
(\citet{1986engleSemiparametricEstimatesRelation},
\citet{1988robinsonRootNConsistent}) which is commonly used in applied
economics. A well known condition in the literature is that \(1 /
\sqrt{n}\)-consistent and asymptotically normal (\(1 / \sqrt{n}\)-CAN) two-step
semiparametric estimators require any first step infinite-dimensional
(henceforth nonparametric) nuisance parameters to be estimated at a rate faster
than \(n^{- \frac{1}{4}}\) where \(n\) is the sample size (see for instance
\citet{1994neweyAsymptoticVarianceSemiparametric} and
\citet{2003chenEstimationSemiparametricModels}). This is mainly because
nonparametric estimators exhibit considerable bias and this condition limits the
effect of this bias on the second estimation step. The \(n^{- \frac{1}{4}}\)
rate condition however is especially restrictive when the covariates are of
higher dimension, due to the curse of dimensionality. Achieving this rate
condition typically requires imposing smoothness and/or Donsker conditions
(often dimension dependent) on unknown nuisance parameters and the estimator in
question. For the ATE, \citet{1998hahnRolePropensityScore} shows that
semiparametrically efficient estimation can be done under regularity conditions
through nonparametric regression adjustments or imputation. In their first step,
series estimators of conditional means and the propensity score are
used. Estimation by kernel methods can also be done, see
\citet{2004imbensNonparametricEstimationAverage} for a review. In all of these,
the use of smoothness restrictions on nonparametric population unknowns is
ubiquitous.
An additional contribution of this paper is to show that the SPEB derived for
the ATE is achievable under the same conditions used for its derivation. We do
this by leveraging the \emph{efficient influence function} to form an estimating
equation (or moment condition) for the ATE. The resulting estimator is the
familiar (and famed) augmented inverse probability weighted (AIPW) estimators
due to \citet{1994robinsEstimationRegressionCoefficients},
\citet{1995robinsAnalysisSemiparametricRegression}, and
\citet{1999scharfsteinAdjustingNonignorableDrop}. We further note that since the
propensity score is known in these experiments, the estimating equation is
linear in the unknown nonparametric component and has the local robustness
property of \citet{2018chernozhukovDoubleDebiasedMachine} and
\citet{2022chernozhukovLocallyRobustSemiparametric}. This reduces
the first order impact of bias in the first stage nonparametric estimates on the
resulting ATE estimator considerably and allows for efficient estimation under
much weaker conditions. The nonparametric unknowns here are conditional means of
the potential outcomes given baseline covariates. We show first that for a
generic estimator of conditional means, a weak \(L_{2}\) consistency
requirement combined with cross-fitting as in
\citet{2018chernozhukovDoubleDebiasedMachine} is sufficient for efficient
estimation of the ATE under CAR. The conditions under which the particular
choice of estimator achieves this \(L_{2}\) consistency are first left abstract
- different conditions apply to kernel estimators, random forests, series
estimators and neural networks for instance. For a given choice of nonparametric
estimator, smoothness conditions or functional form restrictions may be
unavoidable in achieving the \(L_{2}\) consistency property. Next, we establish
that if the particular nonparametric estimator is a cross-fitted Nadaraya-Watson
kernel regression estimator, then efficient estimation is possible under no
additional restrictions on the semiparametric model. Our motivation for using
the cross-fitted Nadaraya-Watson estimator is two-fold. First, the results of
\citet{1980devroyeDistributionFreeConsistencyResults} and
\citet{1980spiegelmanConsistentWindowEstimation} show that the Nadaraya-Watson
estimator is universally \(L_{2}\) consistent if the outcome in the regression
has a finite second moment. Second, the algebra of the cross-fitted
Nadaraya-Watson estimator is convenient in showing negligiblility of remainder
terms.
We provide simulation evidence of the finite sample performance of the feasible
efficient estimator. Our simulations compare this estimator to an infeasible
efficient estimator as well as an estimator that uses information only from the
stratum labels but not the additional baseline covariates. A comparison is also
provided with an imputation estimator of
\citet{1994chengNonparametricEstimationMean} and
\citet{1998hahnRolePropensityScore}. Across all simulations designs, we find
that using the feasible efficient estimator produces an efficiency gain of
13\% in terms of mean squared error reduction in comparison to the estimator
that discards information from baseline covariates beyond strata.
The maximal gain in this comparison over all simulations is a 40\% reduction in
mean squared error. Compared to the imputation estimator, our estimator exhibits
considerably less bias across most simulation designs.
The remainder of the paper is organized as follows. Section
\ref{sec--preliminaries} sets up the assumptions about the underlying population
from which the sample is drawn as well as the sampling process and the treatment
assignment scheme. Section \ref{sec--speb} presents the main result concerning
the semiparametric efficiency bound (\th\ref{thm--speb-car}) after a discussion
about parametric submodels in CAR experiment context. Section
\ref{sec--efficient-estimation} shows that the efficiency bound is achievable
under the same assumptions as used for its derivation via a two-step
semiparametric estimator using the efficient influence function. Section
\ref{sec--sims} provides Monte Carlo evidence of the performance of the
efficient estimator in finite samples. Section \ref{sec--conclusion} concludes
the paper.
\section{Preliminaries}
\label{sec--preliminaries}
This section describes assumptions on the populations of interest as well as the
sampling process which produces the observed data. This is done in two
subsections. Subsection \ref{subsec--CAR-pop} first describes the unobserved
study population of interest through a Neyman-Rubin causal model. Then, a target
\emph{observed population} is described for all CAR experiments considered in
this paper. This target observed population is useful for deriving the
efficiency bound in Section \ref{sec--speb}. Next, Subsection
\ref{subsec--CAR-sample} describes sampling assumptions and provides the two
main examples of sampling schemes in this paper.
Throughout this paper, all random variables and vectors will be defined on a
sufficiently rich underlying probability space \((\Omega, \mathscr{F},
\Pr)\). Independence of random elements is denoted by \(\raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}}\) and
expectations computed against the probability measure \(\Pr\) are denoted by
\(\mathbb{E} [\cdot]\). We will seldom refer to the underlying probability space
but define it nonetheless for clarity. Furthermore, absent any subscripts,
\(\mathbb{E} [\cdot]\) will also denote expected values assuming the ``true''
distribution of underlying data to be defined later on. The \(d\)-dimensional
multivariate normal distribution with mean vector \(\mathbf{m} \in
\mathbb{R}^{d}\) and covariance matrix \(\Sigma \in \mathbb{R}^{d \times d}\) is
denoted \(\mathcal{N} \left( \mathbf{m}, \Sigma \right)\). We denote the
matrix/vector transpose operation by \((\cdot)^{\prime}\). The natural numbers
are denoted by \(\mathbb{N} = \{1, 2, 3, 4, \dots\}\) and for a given
\(\mathcal{S} \in \mathbb{N}\), denote \(\mathbb{N}_{\mathcal{S}} = \{1, \dots,
\mathcal{S}\} = \{s \in \mathbb{N} : 1 \leq s \leq \mathcal{S}\}\).
\subsection{A population framework for CAR experiments.}
\label{subsec--CAR-pop}
We employ a standard binary treatment potential outcomes framework and assume
that samples are drawn from an infinite super-population. The population of
interest is described by a random vector \(W\) taking values in \(\mathbb{R}^{2
+ k}\) for some \(k \in \mathbb{N}\), with \(W^{\prime} = \left(Y (0), Y (1),
Z^{\prime} \right)\). For each \(a \in \{0, 1\}\), \(Y (a)\) is a scalar random
variable representing the potential outcome from receiving treatment \(a\). The
treatment \(a = 1\) can be interpreted as an ``innovation'' whereas \(a = 0\) is
a ``status quo'' or ``control''. \(Z\) is a random \(\mathbb{R}^{k}\)-vector of
baseline covariates. We denote the true distribution of \(W\) by
\(Q_{0}\). Prior to the assignment of treatment, only baseline covariates are
observable. Furthermore, after treatment has been assigned, only the outcome
associated with received treatment is observable so that \((Y (0), Y (1))\) is
not jointly observable. We maintain the following assumptions about the
population distribution \(Q_{0}\).
\begin{assumption}
\th\label{asm--Q}
Let \(\mu_{Z}\) be a \(\sigma\)-finite measure on the Borel sets of
\(\mathbb{R}^{k}\) and for \(a \in \{0, 1\}\), let \(\mu_{a}\) be a
\(\sigma\)-finite measure on the Borel sets of \(\mathbb{R}\). Furthermore, let
\(\mu\) denote the product measure on the Borel sets of \(\mathbb{R}^{2 + k}\)
constructed from \(\mu_{0}\), \(\mu_{1}\) and \(\mu_{Z}\). The true population
distribution, \(Q_{0}\), belongs to a family \(\mathbf{Q}\) of distributions on
\(\mathbb{R}^{2 + k}\) such that for each \(Q \in \mathbf{Q}\),
\begin{enumerate}[label = (\alph*)]
\item \label{asm--dominance}
\(Q\) is dominated by \(\mu\) with Radon-Nikodym density \(q \left( \cdot; Q
\right) = \mathrm{d} Q / \mathrm{d} \mu\).
\item \label{asm--nontriv-var}
The potential outcomes have finite second moments under \(Q\),
i.e. \(\mathbb{E}_{Q} \left[ Y {(a)}^{2} \right] < \infty\) for each \(a \in
\{0, 1\}\), where \(\mathbb{E}_{Q}\) denotes the expected value assuming data
are distributed according to \(Q\).
\end{enumerate}
\end{assumption}
The dominance assumption is made for mathematical convenience and is standard in
the literature on semiparametric efficiency. Additionally, the choices of
dominating measures \(\mu_{0}, \mu_{1}\) and \(\mu_{Z}\) are irrelevant and do
not affect the derivation of the efficiency bound. The finite second moment
condition is used for deriving Gaussian limiting distributions in later
sections. The family of distributions \(\mathbf{Q}\) is a nonparametric family
since it is infinite dimensional. The parameter of interest is the average
treatment effect (ATE), \(\beta_{\ast} : \mathbf{Q} \to \mathbb{R}\) defined by
\begin{equation}
\beta_{\ast} (Q) = \mathbb{E}_{Q} [Y (1) - Y (0)] = \int \left( y_{1} - y_{0}
\right) \; Q (\mathrm{d} y_{0}, \mathrm{d} y_{1}, \mathrm{d} z).
\label{eqn--ate-Q}
\end{equation}
The true value of the ATE will be denoted \(\beta_{0} = \beta_{\ast} \left(
Q_{0} \right)\).
In both observational and experimental studies on the average effect of
treatment, observed outcomes result from the receipt of treatment. That is,
observed outcomes are given by \(Y\) defined by
\begin{equation*}
Y = Y (1) \cdot A + Y (0) \cdot (1 - A)
\end{equation*}
where \(A\) is a Bernoulli random variable denoting the treatment received. In
experimental settings, the conditional distribution of \(A\) given the baseline
covariates \(Z\) is assumed to be fully known to the experimenter. The target
observed population in the special case of a CAR experiment is the distribution
of the vector \(X^{\prime} = \left( Y, A, Z^{\prime} \right)\) which satisfies
\th\ref{asm--car-pop-ssra} below.
\begin{assumption}[Simple Stratified Randomization Assignment]
\th\label{asm--car-pop-ssra}
There exists a \(\mathcal{S} \in \mathbb{N}\), a measurable function
\(\mathbb{S} : \mathbb{R}^{k} \to \mathbb{N}_{\mathcal{S}}\) and a vector \(\pi
= (\pi (1), \dots, \pi (\mathcal{S}))\) with \(\pi (s) \in (0, 1)\) for every
\(s \in \mathbb{N}_{\mathcal{S}}\), which are all known to the experimenter. The
observable random vector \(X^{\prime} = (Y, A, Z^{\prime})\) is constructed from
\(W\) and its components satisfy
\begin{align}
& Y = Y (1) \cdot A + Y (0) \cdot (1 - A),
\label{eqn--pop-obsoutproc} \\
& [(Y (0), Y (1), Z) \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} A] | \mathbb{S} (Z),
\label{eqn--car-pop-A-exog} \\
& [A | \mathbb{S} (Z) = s] \sim \mathrm{Bernoulli} (\pi (s)).
\label{eqn--car-pop-A-cond-Bern}
\end{align}
\end{assumption}
The function \(\mathbb{S}\) is used to construct a (measurable) finite partition
of the support of the covariates \(Z\) and thus, \(\mathcal{S}\) is an upper
bound on the number of strata. The condition in \eqref{eqn--car-pop-A-exog}
requires that assignment to treatment be exogenous to both potential outcomes
and the remaining variation in covariates given stratum labels. The condition in
\eqref{eqn--car-pop-A-cond-Bern} requires that treatment assignment
proportions correspond to the pre-specified target probabilities in \(\pi\). In
finite samples, the covariate adaptive randomization schemes that will be
described in the next subsection all provide different ways to target the
experiment in \th\ref{asm--car-pop-ssra}.
\begin{remark}
\th\label{rem--identification-ate-P}
Let \(\mathbf{P}\) denote the family of distributions for the random vector
\(X\) that is determined by \th\ref{asm--Q} and \th\ref{asm--car-pop-ssra}. Let
\(P_{0}\) denote the distribution in \(\mathbf{P}\) that corresponds to the true
population distribution \(Q_{0} \in \mathbf{Q}\). It is straightforward to show
that the ATE in \eqref{eqn--ate-Q} is nonparametrically identifiable (see
\citet[Definition 3.2]{2007matzkinNonparametricIdentification}) via the map
\(\beta : \mathbf{P} \to \mathbb{R}\)
\begin{equation}
\beta (P) = \mathbb{E}_{P} \left[ \frac{Y \cdot A}{\pi (\mathbb{S} (Z))} -
\frac{Y \cdot (1 - A)}{1 - \pi (\mathbb{S} (Z))} \right].
\label{eqn--ate-P}
\end{equation}
That is, for a given \(Q \in \mathbf{Q}\), if \(P (Q)\) denotes the distribution
in \(\mathbf{P}\) formed from \(Q\) and \th\ref{asm--car-pop-ssra}, then
\(\beta_{\ast} (Q) = \beta (P (Q))\). In particular, the true value of the
average treatment effect can be written as
\begin{equation}
\beta_{0} = \beta_{\ast} \left( Q_{0} \right) = \beta \left( P_{0} \right).
\label{eqn--true-ATE}
\end{equation}
\end{remark}
\subsection{Sampling framework for CAR experiments.}
\label{subsec--CAR-sample}
In this subsection, we describe the assumptions maintained for the sampling
process that produces observed data in a CAR experiment.
\begin{assumption}
\th\label{asm--iid-Q}
\(\mathbf{W} = \left\{ W_{i} : i \in \mathbb{N} \right\}\) is a sequence of
independent and identically distributed (i.i.d.) random \(\mathbb{R}^{2 +
k}\)-vectors with \(W_{i} \sim Q_{0}\) for every \(i \in \mathbb{N}\). The first
\(n \in \mathbb{N}\) elements of \(\mathbf{W}\), denoted by
\(\mathbf{W}_{n}^{\prime} = \left( W_{1}, \dots, W_{n} \right)\), are the
potential outcome and covariate values for observations in the sample.
\end{assumption}
\th\ref{asm--iid-Q} states that potential outcome and covariate values for
observations in the sample are drawn at random from the population distribution
\(Q_{0}\). That is, if both potential outcomes and covariates were all
observable, the experimenter would have a representative sample from the
underlying population. This assumption is maintained in the recent literature on
CAR experiments as well as optimal experimental designs, for instance in
\citet{2018bugniInferenceCovariateAdaptiveRandomization},
\citet{2019bugniInferenceCovariateAdaptiveRandomization},
\citet{2022tabord-meehanStratificationTreesAdaptive},
\citet{2022baiOptimalityMatchedPairDesigns} and
\citet{2021cytrynbaumDesigningRepresentativeBalanced}. Given the baseline
covariates and corresponding strata, the experimenter chooses a vector of
treatment assignments \(\mathbf{A}_{n}^{\prime} = \left( A_{n 1}, \dots, A_{nn}
\right)\) which is a random vector supported in \({\{0, 1\}}^{n}\). The
experimenter has full control over the distribution of treatment assignments and
treatment assignments do not have to be i.i.d. We introduce the following
notation for convenience. For each \(s \in \mathbb{N}_{\mathcal{S}}\) and \(a
\in \{0, 1\}\) the stratum size and stratum treatment group size are
respectively
\begin{equation}
N_{n} (s) = \sum_{i = 1}^{n} \mathbb{I} \left\{ \mathbb{S} \left( Z_{i}
\right) = s \right\} \quad \text{and} \quad N_{n} (a, s) = \sum_{i = 1}^{n}
\mathbb{I} \left\{ A_{n i} = a, \mathbb{S} \left( Z_{i} \right) = s \right\}.
\label{eqn--strat-treat-size}
\end{equation}
We maintain the following assumptions about the observed data in a covariate
adaptive randomized experiment.
\begin{assumption}
\th\label{asm--treat-strat}
Let \(\mathcal{S}\), \(\mathbb{S}\) and \(\pi\) be as in
\th\ref{asm--car-pop-ssra}. For a sample of size \(n \in \mathbb{N}\), the
observed data are \(\mathbf{X}_{n}^{\prime} = \left( X_{n1}, \dots, X_{nn}
\right)\) with individual observations \(X_{n i} = \left( Y_{n i}, A_{n i},
Z_{i} \right)\) for each \(i \in \mathbb{N}_{n}\). The observed outcomes and
treatment assignment mechanism satisfy the following.
\begin{enumerate}[label = (\alph*)]
\item \label{asm--outcomes}
The observed outcomes are given by
\begin{equation}
Y_{n i} = Y_{i} (1) A_{n i} + Y_{i} (0) \left( 1 - A_{n i} \right).
\label{eqn--obsoutproc}
\end{equation}
\item \label{asm--treat-exog}
For every \(n \in \mathbb{N}\), \(\left[\mathbf{W}_{n} \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \mathbf{A}_{n}
\right] | \mathbf{S}_{n}\) where \(\mathbf{S}_{n}^{\prime} = \left( S_{1},
\dots, S_{n} \right)\) and \(S_{i} =
\mathbb{S} \left( Z_{i} \right)\) for each \(i \in \mathbb{N}_{n}\).
\item \label{asm--cov-adapt}
For every \(n \in \mathbb{N}\), the conditional distribution of treatment
assignments given the profile of sample strata
\begin{equation}
\alpha_{n} \left( \mathbf{a}_{n} \middle| \mathbf{s}_{n} \right) := \Pr
\left( \mathbf{A}_{n} = \mathbf{a}_{n} \middle| \mathbf{S}_{n} =
\mathbf{s}_{n} \right) \qquad \forall \mathbf{a}_{n} \in \{0, 1\}^{n},
\forall \mathbf{s}_{n} \in \mathbb{N}_{\mathcal{S}}^{n}
\label{eqn--cond-dist-A-S-sample}
\end{equation}
is completely known to the experimenter and does not depend on the population
distribution \(Q_{0}\).
\item \label{asm--treat-prop}
Under \(Q_{0}\) defined in \th\ref{asm--Q} and \(\alpha_{n} (\cdot | \cdot)\)
chosen by the experimenter above, for each \(s \in \mathbb{N}_{\mathcal{S}}\),
\begin{equation}
\frac{N_{n} (1, s)}{N_{n} (s)} \overset{\mathrm{p}}{\to} \pi (s).
\label{eqn--treat-strat-prop-conv-target}
\end{equation}
\end{enumerate}
\end{assumption}
Denote the distribution of \(\mathbf{X}_{n}\) by \(P_{0, n}\). Note that \(P_{0,
n}\) is determined by the population distribution \(Q_{0}\), the stratification
scheme \(\mathbb{S}\), equation \eqref{eqn--obsoutproc}, and the choice of
randomization scheme. \th\ref{asm--treat-strat} places the same set of
restrictions on \(P_{0, n}\) as in
\citet{2019bugniInferenceCovariateAdaptiveRandomization} on the relationship
between the randomization scheme, the underlying population and the
strata. \th\ref{asm--treat-strat} \ref{asm--outcomes} requires that the observed
outcome for a given observation is exactly the potential outcome associated with
the assigned treatment (as in equation \eqref{eqn--pop-obsoutproc} for the
target experiment population
\(\mathbf{P}\)). \th\ref{asm--treat-strat} \ref{asm--treat-exog} requires that
the treatment assignment be ignorable (or exogenous) given the strata. In
particular, the assignment scheme can only be a function of the stratum labels
and a randomization device exogenous to the information contained in the
potential outcomes and the baseline covariates beyond that afforded by the
strata. This assumption is analogous to \eqref{eqn--car-pop-A-exog} in
\th\ref{asm--car-pop-ssra}. Its main use is to guarantee identification of the
ATE within each stratum and it is additionally a sufficient condition to
identify the overall ATE. \th\ref{asm--treat-strat} \ref{asm--cov-adapt}
requires that the randomization procedure be fully known to the
experimenter. This assumption plays a role in providing a simple
characterization of the joint distribution of the observed data and allows us to
avoid mathematical complications when talking about parametric sub-models during
the discussion of semiparametric efficiency. Finally,
\th\ref{asm--treat-strat} \ref{asm--treat-prop} requires the randomization
procedure to reach the target treatment proportion within a given stratum at
least asymptotically in the sense of convergence in probability.
This is analogous to
\eqref{eqn--car-pop-A-cond-Bern} in \th\ref{asm--car-pop-ssra}. There are a
number of examples of randomization schemes which will satisfy the requirements
imposed by \th\ref{asm--treat-strat}. We provide two examples of such
randomization schemes.
\begin{example}[Simple Stratified Random Assignment (SSRA)]
\th\label{eg--sra}
Let the assignments, \(\left\{ A_{i} \right\}_{i = 1}^{n}\), be i.i.d. Bernoulli
random variables such that \(\Pr \left( A_{i} = 1 | S_{i} = s \right) = \pi
(s)\) and \([(Y_{i} (0), Y_{i} (1), Z_{i}) \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} A_{i}] | S_{i}\). The
observations \(X_{i}^{\prime} = \left( Y_{i}, A_{i}, Z_{i}^{\prime} \right)\)
generated by this process form an i.i.d. sample from \(P_{0}\) defined in
\th\ref{asm--car-pop-ssra}. The conditional mass function \(\alpha_{n} (\cdot |
\cdot)\) here is
\begin{equation}
\alpha_{n} \left( \mathbf{a}_{n} \middle| \mathbf{s}_{n} \right) = \prod_{i =
1}^{n} \pi \left( s_{i} \right)^{a_{i}} \left( 1 - \pi \left( s_{i} \right)
\right)^{1 - a_{i}} \qquad \forall \mathbf{a}_{n} \in \{0, 1\}^{n}, \forall
\mathbf{s}_{n} \in \mathbb{N}_{\mathcal{S}}^{n}, \forall n \in \mathbb{N}.
\label{eqn--cond-dist-A-S-SSRA}
\end{equation}
\th\ref{asm--treat-strat} \ref{asm--treat-prop} can be verified by appealing to
the Strong Law of Large Numbers.
\end{example}
\begin{example}[Stratified Permuted Block Randomization (SPBR)]
\th\label{eg--sbr}
Denote the integer floor function by \(\lfloor \cdot \rfloor\). Within stratum
\(s\), let
\begin{equation*}
N_{n} (1, s) = \left\lfloor \pi (s) N_{n} (s) \right\rfloor \quad \text{and}
\quad \frac{1}{c_{s, n}} = \binom{N_{n} (s)}{N_{n} (1, s)}.
\end{equation*}
There are \(c_{s, n}^{- 1}\) distinct subsets (or blocks) of size \(N_{n} (1,
s)\) from the overall stratum which has size \(N_{n} (s)\). We can choose one of
these blocks uniformly at random, i.e. each distinct block meeting the size
requirements gets assigned a probability mass of \(c_{n
s}\). \th\ref{asm--treat-strat} \ref{asm--treat-exog} can be satisfied by
ensuring the randomization device used to choose the treatment block is
independent to any outcome and covariate information within strata. Since
\(|N_{n} (1, s) - \pi (s) N_{n} (s)| \leq 1\) almost surely for every
\(s \in \mathbb{N}_{\mathcal{S}}\) and \(N_{n} (s) \overset{\mathrm{a.s.}}{\to}
\infty\) if \(Q (\mathbb{S} (Z) = s) > 0\), \th\ref{asm--treat-strat}
\ref{asm--treat-prop} is also satisfied. Additionally, for each \(\mathbf{a}_{n}
\in \{0, 1\}^{n}\), \(\mathbf{s}_{n} \in \mathbb{N}_{\mathcal{S}}^{n}\) and \(n
\in \mathbb{N}\),
\begin{equation}
\alpha_{n} \left( \mathbf{a}_{n} \middle| \mathbf{s}_{n} \right) = \prod_{s =
1}^{\mathcal{S}} \binom{\sum_{i = 1}^{n} \mathbb{I} \left\{ s_{i} = s
\right\}}{\sum_{i = 1}^{n} a_{i} \mathbb{I} \left\{ s_{i} = s \right\}}^{- 1}
\mathbb{I} \left\{ \sum_{i = 1}^{n} a_{i} \mathbb{I} \left\{ s_{i} = s
\right\} = \left\lfloor \pi (s) \sum_{i = 1}^{n} \mathbb{I} \left\{ s_{i} = s
\right\} \right\rfloor \right\}.
\label{eqn--cond-dist-A-S-SPBR}
\end{equation}
\end{example}
Both examples satisfy the requirement imposed by \th\ref{asm--treat-strat}
\ref{asm--treat-prop} for approaching the stratum-specific target proportions,
but they do so at different rates of convergence. The SSRA method provides
convergence to the stratum-specific targets at a \(n^{- \frac{1}{2}}\) rate due
to the Lindeberg-L\'evy Central Limit Theorem. That is,
\begin{equation*}
\left( \frac{N_{n} (1, s)}{N_{n} (s)} - \pi (s) : s = 1, \dots, \mathcal{S}
\right) = O_{\mathrm{p}} \left( \frac{1}{\sqrt{n}} \right) \text{ under SSRA.}
\end{equation*}
It can be shown using the Strong Law of Large Numbers that SPBR achieves the
same targeting property at a faster \(n^{- 1}\) rate, so that
\begin{equation*}
\left( \frac{N_{n} (1, s)}{N_{n} (s)} - \pi (s) : s = 1, \dots, \mathcal{S}
\right) = O_{\mathrm{p}} \left( \frac{1}{n} \right) = o_{\mathrm{p}} \left(
\frac{1}{\sqrt{n}} \right) \text{ under SPBR.}
\end{equation*}
When a randomization scheme achieves this faster convergence property we say
that it achieves ``strong balance''. Numerous other kinds of stratified
randomization techniques have been developed and analyzed in the literature on
randomized control trials --- for instance Efron's biased coin design and Wei's
urn --- see \citet{2018bugniInferenceCovariateAdaptiveRandomization} and
references therein for detailed descriptions. Furthermore, the treatment
assignments are not required to be i.i.d. and hence, the observed outcomes are
all potentially non-i.i.d. For instance, with the SPBR procedure illustrated in
\th\ref{eg--sbr}, assignments are independent across strata, but they exhibit
dependence within strata. Indeed, most (if not all) assignment schemes that
achieve strong balance will result in treatment assignments and observed
outcomes that are non-i.i.d. However, we will show that the same efficiency
bound as that for the i.i.d. sampling scheme in \th\ref{eg--sra} holds for all
CAR schemes that satisfy \th\ref{asm--treat-strat}. Finally while the
assumptions are satisfied by a fairly broad class of randomization schemes,
there are ones that violate them that are used in practice, e.g. the
minimization method of \citet{1975pocockSequentialTreatmentAssignment}.
Developing efficiency theory for these methods may be of interest but they are
not covered in this work.
\section{Semiparametric efficiency bound}
\label{sec--speb}
In this section, we first state the semiparametric efficiency bound for the ATE
and discuss it briefly. We also discuss the derivation of the bound in two
subsections. As stated in the introduction, of the efficiency bounds established
in the literature, the closest one to our setting is that of
\citet{1998hahnRolePropensityScore}. This is derived under i.i.d. observational
data under the assumption of ignorable treatment assignment conditional on
observable covariates (\citet[Section 1.3]{1983rosenbaumCentralRolePropensity}).
The \citet{1998hahnRolePropensityScore} bound does not apply immediately to our
setting since the treatment assignment rule is allowed to depend on the entire
profile of sample strata. For instance, it is not clear \emph{a priori} if the
choice of assignment mechanism will affect the efficiency bound since it can
clearly influence limit behavior of estimators as in
\citet{2018bugniInferenceCovariateAdaptiveRandomization,
2019bugniInferenceCovariateAdaptiveRandomization}. Furthermore, even when
treatment assignments are i.i.d. in a CAR context, the covariates being
conditioned on during assignment are the strata, so it is again not immediate
from the bound in \citet{1998hahnRolePropensityScore} what role the additional
baseline covariates can play in providing efficiency gains. The SPEB derived
here provides an answer to these questions. Before stating our bound, let \(m
\left( a, z; Q \right) = \mathbb{E}_{Q} [Y (a) | Z = z]\) for each \(a
\in \{0, 1\}\), and for brevity, let \(m_{\ast} (a, z) = m \left( a, z;
Q_{0} \right)\) Furthermore, define
\begin{equation}
\mathbb{V}_{\ast} = \mathbb{E} \left[ \frac{\mathrm{Var} [Y (1) | Z]}{\pi
(\mathbb{S} (Z))} + \frac{\mathrm{Var} [Y (0) | Z]}{1 - \pi (\mathbb{S} (Z))}
+ \left\{ m_{\ast} (1, Z) - m_{\ast} (0, Z) - \beta_{0} \right\}^{2} \right].
\label{eqn--speb-car}
\end{equation}
The following theorem establishes that \(\mathbb{V}_{\ast}\) above is the SPEB.
\begin{theorem}
\th\label{thm--speb-car}
Under \th\ref{asm--Q,asm--iid-Q,asm--treat-strat}, the semiparametric efficiency
bound for estimating the ATE from a CAR experiment is \(\mathbb{V}_{\ast}\) in
\eqref{eqn--speb-car}.
\end{theorem}
The proof of \th\ref{thm--speb-car} provides formal justification for
\(\mathbb{V}_{\ast}\) as an efficiency bound by appealing to extensions of
the famed convolution (\citet{1970hajekCharacterizationLimitingDistributions})
and local asymptotic minimax theorems (\citet{1972hajekLocalAsymptoticMinimax})
to semiparametric problems. The efficiency bound in \eqref{eqn--speb-car} is
that of \citet{1998hahnRolePropensityScore} for i.i.d. data, if treatment is
assigned conditional on baseline covariates according to the propensity score
\(\pi (\mathbb{S} (\cdot))\). The intuition for this is that even though
randomization happens conditional on strata, the baseline covariates are
independent to treatment assignment after conditioning on strata
(\th\ref{asm--treat-strat} \ref{asm--treat-exog}). Thus, the minimum possible
limit variance in estimating the ATE can be achieved by utilizing any additional
information the baseline covariates may have about potential outcomes. This can
be better understood by considering what happens if only information from the
strata are used during estimation. For instance, consider the fully saturated
regression estimator of the ATE from
\citet{2019bugniInferenceCovariateAdaptiveRandomization}. This estimator is
constructed from regressing observed outcomes \(Y_{n i}\) on the stratum
indicators and their interactions with treatment status \(A_{n i}\). The
coefficients from this regression are combined to then form the ATE
estimate. The resulting estimator has a ``weighted difference in means''
form:
\begin{equation}
\widehat{\beta}_{n, \mathrm{SAT}} = \sum_{s = 1}^{\mathcal{S}} \frac{N_{n}
(s)}{n} \cdot \left[ \frac{\sum_{i = 1}^{n} Y_{n i} \cdot A_{n i} \cdot
\mathbb{I} \left( S_{i} = s \right)}{N_{n} (1, s)} - \frac{\sum_{i = 1}^{n}
Y_{n i} \cdot \left( 1 - A_{n i} \right) \cdot \mathbb{I} \left( S_{i} = s
\right)}{N_{n} (0, s)} \right].
\label{eqn--est-sat}
\end{equation}
\citet{2019bugniInferenceCovariateAdaptiveRandomization} show that
\(\widehat{\beta}_{n, \mathrm{SAT}}\) is asymptotically normal with mean-zero
and limit variance given by
\begin{equation}
\mathbb{V}_{\mathrm{SAT}} = \mathbb{E} \left[ \frac{\mathrm{Var} [Y (1) |
\mathbb{S} (Z)]}{\pi (\mathbb{S} (Z))} + \frac{\mathrm{Var} [Y (0) |
\mathbb{S} (Z)]}{[1 - \pi (\mathbb{S} (Z))]} + \left\{ \mathbb{E} [Y (1) |
\mathbb{S} (Z)] - \mathbb{E} [Y (0) | \mathbb{S} (Z)] - \beta_{0} \right\}^{2}
\right].
\label{eqn--speb-sat}
\end{equation}
Our results thus establish that the fully saturated regression estimator of the
ATE achieves the SPEB among all estimators that only use information from the
stratum labels. This follows from comparing \eqref{eqn--speb-car} and
\eqref{eqn--speb-sat} and noting that we have \(\mathbb{V}_{\mathrm{SAT}} =
\mathbb{V}_{\ast}\) if \(Z = \mathbb{S} (Z)\) (up to information preserving
relabelling of strata). The latter condition amounts to only using information
given in stratum labels. An additional implication is that when conditional
means or conditional variances of the potential outcomes given baseline
covariates exhibit variation within strata, \(\mathbb{V}_{\ast}\) can
be a strict improvement over \(\mathbb{V}_{\mathrm{SAT}}\) - an observation also
made by \citet{1998hahnRolePropensityScore}. This is confirmed by
Monte Carlo simulation evidence in Section \ref{sec--sims} in which
\(\mathbb{V}_{\ast}\) can be approximately half of \(\mathbb{V}_{\mathrm{SAT}}\)
in some simulation designs (i.e. there is up to a 50\% possible reduction in
asymptotic variance).
We provide a few additional comments on the SPEB in \eqref{eqn--speb-car}. Note
that the choice of assignment mechanism only affects the bound through the
choice of the target assignment proportions in \(\pi (\cdot)\). All else held
equal, the SPBR assignment mechanism in \th\ref{eg--sbr}, or any other procedure
satisfying \th\ref{asm--treat-strat}, will produce the same SPEB as i.i.d. draws
from \(\mathbf{P}\) (\th\ref{eg--sra}). Furthermore,
\citet{1998hahnRolePropensityScore} notes that knowledge of the propensity score
is ancillary to the SPEB for the ATE. In the context of CAR, we will show that
this knowledge is useful for achieving the SPEB during estimation. An important
intermediate product of our derivation for this purpose is the \emph{efficient
influence function}, \(\varphi_{0}\), which is defined by
\begin{equation}
\varphi_{0} (y, a, z) = \frac{a}{\pi (\mathbb{S} (z))} \left[ y - m_{\ast} (1,
z) \right] - \frac{1 - a}{1 - \pi (\mathbb{S} (z))} \left[ y - m_{\ast} (0, z)
\right] + \left[ m_{\ast} (1, z) - m_{\ast} (0, z) - \beta_{0} \right].
\label{eqn--eif}
\end{equation}
\(\varphi_{0}\) is called the efficient influence function because \(\mathbb{E}
\left[ \varphi_{0} (Y, A, Z) \right] = 0\) and its variance is the SPEB in
\eqref{eqn--speb-car}, i.e. \(\mathbb{E} \left[ \varphi_{0} (Y, A, Z)^{2}
\right] = \mathbb{V}_{\ast}\). This function is a key ingredient to achieving
the efficiency bound as illustrated in Section
\ref{sec--efficient-estimation}. In the subsequent two subsections, we provide
some additional discussion of how \eqref{eqn--speb-car} is derived, but
mathematical details are left to the appendix.
\subsection{A product characterization of distribution of observed data}
We establish here that the joint distribution of the observed data
\(\mathbf{X}_{n}\), \(P_{0, n}\), has a product structure even though outcomes
and treatment assignments are potentially non-i.i.d. This product structure will
then imply that \(P_{0, n}\) has a density that also has a product form. To that
end, we first define a dominating measure for a single observation. For Borel
sets \(\mathcal{Y} \subseteq \mathbb{R}\) and \(\mathcal{Z} \subseteq
\mathbb{R}^{k}\) and for \(\mathcal{A} \subseteq \{0, 1\}\), define the measures
\(\nu_{a}\) with \(a \in \{0, 1\}\) and \(\nu\) by
\begin{equation}
\begin{split}
\nu_{a} (\mathcal{Y} \times \mathcal{Z}) =
& \ \mu_{a} (\mathcal{Y}) \cdot \mu_{Z} (\mathcal{Z}), \\
\nu (\mathcal{Y} \times \mathcal{A} \times \mathcal{Z}) =
& \ \nu_{1} (\mathcal{Y} \times \mathcal{Z}) \cdot \mathbb{I} (1 \in
\mathcal{A}) + \nu_{0} (\mathcal{Y} \times \mathcal{Z}) \cdot \mathbb{I}
(0 \in \mathcal{A}).
\end{split}
\label{eqn--nu}
\end{equation}
Let \(q_{a} (y , z; Q)\) denote the marginal density of \((Y (a), Z)\) if the
underlying population has distribution \(Q \in \mathbf{Q}\). Furthermore, let
\(P_{n} (\cdot; Q)\) denote the distribution of the observed data
\(\mathbf{X}_{n}\) if \(\mathbf{W}_{n}\) is an i.i.d. sample from \(Q \in
\mathbf{Q}\). Under this notation, we have \(P_{0, n} \equiv P_{n} \left( \cdot;
Q_{0} \right)\).
\begin{lemma}
\th\label{lem--prod-struct}
Let \(\mathbf{Y}_{n}^{\prime} = \left( Y_{n1}, \dots, Y_{nn} \right)\) and
\(\mathbf{Z}_{n}^{\prime} = \left( Z_{1}, \dots, Z_{n} \right)\). For each
\(\mathbf{y}_{n} \in \mathbb{R}^{n}\), \(\mathbf{a}_{n} \in {\{0, 1\}}^{n}\),
\(\mathbf{s}_{n} \in {\mathbb{N}_{\mathcal{S}}}^{n}\) and
\(\mathbf{z}_{n}^{\prime} = \left( z_{1}, \dots, z_{n} \right) \in \mathbb{R}^{k
\times n}\), denote the event
\begin{equation*}
E_{n} \left( \mathbf{y}_{n}, \mathbf{a}_{n}, \mathbf{z}_{n}, \mathbf{s}_{n}
\right) = \left\{ \mathbf{Y}_{n} \leq \mathbf{y}_{n}, \mathbf{A}_{n} =
\mathbf{a}_{n}, \mathbf{Z}_{n} \leq \mathbf{z}_{n}, \mathbf{S}_{n} =
\mathbf{s}_{n} \right\}.
\end{equation*}
where the inequalities are understood to be element-wise. Then under
\th\ref{asm--Q}, \th\ref{asm--car-pop-ssra}, \th\ref{asm--treat-strat} and
\eqref{eqn--obsoutproc}, if \(\mathbf{W}_{n}\) is an i.i.d. sample from \(Q \in
\mathbf{Q}\),
\begin{equation}
P_{n} \left( E_{n} \left( \mathbf{y}_{n}, \mathbf{a}_{n}, \mathbf{z}_{n},
\mathbf{s}_{n} \right); Q \right) = \ \alpha_{n} \left( \mathbf{a}_{n}
\middle| \mathbf{s}_{n} \right) \times \prod_{i = 1}^{n} \left\{
\begin{array}{l}
Q \left( Y_{i} (1) \leq y_{i}, Z_{i} \leq z_{i}, S_{i} = s_{i}
\right)^{a_{i}} \\
\times Q \left( Y_{i} (0) \leq y_{i}, Z_{i} \leq z_{i}, S_{i} = s_{i}
\right)^{1 - a_{i}}
\end{array} \right\}.
\label{eqn--data-joint-dist-prod}
\end{equation}
Thus, \(P_{n} (\cdot; Q)\) is absolutely continuous against the \(n\)-fold
product measure formed from \(\nu\) in \eqref{eqn--nu} with density
\begin{equation}
p_{n} \left( \mathbf{y}_{n}, \mathbf{a}_{n}, \mathbf{z}_{n}; Q \right) =
\alpha_{n} \left( \mathbf{a}_{n} \middle| \mathbf{s}_{n} \right) \prod_{i =
1}^{n} q_{1} \left( y_{i}, z_{i}; Q \right)^{a_{i}} q_{0}
\left(y_{i}, z_{i}; Q \right)^{1 - a_{i}}.
\label{eqn--data-joint-density-prod}
\end{equation}
where in \eqref{eqn--data-joint-density-prod}, we restrict
\(\mathbf{s}_{n}^{\prime} = \left( \mathbb{S} \left( z_{1} \right), \dots,
\mathbb{S} \left( z_{n} \right) \right)\).
\end{lemma}
The main consequence of \th\ref{lem--prod-struct} is as follows. Let
\(\mathcal{P} = \{\mathbf{P}_{n} : n \in \mathbb{N}\}\) be the sequence of
families of distributions for the observed data \(\mathbf{X}_{n}\) such that for
each \(n \in \mathbb{N}\), every member of \(\mathbf{P}_{n}\) is determined by
an i.i.d. sample from some distribution in \(\mathbf{Q}\) (as defined in
\th\ref{asm--Q}), the observed outcome equation \eqref{eqn--obsoutproc} and a
treatment assignment mechanism that satisfies \th\ref{asm--treat-strat}. Any
member of \(\mathbf{P}_{n}\) must then satisfy the product structure in
\eqref{eqn--data-joint-dist-prod} and \eqref{eqn--data-joint-density-prod}.
Furthermore, different covariate adaptive randomization schemes give rise to
different sequences of families \(\mathcal{P}\) that only vary according to the
choice of the sequence of conditional treatment assignment distributions
\(\left\{ \alpha_{n} (\cdot) : n \in \mathbb{N} \right\}\).
\begin{remark}[Implications for Simple Stratified Random Assignment]
\th\label{rem--prod-struct-ssra}
Consider the family \(\mathbf{P}\) defined by \th\ref{asm--car-pop-ssra} in
\th\ref{rem--identification-ate-P}. An immediate additional consequence of
\th\ref{lem--prod-struct} is that any distribution \(P \in \mathbf{P}\) has a
density against \(\nu\) of the form
\begin{equation*}
p (y, a, z; P) = \left[ q_{1} (y, z) \cdot \pi (\mathbb{S} (z)) \right]^{a}
\left[ q_{0} (y, z) \cdot \pi (1 - \mathbb{S} (z)) \right]^{1 - a}.
\end{equation*}
Recall also that in \th\ref{eg--sra}, we consider a treatment assignment rule
that essentially produces an i.i.d. sample of size \(n\) from a distribution in
\(\mathbf{P}\). In this case, we would of course have that \(\mathbf{P}_{n} =
\mathbf{P}^{n}\) where the latter is the set of all \(n\)-fold product measures
formed from some measure in \(\mathbf{P}\).
\end{remark}
\subsection{Deriving the semiparametric efficiency bound}
In this subsection, we first introduce and discuss definitions of parametric
submodels and regular parametric submodels which are essential concepts to the
derivation of the semiparametric efficiency bound. It is then shown in
\th\ref{lem--parametric-submodel-lan} that the log likelihood ratios in regular
parametric submodels under \(1 / \sqrt{n}\) local alternatives exhibit the local
asymptotic normality (LAN) property of
\citet{1960lecamLocallyAsymptoticallyNormal}. An efficiency bound is
established for parametric submodels
(\th\ref{lem--parametric-submodel-efficiency}). Then, via a discussion on
differentiability of the ATE in regular parametric submodels we provide an
intuitive description of how the parametric efficiency bound is extended to the
semiparametric efficiency bound presented in \th\ref{thm--speb-car}.
\begin{definition}
\th\label{def--parametric-submodel}
A parametric submodel of \(\mathcal{P}\) is a sequence of families of
distributions, \(\mathcal{P}^{0} = \left\{ \mathbf{P}_{n}^{0} : n \in \mathbb{N}
\right\}\) that satisfies the following conditions.
\begin{enumerate}[label=(\alph*)]
\item \label{def--parametric-submodel-subset}
For every \(n \in \mathbb{N}\), \(\mathbf{P}^{0}_{n} \subseteq
\mathbf{P}_{n}\). That is, every member, \(P_{n} \in \mathbf{P}^{0}_{n}\) is
determined by a population distribution from a subfamily \(\mathbf{Q}^{0}
\subseteq \mathbf{Q}\) and satisfies \th\ref{asm--treat-strat}.
\item \label{def--parameteric-submodel-densities}
There is a \(\Theta \subseteq \mathbb{R}^{d}\) (where \(d \in \mathbb{N}\)),
and a map \(\theta \mapsto \left( q_{0} (\cdot; \theta), q_{1} (\cdot; \theta)
\right)\) from \(\Theta\) to \((Y (a), Z)\)-marginal densities of
\(\mathbf{Q}^{0}\) such that every member of \(\mathbf{P}^{0}_{n}\) has a
density (as in \th\ref{lem--prod-struct}) of the form
\begin{align*}
p_{n} \left( \mathbf{y}_{n}, \mathbf{a}_{n}, \mathbf{z}_{n}; \theta \right)
=
& \ \alpha_{n} \left( \mathbf{a}_{n} | \mathbf{s}_{n} \right) \cdot \prod_{i
= 1}^{n} q_{1} \left( y_{i}, z_{i}; \theta \right)^{a_{i}} q_{0} \left(
y_{i}, z_{i}; \theta \right)^{1 - a_{i}}, \\
\text{where } \mathbf{s}_{n}^{\prime} =
& \ \left( \mathbb{S} \left( z_{1} \right), \dots, \mathbb{S} \left( z_{n}
\right) \right)
\end{align*}
Additionally, the map \(\theta \mapsto \left( q_{0} (\cdot; \theta),
q_{1} (\cdot; \theta) \right)\) satisfies the \(Z\)-marginal
restriction: there exists a \(\mu_{Z}\)-density function \(g (\cdot; \theta)\)
(i.e. \(g (z; \theta) \geq 0\) for all \(z \in \mathbb{R}^{k}\) and \(\int g
(z; \theta) \mu_{Z} (\mathrm{d} z) = 1\)) such that
\begin{equation*}
\int_{\mathbb{R}} q_{0} (y, z; \theta) \; \mu_{0} \left( \mathrm{d} y
\right) = \int_{\mathbb{R}} q_{1} (y, z; \theta) \; \mu_{1} \left(
\mathrm{d} y \right) = g (z; \theta) \quad \forall \theta \in \Theta
\text{ for } z \in \mathbb{R}^{k} \ \mu_{Z} \text{-a.e.}
\end{equation*}
We write \(P_{\theta, n}\) for the element of \(\mathbf{P}_{n}^{0}\) whose
density is \(p_{n} (\cdot; \theta)\).
\item \label{def--parameteric-submodel-contains-true}
The population subfamily \(\mathbf{Q}^{0}\) contains the true distribution
\(Q_{0}\) so that \(P_{0, n} \in \mathbf{P}_{n}^{0}\). The parametrization
\(\theta \mapsto p_{n} (\cdot; \theta)\) identifies the true distribution so
that there is a unique \(\theta_{0} \in \Theta\) such that \(P_{0, n} =
P_{\theta_{0}, n}\).
\end{enumerate}
\end{definition}
In \th\ref{def--parametric-submodel}, condition
\ref{def--parametric-submodel-subset} requires that a parametric submodel be a
subset of the overall semiparametric model. Condition
\ref{def--parameteric-submodel-densities} requires that the parametrization
\(\theta \mapsto p_{n} (\cdot; \theta)\) produces densities of the same form as
those in \th\ref{lem--prod-struct}. The additional \(Z\)-marginal restriction
ensures that both densities \(q_{0} (\cdot; \theta)\) and \(q_{1}
(\cdot; \theta)\) can be derived from a common distribution \(Q_{\theta} \in
\mathbf{Q}^{0}\). We do not require the submodel to specify the joint density of
\((Y (0), Y (1), Z)\) since only the marginals with respect to \((Y (0), Z)\)
and \((Y (1), Z)\) are identified by the randomized experiment (essentially due
to \eqref{eqn--data-joint-dist-prod} in \th\ref{lem--prod-struct}). Finally,
\ref{def--parameteric-submodel-contains-true} requires the parametric submodel
to contain the true distribution \(P_{0, n}\). Note that a parametric submodel
does not parameterize the conditional mass function of the treatment
assignments, \(\alpha_{n}\), since it is known and does not depend on any
population unknowns. This is different to the case of observational data
considered in \citet{1998hahnRolePropensityScore} where the propensity score
is unknown and has to be parameterized.
The idea behind the use of parametric submodels is as follows. Since \(\Theta\)
(and therefore \(\mathbf{P}_{n}\)) in \th\ref{def--parametric-submodel} is
finite dimensional, estimation of the ATE restricted to \(\mathcal{P}^{0}\)
cannot be more difficult than estimation within \(\mathcal{P}\). One can imagine
estimation of the ATE by first estimating the nuisance parameter \(\theta_{0}\)
by maximum likelihood and then forming a plug-in estimate of the ATE. If
\(\mathcal{P}^{0}\) is a parametric submodel with a well-defined efficiency
bound, any semiparametric estimator for the ATE that is consistent and
asymptotically normal under the distributions in \(\mathcal{P}\) cannot have
asymptotic variance lower than the efficiency bound under
\(\mathcal{P}^{0}\). Taking the supremum over all parametric submodels, we get a
variance lower bound for the semiparametric model. The statement about
parametric submodels having well-defined efficiency bounds requires some
qualification. Establishing efficiency requires restricting our analysis to
\emph{regular} parametric submodels (in which efficiency bounds are
well-defined) and to \emph{regular} estimators (to rule out the phenomenon of
super-efficiency). We provide the definition of a regular parametric submodel in
our case below. The definition of a regular estimator can be found in
\citet[p. 115 and p. 365]{1998vandervaartAsymptoticStatistics} or
\citet[p. 413]{1996vandervaartWeakConvergenceEmpirical}.
\begin{definition}
\th\label{def--qmd}
Let \(\mathcal{P}^{0}\) be a parametric subfamily of \(\mathcal{P}\) as in
\th\ref{def--parametric-submodel}. \(\mathcal{P}^{0}\) is called \emph{regular}
at \(\theta_{0}\) if the parametrization \(\theta \mapsto \left( q_{0} (\cdot;
\theta), q_{1} (\cdot; \theta) \right)\) satisfies the following.
\begin{enumerate}[label=(\alph*)]
\item \(\Theta\) is an bounded and open subset of \(\mathbb{R}^{d}\).
\item \label{def--qmd-differentiability}
The maps \(\theta \to \sqrt{q_{a} (\cdot; \theta)}\) are differentiable in
the quadratic mean at \(\theta_{0}\). That is, for each \(a \in \{0, 1\}\),
there is a measurable function \(D_{a} : \mathbb{R}^{1 + k} \to
\mathbb{R}^{d}\) such that \(\int D_{a}^{\prime} D_{a} \mathrm{d} \nu_{a} <
\infty\) and
\begin{equation*}
\lim_{\|t\| \to 0} \|t\|^{- 2} \int \left( \sqrt{q_{a} (y, z; \theta_{0} +
t)} - \sqrt{q_{a} \left( y, z; \theta_{0} \right)} - D_{a} (y, z)^{\prime}
\left( \theta - \theta_{0} \right) \right)^{2} \nu_{a} (\mathrm{d} y,
\mathrm{d} z) = 0.
\end{equation*}
\item The information matrix at \(\theta_{0}\), \(\mathcal{I}_{a}\), is
non-singular, where
\begin{equation}
\mathcal{I}_{a} = 4 \int_{\mathbb{R}^{1 + k}} D_{a} (y, z) D_{a} (y,
z)^{\prime} \; \nu_{a} (\mathrm{d} y, \mathrm{d} z).
\label{eqn--parametric-info-mat}
\end{equation}
\end{enumerate}
\end{definition}
The log-likelihood in any parametric submodel is
\begin{equation}
\begin{split}
\ell_{n} (\theta) =
& \ \log p_{n} \left( \mathbf{X}_{n}; \theta \right) = \log \alpha_{n}
\left( \mathbf{A}_{n} \middle| \mathbf{S}_{n} \right) + \sum_{i = 1}^{n}
\ell \left( Y_{n i}, A_{n i}, Z_{i}; \theta \right), \\
\text{where } \ell \left( Y_{n i}, A_{n i}, Z_{i}; \theta \right) =
& \ A_{n i} \log q_{1} \left( Y_{n i} , Z_{i}; \theta \right) +
\left(1 - A_{n i} \right) \log q_{0} \left( Y_{n i}, Z_{i}; \theta
\right).
\end{split}
\label{eqn--sample-loglhood}
\end{equation}
When the parametric submodel is regular, the associated score function at the
truth, \(\theta_{0}\), is
\begin{equation}
\begin{split}
\dot{\ell}_{n} =
& \ \sum_{i = 1}^{n} \dot{\ell} \left( Y_{n i}, A_{n i}, Z_{i} \right), \\
\text{where } \dot{\ell} \left( Y_{n i}, A_{n i}, Z_{i} \right) =
& \ A_{n i} \dot{\ell}_{1} \left( Y_{n i}, Z_{i} \right) + \left( 1 - A_{n
i} \right) \dot{\ell}_{0} \left( Y_{n i}, Z_{i} \right), \\
\dot{\ell}_{a} (y, z) =
& \ 2 \frac{D_{a} (y, z)}{\sqrt{q_{a} \left( y, z; Q_{0}
\right)}} \cdot \mathbb{I} \left\{ q_{a} \left( y, z; Q_{0}
\right) > 0 \right\}.
\end{split}
\label{eqn--sample-score}
\end{equation}
Note that the score in \eqref{eqn--sample-score} is expressed in ``root
density'' terms since it is technically defined by the quadratic mean
differentiability assumption in \th\ref{def--qmd}.\footnote{When the parametric
submodels are such that the densities are continuously differentiable with
respect to \(\theta\), then the usual log-derivative form and
\eqref{eqn--sample-score} coincide. That is, we can write \(\dot{\ell}_{a}
(\cdot) = \nabla_{\theta} \left[ \log q_{a} \right] \left( \cdot; \theta_{0}
\right)\).} Let \(X^{\prime} = (Y, A, Z)\) be as in the CAR target experiment
defined by \th\ref{asm--car-pop-ssra}. Under CAR, the associated Fisher
information matrix will be shown to be
\begin{equation}
\begin{split}
\mathcal{I} =
& \ \mathbb{E} \left[ \dot{\ell} (Y, A, Z) \dot{\ell} (Y, A, Z)^{\prime}
\right] \\
= & \ \sum_{s = 1}^{\mathcal{S}} Q_{0} (\mathbb{S} (Z) = s) \left\{
\begin{array}{l}
\pi (s) \mathbb{E} \left[ \dot{\ell}_{1} (Y (1), Z)
\dot{\ell}_{1} (Y (1), Z)^{\prime} \middle| \mathbb{S}
(Z) = s \right] \\
+ (1 - \pi (s)) \mathbb{E} \left[ \dot{\ell}_{0} (Y (0), Z)
\dot{\ell}_{0} (Y (0), Z)^{\prime} \middle| \mathbb{S} (Z) = s \right]
\end{array}
\right\}.
\end{split}
\label{eqn--strata-info-mat}
\end{equation}
Note that the information matrix, \(\mathcal{I}\) in
\eqref{eqn--strata-info-mat} is related to the marginal information matrices
\(\mathcal{I}_{a}\) for \(a \in \{0, 1\}\) in \eqref{eqn--parametric-info-mat}
but adjusts these for stratification. Furthermore the score \(\dot{\ell}\) and
information matrix \(\mathcal{I}\) both depend on the choice of parametric
submodel since the quadratic mean derivatives \(D_{a}\) depend on the choice of
parametric submodel. \th\ref{lem--parametric-submodel-lan} below shows that the
main consequence of \th\ref{def--qmd} is that the likelihood ratio in any
regular parametric submodel has the usual quadratic expansion and exhibits the
local asymptotic normality (LAN) property of
\citet{1960lecamLocallyAsymptoticallyNormal} under general CAR procedures.
\begin{lemma}
\th\label{lem--parametric-submodel-lan}
Suppose the parametric submodel \(\mathcal{P}^{0}\) of \(\mathcal{P}\) is
regular at the true value \(\theta_{0}\). Then, \(\mathcal{P}^{0}\) has the
local asymptotic normality property of
\citet{1960lecamLocallyAsymptoticallyNormal} at \(\theta_{0} \in \Theta\). In
particular, for \(t \in \mathbb{R}^{d}\) such that \(\theta_{0} + n^{-
\frac{1}{2}} t \in \Theta\) for all \(n \in \mathbb{N}\), let the quadratic
remainder \(R_{n} \left( \theta_{0}, t \right)\) be defined by
\begin{equation}
\ell_{n} \left( \theta_{0} + n^{- \frac{1}{2}} t \right) - \ell_{n}
\left( \theta_{0} \right) = t^{\prime} \frac{1}{\sqrt{n}} \dot{\ell}_{n} -
\frac{1}{2} t^{\prime} \mathcal{I} t + R_{n} \left( \theta_{0}, t \right)
\label{eqn--lr-quad-breakdown}
\end{equation}
where \(\dot{\ell}_{n}\) is as in \eqref{eqn--sample-score} and
\(\mathcal{I}\) is defined as in \eqref{eqn--strata-info-mat}. Let \(G
\sim \mathcal{N} \left( \mathbf{0}_{d}, \mathcal{I} \right)\). Then the
following hold.
\begin{enumerate}[label=(\alph*)]
\item\label{lem--parametric-submodel-lan-wc}
Under \(P_{0, n}\), \(n^{- \frac{1}{2}} \dot{\ell}_{n}
\overset{\mathrm{d}}{\to} G\).
\item\label{lem--parametric-submodel-lan-uan}
For any given \(M > 0\),
\(R_{n} \left( \theta_{0}, t \right) \overset{\mathrm{p}}{\to} 0\) uniformly
over \(\|t\| \leq M\) under \(P_{0, n}\). That is, for any \(\varepsilon >
0\),
\begin{equation}
\lim_{n \to \infty} \sup_{\|t\| \leq M} P_{0, n} \left( \left| R_{n} \left(
\theta_{0}, t \right) \right| > \varepsilon \right) = 0.
\end{equation}
\end{enumerate}
\end{lemma}
The proof of \th\ref{lem--parametric-submodel-lan} in the appendix of this paper
uses the tools developed
\citet{2018bugniInferenceCovariateAdaptiveRandomization} and
\citet{2019bugniInferenceCovariateAdaptiveRandomization} to adapt the proof
of Proposition 2.1.2 in the appendix of
\citet{1998bickelEfficientAdaptiveEstimation} to the context of CAR. In
particular, a partial-sums empirical process argument is used to establish
conclusion \ref{lem--parametric-submodel-lan-wc} in
\th\ref{lem--parametric-submodel-lan}. Using similar arguments, the sample
information matrix is shown to converge in probability to \(\mathcal{I}\). The
remainder \(R_{n} \left( \theta_{0}, t \right)\) summarizes both the error
replacing the sample information matrix with its population (or limiting)
analogue as well as the error inherent in a second-order Taylor expansion. The
latter is shown to be negligible as well. Under the LAN phenomenon,
\th\ref{lem--parametric-submodel-efficiency} below establishes an efficiency
result for estimation of the nuisance parameter \(\theta_{0}\).
\begin{lemma}
\th\label{lem--parametric-submodel-efficiency}
Suppose \(\mathcal{P}^{0}\) is a parametric submodel of
\(\mathcal{P}\) that is regular at \(\theta_{0}\). Let \(\mathcal{I}\) be the
Fisher information matrix defined in \eqref{eqn--strata-info-mat}. Then
\(\mathcal{I}\) is non-singular and \(\mathcal{I}^{- 1}\) is the efficiency
bound for estimation of \(\theta_{0}\). In particular, let \(G \sim \mathcal{N}
\left( \mathbf{0}_{d}, \mathcal{I}^{- 1} \right)\). Then the following hold.
\begin{enumerate}[label=(\alph*)]
\item \label{lem--parametric-submodel-efficiency-convolution}
If \(\widehat{\theta}_{n}\) is a sequence of regular estimators of
\(\theta_{0}\), then
\begin{equation}
\sqrt{n} \left( \widehat{\theta}_{n} - \theta_{0} - \frac{t}{\sqrt{n}}
\right) \overset{\mathrm{d}}{\to} G + U.
\label{eqn--convolution-regular}
\end{equation}
where \(U\) is a random \(\mathbb{R}^{d}\)-vector independent to \(G\) that is
specific to the estimator sequence \(\widehat{\theta}_{n}\).
\item \label{lem--parametric-submodel-efficiency-lam}
Let \(\mathcal{L} : \mathbb{R}^{d} \to \mathbb{R}\) be a bowl-shaped loss
function, i.e. \(\mathcal{L}\) is non-negative, satisfies \(\mathcal{L} (y) =
\mathcal{L} (- y)\) and has convex sublevel sets. If \(\widehat{\theta}_{n}\)
is any sequence of estimators of \(\theta_{0}\), then
\begin{equation}
\lim_{M \to \infty} \liminf_{n \to \infty} \ \sup_{\theta_{n} \in \Theta,
\sqrt{n} \left\| \theta_{n} - \theta_{0} \right\| \leq M}
\mathbb{E}_{P_{\theta_{n}, n}} \left[ \mathcal{L} \left( \sqrt{n} \left(
\widehat{\theta}_{n} - \theta_{n} \right) \right) \right] \geq \mathbb{E}
\left[ \mathcal{L} \left( G \right) \right].
\label{eqn--theta-lam}
\end{equation}
\end{enumerate}
\end{lemma}
\th\ref{lem--parametric-submodel-efficiency} establishes that the usual
Cram\'er-Rao parametric information bound holds for estimation of \(\theta_{0}\)
in regular parametric submodels. This is true even under general CAR procedures
that may result in dependence within observed data.
\th\ref{lem--parametric-submodel-efficiency} is a direct consequence of the LAN
property established in \th\ref{lem--parametric-submodel-lan}. In
\th\ref{lem--parametric-submodel-efficiency},
conclusion \ref{lem--parametric-submodel-efficiency-convolution}
follows from the celebrated convolution theorem of
\citet{1970hajekCharacterizationLimitingDistributions} and
conclusion \ref{lem--parametric-submodel-efficiency-lam} follows from the local
asymptotic minimax theorem of \citet{1972hajekLocalAsymptoticMinimax}.
These are standard tools used to justify the parametric Cram\'er-Rao bound. Note
that in \th\ref{lem--parametric-submodel-efficiency} the random vectors \(G\)
and \(U\) are again specific to the parametric submodel chosen and in the case
of \(U\), also specific to the choice of estimator sequence.
\th\ref{lem--parametric-submodel-efficiency} establishes an efficiency bound for
estimation of the nuisance parameter \(\theta\) whereas the content of
\th\ref{thm--speb-car} is about estimation of the ATE, \(\beta\) in
\eqref{eqn--ate-P}. To relate \th\ref{lem--parametric-submodel-efficiency} to a
version of \th\ref{thm--speb-car}, we have to first show that \(\beta\) is a
pathwise differentiable parameter (see \citet[Definition 3.3.1,
p. 57]{1998bickelEfficientAdaptiveEstimation} or \citet[Section
3]{1990neweySemiparametricEfficiencyBounds}). We do this by considering regular
parametric submodels of the target CAR experiment, \(\mathbf{P}\),
as defined by \th\ref{asm--car-pop-ssra} and
\th\ref{rem--identification-ate-P}. For a regular parametric submodel
\(\mathbf{P}_{0} = \left\{ P_{\theta} : \theta \in \Theta \right\}\) of
\(\mathbf{P}\), the average treatment effect parameter is
\begin{equation*}
\gamma (\theta) := \beta \left( P_{\theta} \right) = \int_{\mathbb{R}^{1 + k}}
y \cdot q_{1} (y, z; \theta) \; \nu_{1} (\mathrm{d} y, \mathrm{d} z) -
\int_{\mathbb{R}^{1 + k}} y \cdot q_{0} (y, z; \theta) \; \nu_{0} (\mathrm{d}
y, \mathrm{d} z).
\end{equation*}
Adapting the pathwise differentiability arguments in
\citet{1998hahnRolePropensityScore} to the context of a CAR experiment, the
derivative (gradient) of \(\gamma (\theta)\) at \(\theta_{0}\) in any given
regular parametric submodel can be written as
\begin{equation}
\nabla \gamma \left( \theta_{0} \right) = \mathbb{E} \left[ \varphi_{0} (Y, A,
Z) \cdot \dot{\ell} (Y, A, Z) \right],
\label{eqn--path-derivative-beta}
\end{equation}
where \(\varphi_{0}\) is the efficient influence function from \eqref{eqn--eif}.
The full derivation of this adaptation is included the appendix for completeness
(\th\ref{lem--tangent-space-P,lem--pathwise-differentiability-ate}). Combining
\th\ref{lem--parametric-submodel-efficiency} and
\eqref{eqn--path-derivative-beta}, it follows via the usual ``delta method
style'' argument that the information bound for estimating the ATE in the
regular parametric submodel \(\mathcal{P}^{0}\) is
\begin{equation*}
\mathbb{V} \left( \mathcal{P}^{0} \right) = \mathbb{E} \left[ \varphi_{0} (Y,
A, Z) \dot{\ell} \left( Y, A, Z \right)^{\prime} \right] \mathbb{E} \left[
\dot{\ell} \left( Y, A, Z \right) \dot{\ell} \left( Y, A, Z \right)^{\prime}
\right]^{- 1} \mathbb{E} \left[\dot{\ell} \left( Y, A, Z \right) \varphi_{0}
(Y, A, Z) \right].
\end{equation*}
The variant of Cauchy-Schwarz inequality for random vectors established in
\citet{1999tripathiMatrixExtensionCauchySchwarz} can be used to conclude that
\(\mathbb{V} \left( \mathcal{P}^{0} \right) \leq \mathbb{E} \left[\varphi_{0}
(Y, A, Z)^{2} \right] = \mathbb{V}_{\ast}\). This establishes that
\(\mathbb{V}_{\ast}\) in \eqref{eqn--speb-car} is an upper bound over all
parametric information bounds for estimation of the ATE. However, one still
needs to show that \(\mathbb{V}_{\ast}\) is in fact a supremum. This final step
is done by establishing that \(\varphi_{0}\) can be approximated by the score
functions of regular parametric submodels.
\section{Semiparametrically efficient estimation of the ATE under CAR}
\label{sec--efficient-estimation}
The efficiency bound derived in \th\ref{thm--speb-car} is a lower bound on
asymptotic variances of regular semiparametric estimators for the ATE. However,
the question remains as to whether it can be achieved. Typically, the conditions
under which estimators can achieve semiparametric efficiency bounds are stronger
than those required to derive the bounds.
\citet{1990ritovAchievingInformationBounds} provide examples for cases
when an efficiency bound can be established and shown to be finite as well as
non-singular, but even consistent estimation (let alone efficient) is impossible
without restricting the semiparametric model. One of their examples is the
partially linear model (\citet{1986engleSemiparametricEstimatesRelation},
\citet{1988robinsonRootNConsistent}). The restrictions they require are in the
form of additional smoothness conditions for the nonparametric nuisance
parameters that appear in the efficient influence function. These smoothness
conditions are to ensure that the nonparametric nuisance parameters can be
estimated consistently with a rate of convergence of at least \(n^{- 1 / 4}\)
(see also \citet{1994neweyAsymptoticVarianceSemiparametric}). Furthermore,
achieving \(n^{- 1 / 4}\)-consistency for a non-parametric estimator becomes
more difficult when there are many covariates due to the curse of
dimensionality, and thus higher order (i.e. stronger) smoothness conditions
are required.
In the context of estimating average treatment effects with the assumption of
selection on observables (or ignorability conditional on covariates) and
i.i.d. data, \citet{1998hahnRolePropensityScore} proposes nonparametric
imputation estimators. These estimators impute the potential outcome \(Y_{i}
(a)\) whenever it is unobserved by using an estimate of the predicted value \(m
\left( a, Z_{i} \right)\), say \(\widehat{m}_{n} \left( a, Z_{i} \right)\). The
difference of the imputed potential outcomes is then averaged to estimate the
ATE. \citet{1998hahnRolePropensityScore} suggests the use of series estimators
for the propensity score \(\Pr (A = 1 | Z = z)\) and the conditional
expectations \(\mathbb{E} [Y \mathbb{I} \{A = a\} | Z = z]\) which then get
combined to construct estimators for \(m_{\ast} (a, \cdot)\). Similar estimators
appear in various parts of the literature on semiparametric estimation of
treatment effects under treatment ignorability - see
\citet{2004imbensNonparametricEstimationAverage} for a review and examples with
kernel estimators for the aforementioned propensity score and conditional
expectations. The requirement of higher order smoothness conditions on the
propensity score and the conditional expectations \(m_{\ast} (a, \cdot)\) are
ubiquitous in these examples.
In our case, the propensity score does not need to be estimated since the target
proportions by strata, \((\pi (1), \dots, \pi (\mathcal{S}))\), are chosen by
the experimenter. Furthermore, we should be able to use sample treatment
proportions by strata in place of true target proportions without any issue
since these form a finite-dimensional additional nuisance parameter. This
feature makes the efficient influence function linear in the nonparametric
nuisance parameters \(m_{\ast} (a, \cdot)\). It is also straightforward
to verify that \(\varphi\) has the Neyman orthogonality or local robustness
property of \citet{2018chernozhukovDoubleDebiasedMachine} and
\citet{2020chernozhukovLocallyRobustSemiparametric} with respect to estimation
of \(m _{\ast}(a, \cdot)\). To see this, denote
\begin{equation}
\begin{split}
\varphi \left( y, a, z; \widetilde{m} (0, \cdot), \widetilde{m} (1, \cdot),
b \right) =
& \ \frac{a}{\pi (\mathbb{S} (z))} \left[ y - \widetilde{m} (1, z)
\right] - \frac{1 - a}{1 - \pi (\mathbb{S} (z))} \left[ y - \widetilde{m}
(0, z) \right] \\
& + \left[ \widetilde{m} (1, z) - \widetilde{m} (0, z) - b \right]
\end{split}
\label{eqn--eif-moment-cond}
\end{equation}
so that the efficient influence function satisfies \(\varphi_{0} (\cdot) =
\varphi \left( \cdot; m_{\ast} (0, \cdot), m_{\ast} (1, \cdot), \beta_{0}
\right)\). If \(m_{\ast} (0, \cdot), m_{\ast} (1, \cdot)\) were known, then we
could estimate \(\beta\) by solving the sample equivalent of the moment
condition
\begin{equation*}
\mathbb{E} \left[ \varphi \left( Y, A, Z; m_{\ast} (0, \cdot), m_{\ast} (1,
\cdot), \beta_{0} \right) \right] = 0.
\end{equation*}
The local robustness property stems from noting that at \(\beta_{0}\),
\begin{equation*}
\frac{d}{d \tau} \mathbb{E} \left[ \varphi \left( Y, A, Z; (1 - \tau) m_{\ast}
(0, \cdot) + \tau \widetilde{m} (0, \cdot), (1 - \tau) m_{\ast} (1, \cdot) +
\tau \widetilde{m} (1, \cdot), \beta_{0} \right) \middle] \right|_{\tau = 0} =
0.
\end{equation*}
This local robustness property allows for \(n^{- \frac{1}{2}}\) consistent and
efficient estimation of the ATE by using plug-in estimates of \(m_{\ast} (a,
\cdot)\) under much weaker conditions. In particular, we do not require
smoothness conditions for \(m_{\ast} (a, \cdot)\) to achieve the efficiency
bound in \eqref{eqn--speb-car}.
In the i.i.d. case (\th\ref{eg--sra}), an ``ideal'' efficient estimator of the
ATE would be \(\widetilde{\beta}^{\ast}_{n} = \beta_{0} + n^{- 1} \sum_{i =
1}^{n} \varphi_{0} \left( Y_{n i}, A_{n i}, Z_{i} \right)\). Since adding
\(b\) to \(\varphi\) in \eqref{eqn--eif-moment-cond} removes it as an argument,
this estimator is
\begin{equation}
\widetilde{\beta}^{\ast}_{n} = \frac{1}{n} \sum_{i = 1}^{n} \left\{
\frac{A_{n i} \left[ Y_{n i} - m_{\ast} \left( 1, Z_{i} \right) \right]}{\pi
\left( \mathbb{S} \left( Z_{i} \right) \right)} - \frac{\left( 1 - A_{n i}
\right) \left[ Y_{n i} - m_{\ast} \left( 0, Z_{i} \right) \right]}{1 - \pi
\left( \mathbb{S} \left( Z_{i} \right) \right)} + m_{\ast} \left( 1, Z_{i}
\right) - m_{\ast} \left( 0, Z_{i} \right) \right\}.
\label{eqn--infeasible-eif-estimator}
\end{equation}
\(\widetilde{\beta}_{n}^{\ast}\) is the familiar AIPW estimator of
\citet{1994robinsEstimationRegressionCoefficients},
\citet{1995robinsAnalysisSemiparametricRegression}, and
\citet{1999scharfsteinAdjustingNonignorableDrop} if \(m_{\ast} (a, \cdot)\) were
known. A feasible estimator must replace these with first stage estimates that
can be computed from the sample. Since it will be convenient later on, we allow
for the estimates of \(m_{\ast} (a, \cdot)\) to vary across observations. We
will also allow for the experimenter to use estimated treatment proportions
rather than the known true ones. That is, let \(\widehat{\pi}_{n} (s)\) be a
consistent estimator for \(\pi (s)\) for each \(s \in
\mathbb{N}_{\mathcal{S}}\). A feasible estimator for \(\beta_{0}\) would be
\begin{equation}
\begin{split}
\widehat{\beta}^{\ast}_{n} =
& \ \frac{1}{n} \sum_{i = 1}^{n} \frac{A_{n i} \left[ Y_{n i} -
\widehat{m}_{n i} \left( 1, Z_{i} \right) \right]}{\widehat{\pi}_{n}
\left( \mathbb{S} \left( Z_{i} \right) \right)} - \frac{1}{n} \sum_{i =
1}^{n} \frac{\left( 1 - A_{n i} \right) \left[ Y_{n i} - \widehat{m}_{n i}
\left( 0, Z_{i} \right) \right]}{1 - \widehat{\pi}_{n} \left( \mathbb{S}
\left( Z_{i} \right) \right)} \\
& + \frac{1}{n} \sum_{i = 1}^{n} \left[ \widehat{m}_{n i} \left( 1, Z_{i}
\right) - \widehat{m}_{n i} \left( 0, Z_{i} \right) \right].
\end{split}
\label{eqn--feasible-eif-estimator}
\end{equation}
It is straightforward to show that \(\widetilde{\beta}^{\ast}_{n}\) achieves the
efficiency bound \(\mathbb{V}_{\ast}\) in \eqref{eqn--speb-car} under the CAR
procedures satisfying our assumptions. Our plug-in estimators \(\widehat{m}_{n
i}\) will be cross-fitted estimators as described in
\citet{2018chernozhukovDoubleDebiasedMachine} (see their DML2 estimators). We
now provide conditions on the estimators \(\widehat{m}_{n i}\), and
\(\widehat{\pi}_{n, a} (s)\) as well as the overall model under which the
difference between \(\widehat{\beta}^{\ast}_{n}\) to
\(\widetilde{\beta}^{\ast}_{n}\) in \eqref{eqn--eif-feas-infeas-diff} are
asymptotically negligible after scaling by \(\sqrt{n}\). We start with
assumptions on a sequence of nonparametric estimators \(\widehat{m} (n, a,
\cdot; \cdot)\) from which \(\widehat{m}_{n i} (a, \cdot)\) will be constructed.
\begin{assumption}
\th\label{asm--mhat-L2}
For \(a \in \{0, 1\}\), \(\widehat{m}\left( n, a, z; y_{1} (a), z_{1}, \dots,
y_{n} (a), z_{n} \right)\) is an estimator sequence satisfying
\begin{equation}
\lim_{n \to \infty} \mathbb{E}_{Q^{n + 1}} \left[ \left( \widehat{m} \left( n,
a, Z; Y_{1} (a), Z_{1}, \dots, Y_{n} (a), Z_{n} \right) - m (a, Z; Q)
\right)^{2} \right] = 0,
\label{eqn--mhat-univ-L2-consistent}
\end{equation}
when \(Z\), \(\left\{ Y_{i} (0), Y_{i} (1), Z_{i} \right\}_{i = 1}^{n}\) are
i.i.d. with distribution \(Q \in \mathbf{Q} \left( \widehat{m} \right)\) with
\(\mathbf{Q} \left( \widehat{m} \right) \subseteq \mathbf{Q}\). In the above,
\(m (a, z; Q) = \mathbb{E}_{Q} \left[ Y (a) | Z = z \right]\) and
\(\mathbb{E}_{Q^{n + 1}}\) is expectation under the product measure \(Q^{n +
1}\). Furthermore, \(Q_{0} \in \mathbf{Q} \left( \widehat{m} \right)\).
\end{assumption}
The estimator \(\widehat{m} (n, a \cdot; \cdot)\) is allowed to use the entire
sample of potential outcomes for \(a\) as well as the covariates. Of course, all
of these are not available - \(\widehat{m} (n, a \cdot; \cdot)\) can only be
computed from the treatment group for treatment \(a\). This will be accounted
for later on when we describe the cross-fitting procedure that leads to
\(\widehat{m}_{n i}\). \th\ref{asm--mhat-L2} requires that \(\widehat{m}\) be
\(L_{2}\) consistent for the true conditional expectation \(m (\cdot; Q)\) when
the data are generated according to \(Q\) which belongs to an appropriately
chosen subset \(\mathbf{Q} \left( \widehat{m} \right) \subseteq
\mathbf{Q}\) which also contains \(Q_{0}\). Note that the \(L_{2}\) convergence
rate of \(\widehat{m}\) can be arbitrarily slow. For our purposes, simple
consistency of this kind is enough. However, \(\mathbf{Q} \left( \widehat{m}
\right)\) can be formed by restrictions that impose a rate condition on
\(\widehat{m}\). For instance, the experimenter might impose
smoothness conditions required in existing literature for the particular
\(\widehat{m}\) to obtain particular rates - e.g. a H\"older or Sobolev ball. If
\(\widehat{m}\) is a series estimator, these can be found for
example in \citet{1997neweyConvergenceRatesAsymptotic} and
\citet{2015belloniSomeNewAsymptotic}. For nonlinear sieves such as
neural networks, one can see for
\citet{1994barronApproximationEstimationBounds},
\citet{1999chenImprovedRatesAsymptotic},
\citet{2020schmidt-hieberNonparametricRegressionUsing} and
\citet{2021farrellDeepNeuralNetworks}. For kernel
estimators with uniform rates, see \citet{1996masryMultivariateLocalPolynomial}
and \citet{2008hansenUniformConvergenceRates}. Other restrictions may be
functional form restrictions, e.g., random forests are known to be \(L_{2}\)
consistent for additive models - see
\citet{2015scornetConsistencyRandomForests}. When \(\mathbf{Q} \left(
\widehat{m} \right) = \mathbf{Q}\), we will say that \(\widehat{m}\) is
\emph{universally \(L_{2}\) consistent}. This is a property enjoyed by the
Nadaraya-Watson Kernel estimator - see for instance
\citet{1980devroyeDistributionFreeConsistencyResults} and
\citet{1980spiegelmanConsistentWindowEstimation}. Next, we state assumptions
that describe the cross-fitting approach to be used here.
\begin{assumption}
\th\label{asm--sample-split}
Let \(J \in \mathbb{N} \setminus \{1\}\) be given. Let \(\mathbf{U}_{n} =
\left\{ U_{n} (a, s) : s \in \mathbb{N}_{\mathcal{S}}, a \in \{0, 1\} \right\}\)
be a set of independent \(\mathrm{Uniform} ([0, 1])\) random variables on
\((\Omega, \mathcal{F}, \Pr)\) such that \(\mathbf{U}_{n} \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \left(
\mathbf{W}_{n}, \mathbf{A}_{n} \right)\). Then,
\begin{equation*}
\left\{ \mathcal{G}_{n, j} (a, s) : j \in \mathbb{N}_{J}, a \in \{0, 1\}, s
\in \mathbb{N}_{\mathcal{S}} \right\}
\end{equation*}
are random subsets of \(\mathbb{N}_{n}\) satisfying the following for each \(s
\in \mathbb{N}_{s}\) and \(a \in \{0, 1\}\).
\begin{enumerate}
\item \label{asm--sample-split-indep}
\(\mathcal{G}_{n 1} (a, s), \dots, \mathcal{G}_{n J} (a, s)\) depend only on
\(U_{n} (a, s)\) and \(N_{n} (a, s)\).\footnote{Formally, \(\mathcal{G}_{n 1}
(a, s), \dots, \mathcal{G}_{n J} (a, s)\) are measureable with respect to the
\(\sigma\)-algebra generated by the pair \(U_{n} (a, s), N_{n} (a, s)\)}
\item \label{asm--sample-split-size}
For each \(s \in \mathbb{N}_{\mathcal{S}}\) and \(a \in \{0, 1\}\),
\(\mathcal{G}_{n 1} (a, s), \dots, \mathcal{G}_{n J} (a, s)\) forms a
partition of \(\left\{ i \in \mathbb{N}_{n} : A_{n i} = a, S_{i} = s
\right\}\) satisfying
\begin{equation}
\begin{split}
\left| \mathcal{G}_{n} (a, s, j) \right| =
& \ \left\lfloor N_{n} (a, s) / J \right\rfloor \quad j = 1, \dots, J -
1. \\
\left| \mathcal{G}_{n} (a, s, J) \right| =
& \ N_{n} (a, s) - (J - 1) \cdot \left\lfloor N_{n} (a, s) / J
\right\rfloor.
\end{split}
\label{eqn--sample-split-size}
\end{equation}
\end{enumerate}
For each \(s \in \mathbb{N}_{\mathcal{S}}\) and \(j \in
\mathbb{N}_{J}\), define the \(j\)\textsuperscript{th} fold within stratum \(s\)
by
\begin{equation}
\mathcal{G}_{n} (s, j) = \mathcal{G}_{n} (1, s, j) \cup \mathcal{G}_{n} (0,
s, j).
\label{eqn--treat-control-combine-sample-split}
\end{equation}
\end{assumption}
\begin{assumption}
\th\label{asm--split-m}
Let \(\left\{ \mathcal{G}_{n j} (s) : j \in \mathbb{N}_{J}, s \in
\mathbb{N}_{\mathcal{S}} \right\}\) be as in
\eqref{eqn--treat-control-combine-sample-split} in
\th\ref{asm--sample-split} and let \(\widehat{m}\) be as in
\th\ref{asm--mhat-L2}. For each \(i \in \mathbb{N}_{n}\), denote
\begin{equation}
\Gamma_{n} (a, i) = \bigcup \left\{ \mathcal{G}_{n j} \left( a, S_{i} \right)
: j \in \mathbb{N}_{J} \text{ such that } i \notin \mathcal{G}_{n j} \left(
S_{i} \right) \right\}.
\label{eqn--estimation-set-m-for-i}
\end{equation}
The estimator \(\widehat{m}_{n i}\) is defined by
\begin{equation}
\widehat{m}_{n i} (a, z) = \widehat{m} \left( \left| \Gamma_{n} (a, i)
\right|, a, z; \left\{ Y_{\iota} (a), Z_{\iota} : \iota \in \Gamma_{n} (a, i)
\right\} \right).
\end{equation}
\end{assumption}
\begin{assumption}
\th\label{asm--weak-balance}
\(\widehat{\pi}_{n} (s)\) is either \(\pi (s)\) or \(N_{n} (s)^{- 1} N_{n} (1,
s)\). Furthermore, if the latter is true, under \(Q_{0}\) defined in
\th\ref{asm--Q} and \(\alpha_{n} (\cdot | \cdot)\) in \th\ref{asm--treat-strat},
for each \(s \in \mathbb{N}_{\mathcal{S}}\),
\begin{equation}
\sqrt{n} \left| \frac{N_{n} (1, s)}{N_{n} (s)} - \pi (s) \right| =
O_{\mathrm{p}} (1).
\label{eqn--weak-balance}
\end{equation}
\end{assumption}
\th\ref{asm--sample-split} first requires the experimenter to split each
treatment group within a given stratum into \(J\) folds in a manner that is
independent to the overall sample using information only about the group
size. These folds are further required to satisfy an ``equal size'' requirement
specific to the treatment group within the stratum. For a given stratum, the two
treatment groups corresponding to fold \(j\) are then combined which produces
\(J\) folds for the overall stratum. We first split treatment groups by strata
into folds to ensure that a given combined fold has units in both treatment and
control groups. \th\ref{asm--split-m} requires that the estimator
\(\widehat{m}_{n i}\) for \(i\)\textsuperscript{th} observation be
constructed from \(\widehat{m}\) in \th\ref{asm--mhat-L2} using observations in
the same stratum as \(i\) but not in the same fold as \(i\). Finally,
\th\ref{asm--weak-balance} requires that the treatment proportion estimator
\(\widehat{\pi}_{n} (s)\) will be either the true treatment proportions or the
sample treatment proportions. For the latter, \th\ref{asm--weak-balance}
requires require consistency at a \(1 / \sqrt{n}\) rate. This corresponds to
requiring that \(\alpha_{n} (\cdot | \cdot)\) in \th\ref{asm--treat-strat} at
least satisfy ``weak balance'' as defined in the paragraph after
\th\ref{eg--sbr}. The key result is \th\ref{thm--efficient-estimation-car}
below.
\begin{theorem}
\th\label{thm--efficient-estimation-car}
Let \th\ref{asm--Q,asm--iid-Q,asm--treat-strat} hold. Then
\begin{equation}
\sqrt{n} \left( \widetilde{\beta}^{\ast}_{n} - \beta_{0} \right) =
\frac{1}{\sqrt{n}} \sum_{i = 1}^{n} \varphi_{0} \left( Y_{n i}, A_{n i}, Z_{i}
\right) \overset{\mathrm{d}}{\to} \mathcal{N} \left( 0, \mathbb{V}_{\ast}
\right).
\label{eqn--eif-gives-speb}
\end{equation}
Suppose in addition that
\th\ref{asm--mhat-L2,asm--sample-split,asm--split-m,asm--weak-balance}
hold. Then \(\widehat{\beta}^{\ast}_{n}\) in \eqref{eqn--feasible-eif-estimator}
satisfies
\begin{equation}
\sqrt{n} \left( \widehat{\beta}^{\ast}_{n} - \beta_{0} \right) = \sqrt{n}
\left( \widetilde{\beta}^{\ast}_{n} - \beta_{0} \right) + o_{p} (1)
\label{eqn--beta-star-efficient}
\end{equation}
so that \(\widehat{\beta}^{\ast}_{n}\) achieves the semiparametric efficiency
bound \(\mathbb{V}_{\ast}\).
\end{theorem}
\th\ref{thm--efficient-estimation-car} first establishes in
\eqref{eqn--eif-gives-speb} that the infeasible ideal estimator
\(\widetilde{\beta}^{\ast}_{n}\) achieves the efficiency bound in
\eqref{eqn--speb-car} under general CAR procedures. This implies that any
estimator of the ATE that is asymptotically linear with influence function
\(\varphi_{0}\) in \eqref{eqn--eif} achieves the semiparametric efficiency
bound. \th\ref{thm--efficient-estimation-car} then establishes in
\eqref{eqn--beta-star-efficient} that \(\widehat{\beta}^{\ast}_{n}\) as defined
by \eqref{eqn--feasible-eif-estimator} has this aforementioned property (under
the additional assumption of this section). The only additional restrictions
placed on the overall semiparametric model \(\mathbf{Q}\) are ones defining
\(\mathbf{Q} \left( \widehat{m} \right)\) specific to the choice of estimator
\(\widehat{m}\) in \th\ref{asm--mhat-L2}. As a further corollary
(\th\ref{cor--efficient-estimation-car} below), we can show that \(\mathbf{Q}
\left( \widehat{m} \right) = \mathbf{Q}\) when \(\widehat{m}\) is the
Nadaraya-Watson kernel regression estimator. Note that the significance of
\(\mathbf{Q} \left( \widehat{m} \right) = \mathbf{Q}\) is that efficient
estimation is possible (pointwise) over the whole of \(\mathbf{Q}\). That is,
the efficiency bound can be achieved under the same conditions used for its
derivation. This is a novel result since the overall semiparametric problem is
non-trivial and involves estimation of (possibly) infinite-dimensional
conditional means. We now first describe the Nadaraya-Watson estimator, state
assumptions on its tuning parameters and then provide the result. In what
follows, we use the normalization that \(0/0 = 0\). Let \(\widehat{m}\)
in \th\ref{asm--mhat-L2} be defined by
\begin{equation}
\widehat{m} \left( n, a, z; y_{1} (a), z_{1}, \dots, Y_{n} (a), z_{n}
\right) = \frac{\sum_{j = 1}^{n} y_{j} \cdot \kappa \left( h_{n}^{- 1}
\left( z_{j} - z \right) \right)}{\sum_{j = 1}^{n} \kappa \left(
h_{n}^{- 1} \left( z_{j} - z \right) \right)}.
\label{eqn--kernel-estimators}
\end{equation}
\th\ref{asm--kernel-bw} below states the assumptions required for the bandwidth
sequence \(\left\{ h_{n} : n \in \mathbb{N} \right\}\) and the kernel function
\(\kappa\).
\begin{assumption}
\th\label{asm--kernel-bw}
The sequence \(\left\{ h_{n} : n \in \mathbb{N} \right\}\) and the function
\(\kappa : \mathbb{R}^{k} \to \mathbb{R}\) satisfy the following.
\begin{enumerate}[label=(\alph*)]
\item
\label{asm--bw}
\(h_{n} > 0\) for each \(n \in \mathbb{N}\), \(\lim_{n \to \infty} h_{n} =
0\) and \(\lim_{n \to \infty} n h_{n}^{k} = \infty\).
\item \label{asm--kernel}
For some \(\underline{\kappa}, \overline{\kappa}, r, R \in (0, \infty)\),
\(\underline{\kappa} \cdot \mathbb{I} (\| u \| \leq r) \leq \kappa (u) \leq
\overline{\kappa} \cdot \mathbb{I} (\| u \| \leq R)\) for each \(u \in
\mathbb{R}^{k}\).
\end{enumerate}
\end{assumption}
\th\ref{asm--kernel-bw} maintains exactly the same conditions required by
\citet{1980devroyeDistributionFreeConsistencyResults} and
\citet{1980spiegelmanConsistentWindowEstimation} for universal \(L_{2}\)
consistency. Note that \th\ref{asm--kernel-bw} \ref{asm--kernel} requires the
kernel function \(\kappa\) to be non-negative, to have compact support and to be
strictly positive in a closed neighborhood of the origin with non-empty
interior. The following result shows that efficient estimation is possible under
exactly the same conditions required to derive the efficiency bound.
\begin{corollary}
\th\label{cor--efficient-estimation-car}
In addition to the conditions of \th\ref{thm--efficient-estimation-car}, suppose
that \(\widehat{m}\) in \th\ref{asm--mhat-L2} is defined by
\eqref{eqn--kernel-estimators} and that \(\left\{ h_{n} : n \in \mathbb{N}
\right\}\) and \(\kappa\) satisfy \th\ref{asm--kernel-bw}. Then \(\mathbf{Q}
\left( \widehat{m} \right) = \mathbf{Q}\) and \(\widehat{\beta}^{\ast}_{n}\) in
\eqref{eqn--beta-star-efficient} achieves the efficiency bound over all of
\(\mathbf{Q}\).
\end{corollary}
The only requirement for the phenomenon in
\th\ref{cor--efficient-estimation-car} to occur is the universal consistency
requirement, i.e. that \(\mathbf{Q} \left( \widehat{m} \right) =
\mathbf{Q}\). This is not special to the Nadaraya-Watson estimator - one can
replace this with a nearest neighbor estimator and replace the conditions in
\th\ref{asm--kernel-bw} with those of
\citet{1977stoneConsistentNonparametricRegression}. As for local polynomial
kernel estimators, these are not universally \(L_{2}\) consistent in their
typical form (see \citet[Section 5.4,
p. 80-81]{2002gyoerfiDistributionFreeTheory}), but adjusted versions can have
the universal \(L_{2}\) consistency property (see
\citet{2002kohlerUniversalConsistencyLocal}). These adjustments introduce a
number of additional tuning parameters beyond the bandwidth. For universal
\(L_{2}\) consistency conditions for series estimators, see \citet[Chapter
10]{2002gyoerfiDistributionFreeTheory}.
We now provide a brief sketch of the main arguments that lead to
\eqref{eqn--beta-star-efficient} in \th\ref{thm--efficient-estimation-car}.
The difference between \(\sqrt{n} \left( \widehat{\beta}_{n}^{\ast} - \beta_{0}
\right)\) and \(\sqrt{n} \left( \widetilde{\beta}_{n}^{\ast} - \beta_{0}
\right)\) and can be summarized as
\begin{equation}
\begin{split}
\sqrt{n} \left( \widehat{\beta}^{\ast}_{n} - \beta_{0} \right) =
& \ \sqrt{n} \left( \widetilde{\beta}^{\ast}_{n} - \beta_{0} \right) + R_{1,
n} - R_{0, n} + \widetilde{R}_{1 n} - \widetilde{R}_{0 n} \\
\text{where} \quad R_{a, n} =
& \ \frac{1}{\sqrt{n}} \sum_{i = 1}^{n} \frac{\mathbb{I} \left\{ A_{n i} = a
\right\} - \widehat{\pi}_{n, a} \left( \mathbb{S} \left( Z_{i} \right)
\right)}{\widehat{\pi}_{n, a} \left( \mathbb{S} \left( Z_{i} \right)
\right)} \left[ \widehat{m}_{n i} \left(a, Z_{i} \right) - m_{\ast} \left(
a, Z_{i} \right) \right]. \\
\text{and} \quad \widetilde{R}_{a, n} =
& \frac{1}{\sqrt{n}} \sum_{i = 1}^{n} \left[ \frac{1}{\widehat{\pi}_{n, a}
\left( \mathbb{S} \left( Z_{i} \right) \right)} - \frac{1}{\pi_{a} \left(
\mathbb{S} \left( Z_{i} \right) \right)} \right] \mathbb{I} \left\{ A_{n
i} = a \right\} \left[ Y_{n i} - m_{\ast} \left( a, Z_{i} \right) \right].
\end{split}
\label{eqn--eif-feas-infeas-diff}
\end{equation}
In the above, \(\pi_{a} (s) = \pi (s)^{a} [1 - \pi (s)]^{1 - a}\) and
\(\widehat{\pi}_{n, a} (s)\) is defined analogously for each \(s \in
\mathbb{N}_{\mathcal{S}}\) and \(a \in \{0, 1\}\). Showing
\eqref{eqn--beta-star-efficient} is a matter of showing that \(R_{a, n}
= o_{\mathrm{p}} (1)\) and \(\widetilde{R}_{a, n} = o_{\mathrm{p}}
(1)\). The remainder term \(\widetilde{R}_{a, n}\) is straightforward to deal
with - a central limit theorem applies to
\begin{equation*}
(1 / \sqrt{n}) \sum_{i = 1}^{n} \mathbb{I} \left\{ A_{n i} = a, \mathbb{S}
\left( Z_{i} \right) = s \right\} \left[ Y_{n i} - m_{\ast} \left( a, Z_{i}
\right) \right]
\end{equation*}
and the difference between estimated and true inverse treatment probabilities is
constant across observations within a given stratum and converging to zero. For
\(R_{a, n}\), we deal with these by first splitting them into a sum of
\(\mathcal{S} \cdot J\) terms each corresponding to a given stratum and a given
fold within that stratum. Each of these terms can be treated separately since we
are trying to prove convergence in probability to zero. The argument then
proceeds by noting that within the stratum-fold remainder each component of the
form \(\widehat{\pi}_{n, a} (s)^{- 1} \left( \mathbb{I} \left\{ A_{n
i} = a \right\} - \widehat{\pi}_{n, a} (s) \right)\) is ``approximately'' mean
zero and has finite variance. Furthermore, for the term \(\widehat{m}_{n i}
\left(a, Z_{i} \right) - m_{\ast} \left( a, Z_{i} \right)\) has second moment
tending to zero. By careful use of the independence between
the evaluation point \(Z_{i}\) and the estimation sample (observations outside
the corresponding fold), the stratum-fold specific version of \(R_{a, n}\) is
shown to be \(o_{\mathrm{p}} (1)\) via Chebychev's inequality. Since the
overall \(R_{a, n}\) is the sum of its stratum-fold specific counterparts and
there are \(\mathcal{S} \cdot J\) (i.e. finitely many) of these, the conclusion
follows from Slutsky's theorem. A rigorous version of these arguments is left to
the appendix.
The key finding here is that knowledge of the propensity score and its finite
support structure reduces the burden for efficient estimation considerably. The
fact that knowledge of the propensity score can reduce this burden was also
noticed by \citet{2021aronowNonparametricIdentificationNot} and
\citet{2018rotheFlexibleCovariateAdjustments}.
\citet{2021aronowNonparametricIdentificationNot}
show that potentially non-linear parametric regression adjustments can improve
estimation accuracy, generalizing \citet{2001yangEfficiencyStudyEstimators} and
\citet{2013linAgnosticNotesRegression} to the nonlinear
case. \citet{2018rotheFlexibleCovariateAdjustments} shows that
semiparametrically efficient estimation is possible under a fairly broad class
of estimators using Donsker conditions.
\citet{2018rotheFlexibleCovariateAdjustments} also shows that
efficient estimation is possible under slightly weaker conditions by using a
leave-one-out locally linear kernel regression estimator by leveraging the
results of \citet{2019rothePropertiesDoublyRobust}. However, both still require
dimension-dependent smoothness conditions on the conditional means to be
estimated. To the best of our knowledge, existing results using nonparametric
estimators in this context all require existence of densities for baseline
covariates that are bounded away from zero on their support and higher order
differentiability on conditional means and possibly also densities. Our results
do not require any of these for efficient estimation.
It should be noted that while \(\widehat{\beta}^{\ast}_{n}\) in
\eqref{eqn--feasible-eif-estimator} achieves the efficiency bound a few issues
prevent it from being usable immediately in the context of statistical
inference. In particular, we have not provided consistent estimators of the
asymptotic variance \(\mathbb{V}_{\ast}\) in \eqref{eqn--speb-car}. It may be
possible to construct these directly from the sample and the nonparametric
estimators \(\widehat{m}_{n i}\). Another possibility might be a bootstrap
procedure to produce confidence intervals for hypothesis tests. Another issue we
have not addressed is data-based choice of tuning parameters for the estimators
\(\widehat{m}_{n i}\) when these are based on the Nadaraya-Watson kernel
estimator. Since our main questions are the efficiency bound and the conditions
under which it can be achieved, we leave the construction of consistent
inference procedures from \(\widehat{\beta}^{\ast}_{n}\) and data-based tuning
parameter selection for future research.
\section{Monte Carlo evidence}
\label{sec--sims}
In this section, we examine the finite sample performance of feasible efficient
estimator \(\widehat{\beta}_{n}^{\ast}\) from
\eqref{eqn--feasible-eif-estimator}. We start by describing the data generating
processes (DGPs) for our simulations. We consider values of \(n \in \{500, 1000,
2000, 4000, 8000\}\) and \(k \in \{1, 5\}\). The DGPs are all of the form
\begin{equation}
Y_{i} (a) = m_{\ast} \left( a, Z_{i} \right) + \sigma \left( a, Z_{i} \right)
\cdot \varepsilon_{i}.
\end{equation}
In all cases, we have \(Z_{i} \sim \mathrm{Uniform} \left([-1, 1]^{k} \right)\)
and \(\varepsilon_{i} \sim \mathcal{N} (0, 1)\). For each DGP, we conduct 5000
Monte Carlo simulations. In each simulation round, we compute
\(\widetilde{\beta}_{n}^{\ast}\) in \eqref{eqn--infeasible-eif-estimator} and
\(\widehat{\beta}_{n}^{\ast}\) in \eqref{eqn--feasible-eif-estimator} as well as
additional estimators of the ATE to be described below. We report mean squared
errors as well as biases of these estimators relative to the true value of the
ATE. For stratification, we stratify on the basis of the first component of the
covariates \(Z_{i}\). We consider \(\mathcal{S} \in \{5, 20\}\) and for each
value of \(\mathcal{S}\), strata are constructed by splitting the interval
\([-1, 1]\) into adjacent segments of length \(2 / \mathcal{S}\). We report
results with constant assignment proportions across all strata as well as
proportions that vary by strata. We present results with SPBR here, and defer
results with SSRA to the appendix. Inspecting both, the reader can see that the
choice of assignment mechanism has no visible impact on the performance of any
of the estimators considered here. For constant proportions, we set assignment
proportions to 1/2 and for varying proportions we use
\begin{equation}
\bm{\pi} =
\begin{cases}
(0.3, 0.4, 0.5, 0.6, 0.7) & \text{if } \mathcal{S} = 5, \\
(0.325, 0.35, 0.375, \dots, 0.75, 0.775, 0.8) & \text{if } \mathcal{S} = 20.
\end{cases}
\end{equation}
We use the boldfaced \(\bm{\pi}\) here to distinguish assignment probabilities
from the mathematical constant \(\pi = 3.14159\dots\), since our choice of
conditional mean functions involve the latter through trigonometric
functions. We consider four different DGP's by choice of \(k\), \(m_{\ast} (a,
\cdot)\) and \(\sigma (a, \cdot)\). In what follows, DGP's 1 and 2 have \(k =
1\) whereas DGP's 3 and 4 have \(k = 5\). In all DGP's, the true value of the
ATE is \(\beta_{0} = 0\). This is done for convenience.
\begin{enumerate}
\item DGP 1: \(m_{\ast} (0, z) = \sin (10 \pi z)\), \(m_{\ast} (1, z) = m_{\ast}
(0, z) + 2 \cos (10 \pi z)\), \(\sigma (0, z) = 1 + |z|\) and \(\sigma (1, z)
= \sqrt{2} \cdot \sigma (0, z)\)
\item DGP 2: \(m_{\ast} (0, z) = \mathrm{sign} (z) \cdot \lfloor 10 z \rfloor /
10\), \(m_{\ast} (1, z) = m_{\ast} (0, z) + 2 m_{\ast} (0, z)^{3}\), \(\sigma
(\cdot)\) as in DGP 1.
\item DGP 3: \(\sigma (0, z) = 1\), \(\sigma (1, z) = \sqrt{2}\), \(m_{\ast} (0,
z) = \lambda \left( z_{5}, z_{4}, z_{3}, z_{2}, z_{1} \right)\) and \(m_{\ast}
(1, z) = \lambda (z) + 2 \sin \left( 2 \pi z_{1} z_{2} \right)\)
\begin{equation*}
\lambda (z) = \cos \left( 2 \pi z_{1} z_{2} \right) + \left( z_{3} + z_{4} -
1 \right)^{2} + \frac{z_{5}}{2}
\end{equation*}
\item DGP 4: \(\sigma (\cdot)\) and \(m_{\ast} (\cdot)\) as in DGP 3, but with
\(\lambda\) defined instead by
\begin{equation*}
\lambda (z) = \ \cos \left( 2 \pi z_{1} z_{2} \right) + \left( z_{3} + z_{4}
- 1 \right)^{2} + \frac{z_{5}}{2} + \mathbb{I} \left( z_{1} \geq 0 \right).
\end{equation*}
\end{enumerate}
In both DGP 1 and DGP 3, \(m_{\ast} (1, \cdot)\) and \(m_{\ast} (0, \cdot)\) are
analytic and exhibit considerable non-linear variation within strata with either
\(\mathcal{S} = 5\) or \(\mathcal{S} = 20\). In DGP 2, \(m_{\ast} (1, \cdot)\)
and \(m_{\ast} (0, \cdot)\) have 19 discontinuities and are exactly fully
saturated regressions with \(\mathcal{S} = 20\) but not \(\mathcal{S} = 5\). DGP
4 adds a jump discontinuity at \(z_{1} = 0\) to the conditional mean functions
of DGP 3.
For comparison, we also examine the finite sample performance of the infeasible
efficient estimator \(\widetilde{\beta}_{n}^{\ast}\), the fully saturated
regression estimator of the ATE \(\widehat{\beta}_{n, \mathrm{SAT}}\) in
\eqref{eqn--est-sat}, as well as the imputation estimator of
\citet{1998hahnRolePropensityScore} using the cross-fitted Nadaraya-Watson
estimator:
\begin{equation}
\widehat{\beta}_{n, \mathrm{IMP}} = \frac{1}{n} \sum_{i = 1}^{n} \left[ Y_{n
i} A_{n i} + \left( 1 - A_{n i} \right) \widehat{m}_{n, i} \left( 1, Z_{i}
\right) \right] - \frac{1}{n} \sum_{i = 1}^{n} \left[ \left( 1 - A_{n i}
\right) Y_{n i} + A_{n i} \widehat{m}_{n, i} \left( 0, Z_{i} \right) \right].
\label{eqn--est-reg}
\end{equation}
The estimator above is an imputation estimator since it imputes the treatment
\(a\) potential outcome for observation \(i\) with the predicted value
\(\widehat{m}_{n i} \left( a, Z_{i} \right)\) whenever that potential outcome is
unobserved. This estimator is also asymptotically linear and with influence
function \(\varphi_{0}\) when the functions \(m_{\ast} (a, \cdot)\) are
sufficiently smooth - see for instance
\citet{1994chengNonparametricEstimationMean}. For our nonparametric estimators
\(\widehat{m}_{n, i} \left( a, \cdot \right)\), we use bandwidths of the form
\(c_{k} n^{- 1 / (4 + k)}\) and the uniform kernel so that
\th\ref{asm--kernel-bw} is satisfied. For DGP 1 and DGP 3, the bandwidths are
rate-optimal in terms of integrated mean-squared error (see
\citet{1982stoneOptimalGlobalRates}) since the conditional mean functions
\(m_{\ast} (a, \cdot)\) are twice continuously differentiable. For \(k = 1\), we
set \(c_{k} = 1 / \sqrt{3}\) and for \(k = 5\), we set \(c_{k} = 3\). In our
results, estimated mean squared errors for \(\widetilde{\beta}_{n}^{\ast}\) act
as an estimate of the semiparametric efficiency bound. The comparison between
\(\widehat{\beta}_{n}^{\ast}\) and \(\widetilde{\beta}_{n}^{\ast}\) illustrates
how far from optimal the feasible estimator can be due to noise in the
estimation of the conditional mean functions. Following the discussion in
Section \ref{sec--speb}, the fully saturated regression estimator
\(\widehat{\beta}_{n, \mathrm{SAT}}\) in \eqref{eqn--est-sat} achieves the
semiparametric efficiency bound among all estimators that use only information
contained in the strata. This is \(\mathbb{V}_{\mathrm{SAT}}\), defined in
\eqref{eqn--speb-sat}. Thus, comparing \(\widehat{\beta}_{n, \mathrm{SAT}}\) to
\(\widetilde{\beta}_{n}^{\ast}\) illustrates the magnitude of efficiency loss
from ignoring information from the covariates (except in DGP 2 with
\(\mathcal{S} = 20\)). Additionally, comparing \(\widehat{\beta}_{n,
\mathrm{SAT}}\) to \(\widehat{\beta}_{n}^{\ast}\) illustrates the possible
tradeoff between finite sample and asymptotic accuracy when estimates of the
nonparametric components are potentially far from the truth. The additional
nonparametric estimator \(\widehat{\beta}_{n, \mathrm{IMP}}\) has the same
influence function as \(\widehat{\beta}_{n}^{\ast}\) but does not directly use
the structure of the influence function. Thus, comparisons between these
estimators illustrates the debiasing effect of using the efficient influence
function directly.
Tables \ref{tab--cons5} and \ref{tab--cons20} present results from simulations
with constant target treatment proportions across strata with \(\mathcal{S} =
5\) and \(\mathcal{S} = 20\) respectively. Tables \ref{tab--var5} and
\ref{tab--var20} present results with varying target treatment proportions
across strata with \(\mathcal{S} = 5\) and \(\mathcal{S} = 20\)
respectively. Some common patterns appear across examples. In the univariate
case (DGP's 1 and 2), \(\widehat{\beta}^{\ast}_{n}\) is fairly close to the
infeasible estimator \(\widetilde{\beta}^{\ast}_{n}\) in terms of mean squared
error. Furthermore, while \(\widehat{\beta}_{n, \mathrm{SAT}}\) is far from
optimal under the nonlinearities in DGP 1 and \(\mathcal{S} = 5\), the distance
is greatly reduced by using finer strata (i.e. \(\mathcal{S} = 20\)).
However, in higher dimensions, both estimators suffer albeit in different
ways. Under DGP's 3 and 4 the feasible optimal estimator
\(\widehat{\beta}^{\ast}_{n}\), despite still being consistent and
asymptotically unbiased, is much slower in its tendency towards the SPEB.
This is mainly due to the curse of dimensionality. In simulations with \(n \in
\{10000, 20000\}\) for instance, \(\widehat{\beta}^{\ast}_{n}\) is closer to
achieving the SPEB, but these sample sizes may be unrealistic for most
experiments. As we can see, \(\widehat{\beta}_{n}^{\ast}\) is also sensitive to
the number of strata - having more strata reduces the effective estimation
sample used for estimation of the conditional mean. This is particularly stark
in the examples with varying treatment proportions - in these, some treatment
groups are going to have small sample sizes within strata by design and the
aforementioned phenomenon is exacerbated. In the multivariate case, increasing
the number of strata does not help reduce the MSE of \(\widehat{\beta}_{n,
\mathrm{SAT}}\) since stratification is done on the basis of a single component
of a continuously distributed covariate. On the other hand, though its
performance in some instances (e.g. \ref{tab--var20}, DGP 3 and 4) leaves much
to be desired, in most cases with DGP 3 and 4, \(\widehat{\beta}^{\ast}_{n}\)
picks up on the nonlinear variation in the conditional means across all
dimensions. To summarize the comparison between \(\widehat{\beta}^{\ast}_{n}\)
and \(\widehat{\beta}_{n, \mathrm{SAT}}\) we find that across all simulation
designs, \(\widehat{\beta}^{\ast}_{n}\) is on average 13\% more efficient
(averaging the quotient of the MSEs). The maximal gain from using
\(\widehat{\beta}^{\ast}_{n}\) is a 40\% reduction in MSE - Table
\ref{tab--var5}, DGP 4 with \(n = 8000\).
For \(\widehat{\beta}_{n, \mathrm{IMP}}\), while it seems to perform well under
\(k = 1\) and \(\mathcal{S} = 20\) the impact of the curse of dimensionality and
its interaction with the lack of debiasing is quite stark as can be seen from
the other cases. Furthermore, the need for undersmoothing to reduce bias for
this estimator is apparent since it seems to sometimes do worse with more
effective observations. For instance, comparing results for DGP 1 in Tables
\ref{tab--cons5} and \ref{tab--cons20}, we see that increasing the number of
strata (i.e. decreasing the effective number of observations per strata)
produces bias reductions without seriously affecting variance for
\(\widehat{\beta}_{n, \mathrm{SAT}}\) or \(\widehat{\beta}_{n,
\mathrm{IMP}}\). However, reducing bandwidth to undersmooth will reduce bias and
inflate the variance of this estimator.
\begin{table}[H]
\centering
\caption{\(\mathcal{S} = 5\) and constant target assignments across strata.}
\begin{tabular}{ccccccccccccc}
\toprule
\toprule
& & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widetilde{\beta}^{\ast}_{n}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}^{\ast}_{n}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}_{n, \mathrm{SAT}}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}_{n, \mathrm{IMP}}\)} \\
\cmidrule{3-4}\cmidrule{6-7}\cmidrule{9-10}\cmidrule{12-13} DGP & \(n\) & MSE & Bias & & MSE & Bias & & MSE & Bias & & MSE & Bias \\
\midrule
\multirow{5}[2]{*}{DGP 1} & 500 & 16.442 & -0.041 & & 19.883 & -0.078 & & 20.931 & -0.071 & & 22.536 & -1.471 \\
& 1000 & 15.731 & 0.024 & & 17.724 & -0.014 & & 19.589 & -0.040 & & 21.316 & -1.661 \\
& 2000 & 16.380 & -0.064 & & 17.888 & -0.087 & & 20.631 & -0.082 & & 22.458 & -1.908 \\
& 4000 & 16.084 & -0.039 & & 17.022 & -0.076 & & 19.891 & -0.085 & & 21.825 & -2.011 \\
& 8000 & 15.749 & 0.025 & & 16.466 & 0.033 & & 20.167 & 0.058 & & 20.990 & -1.955 \\
\midrule
\multirow{5}[2]{*}{DGP 2} & 500 & 14.679 & -0.012 & & 15.139 & -0.020 & & 14.928 & -0.014 & & 15.090 & -0.016 \\
& 1000 & 14.000 & -0.012 & & 14.274 & -0.009 & & 14.235 & -0.005 & & 14.266 & -0.010 \\
& 2000 & 14.604 & -0.058 & & 14.733 & -0.061 & & 14.836 & -0.067 & & 14.749 & -0.060 \\
& 4000 & 14.672 & -0.077 & & 14.786 & -0.087 & & 14.912 & -0.078 & & 14.802 & -0.085 \\
& 8000 & 14.296 & 0.006 & & 14.327 & 0.006 & & 14.511 & 0.000 & & 14.326 & 0.005 \\
\midrule
\multirow{5}[2]{*}{DGP 3} & 500 & 12.593 & -0.011 & & 18.453 & -0.016 & & 24.129 & -0.078 & & 18.291 & -0.685 \\
& 1000 & 13.148 & -0.034 & & 17.397 & -0.007 & & 24.508 & -0.070 & & 18.098 & -0.792 \\
& 2000 & 12.974 & -0.019 & & 16.380 & -0.036 & & 24.091 & -0.008 & & 17.481 & -0.912 \\
& 4000 & 12.813 & -0.044 & & 15.290 & -0.047 & & 23.656 & -0.005 & & 16.836 & -1.015 \\
& 8000 & 13.060 & -0.014 & & 15.155 & -0.030 & & 24.447 & -0.107 & & 16.987 & -1.087 \\
\midrule
\multirow{5}[2]{*}{DGP 4} & 500 & 12.624 & -0.021 & & 18.743 & -0.030 & & 24.668 & -0.090 & & 18.566 & -0.702 \\
& 1000 & 13.120 & -0.040 & & 17.667 & -0.012 & & 25.168 & -0.070 & & 18.351 & -0.795 \\
& 2000 & 12.932 & 0.001 & & 16.449 & -0.014 & & 24.615 & 0.011 & & 17.523 & -0.893 \\
& 4000 & 12.762 & -0.037 & & 15.431 & -0.034 & & 24.227 & 0.018 & & 16.968 & -1.000 \\
& 8000 & 13.006 & -0.020 & & 15.197 & -0.031 & & 25.097 & -0.109 & & 17.076 & -1.090 \\
\bottomrule
\bottomrule
\end{tabular}
\label{tab--cons5}
\end{table}
\begin{table}[H]
\centering
\caption{\(\mathcal{S} = 20\) and constant target assignments across strata.}
\begin{tabular}{ccccccccccccc}
\toprule
\toprule
& & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widetilde{\beta}^{\ast}_{n}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}^{\ast}_{n}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}_{n, \mathrm{SAT}}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}_{n, \mathrm{IMP}}\)} \\
\cmidrule{3-4}\cmidrule{6-7}\cmidrule{9-10}\cmidrule{12-13} DGP & \(n\) & MSE & Bias & & MSE & Bias & & MSE & Bias & & MSE & Bias \\
\midrule
\multirow{5}[2]{*}{DGP 1} & 500 & 16.290 & -0.125 & & 18.864 & -0.126 & & 19.008 & -0.128 & & 18.919 & -0.133 \\
& 1000 & 16.171 & -0.077 & & 17.921 & -0.060 & & 18.268 & -0.058 & & 18.042 & -0.064 \\
& 2000 & 16.582 & -0.151 & & 17.939 & -0.146 & & 18.680 & -0.143 & & 18.225 & -0.149 \\
& 4000 & 15.854 & -0.015 & & 17.104 & 0.002 & & 18.390 & 0.002 & & 17.512 & 0.001 \\
& 8000 & 15.462 & -0.014 & & 16.265 & -0.024 & & 17.811 & -0.020 & & 16.688 & -0.021 \\
\midrule
\multirow{5}[2]{*}{DGP 2} & 500 & 14.549 & -0.102 & & 14.785 & -0.100 & & 14.744 & -0.100 & & 14.769 & -0.108 \\
& 1000 & 14.298 & -0.104 & & 14.436 & -0.104 & & 14.399 & -0.105 & & 14.444 & -0.109 \\
& 2000 & 14.878 & -0.161 & & 14.966 & -0.158 & & 14.933 & -0.162 & & 15.039 & -0.161 \\
& 4000 & 14.354 & -0.035 & & 14.395 & -0.033 & & 14.378 & -0.036 & & 14.530 & -0.034 \\
& 8000 & 13.917 & 0.000 & & 13.954 & -0.001 & & 13.928 & 0.000 & & 14.091 & 0.001 \\
\midrule
\multirow{5}[2]{*}{DGP 3} & 500 & 13.114 & -0.039 & & 24.858 & -0.640 & & 24.637 & -0.038 & & 18.169 & -1.151 \\
& 1000 & 12.669 & -0.049 & & 21.884 & -0.296 & & 24.590 & -0.086 & & 19.145 & -1.079 \\
& 2000 & 12.894 & 0.024 & & 19.287 & -0.114 & & 24.830 & -0.003 & & 18.681 & -1.070 \\
& 4000 & 12.471 & -0.011 & & 16.999 & -0.086 & & 24.358 & -0.046 & & 17.859 & -1.169 \\
& 8000 & 13.402 & 0.013 & & 16.501 & 0.004 & & 24.877 & 0.070 & & 18.007 & -1.088 \\
\midrule
\multirow{5}[2]{*}{DGP 4} & 500 & 13.127 & -0.049 & & 26.421 & -0.797 & & 25.091 & -0.056 & & 19.007 & -1.313 \\
& 1000 & 12.711 & -0.052 & & 22.815 & -0.349 & & 25.092 & -0.086 & & 19.735 & -1.142 \\
& 2000 & 12.893 & 0.041 & & 19.856 & -0.110 & & 25.412 & 0.020 & & 18.990 & -1.073 \\
& 4000 & 12.446 & 0.003 & & 17.231 & -0.080 & & 24.891 & -0.024 & & 18.045 & -1.166 \\
& 8000 & 13.396 & 0.019 & & 16.616 & 0.013 & & 25.588 & 0.077 & & 18.091 & -1.079 \\
\bottomrule
\bottomrule
\end{tabular}
\label{tab--cons20}
\end{table}
\begin{table}[H]
\centering
\caption{\(\mathcal{S} = 5\) and varying target assignments across strata.}
\begin{tabular}{ccccccccccccc}
\toprule
\toprule
& & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widetilde{\beta}^{\ast}_{n}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}^{\ast}_{n}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}_{n, \mathrm{SAT}}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}_{n, \mathrm{IMP}}\)} \\
\cmidrule{3-4}\cmidrule{6-7}\cmidrule{9-10}\cmidrule{12-13} DGP & \(n\) & MSE & Bias & & MSE & Bias & & MSE & Bias & & MSE & Bias \\
\midrule
\multirow{5}[2]{*}{DGP 1} & 500 & 18.224 & 0.050 & & 22.458 & 0.031 & & 22.951 & 0.059 & & 24.469 & -1.342 \\
& 1000 & 17.620 & -0.011 & & 20.586 & -0.055 & & 22.667 & -0.005 & & 24.260 & -1.689 \\
& 2000 & 17.951 & -0.036 & & 19.725 & -0.040 & & 22.685 & -0.034 & & 24.258 & -1.866 \\
& 4000 & 17.347 & -0.039 & & 18.410 & -0.077 & & 21.766 & -0.101 & & 23.362 & -2.011 \\
& 8000 & 17.519 & 0.029 & & 18.221 & 0.003 & & 22.196 & -0.027 & & 23.027 & -1.996 \\
\midrule
\multirow{5}[2]{*}{DGP 2} & 500 & 16.578 & 0.035 & & 17.596 & 0.003 & & 16.942 & 0.047 & & 17.289 & 0.019 \\
& 1000 & 16.054 & -0.038 & & 16.472 & -0.057 & & 16.471 & -0.050 & & 16.364 & -0.059 \\
& 2000 & 16.054 & -0.031 & & 16.248 & -0.021 & & 16.359 & -0.027 & & 16.245 & -0.025 \\
& 4000 & 15.899 & -0.072 & & 16.069 & -0.085 & & 16.134 & -0.067 & & 16.102 & -0.076 \\
& 8000 & 16.179 & 0.046 & & 16.208 & 0.049 & & 16.379 & 0.048 & & 16.244 & 0.049 \\
\midrule
\multirow{5}[2]{*}{DGP 3} & 500 & 13.967 & -0.004 & & 21.786 & -0.046 & & 27.781 & 0.023 & & 20.290 & -0.527 \\
& 1000 & 13.941 & 0.037 & & 20.019 & 0.041 & & 27.937 & 0.045 & & 19.804 & -0.643 \\
& 2000 & 13.738 & -0.049 & & 18.115 & -0.097 & & 27.328 & -0.098 & & 19.272 & -0.955 \\
& 4000 & 13.533 & -0.020 & & 17.139 & -0.050 & & 27.022 & -0.066 & & 18.684 & -1.024 \\
& 8000 & 13.858 & -0.023 & & 16.250 & -0.001 & & 26.544 & 0.013 & & 18.027 & -1.028 \\
\midrule
\multirow{5}[2]{*}{DGP 4} & 500 & 14.112 & -0.011 & & 22.398 & -0.059 & & 28.679 & 0.019 & & 20.594 & -0.392 \\
& 1000 & 13.959 & 0.043 & & 20.355 & 0.046 & & 28.533 & 0.050 & & 19.955 & -0.564 \\
& 2000 & 13.741 & -0.037 & & 18.338 & -0.079 & & 27.870 & -0.081 & & 19.413 & -0.913 \\
& 4000 & 13.599 & -0.016 & & 17.232 & -0.045 & & 27.575 & -0.060 & & 18.777 & -1.012 \\
& 8000 & 13.847 & -0.025 & & 16.363 & -0.002 & & 27.180 & 0.024 & & 18.175 & -1.029 \\
\bottomrule
\bottomrule
\end{tabular}
\label{tab--var5}
\end{table}
\begin{table}[H]
\centering
\caption{\(\mathcal{S} = 20\) and varying target assignments across strata.}
\begin{tabular}{ccccccccccccc}
\toprule
\toprule
& & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widetilde{\beta}^{\ast}_{n}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}^{\ast}_{n}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}_{n, \mathrm{SAT}}\)} & & \multicolumn{2}{c}{\(\sqrt{n} \cdot \widehat{\beta}_{n, \mathrm{IMP}}\)} \\
\cmidrule{3-7}\cmidrule{9-10}\cmidrule{12-13} DGP & \(n\) & MSE & Bias & & MSE & Bias & & MSE & Bias & & MSE & Bias \\
\midrule
\multirow{5}[2]{*}{DGP 1} & 500 & 17.538 & -0.023 & & 20.095 & -0.009 & & 19.849 & -0.020 & & 19.852 & -0.018 \\
& 1000 & 17.204 & -0.124 & & 19.118 & -0.129 & & 19.446 & -0.135 & & 19.259 & -0.112 \\
& 2000 & 17.350 & -0.052 & & 19.123 & -0.050 & & 19.965 & -0.050 & & 19.456 & 0.001 \\
& 4000 & 17.194 & 0.062 & & 18.189 & 0.052 & & 19.246 & 0.043 & & 18.695 & 0.142 \\
& 8000 & 17.317 & 0.083 & & 18.153 & 0.105 & & 19.540 & 0.107 & & 18.505 & 0.233 \\
\midrule
\multirow{5}[2]{*}{DGP 2} & 500 & 16.160 & -0.028 & & 16.506 & -0.031 & & 16.188 & -0.025 & & 16.278 & -0.025 \\
& 1000 & 15.856 & -0.149 & & 15.956 & -0.145 & & 15.874 & -0.149 & & 16.040 & -0.153 \\
& 2000 & 15.556 & -0.046 & & 15.576 & -0.046 & & 15.560 & -0.047 & & 15.683 & -0.048 \\
& 4000 & 15.868 & 0.020 & & 15.916 & 0.022 & & 15.869 & 0.020 & & 16.205 & 0.020 \\
& 8000 & 15.960 & 0.110 & & 15.972 & 0.113 & & 15.959 & 0.111 & & 16.053 & 0.114 \\
\midrule
\multirow{5}[2]{*}{DGP 3} & 500 & 13.420 & -0.090 & & 30.729 & -1.572 & & 26.697 & -0.067 & & 24.945 & 2.900 \\
& 1000 & 13.443 & -0.104 & & 27.120 & -0.862 & & 26.943 & -0.080 & & 30.147 & 3.458 \\
& 2000 & 13.725 & 0.017 & & 23.629 & -0.333 & & 27.433 & -0.006 & & 30.384 & 3.372 \\
& 4000 & 13.581 & -0.092 & & 20.450 & -0.160 & & 27.035 & -0.097 & & 23.979 & 2.354 \\
& 8000 & 13.925 & -0.037 & & 18.690 & -0.020 & & 27.818 & 0.015 & & 19.863 & 1.235 \\
\midrule
\multirow{5}[2]{*}{DGP 4} & 500 & 13.467 & -0.082 & & 33.787 & -1.886 & & 27.377 & -0.058 & & 36.226 & 4.364 \\
& 1000 & 13.486 & -0.122 & & 28.984 & -1.029 & & 27.681 & -0.104 & & 43.866 & 5.022 \\
& 2000 & 13.700 & 0.032 & & 24.703 & -0.361 & & 28.221 & 0.023 & & 42.664 & 4.820 \\
& 4000 & 13.649 & -0.076 & & 21.122 & -0.157 & & 27.870 & -0.084 & & 30.253 & 3.392 \\
& 8000 & 13.821 & -0.024 & & 18.879 & -0.007 & & 28.533 & 0.037 & & 21.760 & 1.827 \\
\bottomrule
\bottomrule
\end{tabular}
\label{tab--var20}
\end{table}
\section{Conclusion}
\label{sec--conclusion}
In this paper, we have characterized the maximal gains in efficiency from using
baseline covariates beyond strata in RCT's covariate adaptive randomization is
used. To do this, we have established a semiparametric efficiency bound under
CAR for such experiments. This is the minimum estimation variance achievable if
one uses all the relevant information the data have to offer about the parameter
of interest (the ATE here). If baseline covariates are used, we use the
information they contain about the ATE fully through differences in conditional
means, averaged over the support of the covariates. With continuous covariates,
or more covariates present than used in stratification, a strict improvement
over fully saturated regression estimates of the ATE is possible. The efficiency
bound established is shown to be achievable under the same conditions used for
its derivation by using a leave one out Nadaraya-Watson kernel regression
estimator. Simulation evidence presented shows that this estimator can indeed
reach the efficiency bound in large samples. However, such improvements are not
guaranteed in finite samples, and reaching the efficiency bound can be slow
especially in the presence of many covariates due to the curse of
dimensionality.
{\printbibliography}
\newpage